BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//WSSP - ECPv6.15.20//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:WSSP
X-ORIGINAL-URL:https://www.wssp.hlrs.de
X-WR-CALDESC:Events for WSSP
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Europe/London
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20170326T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20171029T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20180325T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20181028T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20190331T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20191027T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20200329T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20201025T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20210328T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20211031T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20220327T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20221030T010000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;VALUE=DATE:20210316
DTEND;VALUE=DATE:20210320
DTSTAMP:20260924T175732
CREATED:20221204T235613Z
LAST-MODIFIED:20221205T001529Z
UID:204-1615852800-1616198399@www.wssp.hlrs.de
SUMMARY:31st Workshop on Sustained Simulation
DESCRIPTION:Agenda\nDue to the corona pandemic\, the workshop was hold online. All times are given in Central European Time (CET) \n\n\n\n\n\nTuesday\, March 16th\, 2021\n\n\n\n\n9:15 – 9:30\nWelcome & Introduction\n      Michael Resch\, HLRS\, University of Stuttgart\n\n\n9:30 – 10:00\nThe New Era of Hybrid-Computing on and with SX-Aurora TSUBASA: Vector-Scalar to Vector-Digital Annealing\, to Vector-Quantum Annealing\n      Hiroaki Kobayashi\, Cyberscience Center\, Tohoku University\n      Abstract \n\n\n\n            I will be giving my talk about recent-achievements on our on-going project entitle “R&D of A Quantum-Annealing-Assisted Next Generation HPC Infrastructure and its Applications.”  My talk will start with some performance evaluation results of SX-Aurora TSUBASA as a vector computing platform\, and then go into its hybrid computing using VH-VE together. Finally\, I will discuss two types of hybrid: computing mechanism and hardware platform levels\, which are annealing on SX-Aurora TSUBASA  itself and its on SX-Aurora TSUBASA and Quantum-Annealer. \n \n \n\n\n\n10:00 – 10:30\nHPC in the Next Decade\n      Michael Resch\, HLRS\, University of Stuttgart\n      Abstract \n\n\n\n            The coming ten years of HPC will be dominated by three main developments. First\, the end of Moore’s law is already now showing its impact. Second a new application − Artificial Intelligence − is forcing computing centers to adapt their strategies. Finally\, Quantum Computing is knocking on the door of HPC without providing and hints about the direction in which QC is headed. HLRS has to adapt to these challenges and in the talk we will present challenges and opportunities as well as strategies to cope with both. \n \n \n\n\n\n10:30 – 11:00\nSoftware Methods for Product Virtualization\n      Sabine Roller\, DLR (Deutsches Zentrum für Luft- und Raumfahrt e.V.)\n      Abstract \n\n\n\n            The abstract will be provided soon. \n \n \n\n\n\n11:00 – 11:15\nBreak\n\n\n11:15 – 11:45\nSX-Aurora TSUBASA VE Design\n      Hiroki Asano\, NEC Tokyo\n      Abstract \n\n\n\n            VE (Vector Engine) of SX-Aurora TSUBASA leverages innovative vector processors which have high single-core performance and memory bandwidth. \n            NEC has been developing VE series and launched the latest model VE20 in 2020. \n            At this moment\, our challenge is to achieve better performance for the future VE products through several studies. \n            In this session\, we will introduce one of the studies regarding the memory subsystem. \n \n \n\n\n\n11:45 – 12:15\nFostering HPC Competences in Europe to Support Academia and Industry\n      Bastian Koller\, HLRS\n      Abstract \n\n\n\n            This talk will give an update on the EuroHPC Projects EuroCC and CASTIEL\, which are implementing 33 National Competence Centres in Europe. Also\, it will update on the FF4EuroHPC Projects with insights on the first closed open call for industrial experimentation. \n \n \n\n\n\n12:15\nOpen Meeting\n\n\n\n\nWednesday\, March 17th\, 2021\n\n\n\n\n8:30 – 8:35\nIntroduction of Chair\n\n\n8:35 – 9:05\nExploiting Hybrid Parallelism in LBM Implementation Musubi on Hawk\n      Harald Klimach\, Kannan Masilamani\, Sabine Roller\, University of Siegen\n      Abstract \n\n\n\n            In this contribution we look into the efficiency and scalability of our Lattice Boltzmann implementation Musubi when using OpenMP threads within an MPI parallel computation on Hawk. The Lattice Boltzmann methods enables explicit computation of incompressible flows and the mesh discretization can be automatically generated\, even for complex geometries. The basic Lattice Boltzmann kernel is fairly simple and involves only few floating point operations for each lattice node. A simple loop over all lattice nodes each partition of the MPI parallel setup lends a straight forward loop parallelization with OpenMP. With increased core counts per compute node\, the use of threads on the shared memory nodes is gaining importance\, as it avoids overly small partitions with many outbound communications to neighboring partitions. We briefly discuss the hybrid parallelization of Musubi and investigate how the usage of OpenMP threads affects the performance when running simulations on the Hawk supercomputer at HLRS. \n \n \n\n\n\n9:05 – 9:35\nHybrid Computation on Building Responses for Earthquakes on a VH and VEs of SX-Aurora TSUBASA\n      Mitsuo Yokokawa\, Kobe University\n      Abstract \n\n\n\n            We have earthquakes frequently\, and therefore buildings need to be earthquake resistant. A code for building responses in time for earthquakes was parallelized by using a hybrid execution model on a VH and VEs of the SX-Aurora TSUBASA. Computation performance will be presented. \n \n \n\n\n\n09:35 – 10:05\nForecasting Intensive Care Unit Demand during the COVID-19 Pandemic: A Spatial Age-structured Microsimulation Model\n      Ralf Schneider\, HLRS\, Sebastian Klüsener\, Matthias Rosenbaum-Feldbrügge\n      Abstract \n\n\n\n            Background: The COVID-19 pandemic poses the risk of overburdening health care systems\, and in particular intensive care units (ICUs). Non-pharmaceutical interventions (NPIs)\, ranging from wearing masks to (partial) lockdowns have been implemented as mitigation measures around the globe. However\, especially severe NPIs are used with great caution due to their negative effects on the economy\, social life and mental well-being. Thus\, understanding the impact of the pandemic on ICU demand under alternative scenarios reflecting different levels of NPIs is vital for political decision-making on NPIs. The aim is to support political decision-making by forecasting COVID-19-related ICU demand under alternative scenarios of COVID-19 progression reflecting different levels of NPIs. \n            Methods: In this talk we will present our implementation of a spatial age-structured microsimulation model of the COVID-19 pandemic by extending the Susceptible-Exposed-Infectious-Recovered (SEIR) framework. The model accounts for regional variation in population age structure and in spatial diffusion pathways. In a first step\, we calibrate the model by applying a genetic optimization algorithm against hospital data on ICU patients with COVID-19. In a second step\, we forecast COVID-19-related ICU demand under alternative scenarios of COVID 19 progression reflecting different levels of NPIs. The third step is the automation of the procedure for the provision of weekly forecasts. The automated estimation of the model’s parameters is done by means of Random-Forest regression. \n            Results: In the results section we will show the application of the model to Germany and demonstrate state-level forecasts over a 2-month period\, which can be updated daily based on latest data on the progression of the pandemic.To illustrate the merits of our model\, we present here “forecasts” of ICU demand for different stages of the pandemic during 2020 and 2021. Our forecasts for a quiet summer phase with low infection rates identified quite some variation in potential for relaxing NPIs across the federal states. By contrast\, our forecasts during a phase of quickly rising infection numbers in autumn (second wave) suggested that all federal states should implement additional NPIs. However\, the identified needs for additional NPIs varied again across federal states. In addition\, our model suggests that during large infection waves ICU demand would quickly exceed supply\, if there were no NPIs in place to contain the virus. \n              \n \n \n\n\n\n10:05 – 10:20\nBreak\n\n\n10:20 – 10:50\nHPC Based Analyses of Biofuel Injection in IC Engines and Metall Machining Processes\n      Matthias Meinke\, Tim Wegmann\, Julian Vorspohl\, Daniel Lauwers\, and Wolfgang Schröder\, RWTH Aachen\n      Abstract \n\n\n\n            Some recent engineering applications will be presented\, in which the turbulent flow in technical devices with embedded droplets and particles is simulated with the HPC platform HAWK installed at HLRS. The investigation of the flow field in the internal combustion engine is performed\, to analyze the mixing of air with various injected biofuels. The distribution of the evaporated fuel in the internal combustion engine has a large influence on the emissions of pollutants and the engine efficiency. Since biofuels possess quite different fluid properties\, it is important to accurately predict the concentration of the evaporated fuels for an optimization of the engine performance. The second application is connected to electrical discharge and electro-chemical machining processes\, which are used to manufacture work pieces such as turbine blades of high strength material. In these processes\, fluid flow plays an important role for the transport of removed material and thus on the quality of the final product. The simulations conducted are based on a solver formulated for hierarchical Cartesian meshes\, in which a Lagrangian particle solver is used to track the motion of droplets of the fuel spray or the transport of removed material. Several aspects of the numerical methods\, its parallelization\, dynamic load balancing and the implementation on high performance computing platforms will be presented in this contribution. \n \n \n\n\n\n10:50 – 11:20\nParticle-in-Cell Plasma Simulation of Filamentary Coherent Structures\n      Seiji Ishiguro and Hiroki Hasegawa\, National Institute for Fusion Science\, National Institute of Natural Sciences\n      Abstract \n\n\n\n            We have performed three dimensional particle-in-cell (PIC) plasma simulation to investigate the filamentary coherent structure so called blob/hole.\n            Impurity ion transport by this structure is revealed. The impurity ion profile in the blob/hole structure becomes a dipole structure and this propagates with the blob/hole. \n            The performance of three dimensional PIC code on the new super computer build by NEC SX-Aurora TSUBASA at National Institute for Fusion Science will also be presented. \n \n \n\n\n\n11:20 – 11:50\nComputational Simulation of Chiral Transition and Paramagnetic Current Induced by Paramagnetic Coupling in Chiral Superconductor\n      Hirono Kaneyasu\, University of Hyogo\n      Abstract \n\n\n\n            Assuming the non-unitary chiral superconductivity as a bulk state of Sr₂RuO₄\, we show the field-induced chiral stability generating the paramagnetic current in the eutectic Sr₂RuO₄-Ru by computational simulation of the Ginzburg-Landau equation. The paramagnetic coupling with the chiral magnetization causes the field-induced chiral transition and the paramagnetic current. The field-induced chiral stability consists of the field-dependence of zero-bias anomaly in the tunneling spectroscopy. This good agreement with the experimental result indicates that the non-unitary chiral spin-triplet state is one of candidates for the superconducting state of Sr₂RuO₄\, in addition to the chiral spin-single state as other candidates. The high-performance computing by code optimized for SX-Aurora makes it possible to analyze field dependence and spatial variation of chiral state and supercurrent in more detail. \n \n \n\n\n\n11:50 – 12:20\nDirect Numerical Flow Simulation on Vector and Massively-parallel Supercomputers\n      Johannes Peter\, University of Stuttgart\n      Abstract \n\n\n\n            Direct numerical simulations of turbulent flows require high computational power\, only available on supercomputers such as those provided at the HLRS. At the Institute of Aerodynamics and Gas Dynamics an adapted in-house finite-difference code of high accuracy order is used for the analysis of canonical transitional and turbulent wall-bounded flows and (in-)stability investigations. The talk will present some recent results regarding supersonic mixing of two gases and show performance data for the massively-parallel ‘Hawk’ system and the ‘SX-Aurora’ vector system. \n \n \n\n\n\n12:20\nOpen Meeting\n\n\n\n\nThursday\, March 18th\, 2021\n\n\n\n\n8:30 – 8:35\nIntroduction of Chair\n\n\n8:35 – 9:05\nIntroduction of Cloud Bursting and Seamless Use of SX-Aurora TSUBASA under Job Scheduler.\n      Tatsuyoshi Ohmura\, NEC Tokyo\n      Abstract \n\n\n\n            With the development of AI/BDA\, the number of users of HPC systems is increasing. While the computational power required by users is also increasing On-premise HPC systems are difficult to expand due to limitations in power\, space\, and budget. In this session\, we will introduce cloud bursting\, which is a Job Scheduler function to temporarily expand computational power by using cloud computing resources. In addition\, we will introduce the function that enables the HPC system to be used transparently under Job Scheduler. \n \n \n\n\n\n9:05 – 9:35\nPorting and Optimizing Molecular Docking Simulations on SX-Aurora Vector Engine\n      Erich Focht (NEC Germany)\, Leonardo Solis-Vasquez (TU Darmstadt)\, Andreas Koch (TU Darmstadt)\n      Abstract \n\n\n\n            Molecular docking simulations are widely used in computational drug discovery to predict molecular interactions at close distances. Specifically\, these simulations aim to predict the binding poses between a small molecule and a macromolecular target\, both referred to as ligand and receptor\, respectively. The purpose of drug discovery is to identify ligands that effectively inhibit the harmful function of a certain receptor. In that context\, molecular docking simulations are critical\, by using them\, the time-consuming preliminary tasks consisting of identifying potential drug candidates can be significantly shortened. Subsequent wet lab experiments can be carried out using only a narrowed list of promising ligands\, hence reducing the overall cost of experiments. \n            AutoDock is one of the most widely used software applications for molecular docking simulations. Its main engine is a Lamarckian Genetic Algorithm (LGA)\, which combines a genetic algorithm and a local-search method to explore several molecular poses. The prediction of the best pose is based on the score\, which is a function that evaluates the free energy (kcal/mol) of a ligand-receptor system. AutoDock is characterized by nested loops with variable upper bounds and divergent control structures. Moreover\, the time-intensive score evaluations are typically invoked a couple of million of times within each LGA run. Based on its computation intensity\, AutoDock suffers from long execution runtimes\, which are mainly attributed to its inability to leverage its embarrassing parallelism. In recent years\, an OpenCL-based implementation of AutoDock has been developed to accelerate its executions on a variety of devices including multi-core CPUs\, GPUs\, and even FPGAs. \n            In this work\, we present our experiences porting and optimizing the OpenCL-based AutoDock onto the SX-Aurora Vector Engine. The OpenCL code is composed of a host and device parts that are maintained in the NEC VEO version. As the API functions of OpenCL and VEOffload resemble each other\, porting the host code was very smooth. While the device part was easily ported too\, an extra effort was required to increase the performance on SX-Aurora. For this\, we used hardware-specific techniques that involve: appropriate data types for wider vectors\, leveraging the multiple cores on the SX-Aurora\, pushing outer into inner loops in score calculations and local search and using multi-process VEO to overcome OpenMP limitations in NUMA mode. Our evaluations were done on VE10B and VE20B models and compared to modern multicore CPUs and GPUs. \n \n \n\n\n\n09:35 – 10:05\nEvaluating and Exploiting the Potential of the Second-generation SX-Aurora TSUBASA\n      Hiroyuki Takizawa\, Tohoku University\n      Abstract \n\n\n\n            In October 2020\, we have started the operation of Supercomputer AOBA\, which employs the second-generation SX-Aurora TSUBASA as the main computing resource. In this talk\, I would like to share the performance evaluation results to discuss the potential of the second-generation SX-Aurora TSUBASA through comparison with the previous generations. I will also introduce our recent research activities for making a good use of its performance while keeping the code portable. \n \n \n\n\n\n10:05 – 10:20\nBreak\n\n\n10:20 – 10:50\nAcceleration of Structural Analysis Software FrontISTR on NEC SX-Aurora TSUBASA\n      Toshiaki Hishinuma\, Research Institute for Computational Science Co. Ltd.\n      Abstract \n\n\n\n            Structural analysis using the finite element method (FEM) is widely used in the field of engineering.\n            Recently\, NEC has introduced SX-Aurora TSUBASA\, which has vector accelerator boards (Vector Engine\, VE).\n            One VE has a high-speed memory with a bandwidth of about 1.2 TB/s and eight high-performance vector cores.\n            Each core has three Fused Multiply-Add (FMA) arithmetic units\, each of which can perform 32 double-precision floating point element executions simultaneously.\n            The host CPU is called VH. \n            FrontISTR is one of the highly parallelized open-source FEM software programs for nonlinear structural analysis.\n            This software firstly generates a stiffness matrix using the FEM and then solves linear equations for the sparse matrix generated by FEM.\n            The stiffness matrix generation is not suitable for VE because it cannot process data continuously\, and it involves many integer operations. \n            There is an API for transferring data between VH and VE called Another/Alternative/Awesome VE Offloading (AVEO)\, which can be used to execute compute-intensive portions of an entire program on VE.\n            We accelerate execution speed of the overall structural analysis program by running the linear equation solvers on VE using AVEO.\n            We chose the JAD format as the sparse matrix storage format and the conjugate gradient (CG) method as the linear solver. \n            In this study\, we evaluate accelerated FrontISTR in terms of the following three parts.\n            (1) the generation time of stiffness matrices on VH\, (2) the transfer time of sparse matrices and vectors from VH to VE using AVEO\, and (3) the calculation time of linear equations by the CG method on VE.\n            We describe the effectiveness of accelerated structural analysis execution on NEC SX-Aurora TSUBASA. \n \n \n\n\n\n10:50 – 11:20\nVGL: a High-Performance Graph Processing Framework for the NEC SX-Aurora TSUBASA Vector Architecture\n      Ilya Afanasyev; Moscow State University\n      Abstract \n\n\n\n            Developing efficient graph algorithms implementations is an extremely important problem of modern computer science\, since graphs are frequently used in various real-world applications. Graph algorithms typically belong to the data-intensive class\, and thus using architectures with high-bandwidth memory potentially allows to solve many graph problems significantly faster compared to modern multicore CPUs. Among other supercomputer architectures\, vector systems\, such as the SX family of NEC vector supercomputers\, are equipped with high-bandwidth memory. However\, the highly irregular structure of many real-world graphs makes it extremely challenging to implement graph algorithms on vector systems\, since these implementations are usually bulky and complicated\, and a deep understanding of vector architectures hardware features is required. We present the world first attempt to develop an efficient and simultaneously simple graph processing framework for modern vector systems. Our vector graph library (VGL) framework targets NEC SX-Aurora TSUBASA as a primary vector architecture and provides relatively simple computational and data abstractions. These abstractions incorporate many vector-oriented optimization strategies into a high-level programming model\, allowing quick implementation of new graph algorithms with a small amount of code and minimal knowledge about features of vector systems. The provided comparative performance analysis demonstrates that VGL-based implementations achieve significant acceleration over the existing high-performance frameworks and libraries: up to 14 times speedup over multicore CPUs (Ligra\, Galois\, GAPBS) and up to 3 times speedup compared to NVIDIA GPU (Gunrock\, NVGRAPH) implementations. \n \n \n\n\n\n11:20 – 11:50\nOptimization of the stencil computation considering the architecture of SX-Aurora TSUBASA\n      Kazuhiko Komatsu\, Tohoku University\n      Abstract \n\n\n\n            This presentation introduces optimizations of the stencil computation on SX-Aurora TSUBASA\, which focuses on the differences of bandwidth characteristics of SX-Aurora TSUBASA.\n \n \n\n\n\n11:50 – 12:20\nComputational Simulation of External Aerodynamics: Evaluation of Performance and Scalability\n      Michael Wagner\, DLR (Deutsches Zentrum für Luft- und Raumfahrt e.V.)\n      Abstract \n\n\n\n            This presentation will share our efforts and experiences in evaluating performance and scalability of CODA on current HPC architectures. CODA is a CFD solver for external aircraft aerodynamics developed by DLR\, ONERA\, and Airbus\, and one of the key next-generation engineering applications represented in the European Centre of Excellence for Engineering Applications (EXCELLERAT).\n \n \n\n\n\n12:20\nOpen Meeting\n\n\n\n\nFriday\, March 19th\, 2021\n\n\n\n\n8:30 – 8:35\nIntroduction of Chair\n\n\n8:35 – 9:05\nPacked Mode Vectorization in LLVM for SX-Aurora\n      Simon Moll (NEC Germany)\n      Abstract \n\n\n\n            The abstract will be provided soon.\n \n \n\n\n\n9:05 – 9:35\nAn Energy-aware Cache Control Mechanism for Deep Cache Hierarchy\n      Ryusuke Egawa(Tokyo Denki University)\, Liu Jiaheng(Tohoku University)\n      Abstract \n\n\n\n            To overcome the memory wall problem\, cache hierarchies of modern microprocessors have become deeper and larger as the number of cores increases. Besides\, the power and energy consumption of the deep cache hierarchy become non-negligible. In this talk\, we present a mechanism to improve cache energy efficiency by adapting a cache hierarchy to individual applications and its evaluation results.\n \n \n\n\n\n09:35 – 10:05\nLoose Coupling of Task-based Programming Models with MPI Through Continuations\n      Joseph Schuchart\, HLRS\n      Abstract \n\n\n\n            Using MPI in combination with asynchronous task-based programming models can be a daunting task. Applications typically have to manage a dynamic set of active operations\, fall-back to a fork-join model\, or rely on some middleware to coordinate the interaction between MPI and the task scheduler. In this talk\, I will propose an extension to MPI\, called MPI Continuations\, that provides a callback-based notification mechanism to simplify the usage of MPI inside asynchronous tasks.\n \n \n\n\n\n10:05 – 10:20\nBreak\n\n\n10:20 – 10:50\nFive challenges of new supercomputing system SQUID in Osaka University\n      Susumu Date\, Cybermedia Center\, Osaka University\n      Abstract \n\n\n\n            The abstract will be provided soon.\n \n \n\n\n\n10:50 – 11:20\nBasics on Quantum Computation\n      Thomas Kloss\, University Grenoble\n      Abstract \n\n\n\n            With the announcement of quantum supremacy in 2019\, Google claimed to have solved the first real-world problem out of reach for classical computers. Since then at the latest\, quantum computing has moved into the political and economic spotlight. In this talk I will present some very basic slides on quantum computers and what makes them different of classical computers. I will also show new work which puts Googles supremacy claim into question. \n \n \n\n\n\n11:20 – 11:50\nAI@HLRS: The Past\, the Present\, and the Future\n      Dennis Hoppe\, HLRS\n      Abstract \n\n\n\n            \n        The growth of artificial intelligence (AI) is accelerating. AI has left research and innovation labs\, and nowadays plays a significant role in everyday lives. The impact on society is graspable: autonomous driving cars produced by Tesla\, voice assistants such as Siri\, and AI systems that beat renowned champions in board games like Go. All these advancements are facilitated by powerful computing infrastructures based on HPC and advanced AI-specific hardware\, as well as highly-optimized AI codes. Since several years\, HLRS is engaged in big data and AI-specific activities around HPC. In this talk\, I will give a brief overview about our AI-focused research project CATALYST to engage with researchers and industry\, present selected case studies\, and outline our journey over the last years with respect to the convergence of AI and HPC from both a software and hardware point of view.\n\n          \n \n \n\n\n\n11:50 – 12:00\nBreak\n\n\n12:00 – 13:00\nNEC Vector Engine Performance with Legacy CFD Codes\n      Keith Obenschain\, NRL (United States Naval Research Laboratory)\n      Abstract\n      Abstract \n\n\n\n            \nMany codes that were developed during the vector supercomputing era from the 1970’s to 1990’s are still in use with vector friendly constructs in their codebase. The recently released NEC Vector Engine provides an opportunity to exploit this vector heritage. The NEC Vector engine can potentially provide state of the art performance without a complete rewrite of the codebase.  Given the time and cost required to port or rewrite codes\, this is potentially an attractive solution. This presentation will assess how the NEC Vector engine performance compares with existing architectures using traditional benchmarks\, a legacy CFD program FDL3DI and the effort required to take full advantage of the architecture.\n        \n          \n \n \n\n\n\n13:00\nOpen Meeting
URL:https://www.wssp.hlrs.de/events/31st-workshop-on-sustained-simulation/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ATTACH;FMTTYPE=image/png:https://www.wssp.hlrs.de/wp-content/uploads/2022/09/featured.png
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20181009
DTEND;VALUE=DATE:20181011
DTSTAMP:20260924T175732
CREATED:20221205T170916Z
LAST-MODIFIED:20221205T172127Z
UID:213-1539043200-1539215999@www.wssp.hlrs.de
SUMMARY:28th Workshop on Sustained Simulation
DESCRIPTION:Agenda\n\n\n\n\n\nTuesday\, 9 October 2018\n\n\n\n\n9:00 – 9:15\nIntroduction\n      Michael Resch\, HLRS\, University of Stuttgart\n\n\n9:15 – 9:45\nExperiences with SX-Aurora Tsubasa and its extension for the future\n      Hiroaki Kobayashi\, Cyberscience Center\, Tohoku University\n      Abstract \n\n\n\n            \n        In my talk\, I would like to share with you some experiences with NEC’s New Vector system named SX-Aurora TSUBASA.  I will also present our on-going project\, named Quantum Annealing-Assisted Next generation HPC infrastructure\, with the extension of SX-Aurora TSUBASA for the future.\n        \n          \n \n \n\n\n\n9:45 – 10:15\nUpdate on the HLRS Strategy\n      Michael Resch\, HLRS\n      Abstract \n\n\n\n            \n         This talk will summarize the research and development activities of HLRS and will highlight future activities and directions\n         \n          \n \n \n\n\n\n10:15 – 10:45\nBreak\n\n\n10:45 – 11:15\nStatus of HPC in Siegen\n      Sabine Roller\, Simulationstechnik & Wissenschaftliches Rechnen\, Universität Siegen\n      Abstract \n\n\n\n \n \n\n\n\n11:15 – 11:45\nHPC\, HPDA\, Machine Learning\, Deep Learning….a glance of the evolution of traditional HPC Centers\n      Bastian Koller\, HLRS\n      Abstract \n\n\n\n            \n         In this talk the current evolution of HPC as a single resource offering towards a set of solutions will be presented\, analysed and some thoughts about the future use of such systems will be given\n         \n          \n \n \n\n\n\n11:45 – 12:15\nData Reduction using Singular Value Decomposition (SVD) Algorithm\n      Jing Zhang\, HLRS\n      Abstract \n\n\n\n            \n         This talk will present the Singular Value Decomposition (SVD) algorithm and  give the results for smaller data sets using SVD as a data reduction  algorithm\, which could perform feature extraction in “raw” data.\n         \n          \n \n \n\n\n\n12:15 – 13:15\nLunch\n\n\n13:15 – 13:45\nPorting Climate Models to Aurora TSUBASA\n      Panos Adamidis\, DKRZ\n      Abstract \n\n\n\n            \n         The main interest of the climate models running at DKRZ focuses on two major directions. On one hand\, high resolution grids are being used in order to resolve small-scale physical processes. In this way\, parametrisation and the inherent uncertainty can be avoided \, thus improving significantly climate change projections. On the other hand addressing questions related to climate variability also involves simulations running over long time periods e.g. modeling of complete glacial cycle on grids with coarser resolution.\n  		 Such simulations are computationally very intensive and high sustained performance is vital in order to be able to conduct real world experiments. Both\, single node performance as well as  good scaling capabilities of the soft- and hardware are important.\n  		 The NEC Aurora TSUBASA System promises high sustained performance by combining high floating point performance of vector processors with extremely high memory bandwidth.  The presentation will show first results from our tests with  earth system models on the NEC Aurora system at DKRZ.\n         \n          \n \n \n\n\n\n13:45 – 14:15\nPerformance evaluation and analysis of SX-Aurora TSUBASA\n      Kazuhiko Komatsu\, Cyberscience Center\, Tohoku University\n      Abstract \n\n\n\n            \n         A new vector supercomputer\, SX-Aurora TSUBASA\, has been released. It has a newly developed Vector Engine(VE) processor to achieve a high sustained performance by powerful vector processing and a high memory bandwidth. This presentation examines the basic potential of SX-Aurora TSUBASA through the performance evaluations.\n         \n          \n \n \n\n\n\n14:15 – 14:45\nPerformance of a DNS code on SX-Aurora TSUBASA\n      Mitsuo Yokokawa\, Kobe University\n      Abstract \n\n\n\n           \n         Direct numerical simulations (DNSs) of turbulence at high Reynolds number are very important to understand behavior of turbulent flow and to establish tubulent models. We have carried out large-scale DNSs for more than a decade on the Earth Simulator and the K compute and would like to execute DNSs with larger grid points beyond the present grid points on future supercomputers. In this talk\, the first evaluation results of a DNS code performance on SX-Aurora TSUBASA will be presented. In paticular\, off-loaded I/O performace for checkpoints of instantanious velocity fields from a vector engin (VE) to a vector host (VH) as well as computing performace on VEs will be included.\n         \n          \n \n \n\n\n\n14:45 – 15:15\nOrganizing MPI parallel Simulations\n      Harald Klimach\, Uni Siegen\n      Abstract \n\n\n\n            \n        MPI parallel simulations provide some challenges when dealing with user interaction. We present:\n* a method to obtain configuration settings from Lua scripts in a scalable way\,\n* a strategy to manage logging output during the parallel execution with configurable level of detail\,\n* a concept to deal with errors detected by the application at runtime.\nThese components are put together into a Fortran library to build a basic infrastructure for parallel simulation applications. It relies on Fypp as a pre-processing tool\, which allows the usage Python to generate Fortran code.\n         \n          \n \n \n\n\n\n15:15 – 15:45\nBreak\n\n\n15:45 – 16:15\nAccelerating Heatstroke Risk Simulation on Modern Vector Systems\n      Ryusuke Egawa\, Cyberscience Center\, Tohoku University\n      Abstract \n\n\n\n            \n         TBD\n         \n          \n \n \n\n\n\n16:15 – 16:45\nVector Engine Processor and 2D vector function\n      Shintaro Momose\, NEC Germany\n      Abstract \n\n\n\n            \n         TBD\n         \n          \n \n \n\n\n\n16:45 – 17:15\nAurora SW Update\n      Masashi Ikuta\, NEC Tokyo\n      Abstract \n\n\n\n            \n         TBD\n         \n          \n \n \n\n\n\n17:15 – 17:45\nLimits of sustained performance and remedies\n      Uwe Küster\, NEC Germany\n      Abstract \n\n\n\n            \n         With limited processor frequency any performance increase results from parallelism. But the startup to fill the operating units decreases the effective performance if the executed kernels are not large. We discuss this effect by measurements on Aurora. A remedy could be in implementing bricks of kernels in hardware.\n         \n          \n \n \n\n\n\n19:00 –\nDinner in Goldener Adler\n\n\n\n\nWednesday\, 10 October 2018\n\n\n\n\n9:00 – 9:30\nDeep Neural Networks for Data-Driven Turbulence Models\n      Andrea Beck\, IAG Stuttgart\n      Abstract \n\n\n\n            \n        In this talk\, we present a novel data-based approach to turbulence modelling for Large Eddy Simulation by artificial neural networks. We define the exact closure terms including the discretization operators and generate training data from direct numerical simulations of decaying homogeneous isotropic turbulence. We design and train artificial neural networks based on local convolution filters to predict the underlying unknown non-linear mapping from the coarse grid quantities to the closure terms without a priori assumptions. All investigated networks are able to generalize from the data and learn approximations with a cross correlation of up to 47% and even 73% for the inner elements\, leading to the conclusion that the current training success is data-bound. We further show that selecting both the coarse grid primitive variables as well as the coarse grid LES operator as input features significantly improves training results. Finally\, we construct a stable and accurate LES model from the learned closure terms. Therefore\, we translate the model predictions into a data-adaptive\, pointwise eddy viscosity closure and show that the resulting LES scheme performs well compared to current state of the art approaches. This work represents the starting point for further research into data-driven\, universal turbulence models.\n        \n          \n \n \n\n\n\n9:30 – 10:00\nAn object oriented multiphysics simulation concept for HPC\n      Matthias Meinke\, Lennart Schneiders\, Michael Schlottke-Lakemper\, Wolfgang Schroeder\, Institute of Aerodynamics\, RWTH Aachen Unviersity\, Aachen\, Germany\n      Abstract \n\n\n\n            \n        A simulation framework for a generalized multiphysics simulation concept is introduced which is based on an implementation of various solvers formulated for Cartesian hierarchical meshes. The implementation features a generalized mesh object which communicates with various solvers based on finite-volume or discontinous Galerkin methods. SInce all solution methods share a common mesh\, solution adaptive meshes with dynamic load balancing are straighforward to implement. Interleaved time stepping of the solvers for the different physics allow an efficient implemention on HPC systems. Examples are presented for the coupling of various solution methods for flow simulations\, acoustic fields and level sets used for tracking moving surfaces.\n        \n          \n \n \n\n\n\n10:00 – 10:30\nMoving geometries in high-order discontinuous Galerkin discretization\n      Neda Ebrahimi Pour\, Uni Siegen\n      Abstract \n\n\n\n           \n         Representing geometries in high-order schemes is a crucial task with special requirements. An attractive solution to this problem is the employment of penalizing terms to represent the geometry within elements. This approach also allows for a convenient movement of obstacles through flows for example\, as it avoids the need for expensive remeshing and interpolations. We present this concept in our high-order discontinuous Galerkin solver Ateles and show first results for compressible flows.\n         \n          \n \n \n\n\n\n10:30 – 11:00\nBreak\n\n\n11:00 – 11:30\nTowards Performance and Power Model for Multi-Core Processors with DVFS\n      Dmitry Khabi\, HLRS\n      Abstract \n\n\n\n            \n         A significant part of homogeneous supercomputer consists of an extensive number of general-purpose processors (CPU)\,which are connected with each other over a high-performance network. The degree of parallelism of these high-performance computing (HPC)platforms is generally limited to the number of processor cores. The scalability and performance enhancement almost exclusively stems fromthe growing number of CPU cores\, although that number no longer meets the constantly expanding HPC requirements.The growing number of CPU cores is to be seen not only as a simple increasing number of cores but also as the additionaloverhead in a distribution of the rest of the hardware resoursers\, such as the network\, last level cache\, memory channels\, power budget\, etc.\,between the growing number of the cores and hence also between the parallel processes and threads. Particulary in combination with the capabilitiesof the processor to change the operating voltage and frequency\, so-called “Dynamic Voltage and Frequency Scaling (DVFS)”\,the analysis of the scalability and energy efficienty of a multi-core processor even on the basis of the existing models\,such as “Roofline Performance Model” or “Execution-Cache-Memory Model” (ECM)\, is additionally complicated.The performance and power dissipation of CPU and DRAM are in complex interaction with the number of the active cores and the CPU frequency.This talk presents an extension of ECM model (hereinafter referred as DTM – “Data Transfer Model”)\, which describes the performance taking intoconsideration the various frequencies of the hardware components. The evaluation of the model using the “STREAM” kernels (with temporally memory access)is performed on different hardware architectures.\n         \n          \n \n \n\n\n\n11:30 – 12:00\nA method to reduce load imbalances in simulations of phase change processes with FS3D\n      Johannes Müller\, Martin Reitzle\, Institute of Aerospace Thermodynamics (ITLR)\, University of Stuttgart\, Philipp Offenhäuser\, High-Performance Computing Center Stuttgart (HLRS)\n      Abstract \n\n\n\n           \n         Numerical simulations of phase change processes require a precise reconstruction of the interface between two phases. Based on the Volume of Fluid (VoF) method for multiphase flows the height function technique is able to reconstruct the sharp interface accurately and enables simulations with complex interface deformations. But this calculations increase the computational load for cells containing the phase interface. An equidistant domain decomposition leads to an imbalanced workload distribution. In order to perform investigations with a high spatial and temporal resolution\, it is necessary to use the available HPC resources efficiently. The challenge of parallelization is to distribute the workload homogeneously among the cores. For simulations of the solidification processes with the multiphase code Free Surface 3D (FS3D)\, a load balanced domain decomposition is presented. The first part is the decomposition of the structured computational domain by recursive bisection. The second part is the corresponding process communication\, which enables a nearest neighbor communication through non-blocking MPI calls. The transport of the diagonal element is realized via a communication sequence and thus an exchange of small amounts of data is avoided. A measure for the load imbalance is presented based on test cases. Finally\, advantages and limitations of load balancing are discussed based on the tracing of the calculations for one timestep.\n         \n          \n \n \n\n\n\n12:00 – 12:30\nOptimization and Parallelization of a Phase-Field Solver to Investigate Sintering Processes\n      J. Hötzer\, H. Hierl\, M. Seiz\, M. Kellner\, B. Nestler\, IAM-CMS / KIT\n      Abstract \n\n\n\n            \n         Simulations allow to improve the development of new high performance materials with tailored microstructures and defined properties. The process of sintering is of high interest in order to produce defined ceramic materials as needed for a broad range of applications e.g. healthcare\, electronics\, automotive and aerospace. The phase-field method allows to efficiently investigate the microstructure evolution in large scale 3D domains during the sintering process. A phase-field model based on the grand potential approach is implemented in the massively parallel phase-field solver framework PACE3D. It is optimized on various levels starting from the model and parameters down to the hardware. The solver allows to resolve and calculate an arbitrary number of individual particles in the green body by using a local reduction technique and a material class based parametrization concept. The evolution equations for the phase-fields and the concentration are explicitly vectorized using vector intrinsics. Performance results on a single core\, single node and with up to 96100 processes on the German supercomputers Hazel Hen\, SuperMuc and ForHLR II are shown and discussed. Besides an optimized voxel based format to store the simulation data for checkpointing with MPI-IO\, an efficient and reduced mesh based output is used.\n         \n          \n \n \n\n\n\n12:30 – 13:30\nLunch\n\n\n13:30 – 14:00\nThe potential of MPI shared-memory model for supporting hybrid communication scheme\n      Huan Zhou\, HLRS\n      Abstract \n\n\n\n            \n         This talk first brings the MPI shared-memory model forward and then illustrates the usages of this model via two use cases. One is the hybrid RMA. It is fully employed in the DART-MPI. The other is the hybrid collectives (to be specific\, allgather operation). These two use cases highlight the performance benefit brought by the appliance of MPI shared-memory model. However\, extra synchronization operations should be added  to guarantee a deterministic behaviour.\n         \n          \n \n \n\n\n\n14:00 – 14:30\nFine-Grained Synchronization using Global Task Dependencies in DASH\n      J. Schuchart\, J.Gracia\, HLRS\n      Abstract \n\n\n\n            \n         The current usage of MPI communication operations leads to a global synchronization across many processes and compute nodes. The problem becomes more severe when combining MPI with a thread-parallel programming model such as OpenMP: synchronization latencies are paid manyfold by all threads within an MPI processes. We present ongoing work to address this problem by implementing a task-based programming model which allows to express dependencies across MPI processes. This kind of fine-grained synchronization can replace global MPI synchronization in many cases and thus result in substantially improved communication efficiency.\n         \n          \n \n \n\n\n\n14:30 – 15:00\nAutomatic Parameter Tuning for Efficient Checkpointing\n      Hiroyuki Takizawa\, Muhammad Alfian Amrizal\, Kazuhiko Komatsu\, and Ryusuke Egawa\, Graduate School of Information Sciences\, Tohoku University\n      Abstract \n\n\n\n            \n         One of the most intensive I/O operations of a scientific simulation is so-called checkpointing\, which is to save the state of a running simulation into a checkpoint file so that the simulation can be resumed from the file upon a system failure. Generally\, it is difficult to increase the I/O performance of a system at the same pace as its computational performance. Moreover\, it is necessary for a future system to perform checkpointing more frequently because the system will consist of more hardware components and hence the probability of encountering a system failure during the simulation will significantly increase. As a result\, the overhead for checkpointing will be relatively growing\, and could dominate the total simulation time. In this talk\, therefore\, I will discuss the possibility and potential benefit of employing automatic parameter tuning for reducing the checkpointing overheads.\n         \n          \n \n \n\n\n\n15:00 – 15:30\nBreak\n\n\n15:30 – 16:00\nNEC SX-Aurora TSUBASA and the LLVM ecosystem\n      Simon Moll\, Compiler Design Lab\, Saarland University\n      Abstract \n\n\n\n            \n         The NEC SX-Aurora TSUBASA is a high-performance vector CPU for sustained simulation performance. The existing compiler toolchain for the SX-Aurora is comprehensive but also proprietary restricting its use in research and confining its development to internal teams at NEC. In recent years\, the open source LLVM compiler infrastructure has seen significant support and contributions by major players such as NVIDIA\, AMD\, ARM\, Intel\, Apple and Google. These employ LLVM in their official toolchains\, GPU driver stacks and mission-critical infrastructure. Likewise\, many compiler research labs have adopted LLVM for its accessibility\, robustness and permissive license. Recently\, the LLVM community has been discussing an extension for scalable vector architectures (LLVM-SVE)\, which feature an active vector length just as the SX-Aurora does.  In this talk\, we will discuss the potential of LLVM for the NEC SX-Aurora. The Compiler Design Lab at Saarland University is working with NEC on an LLVM-SVE backend for the SX-Aurora.\n         \n          \n \n \n\n\n\n16:00 – 16:30\nApproach to provide supercomputer storage I/O information toward users\n      Tsuyoshi Nakagawa\, JAMSTEC\n      Abstract \n\n\n\n            \n         The role of supercomputer storage on JAMSTEC becomes increasingly important not only for simulation but also for data driven science. The information such as the storage I/O performance and the file system characteristics may improve the user’s availability and make new scientific knowledge and innovation. In this talk\, we introduce our approach to provide I/O information toward users.\n         \n          \n \n \n\n\n\n16:30 – 17:00\nJob Scheduler Simulator Extension For Evaluating Queue Mapping to Computing Node\n      Susumu Date\, Yuki Matsui\, Yasuhiro Watashiba\, Tatashi Yoshikawa\, Shinji Shimojo\, Cybermedia Center\, Osaka University\n      Abstract \n\n\n\n            \n        TBD\n        \n          \n \n \n\n\n\n17:00 – 17:30\nFarewell
URL:https://www.wssp.hlrs.de/events/28th-workshop-on-sustained-simulation/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ATTACH;FMTTYPE=image/png:https://www.wssp.hlrs.de/wp-content/uploads/2022/09/featured.png
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
END:VCALENDAR