BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//WSSP - ECPv6.15.20//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:WSSP
X-ORIGINAL-URL:https://www.wssp.hlrs.de
X-WR-CALDESC:Events for WSSP
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Europe/London
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20210328T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20211031T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20220327T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20221030T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20230326T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20231029T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20240331T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20241027T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0000
TZOFFSETTO:+0100
TZNAME:BST
DTSTART:20250330T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0100
TZOFFSETTO:+0000
TZNAME:GMT
DTSTART:20251026T010000
END:STANDARD
END:VTIMEZONE
BEGIN:VTIMEZONE
TZID:Europe/Berlin
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20240331T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20241027T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20250330T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20251026T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20260329T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20261025T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20270328T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20271031T010000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;VALUE=DATE:20220523
DTEND;VALUE=DATE:20220525
DTSTAMP:20260924T234708
CREATED:20220926T181045Z
LAST-MODIFIED:20221204T144319Z
UID:71-1653264000-1653436799@www.wssp.hlrs.de
SUMMARY:33rd Workshop on Sustained Simulation
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST). An additional live stream will be available. Please contact Mr. Johannes Gebert (gebert@hlrs.de) if you would like to participate. \n\n\n\n\n\nMonday\, May 23rd\, 2022\n\n\n\n\n10:00 – 10:15\nWelcome & Introduction\n      Michael Resch\, HLRS\, University of Stuttgart\n\n\n10:15 – 10:45\nReaggregation of Disaggregation:\n      A Smart Approach to the Optimized Architecture and Platform Design\n      Hiroaki Kobayashi\, Tohoku University\n\n\n10:45 – 11:15\nCombinatorial Clustering for a Material Informatics Application using Aurora Vector Annealing\n      Kazuhiko Komatsu\, Tohoku University\n      Abstract \n\n\n\n            Due to the recent advancement of data science\, such as machine learning and big-data analysis\, the approach using data science techniques has attracted attention even to develop new materials\, called material informatics. In material informatics\, clustering is one of the essential data processing techniques to understand thermophysical properties. To improve clustering accuracy\, this presentation gives an Ising-based clustering method using aurora annealing machine for a material informatics application. \n \n \n\n\n\n11:15 – 11:45\nHPC Refactoring Catalog : Updates\n      Ryusuke Egawa\, Tokyo Denki University\n      Abstract \n\n\n\n            Aiming to support smooth code migration and optimization among HPC systems\, HPC Refactoring\, a database of system-aware code optimization patterns\, was designed in 2015. This talk introduces an overview and updates/future plans for improving the HPC refactoring catalog. \n \n \n\n\n\n11:45 – 13:15\nLunch\n\n\n13:15 – 13:45\nConcept of a File Tracing Mechanism for Research Data Management in High Performance Computing System\n      Yuta Namiki\, Takeo Hosomi\, AKihiro Yamashita\, Susumu Date\, Joint Research Laboratory for Integrated Infrastructure of High Performance Computing and Data Analysis\, Cybermedia Center\, Osaka University\n      Abstract \n\n\n\n            Research data management has come to take a role of great importance for reproducible and reusable research. In the recent academic research scene\, researchers and scientists leverage HPC systems as a base for processing and analyzing a large amount of data through numerical analysis and computer simulations. Despite that\, however\, the data produced and analyzed on HPC systems are not managed due to the lack of a functionality that allows us to understand how data is produced and processed in the system. From the perspective\, by focusing on the lineage of data\, which shows the origins and history of data for assuring that the data is produced on HPC systems. we have prototyped a mechanism that generates the lineage by tracing file access operations of user programs in the kernel. In this presentation\, we report the interim result we have achieved so far with future issues. \n \n \n\n\n\n13:45 – 14:15\nMitigation of Aeroacoustic Noise based on Simulations on HPC Systems\n      Matthias Meinke\, Ansgar Niemoeller\, Miro Gondrum\, Moritz Waldmann\, and Wolfgang Schroeder\, AIA\, RWTH Aachen\n      Abstract \n\n\n\n            The numerical simulation of aeroacoustic sound is important for an improved understanding of noise generation mechanisms and the design of noise mitigation strategies. In this paper\, the performance of two direct coupled two-step CFD/CAA methods implemented on HPC hardware are discussed. For the flow field either a finite-volume method for the solution of the Navier-Stokes equations or a lattice Boltzmann method is coupled to a discontinuous Galerkin method for the solution of the acoustic pertubation equations. The coupling takes advantage of a joint Cartesian mesh allowing for the exchange of the acoustic sources without MPI communication. An immersed boundary treatment of the acoustic scatttering from solid bodies by a novel solid wall formulation is implemented and validated in the DG method. Results for the case of a spinning vortex pair and the low Reynolds number unsteady flow around a circular cylinder show that a solution with comparable accuracy is obtained for the two direct hybrid methods when using identical mesh resolution. Finally\, results of a large scale application\, i.e.\, the noise prediction for a nose landing gear are presented. \n \n \n\n\n\n14:15 – 14:45\nManagement of data flows between Cloud\, HPC and IoT/Edge\n      Kamil Tokamov\, HLRS\, University of Stuttgart\n      Abstract \n\n\n\n            The components of heterogeneous applications are deployed across various execution platforms and utilise the capabilities of the platforms. As such\, one component can utilise HPC resources for better performance in batch computation\, while another – Cloud resources\, for better scalability and elasticity. Furthermore\, this is also a possibility for processing on Edge devices. The usage of such a hybrid setup\, where dependent components of the applications are deployed across various platforms\, might require flexible and adaptive data transfers from one platform to another. This work presents a data management framework\, based on the Apache NiFi dataflow management system and developed in the scope of the SODALITE EU project. This framework enables scalable data transfer between any of GridFTP (a file transfer protocol dominant in HPC)\, HTTP\, S3-compatible and data streaming (such as MQTT) endpoints. \n \n \n\n\n\n14:45 – 15:15\nBreak\n\n\n15:15 – 15:45\n\n      Sabine Roller\, DLR (Deutsches Zentrum fÃ¼r Luft- und Raumfahrt e.V.)\n      Abstract \n\n\n\n            The abstract will be provided soon. \n \n \n\n\n\n15:45 – 16:15\nIntegration of parallel HDF5 I/O in a large scale computational fluid dynamics solver\n      Tobias Gibis\, University of Stuttgart\n      Abstract \n\n\n\n            As the number of cores in massively parallel computer systems increases\, I/O strategies must be adapted as to not present bottlenecks. The “read and write one file per core” strategy\, while efficient for smaller earlier computational architectures\, leads to an unmanageable amount of files and is poorly suited for Lustre filesystems. With the introduction of Hawk at HLRS\, an urgent need had arisen to develop a new framework based on MPI-I/O in which all cores write simultaneously to a common file. To this end\, a new I/O framework based on HDF5 was implemented in the IAG in-house CFD code NS3D. The associated talk will discuss selected design decisions and challenges encountered\, escpecially regarding adequate I/O performance.\n          \n \n \n\n\n\n16:15 – 16:45\nHigh Performance Object-Oriented Data Processing Workflows for Researchers and Scientists\n      Jason Appelbaum \, University of Stuttgart\n      Abstract \n\n\n\n            As computing capacity increases\, datasets generated by HPC applications grow in size as well. The researchers and scientists who use such datasets for their work require scalability and efficiency in their data-processing workflows\, but still prioritize utility and practicality. Typically such researchers are self-taught\, intermediate-level programmers working collaboratively with others\, in which case object-oriented languages such as python offer a greatly reduced barrier to entry and improved code maintainability within their research groups. Research progress is accelerated by flexible\, easy access to large datasets for comparison and testing. The merits of a high-performance yet flexible approach utilizing HDF5 and MPI\, wrapped by h5py and mpi4py in python\, will be discussed. Examples taking advantage of collective I/O and parallel processing for the analysis of direct numerical simulation datasets will be presented\, along with performance metrics. Additionally\, techniques for ‘big data’ visualization using Paraview\, HDF5 and XDMF will be showcased. \n \n \n\n\n\n18:30\nDinner (Registration closed)\n\n\n\n\nTuesday\, May 24th\, 2022\n\n\n\n\n09:00 – 09:45\nKeynote\n      Jack Dongarra \, University of Tennessee\n\n\n09:45 – 10:15\nA ML-Based Approach to Automatic Selection of Compiler and its Option Flags\n      Hiroyuki Takizawa\, Tohoku University\n      Abstract \n\n\n\n            Today\, one HPC platform could have multiple compilers\, each of which provides a lot of option flags. Those compilers have different optimization capabilities\, and target at even different processors on a heterogeneous computing system such as NEC SX-Aurora TSUBASA. Thus\, it could be challenging to select an appropriate build configuration such as the best available compiler and its option flags for each application code. In this talk\, I will introduce our ongoting work to use machine learning for predicting an appropriate build configuration from performance counter values. \n \n \n\n\n\n10:15 – 11:00\nBreak\n\n\n11:00 – 11:30\nSpeeding up k-nearest neighbors search with space-filling curve.\n      Masashi Kotera\, Sourav Saha\, Takuya Araki\, NEC Corporation\n      Abstract \n\n\n\n            k-nearest neighbors search (k-NN) is a useful algorithm that can be used for classification and regression\, but naive k-NN is slow because it requires scanning all the training data at prediction time. There is a method of solving this problem for low dimensional data\, such as dividing the search space Like kd-tree. However\, algorithms using tree structures are difficult to vectorize because they require recursion. In this talk\, we talk about implementing the speed-up method for k-NN using z-curve\, a kind of space-filling curve; the calculation of z-curve is easy to vectorize\, and z-curve enables efficient range search that can be used to implement k-NN. \n \n \n\n\n\n11:30 – 12:00\nHeterogeneous Computing with SX-Aurora TSUBASA Vector Engine\n      Ryota Ishihara\, NEC Corporation\n      Abstract \n\n\n\n            SX-Aurora TSUBASA supports a variety of execution models in heterogeneous environments including various computational resources such as GPU and x86. User can select appropriate computational resources according to characteristics of each of applications in the executions. In this session\, we will introduce the functions provided in each execution model and how to use them. \n \n \n\n\n\n12:00 – 13:30\nLunch\n\n\n13:30 – 14:00\nTrends\n      Michael Resch\, HLRS\, University of Stuttgart\n\n\n14:00 – 14:30\nSearching a roadmap to solve partial differential equations with quantum machine learning\n      Markus Mieth\, Pia Siegl\, DLR (German Aerospace Center)\n      Abstract \n\n\n\n            Our research aims to evaluate the potential of quantum computing to solve partial differential equations (PDEs) in the context of aerospace engineering. Established algorithms and methods for PDEs often rely on discretization in time and space. For a reasonable accuracy they come with high computational costs in terms of time and memory space. Machine learning approaches are studied as an alternative. One promising concept is the physical informed neural network (PINN) [1]. Here\, the PDE is directly included into the loss function such that no data or only a limited amount is needed for training. In our approach\, we exchange the classical neural network of the PINN with a trainable quantum circuit\, while the optimization still runs on a classical computer. While we can already approximate simple functions successfully with known strategies [2\, 3]\, more complex PDEs are hard to solve. Our work focuses on the search of problem-oriented quantum circuits and data encoding strategies\, which increase the expressibility of the quantum model and allow for the approximation of more complex functions. \n \n \n\n\n\n14:30 – 15:00\nPrediction of Bio-Hybrid Fuel Injection and Mixture Formation in an Internal Combustion Engine\n      Tim Wegmann\, Matthias Meinke\, Wolfgang Schroeder\, AIA\, RWTH Aachen\n      Abstract \n\n\n\n            For an efficient\, stable and low emission combustion of novel e-fuels in piston engines\, the fuel distribution at start of ignition plays a crucial rule. The injection system and fuel properties define the initial fuel vapor distribution. The subsequent fuel-air mixing depends on the convection\, turbulence intensities\, and the formation and break-up of large-scale flow structures\, in particular the tumble vortex. Large-eddy simulations (LES) with high mesh resolution are necessary to accurately predict all involved scales for the mixing process. In this study\, numerical analyses of the liquid fuel injection and the fuel-air mixing in a piston engine are performed. LES are conducted using a hierachical unstructured Cartesian mesh method with an efficient four-way coupling of the spray droplets with the gas phase. Due to the large number of spray droplets\, a Lagrangian Particle Tracking (LPT) algorithm is used to accurately predict the liquid spray propagation and evaporation. The spray model is based on a KHRT-breakup formulation. The Navier-Stokes equations are solved for compressible flow using a finite-volume method\, where boundary surfaces are represented by a conservative cut-cell method. The hierarchical Cartesian mesh ensures efficient use of high performance computing platforms through solution adaptive refinement and dynamic load balancing. \n \n \n\n\n\n15:00 – 15:30\nBreak\n\n\n15:30 – 16:00\nQuantum machine learning for data analysis\n      Li Zhong\, HLRS\, University of Stuttgart\n      Abstract \n\n\n\n            Fault-tolerant quantum computers have been proven to be able to improve machine learning through speed-ups in computation orimproved model scalability. Therefore\, research at the junction of the two fields has garnered an increasing amount of interest\, which has led to the rapid development of quantum deep learning and quantum-inspired deep learning techniques. In thiswork\, we will demonstrate how quantum computers and quantum algorithms can be leveraged for image processing through quantum-inspired deep neural networks. \n \n \n\n\n\n16:00\nOpen End
URL:https://www.wssp.hlrs.de/events/33rd-workshop-on-sustained-simulation/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ATTACH;FMTTYPE=image/png:https://www.wssp.hlrs.de/wp-content/uploads/2022/09/featured.png
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=Europe/London:20230113T080000
DTEND;TZID=Europe/London:20230113T170000
DTSTAMP:20260924T234708
CREATED:20230113T120530Z
LAST-MODIFIED:20230113T120530Z
UID:219-1673596800-1673629200@www.wssp.hlrs.de
SUMMARY:26th Workshop on Sustained Simulation
DESCRIPTION:Agenda\n\n\n\n\n\nTuesday\, 10 October 2017\n\n\n\n\n9:00 – 9:15\nIntroduction\n      Michael Resch\, HLRS\, University of Stuttgart\n\n\n9:15 – 9:45\nTwo-Year Experiences with Vector Supercomputer SX-ACE and Design Space Exploration of the Next Generation Vector System\n      Hiroaki Kobayashi\, Cyberscience Center\, Tohoku University\n      Abstract \n\n\n\n             \n          In this talk\, I will be presenting two-year experiences with our brand-new vector-parallel supercomputer SX-ACE.  In particular\, we will show you operation statistics\, applications developed on SX-ACE\, and some case study of program tuning to exploit its potential.   In addition\, I will describe the future plan for supercomputing resource installation and deployment at Tohoku University and make some discussion on the design space exploration of the future vector system. \n            \n          \n \n \n\n\n\n9:45 – 10:15\nSiVeGCS – The Future of German Supercomputing\n      Michael Resch\, HLRS\n      Abstract \n\n\n\n            \n         This talk will summarize the research and development activities of HLRS and will highlight future activities and directions\n         \n          \n \n \n\n\n\n10:15 – 10:45\nBreak\n\n\n10:45 – 11:15\nJAMSTEC Next Scalar Supercomputer System\n      Ken’ichi Itakura\, JAMSTEC\n      Abstract\n\n\n11:15 – 11:45\nOCTOPUS: a new supercomputing service of Osaka University\n      Susumu Date\, Cybermedia Center\, Osaka University\n      Abstract\n\n\n11:45 – 12:15\nThe Brand-new Vector Supercomputer\, Aurora\n      Shintaro Momose\, NEC\n      Abstract\n\n\n12:15 – 13:15\nLunch\n\n\n13:15 – 13:45\nA Multiple-layer Bypass Mechanism for Energy Efficient Computing\n      Ryusuke Egawa\, Masayuki Sato\, Ryoma Saito\, Hiroaki Kobayashi\, Cyberscience Center\, Tohoku University\n      Abstract\n\n\n13:45 – 14:15\nCoupling Strategies for Multiphysics Simulations on Hierarchical Cartesian Meshes\n      Matthias Meinke\, Michael Schlottke\, Ansgar Niemöller\, Institute of Aerodynamics\, RWTH Aachen University\n      Abstract\n\n\n14:15 – 14:45\nLocally Linearized Euler Equations in Discontinuous Galerkin with Legendre Polynomials\n      H. Klimach\, M. Gaida\, S. Roller\, Simulationstechnik & Wissenschaftliches Rechnen\, Universität Siegen\n      Abstract\n\n\n14:45 – 15:15\nUnveiling Insight on Fluid Systems in a Diverse Environment using CFD\n      Manuel Hasert\, Festo AG & Co. KG\n      Abstract\n\n\n15:15 – 15:45\nbreak\n\n\n15:45 – 16:15\nHigh-fidelity Simulation of Helicopter Phenomena HPC aspects in advanced engineering applications\n      Manuel Keßler\, Institute for Aero and Gasdynamics University of Stuttgart\n      Abstract\n\n\n16:15 – 16:45\nHighly portable CFD solutions for heterogeneous computing on unstructured meshes\n      A.V. Gorobets\, S.A.Soukov\, Keldysh Institute of Applied Mathematics of RAS\, Moscow\, Russia\n      P.B. Bogdanov\, Scientific Research Institute of System Development of RAS\, Moscow\, Russia\n      X. Alvarez\, F.X.Trias\, Heat and Mass Transfer Technological Center of UPC\, Barcelona\, Spain\n      Abstract \n      \n\n\n16:45 – 17:15\nNumerical modelling of phase change processes in clouds – challenges and approaches\n      Martin Reitzle\, Bernhard Weigand\, Institute of Aerospace Thermodynamics\, University of Stuttgart\n      Abstract\n\n\n17:15 – 17:45\nA dynamic load-balancing strategy for large scale CFD-applications\n      Philipp Offenhäuser\, HLRS\n      Abstract\n\n\n19:00 – 21:00\nDinner in Goldener Adler\n\n\n\n\nWednesday\, 11 October 2017\n\n\n\n\n9:00 – 9:30\nAPI Extension and Resource Manager Integration for Malleable MPI Applications\n      Isaias Alberto Compres Urena\, Institute of Informatics\, Technical University of Munich\n      Abstract\n\n\n9:30 – 10:00\nPerformance and Quality Analysis of Interpolation Methods for Coupling\n      N. Ebrahimi-Pour\, S. Roller\, Simulationstechnik & Wissenschaftliches Rechnen\, Universität Siegen\n      Abstract\n\n\n10:00 – 10:30\nTowards Realizing a Dynamic and MPI Application-aware Interconnect with SDN\n      Keichi Takahashi\, Cybermedia Center\, Osaka University\n      Abstract\n\n\n10:30 – 11:00\nBreak\n\n\n11:00 – 11:30\nFEniCS HPC: An automated predictive high-performance framework for multiphysics simulations\n      Niclas Jansson\, Department of High Performance Computing and Visualization\, School of Computer Science and Communication\, KTH Royal Institute of Technology\n      Abstract\n\n\n11:30 – 12:00\nAutomated derivation and parallel execution of finite difference models on CPUs\, GPUs and Intel Xeon Phi processors using code generation techniques\n      Christian T. Jacobs\, Satya P. Jammy\, David J. Lusher\, Neil D. Sandham\, Engineering and the Environment at the University of Southampton\n      Abstract \n      \n\n\n12:00 – 12:30\nPerformance tuning of Ateles using Xevolver\n      Kazuhiko Komatsu\, Cyberscience Center\, Tohoku University\n      Abstract\n\n\n12:30 – 13:30\nLunch\n\n\n13:30 – 14:00\nvTorque – Introducing virtualization capabilities to Torque\n      Nico Struckmann\, HLRS\n      Abstract\n\n\n14:00 -14:30\nOptimised scheduling mechanisms for Virtual Machine deployment in Cloud infrastructures\n      Michael Gienger\, HLRS\n      Abstract\n\n\n14:30 – 15:00\nTo be defined\n      Christopher L. Barrett\, Biocomplexity Institute\, Virginia Tech\n      Abstract\n\n\n15:00 – 15:30\nBreak\n\n\n15:30 – 16:00\nSoftware for agent based social simulation in the distributed HPC environments\n      Sergiy Gogolenko\, HLRS\n      Abstract\n\n\n16:00 – 16:30\nA parallel solver for a linear system with a symmetric sparse matrix by one-dissection ordering\n      Mitsuo Yokokawa\, Tomoki Nakano\, Kobe University\n      Takeshi Fukaya\, Hokkaido University\n      Yusaku Yamamoto\, The University of Electro-Communications\n      Abstract\n\n\n16:30 – 17:00\nVistle\, a scalable visualization system for immersive virtual environments\n      Martin Aumüller\, Uwe Wössner\, HLRS\n      Abstract\n\n\n17:00 – 17:30\nFarewell
URL:https://www.wssp.hlrs.de/events/26th-workshop-on-sustained-simulation/
LOCATION:Baden-Württemberg
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20230413
DTEND;VALUE=DATE:20230415
DTSTAMP:20260924T234708
CREATED:20230203T095652Z
LAST-MODIFIED:20230704T074522Z
UID:229-1681344000-1681516799@www.wssp.hlrs.de
SUMMARY:35th Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \nThe registration is closed. \n\n\n\n\n\nThursday\, April 13\, 2022\n\n\n\n\n10:00 – 10:15\nWelcome & Introduction\nMichael Resch\, HLRS\, University of Stuttgart\n\n\n10:15 – 10:45\nLessons Learned from A Quantum-Annealing Assisted HPC R&D Project\nHiroaki Kobayashi\, Tohoku University\n\n\n10:45 – 11:15\nNEC’s Quantum Computing Strategy\, Technology\, and Use Cases\nShintaro Momose\, NEC Corporation \n\n\n\nThis presentation consists of two parts\, discussing SX-Aurora TSUBASA vector supercomputer and introducing simulated annealer working on SX-Aurora TSUBASA called Aurora Vector Annealing. The first half of the presentation shows the vector architecture of SX-Aurora TSUBASA\, especially its latest vector processors having the highest-level memory  bandwidth. Sustained performance and power efficiency are also discussed\, as well as NEC’s future plans and roadmap. The second half of the presentation shows NEC’s quantum computing strategies and their products to provide higher sustained performance in the annealing/optimization fields. NEC developed the Aurora Vector Annealing as a simulated annealer and has a strong business relationship with D-Wave providing a quantum annealer. NEC aims at solving various social issues by using the quantum/simulated annealing technologies and by developing a hybrid platform with supercomputer and quantum/simulated annealer to provide much higher sustained performance. \n\n\n\n\n\n\n11:15 – 11:45\nTowards Science DMZ based on Accelerated ONION using DTN\nSusumu Date\, Osaka University\nThe speaker introduces what is happening in Osaka University towards the promotion and advancement of Data-driven Scientific Research. In this talk\, the experience of using DTN is first introduced and then a future direction of compute infrastructure composed of supercomputers( SQUID and OCTOPUS ) and data infrastructure is explained based on the experience.\n\n\n11:45 – 13:15\nLunch\n\n\n13:15 – 13:45\nSPH-EXA: A Framework for Scalable\, Flexible\, and Extensible Astrophysical and Cosmological Simulations – Slides\nFlorina M. Ciorba\, Department of Mathematics and Computer Science\, University of Basel \n\n\n\nSPH-EXA is a highly scalable and extensible simulation framework for astrophysical and cosmological simulations. It is codesigned by computational (astrophysicists and cosmologists) and computer scientists (high-performance computing) to achieve scientific advances in astrophysics\, cosmology\, and high-performance computing\, for highly scalable simulations using Smoothed Particle Hydrodynamics.  SPH-EXA includes highly optimized and parallelized hydrodynamics and gravity solvers. It supports new particle types and particle data fields that can be combined with custom-made\, simulation-derived observable properties and in-situ data analysis.  It requires minimal software dependencies\, provides scalable parallelization and communication support\, and portability with optimizations for recent CPUs and GPUs. This design relieves potential users from architectural specifics and performance concerns for performing and scaling up simulations with SPH-EXA. This talk will describe the SPH-EXA framework\, its approach to domain decomposition\, parallelization\, gravity and hydrodynamics solvers\, and their flexible use as pluggable components. A scalable initial conditions generator\, test cases and simulations will also be presented.  SPH-EXA is extensible with additional physical effects through the use of propagators\, whereby the pluggable framework components are easily customized or implemented from scratch using the provided building blocks in an abstract\, efficient\, and scalable way. \n\n\n\n\n\n\n13:45 – 14:15\nPrediction and Mitigation of Aeroacoustic Noise on HPC Systems\nMatthias Meinke\, Ansgar Niemöller\, Miro Gondrum\, Zhe Yang\, Wolfgang Schröder\, Institute of Aerodynamics\, RWTH Aachen University\, Germany\n\n\n14:15 – 14:45\nReal-time flood inudndation simulation on SX-Aurora TSUBASA\nHiroyuki Takizawa\, Yoichi Shimomura\, Akihiro Musa\, Yoshihiko Sato\, Atuhiko Konja\, Guoqing Cui\, Rei Aoyagi\, and Keichi Takahashi\, Tohoku University\nA real-time flood inundation simulation based on the Rainfall-Runoff Inundation (RRI) model is memory-intensive\, and SX-Aurora TSUBASA is hence promising to efficiently execute the simulation in time due to high sustained memory bandwidth provided by vector processors. This talk will report that the real-time simulation has successfully been migrated and optimized for SX-Aurora TSUBASA. Assuming a shared computing system such as Supercomputer AOBA installed at Tohoku University\, a resource demand estimation method is developed to minimize the amount of shared computing resources used for prediction in order to reduce the impact on other users sharing the system. The evaluation results show that SX-Aurora TSUBASA with only 32 cores can meet the real-time simulation requirement of simulating 7-hour flood inundation for the Tohoku region of Japan within 20 minutes\, and also the resource demand estimation method can adaptively adjust the computing resource amount used for the real-time simulation.\n\n\n14:45 – 15:15\nBreak\n\n\n15:15 – 15:45\nCompetence Centres and Centres of Excellence within the European Strategy – Slides\nBastian Koller\, HLRS\, University of Stuttgart\n\n\n15:45 – 16:15\nPower capping in high performance computing – experiences and prospects\nPawel Czarnul\, Adam Krzywaniak and Jerzy Proficz\, Gdańsk University of TechnologyIn this work we investigate usage\, limitations and prospects of power capping in high performance computing (HPC). Specifically\, we discuss APIs for modern CPUs\, GPUs and present\, as an illustration\, new unpublished data showing performance and energy characteristics for selected parallel OpenMP applications under power caps executed on a dual socket Intel Skylake-X system. These APIs can be used within algorithms: deriving configurations for which selected performance-energy goals (such as EDP\, EDS) are optimized for non-trivial i.e. non-default power caps; and also allowing minimization of execution times under power caps. We discuss various factors of interest in the future in power capping aware HPC such as other metrics not considered so far\, applicability and accuracy of measurement methods: using filters\, hardware vs software methods\, conditions and use cases for particular methods. We describe future works and areas including scenarios that can benefit from power capping in HPC.\n\n\n16:15 – 16:45\nContainerization for DLR HPC applications\nSabine Roller\, Deutsches Zentrum für Luft- und Raumfahrt\, Institut für Softwaremethoden zur Produkt-Virtualisierung\n\n\n18:30\nDinner\n\n\n\n\nFriday\, April 14\, 2023\n\n\n\n\n09:00 – 09:45\nKeynote – Sustaining Simulation Performance in the US Exascale Computing Project – Slides\nHartwig Anzt \, University of Tennessee\, Knoxville\nThe US Exascale Computing Project (ECP) has the goal to deliver a capable exascale computing ecosystem to provide breakthrough modeling and simulation solutions to address the most critical challenges in scientific discovery\, energy assurance\, economic  competitiveness\, and national security. This requires providing scientific computing applications with a software stack that allows them to perform on the leadership supercomputers. In this talk\, we discuss the impact of the ECP hardware landscape on software design and how the Ginkgo math library responds to the ECP application requirements and helps to achieve the simulation performance goals.\n\n\n09:45 – 10:15\nHPC and AI at HLRS – Slides\nMichael Resch\, HLRS\, University of Stuttgart \n\n\n\nHLRS was created to support users with High-Performance Computing. Over the last years\, however\, Artificial Intelligence has become an important topic. In this talk we look into the changes that AI drives in computer simulation. We will explore the needs of AI users. Furthermore\, we will have a first look at the needs of traditional computer simulation and t the potential of AI for such simulations. \n  \n\n\n\n\n\n\n10:15 – 11:00\nBreak\n\n\n11:00 – 11:30\nTowards Building a Digital Twin of Job Scheduling\nTatsuyoshi Ohmura\, NEC Corporation \n\n\n\nA job scheduler\, which is the core component of the HPC system\, has several parameters. Tuning these parameters can improve the efficiency of system operation\, but this requires knowledge and experience and places a burden on\nthe system operator. To reduce this burden\, we are studying the digital twin of the job scheduler. To realize the digital twin\, we have developed a job scheduler simulator. We demonstrate examples of the use of the simulator. \n\n\n\n\n\n\n11:30 – 12:00\nScalable Cluster Administration with LXC³ – Slides\nErich Focht\, NEC Corporation\nThe LXC^3 Cluster Command and Control tools have evolved at NEC Germany since two decades\, going through various changes to adapt to continuously changing requirements. The talk discusses the design choices of LXC^3-neo which runs as a pool of micro-services orchestrated by docker swarm\, its scalability and limitations seen at customer sites. The most recent developments move the cluster management stack even closer to methods used in cloud systems management\, simplifying node image handling and management network setup.\n\n\n12:00 – 13:30\nLunch\n\n\n13:30 – 14:00\nA feasibility study of quantum annealing for the next-generation computing infrastructure – Slides\nKazuhiko Komatsu\, Tohoku University\nThis presentation introduces a new project\, feasibility study of quantum computing for the next-generation computing infrastructure\, and shows an early evaluation of annealing machines.\n\n\n14:00 – 14:30\nEnergy Efficiency and Renewable Energy for Distributed High-Performance Computing\nChristoph Niethammer\, HLRS\, University of Stuttgart\n\n\n14:30 – 15:00\nA new framework for calibrating COVID-19 SEIR models with spatial-/time-varying coefficients using genetic and sliding window algorithms – Slides\nHuan Zhou\, HLRS\, University of Stuttgart \n\n\n\nA susceptible-exposed-infected-removed (SEIR) model assumes spatial-/time-varying coefficients to model the effect of non-pharmaceutical interventions (NPIs) on the regional and temporal distribution of COVID-19 disease epidemics. A significant challenge in using such models is their fast and accurate calibration to observed data from geo-referenced hospitalized data\, i.e.\, efficient estimation of the spatial-/time-varying parameters. In this work\, a new calibration framework is proposed towards optimizing the spatial-/time-varying parameters of SEIR model. We also devise a method for combing the overlapping sliding window technique (OSW) with a genetic algorithm (GA) calibration routine to automatically search the segmented parameter space. Parallelized GA is used to reduce the computational burden. Our framework abstracts the implementation complexity of the method away from the user. It provides high-level APIs for setting up a customized calibration system and consuming the optimized values of parameters. We evaluated the application of our method on the calibration of a spatial age-structured microsimulation model (CoSMic) using a single objective function that comprises observed COVID-19-related ICU demand. The results reflect the effectiveness of the proposed method towards estimating the parameters in a changing environment. \n\n\n\n\n\n\n15:00\nOpen End\n\n\n\n\n\n\n\n\n\n\n\n\n 
URL:https://www.wssp.hlrs.de/events/35th-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20240617
DTEND;VALUE=DATE:20240619
DTSTAMP:20260924T234708
CREATED:20240527T091834Z
LAST-MODIFIED:20240618T075544Z
UID:292-1718582400-1718755199@www.wssp.hlrs.de
SUMMARY:37th Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \n\n\n\n\n\nMonday\, June 17th \, 2024\n\n\n\n\n10:00 – 10:15\nWelcome & Introduction\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n10:15 – 10:45\nOperational experience of the latest-generation SX-Aurora TSUBASA system\, AOBA-S\nHiroyuki Takizawa\, Cyberscience Center\, Tohoku University \nTohoku University Cyberscience Center started operation of the world’s largest SX-Aurora TSUBASA system named AOBA-S in August 2023. This talk reports the experience of operating the AOBA-S system while showing some performance evaluation results and discussions. Several important applications have already been optimized for AOBA-S\, and the performance evaluation results clearly suggest the potential of the latest-generation vector engines adopted in AOBA-S. This talk also introduces our research projects that have recently started in collaboration with AOBA users.\n\n\n10:45 – 11:15\nPerformance of a direct numerical simulation code for isothermal compressible turbulence on the SX-Aurora TSUBASA\nMitsuo Yokokawa\, Kobe University \nA direct numerical simulation code for compressible turbulent flows under isothermal conditions in a box with periodic boundary conditions was developed. A finite difference method is used for a discretization of the governing equations. In paticular\, an eighth-order compact difference scheme was used for the covective terms and the Mattor’s method\, which is a parallel solver for a linear system with tridiagonal matrix\, was used to compute the first-order derivative of the covective terms. Performance of the DNS code was measured on the SX-Aurora TSUBASA.\n\n\n11:15 – 11:45\nAim and Strategy of mdx2\, IaaS-typed Computing Infrastructure\nSusumu Date\, Osaka University \nThe Cybermedia Center at Osaka University has installed an IaaS-typed Computing Infrastructure in March 2024 and will soon start the IaaS service. In this talk\, the speaker will start with the background and show the aim and strategy of this system installation from now on as well as overview the structure of the system.\n\n\n11:45 – 12:15\nImproving Efficiency of Monte Carlo Method via Code Intrinsic framework\nQifeng Pan\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nThe Monte Carlo (MC) method is widely used in many engineering fields\, especially in uncertainty quantification\, due to its robustness and simplicity. However\, the existing execution pattern of MC suffers from low efficiency and scaling problems in high-performance computing (HPC). In this talk\, the speaker will introduce the code intrinsic framework designed to tackle the efficiency problem of MC in HPC. The basic idea of the code intrinsic framework is to reduce the redundant calculations of MC and increase the code vectorization rate. Numerical results show that performance improvements can be achieved on various platforms\, including the Intel and SX-Aurora TSUBASA machines.\n\n\n\n12:15 – 13:15\nLunch Break\n\n\n\n13:15 – 13:45\nDEPO Meets Mechanics: A Case Study on Dynamic Power Capping for Energy Efficiency\nJohannes Gebert\, Jonathan Schäfer\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nMoore’s law is slowing down despite HPC centers’ increasing energy consumption. For cost and climate reasons\, developing techniques for reducing energy usage while at least maintaining performance is sensible. DEPO\, a software-agnostic\, node-level\, and dynamic power-capping approach by Krzywaniak\, Czarnul\, and Proficz (2022) promises to achieve these goals. We present their approach to real-world challenges based on an FEA. The Direct Tensor Computation (DTC) of Ralf Schneider and Johannes Gebert is run with DEPO to demonstrate the tools’ applicability. We explore different points in the input configuration space of the application and investigate the impact on energy consumption under power capping. Furthermore\, we present and discuss different ways to continue and expand this research.\n\n\n\n13:45 – 14:15\nConnecting Software Methods and Data Driven Methods\nSabine Roller\, German Aerospace Center\, TU Dresden\n\n\n14:15 – 14:45\nEvaluating a Real-Time Lossy Array Compression Algorithm for Computer Simulations\nDarjan Krijan\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nComputer Simulations that were previously regarded as CPU-bound become gradually memory-bound as the growth in memory bandwidth cannot keep up with the much higher advancements in raw computational power. This imbalance is quantified with a relative factor of approximately5.1 per decade since the 1990s\, where a rise in memory bandwidth is met with a 5.1-times increase in relative computing power. In practical terms\, comparing a NEC SX-4 from 1994 that operated at a balanced computational intensity of 0.125 FLOP/Byte with an Intel Ponte Vecchio accelerator from 2023 that operates at 15.9 FLOP/Byte shows a factor of 127 in the described imbalance. Mixed precision approaches that were traditionally used to speed up the throughput of calculations inside the CPU core on a cache/register level now provide speedup due to less demanded memory bandwidth. Approaches to compress arrays in a lossless or lossy manner to reduce memory bandwidth were implemented in LLNL’s zfp library\, although it is not able to process the data in real-time. A similar approach targeting a real-time lossy array compression (RTLAC) algorithm is currently being evaluated for use in highly memory-bound computer simulations at HLRS.\n\n\n14:45 – 15:15\nCoffee Break\n\n\n15:15 – 15:45\nDirect Numerical Simulation of Turbulent Boundary Layers – On the Road to High Reynolds Numbers\nChristoph Wenzel\, Institute of Aerodynamics and Gas Dynamics\, University of Stuttgart\n\n\n\n15:45 – 16:15\nDatenmanagement for HPC Workloads at Scale: Case Studies on Data Structure and Compression.\nGregror Weiß\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n16:15 – 16:45\nCloud resolving global weather simulation with the Model for Prediction Across Scales (MPAS) – A case study\nThomas Schwitalla\, University of Hohenheim\n\n\n\n16:45 – 17:15\nThe Quest for Sustained Performance on Heterogeneous Exascale Architectures with the Climate Model ICON\nPanagiottis Adamidis\, Deutsches Klimarechenzentrum Hamburg\n\n\n\n18:30\nDinner\n\n\n\n\nTuesday\, June 18\, 2024\n\n\n\n\n09:00 – 09:30\nChallenges for HLRS ahead\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nThis talk summarizes the challenges that lay ahead for HLRS in the coming decade. It looks into the challenges that we face when changing technology from CPU to GPU. It will also address the issue of AI as a new user community in HPC. \n\n\n09:30 – 10:00\nAnalyses of Turbomachinery and Heat Transfer Cases using HPC Systems\nMatthias Meinke\, Institute of Aerodynamics\, RWTH Aachen University \nThis presentation will highlight computational methodologies and results from two industrial applications related to the field of turbomachinery and steel casting. Details of the computational methods featuring\, adaptive mesh refinement\, dynamic load balancing and multigrid methods will be presented together with the required HPC resources.\n\n\n10:00 – 10:30\nAlgorithmic Differentiation of Geometric Modelling Libraries Aimed at Gradient-Based Shape Optimization\nMladen Banovic\, German Aerospace Center\, TU Dresden\n\n\n\n10:30 – 11:00\nCoffee Break\n\n\n\n11:00 – 11:30\nML for Computational Science: Machine Learning Models to Accelerate Simulation Science on HPC\nMakoto Takamoto\, NEC Germany\n\n\n\n11:30 – 12:00\nA Constraint Partition Method for Combinatorial Optimization Problems\nKazuhiko Komatsu\, Cyberscience Center\, Tohoku University \nIn recent years\, Ising machines have attracted much attention due to their potential in solving combinatorial optimization problems that are challenging for conventional computers. For optimization problems formulated into quadratic unconstrained binary optimization (QUBO) problems with constraints\, known as constraint problems\, an objective function and constraint function\, along with penalty coefficients\, are combined into a single Hamiltonian of a QUBO problem. However\, when solving constraint problems\, solution accuracy typically degrades by excessively large penalty coefficients avoiding constraint violations. As a result\, minimizing the objective function becomes challenging. To solve this issue\, the presentation introduces a method that partitions constraint functions and reduces penalty coefficient values. Performance evaluation using the traveling salesperson problem (TSP) with one-hot constraints illustrates that the proposed method enhances solution accuracy compared to conventional approaches.\n\n\n12:00 – 12:30\nAn analyse of Kernels Performances when porting a CFD code to Nvidia GPUs using different Programming Models\nPaul Saumet\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n12:30 – 13:30\nLunch\n\n\n13:30 – 14:00\nAI for HPC: Optimising Operations\nRishabh Saxena\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nIn the last few years\, Artificial Intelligence and Machine Learning have been dominant topics in the general field of computer science\, and the scientific community as well. Since one of the core principles of machine learning-based algorithms requires a huge amount of data being processed on a large scale\, it is expectant that high performance hardware would be needed for such computations. In this context\, HPC can provide a suitable platform for various aspects of the ML pipeline\, from data pre-processing to the deployment of models in production environment. In this talk\, we will look at how HLRS is working towards optimizing its systems for AI and ML workloads\, what are the aspects of ML pipelines that are relevant for HPC\, how traditional HPC workloads\, like simulations\, can be integrated into ML pipelines\, and what is the outlook for AI on HPC in the near future at HLRS.\n\n\n14:00 – 14:30\nModelling of a Offshore Wind Park with OpenFOAM\nFlavio Galeazzo\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n14:30 – 15:00\n HANAMI: Advancing Supercomputing Collaboration Between Europe and Japan\nSophia Honisch\, High-Performance Computing Center Stuttgart\, University of Stuttgart \nThe HANAMI project\, a strategic alliance between Europe and Japan\, aims to innovate high-performance computing (HPC) applications for next-generation supercomputers across various scientific domains. This collaboration focuses on enhancing simulation capabilities in environmental sciences\, biomedicine\, and materials science. Aligned with the EuroHPC Joint Undertaking\, HANAMI will port existing code\, evaluate application performance on new architectures\, and facilitate access to advanced supercomputing resources like Fugaku and EuroHPC systems.\n\n\n15:00 – 15:45\n\nFarewell\n\n\n\n  \n  \n 
URL:https://www.wssp.hlrs.de/events/37th-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20250527
DTEND;VALUE=DATE:20250529
DTSTAMP:20260924T234708
CREATED:20250423T112900Z
LAST-MODIFIED:20250528T135251Z
UID:346-1748304000-1748476799@www.wssp.hlrs.de
SUMMARY:39th Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \n\n\n\n\n\nTuesday\, May 27th\, 2025\n\n\n\n\n09:15 – 09:30\nWelcome & Introduction\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n09:30 – 10:00\nResearch and user support activities at Tohoku University Cyberscience Center\nHiroyuki Takizawa\,Cyberscience Center\, Tohoku University\nTohoku University Cyberscience Center has been operating vector supercomputers and assisting users in fully utilizing their potential. This presentation will report on our recent efforts to help users optimize their code for our computing system\, AOBA\, which is powered by the latest generation of NEC SX-Aurora TSUBASA\, the most powerful vector supercomputer. At both the system operation and research levels\, we are continuously exploring effective methods to make the most of vector computing technologies. Performance evaluation results demonstrate that the SX-Aurora TSUBASA achieves high sustained performance for memory-intensive applications without requiring special programming models or languages.\n\n\n10:00 – 10:30\nThe Future of HLRS\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIn this talk we will look at the new developments in HPC and what impact they will have on HLRS. We will explore the role of AI for an HPC center and see how this will change services and operation of HLRS.\n\n\n10:30 – 11:00\nMusing on performance sustainability in the age of AI\, superchips\, APUs\, increasing TCOs and NetZero impact\nSadaf Alam\, Bristol Center of Supercomputing\, University of Bristol\nSupercomputing ecosystems have experienced considerable shifts since the early 2020 across applications and technology domains. HPC and supercomputing resources are increasingly being allocated to AI—a domain where the frequency of hardware changes and especially software stacks updates are considerably different compared to classic modelling and simulation HPC applications. Recently\, the fastest reported HPL system on November 2024 Top500 list is based on an APU\, or accelerated processor unit. This talks overviews Isambard-AI\, a UK national AI RR or research resource comprising Nvidia Arm-GPU superchip called GH200\, its software stack for AI and HPC\, and its sustainability credentials using the Modular Data Centre (MDC) solution to manage Total Cost of Ownership (TCO) and NetZero impact.\n\n\n11:00 – 11:30\nCoffee Break\n\n\n11:30 – 12:00\nFuture Computing at HLRS\nJohannes Gebert\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIncreasing the computational capabilities enables larger and more detailed simulations\, even after decades of development. While the high-performance computer’s (HPC) performance increases\, the physical\, hence technological limits are on the horizon. Specialized processors may focus on certain mathematical operations in widespread software stacks\, but they also promise to accelerate the performance of potentially heterogeneous systems. Hardware improvements need software environments to allow for reasonably quick ports of existing software stacks to these new\, accelerating devices.\nAt HLRS\, we contribute to the HPC community by investigating computing devices’ capabilities\, testing the software stack\, and sharing our experiences and improvements on hardware architectures.\nIn this talk\, we will give an overview of the future computing program at HLRS\, goals\, and strategy. We will outline current projects and present initial results.\n\n\n12:00 – 12:30\n Introduction and challenge of new supercomputing system towards the Open Science era\nSusumu Date\, University of Osaka\nOsaka University has been working on the procurement of a supercomputing system after OCTOPUS. We will plan to complete the installation of it and start the operation in September 2025. In this talk the speaker introduce and explain the specification of the new supercomputing system as well as a challenge which we have faced for realizing supercomputing system in the Open Science era.\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nFirst Results of GPU Porting Activities for Aeroacoustic Prediction\nMethods\nMatthias Meinke\, Institute of Aerodynamics\, RWTH Aachen University\ntbd\n\n\n14:00 – 14:30\nFrom PDE to x in NeoFOAM: An overview of the NeoFOAM project\nGregor Olenik\, TU München\nNeoFOAM is a platform portable implementation of OpenFOAMs core algorithms and data structures. It aims to bring modern software development methods to existing simulation workflows\, leveraging C++20 compliant code\, being extensively unit-tested\, hardware vendor agnostic\, and extensible via plugins. This talk discusses the architecture and design choices behind NeoFOAM using neoIcoFoam as an illustrative example.\n\n\n14:30 – 15:00\nAccelerating the FlowSimulator: Speeding up the HPC codes used at DLR \nImmo Huismann\, DLR Dresden\nThis contribution summarizes activities performed at German Aerospace Center (DLR) to speed up the used high-performance computing (HPC) codes.\nIt starts out from the overarching goals that are to be solved via HPC\, showcases the current distribution of HPC usage at DLR and\, thereafter\, demonstrates multiple case studies for analysing and accelerating the specific codes.\nThe list of case studies for performance analysis includes\, but is not limited to\, one of the CFD software “CFD by ONERA\, DLR and Airbus” (CODA) and one of an industrial-grade aeroelastic toolchain.\nRuntimes\, profiles\, and traces are shown\, and current actions to address the underlying bottlenecks discussed.\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nSustainable research software for accessible high-performance computing\nMichael Schlottke-Lakemper\, High-Performance Scientific Computing\, University of Augsburg\nModern supercomputers are becoming increasingly heterogeneous\, incorporating hardware components from multiple vendors. At the same time\, high-performance computing software development has grown more collaborative\, often uniting research groups and institutions across different regions. Coupled with the constant addition of new features and performance optimizations\, this raises critical questions about sustainability: How can we handle hardware complexity\, coordinate diverse development teams\, and maintain evolving research software within an academic environment while still writing energy-efficient\, high-performance code?In this talk\, we present some of our strategies for tackling these challenges through the Trixi Framework. We introduce Trixi.jl\, a high-order numerical simulation environment for conservation laws built in Julia\, along with its spin-off packages (TrixiShallowWater.jl\, TrixiAtmo.jl) and its sister project\, TrixiParticles.jl. We then discuss how we handle software architectures\, automation\, code reuse\, and organizational practices to balance extensibility with accessible high performance on heterogeneous systems. Finally\, we point out remaining open questions and outline plans for future development.\n\n\n16:00 – 16:30\nInvestigating the Cerebras CS2 Chip: Mathematical and Software Engineering Goals\, Methods and Initial Results\nJonathan Schäfer\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIn a world of ever-increasing computational demand and Moore’s law slowing down for conventional chips\, interest in alternative hardware for specialized mathematical operations rises. An instance of an interesting new hardware concept is the Cerebras CS2 chip with around 850000 cores\, which is therefore called a “supercomputer on a chip”. As part of the new Future Computing Group led by Johannes Gebert and in collaboration with the Computational Mathematics Group led by Prof. Hartwig Anzt at the Technical University of Munich\, HLRS investigates the new chip from various angles\, namely from a mathematical\, hardware-oriented\, software engineering\, and user experience point of view. We present a roadmap for research as well as initial findings.\n\n\n16:30 – 17:00\nReproducible and Performance-Optimized Environments for Large-Scale Machine Learning Applications\nFelix Ruhnke\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nWith the increasing size of machine learning models for solving complex problems\, the demand for computational resources is rising significantly. This leads to a stronger convergence between the disciplines of machine learning and High-Performance Computing. The growing performance capabilities of modern machine learning models open up new application areas\, including safety-critical domains such as medical technology. At the same time\, the scientific community is increasingly drawing attention to a reproducibility crisis in the field of machine learning. The intersection of these three developments underscores the necessity of thoroughly investigating reproducible and performance-optimized environments for large-scale machine learning applications. In this work\, various implementations of the Message Passing Interface in containerized environments were developed to examine the impact of different communication modes on the numerical reproducibility of results. For this purpose\, benchmark tests were conducted to optimize polynomials of varying degrees on a single node with a varying number of Graphics Processing Units. The gradient aggregation algorithms Average and Adaptive Summation were employed. The results showed that variations in communication models using the Average algorithm had no impact on numerical reproducibility. In contrast\, tests using the Adaptive Summation algorithm with varying communication modes resulted in non-reproducible outcomes.Furthermore\, it was observed that\, due to the non-associative nature of floating-point arithmetic operations and the varying execution order of computations during parallel training of machine learning models\, deviations of over 7% in the number of iterations until convergence can occur. Additionally\, the investigations revealed a highly sensitive convergence behavior of the models during training concerning configuration changes\, emphasizing the need for careful selection of the computing environment and precise hyperparameter adjustment.\n\n\n17:00 – 17:30\nDevelopment of dynamic resource assignment for effective system usage\nMasatoshi Kawai \, Tohoku University\nRecently\, energy efficiency as well as improving parallel performance of applications has become important for the operation and use of supercomputers. In this presentation\, we will introduce a developing platform that provides dynamic resource assignments for improving parallel performance and energy consumption.\n\n\n\n\nWednesday\, May 28\, 2025\n\n\n\n\n09:30 – 10:00\nHammerHAI – The German AI Factory for Engineering\, Global Challenges and Industry\nBastian Koller\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThis talk will provide insight into the German AI Factory HammerHAI\, which resulted from a European Key Initiative and which started service operations in Q1/25\n\n\n10:00 – 10:30\nLeveraging Cloud-Native Supercomputing for AI Workflows\nDennis Hoppe\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThis talk explores the transformative potential of cloud-native supercomputing concepts and their impact on AI workflows. While powerful and GPU-rich\, traditional High-Performance Computing systems present unique challenges when executing AI tasks. A clear paradigm shift is underway\, moving from a system-centric to a user-centric approach\, where usability\, accessibility\, and flexibility become key design criteria. Using the AI Factory HammerHAI as a concrete example\, the presentation will demonstrate how HammerHAI aims to lower these barriers by integrating cloud-native concepts like containerisation and orchestration\, thereby optimising supercomputing resources specifically for the practical needs of the AI community.\n\n\n10:30 – 11:00\nCoffee Break\n\n\n11:00 – 11:30\nNew Brand “NEC BluStellar” Use Case – Research Information Infrastructure(RII)\nFutoshi Tabata\, NEC Japan\ntbd\n\n\n11:30 – 12:00\nIntroduction of Confidential Computing in HPC: Usage Model\, Use Cases and Challenges\nKamil Tokmakov\, NEC\nHPC data centres accommodate users from various domains\, each with varying security requirements. Sensitive data processing\, such as for medical use cases or intellectual property protection\, requires stricter security measures\, including encryption across all data states: at-rest\, in-motion and in-use. While parallel file systems found in HPC\, such as GPFS and Lustre\, already offer encryption of data at-rest and in-motion\, encryption keys and sensitive data are not yet fully protected in memory. Confidential computing addresses such protection of data in-use by performing computations in the trusted execution environments\, where data and code are secured at the hardware level. This talk introduces confidential computing in the context of HPC\, covering its use cases and challenges.\n\n\n12:00 – 12:30\nNEC SX-Aurora TSUBASA – Our best friend for a long time\nChristoph Wenzel\, Institut für Aero- und Gasdynamik (IAG)\, University of Stuttgart\nFor a long time\, the NEC SX-Aurora TSUBASA has been a valuable addition to the HPE Apollo (Hawk) system at HLRS. Even though its size was not sufficient for large-scale production runs for the fundamental research on turbulent boundary layers with direct numerical simulation (DNS)\, Aurora has still played a central role in our research pipeline. In this talk\, our group’s experience working with the SX-Aurora platform will be presented\, highlighting its integration into our workflows. Additionally\, a performance study of our in-house DNS code NS3D on Aurora will be presented\, providing insights into computational efficiency achieved on Aurora.\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nSuper-resolution Reconstruction of Three-dimensional Vorticity Fields by Latent Diffusion Models\nMitsuo Yokokawa\, Tohoku University\ntbd\n\n\n14:00 – 14:30\nEnabling AMD APU Support for CalculiX\nCrunchiX: A Port to the Instinct MI300A\nBenjamin Schnabel\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nCalculiX CrunchiX is an open-source finite-element analysis (FEA) application featuring both implicit and explicit solvers\, written in C\, C++ and Fortran 77 and maintained by Guido Dhondt since 1998. Its flexible architecture allows it to interface with a variety of sparse linear solvers-such as an iterative Cholesky solver\, SPOOLES\, or intel oneMKL PARDISO-to solve structural mechanics problems. Recently\, heterogeneous computing with GPUs has emerged as a key strategy to accelerate the solution of linear systems of equations in scientific applications. While CalculiX already supports NVIDIA GPUs through the PaStiX solver and the CUDA library\, there is currently no counterpart for AMD architectures. In this work\, we describe the design\, implementation\, and optimization of a new backend for CalculiX that targets the AMD Instinct MI300A APUs.\nOur implementation is deployed and benchmarked on HLRS’s newest flagship supercomputer HPE Cray EX4000 system (Hunter).\n\n\n14:30 – 15:00\nPower and Performance\nNico Formanek\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThe biggest driver of computing performance are hardware improvements. Even though the relative energy efficiency (e.g. FLOPS/watt) is improving at an almost exponential rate this does not translate in any reduction in absolute energy consumption. Rebound effects like this have been studied for centuries in economics starting from Jevons (1865) but there is still no consensus if economic growth can be decoupled from absolute energy consumption. Here I will argue that we face a similar problem in computing\, i.e. that performance cannot be decoupled from absolute energy input. This in turn casts doubt on the feasibility of sustainability efforts like the GREENER (2023) principles. I will close by evaluating several accounts of why we still could want to improve performance even in the light of such hard tradeoffs.\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nTowards scientific foundation models\nSteffen Staab\, University of Stuttgart\nFoundation models are machine-learned models that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. Foundation models have been used successfully for question answering and text generation (ChatGPT)\, image understanding (Clip\, VIT)\, or image generation. Recently\, the basic idea underlying foundation models been considered for learning scientific foundation models that capture expectations about partial differential equations. Existing scientific foundation models have still been very much limited wrt. the type of PDEs or differential operators . In this talk\, I present some of our recent work on paving the way towards scientific foundation models that aims at making them more robust and better generalisable.\n\n\n16:00 – 16:30\nQuantum Computing Status -Technologies\, Benchmarks and Use Cases in Academic and Industry Fields\nShintaro Momose\, NEC Japan\ntbd\n\n\n16:30 – 16:45\nFarewell\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n 
URL:https://www.wssp.hlrs.de/events/39th-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20260423
DTEND;VALUE=DATE:20260425
DTSTAMP:20260924T234708
CREATED:20260419T174509Z
LAST-MODIFIED:20260525T091232Z
UID:400-1776902400-1777075199@www.wssp.hlrs.de
SUMMARY:41st Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \n\n\n\n\n\nThursday\, April 23rd\, 2026\n\n\n\n\n09:00\nRegistration Desk Opens\n\n\n09:45 – 10:00\nWelcome & Introduction\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n10:00 – 10:30\nA brief case study of “AI for Science” at the Cyberscience Center\nHiroyuki Takizawa\,Cyberscience Center\, Tohoku University\n\n\n\n10:30 – 11:00\nHLRS – Status and Outlook\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n11:00 – 11:30\nCoffee Break\n\n\n11:30 – 12:00\nOne Year of HUNTER from a User Perspective: Progress\, Performance\, and Patience with NS3D\nChristoph Wenzel\, Institute of Aerodynamics and Gas Dynamics\, University of Stuttgart\n\n\n\n12:00 – 12:30\nD3 Center’s strategy and project for supporting academic research in the “AI for Science” era\nSusumu Date\, The University of Osaka\n\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nImpact of GPU Virtualization on LLM Inference Performance: A Comparative Study of Bare Metal\, vGPU\nQifeng Pan\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n14:00 – 14:30\nThoughts on current and future HPC&AI installations\nSabine Roller\, Institute of Software Methods for Product Virtualization\, German Aerospace Center (DLR e.V.)\n\n\n\n14:30 – 15:00\nAeroacoustic Optimization of a Chevron Nozzle on the Hunter HPC System\nMatthias Meinke\, Institute of Aerodynamics\, RWTH Aachen University\n\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nEvaluating a Real-Time Lossy Array Compression Algorithm for a Lattice Boltzmann Solver\nDarjan Krijan\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nComputer simulations that were previously regarded as CPU-bound become gradually memory-bound as the growth in memory bandwidth cannot keep up with the much higher advancements in raw computing power. This imbalance is quantified with a relative factor of approximately 5.1 per decade since the 1990s\, where a rise in memory bandwidth is met with a 5.1-times increase in relative computing power. In practical terms\, comparing a NEC SX-4 from 1994 that operated at a balanced arithmetic intensity of 0.125 FLOP/Byte with an Intel Ponte Vecchio accelerator from 2023 that operates at 15.9 FLOP/Byte shows a factor of 127 in the described imbalance. Mixed-precision approaches that were traditionally used to speed up the throughput of calculations on a CPU core level now provide speedup due to less demanded memory bandwidth. Approaches to compress arrays in a lossless or lossy manner to reduce memory bandwidth were implemented in LLNL’s zfp library\, although it is not able to process the data in real-time. In this work\, a similar approach targeting a real-time lossy array compression (RTLAC) algorithm utilizing known value ranges of variables was developed and applied to a Lattice Boltzmann method for CFD simulations. Here\, the main simulation variables are within certain bounds\, making them ideal candidates for the RTLAC approach. Accuracy and performance results of the algorithm implemented in the m-AIA solver framework with multiple compression sizes will be presented.\n\n\n\n16:00 – 16:30\nFlowSimulator: A framework for multidisciplinary simulations for virtual aircraft\nJulian Braun\, Institute of Software Methods for Product Virtualization\, German Aerospace Center (DLR e.V.)\nMultidisciplinary simulations are essential for modern aerospace engineering. Applications include static aeroelastic coupling\, flutter analysis\, time-resolved aeroelastic simulations\, and gradient-based shape optimization. Highly specialized codes are able to compute high-fidelity solutions in their respective domain\, but external frameworks are needed to establish the coupling between them. This talk presents FlowSimulator and its approach for flexible and performant in-memory coupling. FlowSimulator is an environment which contains a variety of independent codes such as CFD for ONERA\, DLR and Airbus (CODA) and the structure solver b2000++pro. It follows a layered approach: users orchestrate their own workflows in Python\, while data exchange between simulation codes is realized on a C++ layer through MPI-parallel data structures for grids and simulation data. This allows for rapid prototyping as well as for performant\, parallel data exchange. Following a discussion of the design\, an outlook on fluid–structure interaction between high-order codes will be provided.\n\n\n\n16:30 – 17:00\nMPPI – Type safe C++ Datatypes for MPI\nMike Söhner\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nMPI provides a flexible C-API to communicate data of various types between a set of distributed processes over high-speed interconnects in HPC systems. Data buffers are described using MPI-Datatypes\, which specify the type and layout of the data to be transmitted. To construct these datatypes\, users must manually describe the memory layout of buffer elements via the MPI-API. However\, modern applications are typically written in object-oriented C++\, which offers significant advantages over C\, including type safety and metaprogramming capabilities.\nIn this work\, we introduce a new C++-API and datatype engine that leverage C++ language features such as concepts\, ranges\, and the upcoming reflection to extract the necessary datatype information for the user at compile-time. This approach simplifies the user’s work\, enhances code safety by eliminating manual datatype construction and offers previously unavailable possibilities. Our measurements demonstrate that this interface introduces no performance overhead and\, in some cases\, even improves performance.\n\n\n\n\n\nFriday\, April 24th\, 2026\n\n\n\n\n09:00\nRegistration Desk Opens\n\n\n09:30 – 10:00\nToward Anomaly Prediction in HPC Systems\nRyusuke Egawa\, School of Engineering\, Tokyo Denki University\n\n\n\n10:00 – 10:30\nRISC-V Vector Architecture & Programming Model\nFredrik Unger\, Openchip & Software Technologies\nIntroduction to RISC-V Vector architecture and programming model. A view on how the vector architecture differentiates from traditional scalar processors and GPU-based accelerators\, what the RISC-V open ISA with vector extensions brings to the HPC ecosystem\, and basic differences with respect to the programming model.\n\n\n\n10:30 – 11:00\nCoffee Break\n\n\n11:00 – 11:30\nFuture Computing: Researching and Working with the Cerebras Wafer Scale Engine\nJonathan Schäfer\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nWe present general remarks on researching and working with the Cerebras Wafer-Scale Engine (WSE). We discuss the programming model\, current efforts worldwide and in our group to implement non-AI workloads on the WSE\, and some preliminary and submitted results.\n\n\n\n11:30 – 12:00\nPatient-Specific Hemodynamic Simulations with SPH: Challenges and Implementation Pipeline\nNiklas Neher\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n12:00 – 12:30\nTowards Energy‑Aware HPC: Measuring Efficiency Across Heterogeneous Hardware with HWS\nDirk Pflüger\, Scientific Computing – Institute of Parallel and Distributed Systems\, University of Stuttgart\nEnergy efficiency has become a critical aspect of sustained high‑performance computing. Yet for developers\, obtaining reliable energy metrics remains challenging due to mixed hardware environments\, varying vendor interfaces\, and limited portability of measurement tools. In this talk\, we present HWS\, our HardWare Sampling library that offers uniform\, low‑overhead access to performance and power data acrossCPUs and GPUs. HWS enables detailed energy efficiency analysis in heterogeneous HPC systems. We further show comparative results from benchmark applications implemented in SYCL\, illustrating how the combination of HWS and SYCL supports energy‑aware optimizations and fair cross‑architecture performance evaluations.\n\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nEnhancing Research with Provenance Management for HPC and AI Systems\nYosuke Taira\, NEC Corporation\n\n\n\n14:00 – 14:30\nTools for Rapid and Efficient I/O for Exascale\nPatrick Vogler\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n14:30\nFarewell\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n 
URL:https://www.wssp.hlrs.de/events/41st-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Benjamin Schnabel":MAILTO:schnabel@hlrs.de
END:VEVENT
END:VCALENDAR