BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//WSSP - ECPv6.15.20//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://www.wssp.hlrs.de
X-WR-CALDESC:Events for WSSP
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Europe/Berlin
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20240331T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20241027T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20250330T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20251026T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20260329T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20261025T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20270328T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20271031T010000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;VALUE=DATE:20250527
DTEND;VALUE=DATE:20250529
DTSTAMP:20260923T202256
CREATED:20250423T112900Z
LAST-MODIFIED:20250528T135251Z
UID:346-1748304000-1748476799@www.wssp.hlrs.de
SUMMARY:39th Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \n\n\n\n\n\nTuesday\, May 27th\, 2025\n\n\n\n\n09:15 – 09:30\nWelcome & Introduction\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n09:30 – 10:00\nResearch and user support activities at Tohoku University Cyberscience Center\nHiroyuki Takizawa\,Cyberscience Center\, Tohoku University\nTohoku University Cyberscience Center has been operating vector supercomputers and assisting users in fully utilizing their potential. This presentation will report on our recent efforts to help users optimize their code for our computing system\, AOBA\, which is powered by the latest generation of NEC SX-Aurora TSUBASA\, the most powerful vector supercomputer. At both the system operation and research levels\, we are continuously exploring effective methods to make the most of vector computing technologies. Performance evaluation results demonstrate that the SX-Aurora TSUBASA achieves high sustained performance for memory-intensive applications without requiring special programming models or languages.\n\n\n10:00 – 10:30\nThe Future of HLRS\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIn this talk we will look at the new developments in HPC and what impact they will have on HLRS. We will explore the role of AI for an HPC center and see how this will change services and operation of HLRS.\n\n\n10:30 – 11:00\nMusing on performance sustainability in the age of AI\, superchips\, APUs\, increasing TCOs and NetZero impact\nSadaf Alam\, Bristol Center of Supercomputing\, University of Bristol\nSupercomputing ecosystems have experienced considerable shifts since the early 2020 across applications and technology domains. HPC and supercomputing resources are increasingly being allocated to AI—a domain where the frequency of hardware changes and especially software stacks updates are considerably different compared to classic modelling and simulation HPC applications. Recently\, the fastest reported HPL system on November 2024 Top500 list is based on an APU\, or accelerated processor unit. This talks overviews Isambard-AI\, a UK national AI RR or research resource comprising Nvidia Arm-GPU superchip called GH200\, its software stack for AI and HPC\, and its sustainability credentials using the Modular Data Centre (MDC) solution to manage Total Cost of Ownership (TCO) and NetZero impact.\n\n\n11:00 – 11:30\nCoffee Break\n\n\n11:30 – 12:00\nFuture Computing at HLRS\nJohannes Gebert\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIncreasing the computational capabilities enables larger and more detailed simulations\, even after decades of development. While the high-performance computer’s (HPC) performance increases\, the physical\, hence technological limits are on the horizon. Specialized processors may focus on certain mathematical operations in widespread software stacks\, but they also promise to accelerate the performance of potentially heterogeneous systems. Hardware improvements need software environments to allow for reasonably quick ports of existing software stacks to these new\, accelerating devices.\nAt HLRS\, we contribute to the HPC community by investigating computing devices’ capabilities\, testing the software stack\, and sharing our experiences and improvements on hardware architectures.\nIn this talk\, we will give an overview of the future computing program at HLRS\, goals\, and strategy. We will outline current projects and present initial results.\n\n\n12:00 – 12:30\n Introduction and challenge of new supercomputing system towards the Open Science era\nSusumu Date\, University of Osaka\nOsaka University has been working on the procurement of a supercomputing system after OCTOPUS. We will plan to complete the installation of it and start the operation in September 2025. In this talk the speaker introduce and explain the specification of the new supercomputing system as well as a challenge which we have faced for realizing supercomputing system in the Open Science era.\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nFirst Results of GPU Porting Activities for Aeroacoustic Prediction\nMethods\nMatthias Meinke\, Institute of Aerodynamics\, RWTH Aachen University\ntbd\n\n\n14:00 – 14:30\nFrom PDE to x in NeoFOAM: An overview of the NeoFOAM project\nGregor Olenik\, TU München\nNeoFOAM is a platform portable implementation of OpenFOAMs core algorithms and data structures. It aims to bring modern software development methods to existing simulation workflows\, leveraging C++20 compliant code\, being extensively unit-tested\, hardware vendor agnostic\, and extensible via plugins. This talk discusses the architecture and design choices behind NeoFOAM using neoIcoFoam as an illustrative example.\n\n\n14:30 – 15:00\nAccelerating the FlowSimulator: Speeding up the HPC codes used at DLR \nImmo Huismann\, DLR Dresden\nThis contribution summarizes activities performed at German Aerospace Center (DLR) to speed up the used high-performance computing (HPC) codes.\nIt starts out from the overarching goals that are to be solved via HPC\, showcases the current distribution of HPC usage at DLR and\, thereafter\, demonstrates multiple case studies for analysing and accelerating the specific codes.\nThe list of case studies for performance analysis includes\, but is not limited to\, one of the CFD software “CFD by ONERA\, DLR and Airbus” (CODA) and one of an industrial-grade aeroelastic toolchain.\nRuntimes\, profiles\, and traces are shown\, and current actions to address the underlying bottlenecks discussed.\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nSustainable research software for accessible high-performance computing\nMichael Schlottke-Lakemper\, High-Performance Scientific Computing\, University of Augsburg\nModern supercomputers are becoming increasingly heterogeneous\, incorporating hardware components from multiple vendors. At the same time\, high-performance computing software development has grown more collaborative\, often uniting research groups and institutions across different regions. Coupled with the constant addition of new features and performance optimizations\, this raises critical questions about sustainability: How can we handle hardware complexity\, coordinate diverse development teams\, and maintain evolving research software within an academic environment while still writing energy-efficient\, high-performance code?In this talk\, we present some of our strategies for tackling these challenges through the Trixi Framework. We introduce Trixi.jl\, a high-order numerical simulation environment for conservation laws built in Julia\, along with its spin-off packages (TrixiShallowWater.jl\, TrixiAtmo.jl) and its sister project\, TrixiParticles.jl. We then discuss how we handle software architectures\, automation\, code reuse\, and organizational practices to balance extensibility with accessible high performance on heterogeneous systems. Finally\, we point out remaining open questions and outline plans for future development.\n\n\n16:00 – 16:30\nInvestigating the Cerebras CS2 Chip: Mathematical and Software Engineering Goals\, Methods and Initial Results\nJonathan Schäfer\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nIn a world of ever-increasing computational demand and Moore’s law slowing down for conventional chips\, interest in alternative hardware for specialized mathematical operations rises. An instance of an interesting new hardware concept is the Cerebras CS2 chip with around 850000 cores\, which is therefore called a “supercomputer on a chip”. As part of the new Future Computing Group led by Johannes Gebert and in collaboration with the Computational Mathematics Group led by Prof. Hartwig Anzt at the Technical University of Munich\, HLRS investigates the new chip from various angles\, namely from a mathematical\, hardware-oriented\, software engineering\, and user experience point of view. We present a roadmap for research as well as initial findings.\n\n\n16:30 – 17:00\nReproducible and Performance-Optimized Environments for Large-Scale Machine Learning Applications\nFelix Ruhnke\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nWith the increasing size of machine learning models for solving complex problems\, the demand for computational resources is rising significantly. This leads to a stronger convergence between the disciplines of machine learning and High-Performance Computing. The growing performance capabilities of modern machine learning models open up new application areas\, including safety-critical domains such as medical technology. At the same time\, the scientific community is increasingly drawing attention to a reproducibility crisis in the field of machine learning. The intersection of these three developments underscores the necessity of thoroughly investigating reproducible and performance-optimized environments for large-scale machine learning applications. In this work\, various implementations of the Message Passing Interface in containerized environments were developed to examine the impact of different communication modes on the numerical reproducibility of results. For this purpose\, benchmark tests were conducted to optimize polynomials of varying degrees on a single node with a varying number of Graphics Processing Units. The gradient aggregation algorithms Average and Adaptive Summation were employed. The results showed that variations in communication models using the Average algorithm had no impact on numerical reproducibility. In contrast\, tests using the Adaptive Summation algorithm with varying communication modes resulted in non-reproducible outcomes.Furthermore\, it was observed that\, due to the non-associative nature of floating-point arithmetic operations and the varying execution order of computations during parallel training of machine learning models\, deviations of over 7% in the number of iterations until convergence can occur. Additionally\, the investigations revealed a highly sensitive convergence behavior of the models during training concerning configuration changes\, emphasizing the need for careful selection of the computing environment and precise hyperparameter adjustment.\n\n\n17:00 – 17:30\nDevelopment of dynamic resource assignment for effective system usage\nMasatoshi Kawai \, Tohoku University\nRecently\, energy efficiency as well as improving parallel performance of applications has become important for the operation and use of supercomputers. In this presentation\, we will introduce a developing platform that provides dynamic resource assignments for improving parallel performance and energy consumption.\n\n\n\n\nWednesday\, May 28\, 2025\n\n\n\n\n09:30 – 10:00\nHammerHAI – The German AI Factory for Engineering\, Global Challenges and Industry\nBastian Koller\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThis talk will provide insight into the German AI Factory HammerHAI\, which resulted from a European Key Initiative and which started service operations in Q1/25\n\n\n10:00 – 10:30\nLeveraging Cloud-Native Supercomputing for AI Workflows\nDennis Hoppe\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThis talk explores the transformative potential of cloud-native supercomputing concepts and their impact on AI workflows. While powerful and GPU-rich\, traditional High-Performance Computing systems present unique challenges when executing AI tasks. A clear paradigm shift is underway\, moving from a system-centric to a user-centric approach\, where usability\, accessibility\, and flexibility become key design criteria. Using the AI Factory HammerHAI as a concrete example\, the presentation will demonstrate how HammerHAI aims to lower these barriers by integrating cloud-native concepts like containerisation and orchestration\, thereby optimising supercomputing resources specifically for the practical needs of the AI community.\n\n\n10:30 – 11:00\nCoffee Break\n\n\n11:00 – 11:30\nNew Brand “NEC BluStellar” Use Case – Research Information Infrastructure(RII)\nFutoshi Tabata\, NEC Japan\ntbd\n\n\n11:30 – 12:00\nIntroduction of Confidential Computing in HPC: Usage Model\, Use Cases and Challenges\nKamil Tokmakov\, NEC\nHPC data centres accommodate users from various domains\, each with varying security requirements. Sensitive data processing\, such as for medical use cases or intellectual property protection\, requires stricter security measures\, including encryption across all data states: at-rest\, in-motion and in-use. While parallel file systems found in HPC\, such as GPFS and Lustre\, already offer encryption of data at-rest and in-motion\, encryption keys and sensitive data are not yet fully protected in memory. Confidential computing addresses such protection of data in-use by performing computations in the trusted execution environments\, where data and code are secured at the hardware level. This talk introduces confidential computing in the context of HPC\, covering its use cases and challenges.\n\n\n12:00 – 12:30\nNEC SX-Aurora TSUBASA – Our best friend for a long time\nChristoph Wenzel\, Institut für Aero- und Gasdynamik (IAG)\, University of Stuttgart\nFor a long time\, the NEC SX-Aurora TSUBASA has been a valuable addition to the HPE Apollo (Hawk) system at HLRS. Even though its size was not sufficient for large-scale production runs for the fundamental research on turbulent boundary layers with direct numerical simulation (DNS)\, Aurora has still played a central role in our research pipeline. In this talk\, our group’s experience working with the SX-Aurora platform will be presented\, highlighting its integration into our workflows. Additionally\, a performance study of our in-house DNS code NS3D on Aurora will be presented\, providing insights into computational efficiency achieved on Aurora.\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nSuper-resolution Reconstruction of Three-dimensional Vorticity Fields by Latent Diffusion Models\nMitsuo Yokokawa\, Tohoku University\ntbd\n\n\n14:00 – 14:30\nEnabling AMD APU Support for CalculiX\nCrunchiX: A Port to the Instinct MI300A\nBenjamin Schnabel\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nCalculiX CrunchiX is an open-source finite-element analysis (FEA) application featuring both implicit and explicit solvers\, written in C\, C++ and Fortran 77 and maintained by Guido Dhondt since 1998. Its flexible architecture allows it to interface with a variety of sparse linear solvers-such as an iterative Cholesky solver\, SPOOLES\, or intel oneMKL PARDISO-to solve structural mechanics problems. Recently\, heterogeneous computing with GPUs has emerged as a key strategy to accelerate the solution of linear systems of equations in scientific applications. While CalculiX already supports NVIDIA GPUs through the PaStiX solver and the CUDA library\, there is currently no counterpart for AMD architectures. In this work\, we describe the design\, implementation\, and optimization of a new backend for CalculiX that targets the AMD Instinct MI300A APUs.\nOur implementation is deployed and benchmarked on HLRS’s newest flagship supercomputer HPE Cray EX4000 system (Hunter).\n\n\n14:30 – 15:00\nPower and Performance\nNico Formanek\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nThe biggest driver of computing performance are hardware improvements. Even though the relative energy efficiency (e.g. FLOPS/watt) is improving at an almost exponential rate this does not translate in any reduction in absolute energy consumption. Rebound effects like this have been studied for centuries in economics starting from Jevons (1865) but there is still no consensus if economic growth can be decoupled from absolute energy consumption. Here I will argue that we face a similar problem in computing\, i.e. that performance cannot be decoupled from absolute energy input. This in turn casts doubt on the feasibility of sustainability efforts like the GREENER (2023) principles. I will close by evaluating several accounts of why we still could want to improve performance even in the light of such hard tradeoffs.\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nTowards scientific foundation models\nSteffen Staab\, University of Stuttgart\nFoundation models are machine-learned models that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. Foundation models have been used successfully for question answering and text generation (ChatGPT)\, image understanding (Clip\, VIT)\, or image generation. Recently\, the basic idea underlying foundation models been considered for learning scientific foundation models that capture expectations about partial differential equations. Existing scientific foundation models have still been very much limited wrt. the type of PDEs or differential operators . In this talk\, I present some of our recent work on paving the way towards scientific foundation models that aims at making them more robust and better generalisable.\n\n\n16:00 – 16:30\nQuantum Computing Status -Technologies\, Benchmarks and Use Cases in Academic and Industry Fields\nShintaro Momose\, NEC Japan\ntbd\n\n\n16:30 – 16:45\nFarewell\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n 
URL:https://www.wssp.hlrs.de/events/39th-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Johannes Gebert":MAILTO:gebert@hlrs.de
END:VEVENT
BEGIN:VEVENT
DTSTART;VALUE=DATE:20260423
DTEND;VALUE=DATE:20260425
DTSTAMP:20260923T202256
CREATED:20260419T174509Z
LAST-MODIFIED:20260525T091232Z
UID:400-1776902400-1777075199@www.wssp.hlrs.de
SUMMARY:41st Workshop on Sustained Simulation Performance
DESCRIPTION:Agenda\nAll times are given in Central European Summer Time (CEST).  \n\n\n\n\n\nThursday\, April 23rd\, 2026\n\n\n\n\n09:00\nRegistration Desk Opens\n\n\n09:45 – 10:00\nWelcome & Introduction\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n10:00 – 10:30\nA brief case study of “AI for Science” at the Cyberscience Center\nHiroyuki Takizawa\,Cyberscience Center\, Tohoku University\n\n\n\n10:30 – 11:00\nHLRS – Status and Outlook\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n11:00 – 11:30\nCoffee Break\n\n\n11:30 – 12:00\nOne Year of HUNTER from a User Perspective: Progress\, Performance\, and Patience with NS3D\nChristoph Wenzel\, Institute of Aerodynamics and Gas Dynamics\, University of Stuttgart\n\n\n\n12:00 – 12:30\nD3 Center’s strategy and project for supporting academic research in the “AI for Science” era\nSusumu Date\, The University of Osaka\n\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nImpact of GPU Virtualization on LLM Inference Performance: A Comparative Study of Bare Metal\, vGPU\nQifeng Pan\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n14:00 – 14:30\nThoughts on current and future HPC&AI installations\nSabine Roller\, Institute of Software Methods for Product Virtualization\, German Aerospace Center (DLR e.V.)\n\n\n\n14:30 – 15:00\nAeroacoustic Optimization of a Chevron Nozzle on the Hunter HPC System\nMatthias Meinke\, Institute of Aerodynamics\, RWTH Aachen University\n\n\n\n15:00 – 15:30\nCoffee Break\n\n\n15:30 – 16:00\nEvaluating a Real-Time Lossy Array Compression Algorithm for a Lattice Boltzmann Solver\nDarjan Krijan\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nComputer simulations that were previously regarded as CPU-bound become gradually memory-bound as the growth in memory bandwidth cannot keep up with the much higher advancements in raw computing power. This imbalance is quantified with a relative factor of approximately 5.1 per decade since the 1990s\, where a rise in memory bandwidth is met with a 5.1-times increase in relative computing power. In practical terms\, comparing a NEC SX-4 from 1994 that operated at a balanced arithmetic intensity of 0.125 FLOP/Byte with an Intel Ponte Vecchio accelerator from 2023 that operates at 15.9 FLOP/Byte shows a factor of 127 in the described imbalance. Mixed-precision approaches that were traditionally used to speed up the throughput of calculations on a CPU core level now provide speedup due to less demanded memory bandwidth. Approaches to compress arrays in a lossless or lossy manner to reduce memory bandwidth were implemented in LLNL’s zfp library\, although it is not able to process the data in real-time. In this work\, a similar approach targeting a real-time lossy array compression (RTLAC) algorithm utilizing known value ranges of variables was developed and applied to a Lattice Boltzmann method for CFD simulations. Here\, the main simulation variables are within certain bounds\, making them ideal candidates for the RTLAC approach. Accuracy and performance results of the algorithm implemented in the m-AIA solver framework with multiple compression sizes will be presented.\n\n\n\n16:00 – 16:30\nFlowSimulator: A framework for multidisciplinary simulations for virtual aircraft\nJulian Braun\, Institute of Software Methods for Product Virtualization\, German Aerospace Center (DLR e.V.)\nMultidisciplinary simulations are essential for modern aerospace engineering. Applications include static aeroelastic coupling\, flutter analysis\, time-resolved aeroelastic simulations\, and gradient-based shape optimization. Highly specialized codes are able to compute high-fidelity solutions in their respective domain\, but external frameworks are needed to establish the coupling between them. This talk presents FlowSimulator and its approach for flexible and performant in-memory coupling. FlowSimulator is an environment which contains a variety of independent codes such as CFD for ONERA\, DLR and Airbus (CODA) and the structure solver b2000++pro. It follows a layered approach: users orchestrate their own workflows in Python\, while data exchange between simulation codes is realized on a C++ layer through MPI-parallel data structures for grids and simulation data. This allows for rapid prototyping as well as for performant\, parallel data exchange. Following a discussion of the design\, an outlook on fluid–structure interaction between high-order codes will be provided.\n\n\n\n16:30 – 17:00\nMPPI – Type safe C++ Datatypes for MPI\nMike Söhner\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nMPI provides a flexible C-API to communicate data of various types between a set of distributed processes over high-speed interconnects in HPC systems. Data buffers are described using MPI-Datatypes\, which specify the type and layout of the data to be transmitted. To construct these datatypes\, users must manually describe the memory layout of buffer elements via the MPI-API. However\, modern applications are typically written in object-oriented C++\, which offers significant advantages over C\, including type safety and metaprogramming capabilities.\nIn this work\, we introduce a new C++-API and datatype engine that leverage C++ language features such as concepts\, ranges\, and the upcoming reflection to extract the necessary datatype information for the user at compile-time. This approach simplifies the user’s work\, enhances code safety by eliminating manual datatype construction and offers previously unavailable possibilities. Our measurements demonstrate that this interface introduces no performance overhead and\, in some cases\, even improves performance.\n\n\n\n\n\nFriday\, April 24th\, 2026\n\n\n\n\n09:00\nRegistration Desk Opens\n\n\n09:30 – 10:00\nToward Anomaly Prediction in HPC Systems\nRyusuke Egawa\, School of Engineering\, Tokyo Denki University\n\n\n\n10:00 – 10:30\nRISC-V Vector Architecture & Programming Model\nFredrik Unger\, Openchip & Software Technologies\nIntroduction to RISC-V Vector architecture and programming model. A view on how the vector architecture differentiates from traditional scalar processors and GPU-based accelerators\, what the RISC-V open ISA with vector extensions brings to the HPC ecosystem\, and basic differences with respect to the programming model.\n\n\n\n10:30 – 11:00\nCoffee Break\n\n\n11:00 – 11:30\nFuture Computing: Researching and Working with the Cerebras Wafer Scale Engine\nJonathan Schäfer\, High-Performance Computing Center Stuttgart\, University of Stuttgart\nWe present general remarks on researching and working with the Cerebras Wafer-Scale Engine (WSE). We discuss the programming model\, current efforts worldwide and in our group to implement non-AI workloads on the WSE\, and some preliminary and submitted results.\n\n\n\n11:30 – 12:00\nPatient-Specific Hemodynamic Simulations with SPH: Challenges and Implementation Pipeline\nNiklas Neher\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n12:00 – 12:30\nTowards Energy‑Aware HPC: Measuring Efficiency Across Heterogeneous Hardware with HWS\nDirk Pflüger\, Scientific Computing – Institute of Parallel and Distributed Systems\, University of Stuttgart\nEnergy efficiency has become a critical aspect of sustained high‑performance computing. Yet for developers\, obtaining reliable energy metrics remains challenging due to mixed hardware environments\, varying vendor interfaces\, and limited portability of measurement tools. In this talk\, we present HWS\, our HardWare Sampling library that offers uniform\, low‑overhead access to performance and power data acrossCPUs and GPUs. HWS enables detailed energy efficiency analysis in heterogeneous HPC systems. We further show comparative results from benchmark applications implemented in SYCL\, illustrating how the combination of HWS and SYCL supports energy‑aware optimizations and fair cross‑architecture performance evaluations.\n\n\n\n12:30 – 13:30\nLunch Break\n\n\n13:30 – 14:00\nEnhancing Research with Provenance Management for HPC and AI Systems\nYosuke Taira\, NEC Corporation\n\n\n\n14:00 – 14:30\nTools for Rapid and Efficient I/O for Exascale\nPatrick Vogler\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n14:30\nFarewell\nMichael Resch\, High-Performance Computing Center Stuttgart\, University of Stuttgart\n\n\n\n 
URL:https://www.wssp.hlrs.de/events/41st-workshop-on-sustained-simulation-performance/
LOCATION:HLRS\, Nobelstraße 19\, Stuttgart\, Baden-Württemberg\, 70569\, Germany
ORGANIZER;CN="Mr. Benjamin Schnabel":MAILTO:schnabel@hlrs.de
END:VEVENT
END:VCALENDAR