Hardware Modularity for Practical Heterogeneous HPC

ISC 2025 Modularity Workshop

ISC High PerformanceDate and location: Friday, June 13th.

Time: 9am to 1pm

Location: Hamburg, Germany, CCH, Hall X6 – 1st floor

Description

Hardware specialization is a widely accepted strategy to preserve performance scaling for HPC. However, designing and procuring specialized hardware or even entire systems requires substantial investments. This means that many HPC workloads that do not have large commercial financial backing risk falling behind because they can no longer piggyback on the advancement of general-purpose hardware. This realization motivates modular HPC systems where specialized hardware can be easily generated and integrated into production systems.

This workshop will engage and motivate the broad HPC community towards a modular heterogeneous future HPC system where custom hardware can be developed and integrated into future HPC systems through open standards and synergistic technologies. This is truly a diverse topic that offers multiple prime opportunities for co-design across different technologies, making community collaboration important; therefore, we invite everyone broadly interested in this topic to attend. This workshop will feature established community leaders talking about various important topics towards this vision, and ample interactive sessions for community engagement. Invited talks will provide an overview of why this grand goal is important, what challenges remain and what technologies promise to provide solutions, how different technologies can inter-operate to maximum high-level impact, and the promise of open standards.

Organizers

  • George Michelogiannakis, LBNL
  • Patricia Gonzalez-Guerrero, LBNL
  • John Shalf, LBNL

Organizing committee

  • Kentaro Sano, RIKEN
  • Cliff Grossner, OCP
  • Kazumoto Yoshii, ANL
  • Estela Suarez, SiPEARL and Juelich Supercomputing Centre
  • Galen Shipman, LANL
  • Jens Domke, RIKEN
  • Nick Brown, University of Edinburgh

Register here

Schedule

TIME TOPIC SPEAKER

9:00am – 9:10am

Welcome

Organizers

9:10am – 9:25am

RISC-V CPU cores for accelerators, and how to drive efficiency across use cases Felix LeClair

9:25am – 10:10am

(Keynote) End-to-end Modularity, spanning from modular data centres to federated AI supercomputing services Sadaf Alam

10:10am – 10:35am

Open Chiplet Economies for AI and HPC Hardware Specialization Raul Alvarez

10:35am – 11:00am

How can we program modular heterogeneous future HPC systems? Nick Brown
11:00am – 11:30am Break Break

11:30am – 11:45am

Chiplets Modularity for HPC and AI

Patricia Gonzalez-Guerrero

11:45am – 12:00pm

RIKEN CGRA for HPC and AI

Kentaro Sano

12:00pm – 1:00pm

Open discussion  

Talk Abstracts

  1. End-to-end Modularity, spanning from modular data centres to federated AI supercomputing services

The landscape of hardware and software is evolving quickly, driven largely by the changing needs of diverse AI users and community-driven platforms. Isambard-AI is one of UK’s national AI research Resource (AIRR) providing large-scale compute capacity to AI researchers and innovators.  Its core architectural principles include sustainability, accessibility, and adaptability to enable the rapid deployment of AI services to a wide range of evolving communities.   I will discuss the Isambard Park Modular Data Centre (MDC) solution’s emphasis on modularity and interoperability, leveraging a sustainable, Lego-style architecture for scalable development.  Isambard-AI services promote interoperability and seamless identity federation across academic and research organisations worldwide, significantly lowering barriers to access by utilising open standards and protocols via a modular framework for identity and access management (IAM).  Furthermore, by combining cloud and HPC, the resulting Isambard-AI software stacks deliver a modular, adaptable, and scalable framework to address diverse and evolving demands, for workloads and workflows as well as emerging hardware technologies.  Finally, I will introduce OpenCHAMI, a community-driven Open Composable Heterogeneous Adaptable Management Infrastructure, that targets system management and provisioning designed to bring cloud-like, highly heterogeneous, scalable infrastructure management flexibility to supercomputing environments.

Author: Dr Sadaf Alam is chief technology officer (CTO) for Bristol Centre for Supercomputing (BriCS), home to Isambard 3 and Isambard-AI, part of the national AI Research Resource (AIRR). She is also director of strategy and academia in the Advanced Computing Research Centre. Across both roles, she is responsible for digital transformation of research computing and data services. Prior to joining Bristol, Alam was the CTO at CSCS, the Swiss National Supercomputing Centre. She was chief architect for two generations of the Piz Daint innovative flagship supercomputing facilities and the MeteoSwiss operational weather forecasting platforms. From 2004 to 2009, Dr Alam was a computer scientist at Oak Ridge National Laboratory (ORNL) and a staff scientist at the ORNL Leadership Computing Facility (OLCF). She studied computer science at the University of Edinburgh, UK, where she received her PhD. 

  1. RISC-V CPU cores for accelerators, and how to drive efficiency across use cases

A flash talk from Tenstorrents lead HPC engineer Felix LeClair exploring the design decisions and requirements that led to the inclusion of 5 in order RISC-V cores within each accelerator core.

Author: Lead HPC Engineer at Tenstorrent, Felix LeClair has ownership of the 2,3,5, and 10 year HPC Hardware Software Product Roadmap. Prior to Tenstorent, he worked on global scale meteorological forecasting, optimized mixed and reduced precision BLAS routines, and reduced precision computational fluid dynamics.

  1. Project of RIKEN CGRA for HPC and AI

At Processor research team in RIKEN Center for Computational Science (R-CCS), we have been researching reconfigurable HPC with FPGAs and reconfigurable accelerators such as a coarse-grained reconfigurable array (CGRA) for general-purpose and/or domain-specific computing. Although FPGAs allow us dataflow computing and specialization of circuits that are promising for scalable and power-efficient computing, they suffer from low productivity in designing hardware, an overhead of area and frequency, and long compilation time as well as insufficient performance compared with GPUs for general HPC applications. In this talk, I introduce our research project on RIKEN CGRA for HPC and AI where we design and implement a hawdware module for data-flow computing.

Author: Kentaro Sano is the leader of the processor research team and the advanced AI device development unit at RIKEN Center for Computational Science (R-CCS) since 2017, responsible for research and development of future processors and systems for HPC and AI. He is also a visiting professor with an advanced computing system laboratory at Tohoku University. He received his Ph.D. from the graduate school of information sciences, Tohoku University, in 2000. From 2000 until 2018, he was a Research Associate and an Associate Professor at Tohoku University. He was a visiting researcher at the Department of Computing, Imperial College, London, and Maxeler Technology corporation in 2006 and 2007. Nowadays he leads the architecture research group in the feasibility study project for the next-generation supercomputer development in Japan. His research interests include data-driven and spatial-parallel processor architectures such as a coarse-grain reconfigurable array (CGRA), FPGA-based high-performance reconfigurable computing, high-level synthesis compilers and tools for reconfigurable custom computing machines, and system architectures for next-generation supercomputing based on the data-flow computing model.

  1. The elephant in the room: How can we program modular heterogeneous future HPC systems?

Whilst specialised hardware has the potential to provide significant performance and energy efficiency benefits, the programming of such technologies is far from simple. Often involving new concepts and APIs, this requires at the very least expertise of the target hardware and significant changes to a code base – and not infrequently entire rewriting of kernels in new languages. These challenges become even more severe as we look to adopt modular, heterogeneous, future HPC systems which combine and integrate a range of custom hardware. In short, programmability challenges are a major barrier to the HPC community leveraging specialised hardware today, and if we don’t address this issue it will likely be a blocker for such as heterogeneous future.

In this talk I will describe work done in the compiler space, using MLIR, to address these programmability challenges. By leveraging a composable compiler ecosystem, it is possible to more easily drive optimisations for these future architectures from existing languages and programming models such as OpenMP. 

Author: Dr Nick Brown is a Senior Research Fellow at EPCC, the University of Edinburgh. His main interest is in the role that novel hardware can play in future supercomputers, and is specifically motivated by the grand-challenge of how we can ensure scientific programmers are able to effectively exploit such technologies without extensive hardware/architecture expertise. Combining novel algorithmic techniques for new hardware, programming language & library design, and compilers. He is chair of the RISC-V HPC SIG, leads EPCC’s RISC-V testbed, and has organised and led the RISC-V series of workshops at SC, ISC and HPC Asia conferences.

                 V.  Open Chiplet Economies for AI and HPC Hardware Specialization

The increase in demand for compute performance at scale has never been greater, and the current silicon supply chain delivering monolithic and generic SoCs can no longer deliver the generation over generation increase in performance needed to satiate this demand. Chiplets offer the promise of diversity to better match workloads to computational infrastructure, creating large-scale high-performance computers with pools of heterogenous processors that can be dynamically composed to provide virtual compute nodes specialized for particular workloads. Unfortunately, Chiplet technology today is mostly used in proprietary settings by a few large SoC suppliers, limiting the ability of the market to innovate. The Open Compute Project (OCP) Community believes standardized chiplet Interfaces and modular form factors will allow diverse chiplets to be interchanged and reused enabling innovation in specialized silicon processing elements by smaller companies, which is necessary to meet the future demands of high-performance computing. This talk will cover the ongoing work at the OCP Open Chiplet Economy Project focused on enabling a new silicon supply chain with an open chiplet marketplace intended to foster the innovation and emerging market for specialized chiplet-based System in Package (SiP) SoCs, making hardware diversification cost effective rather than cost prohibitive. 

Author: Raúl has a rich background in the IT and data center market, with extensive experience in Data center operations. Prior to joining OCP as a Staff Member, Raúl has been participating in the community, with contributions to immersion cooling technologies and Data Center Facilities. He has led various workstreams and co-developed essential guidelines, including Immersion Requirements and Liquid Cooling Facilities. Raúl is also co-leading the OCP European Regional Community project, focusing on outreach and fostering community growth. His expertise has made a meaningful impact in the field, as he actively engages with the data center community to drive innovation and collaboration.

  1. Chiplets Modularity for HPC and AI

Specialized computing and memory can drive a 1000x performance increase in High-Performance Computing (HPC). However, the cost of developing specialized hardware for HPC alone is prohibitive. As part of the Open Compute Project (OCP), we launched a new workstream, run by volunteers, to specify integration strategies to deliver specialization and heterogeneity for HPC and AI effectively. By streamlining chiplet modularity through open standards, we can significantly reduce the development costs of specialized hardware. These standards will facilitate the integration of chiplets from multiple vendors, promoting hardware IP reusability for both HPC and AI. We will define form factors and protocol standards and explore technologies that could accelerate scientific workloads. The ultimate goal is to produce a specification for standardized chiplet form factors and supporting technologies that enable seamless integration of chiplets and related hardware blocks into a System in Package (SiP), thereby unlocking the next generation of performance growth for scientific and engineering applications.

Author: Dr. Patricia Gonzalez-Guerrero is a Research Scientist at Lawrence Berkeley National Laboratory. Her work spans ultra-low-power digital and mixed-signal SoC/ASIC/VLSI design for conventional and non-conventional forms of processing; Superconducting computing; Quantum readout and control; FPGAs/RISCV for exploration and evaluation of high-performance computing architectures; and hardware specialization through chiplets. She has three best paper awards and two awarded patents and is currently co-leading the HPC&AI chiplets modularity workstream in the Open Compute Project.