
As a follow-up to our recent Open Systems for AI – IT Infrastructure Workshops, we now turn our focus to Systems Management. This workshop will bring together practitioners and experts to explore how open standardization and tools can enable reliable, scalable, and energy-efficient AI clusters. We’ll dive into the frameworks, orchestration, telemetry, and automation needed to manage complex AI/ML systems across diverse environments.
| Start Time PST |
End Time PST |
Title | Speakers |
|---|---|---|---|
| 8:00 | 8:10 | Welcome & Introduction | John Leung (Intel) Rob Coyle (OCP) |
| 8:10 | 8:30 | UALink Switch Management with the UALink-SAI | Justin King (AMD) Ayyappa Nuthalapati (Marvell) |
| 8:30 | 8:50 | AI Fleet learnings and standardization approach to minimize fleet downtime and Improve Margins | Rama Bhimanadhuni (Microsoft) |
| 8:50 | 9:10 | FPGA based liquid cooling leak detection | Munir Ahmad (Lattice Semiconductor) |
| 9:10 | 9:30 | Deterministic Chiplet Boot Flow and Debug Management for Scale-Independent AI/ML Systems | Paul Borrill (Deadealus) |
| 9:30 | 9:50 | RAS API as a common IB/OOB interface by GPUs and other system components | Yogesh Varma |
| 9:40 | 10:00 | Redfish Telemetry Enhancements | Jeff Autor (Vertiv) |
| 10:00 | 10:20 | Liquid Leak Detection | Sean McDaniel (Chemelex) |
| 10:20 | 10:30 | Closing and Next Steps | John Leung (Intel) Rob Coyle (OCP) |
