The Open Systems for AI – Systems Management Workshop #1 focused on advancing data center automation and fault management through discussions of Redfish standards and fault detection processes. Presentations covered various aspects of Redfish implementation, including power management modeling, thermal equipment support, and the development of standardized interfaces and tools for error analysis and mitigation. The workshop concluded with discussions on creating scalable solutions for data center diagnostics and debug workstream, emphasizing the need for standardized tools across complex AI systems.
Workshop Agenda
| Start | End | Title | Speakers |
|---|---|---|---|
| 8:00 | 8:10 | Welcome and Introduction [Slides] | John Leung (Intel) Rob Coyle (OCP) |
| 8:10 | 8:30 | Redfish Location Resource Part Location Context & Service Label [Slides] | Justin York (Google) |
| 8:30 | 8:50 | Hardware Fault Management Framework as a system level framework that provides a fault management structure used by the GPU and other system components [Slides] | Drew Walton (Microsoft) |
| 8:50 | 9:10 | Power Management [Slides] | Mike Rainer (Dell) |
| 9:10 | 9:30 | Cooling Management [Slides] | Jeff Autor (Vertiv) |
| 9:30 | 9:50 | SPDM = Security Protocol and Data Model [Slides] | Jiewen Yao (Intel) |
| 9:50 | 10:10 | Datacenter diagnostics and Debug [Slides] | Marko Bartscherer (Intel) |
| 10:10 | 10:20 | Closing and Next Steps | Rob Coyle (OCP) |
