Hardware Management/Hardware Fault Management
Jump to navigation
Jump to search
You are not logged in. Your IP address will be publicly visible if you make any edits. If you
create an account
, your edits will be attributed to your username, along with other benefits.
Anti-spam check. Do
fill this in!
==Project Overview== The Hardware Fault Management subproject’s goal is to address the Large scale data center pain-points in managing hardware faults effectively. Objective of the subproject: to define hardware fault management solution such that it - Minimizes performance degradation - Reduces No Trouble Found (NTF) in RMA - Scales the solution across heterogeneous fleet Activities of the sub-project include: - Develop running list of pain-points related to HW fault handling in the fleet - Develop hardware fault management solution best practices to address the above pain-points - Hardware Error reporting format Standardization - Define hardware error classification - OCP Fault Management Infrastructure Proposal - Standardizing system behavior under hardware failures (future topic) - Provide reference and guidance on system hardware failure management (future topic)
Please note that all contributions to OpenCompute may be edited, altered, or removed by other contributors. If you do not want your writing to be edited mercilessly, then do not submit it here.
You are also promising us that you wrote this yourself, or copied it from a public domain or similar free resource (see
Do not submit copyrighted work without permission!
(opens in new window)
Retrieved from "
Not logged in
Help about MediaWiki
What links here