DocumentCode :
2013291
Title :
RAS Modeling of an HPC Switch System
Author :
Tang, Dong ; Bryson, William ; Elling, Richard
fYear :
2008
fDate :
15-17 Dec. 2008
Firstpage :
81
Lastpage :
86
Abstract :
The high end of high performance computing (HPC) systems is now moving toward petascale deployments, delivering petaflops of computational capacity and petabytes of storage capacity. Interconnection of the sheer number of server nodes in an HPC system plays a vital role in the developments. InfiniBand has emerged as a compelling interconnect technology, and provides more scalability and significantly better cost- performance than any other known protocols. This paper presents a reliability, availability, and serviceability (RAS) modeling and analysis of the Sun Datacenter Switch 3456 system, the world´s largest standards-based InfiniBand switch, with direct capacity to host up to 3,456 server nodes, against hardware faults. The results show that the system reliability, in terms of connectivity between the server nodes physically connected to the switch, is high for configurations with redundant ports. The study also shows that practicing deferred repair strategies can significantly reduce unscheduled service events and system downtime. Further, the study identifies optimal service strategies by a tradeoff analysis on reliability and availability.
Keywords :
multiprocessor interconnection networks; parallel machines; telecommunication switching; HPC switch system; InfiniBand switch; RAS modeling; Sun Datacenter Switch 3456 system; availability; connectivity; high performance computing; interconnect technology; reliability; server nodes; serviceability; Availability; Connectors; Cooling; Delay; Fabrics; High performance computing; Petascale computing; Quality of service; Sun; Switches; Availability; High Performance Computing; Markov Model; Reliability; Serviceability;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Dependable Computing, 2008. PRDC '08. 14th IEEE Pacific Rim International Symposium on
Conference_Location :
Taipei
Print_ISBN :
978-0-7695-3448-0
Electronic_ISBN :
978-0-7695-3448-0
Type :
conf
DOI :
10.1109/PRDC.2008.19
Filename :
4725282
Link To Document :
بازگشت