Title :
A Runtime Fault Detection Method for HPC Cluster
Author :
Linping, Wu ; Hongbing, Luo ; Jianfeng, Zhan ; Dan, Meng
Author_Institution :
HPCC, Inst. of Appl. Phys. & Comput. Mathematic, Beijing, China
Abstract :
As the number of nodes keeps increasing, faults have become commonplace for HPC cluster. For fast recovery from faults, the fault detection method is necessary. Based on the usage patterns of HPC cluster, a automatic runtime fault detection mechanism is proposed in this paper: First, the normal activities for nodes in HPC cluster are modeled using runtime state by clustering analysis, Second, the fault detection process is implemented by comparing the current runtime state of nodes with normal activity models. A fault alarm is made immediately when the current runtime state deviates from the normal activity models. In the experiments, the faults are simulated by fault injection methods and the experimental results show that the runtime fault detection method in this paper can detect faults with high accuracy.
Keywords :
fault diagnosis; parallel processing; pattern clustering; HPC cluster analysis; automatic runtime fault detection mechanism; fault alarm; fault injection method; fault recovery; normal activity model; runtime state; Aging; Computational modeling; Fault detection; Runtime; Software; System recovery; Vectors; HPC cluster; Runtime fault detection; SOM;
Conference_Titel :
Parallel and Distributed Computing, Applications and Technologies (PDCAT), 2011 12th International Conference on
Conference_Location :
Gwangju
Print_ISBN :
978-1-4577-1807-6
DOI :
10.1109/PDCAT.2011.9