DocumentCode :
2866865
Title :
A Runtime Fault Detection Method for HPC Cluster
Author :
Linping, Wu ; Hongbing, Luo ; Jianfeng, Zhan ; Dan, Meng
Author_Institution :
HPCC, Inst. of Appl. Phys. & Comput. Mathematic, Beijing, China
fYear :
2011
fDate :
20-22 Oct. 2011
Firstpage :
68
Lastpage :
72
Abstract :
As the number of nodes keeps increasing, faults have become commonplace for HPC cluster. For fast recovery from faults, the fault detection method is necessary. Based on the usage patterns of HPC cluster, a automatic runtime fault detection mechanism is proposed in this paper: First, the normal activities for nodes in HPC cluster are modeled using runtime state by clustering analysis, Second, the fault detection process is implemented by comparing the current runtime state of nodes with normal activity models. A fault alarm is made immediately when the current runtime state deviates from the normal activity models. In the experiments, the faults are simulated by fault injection methods and the experimental results show that the runtime fault detection method in this paper can detect faults with high accuracy.
Keywords :
fault diagnosis; parallel processing; pattern clustering; HPC cluster analysis; automatic runtime fault detection mechanism; fault alarm; fault injection method; fault recovery; normal activity model; runtime state; Aging; Computational modeling; Fault detection; Runtime; Software; System recovery; Vectors; HPC cluster; Runtime fault detection; SOM;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Parallel and Distributed Computing, Applications and Technologies (PDCAT), 2011 12th International Conference on
Conference_Location :
Gwangju
Print_ISBN :
978-1-4577-1807-6
Type :
conf
DOI :
10.1109/PDCAT.2011.9
Filename :
6118961
Link To Document :
بازگشت