DocumentCode
3229594
Title
A faster checkpointing and recovery algorithm with a hierarchical storage approach
Author
Gao, Wen ; Chen, Mingyu ; Nanya, Takashi
Author_Institution
Inst. of Comput. Technol., Chinese Acad. of Sci.
fYear
2005
fDate
1-1 July 2005
Lastpage
402
Abstract
Fault tolerance is an inevitable part of cluster operating system. In Score cluster system, it provides coordinated checkpointing, rollback recovery mechanism and watch-dog timer detector for fault tolerance. In the checkpointing algorithm in Score, disk write is the bottleneck. To eliminate disk write overhead, this paper proposes a new diskless checkpointing and rollback recovery algorithm. Since the proposed algorithm does not need to calculate parity and write the checkpointing data into disk, it is analyzed to be a faster checkpointing algorithm than the original one. Based on comparison, the recovery time of the proposed algorithm is also less. However, the cluster can not tolerant multiple transient failure using this diskless checkpointing algorithm. To compensate this drawback, a hierarchical storage strategy is adopted. An experimental result shows that this diskless algorithm with a hierarchical storage approach is fast and effective
Keywords
checkpointing; fault tolerant computing; operating systems (computers); storage management; cluster operating system; diskless checkpointing; fault tolerance; hierarchical storage; rollback recovery; watch-dog timer detector; Algorithm design and analysis; Checkpointing; Clustering algorithms; Computers; Concurrent computing; Detectors; Fault detection; Fault tolerant systems; High performance computing; Operating systems;
fLanguage
English
Publisher
ieee
Conference_Titel
High-Performance Computing in Asia-Pacific Region, 2005. Proceedings. Eighth International Conference on
Conference_Location
Beijing
Print_ISBN
0-7695-2486-9
Type
conf
DOI
10.1109/HPCASIA.2005.2
Filename
1592295
Link To Document