DocumentCode
3013042
Title
Checkpointing in hybrid distributed systems
Author
Cao, Jiannong ; Chen, Yifeng ; Zhang, Kang ; He, Yanxiang
Author_Institution
Dept. of Comput., Hong Kong Polytech. Univ., China
fYear
2004
fDate
10-12 May 2004
Firstpage
136
Lastpage
141
Abstract
To provide fault tolerance to computer systems suffering from transient faults, checkpointing and rollback recovery is widely-used. Among other techniques, two primary checkpointing schemes have been proposed: independent and coordinated schemes. However, most existing work addresses only the need to employ a single checkpointing and rollback recovery scheme to a target system. In this paper, issues are discussed and a new algorithm is developed to address the need of integrating independent and coordinated checkpointing schemes for applications running in a hybrid distributed environment containing multiple heterogeneous subsystems. The required changes to the original checkpointing schemes for each subsystem and the overall prevented unnecessary rollbacks for the integrated system are presented. Also described is an algorithm for collecting garbage checkpoints in the combined hybrid system.
Keywords
distributed processing; fault tolerant computing; protocols; storage management; system recovery; checkpointing; computer systems; fault tolerance; garbage checkpoint collection; hybrid distributed systems; integrated system; multiple heterogeneous subsystems; rollback recovery; transient faults; Checkpointing; Computer science; Fault detection; Fault tolerant systems; Grid computing; Helium; Large-scale systems; Message passing; Resource management; Web services;
fLanguage
English
Publisher
ieee
Conference_Titel
Parallel Architectures, Algorithms and Networks, 2004. Proceedings. 7th International Symposium on
ISSN
1087-4089
Print_ISBN
0-7695-2135-5
Type
conf
DOI
10.1109/ISPAN.2004.1300471
Filename
1300471
Link To Document