DocumentCode
2999521
Title
Distributed Virtual Diskless Checkpointing: A Highly Fault Tolerant Scheme for Virtualized Clusters
Author
Eckart, Ben ; He, Xubin ; Wu, Chentao ; Aderholdt, Ferrol ; Han, Fang ; Scott, Stephen
Author_Institution
Robot. Inst., Carnegie Mellon Univ., Pittsburgh, PA, USA
fYear
2012
fDate
21-25 May 2012
Firstpage
1120
Lastpage
1127
Abstract
Today´s high-end computing systems are facing a crisis of high failure rates due to increased numbers of components. Recent studies have shown that traditional fault tolerant techniques incur overheads that more than double execution times on these highly parallel machines. Thus, future high-end computing must be able to provide adequate fault tolerance at an acceptable cost or the burdens of fault management will severely affect the viability of such systems. Cluster virtualization offers a potentially unique solution for fault management, but brings significant overhead, especially for I/O. In this paper, we propose a novel diskless check pointing technique on clusters of virtual machines. Our technique splits Virtual Machines into sets of orthogonal RAID systems and distributes parity evenly across the cluster, similar to a RAID-5 configuration, but using VM images as data elements. Our theoretical analysis shows that our technique significantly reduces the overhead associated with check pointing by removing the disk I/O bottleneck.
Keywords
RAID; checkpointing; distributed processing; fault tolerant computing; parallel machines; virtual machines; virtualisation; VM image; cluster virtualization; data element; disk I/O bottleneck removal; distributed virtual diskless checkpointing; failure rate; fault management; fault tolerant scheme; high-end computing system; orthogonal RAID system; parallel machine; virtual machine; virtualized cluster; Checkpointing; Fault tolerance; Fault tolerant systems; Hardware; Libraries; Program processors; Virtual machining; cluster virtualization; diskless checkpointing; fault tolerance; live migration;
fLanguage
English
Publisher
ieee
Conference_Titel
Parallel and Distributed Processing Symposium Workshops & PhD Forum (IPDPSW), 2012 IEEE 26th International
Conference_Location
Shanghai
Print_ISBN
978-1-4673-0974-5
Type
conf
DOI
10.1109/IPDPSW.2012.136
Filename
6270762
Link To Document