• DocumentCode
    2320831
  • Title

    Checkpointing Orchestration: Toward a Scalable HPC Fault-Tolerant Environment

  • Author

    Jin, Hui ; Ke, Tao ; Chen, Yong ; Sun, Xian-He

  • Author_Institution
    Illinois Inst. of Technol., Chicago, IL, USA
  • fYear
    2012
  • fDate
    13-16 May 2012
  • Firstpage
    276
  • Lastpage
    283
  • Abstract
    Check pointing is widely used in technical computing. However, the overhead of check pointing is a subject of increasing in concern in recent years, especially for large-scale parallel computer systems. In these systems, check pointing generates a huge number of concurrent I/O writes. The burst of writes plus the worsening I/O-wall problem often leads to network and I/O congestion, and makes the overall system performance painfully slow. Recognizing contention as a dominant performance factor, in this paper we propose a systematic approach named check pointing orchestration to reduce write contention, which combines the marshaling of concurrent checkpoint requests and the adopting of vertical data access in coordination. A prototype of the proposed check pointing orchestration approach has been implemented at the system-level under Open MPI over the PVFS2 file system. Extensive experiments based on NPB benchmarks have been conducted to verify the design and implementation. Experimental results show that check pointing orchestration reduced the check pointing cost at a degree of more than 30%. Check pointing cost was halved for 4 out of 5 the C class NPB benchmarks.
  • Keywords
    application program interfaces; checkpointing; message passing; software fault tolerance; C class NPB benchmarks; I-O congestion; I-O-wall problem; Open MPI; PVFS2 file system; checkpointing orchestration approach; concurrent I-O writes; concurrent checkpoint requests; dominant performance factor; high-performance computing; large-scale parallel computer systems; scalable HPC fault-tolerant environment; technical computing; vertical data access; write contention reduction; Bandwidth; Benchmark testing; Checkpointing; Fault tolerance; Fault tolerant systems; File systems; Servers; Checkpointing; Fault Tolerance; Parallel File System;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Cluster, Cloud and Grid Computing (CCGrid), 2012 12th IEEE/ACM International Symposium on
  • Conference_Location
    Ottawa, ON
  • Print_ISBN
    978-1-4673-1395-7
  • Type

    conf

  • DOI
    10.1109/CCGrid.2012.61
  • Filename
    6217432