• DocumentCode
    3443906
  • Title

    Efficient checkpointing over local area networks

  • Author

    Ziv, Avi ; Bruck, Jehoshua

  • Author_Institution
    Inf. Syst. Lab., Stanford Univ., CA, USA
  • fYear
    1994
  • fDate
    12-14 Jun 1994
  • Firstpage
    30
  • Lastpage
    35
  • Abstract
    Parallel and distributed computing on clusters of workstations is becoming very popular as it provides a cost effective way for high performance computing. In these systems, the bandwidth of the communication subsystem (using Ethernet technology) is about an order of magnitude smaller compared to the bandwidth of the storage subsystem. Hence, storing a state in a checkpoint is much more efficient than comparing states over the network. In this paper we present a novel checkpointing approach that enables efficient performance over local area networks. The main idea is that we use two types of checkpoints: compare-checkpoints (comparing the states of the redundant processes to detect faults) and store-checkpoints (where the state is only stored). The store-checkpoints reduce the rollback needed after a fault is detected, without performing many unnecessary comparisons. As a particular example of this approach we analyzed the DMR checkpointing scheme with store-checkpoints. Our main result is that the overhead of the execution time can be significantly reduced when store-checkpoints are introduced. We have implemented a prototype of the new DMR scheme and run it on workstations connected by a LAN. The experimental results we obtained match the analytical results and show that in some cases the overhead of the DMR checkpointing schemes over LAN´s can be improved by as much as 20%
  • Keywords
    fault tolerant computing; local area networks; Ethernet technology; checkpointing; communication subsystem; compare-checkpoints; high performance computing; local area networks; rollback; storage subsystem; store-checkpoints; Bandwidth; Checkpointing; Costs; Distributed computing; Ethernet networks; Fault detection; High performance computing; Local area networks; Prototypes; Workstations;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Fault-Tolerant Parallel and Distributed Systems, 1994., Proceedings of IEEE Workshop on
  • Conference_Location
    College Station, TX
  • Print_ISBN
    0-8186-6807-5
  • Type

    conf

  • DOI
    10.1109/FTPDS.1994.494471
  • Filename
    494471