• DocumentCode
    2050784
  • Title

    Towards building a highly-available cluster based model for high performance computing

  • Author

    Boukerche, Azzedine ; Al-Shaikh, Raed ; Notare, M.S.M.

  • Author_Institution
    SITE, Ottawa Univ., Ont., Canada
  • fYear
    2006
  • fDate
    25-29 April 2006
  • Abstract
    In recent years, we have witnessed a growing interest in high performance computing (HPC) using a cluster of workstations. However, many challenges remain to be resolved before these systems become dependable. One of the challenges in a clustered environment is to keep system failure to the minimum level and while achieving the highest possible level of system availability. High-availability (HA) computing attempts to avoid the problems of unexpected failures through active redundancy and preemptive measures. In this paper, we propose to build HA-clusters based model for high performance computing. Our model is based on combination of both HPC and HA concepts, we also propose to investigate further the hardware and the management layers of the HA-HPC cluster design, and the parallel-applications layer (i.e. FT-MPI implementations). In this work, we focus upon the latter layer. We discuss our model, and present our simulation experiments we have carried out to evaluate our proposed model.
  • Keywords
    workstation clusters; high performance computing; high-availability computing; highly-available cluster model; workstation cluster; Availability; Computational modeling; Distributed computing; Fault tolerance; Hardware; High performance computing; Laboratories; Redundancy; Supercomputers; Workstations;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel and Distributed Processing Symposium, 2006. IPDPS 2006. 20th International
  • Conference_Location
    Rhodes Island
  • Print_ISBN
    1-4244-0054-6
  • Type

    conf

  • DOI
    10.1109/IPDPS.2006.1639633
  • Filename
    1639633