• DocumentCode
    2674254
  • Title

    Fault-tolerant message switching based on wormhole switching and backtracking

  • Author

    Sueishi, Manabu ; Kitakami, Masato ; Ito, Hideo

  • Author_Institution
    Graduate Sch. of Sci. & Technol., Chiba Univ., Japan
  • fYear
    2004
  • fDate
    3-5 March 2004
  • Firstpage
    183
  • Lastpage
    190
  • Abstract
    Parallel computers are now popularly applied to applications where many calculations are required. In a NO Remote memory Access model (NORA) parallel computer, many processors are connected by communication links and calculation results are obtained by communications among processors. The message switching method, which controls message transmission in the parallel computer, is one of the most important parameters to improve the performance of the parallel computer. Since parallel computers include many processors, its failure rate is very high and many fault-tolerant switching methods have been proposed. The existing methods have problems, however, such as low communication throughput, low fault-tolerant capability, and large hardware overhead. We propose fault-tolerant switching by improving wormhole switching. The proposed method inserts dummy flits, having no information, after the header flit, the first flit of the packet. By overwriting the header flit to the dummy flit, backtracking is implemented without hardware overhead. Computer simulation says that in a 16 by 16 2D torus, for example, the throughput of the proposed method is almost equal to that of existing methods which require large hardware overhead if the number of the faulty nodes is less then 40.
  • Keywords
    backtracking; fault tolerant computing; message switching; multiprocessor interconnection networks; parallel machines; NO Remote memory Access model; NORA parallel computer; fault-tolerant message switching method; large hardware overhead; low communication throughput; low fault-tolerant capability; wormhole switching; Communication switching; Computer aided instruction; Concurrent computing; Delay; Error analysis; Fault tolerance; Hardware; Indium tin oxide; Switching circuits; Throughput;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Dependable Computing, 2004. Proceedings. 10th IEEE Pacific Rim International Symposium on
  • Print_ISBN
    0-7695-2076-6
  • Type

    conf

  • DOI
    10.1109/PRDC.2004.1276569
  • Filename
    1276569