• DocumentCode
    1858650
  • Title

    Cache Accurate Time Skewing in Iterative Stencil Computations

  • Author

    Strzodka, Robert ; Shaheen, Mohammed ; Paja, Dawid ; Seidel, Hans-Peter

  • Author_Institution
    Max Planck Inst. Inf., Saarbrucken, Germany
  • fYear
    2011
  • fDate
    13-16 Sept. 2011
  • Firstpage
    571
  • Lastpage
    581
  • Abstract
    We present a time skewing algorithm that breaks the memory wall for certain iterative stencil computations. A stencil computation, even with constant weights, is a completely memory-bound algorithm. For example, for a large 3D domain of 5003 doubles and 100 iterations on a quad-core Xeon X5482 3.2GHz system, a hand-vectorized and parallelized naive 7-point stencil implementation achieves only 1.4 GFLOPS because the system memory bandwidth limits the performance. Although many efforts have been undertaken to improve the performance of such nested loops, for large data sets they still lag far behind synthetic benchmark performance. The state-of-art automatic locality optimizer PluTo achieves 3.7 GFLOPS for the above stencil, whereas a parallel benchmark executing the inner stencil computation directly on registers performs at 25.1 GFLOPS. In comparison, our algorithm achieves 13.0 GFLOPS (52% of the stencil peak benchmark).We present results for 2D and 3D domains in double precision including problems with gigabyte large data sets. The results are compared against hand-optimized naive schemes, PluTo, the stencil peak benchmark and results from literature. For constant stencils of slope one we break the dependence on the low system bandwidth and achieve at least 50% of the stencil peak, thus performing within a factor two of an ideal system with infinite bandwidth (the benchmark runs on registers without memory access). For large stencils and banded matrices the additional data transfers let the limitations of the system bandwidth come into play again, however, our algorithm still gains a large improvement over the other schemes.
  • Keywords
    cache storage; electronic data interchange; iterative methods; mathematics computing; multiprocessing systems; natural sciences computing; PluTo; automatic locality optimizer; cache accurate time skewing; completely memory bound algorithm; data transfers; hand vectorized naive 7-point stencil implementation; iterative stencil computations; parallelized naive 7-point stencil implementation; quad core Xeon X5482; stencil peak benchmark; system memory bandwidth limits; Bandwidth; Cats; Diamond-like carbon; Instruction sets; Synchronization; Three dimensional displays; Tiles; banded matrix; memory bound; memory wall; stencil; temporal blocking; time skewing; wavefront;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel Processing (ICPP), 2011 International Conference on
  • Conference_Location
    Taipei City
  • ISSN
    0190-3918
  • Print_ISBN
    978-1-4577-1336-1
  • Electronic_ISBN
    0190-3918
  • Type

    conf

  • DOI
    10.1109/ICPP.2011.47
  • Filename
    6047225