• DocumentCode
    3756332
  • Title

    Intra-Clustering: Accelerating On-chip Communication for Data Parallel Architectures

  • Author

    Wen Yuan;Rahul Boyapati;Lei Wang;Hyunjun Jang;Yuho Jin;Ki Hwan Yum;Eun Jung Kim

  • Author_Institution
    Samsung Austin Res. Center, Austin, TX, USA
  • fYear
    2015
  • Firstpage
    55
  • Lastpage
    60
  • Abstract
    Modern computation workloads contain abundant Data Level Parallelism (DLP), which requires specialized data parallel architectures, such as Graphics Processing Units (GPUs). With parallel programming models, such as CUDA and OpenCL, GPUs are easily to be programmed for non-graphics applications, and therefore become a cost effective approach for data parallel architectures. The large quantity of available parallelism places a heavy stress on the memory system as the limited number of pins confines the number of memory controllers on the chip. This creates a potential bottleneck for performance scalability of the GPUs. To accelerate communication with the memory system, we propose the Intra-Clustering on-chip network for data parallel architectures, which is built upon a traditional two-dimensional electrical mesh network with memory controllers connected through a nanophotonic ring and compute cores grouped into different clusters. Our evaluations with CUDA benchmarks show that the Intra-Clustering architecture can improve communication delay by an average of 17% (up to 32%) and IPC by an average of 5% (up to 11.5%).
  • Keywords
    "Routing","Graphics processing units","Benchmark testing","Optical waveguides","Parallel architectures","Delays"
  • Publisher
    ieee
  • Conference_Titel
    Computer Architecture and High Performance Computing Workshop (SBAC-PADW), 2015 International Symposium on
  • Type

    conf

  • DOI
    10.1109/SBAC-PADW.2015.15
  • Filename
    7423181