• DocumentCode
    2989510
  • Title

    Exploiting concurrent kernel execution on graphic processing units

  • Author

    Wang, Lingyuan ; Huang, Miaoqing ; El-Ghazawi, Tarek

  • fYear
    2011
  • fDate
    4-8 July 2011
  • Firstpage
    24
  • Lastpage
    32
  • Abstract
    Graphics processing units (CPUs) have been accepted as a powerful and viable coprocessor solution in high-performance computing domain. In order to maximize the benefit of CPUs for a multicore platform, a mechanism is needed for CPU threads in a parallel application to share this computing resource for efficient execution. NVIDIA´s Fermi architecture pioneers the feature of concurrent kernel execution; however, only kernels of the same thread context can execute in parallel. In order to get the best use of a GPU device in a multi-threaded application environment, this paper explores the techniques to effectively share a context, i.e., context funneling, which could be done either manually at application level, or automatically at the GPU runtime starting from CUDA v4.0. For synthetic microbenchmark tests, we find that both funneling mechanisms are more capable of exploring the benefit of concurrent kernel execution than traditional context switching, therefore improving the overall application performance. We also find that the manual funneling mechanism provides the highest performance and more explicit control, while CUDA v4.0 provides better productivity with good performance. Finally, we assess the impact of such techniques on a compact application benchmark, SSCA#3 - SAR sensor processing.
  • Keywords
    computer graphic equipment; concurrency control; coprocessors; multi-threading; multiprocessing systems; parallel processing; resource allocation; CPU threads; CUDA v4.0; Fermi architecture; NVIDIA; SAR sensor processing; SSCA#3; compact application benchmark; computing resource sharing; concurrent kernel execution; context funneling; coprocessor solution; graphic processing units; high-performance computing domain; multicore platform; multithreaded application environment; parallel application; synthetic microbenchmark tests; Context; Graphics processing unit; Instruction sets; Kernel; Manuals; Programming; Switches; Concurrent kernel execution; GPU computing; Multi-threaded programming;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    High Performance Computing and Simulation (HPCS), 2011 International Conference on
  • Conference_Location
    Istanbul
  • Print_ISBN
    978-1-61284-380-3
  • Type

    conf

  • DOI
    10.1109/HPCSim.2011.5999803
  • Filename
    5999803