• DocumentCode
    1877187
  • Title

    High Performance Matrix Multiplication on General Purpose Graphics Processing Units

  • Author

    Wu, Fan ; Cabral, Miguel ; Brazelton, Jessica

  • Author_Institution
    Comput. Sci. Dept., Tuskegee Univ., Tuskegee, AL, USA
  • fYear
    2010
  • fDate
    10-12 Dec. 2010
  • Firstpage
    1
  • Lastpage
    4
  • Abstract
    In recent years, there has been significant interest from both academia and industry in applying commodity graphics processing units (GPUs) toward general computing problems. The nVidia CUDA programming model provides a straightforward means of describing inherently parallel computations. In this paper, we present our GPU-based matrix multiplication with high performance on General Purpose Graphics Processing Unit (GPGPUs). We implemented our algorithm using nVidia CUDA API and compared its performance with an optimized CPU-implementation on a high-end AMD Opteron Dual Core CPU. Our experimental results show that a significant performance improvement over CPU-based algorithm and the maximum observed speedups are about 100 times.
  • Keywords
    coprocessors; matrix multiplication; AMD Opteron Dual Core CPU; general purpose graphics processing units; matrix multiplication; nVidia CUDA programming; parallel computations; Central Processing Unit; Computer architecture; Graphics processing unit; Instruction sets; Kernel; Programming;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computational Intelligence and Software Engineering (CiSE), 2010 International Conference on
  • Conference_Location
    Wuhan
  • Print_ISBN
    978-1-4244-5391-7
  • Electronic_ISBN
    978-1-4244-5392-4
  • Type

    conf

  • DOI
    10.1109/CISE.2010.5677044
  • Filename
    5677044