• DocumentCode
    1956813
  • Title

    Real-time GPU-based software beamformer designed for advanced imaging methods research

  • Author

    Yiu, Billy Y S ; Tsang, Ivan K H ; Yu, Alfred C H

  • Author_Institution
    Med. Eng. Program, Univ. of Hong Kong, Hong Kong, China
  • fYear
    2010
  • fDate
    11-14 Oct. 2010
  • Firstpage
    1920
  • Lastpage
    1923
  • Abstract
    High computational demand is known to be a technical hurdle for real-time implementation of advanced methods like synthetic aperture imaging (SAI) and plane wave imaging (PWI) that work with the pre-beamform data of each array element. In this paper, we present the development of a software beamformer for SAI and PWI with real-time parallel processing capacity. Our beamformer design comprises a pipelined group of graphics processing units (GPU) that are hosted within the same computer workstation. During operation, each available GPU is assigned to perform demodulation and beamforming for one frame of pre-beamform data acquired from one transmit firing (e.g. point firing for SAI). To facilitate parallel computation, the GPUs have been programmed to treat the calculation of depth pixels from the same image scanline as a block of processing threads that can be executed concurrently, and it would repeat this process for all scanlines to obtain the entire frame of image data - i.e. low-resolution image (LRI). To reduce processing latency due to repeated access of each GPU´s global memory, we have made use of each thread block´s fast-shared memory (to store an entire line of pre-beamform data during demodulation), created texture memory pointers, and utilized global memory caches (to stream repeatedly used data samples during beamforming). Based on this beamformer architecture, a prototype platform has been implemented for SAI and PWI, and its LRI processing throughput has been measured for test datasets with 40 MHz sampling rate, 32 receive channels, and imaging depths between 5-15 cm. When using two Fermi-class GPUs (GTX-470), our beamformer can compute LRIs of 512-by-255 pixels at over 3200 fps and 1300 fps respectively for imaging depths of 5 cm and 15 cm. This processing throughput is roughly 3.2 times higher than a Tesla-class GPU (GTX-275).
  • Keywords
    biomedical ultrasonics; coprocessors; demodulation; image processing; medical image processing; parallel processing; software architecture; ultrasonic imaging; Fermi-class GPU; beamformer architecture; beamformer design; demodulation; frequency 40 MHz; global memory cache; graphics processing unit; low-resolution image; parallel computation; plane wave imaging; real-time GPU; real-time parallel processing capacity; software beamformer architecture; synthetic aperture imaging; ultrasound imaging; Array signal processing; Graphics processing unit; Imaging; Instruction sets; Pixel; Real time systems; Ultrasonic imaging; graphics processing units; parallel processing; plane wave imaging; software beamformer; synthetic aperture imaging;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Ultrasonics Symposium (IUS), 2010 IEEE
  • Conference_Location
    San Diego, CA
  • ISSN
    1948-5719
  • Print_ISBN
    978-1-4577-0382-9
  • Type

    conf

  • DOI
    10.1109/ULTSYM.2010.5935689
  • Filename
    5935689