مرکز منطقه ای اطلاع رساني علوم و فناوري - Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

DocumentCode :

85141

Title :

Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

Author :

Khayyat, Ahmad ; Manjikian, Naraig

Author_Institution :

Dept. of Comput. Eng., King Fahd Univ. of Pet. & Miner., Dhahran, Saudi Arabia

Volume :

Issue :

fYear :

2014

fDate :

Spring 2014

Firstpage :

Lastpage :

Abstract :

This paper considers blocking and scheduling for the design and implementation of field-programmable gate array (FPGA)-based floating-point parallel matrix multiplication in the presence of a memory hierarchy. For high performance, on-chip memory holds data that are reused when the computation is divided into blocks, and multiple arithmetic units perform independent operations within each block in parallel. The first contribution of this paper is a detailed analysis of the design space to characterize performance based on the amount of on-chip memory used and the approaches considered for blocking and scheduling of the computation. A comparison is also made to prior work with a unified view. The second contribution is a flexible high-performance implementation for the Altera Stratix IV EP4SGX530C2 FPGA with an interface to external double-data-rate synchronous dynamic RAM (DDR2 SDRAM) memory. Various configuration options support optimization of different objectives, and the resulting configurations have been verified in simulation and in hardware. For double-precision floating-point, a performance of 16 giga-floating-point operations per second (GFLOPS) is achievable with 64 arithmetic units at 160 MHz.

Keywords :

DRAM chips; SRAM chips; field programmable gate arrays; floating point arithmetic; logic design; matrix multiplication; scheduling; Altera Stratix IV EP4SGX530C2 FPGA; DDR2 SDRAM memory; FPGA-based floating-point matrix multiplication; GFLOPS; blocking analysis; design space analysis; double-precision floating-point; external double-data-rate synchronous dynamic RAM memory; field-programmable gate array; frequency 160 MHz; giga-floating-point operations per second; memory hierarchy; multiple arithmetic units; on-chip memory; scheduling analysis; Field programmable gate arrays; Memory management; Parallel processing; SDRAM; Schedules; System-on-chip; Accelerator architectures; floating-point arithmetic; matrices; parallel architectures; reconfigurable logic;

fLanguage :

English

Journal_Title :

Electrical and Computer Engineering, Canadian Journal of

Publisher :

ieee

ISSN :

0840-8688

Type :

jour

DOI :

10.1109/CJECE.2014.2317983

Filename :

6850120

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=85141