• DocumentCode
    2445519
  • Title

    Correlation Based File Prefetching Approach for Hadoop

  • Author

    Dong, Bo ; Zhong, Xiao ; Zheng, Qinghua ; Jian, Lirong ; Liu, Jian ; Qiu, Jie ; Li, Ying

  • Author_Institution
    MOE KLINNS Lab., Xi´´an Jiaotong Univ., Xi´´an, China
  • fYear
    2010
  • fDate
    Nov. 30 2010-Dec. 3 2010
  • Firstpage
    41
  • Lastpage
    48
  • Abstract
    Hadoop Distributed File System (HDFS) has been widely adopted to support Internet applications because of its reliable, scalable and low-cost storage capability. Blue Sky, one of the most popular e-Learning resource sharing systems in China, is utilizing HDFS to store massive courseware. However, due to the inefficient access mechanism of HDFS, access latency of reading files from HDFS significantly impacts the performance of processing user requests. This paper introduces a two-level correlation based file prefetching approach, taking the characteristics of HDFS into consideration, to improve performance by reducing access latency. Four placement patterns to store prefetched data are presented, with policies to achieve trade-off between performance and efficiency of HDFS prefetching. Moreover, a dynamic replica selection algorithm is investigated to improve the efficiency of HDFS prefetching. The proposed prefetching approach has been implemented in Blue Sky, and experimental results prove that correlation based file prefetching can significantly reduce access latency therefore improve performance of Hadoop-based Internet applications.
  • Keywords
    Internet; correlation methods; distributed databases; network operating systems; storage management; BlueSky; China; Hadoop distributed file system; Hadoop-based Internet applications; access latency; correlation based file prefetching approach; courseware; dynamic replica selection algorithm; e-Learning resource sharing systems; placement patterns; Correlation; File systems; Heuristic algorithms; Internet; Prefetching; Servers; Throughput; Hadoop distributed file system; cloud storage; file correlation; prefetching;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Cloud Computing Technology and Science (CloudCom), 2010 IEEE Second International Conference on
  • Conference_Location
    Indianapolis, IN
  • Print_ISBN
    978-1-4244-9405-7
  • Electronic_ISBN
    978-0-7695-4302-4
  • Type

    conf

  • DOI
    10.1109/CloudCom.2010.60
  • Filename
    5708432