• DocumentCode
    560177
  • Title

    On the duality of data-intensive file system design: Reconciling HDFS and PVFS

  • Author

    Tantisiriroj, Wittawat ; Patil, Swapnil ; Gibson, Garth ; Son, Seung Woo ; Lang, Samuel J. ; Ross, Robert B.

  • fYear
    2011
  • fDate
    12-18 Nov. 2011
  • Firstpage
    1
  • Lastpage
    12
  • Abstract
    Data-intensive applications fall into two computing styles: Internet services (cloud computing) or high-performance computing (UPC). In both categories, the underlying file system is a key component for scalable application performance. In this paper, we explore the similarities and differences between PVFS, a parallel file system used in UPC at large scale, and HDFS, the primary storage system used in cloud computing with Hadoop. We integrate PVFS into Hadoop and compare its performance to HDFS using a set of data-intensive computing benchmarks. We study how HDFS-specific optimizations can be matched using PVFS and how consistency, durability, and persistence tradeoffs made by these file systems affect application performance. We show how to embed multiple replicas into a PVFS file, including a mapping with a complete copy local to the writing client, to emulate HDFS´s file layout policies. We also highlight implementation issues with HDFS´s dependence on disk bandwidth and benefits from pipelined replication.
  • Keywords
    cloud computing; file organisation; HDFS; Internet services; PVFS; UPC; cloud computing; data-intensive file system design; high-performance computing; Cloud computing; Distributed databases; Layout; Prefetching; Semantics; Servers; Web and internet services; HDFS; Hadoop; PVFS; cloud computing; file systems;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    High Performance Computing, Networking, Storage and Analysis (SC), 2011 International Conference for
  • Conference_Location
    Seatle, WA
  • Electronic_ISBN
    978-1-4503-0771-0
  • Type

    conf

  • Filename
    6114443