• DocumentCode
    3462379
  • Title

    OAMS: A Highly Reliable Metadata Service for Big Data Storage

  • Author

    Jiang Zhou ; Jing Guo ; Weiping Wang ; Cuilan Du ; Xiaoyan Gu ; Dan Meng

  • Author_Institution
    Inst. of Comput. Technol., Comput. Applic. Res. Center, Grad. Univ. of Chinese Acad. of Sci., Beijing, China
  • fYear
    2013
  • fDate
    3-5 Dec. 2013
  • Firstpage
    1287
  • Lastpage
    1294
  • Abstract
    As application requirements increase in quantity and popularity, big data storage is becoming an important technology which data centers and Internet companies depend on. The cluster file system with centralized metadata management often encounters planned or unplanned downtime which requires higher reliability for metadata service. Current paradigms use the backup server to take over as the primary when the latter is in the case of failures. But if the backup crashes, the file system is still in an unreliable state. In this paper, we present a novel primary-backup policy (OAMS), which ensures the availability of metadata service in cluster file system. Different from traditional paradigms, OAMS employs multiple standbys to tolerate the single point of failure. It is based on the built-in shared storage pool for metadata synchronization and a series of protocols for active election, active-standby switching and etc. By using a prepared, automatic state transition among metadata servers, OAMS achieves an automatic recovery in the form of hot standby. It also supports server self-recovery and dynamical addition for standbys at runtime. Evaluation results show that OAMS obviously improves the reliability of metadata service while the average performance degradation is below 8%.
  • Keywords
    Big Data; back-up procedures; meta data; software reliability; storage management; Internet company; OAMS; active-standby switching; automatic recovery; automatic state transition; average performance degradation; backup server; big data storage; built-in shared storage pool; centralized metadata management; cluster file system; data centers; metadata servers; metadata service reliability; metadata synchronization; primary-backup policy; server self-recovery; unplanned downtime; Computer crashes; Manuals; Protocols; Reliability; Servers; Switches; Synchronization; big data; cluster file system; high reliability; metadata service; primary-backup paradigm;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computational Science and Engineering (CSE), 2013 IEEE 16th International Conference on
  • Conference_Location
    Sydney, NSW
  • Type

    conf

  • DOI
    10.1109/CSE.2013.191
  • Filename
    6755373