• DocumentCode
    3605578
  • Title

    Reliability of Erasure Coded Storage Systems: A Combinatorial-Geometric Approach

  • Author

    Vaishampayan, Vinay A. ; Campello, Antonio

  • Author_Institution
    AT&T Shannon Lab., Florham Park, NJ, USA
  • Volume
    61
  • Issue
    11
  • fYear
    2015
  • Firstpage
    5795
  • Lastpage
    5809
  • Abstract
    We consider the probability of data loss, or equivalently, the reliability function for an erasure coded distributed data storage system under worst case conditions. Data loss in an erasure coded system depends on probability distributions for the disk repair duration and the disk failure duration. In previous works, the data loss probability of such systems has been studied under the assumption of exponentially distributed disk failure and disk repair durations, using well-known analytic methods from the theory of Markov processes. These methods lead to an estimate of the integral of the reliability function. Here, we address the problem of directly calculating the data loss probability for general repair and failure duration distributions. A closed limiting form is developed for the probability of data loss, and it is shown that the probability of the event that a repair duration exceeds a failure duration is sufficient for characterizing the data loss probability. For the case of constant repair duration, we develop an expression for the conditional data loss probability given the number of failures experienced by a each node in a given time window. We do so by developing a geometric approach that relies on the computation of volumes of a family of polytopes that are related to the code. An exact calculation is provided, and an upper bound on the data loss probability is obtained by posing the problem as a set avoidance problem. Theoretical calculations are compared with simulation results.
  • Keywords
    Markov processes; combinatorial mathematics; geometric codes; reliability theory; statistical distributions; Markov process; combinatorial-geometric approach; conditional data loss probability; erasure coded storage system reliability; probability distribution; reliability function integral estimation; set avoidance problem; Distributed databases; Limiting; Maintenance engineering; Probability; Reliability theory; Upper bound; Data Loss Probability; Erasure Code; Error Analysis; Fault Tolerant Systems; MDS Code; Polytopes; Reliability; Renewal Theory;
  • fLanguage
    English
  • Journal_Title
    Information Theory, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    0018-9448
  • Type

    jour

  • DOI
    10.1109/TIT.2015.2477401
  • Filename
    7247711