• DocumentCode
    174803
  • Title

    Computing Defects per Million in Cloud Caused by Virtual Machine Failures with Replication

  • Author

    Mondal, Subrota K. ; Muppala, Jogesh K. ; Machida, Fumio ; Trivedi, Kishor S.

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Hong Kong Univ. of Sci. & Technol., Kowloon, China
  • fYear
    2014
  • fDate
    18-21 Nov. 2014
  • Firstpage
    161
  • Lastpage
    168
  • Abstract
    Virtual machines (VM) are used in cloud computing systems to handle user requests for service. A typical user request goes through several cloud service provider specific processing steps from the instant it is submitted until the service is completed. In the process of providing the service, VM failures cause the user´s request to be dropped. To mitigate the adverse impact of VM failure, replication mechanisms, either using cold, warm or hot replication, can be used. In this paper, we model the system behavior with a structure-state process to characterize the failure-recovery behavior of a VM in a cloud that uses one of the aforementioned replication schemes. We use a service-oriented dependability metric called Defects Per Million (DPM), defined as the number of user requests dropped out of a million. The structure-state process approach is used to analyze the job completion time distribution and subsequently we compute the DPM by counting the number of requests exceed the specified deadline. The effectiveness of replication schemes are demonstrated through numerical results.
  • Keywords
    cloud computing; software fault tolerance; virtual machines; DPM; VM; cloud computing systems; cloud service provider specific processing steps; cold replication; defects per million; hot replication; job completion time distribution; replication mechanisms; structure-state process approach; user requests; virtual machine failures; warm replication; Availability; Computational modeling; Equations; Mathematical model; Measurement; Transforms; Virtual machining; Cloud; Defects Per Million; Fault tolerance; Job completion time; Replication;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Dependable Computing (PRDC), 2014 IEEE 20th Pacific Rim International Symposium on
  • Conference_Location
    Singapore
  • Print_ISBN
    978-1-4799-6473-4
  • Type

    conf

  • DOI
    10.1109/PRDC.2014.29
  • Filename
    6974785