Title :
Automatic reconfiguration in the presence of failures
Author :
Cristian, Flaviu
Author_Institution :
Dept. of Comput. Sci. & Eng., California Univ., San Diego, La Jolla, CA, USA
fDate :
3/1/1993 12:00:00 AM
Abstract :
The paper describes a new kind of distributed system service, the availability management service, responsible for ensuring that the critical services of a distributed system remain continuously available to users despite arbitrary numbers of concurrent node removals and node restarts caused by failures, maintenance, and growth. It stresses the main ideas behind this new service, and outlines a simple design that depends on the existence of synchronous membership and atomic broadcast group communication services. Extensions of this initial design to deal with asynchronous group communication services are also briefly discussed
Keywords :
distributed processing; software reliability; atomic broadcast group communication services; automatic reconfiguration; availability management service; concurrent node removals; critical services; distributed system service; failures; node restarts; synchronous membership;
Journal_Title :
Software Engineering Journal