DocumentCode
3350932
Title
Runtime system for autonomic rescheduling of MPI programs
Author
Du, Cong ; Ghosh, Sumonto ; Shankar, Shashank ; Sun, Xian-He
Author_Institution
Dept. of Comput. Sci., Illinois Inst. of Technol., Chicago, IL, USA
fYear
2004
fDate
15-18 Aug. 2004
Firstpage
4
Abstract
Intensive research has been conducted on dynamic job scheduling, which dynamically allocates jobs to computing systems. However, most of the existing work is limited to redistribute independent tasks or at the algorithm design level. There is no runtime system available to support automatic redistribution of a running process in a heterogeneous network environment. In this study, we present the design and implementation of a system that dynamically reschedules running processes over a network of computing resources via automatic decision-making and process migration. The system is implemented on top of MPI-2 and HPCM (high performance computing mobility) middleware. Experimental and analytical results show that the runtime system works well. It makes dynamic rescheduling of running tasks possible and improves system performance considerably. While the implementation is for MPI programs and using HPCM, the design of the system is general and can be extended to other distributed environments as well.
Keywords
decision making; dynamic scheduling; grid computing; message passing; middleware; MPI programs; automatic decision-making; autonomic rescheduling; distributed environment; dynamic job scheduling; heterogeneous network environment; high performance computing mobility; middleware; process migration; runtime system; system performance; Computer networks; Decision making; Distributed computing; Dynamic scheduling; Message passing; Middleware; Processor scheduling; Resource management; Runtime environment; Sun;
fLanguage
English
Publisher
ieee
Conference_Titel
Parallel Processing, 2004. ICPP 2004. International Conference on
ISSN
0190-3918
Print_ISBN
0-7695-2197-5
Type
conf
DOI
10.1109/ICPP.2004.1327898
Filename
1327898
Link To Document