DocumentCode
3134681
Title
Kepler: an extensible system for design and execution of scientific workflows
Author
Altintas, Ilkay ; Berkley, Chad ; Jaeger, Efrat ; Jones, Matthew ; Ludascher, Bertram ; Mock, Steve
Author_Institution
San Diego Supercomput. Center, California Univ., San Diego, CA, USA
fYear
2004
fDate
21-23 June 2004
Firstpage
423
Lastpage
424
Abstract
Most scientists conduct analyses and run models in several different software and hardware environments, mentally coordinating the export and import of data from one environment to another. The Kepler scientific workflow system provides domain scientists with an easy-to-use yet powerful system for capturing scientific workflows (SWFs). SWFs are a formalization of the ad-hoc process that a scientist may go through to get from raw data to publishable results. Kepler attempts to streamline the workflow creation and execution process so that scientists can design, execute, monitor, re-run, and communicate analytical procedures repeatedly with minimal effort. Kepler is unique in that it seamlessly combines high-level workflow design with execution and runtime interaction, access to local and remote data, and local and remote service invocation. SWFs are superficially similar to business process workflows but have several challenges not present in the business workflow scenario. For example, they often operate on large, complex and heterogeneous data, can be computationally intensive and produce complex derived data products that may be archived for use in reparameterized runs or other workflows. Moreover, unlike business workflows, SWFs are often dataflow-oriented as witnessed by a number of recent academic systems (e.g., DiscoveryNet, Taverna and Triana) and commercial systems (Scitegic/Pipeline-Pilot, Inforsense). In a sense, SWFs are often closer to signal-processing and data streaming applications than they are to control-oriented business workflow applications.
Keywords
data flow computing; data handling; database management systems; scientific information systems; workflow management software; DiscoveryNet; Inforsense; Kepler scientific workflow system; Scitegic/Pipeline-Pilot; Taverna; Triana; complex data; data access; data export; data import; data products; data streaming; extensible system; hardware environment; heterogeneous data; high-level workflow design; large data; runtime interaction; scientific analysis; scientific workflow design; scientific workflow execution; service invocation; software environment; workflow creation; Biological system modeling; Business; Java; Plugs; Power system modeling; Prototypes; Runtime; Supercomputers; Web services; Yarn;
fLanguage
English
Publisher
ieee
Conference_Titel
Scientific and Statistical Database Management, 2004. Proceedings. 16th International Conference on
ISSN
1099-3371
Print_ISBN
0-7695-2146-0
Type
conf
DOI
10.1109/SSDM.2004.1311241
Filename
1311241
Link To Document