• DocumentCode
    3134681
  • Title

    Kepler: an extensible system for design and execution of scientific workflows

  • Author

    Altintas, Ilkay ; Berkley, Chad ; Jaeger, Efrat ; Jones, Matthew ; Ludascher, Bertram ; Mock, Steve

  • Author_Institution
    San Diego Supercomput. Center, California Univ., San Diego, CA, USA
  • fYear
    2004
  • fDate
    21-23 June 2004
  • Firstpage
    423
  • Lastpage
    424
  • Abstract
    Most scientists conduct analyses and run models in several different software and hardware environments, mentally coordinating the export and import of data from one environment to another. The Kepler scientific workflow system provides domain scientists with an easy-to-use yet powerful system for capturing scientific workflows (SWFs). SWFs are a formalization of the ad-hoc process that a scientist may go through to get from raw data to publishable results. Kepler attempts to streamline the workflow creation and execution process so that scientists can design, execute, monitor, re-run, and communicate analytical procedures repeatedly with minimal effort. Kepler is unique in that it seamlessly combines high-level workflow design with execution and runtime interaction, access to local and remote data, and local and remote service invocation. SWFs are superficially similar to business process workflows but have several challenges not present in the business workflow scenario. For example, they often operate on large, complex and heterogeneous data, can be computationally intensive and produce complex derived data products that may be archived for use in reparameterized runs or other workflows. Moreover, unlike business workflows, SWFs are often dataflow-oriented as witnessed by a number of recent academic systems (e.g., DiscoveryNet, Taverna and Triana) and commercial systems (Scitegic/Pipeline-Pilot, Inforsense). In a sense, SWFs are often closer to signal-processing and data streaming applications than they are to control-oriented business workflow applications.
  • Keywords
    data flow computing; data handling; database management systems; scientific information systems; workflow management software; DiscoveryNet; Inforsense; Kepler scientific workflow system; Scitegic/Pipeline-Pilot; Taverna; Triana; complex data; data access; data export; data import; data products; data streaming; extensible system; hardware environment; heterogeneous data; high-level workflow design; large data; runtime interaction; scientific analysis; scientific workflow design; scientific workflow execution; service invocation; software environment; workflow creation; Biological system modeling; Business; Java; Plugs; Power system modeling; Prototypes; Runtime; Supercomputers; Web services; Yarn;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Scientific and Statistical Database Management, 2004. Proceedings. 16th International Conference on
  • ISSN
    1099-3371
  • Print_ISBN
    0-7695-2146-0
  • Type

    conf

  • DOI
    10.1109/SSDM.2004.1311241
  • Filename
    1311241