DocumentCode
2446927
Title
Efficient Metadata Generation to Enable Interactive Data Discovery over Large-Scale Scientific Data Collections
Author
Pallickara, Sangmi Lee ; Pallickara, Shrideep ; Zupanski, Milija ; Sullivan, Stephen
Author_Institution
Dept. of Comput. Sci., Colorado State Univ., Fort Collins, CO, USA
fYear
2010
fDate
Nov. 30 2010-Dec. 3 2010
Firstpage
573
Lastpage
580
Abstract
Discovering the correct dataset efficiently is critical for computations and effective simulations in scientific experiments. In contrast to searching web documents over the Internet, massive binary datasets are difficult to browse or search. Users must select a reliable data publisher from the large collection of data services available over the Internet. Once a publisher is selected, the user must then discover the dataset that matches the computation´s needs, among tens of thousands of large data packages that are available. Some of the data hosting services provide advanced data search interfaces but their search scope is often limited to local datasets. Because scientific datasets are often encoded as binary data formats, querying or validating missing data over hundreds of Megabytes of a binary file involves a compute intensive decoding process. We have developed a system, GLEAN, that provides an efficient data discovery environment for users in scientific computing. Fine-grained metadata is automatically extracted to provide a micro view and profile of the large dataset to the users. We have used the Granules cloud runtime to orchestrate the MapReduce computations that extract metadata from the datasets. Here we focus on the overall architecture of the system and how it enables efficient data discovery. We applied our framework to a data discovery application in the atmospheric science domain. This paper includes a performance evaluation with observational datasets.
Keywords
Web services; data analysis; data mining; data visualisation; meta data; scientific information systems; GLEAN; Internet; MapReduce computations; Web document search; binary data; data decoding; data publisher; data querying; data search interfaces; data services; granules cloud; interactive data discovery; large scale scientific data; metadata extraction; metadata generation; scientific computing; Atmospheric modeling; Catalogs; Computational modeling; Databases; Decoding; Feature extraction; Processor scheduling; atmospheric sciences; cloud computing; data discovery; large-scale datasets; metadata;
fLanguage
English
Publisher
ieee
Conference_Titel
Cloud Computing Technology and Science (CloudCom), 2010 IEEE Second International Conference on
Conference_Location
Indianapolis, IN
Print_ISBN
978-1-4244-9405-7
Electronic_ISBN
978-0-7695-4302-4
Type
conf
DOI
10.1109/CloudCom.2010.99
Filename
5708502
Link To Document