DocumentCode
1814964
Title
Grid collector: an event catalog with automated file management
Author
Wu, Kesheng ; Wei-Ming Zlang ; Sim, Alexander ; Gu, Junmin ; Shoshani, Arie
Author_Institution
Lawrence Berkeley Nat. Lab, CA, USA
Volume
2
fYear
2003
fDate
19-25 Oct. 2003
Firstpage
848
Abstract
High Energy Nuclear Physics (HENP) experiments such as STAR at BNL and ATLAS at CERN produce large amounts of data that are stored as files on mass storage systems in computer centers. In these files, the basic unit of data is an event. Analysis is typically performed on a selected set of events. The files containing these events have to be located, copied from mass storage systems to disks before analysis, and removed when no longer needed. These file management tasks are tedious and time consuming. Typically, all events contained in the files are read into memory before a selection is made. Since the time to read the events dominate the overall execution time, reading the unwanted event needlessly increases the analysis time. The Grid Collector is a set of software modules that works together to address these two issues. It automates the file management tasks and provides "direct" access to the selected events for analyses. It is currently integrated with the STAR analysis framework. The users can select events based on tags, such as, "production date between March 10 and 20, and the number of charged tracks > 100:" The Grid Collector locates the files containing relevant events, transfers the files across the Grid if necessary, and delivers the events to the analysis code through the familiar iterators. There has been some research efforts to address the file management issues, the Grid Collector is unique in that it addresses the event access issue together with the file management issues. This makes it more useful to a large varieties of users.
Keywords
data acquisition; data analysis; file organisation; high energy physics instrumentation computing; ATLAS; BNL; CERN; STAR; automated file management; data analysis; event catalog; file management tasks; grid collector; mass storage systems; Energy management; Energy storage; Information analysis; Laboratories; Nuclear physics; Performance analysis; Physics computing; Production; Software performance; Storage automation;
fLanguage
English
Publisher
ieee
Conference_Titel
Nuclear Science Symposium Conference Record, 2003 IEEE
ISSN
1082-3654
Print_ISBN
0-7803-8257-9
Type
conf
DOI
10.1109/NSSMIC.2003.1351830
Filename
1351830
Link To Document