Title :
Implementing web data extraction and making Mashup with Xtractorz
Author :
Gultom, Rudy A G ; Sari, Riri Fitri ; Budiardjo, Bagio
Author_Institution :
Dept. of Electr. Eng., Univ. of Indonesia, Depok, Indonesia
Abstract :
Implementing web data extraction means we can directly extract data from various web pages, where they mostly formed in an unstructured HTML format, into a new structured format such as XML or XHTML. In this paper we review the implementation of web data extraction and stages in making a Mashup. We implement web data extraction by visually extract targeted data from data sources (web pages). Afterward, we combined web data extraction with the stages of making a Mashup, e.g. data retrieval, data source modeling, data cleaning/ filtering, data integration and data visualization. Problems arise in querying data sources due to unstructured contents of web pages (HTML), we cannot directly extract data into a new structured form. To address this problem, we propose a system, called Xtractorz, that can perform web data extraction in a Mashup format. We provide a fully visual and interactive user interface with new technique and approach using PHP and AJAX as the programming languages, and MySQL as the Data Repository. Furthermore, Xtractorz enables the user to conduct their job without the need to write a script or program or even without any knowledge of computer programming. The test results shows that Xtractorz requires less number of steps in making a Mashup compared with RoboMaker and Karma.
Keywords :
data handling; Mashup; Web data extraction; XHTML; XML; Xtractorz; data cleaning/; data filtering; data integration; data retrieval; data sources; data visualization; Cleaning; Data mining; Data visualization; HTML; Information filtering; Information filters; Information retrieval; Mashups; Web pages; XML; AJAX; DOM tree; HTML; Making Mashup; Mashup Stages; MySQL; PHP; Web Data Extraction; XML;
Conference_Titel :
Advance Computing Conference (IACC), 2010 IEEE 2nd International
Conference_Location :
Patiala
Print_ISBN :
978-1-4244-4790-9
Electronic_ISBN :
978-1-4244-4791-6
DOI :
10.1109/IADCC.2010.5422921