DocumentCode
1626953
Title
Implementing web data extraction and making Mashup with Xtractorz
Author
Gultom, Rudy A G ; Sari, Riri Fitri ; Budiardjo, Bagio
Author_Institution
Dept. of Electr. Eng., Univ. of Indonesia, Depok, Indonesia
fYear
2010
Firstpage
385
Lastpage
393
Abstract
Implementing web data extraction means we can directly extract data from various web pages, where they mostly formed in an unstructured HTML format, into a new structured format such as XML or XHTML. In this paper we review the implementation of web data extraction and stages in making a Mashup. We implement web data extraction by visually extract targeted data from data sources (web pages). Afterward, we combined web data extraction with the stages of making a Mashup, e.g. data retrieval, data source modeling, data cleaning/ filtering, data integration and data visualization. Problems arise in querying data sources due to unstructured contents of web pages (HTML), we cannot directly extract data into a new structured form. To address this problem, we propose a system, called Xtractorz, that can perform web data extraction in a Mashup format. We provide a fully visual and interactive user interface with new technique and approach using PHP and AJAX as the programming languages, and MySQL as the Data Repository. Furthermore, Xtractorz enables the user to conduct their job without the need to write a script or program or even without any knowledge of computer programming. The test results shows that Xtractorz requires less number of steps in making a Mashup compared with RoboMaker and Karma.
Keywords
data handling; Mashup; Web data extraction; XHTML; XML; Xtractorz; data cleaning/; data filtering; data integration; data retrieval; data sources; data visualization; Cleaning; Data mining; Data visualization; HTML; Information filtering; Information filters; Information retrieval; Mashups; Web pages; XML; AJAX; DOM tree; HTML; Making Mashup; Mashup Stages; MySQL; PHP; Web Data Extraction; XML;
fLanguage
English
Publisher
ieee
Conference_Titel
Advance Computing Conference (IACC), 2010 IEEE 2nd International
Conference_Location
Patiala
Print_ISBN
978-1-4244-4790-9
Electronic_ISBN
978-1-4244-4791-6
Type
conf
DOI
10.1109/IADCC.2010.5422921
Filename
5422921
Link To Document