• DocumentCode
    1626953
  • Title

    Implementing web data extraction and making Mashup with Xtractorz

  • Author

    Gultom, Rudy A G ; Sari, Riri Fitri ; Budiardjo, Bagio

  • Author_Institution
    Dept. of Electr. Eng., Univ. of Indonesia, Depok, Indonesia
  • fYear
    2010
  • Firstpage
    385
  • Lastpage
    393
  • Abstract
    Implementing web data extraction means we can directly extract data from various web pages, where they mostly formed in an unstructured HTML format, into a new structured format such as XML or XHTML. In this paper we review the implementation of web data extraction and stages in making a Mashup. We implement web data extraction by visually extract targeted data from data sources (web pages). Afterward, we combined web data extraction with the stages of making a Mashup, e.g. data retrieval, data source modeling, data cleaning/ filtering, data integration and data visualization. Problems arise in querying data sources due to unstructured contents of web pages (HTML), we cannot directly extract data into a new structured form. To address this problem, we propose a system, called Xtractorz, that can perform web data extraction in a Mashup format. We provide a fully visual and interactive user interface with new technique and approach using PHP and AJAX as the programming languages, and MySQL as the Data Repository. Furthermore, Xtractorz enables the user to conduct their job without the need to write a script or program or even without any knowledge of computer programming. The test results shows that Xtractorz requires less number of steps in making a Mashup compared with RoboMaker and Karma.
  • Keywords
    data handling; Mashup; Web data extraction; XHTML; XML; Xtractorz; data cleaning/; data filtering; data integration; data retrieval; data sources; data visualization; Cleaning; Data mining; Data visualization; HTML; Information filtering; Information filters; Information retrieval; Mashups; Web pages; XML; AJAX; DOM tree; HTML; Making Mashup; Mashup Stages; MySQL; PHP; Web Data Extraction; XML;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Advance Computing Conference (IACC), 2010 IEEE 2nd International
  • Conference_Location
    Patiala
  • Print_ISBN
    978-1-4244-4790-9
  • Electronic_ISBN
    978-1-4244-4791-6
  • Type

    conf

  • DOI
    10.1109/IADCC.2010.5422921
  • Filename
    5422921