• DocumentCode
    1811862
  • Title

    Web Content Extraction based on Webpage Layout Analysis

  • Author

    Fu, Lei ; Meng, Yao ; Xia, YingJu ; Yu, Hao

  • Author_Institution
    Fujitsu R&D Center CO., Ltd., Beijing, China
  • fYear
    2010
  • fDate
    24-25 July 2010
  • Firstpage
    40
  • Lastpage
    43
  • Abstract
    For web content extraction task, researchers have proposed many different methods, such as wrapper-based method, DOM tree rule-based method, machine learning-based method and so on. To some extent, all these methods ignore the layout information of the webpage, although the layout information such as the spatial and visual cues often plays a very important role in the process of locating the main content of the webpage when browsing. As a consequence, these methods often throw part of the main content away when extracting content from the webpage. In this paper, we present a method which combines webpage layout analysis with DOM tree rule-base method, it can make full use of the advantages of the two methods. It uses the layout information to guide the extraction work with a global view and can gain a better performance than the traditional methods.
  • Keywords
    Web design; information retrieval; knowledge based systems; learning (artificial intelligence); trees (mathematics); DOM tree rule-based method; Web content extraction task; Webpage layout analysis; document object model; layout information; machine learning-based method; wrapper-based method; Algorithm design and analysis; Bismuth; Data mining; HTML; Layout; Learning systems; Particle separators; DOM tree rule-based method; web content extraction; webpage layout analysis;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Information Technology and Computer Science (ITCS), 2010 Second International Conference on
  • Conference_Location
    Kiev
  • Print_ISBN
    978-1-4244-7293-2
  • Electronic_ISBN
    978-1-4244-7294-9
  • Type

    conf

  • DOI
    10.1109/ITCS.2010.16
  • Filename
    5557336