Title :
Mining Web informative structures and contents based on entropy analysis
Author :
Kao, Hung-Yu ; Lin, Shian-Hua ; Ho, Jan-Ming ; Chen, Ming-Syan
Author_Institution :
Dept. of Electr. Eng., Nat. Taiwan Univ., Taipei, Taiwan
Abstract :
We study the problem of mining the informative structure of a news Web site that consists of thousands of hyperlinked documents. We define the informative structure of a news Web site as a set of index pages (or referred to as TOC, i.e., table of contents, pages) and a set of article pages linked by these TOC pages. Based on the Hyperlink Induced Topics Search (HITS) algorithm, we propose an entropy-based analysis (LAMIS) mechanism for analyzing the entropy of anchor texts and links to eliminate the redundancy of the hyperlinked structure so that the complex structure of a Web site can be distilled. However, to increase the value and the accessibility of pages, most of the content sites tend to publish their pages with intrasite redundant information, such as navigation panels, advertisements, copy announcements, etc. To further eliminate such redundancy, we propose another mechanism, called InfoDiscoverer, which applies the distilled structure to identify sets of article pages. InfoDiscoverer also employs the entropy information to analyze the information measures of article sets and to extract informative content blocks from these sets. Our result is useful for search engines, information agents, and crawlers to index, extract, and navigate significant information from a Web site. Experiments on several real news Web sites show that the precision and the recall of our approaches are much superior to those obtained by conventional methods in mining the informative structures of news Web sites. On the average, the augmented LAMIS leads to prominent performance improvement and increases the precision by a factor ranging from 122 to 257 percent when the desired recall falls between 0.5 and 1. In comparison with manual heuristics, the precision and the recall of InfoDiscoverer are greater than 0.956.
Keywords :
Web sites; data mining; information retrieval; publishing; search engines; text analysis; HITS; InfoDiscoverer; LAMIS; TOC pages; Web crawlers; Web informative structure mining; advertisements; anchor texts; article pages; copy announcements; entropy analysis; entropy information; entropy-based analysis mechanism; hyperlink induced topics search; hyperlinked documents; index pages; information agents; information extraction; information measures; informative content blocks; intrasite redundant information; navigation panels; news Web site; search engines; table of contents; Algorithm design and analysis; Computer Society; Data mining; Entropy; IEEE news; Information analysis; Internet; Navigation; Search engines; Web pages;
Journal_Title :
Knowledge and Data Engineering, IEEE Transactions on
DOI :
10.1109/TKDE.2004.1264821