Title of article :
Geometric algorithms and experiments for automated document structuring
Author/Authors :
Rus، نويسنده , , D. and Summers، نويسنده , , K.، نويسنده ,
Issue Information :
روزنامه با شماره پیاپی سال 1997
Pages :
29
From page :
55
To page :
83
Abstract :
We present and analyze algorithms for the automated segmentation and classification of layout structures in electronic documents. The key idea is to use the patterns in the distribution of white space in a document to recognize and interpret its components. The segmentation algorithm divides the document into a hierarchy of logical elements; the classification algorithms classify these divisions as base-text, tables, indented lists, polygonal drawings, and graphs. We present experimental data and discuss an information access application. Our methodology allows the automatic markup of documents (for instance in the sgml format) and the creation of multilevel indices and browsing tools for electronic libraries.
Keywords :
document analysis , information access , Document structure , Information capture
Journal title :
Mathematical and Computer Modelling
Serial Year :
1997
Journal title :
Mathematical and Computer Modelling
Record number :
1590874
Link To Document :
بازگشت