DocumentCode :
2100520
Title :
Automated detection and segmentation of table of contents page and index pages from document images
Author :
Mandal, S. ; Chowdhury, S.P. ; Das, A.K. ; Chanda, Bhabatosh
Author_Institution :
CST Dept., B. E. Coll., Howrah, India
fYear :
2003
fDate :
17-19 Sept. 2003
Firstpage :
213
Lastpage :
218
Abstract :
The requirement of identifying and segmenting the table of contents (TOC) and index pages in the development of a digital library is obvious. A digital document library is created to provide a non-labour intensive, cheap and flexible way of storing, representing and managing paper documents in electronic form to facilitate indexing, viewing, printing and extracting the intended portions. Information from the TOC and index pages is extracted to use in a document database for effective retrieval of the required pieces of information. We present fully automatic identification and segmentation of TOC and index pages from a scanned document.
Keywords :
digital libraries; document image processing; image classification; image segmentation; library automation; object detection; text analysis; automated page detection; automated page segmentation; digital document library; digital library; document database; document images; index pages; information extraction; information retrieval; table of contents page;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Image Analysis and Processing, 2003.Proceedings. 12th International Conference on
Print_ISBN :
0-7695-1948-2
Type :
conf
DOI :
10.1109/ICIAP.2003.1234052
Filename :
1234052
Link To Document :
بازگشت