• DocumentCode
    855809
  • Title

    A prototype document image analysis system for technical journals

  • Author

    Nagy, George ; Seth, Sharad ; Viswanathan, Mahesh

  • Author_Institution
    Dept. of Electr.-Comput.-Syst. Eng., Rensselaer Polytech. Inst., Troy, NY, USA
  • Volume
    25
  • Issue
    7
  • fYear
    1992
  • fDate
    7/1/1992 12:00:00 AM
  • Firstpage
    10
  • Lastpage
    22
  • Abstract
    Gobbledoc, a system providing remote access to stored documents, which is based on syntactic document analysis and optical character recognition (OCR), is discussed. In Gobbledoc, image processing, document analysis, and OCR operations take place in batch mode when the documents are acquired. The document image acquisition process and the knowledge base that must be entered into the system to process a family of page images are described. The process by which the X-Y tree data structure converts a 2-D page-segmentation problem into a series of 1-D string-parsing problems that can be tackled using conventional compiler tools is also described. Syntactic analysis is used in Gobbledoc to divide each page into labeled rectangular blocks. Blocks labeled text are converted by OCR to obtain a secondary (ASCII) document representation. Since such symbolic files are better suited for computerized search than for human access to the document content and because too many visual layout clues are lost in the OCR process (including some special characters), Gobbledoc preserves the original block images for human browsing. Storage, networking, and display issues specific to document images are also discussed.<>
  • Keywords
    computerised picture processing; data structures; document image processing; optical character recognition; ASCII document representation; Gobbledoc; X-Y tree data structure; compiler tools; knowledge base; optical character recognition; prototype document image analysis system; syntactic document analysis; technical journals; Character recognition; Humans; Image analysis; Image converters; Image processing; Image storage; Optical character recognition software; Prototypes; Text analysis; Tree data structures;
  • fLanguage
    English
  • Journal_Title
    Computer
  • Publisher
    ieee
  • ISSN
    0018-9162
  • Type

    jour

  • DOI
    10.1109/2.144436
  • Filename
    144436