Title :
Handwritten Numeral Databases of Indian Scripts and Multistage Recognition of Mixed Numerals
Author :
Bhattacharya, Ujjwal ; Chaudhuri, Bidyut B.
Author_Institution :
Comput. Vision & Pattern Recognition Unit, Indian Stat. Inst., Kolkata
fDate :
3/1/2009 12:00:00 AM
Abstract :
This article primarily concerns the problem of isolated handwritten numeral recognition of major Indian scripts. The principal contributions presented here are (a) pioneering development of two databases for handwritten numerals of two most popular Indian scripts, (b) a multistage cascaded recognition scheme using wavelet based multiresolution representations and multilayer perceptron classifiers and (c) application of (b) for the recognition of mixed handwritten numerals of three Indian scripts Devanagari, Bangla and English. The present databases include respectively 22,556 and 23,392 handwritten isolated numeral samples of Devanagari and Bangla collected from real-life situations and these can be made available free of cost to researchers of other academic Institutions. In the proposed scheme, a numeral is subjected to three multilayer perceptron classifiers corresponding to three coarse-to-fine resolution levels in a cascaded manner. If rejection occurred even at the highest resolution, another multilayer perceptron is used as the final attempt to recognize the input numeral by combining the outputs of three classifiers of the previous stages. This scheme has been extended to the situation when the script of a document is not known a priori or the numerals written on a document belong to different scripts. Handwritten numerals in mixed scripts are frequently found in Indian postal mails and table-form documents.
Keywords :
handwritten character recognition; image recognition; image resolution; multilayer perceptrons; natural language processing; visual databases; Bangla; Devanagari; Indian scripts; coarse-to-fine resolution levels; handwritten numeral databases; isolated handwritten numeral recognition; mixed numerals; multilayer perceptron; multistage cascaded recognition scheme; multistage recognition; wavelet-based multiresolution representations; Handwriting analysis; Optical character recognition; Algorithms; Automatic Data Processing; Databases, Factual; Handwriting; Image Enhancement; Image Interpretation, Computer-Assisted; India; Information Storage and Retrieval; Pattern Recognition, Automated; Reading; Reference Values;
Journal_Title :
Pattern Analysis and Machine Intelligence, IEEE Transactions on
DOI :
10.1109/TPAMI.2008.88