DocumentCode :
3586045
Title :
Lexico-syntactic normalization model for noisy SMS text
Author :
Jose, Greety ; Raj, Nisha S.
Author_Institution :
Dept. of Comput. Sci., SCMS Sch. of Eng. & Technol., Ernakulam, India
fYear :
2014
Firstpage :
163
Lastpage :
168
Abstract :
Today, digital mediated interactions and communications being an important constituent. The expeditious growth of electronic communications such as Emails, micro blogs, SMS and chats etc has fabricated extensively noisy forms of text. It predominantly in young urbanites. The tremendous growth of noises in text are due to a variety of factors, such as the small number of characters allowed per text messages (160 characters is allowed per SMS and 140 characters allowed per tweets), inventing new abbreviations, using non standard orthographic forms, phonetic substitution etc. In this paper we introduce a lexico-syntactic normalization model for cleaning the noisy texts. The normalization is based on the channelized database and a user feedback system. The syntactic analysis of sentences is based on a bottom up parser. The model will capture the user interaction for improving the model accuracy. Precursory evaluation shows that the channel model will normalize the noisy word to their standard peer with better accuracy. The sentence validation achieved 95.7% accuracy.
Keywords :
computer mediated communication; electronic messaging; grammars; language translation; natural language processing; bottom up parser; channelized database; digital mediated communication; digital mediated interaction; electronic communication; lexico-syntactic normalization model; noisy SMS text; user feedback system; Computational modeling; Databases; Dictionaries; Natural language processing; Noise measurement; Standards; Syntactics; Lexical Normalization; Machine Translation; Natural Language Processing; Noisy words; Non-noisy word; Parser; SMS; Social Media; Text Normalization;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Electronics,Communication and Computational Engineering (ICECCE), 2014 International Conference on
Type :
conf
DOI :
10.1109/ICECCE.2014.7086652
Filename :
7086652
Link To Document :
بازگشت