DocumentCode
2904882
Title
Statistical study on a literary Romanian corpus for the beginning and ending of the words
Author
Mitrea, Adrian ; Vlad, Adriana ; Luca, Adrian
Author_Institution
Fac. of Electron., Telecommun. & Inf. Technol., Politeh. Univ. of Bucharest, Bucharest, Romania
fYear
2012
fDate
21-23 June 2012
Firstpage
81
Lastpage
84
Abstract
The paper attempts to investigate the statistical structure of letters and of letter digrams with which the words begin and end, as well as of trigrams that link two successive words. The investigation is carried out on a printed Romanian literary corpus summing up about 12.5 million words. The impact of the orthography and punctuation marks in the language model assigned to the beginning and to the ending of words is considered.
Keywords
grammars; natural language processing; word processing; language model; letter digrams; letters statistical structure; orthography marks; printed Romanian literary corpus; punctuation marks; successive words; trigrams; Analytical models; Artificial intelligence; Educational institutions; Joining processes; Mathematical model; Pragmatics; Probability; m-grams by which words begin and finish; mathematics of natural language; orthography and punctuation marks;
fLanguage
English
Publisher
ieee
Conference_Titel
Communications (COMM), 2012 9th International Conference on
Conference_Location
Bucharest
Print_ISBN
978-1-4577-0057-6
Type
conf
DOI
10.1109/ICComm.2012.6262546
Filename
6262546
Link To Document