• DocumentCode
    2904882
  • Title

    Statistical study on a literary Romanian corpus for the beginning and ending of the words

  • Author

    Mitrea, Adrian ; Vlad, Adriana ; Luca, Adrian

  • Author_Institution
    Fac. of Electron., Telecommun. & Inf. Technol., Politeh. Univ. of Bucharest, Bucharest, Romania
  • fYear
    2012
  • fDate
    21-23 June 2012
  • Firstpage
    81
  • Lastpage
    84
  • Abstract
    The paper attempts to investigate the statistical structure of letters and of letter digrams with which the words begin and end, as well as of trigrams that link two successive words. The investigation is carried out on a printed Romanian literary corpus summing up about 12.5 million words. The impact of the orthography and punctuation marks in the language model assigned to the beginning and to the ending of words is considered.
  • Keywords
    grammars; natural language processing; word processing; language model; letter digrams; letters statistical structure; orthography marks; printed Romanian literary corpus; punctuation marks; successive words; trigrams; Analytical models; Artificial intelligence; Educational institutions; Joining processes; Mathematical model; Pragmatics; Probability; m-grams by which words begin and finish; mathematics of natural language; orthography and punctuation marks;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Communications (COMM), 2012 9th International Conference on
  • Conference_Location
    Bucharest
  • Print_ISBN
    978-1-4577-0057-6
  • Type

    conf

  • DOI
    10.1109/ICComm.2012.6262546
  • Filename
    6262546