• DocumentCode
    3102447
  • Title

    Challenges in Developing Persian Corpora from Online Resources

  • Author

    Ghayoomi, Masood ; Momtazi, Saeedeh

  • Author_Institution
    Dept. of Comput. Linguistics, Saarland Univ., Saarbrucken, Germany
  • fYear
    2009
  • fDate
    7-9 Dec. 2009
  • Firstpage
    108
  • Lastpage
    113
  • Abstract
    Persian is one of the Indo-European languages which has borrowed its script from Arabic, a member of Semitic language family. Since Persian and Arabic scripts are so similar, problems arise when we want to process an electronic text. In this paper, some of the common problems faced experimentally in developing a corpus for Persian from on-line materials are discussed. The sources of the problems are the Persian script itself; mixture with the Arabic script; Persian orthography; the typists´ typing styles; and mixing Persian code pages with Arabic code pages in operating systems.
  • Keywords
    natural language processing; text analysis; Persian corpora; Semitic language family; electronic text processing; online resources; Books; Computational linguistics; Mood; Natural languages; Operating systems; Spatial databases; Speech; Telephony; Web pages; Writing;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Asian Language Processing, 2009. IALP '09. International Conference on
  • Conference_Location
    Singapore
  • Print_ISBN
    978-0-7695-3904-1
  • Type

    conf

  • DOI
    10.1109/IALP.2009.31
  • Filename
    5380780