DocumentCode
3102447
Title
Challenges in Developing Persian Corpora from Online Resources
Author
Ghayoomi, Masood ; Momtazi, Saeedeh
Author_Institution
Dept. of Comput. Linguistics, Saarland Univ., Saarbrucken, Germany
fYear
2009
fDate
7-9 Dec. 2009
Firstpage
108
Lastpage
113
Abstract
Persian is one of the Indo-European languages which has borrowed its script from Arabic, a member of Semitic language family. Since Persian and Arabic scripts are so similar, problems arise when we want to process an electronic text. In this paper, some of the common problems faced experimentally in developing a corpus for Persian from on-line materials are discussed. The sources of the problems are the Persian script itself; mixture with the Arabic script; Persian orthography; the typists´ typing styles; and mixing Persian code pages with Arabic code pages in operating systems.
Keywords
natural language processing; text analysis; Persian corpora; Semitic language family; electronic text processing; online resources; Books; Computational linguistics; Mood; Natural languages; Operating systems; Spatial databases; Speech; Telephony; Web pages; Writing;
fLanguage
English
Publisher
ieee
Conference_Titel
Asian Language Processing, 2009. IALP '09. International Conference on
Conference_Location
Singapore
Print_ISBN
978-0-7695-3904-1
Type
conf
DOI
10.1109/IALP.2009.31
Filename
5380780
Link To Document