• DocumentCode
    3481903
  • Title

    An analysis of the effect of training data variation in English-Persian Statistical Machine Translation

  • Author

    Mohaghegh, Mahsa ; Sarrafzadeh, Abdolhossein

  • Author_Institution
    IIMS, Massey Univ., Auckland, New Zealand
  • fYear
    2009
  • fDate
    15-17 Dec. 2009
  • Firstpage
    105
  • Lastpage
    109
  • Abstract
    Globalization has made machine translation an attractive area of research and development. As technology opens up e-commerce opportunities, companies must overcome language barriers to reach new potential customers and partners. Web2.0 with tools like Google Translate makes the web more accessible. Statistical Machine Translation has been used for translation between many language pairs contributing to its popularity in recent years. It has however not been used for the English/Persian pair. This paper presents the first such attempt and describes the problems faced in creating a corpus and building a base line system. Our experience with the construction of a parallel corpus during this study and the problems encountered especially with the process of alignment are discussed. The prototype constructed and its evaluation is described and results analyzed. In the final part of the paper, conclusions are drawn and work planned for the future is discussed.
  • Keywords
    Internet; globalisation; language translation; natural language processing; research and development; statistical analysis; English-Persian statistical machine translation; Google Translate; Web 2.0; data variation training; e-commerce; globalization; parallel corpus; research and development; Buildings; Globalization; HTML; Humans; Natural languages; Probability; Prototypes; Research and development; Surface-mount technology; Training data;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Innovations in Information Technology, 2009. IIT '09. International Conference on
  • Conference_Location
    Al Ain
  • Print_ISBN
    978-1-4244-5698-7
  • Type

    conf

  • DOI
    10.1109/IIT.2009.5413782
  • Filename
    5413782