• DocumentCode
    3402239
  • Title

    Rule based stemmer in Urdu

  • Author

    Gupta, V. ; Joshi, Niranjan ; Mathur, Iti

  • Author_Institution
    IES, IPS Acad., Indore, India
  • fYear
    2013
  • fDate
    20-22 Sept. 2013
  • Firstpage
    129
  • Lastpage
    132
  • Abstract
    Urdu is a combination of several languages like Arabic, Hindi, English, Turkish, Sanskrit etc. It has a complex and rich morphology. This is the reason why not much work has been done in Urdu language processing. Stemming is used to convert a word into its respective root form. In stemming, we separate the suffix and prefix from the word. It is useful in search engines, natural language processing and word processing, spell checkers, word parsing, word frequency and count studies. This paper presents a rule based stemmer for Urdu. The stemmer that we have discussed here is used in information retrieval. We have also evaluated our results by verifying it with a human expert.
  • Keywords
    information retrieval; knowledge based systems; natural language processing; Urdu language processing; count studies; information retrieval; natural language processing; prefix; rule based stemmer; search engines; spell checkers; stemming; suffix; word frequency; word parsing; word processing; Accuracy; Communications technology; Computers; Conferences; Educational institutions; Morphology; Complex and Rich morphology; Rule Based Stemmer; Stemming; Urdu;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computer and Communication Technology (ICCCT), 2013 4th International Conference on
  • Conference_Location
    Allahabad
  • Print_ISBN
    978-1-4799-1569-9
  • Type

    conf

  • DOI
    10.1109/ICCCT.2013.6749615
  • Filename
    6749615