• Title of article

    A Hybrid Approach for Urdu Sentence Boundary Disambiguation

  • Author/Authors

    Rehman, Zobia COMSATS Institute of Information Technology - Department of Computer Science, Pakistan , Anwar, Waqas COMSATS Institute of Information Technology - Department of Computer Science, Pakistan

  • From page
    250
  • To page
    255
  • Abstract
    Sentence boundary identification is a preliminary step for preparing a text document for Natural Language Processing tasks, e.g., machine translation, POS tagging, text summarization and etc. We present a hybrid approach for Urdu sentence boundary disambiguation comprising of unigram statistical model and rule based algorithm. After implementing this approach, we obtained 99.48% precision, 86.35% recall and 92.45% F1-Measure while keeping training and testing data different from each other, and with same training and testing data, we obtained 99.36% precision, 96.45% recall and 97.89% F1-Measure.
  • Keywords
    Sentence boundary disambiguation , and unigram model.
  • Journal title
    The International Arab Journal of Information Technology (IAJIT)
  • Journal title
    The International Arab Journal of Information Technology (IAJIT)
  • Record number

    2543836