DocumentCode
3659966
Title
Capturing the Common Syntactical Rules for the Holy Quran: A Data Mining Approach
Author
Mahmoud Hasan Alsaheb;Dia Eddin Mohammad Asad AbuZeina
Author_Institution
Coll. of IT &
fYear
2013
Firstpage
670
Lastpage
680
Abstract
This paper presents a novel approach to capture the common syntactical rules for the Holy Quran . By syntactical rules, we mean the common relationships between the words´ tags that highly show up in the Quran. Arabic, like other language, has a number of tags which include nouns, verbs, and pronouns with a number of sub-types of each one of them. In this paper we used data mining approach to extract the common syntactical rules which will be offered to the natural language processing applications. Stanford part of speech tagger (29 tags) will be used to tag the Quran words. Then, the data mining too called WEKA (PredictiveApriori algorithm) will be used to find the famous syntactical rules. The extracted syntactical rules have a property that it is not necessary to have adjacent words tags. That is, long distance relation. The most common syntactical rule found is: tag1=RP tag2=NN tag3=WP 91 ⇒ tag4=VBD 90 acc:(0.97912)Which can be seen in the Quran sentence. This phrase (which is part of an ayah) appeared in 89 ayahs in 20 different surahs; the study used Mushaf Al-Madinah Al-Munawwarah (published by the King Fahd Complex for Printing the Holy Quran ).
Keywords
"Data mining","Speech","Computers","Natural language processing","Prediction algorithms","Printing"
Publisher
ieee
Conference_Titel
Advances in Information Technology for the Holy Quran and Its Sciences (32519), 2013 Taibah University International Conference on
Type
conf
DOI
10.1109/NOORIC.2013.105
Filename
7277302
Link To Document