Title :
The refined MI: A significant improvement to mutual information
Author :
Alrabiah, Maha ; Al-Salman, AbdulMalik ; Atwell, Eric
Author_Institution :
Comput. Sci. Dept., King Saud Univ., Riyadh, Saudi Arabia
Abstract :
Distributional lexical semantics is an empirical approach that is mainly concerned with modeling words´ meanings using word distribution statistics gathered from very large corpora. It is basically built on the Distributional Hypothesis by Zellig Harris in 1970, which states that the difference in words´ meanings is associated with the difference in their distribution in text. This difference in meaning originates from two kinds of relations between words, which are syntagmatic and paradigmatic relations. Syntagmatic relations are linear combinatorial relations that are established between words that co-occur together in sequential text; while paradigmatic relations are substitutional relations that are established between words that occur in the same context, share neighboring words, but do not co-occur in the same text. In this paper, we present a new association measure, the Refined MI, for measuring syntagmatic relations between words. In addition, an experimental study to evaluate the performance of the proposed measure is presented. The measure showed outstanding results in identifying significant co-occurrences from Classical Arabic text.
Keywords :
natural language processing; text analysis; Arabic text; distributional lexical semantics; mutual information; paradigmatic relations; refined MI; syntagmatic relations; word distribution statistics; Context; Educational institutions; Frequency conversion; Frequency measurement; Mutual information; Pragmatics; Semantics; Classical Arabic; Mutual Information; association measures; distributional semantics;
Conference_Titel :
Asian Language Processing (IALP), 2014 International Conference on
Conference_Location :
Kuching
DOI :
10.1109/IALP.2014.6973512