Markov random fields on graphs for natural languages

Author

O´Sullivan, Joseph A. ; Mark, Kevin ; Miller, Michael I.

Author_Institution

Dept. of Electr. Eng., Washington Univ., St. Louis, MO, USA

fYear

1994

fDate

27-29 Oct 1994

Firstpage

47

Abstract

The use of model-based methods for data compression for English dates back at least to Shannon´s Markov chain (n-gram) models, where the probability of the next word given all previous words equals the probability of the next word given the previous n-1 words. A second approach seeks to model the hierarchical nature of language via tree graph structures arising from a context-free language (CFL). Neither the n-gram nor the CFL models approach the data compression predicted by the entropy of English as estimated by Shannon and Cover and King. This paper presents two models that incorporate the benefits of both the n-gram model and the tree-based models. In either case the neighborhood structure on the syntactic variables is determined by the tree while the neighborhood structure of the words is determined by the n-gram and the parent syntactic variable (preterminal) in the tree, Having both types of neighbors for the words should yield decreased entropy of the model and hence fewer bits per word in data compression. To motivate estimation of model parameters, some results in estimating parameters for random branching processes is reviewed

Keywords

Markov processes; context-free languages; data compression; entropy; natural languages; probability; random processes; speech processing; trees (mathematics); English; Markov random fields; Shannon´s Markov chain models; context-free language; data compression; entropy; graphs; model-based methods; n-gram model; natural languages; neighborhood structure; parameter estimation; parent syntactic variable; preterminal; probability; random branching processes; syntactic variables; tree graph structures; tree-based models; words; Context modeling; Data compression; Eigenvalues and eigenfunctions; Entropy; Markov random fields; Natural languages; Parameter estimation; Predictive models; Stochastic processes; Tree graphs;

fLanguage

English

Publisher

ieee

Conference_Titel

Information Theory and Statistics, 1994. Proceedings., 1994 IEEE-IMS Workshop on

Conference_Location

Alexandria, VA

Print_ISBN

0-7803-2761-6

Type

conf

DOI

10.1109/WITS.1994.513880

Filename

513880