Title :
Towards integrated machine translation using structural alignment from syntax-augmented synchronous parsing
Author :
Bing Xiang;Bowen Zhou;Martin ?mejrek
Author_Institution :
IBM T. J. Watson Research Center, Yorktown Heights, NY 10598, USA
Abstract :
In current statistical machine translation, IBM model based word alignment is widely used as a starting point to build phrase-based machine translation systems. However, such alignment model is separated from the rest of machine translation pipeline and optimized independently. Furthermore, structural information is not taken into account in the alignment model, which sometimes leads to incorrect alignments. In this paper, we present a novel method to connect a re-alignment model with a translation model in an integrated framework. We conduct bilingual chart parsing based on syntax-augmented synchronous context-free grammar. A Viterbi derivation tree is generated for each sentence pair with multiple features employed in a log-linear model. A new word alignment is created under the structural constraint from the Viterbi tree. Extensive experiments are conducted in a Farsi-to-English translation task in conversational speech domain and also a German-to-English translation task in text domain. Systems trained on the new alignment provide significant higher BLEU scores compared to a state-of-the-art baseline.
Keywords :
"Viterbi algorithm","Hidden Markov models","Context modeling","Pipelines","Surface-mount technology","Natural languages","Solids","Joining processes","Speech analysis"
Conference_Titel :
Automatic Speech Recognition & Understanding, 2009. ASRU 2009. IEEE Workshop on
Print_ISBN :
978-1-4244-5478-5
DOI :
10.1109/ASRU.2009.5372892