Title
Topic-Based Coherence Modeling for Statistical Machine Translation
Abstract
Coherence that ties sentences of a text into a meaningfully connected structure is of great importance to text generation and translation. In this paper, we propose topic-based coherence models to produce coherence for document translation, in terms of the continuity of sentence topics in a text. We automatically extract a coherence chain for each source text to be translated. Based on the extracted source coherence chain, we adopt a maximum entropy classifier to predict the target coherence chain that defines a linear topic structure for the target document. We build two topic-based coherence models on the predicted target coherence chain: 1) a word level coherence model that helps the decoder select coherent word translations and 2) a phrase level coherence model that guides the decoder to select coherent phrase translations. We integrate the two models into a state-of-the-art phrase-based machine translation system. Experiments on large-scale training data show that our coherence models achieve substantial improvements over both the baseline and models that are built on either document topics or sentence topics obtained under the assumption of direct topic correspondence between the source and target side. Additionally, further evaluations on translation outputs suggest that target translations generated by our coherence models are more coherent and similar to reference translations than those generated by the baseline.
Year
DOI
Venue
2015
10.1109/TASLP.2015.2395254
IEEE/ACM Transactions on Audio, Speech & Language Processing
Keywords
DocType
Volume
text generation,maximum entropy classifier,linear topic structure,large-scale training data,text translation,coherent word translations,topic-based coherence modeling,statistical analysis,target coherence chain,source text,pattern classification,statistical machine translation,extracted source coherence chain,coherent phrase translations,phrase-based machine translation system,topic modeling,coherence chain,language translation,text coherence,word level coherence model,phrase level coherence model,natural language processing,statistical machine translation (smt),sentence topics continuity,text analysis,document translation,predictive models,data models,training data,nist,coherence,hidden markov models,natural,decoding
Journal
23
Issue
ISSN
Citations 
3
2329-9290
9
PageRank 
References 
Authors
0.45
28
3
Name
Order
Citations
PageRank
Deyi Xiong184567.74
Min Zhang21849157.00
Xing Wang35810.07