Logo image
Phrase Correspondence Extraction in Bilingual Corpora Based on Phrase Translation Model
Thesis

Phrase Correspondence Extraction in Bilingual Corpora Based on Phrase Translation Model

FAN, WEI-GENG
Masters, 國立清華大學, 資訊系統與應用研究所
2004

Abstract

機器翻譯 片語翻譯 machine translation phrase translation
We introduce a method to extract the phrase correspondence from a given bilingual sentence in a bilingual corpus by using a phrase-based translation model. In our approach, the bilingual source and target sentences are transformed into a set of phrase pairs with the purpose of selecting the phrase correspondences with maximum phrase translation probability, which include lexical translation probability (LTP) and phrase alignment probability (PAP). The method involves estimating PAP from phrases hand tagged with phrase alignment information, learning LTP from PAP and a bilingual phrase lexicon, and iteratively co-training LTP and PAP. At run time, we generate a set of candidate translations from the target sentence for each phrase in the source sentence, calculate probabilities of candidates by using the phrase-based translation model, and determine the candidate with maximum probability as the correspondence. We implemented the proposed method by using bilingual phrase lexicon of National Institute for Compilation and Translation and the parallel corpus of Sinorama magazine. Comparative evaluation on phrases in a set of bilingual sentences randomly chosen from the parallel corpus shows that our model outperforms the IBM Model 4. Experimental results prove that the proposed method could efficiently improve the performance of extracting phrase correspondences in bilingual corpora and further provide valuable phrase-level translation knowledge for work on machine translation and computer assisted translation systems.

Metrics

1 Record Views

Details

Logo image