Logo image
Using Punctuation Marks for Bilingual Sentence Alignment
Thesis

Using Punctuation Marks for Bilingual Sentence Alignment

Kevin Chih-Cheng Yeh
Masters, 國立清華大學, 資訊工程學系
2002

Abstract

句子對應 標點符號 機器翻譯 雙語平行語料 機率式模型 同源詞 Sentence Alignment Punctuation Marks Machine Translation Bilingual Parallel Corpus Probabilistic Model Cognate
In this paper, we present a new approach to aligning English and Chinese sentences in parallel corpora based solely on punctuation marks. Although the length-based approach produces high accuracy rates of sentence alignment for clean parallel corpora written in two Western languages such as French and English or German and English, it is not fair as well for parallel corpora written in two disparate languages such as Chinese and English. It is possible to use cognates on top of length-based approach to increase alignment accuracy. However, cognates do not exist between two disparate languages, therefore limiting the applicability of cognate-based approach. In this paper, we examine the feasibility of exploiting soft, ordered matching punctuation marks in two languages for high accuracy sentence alignment. We experimented with an implementation of the proposed method on the parallel corpus of Chinese-English Sinorama Magazine Corpus and Scientific American Magazine Corpus with satisfactory results. We have carried out experiments on sentence alignment using our method and comparing with the length-based method. We evaluated the results based on precision and recall rates with good results. We also demonstrated that the method is applicable to other language pairs such as English and Japanese with minimal additional effort.

Metrics

1 Record Views

Details

Logo image