Abstract
In this paper, we present a new approach to aligning English and Chinese sentences in parallel corpora based solely on punctuation marks. Although the length-based approach produces high accuracy rates of sentence alignment for clean parallel corpora written in two Western languages such as French and English or German and English, it is not fair as well for parallel corpora written in two disparate languages such as Chinese and English. It is possible to use cognates on top of length-based approach to increase alignment accuracy. However, cognates do not exist between two disparate languages, therefore limiting the applicability of cognate-based approach. In this paper, we examine the feasibility of exploiting soft, ordered matching punctuation marks in two languages for high accuracy sentence alignment. We experimented with an implementation of the proposed method on the parallel corpus of Chinese-English Sinorama Magazine Corpus and Scientific American Magazine Corpus with satisfactory results. We have carried out experiments on sentence alignment using our method and comparing with the length-based method. We evaluated the results based on precision and recall rates with good results. We also demonstrated that the method is applicable to other language pairs such as English and Japanese with minimal additional effort.