Abstract
Written in LangGeh orthography, the alignment of parallel sentences in Taiwanese and in Mandarin has been studied (Lin 2009). By substituting a few common words in Taiwanese with their counterparts in Mandarin, the LCS (longest common subsequence) algorithm is able to give about 70% recall rate while keeps those aligned highly correct (it actually was perfectly correct in the experiment). This thesis continues the study on alignment by constructing sausage nets from Taiwanese sentences and from Mandarin sentences using various parallel dictionaries, and then applying the LCS algorithm. The sausage net approach gives in 85%~90% recall rates on various corpora while still retaining nearly perfect correctness for those marked aligned.