Abstract
“LangGeh” is a new orthography for languages using Chinese character such as Taiwanese or Mandarin. Similar to word separation in English orthography, LangGeh proposes simple phrase separation. Based on LangGeh, We build “LangGeh 09 parallel alignment corpus” and extract a Taiwanese-Mandarin “phrase” dictionary from the parallel corpus. We compare two methods for the extraction of the bilingual collocation dictionary. The first method uses a criterion based on high association, while the second method is based on alignment of the sentences in the parallel corpus. Compared to other language pair such as English and French, there are at least two common characteristics between Taiwanese and Mandarin: they have many common phrases, and their word orders are similar. This paper demonstrates how the common characteristics can be utilized.