Abstract
“LangGeh” is a new orthography for languages using Chinese character such as Taiwanese or Mandarin. Similar to word separation in English orthography, LangGeh proposes simple phrase separation. Based on LangGeh, We build a Taiwanese-Mandarin parallel corpus and use it to study the translation between Taiwanese and Mandarin using the statistical machine translation framework of Brown et. al. (1990, 1993). There are at least two common characteristics between Taiwanese and Mandarin that one can utilize in translation: many common phrases and word orders are similar. We simplify the translation framework using the concept of “Sausage Phrase”. It has the advantage of being conceptual simple and easy to calculate.