Logo image
Learning to Find Translations and Transliterations on the Web based on Conditional Random Fields
Thesis

Learning to Find Translations and Transliterations on the Web based on Conditional Random Fields

Chang, Chee
Masters, 國立清華大學, 資訊系統與應用研究所
2012

Abstract

機器翻譯 跨語言資訊擷取 維基百科 conditional random fields
In recent years, state-of-the-arts cross-linguistic systems have been based on parallel cor- pora. However, it is difficult at times to find translations of a certain technical term or named entity even with a very large parallel corpus. In this paper, we present a new method for learning to find translations on the Web for a given term. In our approach, we use a small set of terms and translations to obtain mixed-code snippets returned by a search engine. We then automatically annotate the data with translation tags, automati- cally generate features to augment the tagged data, and automatically train a conditional random fields model for identifying translations. At runtime, we obtain mixed-code web- pages containing the given term, and run the model to extract translations as output. Pre- liminary experiments and evaluation results show our method cleanly combines various features, resulting in a system which outperforms previous work.

Metrics

1 Record Views

Details

Logo image