Abstract
Because of the trend of globalization, organizations and individuals often generate, acquire, and then archive documents written in different languages (i.e., poly-lingual documents). If organizations or individuals have already organized poly-lingual documents into their categories and would like to use this set of preclassified poly-lingual documents as training documents for constructing text categorization models that can classify newly arrived poly-lingual documents into appropriate categories, the organizations and individuals face the poly-lingual text categorization (PLTC) problem. Poly-lingual text categorization (PLTC) refers to the automatic learning of a text categorization model(s) from a set of preclassified training documents written in different languages and the subsequent assignment of unclassified poly-lingual documents to predefined categories on the basis of the induced text categorization model(s).Many text categorization techniques have been proposed in the literature; however, most of them deal with monolingual documents. In this study, we propose a cost-sensitive poly-lingual text categorization (CS-PLTC) technique that involves inclusion of translated documents to expand the training size for PLTC and use of cost-sensitive learning to reflect different qualities of training documents. Using the existing feature-reinforcement-based PLTC (FR-PLTC) technique as performance benchmarks, our empirical evaluation results show that our proposed CS-PLTC technique outperforms than the benchmark technique in both English and Chinese corpora. Keywords: Text mining, Text categorization, Poly-lingual text categorization, Document translation, Cost-sensitive learning, statistical-based bilingual thesaurus