Abstract
In this paper, we propose a new method for extracting bilingual collocation from a parallel corpus. The method integrates statistical and linguistic information for effective extraction of bilingual collocations. The method involves first obtaining an extended list of distinct English collocations from a very large monolingual corpus, identifying the collocation instances in a parallel corpus, and extracting translation equivalent of the collocations based on word alignment information. At run time, collocations in the parallel corpus are identified and aligned to the translation equivalent. Experimental results show our method is efficient for learning translation memory of collocations. We applied the method to develop a collocational concordancer, TANGO, which showed great potential for applications in Computer Assisted Language Leaning and Computer Assisted Translation.