Abstract
We introduce a method for learning to estimate the similarity between collocations. In our approach, collocate pairs under certain headword are transformed into thesaurus-based and distributional similarity features from multiple sources. The method involves automatically generating similarity features, including WordNet-based features, n-gram based features, translation based features and headword-sensitive features for predicting the similarity between given collocate pairs. We present a similarity estimation prototype, ColloSim, that applies the method to collocations from a collocation dictionary. Evaluation on a set of collocates and their headword show that the method achieve reasonable good performance comparable to state-of-the-arts. Our method supports estimating the similarity of headword-sensitive collocations, resulting in additional improvement of the accuracy in predicting semantic similarity between collocate pairs.