Logo image
A Hybrid Learning-based Method for Estimating Word Similarity in Collocation Clustering
Thesis

A Hybrid Learning-based Method for Estimating Word Similarity in Collocation Clustering

To, Wei-Jen
Masters, 國立清華大學, 資訊工程學系所
2016

Abstract

搭配詞相似度 相似字擷取 機器學習 Collocation Similarity Synonym Retrieval Machine Learning
We introduce a method for learning to estimate the similarity between collocations. In our approach, collocate pairs under certain headword are transformed into thesaurus-based and distributional similarity features from multiple sources. The method involves automatically generating similarity features, including WordNet-based features, n-gram based features, translation based features and headword-sensitive features for predicting the similarity between given collocate pairs. We present a similarity estimation prototype, ColloSim, that applies the method to collocations from a collocation dictionary. Evaluation on a set of collocates and their headword show that the method achieve reasonable good performance comparable to state-of-the-arts. Our method supports estimating the similarity of headword-sensitive collocations, resulting in additional improvement of the accuracy in predicting semantic similarity between collocate pairs.

Metrics

1 Record Views

Details

Logo image