Logo image
局部最長連續共同子序列與新詞組收集
Thesis

局部最長連續共同子序列與新詞組收集

謝博行
Masters, 國立清華大學, 統計學研究所
2012

Abstract

未知詞 新詞組 局部最長共同子序列 Unknown word New phrase Locally longest common consecutive subsequence
Adapting from the well-known longest common subsequence (LCS) algorithm, we propose an efficient algorithm that is capable of extracting locally longest consecutive common subsequence (LLCCS) from one or two different articles. Further processing on the extracted subsequence makes them closer to syntatical phrases/words. With world wide web full of adundant articles, we hope this is an efficient way to enrich the entries of Chinese lexicon.

Metrics

1 Record Views

Details

Logo image