Logo image
Unsupervised corpus distillation for represented indicator measurement on focus species detection
期刊文章   同儕審查

Unsupervised corpus distillation for represented indicator measurement on focus species detection

Chih-Hsuan WeiHung-Yu Kao
International Journal of Data Mining and Bioinformatics, 卷.8(4), 頁碼.413-426
2013
PMID: 24400519

摘要

Document classification Focus species identification Represented indicator measurement Information Systems Biochemistry Genetics and Molecular Biology (all) Library and Information Sciences
The gene ambiguity with the highest dimension is the species with which an entity is associated in biomedical text mining. Furthermore, one of the bottlenecks in gene normalisation is focus species detection. This study presents a method which is robust for all types of articles, particularly those without explicit species mentions. Since our method requires a training corpus, we developed an iterative distillation method to extend the corpus. Unsupervised corpus is therefore helpful for the detection of focus species. In experiments, the proposed method achieved a high accuracy of 85.64% and 84.32% in datasets with and without species mentions respectively. Copyright © 2013 Inderscience Enterprises Ltd.

相關連結

指標

1 檢視次數

詳細資料

Logo image