Logo image
Represented indicator measurement and corpus distillation on focus species detection
Conference paper

Represented indicator measurement and corpus distillation on focus species detection

Chih-Hsuan Wei and Hung-Yu Kao
Proceedings - 2010 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2010, pp.657-662
2010

Abstract

Biomedical Engineering Health Informatics
In extraction of information from the biomedical literature, name disambiguation of domain-specific entities, such as proteins, is one of the most important issues. The entity ambiguity with the highest dimension is the species to which an entity is associated with. Furthermore, one of the bottlenecks in inter-species gene name normalization is species disambiguation. To enhance the performance of species disambiguation, the detection of focus species detection remains a substantial challenge. This study presents a method addressing this issue. The results present evaluations of all articles from the BioCreaTive I&II GN task. Our method is robust for all types of articles, particularly those without explicit species entity information. Since our method requires a training corpus to be the indicator vector, we developed an iterative corpus distillation method to extend the corpus. In the conducted experiments, the proposed method achieved a high accuracy of 85.64% and 84.32% without species entity information. ©2010 IEEE.

Metrics

1 Record Views

Details

Logo image