Abstract
Dysphonia may originate from various vocal cord disorders (VCDs), and determination of its etiology usually requires image-based examination. However, such examination is invasive and not available in rural areas. Therefore, we aim to evaluate whether the voice of VCD patients may provide cues that could be utilized for screening purposes. Presently, a voice dataset containing entries of vocal fold atrophy, paralysis, benign organic lesions, and laryngeal cancer was prepared and support vector machine and neural network models were trained for VCD classification. Features were first extracted from/a/, and performance on the present dataset was compared against the Saarbruecken voice database. Next, the utterances of counting from 1 to 5 were processed, and “4”, pronounced as/s1/in Mandarin, was found most suitable for classifying non-cancer VCDs. Finally, we demonstrated that inclusion of the subjective GRBAS scale consistently raised the classification accuracy.