Logo image
EXPLOITING ANNOTATORS' TYPED DESCRIPTION OF EMOTION PERCEPTION TO MAXIMIZE UTILIZATION OF RATINGS FOR SPEECH EMOTION RECOGNITION
Conference paper

EXPLOITING ANNOTATORS' TYPED DESCRIPTION OF EMOTION PERCEPTION TO MAXIMIZE UTILIZATION OF RATINGS FOR SPEECH EMOTION RECOGNITION

Huang-Cheng Chou, Wei-Cheng Lin, Chi-Chun Lee and Carlos Busso
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, Vol.2022-May, pp.7717-7721
2022

Abstract

Distribution-label learning Emotion recognition Multi-label learning Soft-label learning Software Signal Processing Electrical and Electronic Engineering
The decision of ground truth for speech emotion recognition (SER) is still a critical issue in affective computing tasks. Previous studies on emotion recognition often rely on consensus labels after aggregating the classes selected by multiple annotators. It is common for a perceptual evaluation conducted to annotate emotional corpora to include the class “other,” allowing the annotators the opportunity to describe the emotion with their own words. This practice provides valuable emotional information, which, however, is ignored in most emotion recognition studies. This paper utilizes easy-accessed natural language processing toolkits to mine the sentiment of these typed descriptions, enriching and maximizing the information obtained from the annotators. The polarity information is combined with primary and secondary annotations provided by individual evaluators under a label distribution framework, creating a complete representation of the emotional content of the spoken sentences. Finally, we train multitask learning SER models with existing learning methods (soft-label, multi-label, and distribution-label) to show the performance of the novel ground truth in the MSP-Podcast corpus.

Metrics

1 Record Views

Details

Logo image