Logo image
Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition Systems
Conference paper

Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition Systems

Haibin Wu, Huang-Cheng Chou, Kai-Wei Chang, Lucas Goncalves, Jiawei Du, Jyh-Shing Roger Jang, Chi-Chun Lee and Hung-Yi Lee
Proceedings of 2024 IEEE Spoken Language Technology Workshop, SLT 2024, pp.510-517
2024

Abstract

Ambiguity of emotion Multi-label classification Reproducibility Speech emotion recognition Subjectivity of emotion perception Computer Vision and Pattern Recognition Hardware and Architecture Media Technology Instrumentation Linguistics and Language
Speech emotion recognition (SER) is an essential technology for human-computer interaction systems. However, the previous study reveals that 80.77% of SER papers yield results that cannot be reproduced on the well-known IEMOCAP dataset. The main reason for reproducibility challenges is that the database did not provide standard data splits (e.g., train, development, and test sets). Prior papers could define its partition, but they did not provide details of the partition or source code for processing the partition. Therefore, this work aims to make SER open and reproducible to everyone. We develop the EMO-SUPERB, shorted for EMOtion Speech Universal PERformance Benchmark, including a user-friendly codebase to leverage 16 state-of-the-art (SOTA) speech self-supervised learning models for exhaustive evaluation plus one SOTA SER model across 6 open-source SER datasets in English and Chinese. We make all resources open-source to facilitate future developments in SER. Researchers can easily upload their systems or datasets to EMO-SUPERB, and we name the project 'Open-Emotion'.

Metrics

1 Record Views

Details

Logo image