Logo image
An Inter-Speaker Fairness-Aware Speech Emotion Regression Framework
Conference paper

An Inter-Speaker Fairness-Aware Speech Emotion Regression Framework

Hsing-Hang Chou, Woan-Shiuan Chien, Ya-Tse Wu and Chi-Chun Lee
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp.3190-3194
2024

Abstract

fairness privacy speech emotion recognition Language and Linguistics Human-Computer Interaction Signal Processing Software Modeling and Simulation
Speech emotion recognition (SER) helps to achieve better human-to-machine interactions in voice technologies. Recent studies have pointed out critical fairness issues in the SER. While there are efforts in building fair SER, most of the works focus on fairness between demographic groups and rely on these broad categorical attributes to build a fair SER. In this paper, we instead focus on the fairness learning among individual speakers, which is rarely discussed yet much more intuitively appealing in constructing a fair SER model. To reduce the reliance on knowing speaker IDs, we perform unsupervised clustering on the utterance embeddings from a pre-trained speaker verification model that puts utterances with different characteristics into clusters that roughly represent the true speaker index. Our evaluation demonstrates that with these cluster IDs, we can construct a fairness-aware SER model at an individual speakerlevel without knowing speaker IDs upfront.

Metrics

1 Record Views

Details

Logo image