Logo image
Speech representation learning for emotion recognition using end-to-end ASR with factorized adaptation
Conference paper

Speech representation learning for emotion recognition using end-to-end ASR with factorized adaptation

Sung-Lin Yeh, Yun-Shao Lin and Chi-Chun Lee
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, Vol.2020-October, pp.536-540
2020

Abstract

Acoustic representation Domain adaptation End-to-end ASR Speech emotion recognition Language and Linguistics Human-Computer Interaction Signal Processing Software Modeling and Simulation
Developing robust speech emotion recognition (SER) systems is challenging due to small-scale of existing emotional speech datasets. However, previous works have mostly relied on handcrafted acoustic features to build SER models that are difficult to handle a wide range of acoustic variations. One way to alleviate this problem is by using speech representations learned from deep end-to-end models trained on large-scale speech database. Specifically, in this paper, we leverage an end-to-end ASR to extract ASR-based representations for speech emotion recognition. We further devise a factorized domain adaptation approach on the pre-trained ASR model to improve both the speech recognition rate and the emotion recognition accuracy on the target emotion corpus, and we also provide an analysis in the effectiveness of representations extracted from different ASR layers. Our experiments demonstrate the importance of ASR adaptation and layer depth for emotion recognition.

Metrics

1 Record Views

Details

Logo image