Logo image
Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
會議論文

Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Jing-Tong Tzeng, Su Bo-Hao, Wu Ya-Tse, Hsing-Hang Chou 和 Chi-Chun Lee
arXiv.org
Cornell University Library, arXiv.org
Interspeech 2025 (Rotterdam, The Netherlands, 17/08/2025–21/08/2025)
25/09/2025

摘要

Emotion recognition Speech recognition Emotions Machine Learning
In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.

相關連結

指標

1 檢視次數

詳細資料

Logo image