Speech Emotion Recognition (SER) systems often face challenges of real-world noises, which limits their robustness outside controlled environments. To tackle this, the recent bimodal approach fusing audio with textual inputs has been in the spotlight of researchers. However, prior work predominantly relied on transcripts or coarse scene descriptions, which offer limited semantic depth. In this work, we introduce reasoning-driven captions generated by Mellow, a small audio language model, for context-aware textual encoding of high-order semantic information. These reasoning-based captions capture contextualized cues beyond lexical transcripts, thus providing balanced emotional grounding. Our experiments demonstrate that reasoning-based captions consistently improve SER performance under noisy conditions, particularly at low signal-to-noise ratios, where conventional transcripts mainly benefit valence but compromise arousal and dominance. In contrast, our proposed reasoning-rich captions achieve robust and balanced prediction across arousal, valence, and dominance, setting a new direction for noise-resilient multimodal SER.
- Reasoning Driven Captions to Assist Noise Robust Speech Emotion Recognition
- Snehit B. Chunarkar (Author)Chi-Chun Lee (Author) - National Tsing Hua University, College of Semiconductor Research
- ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (Barcelona, Spain, 03/05/2026–08/05/2026)
- Conference paper
- English