Audio-visual emotion recognition (AVER) often performs well under ideal conditions but faces significant challenges in scenarios with missing modalities (e.g., missing frames of audio and/or video). Addressing these challenges is crucial for the effective deployment of AVER systems in human-computer interaction (HCI) applications, where robustness can significantly impact user experience. This study introduces a novel approach that enhances AVER robustness by leveraging a decoder-like summarizer structure. This structure processes audio and visual content and generates contextual summaries that effectively capture emotional cues even when modalities are degraded. To enhance system resilience against missing modalities, we integrate modality dropout during training, enabling the summarizer to adaptively handle these scenarios. We define the context summary length as the number of learnable query tokens used in the summarizer, a fixed hyperparameter in our model. We analyze how varying context summary lengths affect performance, identifying an optimal balance between compression and expressiveness. In addition to improving robustness, we systematically evaluate model calibration across emotions in current state-of-the-art (SOTA) AVER methods. Our experiments on the MSP-IMPROV and CREMA-D databases demonstrate that our model achieves superior performance across macro-, micro-, and weighted-F1 scores, both under ideal conditions and in scenarios with modality losses. Additionally, we conduct ablation studies to assess the impact of different context lengths on our summarizer structure in terms of overall AVER performance.
- Contextual Attention for Robust Audio-Visual Emotion Recognition
- Lucas Goncalves - University of DallasHuang-Cheng Chou - National Tsing Hua UniversityAli N. Salman - The University of Texas at DallasChi-Chun Lee - National Tsing Hua University, College of Semiconductor ResearchCarlos Busso - Carnegie Mellon University
- IEEE
- 12
- CNS-2016719 / National Science Foundation
- Journal article
- 2026
- IEEE open journal of signal processing, Vol.7, pp.42-53
- English