Abstract
Continuous emotion recognition aims to recognize human emotion from audio-visual sequences. Continuous emotion labels are noisy because the annotators cannot accurately estimate the continuous emotion while watching the audio-visual sequences in real-time. Most existing methods exclude the label noises and bridge the gap between real emotion and their model with manual and hand-crafted designs. However, these manual designs are not beneficial to the automation of emotion understanding. The purpose of this work is to purify the noisy emotion labels and also to bridge the gap between emotion annotators and emotion regressors automatically. We propose a jointly-optimized model of emotion regressor and common label bias estimation with a label noise measurement by feature-label relationships. We also empower this model with deep emotion features. The proposed method is capable of jointly emotion recognition, label purification, and label bias compensation under minimal human interventions. The results of our study outperform various models under fair comparisons, and are comparable to the State-of-the-Arts on the well-known AVEC 2012 dataset.