Abstract
Human communication and interaction is a multimodal structure through voice, facial expression, body action, which are able to convey the people’s affect. In different moment of interaction, people feel different emotion. The expressions of emotion c be annotated by global emotion and continuous emotion. Continuous emotion is very complex and informative, but continuous and global emotion are both important. Therefore, we perform time encoding framework at behavior signal, body language and audio, to track continuous emotion, and focus on three dimensions of emotion which are activation, valence and dominance. During the training and testing process, we used two machine learning algorithms, support vector regression and sequence to sequence learning. Moreover, we extend our framework by two direction. First, we concentrate on continuous emotion tracking. Secondly, we use continuous emotion to predict the global affect. Comparing the previous study, continuous emotion tracking results achieve better performance of our system, and it also effectively help the global affect recognition. Interestingly, the continuous emotion tracking results bring the human perception mechanism of thin-slicing and further improve the global affect correlation.