Abstract
Research on creating friendly human interfaces between a human and a computer, or between human in distant locations flourished lately, partly because of the advances in computer, multimedia and internet technologies. One such style involves using an avatar, others might use a synthetic animated human face to provide an effective and efficient “face-to-face” multimodal communication channel in the distributed collaboration environments. Most of them adopt a real-time speech-driven face animation technique to avoid the need of directly transmitting the much larger-sized video data in order to meet the real time interpersonal communication requirement. The audio-to-visual conversion plays an important role in such real-time speech-driven face animation systems.In this thesis, we focus on the study of deriving the lip movement of a human user from its corresponding speech signal used in a speech-driven facial expression animation system. Two methods are proposed to design an audio to visual system for single user case, namely, the GRBF and ART2. The GRBF can improve the efficiency of VQ based audio to visual conversion system. Adaptive resonant theory 2 (ART 2) is extended from previous ART model with the capability of handling real input signals. The ART works like human’s memory. It has the ability to learn new thing fast without forgetting things learned in the past. An improvement in learning speed of ART2 over GMM with comparable error rate is observed in our experiments. However, the size of model parameter of ART2 is its disadvantage. In the multi-users case, a framework utilizing a reference ART2 audio-to-visual conversion model and an audio-adapting and visual learning mechanism is proposed to handle multi-user adaptation. Since the reference ART2 model is used for every user, only the incremental differences between the new user and the reference model need to be transmitted, the size disadvantage of ART2 is partially overcome in the multi-user adaptation case. Experiments supported the suitability of ART2 to the audio-visual conversion problem.