Abstract
Analysis and synthesis of facial expressions are two core technologies in construction of a virtual conferencing system for face-to-face communications. In this thesis, we explore the possible integration of researches from different fields, namely, computer vision and computer graphics, 2D video coding and 3D model-based coding, expression analysis and speech analysis, for providing better solutions for facial expression analysis and synthesis for use in virtual conferencing systems. With the integration of computer vision and computer graphics, the head modeling, head pose estimation, and facial expression analysis can be performed in an analysis-by-synthesis approach. Thus, the head modeling from monocular image sequences can be achieved with a more complete feature set composed of salient facial feature points and head profiles; the head pose estimation can be performed with high accuracy under various lighting conditions; and expression analysis can be operated under different head poses with the assistance of user-customized facial model. With the integration of 2D video coding and 3D model-based coding, expression synthesis can achieve better visual quality while merits of 3D model-based coding, the coding efficiency and 3D rendering from arbitrary viewpoints, are well preserved for use in a multiuser virtual conferencing system. Finally, the concept of integration of expression analysis and speech analysis is presented, which can lead to a lower complexity solution to expression analysis and a much lower bandwidth requirement solution for expressive expression synthesis. Prototypes of virtual conferencing applications for point-to-point and multipoint audio-visual communications are also realized to verify the usefulness of the proposed methods.