Abstract
Avatars are usually employed in virtual conferencing system as user’s agent for face-to-face talking with each other in a 3-D virtual space. In previous research, system architecture is established by integrating a 3-D model-based coder and a 2-D video coder under very low bit-rate video coding. It makes the avatars to have more realistic expression as users’ representatives to form the video realistic avatars.In this thesis, we develop a new error concealment scheme for the 2-D video coder as the enhancement layer, which is expected to be more suitable in the virtual conferencing system. With the proposed scheme, we can alleviate the defect caused by the loss of the enhancement layer and reconstruct a vivid 3D virtual facial model while inefficient bandwidth is available.First, we extract the facial animation parameters (FAP) and texture difference, which is the difference between expressive facial image and neutral facial image synthesized by 3-D model-based coding, from the input sequence. Then, we define the region of interest (ROI) by using facial animation tables (FAT) to separate the texture differences of the whole face into several partial texture differences in each ROI. Besides, each ROI are affected by specific FAP groups, that is the partial texture difference in ROI are also affected.Thus, we propose a VQ-like algorithm to cluster the FAPs and its partial texture differences of each ROI according to the minimum quantization error, which is determined by the differences between synthetic facial image with partial updated by the partial texture difference and the real facial image. With the codevectors of each cluster, we can predict the partial texture difference by the nearest one to the received FAPs. Finally, a predicted texture map is generated by partial texture difference, and then we can conceal the enhancement layer by using these predicted texture map.