Abstract
Visual content retrieval is an emerging technique to access multimedia information. However, this technique suffers from the semantic gap problem, which is the mismatch between low-level visual features and high-level human concepts. A way to narrow down the semantic gap is to incorporate multi-modal interactions into the retrieval technique. Multi-modal interactions provide users with a combination of multiple interacting modalities, such as relevance feedback, various input and output interfaces. Therefore users can select a convenient and appropriate mode to specify query conditions and observe search results during the retrieval process.In this study, we propose a multi-modal interaction framework in visual content retrieval. Several novel interacting modalities are presented, and associated indexing and matching algorithms are devised. Two types of multimedia information, namely, texture images and human motion, are investigated and demonstrated. In our texture image retrieval system, users can specify linguistic terms, together with visual examples to find desired texture images. Besides, an automatic annotation and indexing is devised based on these linguistic terms. Further, a personalization mechanism is developed to cope with the human subjectivity to linguistic terms. In our human motion retrieval system, users can choose appropriate input modes, including text, stickman, images, and motion clips, to specify their queries. Later, they can observe retrieval results via graphics images or animation video. Moreover, we design efficient and effective indexing and matching algorithms to decrease retrieval time and improve retrieval accuracy. In particular for the image input mode, we present a novel model-based approach to reconstruct a 3D human posture from a single image. Our experimental results reveal a promising direction of the proposed multi-modal interaction framework in visual content retrieval.