Logo image
AUXILIARY BI-LEVEL GRAPH REPRESENTATION FOR CROSS-MODAL IMAGE-TEXT RETRIEVAL
Conference paper

AUXILIARY BI-LEVEL GRAPH REPRESENTATION FOR CROSS-MODAL IMAGE-TEXT RETRIEVAL

Xian Zhong, Zhengwei Yang, Mang Ye, Wenxin Huang, Jingling Yuan and Chia-Wen Lin
Proceedings - IEEE International Conference on Multimedia and Expo
2021

Abstract

Cross Modal Graph Convolution Image-text Retrieval Scene Graph Computer Networks and Communications Computer Science Applications
Image-text retrieval is one of the most common tasks in multi-modal retrieval. It suffers from the problem of information imbalance between modalities, which is so-called modality gap. It remains challenging because prior methods cannot bridge the gap reasonably. With the help of scene graph, we start by designing an auxiliary bi-level graph representation (ABGR) pipeline that can fully mine the potential information and reduce the information redundancy. By doing so, each modality will be represented by lexical word graph that carries the main content of the information. Specifically, we design a graph feature enhancement (GFE) module to embed the graph-structured information in a common subspace while exploring the relationship between lexical words. As a result, a better representation for both image and text can be obtained, which helps us to evaluate the similarity between images and texts more reasonably. Experimental results conducted on two benchmark datasets Flickr30K and MS-COCO demonstrate the effectiveness of our proposed model for cross-modal retrieval task.

Metrics

1 Record Views

Details

Logo image