Logo image
Attention-guided image captioning with adaptive global and local feature fusion
期刊文章   同儕審查

Attention-guided image captioning with adaptive global and local feature fusion

Xian Zhong, Guozhang Nie, Wenxin Huang, Wenxuan Liu, Bo MaChia-Wen Lin
Journal of Visual Communication and Image Representation, 卷.78, 103138
07/2021

摘要

Adaptive attention Encoder-decoder Image captioning Spatial information Signal Processing Media Technology Computer Vision and Pattern Recognition Electrical and Electronic Engineering
Although attention mechanisms are exploited widely in encoder-decoder neural network-based image captioning framework, the relation between the selection of salient image regions and the supervision of spatial information on local and global representation learning was overlooked, thereby degrading captioning performance. Consequently, we propose an image captioning scheme based on adaptive spatial information attention (ASIA), extracting a sequence of spatial information of salient objects in a local image region or an entire image. Specifically, in the encoding stage, we extract the object-level visual features of salient objects and their spatial bounding-box. We obtain the global feature maps of an entire image, which are fused with local features and the fused features are fed into the LSTM-based language decoder. In the decoding stage, our adaptive attention mechanism dynamically selects the corresponding image regions specified by an image description. Extensive experiments conducted on two datasets demonstrate the effectiveness of the proposed method.

相關連結

指標

1 檢視次數

詳細資料

Logo image