Logo image
Video captioning via sentence augmentation and spatio-temporal attention
Conference paper   Peer reviewed

Video captioning via sentence augmentation and spatio-temporal attention

Tseng-Hung Chen, Kuo-Hao Zeng, Wan-Ting Hsu and Min Sun
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Vol.10116 LNCS, pp.269-286
2017

Abstract

Theoretical Computer Science Computer Science (all)
Generating video descriptions has many important applications such as human-robot interaction, video indexing, video sum-marization and assisting for the visually impaired. Many significant breakthroughs in deep learning and releases of large-scale open-domain video description datasets allow us to explore this task more effectively. Recently, Venugopalan et al. (S2VT) propose to caption a video via the technique on machine translation. We propose tracklet attention method to capture spatio-temporal information in the decoding phase and reserve the encoding phase similar to S2VT to retain the technique on machine translation. On the other hand, labels for video captioning are expensive and scarce, and training corpus is hard to completely cover rare words presenting in testing set. Hence, we propose to use sentence augmentation method to enrich our training corpus. Finally, we conduct experiments to demonstrate that tracklet attention and sentence augmentation improve the performance of S2VT on the validation set of Microsoft Research Video to Text dataset (MSR-VTT). In addition, we also achieve the state-of-the-art performance on Video Titles in the Wild dataset (VTW) by applying tracklet attention.

Metrics

1 Record Views

Details

Logo image