Logo image
Generating Cross-domain Visual Description via Adversarial Learning
Thesis

Generating Cross-domain Visual Description via Adversarial Learning

Chen, Tseng-Hung
Masters, 國立清華大學, 電機工程學系所
2016

Abstract

深度學習 圖像字幕生成 遷移學習 電腦視覺 對抗式訓練 增強學習 Deep Learning Image Captioning Transfer Learning Computer Vision Adversarial Training Reinforcement Learning
Impressive image captioning results are achieved in domains with plenty of training image and sentence pairs (e.g., MSCOCO). However, transferring to a target domain with significant domain shifts but no paired training data (referred to as cross-domain image captioning) remains largely unexplored. We propose a novel adversarial training procedure to leverage unpaired data in the target domain. Two critic networks are introduced to guide the captioner, namely domain critic and multi-modal critic. The domain critic assesses whether the generated sentences are indistinguishable from sentences in the target domain. The multi-modal critic assesses whether an image and its generated sentence are a valid pair. During training, the critics and captioner act as adversaries -- captioner aims to generate indistinguishable sentences, whereas critics aim at distinguishing them. The assessment improves the captioner through policy gradient updates. During inference, we further propose a novel critic-based planning method to select high-quality sentences without additional supervision (e.g., tags). To evaluate, we use MSCOCO as the source domain and four other datasets (CUB-200-2011, Oxford-102, TGIF, and Flickr30k) as the target domains. Our method consistently performs well on all datasets. Utilizing the learned critic during inference further boosts the overall performance in CUB-200 and Oxford-102. Furthermore, we extend our method to the task of video captioning. We observe improvements for the adaptation between large-scale video captioning datasets such as MSR-VTT, M-VAD and MPII-MD.

Metrics

1 Record Views

Details

Logo image