Learning Deep Structure-Preserving Image-Text Embeddings

2015-11-19CVPR 2016Unverified0· sign in to hype

Liwei Wang, Yin Li, Svetlana Lazebnik

Unverified — Be the first to reproduce this paper.

Abstract

This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a large margin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and text-to-image retrieval. Our method achieves new state-of-the-art results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset.

Tasks

Image Retrieval Image to text Metric Learning Phrase Grounding Retrieval Sentence

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
Flickr30K 1K test	SPE	R@1	29.7	—	Unverified

Learning Deep Structure-Preserving Image-Text Embeddings

Abstract

Tasks

Benchmark Results

Reproductions