General Image Descriptors for Open World Image Retrieval using ViT CLIP

2022-10-20Code Available1· sign in to hype

Marcos V. Conde, Ivan Aerlic, Simon Jégou

Code Available — Be the first to reproduce this paper.

Code

github.com/ivanaer/g-universal-clip
OfficialIn paperpytorch★ 43

Abstract

The Google Universal Image Embedding (GUIE) Challenge is one of the first competitions in multi-domain image representations in the wild, covering a wide distribution of objects: landmarks, artwork, food, etc. This is a fundamental computer vision problem with notable applications in image retrieval, search engines and e-commerce. In this work, we explain our 4th place solution to the GUIE Challenge, and our "bag of tricks" to fine-tune zero-shot Vision Transformers (ViT) pre-trained using CLIP.

Tasks

Image Retrieval Retrieval Zero-Shot Image Classification Zero-shot Image Retrieval Zero-Shot Learning

General Image Descriptors for Open World Image Retrieval using ViT CLIP

Code

Abstract

Tasks

Reproductions