SOTAVerified

Image-to-Text Retrieval

Image-text retrieval is the process of retrieving relevant images based on textual descriptions or finding corresponding textual descriptions for a given image. This task is interdisciplinary, combining techniques from computer vision, and natural language processing. The primary challenge lies in bridging the semantic gap — the difference between how visual data is represented in images and how humans describe that information using language. To address this, many methods focus on learning a shared embedding space where both images and text can be represented in a comparable way, allowing their similarities to be measured and facilitating more accurate retrieval.

Source: Extending CLIP for Category-to-Image Retrieval in E-commerce

Papers

Showing 1–10 of 59 papers

TitleStatusHype
Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration—0
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision ModelsCode1
Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution—0
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs—0
DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation—0
ABC: Achieving Better Control of Multimodal Embeddings using VLMs—0
Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation—0
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding—0
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization—0
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization—0
Show:102550
← PrevPage 1 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1OscarRecall@1099.8—Unverified
2OscarRecall@1098.3—Unverified
3Unicoder-VLRecall@1097.2—Unverified
4BLIP-2 (ViT-G, fine-tuned)Recall@185.4—Unverified
5ONE-PEACE (ViT-G, w/o ranking)Recall@184.1—Unverified
6BLIP-2 (ViT-L, fine-tuned)Recall@183.5—Unverified
7DVSARecall@1074.8—Unverified
8IAISRecall@167.78—Unverified
9CLIP (zero-shot)Recall@158.4—Unverified
10FLAVA (ViT-B, zero-shot)Recall@142.74—Unverified
#ModelMetricClaimedVerifiedStatus
1InternVL-G-FT (finetuned, w/o ranking)Recall@197.9—Unverified
2BLIP-2 ViT-G (zero-shot, 1K test set)Recall@197.6—Unverified
3ONE-PEACE (finetuned, w/o ranking)Recall@197.6—Unverified
4InternVL-C-FT (finetuned, w/o ranking)Recall@197.2—Unverified
5BLIP-2 ViT-L (zero-shot, 1K test set)Recall@196.9—Unverified
6ERNIE-ViL 2.0Recall@196.1—Unverified
7ALBEFRecall@195.9—Unverified
8UNITERRecall@187.3—Unverified
9GSMNRecall@176.4—Unverified
10LGSGMRecall@171—Unverified
#ModelMetricClaimedVerifiedStatus
1BLIP2 FlanT5-XXL (Text-only FT)Specificity94—Unverified
2BLIP2 FlanT5-XXL (Fine-tuned)Specificity84—Unverified
3BLIP2 FlanT5-XL (Fine-tuned)Specificity81—Unverified
4BLIP LargeSpecificity77—Unverified
5CoCa ViT-L-14 MSCOCOSpecificity72—Unverified
6BLIP2 FlanT5-XXL (Zero-shot)Specificity71—Unverified
7CLIP ViT-L/14Specificity70—Unverified
#ModelMetricClaimedVerifiedStatus
1ERNIE-ViL2.0Recall@133.7—Unverified
2CMCLRecall@120.3—Unverified
3ERNIE-ViL2.0Recall@119—Unverified
#ModelMetricClaimedVerifiedStatus
1FETA's CLIP-MIL (Many-Shot Image-to-text)R@135.5—Unverified
2FETA's CLIP-MIL (Many-Shot Image-to-text)R@129—Unverified
#ModelMetricClaimedVerifiedStatus
1CMCLRecall@136.1—Unverified
2CMCLRecall@136—Unverified
#ModelMetricClaimedVerifiedStatus
1SigLIP (ViT-L, zero-shot)Recall@170.6—Unverified
#ModelMetricClaimedVerifiedStatus
1GeoRSCLIP-FTImage to Text Recall@122.14—Unverified