SOTAVerified

Phrase Grounding

Given an image and a corresponding caption, the Phrase Grounding task aims to ground each entity mentioned by a noun phrase in the caption to a region in the image.

Source: Phrase Grounding by Soft-Label Chain Conditional Random Field

Papers

Showing 26–50 of 88 papers

TitleStatusHype
PG-Video-LLaVA: Pixel Grounding Large Video-Language ModelsCode2
Augment the Pairs: Semantics-Preserving Image-Caption Pair Augmentation for Grounding-Based Vision and Language ModelsCode0
Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeCode0
Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive LearningCode1
Box-based Refinement for Weakly Supervised and Unsupervised Localization TasksCode0
A Joint Study of Phrase Grounding and Task Performance in Vision and Language ModelsCode0
A Survey on Interpretable Cross-modal ReasoningCode1
Catalog Phrase Grounding (CPG): Grounding of Product Textual Attributes in Product Images for e-commerce Vision-Language Applications—0
Kosmos-2: Grounding Multimodal Large Language Models to the WorldCode1
Read, look and detect: Bounding box annotation from image-caption pairs—0
ELVIS: Empowering Locality of Vision Language Pre-training with Intra-modal Similarity—0
CAVL: Learning Contrastive and Adaptive Representations of Vision and Language—0
Trade-offs in Fine-tuned Diffusion Models Between Accuracy and InterpretabilityCode0
LIMITR: Leveraging Local Information for Medical Image-Text Representation—0
Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection—0
Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment—0
Learning to Exploit Temporal Structure for Biomedical Vision-Language ProcessingCode0
Similarity Maps for Self-Training Weakly-Supervised Phrase GroundingCode0
DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingCode1
Extending Phrase Grounding with Pronouns in Visual DialoguesCode0
Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding—0
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection networkCode3
What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsCode1
Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneCode1
GLIPv2: Unifying Localization and Vision-Language UnderstandingCode4
Show:102550
← PrevPage 2 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GLIPv2R@187.7—Unverified
2FIBER-BR@187.4—Unverified
3GLIPR@187.1—Unverified
4PEVLR@184.4—Unverified
5MDETR-ENB5R@184.3—Unverified
6DIGNR@178.73—Unverified
7LCMCGR@176.74—Unverified
8Soft-Label Chain CRF (SL-CCRF)R@174.69—Unverified
9DDPN (ResNet-101)R@173.3—Unverified
10VisualBERTR@171.33—Unverified
#ModelMetricClaimedVerifiedStatus
1GBS Ensemble + 12-in-1Pointing Game Accuracy85.9—Unverified
2GbS Ensemble MS-COCOPointing Game Accuracy75.6—Unverified
3COCO_ELMo_PNASNetPointing Game Accuracy69.19—Unverified
#ModelMetricClaimedVerifiedStatus
1Fiber-BR@187.1—Unverified
2PEVLR@184.1—Unverified
3VisualBERTR@170.4—Unverified
#ModelMetricClaimedVerifiedStatus
1VG_BiLSTM_VGGPointing Game Accuracy62.76—Unverified
2GbS Ensemble MS-COCOPointing Game Accuracy58.21—Unverified
3MCBAccuracy28.91—Unverified
#ModelMetricClaimedVerifiedStatus
1GbS VGPointing Game Accuracy55.91—Unverified
2VG_ELMo_PNASNetPointing Game Accuracy55.16—Unverified
3GbS Ensemble MS-COCOPointing Game Accuracy54.55—Unverified