SOTAVerified

Visual Grounding

Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence, or even a multi-round dialogue. There are three main challenges in VG:

  • What is the main focus in a query?
  • How to understand an image?
  • How to locate an object?

Papers

Showing 551–571 of 571 papers

TitleStatusHype
Visual Coreference Resolution in Visual Dialog using Neural Module NetworksCode0
Interpretable Visual Question Answering by Visual Grounding from Attention Supervision Mining—0
Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search—0
Visually grounded cross-lingual keyword spotting in speech—0
Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction—0
Finding "It": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos—0
Visual Grounding via Accumulated Attention—0
Rethinking Diversified and Discriminative Proposal Generation for Visual GroundingCode0
Finding beans in burgers: Deep semantic-visual embedding with localizationCode0
Learning Unsupervised Visual Grounding Through Semantic Self-Supervision—0
Interactive Reinforcement Learning for Object Grounding via Self-Talking—0
Improving Visually Grounded Sentence Representations with Self-Attention—0
Self-view Grounding Given a Narrated 360° VideoCode0
Visual Reference Resolution using Attention Memory for Visual Dialog—0
Weakly-supervised Visual Grounding of Phrases with Linguistic Structures—0
Learning Two-Branch Neural Networks for Image-Text Matching TasksCode0
Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation—0
Revisiting Visual Question Answering BaselinesCode0
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual GroundingCode0
Visual Word2Vec (vis-w2v): Learning Visually Grounded Word Embeddings Using Abstract ScenesCode0
Grounding of Textual Phrases in Images by ReconstructionCode0
Show:102550
← PrevPage 12 of 12Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)95.3—Unverified
2mPLUG-2Accuracy (%)92.8—Unverified
3X2-VLM (large)Accuracy (%)92.1—Unverified
4XFM (base)Accuracy (%)90.4—Unverified
5X2-VLM (base)Accuracy (%)90.3—Unverified
6X-VLM (base)Accuracy (%)89—Unverified
7HYDRAIoU61.7—Unverified
8HYDRAIoU61.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)92—Unverified
2mPLUG-2Accuracy (%)86.05—Unverified
3X2-VLM (large)Accuracy (%)81.8—Unverified
4XFM (base)Accuracy (%)79.8—Unverified
5X2-VLM (base)Accuracy (%)78.4—Unverified
6X-VLM (base)Accuracy (%)76.91—Unverified
#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)93.4—Unverified
2mPLUG-2Accuracy (%)90.33—Unverified
3X2-VLM (large)Accuracy (%)87.6—Unverified
4XFM (base)Accuracy (%)86.1—Unverified
5X2-VLM (base)Accuracy (%)85.2—Unverified
6X-VLM (base)Accuracy (%)84.51—Unverified