SOTAVerified

Visual Grounding

Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence, or even a multi-round dialogue. There are three main challenges in VG:

  • What is the main focus in a query?
  • How to understand an image?
  • How to locate an object?

Papers

Showing 301–325 of 571 papers

TitleStatusHype
Improved Visual Grounding through Self-Consistent Explanations—0
Improving Visually Grounded Sentence Representations with Self-Attention—0
Individuation in Neural Models with and without Visual Grounding—0
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention—0
Interactive Reinforcement Learning for Object Grounding via Self-Talking—0
Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction—0
Interpretable Visual Question Answering by Visual Grounding from Attention Supervision Mining—0
Interpretable Visual Question Answering via Reasoning Supervision—0
INVIGORATE: Interactive Visual Grounding and Grasping in Clutter—0
I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs—0
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding—0
Knowledge Supports Visual Language Grounding: A Case Study on Colour Terms—0
Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving—0
Language learning using Speech to Image retrieval—0
LanguageRefer: Spatial-Language Model for 3D Visual Grounding—0
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering—0
Learning from Synthetic Data for Visual Grounding—0
Visually Consistent Hierarchical Image Classification—0
Learning Language Structures through Grounding—0
Learning to Compose and Reason with Language Tree Structures for Visual Grounding—0
Learning to Ground VLMs without Forgetting—0
Learning Unsupervised Visual Grounding Through Semantic Self-Supervision—0
Learning Visual Grounding from Generative Vision and Language Model—0
Learning with Difference Attention for Visually Grounded Self-supervised Representations—0
Less is More: Generating Grounded Navigation Instructions from Landmarks—0
Show:102550
← PrevPage 13 of 23Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)95.3—Unverified
2mPLUG-2Accuracy (%)92.8—Unverified
3X2-VLM (large)Accuracy (%)92.1—Unverified
4XFM (base)Accuracy (%)90.4—Unverified
5X2-VLM (base)Accuracy (%)90.3—Unverified
6X-VLM (base)Accuracy (%)89—Unverified
7HYDRAIoU61.7—Unverified
8HYDRAIoU61.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)92—Unverified
2mPLUG-2Accuracy (%)86.05—Unverified
3X2-VLM (large)Accuracy (%)81.8—Unverified
4XFM (base)Accuracy (%)79.8—Unverified
5X2-VLM (base)Accuracy (%)78.4—Unverified
6X-VLM (base)Accuracy (%)76.91—Unverified
#ModelMetricClaimedVerifiedStatus
1Florence-2-large-ftAccuracy (%)93.4—Unverified
2mPLUG-2Accuracy (%)90.33—Unverified
3X2-VLM (large)Accuracy (%)87.6—Unverified
4XFM (base)Accuracy (%)86.1—Unverified
5X2-VLM (base)Accuracy (%)85.2—Unverified
6X-VLM (base)Accuracy (%)84.51—Unverified