SOTAVerified

Phrase Grounding

Given an image and a corresponding caption, the Phrase Grounding task aims to ground each entity mentioned by a noun phrase in the caption to a region in the image.

Source: Phrase Grounding by Soft-Label Chain Conditional Random Field

Papers

Showing 1–25 of 88 papers

TitleStatusHype
Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models—0
Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic StructuresCode0
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data—0
Progressive Local Alignment for Medical Multimodal Pre-training—0
Anatomical grounding pre-training for medical phrase groundingCode0
VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human FeedbackCode0
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding—0
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension—0
Towards Visual Grounding: A SurveyCode3
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation—0
Context-Infused Visual Grounding for ArtCode0
Transformer with Controlled Attention for Synchronous Motion CaptioningCode0
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data—0
A Lightweight Modular Framework for Low-Cost Open-Vocabulary Object Detection TrainingCode0
CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting—0
Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational AgentsCode0
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM—0
Zero-Shot Medical Phrase Grounding with Off-the-shelf Diffusion ModelsCode0
MedRG: Medical Report Grounding with Multi-modal Large Language Model—0
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring—0
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training—0
How to Understand "Support"? An Implicit-enhanced Causal Inference Approach for Weakly-supervised Phrase Grounding—0
Phrase Grounding-based Style Transfer for Single-Domain Generalized Object Detection—0
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement—0
An Open and Comprehensive Pipeline for Unified Object Grounding and DetectionCode1
Show:102550
← PrevPage 1 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GLIPv2R@187.7—Unverified
2FIBER-BR@187.4—Unverified
3GLIPR@187.1—Unverified
4PEVLR@184.4—Unverified
5MDETR-ENB5R@184.3—Unverified
6DIGNR@178.73—Unverified
7LCMCGR@176.74—Unverified
8Soft-Label Chain CRF (SL-CCRF)R@174.69—Unverified
9DDPN (ResNet-101)R@173.3—Unverified
10VisualBERTR@171.33—Unverified
#ModelMetricClaimedVerifiedStatus
1GBS Ensemble + 12-in-1Pointing Game Accuracy85.9—Unverified
2GbS Ensemble MS-COCOPointing Game Accuracy75.6—Unverified
3COCO_ELMo_PNASNetPointing Game Accuracy69.19—Unverified
#ModelMetricClaimedVerifiedStatus
1Fiber-BR@187.1—Unverified
2PEVLR@184.1—Unverified
3VisualBERTR@170.4—Unverified
#ModelMetricClaimedVerifiedStatus
1VG_BiLSTM_VGGPointing Game Accuracy62.76—Unverified
2GbS Ensemble MS-COCOPointing Game Accuracy58.21—Unverified
3MCBAccuracy28.91—Unverified
#ModelMetricClaimedVerifiedStatus
1GbS VGPointing Game Accuracy55.91—Unverified
2VG_ELMo_PNASNetPointing Game Accuracy55.16—Unverified
3GbS Ensemble MS-COCOPointing Game Accuracy54.55—Unverified