SOTAVerified

Referring Expression

Referring expressions places a bounding box around the instance corresponding to the provided description and image.

Papers

Showing 226–250 of 364 papers

TitleStatusHype
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training—0
Cops-Ref: A new Dataset and Task on Compositional Referring Expression Comprehension—0
Corpus-based Referring Expressions Generation—0
CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding—0
Creating Training Corpora for NLG Micro-Planners—0
Decoding Strategies for Neural Referring Expression Generation—0
Decoupling Pragmatics: Discriminative Decoding for Referring Expression Generation—0
Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model—0
Differentiated Relevances Embedding for Group-based Referring Expression Comprehension—0
DisCLIP: Open-Vocabulary Referring Expression Generation—0
Discovering User Groups for Natural Language Generation—0
Dual Convolutional LSTM Network for Referring Image Segmentation—0
DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension—0
Dynamic Graph Attention for Referring Expression Comprehension—0
Dynamic Inference With Grounding Based Vision and Language Models—0
Easy Things First: Installments Improve Referring Expression Generation for Objects in Photographs—0
End-to-End Neural Context Reconstruction in Chinese Dialogue—0
Event versus entity co-reference: Effects of context and form of referring expression—0
Exploring Spatial Language Grounding Through Referring Expressions—0
Exploring the Behavior of Classic REG Algorithms in the Description of Characters in 3D Images—0
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation—0
FindIt: Generalized Localization with Natural Language Queries—0
FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis—0
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes—0
Fully and Weakly Supervised Referring Expression Segmentation with End-to-End Learning—0
Show:102550
← PrevPage 10 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Random[email protected]14.6—Unverified