SOTAVerified

Referring Expression

Referring expressions places a bounding box around the instance corresponding to the provided description and image.

Papers

Showing 101150 of 364 papers

TitleStatusHype
VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsCode1
A Fast and Accurate One-Stage Approach to Visual GroundingCode1
Relationship-Embedded Representation Learning for Grounding Referring ExpressionsCode1
Generating Easy-to-Understand Referring Expressions for Target IdentificationsCode1
Colors in Context: A Pragmatic Neural Model for Grounded Language UnderstandingCode1
Modeling Context in Referring ExpressionsCode1
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval0
Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language ModelsCode0
Referring Expression Instance Retrieval and A Strong End-to-End Baseline0
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation0
Synthetic Visual Genome0
Refer to Anything with Vision-Language Prompts0
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes0
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning0
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions0
Improving Contrastive Learning for Referring Expression CountingCode0
Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model0
WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationCode0
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models0
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation0
Vision-Language Models Are Not Pragmatically Competent in Referring Expression GenerationCode0
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation0
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation0
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target GranularitiesCode0
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote SensingCode0
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding0
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing0
Cognitive Disentanglement for Referring Multi-Object Tracking0
Exploring Spatial Language Grounding Through Referring Expressions0
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities0
FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis0
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks0
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension0
Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding0
DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension0
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension0
Instance-Aware Generalized Referring Expression Segmentation0
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation0
SegLLM: Multi-round Reasoning Segmentation0
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal ModelsCode0
Grounding Language in Multi-Perspective Referential CommunicationCode0
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension GuidingCode0
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression0
A Lightweight Modular Framework for Low-Cost Open-Vocabulary Object Detection TrainingCode0
Revisiting Multi-Modal LLM Evaluation0
MaskInversion: Localized Embeddings via Optimization of Explainability Maps0
Look Hear: Gaze Prediction for Speech-directed Human Attention0
Learning Visual Grounding from Generative Vision and Language Model0
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge0
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation0
Show:102550
← PrevPage 3 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1RandomAcc@0.5m14.6Unverified