SOTAVerified

Referring Expression Comprehension

Papers

Showing 126–150 of 167 papers

TitleStatusHype
Give Me Something to Eat: Referring Expression Comprehension with Commonsense Knowledge—0
Giving Commands to a Self-driving Car: A Multimodal Reasoner for Visual Grounding—0
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models—0
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models—0
Cops-Ref: A new Dataset and Task on Compositional Referring Expression Comprehension—0
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension—0
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension—0
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training—0
Video Referring Expression Comprehension via Transformer with Content-conditioned Query—0
Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding—0
Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving—0
Language-Mediated, Object-Centric Representation Learning—0
Text-driven Affordance Learning from Egocentric Vision—0
Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection—0
Leveraging Non-Specialists for Accurate and Time Efficient AMR Annotation—0
Learning Visual Grounding from Generative Vision and Language Model—0
Lite-MDETR: A Lightweight Multi-Modal Detector—0
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge—0
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects—0
M^2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension—0
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression—0
ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments—0
MaskInversion: Localized Embeddings via Optimization of Explainability Maps—0
A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension—0
Compositional Zero-Shot Learning for Attribute-Based Object Reference in Human-Robot Interaction—0
Show:102550
← PrevPage 6 of 7Next →

No leaderboard results yet.