SOTAVerified

Referring Expression Segmentation

The task aims at labeling the pixels of an image or video that represent an object instance referred by a linguistic expression. In particular, the referring expression (RE) must allow the identification of an individual object in a discourse or scene (the referent). REs unambiguously identify the target instance.

Papers

Showing 51100 of 145 papers

TitleStatusHype
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image SegmentationCode1
MDETR -- Modulated Detection for End-to-End Multi-Modal UnderstandingCode1
Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationCode1
MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationCode1
Multi-Attention Network for Compressed Video Referring Object SegmentationCode1
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsCode1
Multi-task Collaborative Network for Joint Referring Expression Comprehension and SegmentationCode1
Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsCode1
OCID-Ref: A 3D Robotic Dataset with Embodied Language for Clutter Scene GroundingCode1
OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationCode1
PhraseCut: Language-based Image Segmentation in the WildCode1
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?Code1
PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationCode1
Image Segmentation Using Text and Image PromptsCode1
Towards Robust Referring Video Object Segmentation with Cyclic Relational ConsensusCode1
Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationCode1
Referring Image Segmentation Using Text SupervisionCode1
Referring Image Segmentation via Cross-Modal Progressive ComprehensionCode1
Referring Transformer: A One-step Approach to Multi-task Visual GroundingCode1
RefVOS: A Closer Look at Referring Expressions for Video Object SegmentationCode1
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationCode1
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression SegmentationCode1
SeqTR: A Simple yet Universal Network for Visual GroundingCode1
SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationCode1
Spectrum-guided Multi-granularity Referring Video Object SegmentationCode1
SynthRef: Generation of Synthetic Referring Expressions for Object SegmentationCode1
Temporally Consistent Referring Video Object Segmentation with Hybrid MemoryCode1
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression SegmentationCode1
URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale BenchmarkCode1
ViLLa: Video Reasoning Segmentation with Large Language ModelCode1
Vision-Language Transformer and Query Generation for Referring SegmentationCode1
3D-GRES: Generalized 3D Referring Expression SegmentationCode1
Segmentation from Natural Language ExpressionsCode0
CLEVR-Ref+: Diagnosing Visual Reasoning with Referring ExpressionsCode0
MAttNet: Modular Attention Network for Referring Expression ComprehensionCode0
InstructSeq: Unifying Vision Tasks with Instruction-conditioned Multi-modal Sequence GenerationCode0
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context ModelingCode0
Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryCode0
Comprehensive Multi-Modal Interactions for Referring Image SegmentationCode0
Referring Expression Object Segmentation with Caption-Aware ConsistencyCode0
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target GranularitiesCode0
Exploring Modulated Detection Transformer as a Tool for Action Recognition in VideosCode0
Learning To Segment Every Referring Object Point by PointCode0
Cross-Modal Self-Attention Network for Referring Image SegmentationCode0
Bring Adaptive Binding Prototypes to Generalized Referring Expression SegmentationCode0
Expression Prompt Collaboration Transformer for Universal Referring Video Object SegmentationCode0
Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional FiltersCode0
Referring Image Segmentation via Recurrent Refinement NetworksCode0
Towards Omni-supervised Referring Expression SegmentationCode0
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context UnderstandingCode0
Show:102550
← PrevPage 2 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1DeRIS-LOverall IoU85.41Unverified
2HyperSegOverall IoU84.8Unverified
3PSALMOverall IoU83.6Unverified
4MLCD-Seg-7BOverall IoU83.6Unverified
5HIPIEOverall IoU82.8Unverified
6EVF-SAMOverall IoU82.4Unverified
7UNINEXT-HOverall IoU82.19Unverified
8UniLSeg-100Overall IoU81.74Unverified
9DETRISOverall IoU81Unverified
10C3VGOverall IoU80.89Unverified
#ModelMetricClaimedVerifiedStatus
1DeRIS-LOverall IoU86.49Unverified
2HyperSegOverall IoU85.7Unverified
3MLCD-Seg-7BOverall IoU85.3Unverified
4EVF-SAMOverall IoU84.2Unverified
5HyperSegOverall IoU83.5Unverified
6C3VGOverall IoU83.18Unverified
7MLCD-Seg-7BOverall IoU82.9Unverified
8DeRIS-LOverall IoU82.34Unverified
9DETRISOverall IoU81.9Unverified
10MaskRIS (Swin-B, combined DB)Overall IoU80.64Unverified
#ModelMetricClaimedVerifiedStatus
1MPG-SAM 2J&F73.9Unverified
2VRS-HQ (Chat-UniVi-13B)J&F71Unverified
3GLEE-ProJ&F70.6Unverified
4UNINEXT-HJ&F70.1Unverified
5ReferDINO (Swin-B)J&F69.3Unverified
6MUTRJ&F68.4Unverified
7VLP (VLMo-L)J&F67.6Unverified
8UniRef-L (Swin-L)J&F67.4Unverified
9HTR (Pre-training)J&F67.1Unverified
10DsHmp (Video-Swin-Base)J&F67.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeRIS-LMean IoU78.59Unverified
2MLCD-Seg-7BOverall IoU75.6Unverified
3HyperSegOverall IoU75.2Unverified
4EVF-SAMOverall IoU71.9Unverified
5DETRISOverall IoU70.2Unverified
6C3VGOverall IoU68.95Unverified
7UniLSeg-100Overall IoU68.15Unverified
8UniLSeg-20Overall IoU66.99Unverified
9UNINEXT-HOverall IoU66.22Unverified
10GROUNDHOGOverall IoU64.9Unverified
#ModelMetricClaimedVerifiedStatus
1HINetIoU overall0.68Unverified
2RefVOSIoU overall0.67Unverified
3ClawCraneNetIoU overall0.64Unverified
4CMSA+CFSAIoU overall0.62Unverified
5RefVOSIoU overall0.6Unverified
6SgMg (Video-Swin-B)AP0.59Unverified
7SOC (Video-Swin-B)AP0.57Unverified
8ReferFormer (Video-Swin-B)AP0.55Unverified
9SOC (Video-Swin-T)AP0.5Unverified
10MANETAP0.47Unverified