SOTAVerified

Referring Expression Segmentation

The task aims at labeling the pixels of an image or video that represent an object instance referred by a linguistic expression. In particular, the referring expression (RE) must allow the identification of an individual object in a discourse or scene (the referent). REs unambiguously identify the target instance.

Papers

Showing 5175 of 145 papers

TitleStatusHype
Referring Image Segmentation Using Text SupervisionCode1
Spectrum-guided Multi-granularity Referring Video Object SegmentationCode1
Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationCode1
OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationCode1
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationCode1
SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationCode1
Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationCode1
Advancing Referring Expression Segmentation Beyond Single ImageCode1
Zero-shot Referring Image Segmentation with Global-Local Context FeaturesCode1
PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationCode1
Multi-Attention Network for Compressed Video Referring Object SegmentationCode1
Towards Robust Referring Video Object Segmentation with Cyclic Relational ConsensusCode1
Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationCode1
SeqTR: A Simple yet Universal Network for Visual GroundingCode1
Local-Global Context Aware Transformer for Language-Guided Video SegmentationCode1
Image Segmentation Using Text and Image PromptsCode1
LAVT: Language-Aware Vision Transformer for Referring Image SegmentationCode1
CRIS: CLIP-Driven Referring Image SegmentationCode1
End-to-End Referring Video Object Segmentation with Multimodal TransformersCode1
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsCode1
Vision-Language Transformer and Query Generation for Referring SegmentationCode1
SynthRef: Generation of Synthetic Referring Expressions for Object SegmentationCode1
Referring Transformer: A One-step Approach to Multi-task Visual GroundingCode1
Cross-Modal Progressive Comprehension for Referring SegmentationCode1
MDETR -- Modulated Detection for End-to-End Multi-Modal UnderstandingCode1
Show:102550
← PrevPage 3 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1DeRIS-LOverall IoU85.41Unverified
2HyperSegOverall IoU84.8Unverified
3PSALMOverall IoU83.6Unverified
4MLCD-Seg-7BOverall IoU83.6Unverified
5HIPIEOverall IoU82.8Unverified
6EVF-SAMOverall IoU82.4Unverified
7UNINEXT-HOverall IoU82.19Unverified
8UniLSeg-100Overall IoU81.74Unverified
9DETRISOverall IoU81Unverified
10C3VGOverall IoU80.89Unverified
#ModelMetricClaimedVerifiedStatus
1DeRIS-LOverall IoU86.49Unverified
2HyperSegOverall IoU85.7Unverified
3MLCD-Seg-7BOverall IoU85.3Unverified
4EVF-SAMOverall IoU84.2Unverified
5HyperSegOverall IoU83.5Unverified
6C3VGOverall IoU83.18Unverified
7MLCD-Seg-7BOverall IoU82.9Unverified
8DeRIS-LOverall IoU82.34Unverified
9DETRISOverall IoU81.9Unverified
10MaskRIS (Swin-B, combined DB)Overall IoU80.64Unverified
#ModelMetricClaimedVerifiedStatus
1MPG-SAM 2J&F73.9Unverified
2VRS-HQ (Chat-UniVi-13B)J&F71Unverified
3GLEE-ProJ&F70.6Unverified
4UNINEXT-HJ&F70.1Unverified
5ReferDINO (Swin-B)J&F69.3Unverified
6MUTRJ&F68.4Unverified
7VLP (VLMo-L)J&F67.6Unverified
8UniRef-L (Swin-L)J&F67.4Unverified
9HTR (Pre-training)J&F67.1Unverified
10DsHmp (Video-Swin-Base)J&F67.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeRIS-LMean IoU78.59Unverified
2MLCD-Seg-7BOverall IoU75.6Unverified
3HyperSegOverall IoU75.2Unverified
4EVF-SAMOverall IoU71.9Unverified
5DETRISOverall IoU70.2Unverified
6C3VGOverall IoU68.95Unverified
7UniLSeg-100Overall IoU68.15Unverified
8UniLSeg-20Overall IoU66.99Unverified
9UNINEXT-HOverall IoU66.22Unverified
10GROUNDHOGOverall IoU64.9Unverified
#ModelMetricClaimedVerifiedStatus
1HINetIoU overall0.68Unverified
2RefVOSIoU overall0.67Unverified
3ClawCraneNetIoU overall0.64Unverified
4CMSA+CFSAIoU overall0.62Unverified
5RefVOSIoU overall0.6Unverified
6SgMg (Video-Swin-B)AP0.59Unverified
7SOC (Video-Swin-B)AP0.57Unverified
8ReferFormer (Video-Swin-B)AP0.55Unverified
9SOC (Video-Swin-T)AP0.5Unverified
10MANETAP0.47Unverified