SOTAVerified

Open Vocabulary Semantic Segmentation

Papers

Showing 1–50 of 113 papers

TitleStatusHype
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation—0
ReME: A Data-Centric Framework for Training-Free Open-Vocabulary SegmentationCode1
Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationCode1
A Survey on Training-free Open-Vocabulary Semantic Segmentation—0
Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic SegmentationCode1
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual ReasoningCode1
DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation—0
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic SegmentationCode1
FLOSS: Free Lunch in Open-vocabulary Semantic SegmentationCode1
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability—0
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians—0
Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments—0
Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic SegmentationCode1
LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationCode1
Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive ReinforcementCode4
From Open-Vocabulary to Vocabulary-Free Semantic Segmentation—0
Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation—0
Efficient Redundancy Reduction for Open-Vocabulary Semantic SegmentationCode1
Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models—0
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationCode1
Test-Time Optimization for Domain Adaptive Open Vocabulary SegmentationCode0
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space—0
Dual Semantic Guidance for Open Vocabulary Semantic Segmentation—0
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic SegmentationCode1
OVGaussian: Generalizable 3D Gaussian Segmentation with Open VocabulariesCode0
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language AlignmentCode0
Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation—0
MaskCLIP++: A Mask-Based CLIP Fine-tuning Framework for Open-Vocabulary Image SegmentationCode1
VLMs meet UDA: Boosting Transferability of Open Vocabulary Segmentation with Unsupervised Domain Adaptation—0
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationCode2
Multi-robot autonomous 3D reconstruction using Gaussian splatting with Semantic guidance—0
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation—0
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation—0
HyperSeg: Towards Universal Visual Segmentation with Large Language ModelCode2
Effective SAM Combination for Open-Vocabulary Semantic Segmentation—0
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationCode1
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic SegmentationCode1
ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural EnhancementsCode1
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic SegmentationCode2
InvSeg: Test-Time Prompt Inversion for Semantic Segmentation—0
3D Vision-Language Gaussian Splatting—0
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments—0
SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesCode3
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels—0
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise—0
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation—0
OVOSE: Open-Vocabulary Semantic Segmentation in Event-Based CamerasCode0
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic SegmentationCode2
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary SegmentationCode2
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary SegmentationCode2
Show:102550
← PrevPage 1 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1HyperSegmIoU64.6—Unverified
2SILCmIoU63.5—Unverified
3CAT-SegmIoU63.3—Unverified
4MaskCLIP++mIoU62.5—Unverified
5CLIPSelfmIoU62.3—Unverified
6UMG-CLIP-L/14mIoU61—Unverified
7SEDmIoU60.6—Unverified
8Mask-AdaptermIoU60.4—Unverified
9EBSeg-LmIoU60.2—Unverified
10MAFT+mIoU59.4—Unverified
#ModelMetricClaimedVerifiedStatus
1UMG-CLIP-E/14mIoU38.2—Unverified
2MaskCLIP++mIoU38.2—Unverified
3Mask-AdaptermIoU38.2—Unverified
4CAT-SegmIoU37.9—Unverified
5SILCmIoU37.7—Unverified
6UMG-CLIP-L/14mIoU36.1—Unverified
7MAFT+mIoU36.1—Unverified
8OVSeg + OpenDASmIoU35.8—Unverified
9SEDmIoU35.2—Unverified
10CLIPSelfmIoU34.5—Unverified
#ModelMetricClaimedVerifiedStatus
1UMG-CLIP-E/14mIoU17.3—Unverified
2MaskCLIP++mIoU16.8—Unverified
3Mask-AdaptermIoU16.2—Unverified
4CAT-SegmIoU16—Unverified
5UMG-CLIP-L/14mIoU15.4—Unverified
6MAFT+mIoU15.1—Unverified
7SILCmIoU15—Unverified
8PosSAMmIoU14.9—Unverified
9FC-CLIPmIoU14.8—Unverified
10SCANmIoU14—Unverified
#ModelMetricClaimedVerifiedStatus
1UMG-CLIP-L/14mIoU97.9—Unverified
2SILCmIoU97.6—Unverified
3SCANmIoU97.2—Unverified
4CAT-SegmIoU97—Unverified
5MaskCLIP++mIoU96.8—Unverified
6MAFT+mIoU96.5—Unverified
7EBSeg-LmIoU96.4—Unverified
8FC-CLIPmIoU95.4—Unverified
9OVSeg Swin-BmIoU94.5—Unverified
10HyperSegmIoU92.1—Unverified
#ModelMetricClaimedVerifiedStatus
1SILCmIoU25.8—Unverified
2UMG-CLIP-E/14mIoU25.2—Unverified
3MaskCLIP++mIoU23.9—Unverified
4CAT-SegmIoU23.8—Unverified
5UMG-CLIP-L/14mIoU23.2—Unverified
6Mask-AdaptermIoU22.7—Unverified
7SEDmIoU22.6—Unverified
8MAFT+mIoU21.6—Unverified
9EBSeg-LmIoU21—Unverified
10FC-CLIPmIoU18.2—Unverified
#ModelMetricClaimedVerifiedStatus
1POMPHIoU39.1—Unverified
2ZSSegHIoU37.8—Unverified
3ZegFormerHIoU34.8—Unverified
4TTD (TCL)mIoU23.7—Unverified
5LaVGmIoU23.2—Unverified
6CLIP Surgery (original CLIP without any fine-tuning)mIoU21.9—Unverified
7TTD (MaskCLIP)mIoU19.4—Unverified
#ModelMetricClaimedVerifiedStatus
1FC-CLIPmIoU56.2—Unverified
2SimSegmIoU34.5—Unverified
3TTD (TCL)mIoU32—Unverified
4CLIP Surgery (CLIP without any fine-tuning)mIoU31.4—Unverified
5TTD (MaskCLIP)mIoU27—Unverified
#ModelMetricClaimedVerifiedStatus
1UMG-CLIP-E/14mIoU85.4—Unverified
2CAT-SegmIoU82.5—Unverified
3SILCmIoU82.5—Unverified
4FC-CLIPmIoU81.8—Unverified
#ModelMetricClaimedVerifiedStatus
1SkySense-OmIoU-43.9—Unverified
2SegEarth-OVmIoU-21.7—Unverified
#ModelMetricClaimedVerifiedStatus
1PACLmIoU38.8—Unverified
#ModelMetricClaimedVerifiedStatus
1SkySense-OmIoU8.3—Unverified
#ModelMetricClaimedVerifiedStatus
1SkySense-OmIoU54.1—Unverified
#ModelMetricClaimedVerifiedStatus
1SkySense-OmIoU30.89—Unverified
#ModelMetricClaimedVerifiedStatus
1SkySense-OmIoU32.12—Unverified