SOTAVerified

Video Semantic Segmentation

The goal of video semantic segmentation is to assign a predefined class to each pixel in all frames of a video. This requires the model not only to predict accurate segmentation masks but also to ensure that these masks remain temporally consistent across frames. This task has broad applications in areas such as autonomous driving, medical video analysis, and AR/VR.

Papers

Showing 51–75 of 895 papers

TitleStatusHype
Leveraging Motion Information for Better Self-Supervised Video Correspondence Learning—0
Investigation of Frame Differences as Motion Cues for Video Object Segmentation—0
Open-World Skill Discovery from Unsegmented Demonstrations—0
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation—0
Rethinking Few-Shot Medical Image Segmentation by SAM2: A Training-Free Framework with Augmentative Prompting and Dynamic Matching—0
Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object SegmentationCode2
Parameter-free Video Segmentation for Vision and Language Understanding—0
BST: Badminton Stroke-type Transformer for Skeleton-based Action Recognition in Racket SportsCode1
An Analysis of Data Transformation Effects on Segment Anything 2—0
Deep learning approaches to surgical video segmentation and object detection: A Scoping Review—0
Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field—0
Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation—0
SASVi - Segment Any Surgical VideoCode1
Wandering around: A bioinspired approach to visual attention through object motion sensitivityCode0
HD-EPIC: A Highly-Detailed Egocentric Video Dataset—0
Efficient Portrait Matte Creation With Layer Diffusion and Connectivity Priors—0
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations—0
MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationCode1
Efficient Frame Extraction: A Novel Approach Through Frame Similarity and Surgical Tool Tracking for Video SegmentationCode0
Few-shot Structure-Informed Machinery Part Segmentation with Foundation Models and Graph Neural NetworksCode1
Learning Motion and Temporal Cues for Unsupervised Video Object SegmentationCode1
EdgeTAM: On-Device Track Anything ModelCode4
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video CaptioningCode1
Static Segmentation by Tracking: A Frustratingly Label-Efficient Approach to Fine-Grained Segmentation—0
Multi-Context Temporal Consistent Modeling for Referring Video Object SegmentationCode0
Show:102550
← PrevPage 3 of 36Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TMANet-50mIoU80.3—Unverified
2TDNet-50 [9]mIoU79.9—Unverified
3DeltaDist-DDRNet-39mIoU79.9—Unverified
4PSPNet-101 [20]mIoU79.7—Unverified
5PSPNet-50 [20]mIoU78.1—Unverified
6LVS [12]mIoU76.8—Unverified
7GRFP [15]mIoU73.6—Unverified
8FCN-50 [14]mIoU70.1—Unverified
9DFF [22]mIoU69.2—Unverified
#ModelMetricClaimedVerifiedStatus
1TMANet-50Mean IoU76.5—Unverified
2ETC-MobileNetMean IoU76.3—Unverified
3TDNet-50Mean IoU76.2—Unverified
4PSPNet-50Mean IoU76—Unverified
5NetwarpMean IoU74.7—Unverified
6GRFPMean IoU67.1—Unverified
#ModelMetricClaimedVerifiedStatus
1DVIS++(VIT-L)mIoU63.8—Unverified
2UniVS(Swin-L)mIoU59.8—Unverified
3Tube-Link(Swin-large)mIoU59.6—Unverified
4MRCFA(MiT-B5)mIoU49.9—Unverified
5CFFM(MiT-B5)mIoU49.3—Unverified
#ModelMetricClaimedVerifiedStatus
1WaSR-T (ResNet-101)Q60.1—Unverified
2TMANet (ResNet-50)Q57.5—Unverified
3CSANet (ResNet-101)Q49.1—Unverified
#ModelMetricClaimedVerifiedStatus
1MVNet(DeepLabV3)mIoU54.52—Unverified
2MVNet(PSPNet)mIoU54.36—Unverified
3MVNet(FCN)mIoU53.9—Unverified