SOTAVerified

Video Semantic Segmentation

The goal of video semantic segmentation is to assign a predefined class to each pixel in all frames of a video. This requires the model not only to predict accurate segmentation masks but also to ensure that these masks remain temporally consistent across frames. This task has broad applications in areas such as autonomous driving, medical video analysis, and AR/VR.

Papers

Showing 651–700 of 895 papers

TitleStatusHype
Is Segment Anything Model 2 All You Need for Surgery Video Segmentation? A Systematic Evaluation—0
Is Two-shot All You Need? A Label-efficient Approach for Video Segmentation in Breast Ultrasound—0
Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation Easier—0
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation—0
Joint Tracking and Segmentation of Multiple Targets—0
JOTS: Joint Online Tracking and Segmentation—0
Key Instance Selection for Unsupervised Video Object Segmentation—0
Immersive Human-Machine Teleoperation Framework for Precision Agriculture: Integrating UAV-based Digital Mapping and Virtual Reality Control—0
Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment—0
Learning a Fast 3D Spectral Approach to Object Segmentation and Tracking over Space and Time—0
Learning a Weakly-Supervised Video Actor-Action Segmentation Model with a Wise Selection—0
Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation—0
Learning Keypoints for Multi-Agent Behavior Analysis using Self-Supervision—0
Learning Pixel Trajectories with Multiscale Contrastive Random Walks—0
Learning Position and Target Consistency for Memory-based Video Object Segmentation—0
Learning Referring Video Object Segmentation from Weak Annotation—0
Learning Representations from Audio-Visual Spatial Alignment—0
Learning Spatial-Semantic Features for Robust Video Object Segmentation—0
Learning the What and How of Annotation in Video Object Segmentation—0
Learning to Adapt to Online Streams with Distribution Shifts—0
Learning to Better Segment Objects from Unseen Classes with Unlabeled Videos—0
Learning To Segment Dominant Object Motion From Watching Videos—0
Learning to Segment Human by Watching YouTube—0
Learning to Segment Moving Objects in Videos—0
Learning to Segment Referred Objects from Narrated Egocentric Videos—0
Learning to Sort Image Sequences via Accumulated Temporal Differences—0
Learning to Track Any Object—0
Learning Video Object Segmentation with Visual Memory—0
Leveraging Motion Information for Better Self-Supervised Video Correspondence Learning—0
Lifelong Learning Using a Dynamically Growing Tree of Sub-networks for Domain Generalization in Video Object Segmentation—0
LIP: Learning Instance Propagation for Video Object Segmentation—0
Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation—0
Look Before You Match: Instance Understanding Matters in Video Object Segmentation—0
LooseCut: Interactive Image Segmentation with Loosely Bounded Boxes—0
Low-Latency Video Semantic Segmentation—0
LSM: Learning Subspace Minimization for Low-level Vision—0
LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS—0
LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation—0
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking—0
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology—0
MAIN: Multi-Attention Instance Network for Video Segmentation—0
MAS for video objects segmentation and tracking based on active contours and SURF descriptor—0
Mask Propagation Network for Video Object Segmentation—0
MaskRNN: Instance Level Video Object Segmentation—0
Maximal Cliques on Multi-Frame Proposal Graph for Unsupervised Video Object Segmentation—0
MEDIAPI-SKEL - A 2D-Skeleton Video Database of French Sign Language With Aligned French Subtitles—0
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer—0
MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation—0
Memory Aggregation Networks for Efficient Interactive Video Object Segmentation—0
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation—0
Show:102550
← PrevPage 14 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TMANet-50mIoU80.3—Unverified
2TDNet-50 [9]mIoU79.9—Unverified
3DeltaDist-DDRNet-39mIoU79.9—Unverified
4PSPNet-101 [20]mIoU79.7—Unverified
5PSPNet-50 [20]mIoU78.1—Unverified
6LVS [12]mIoU76.8—Unverified
7GRFP [15]mIoU73.6—Unverified
8FCN-50 [14]mIoU70.1—Unverified
9DFF [22]mIoU69.2—Unverified
#ModelMetricClaimedVerifiedStatus
1TMANet-50Mean IoU76.5—Unverified
2ETC-MobileNetMean IoU76.3—Unverified
3TDNet-50Mean IoU76.2—Unverified
4PSPNet-50Mean IoU76—Unverified
5NetwarpMean IoU74.7—Unverified
6GRFPMean IoU67.1—Unverified
#ModelMetricClaimedVerifiedStatus
1DVIS++(VIT-L)mIoU63.8—Unverified
2UniVS(Swin-L)mIoU59.8—Unverified
3Tube-Link(Swin-large)mIoU59.6—Unverified
4MRCFA(MiT-B5)mIoU49.9—Unverified
5CFFM(MiT-B5)mIoU49.3—Unverified
#ModelMetricClaimedVerifiedStatus
1WaSR-T (ResNet-101)Q60.1—Unverified
2TMANet (ResNet-50)Q57.5—Unverified
3CSANet (ResNet-101)Q49.1—Unverified
#ModelMetricClaimedVerifiedStatus
1MVNet(DeepLabV3)mIoU54.52—Unverified
2MVNet(PSPNet)mIoU54.36—Unverified
3MVNet(FCN)mIoU53.9—Unverified