SOTAVerified

Video Semantic Segmentation

The goal of video semantic segmentation is to assign a predefined class to each pixel in all frames of a video. This requires the model not only to predict accurate segmentation masks but also to ensure that these masks remain temporally consistent across frames. This task has broad applications in areas such as autonomous driving, medical video analysis, and AR/VR.

Papers

Showing 601–650 of 895 papers

TitleStatusHype
FODVid: Flow-guided Object Discovery in Videos—0
FOMTrace: Interactive Video Segmentation By Image Graphs and Fuzzy Object Models—0
FoodMem: Near Real-time and Precise Food Video Segmentation—0
Fully Automated 2D and 3D Convolutional Neural Networks Pipeline for Video Segmentation and Myocardial Infarction Detection in Echocardiography—0
Fully Connected Object Proposals for Video Segmentation—0
Fully Hyperbolic Convolutional Neural Networks—0
Fully Transformer-Equipped Architecture for End-to-End Referring Video Object Segmentation—0
FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos—0
FusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos—0
FVOS for MOSE Track of 4th PVUW Challenge: 3rd Place Solution—0
Gamifying Video Object Segmentation—0
GaussianCut: Interactive segmentation via graph cut for 3D Gaussian Splatting—0
GenDeF: Learning Generative Deformation Field for Video Generation—0
Generalized Product-of-Experts for Learning Multimodal Representations in Noisy Environments—0
Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in Videos—0
Generative Video Propagation—0
Geodesic Distance Histogram Feature for Video Segmentation—0
Geometric Algebra Planes: Convex Implicit Neural Volumes—0
Geometric Context from Videos—0
Global Motion Understanding in Large-Scale Video Object Segmentation—0
Global Optimality Guarantees for Nonconvex Unsupervised Video Segmentation—0
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation—0
Grouping-Based Low-Rank Trajectory Completion and 3D Reconstruction—0
Guess What Moves: Unsupervised Video and Image Segmentation by Anticipating Motion—0
HD-EPIC: A Highly-Detailed Egocentric Video Dataset—0
Hierarchical interaction network for video object segmentation from referring expressions—0
Hierarchical Reinforcement Learning Based Video Semantic Coding for Segmentation—0
Hierarchical Spatiotemporal Transformers for Video Object Segmentation—0
Hierarchical Video Representation with Trajectory Binary Partition Tree—0
High Fidelity Interactive Video Segmentation Using Tensor Decomposition Boundary Loss Convolutional Tessellations and Context Aware Skip Connections—0
Highway Driving Dataset for Semantic Video Segmentation—0
Tamed Warping Network for High-Resolution Semantic Video Segmentation—0
HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation—0
Human Instance Segmentation and Tracking via Data Association and Single-stage Detector—0
Image Segmentation by Uniform Color Clustering Approach and Benchmark Results—0
I-MPN: Inductive Message Passing Network for Efficient Human-in-the-Loop Annotation of Mobile Eye Tracking Data—0
Improved Image Boundaries for Better Video Segmentation—0
Improving Streaming Video Segmentation with Early and Mid-Level Visual Processing—0
Improving Unsupervised Video Object Segmentation with Motion-Appearance Synergy—0
Improving Unsupervised Video Object Segmentation via Fake Flow Generation—0
In defense of OSVOS—0
Instance Embedding Transfer to Unsupervised Video Object Segmentation—0
Instance-Level Video Segmentation From Object Tracks—0
Monocular Instance Motion Segmentation for Autonomous Driving: KITTI InstanceMotSeg Dataset and Multi-task Baseline—0
Interactive Video Object Segmentation in the Wild—0
InterRVOS: Interaction-aware Referring Video Object Segmentation—0
Investigation of Frame Differences as Motion Cues for Video Object Segmentation—0
ISAR: A Benchmark for Single- and Few-Shot Object Instance Segmentation and Re-Identification—0
ISEC: Iterative over-Segmentation via Edge Clustering—0
Is SAM 2 Better than SAM in Medical Image Segmentation?—0
Show:102550
← PrevPage 13 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TMANet-50mIoU80.3—Unverified
2TDNet-50 [9]mIoU79.9—Unverified
3DeltaDist-DDRNet-39mIoU79.9—Unverified
4PSPNet-101 [20]mIoU79.7—Unverified
5PSPNet-50 [20]mIoU78.1—Unverified
6LVS [12]mIoU76.8—Unverified
7GRFP [15]mIoU73.6—Unverified
8FCN-50 [14]mIoU70.1—Unverified
9DFF [22]mIoU69.2—Unverified
#ModelMetricClaimedVerifiedStatus
1TMANet-50Mean IoU76.5—Unverified
2ETC-MobileNetMean IoU76.3—Unverified
3TDNet-50Mean IoU76.2—Unverified
4PSPNet-50Mean IoU76—Unverified
5NetwarpMean IoU74.7—Unverified
6GRFPMean IoU67.1—Unverified
#ModelMetricClaimedVerifiedStatus
1DVIS++(VIT-L)mIoU63.8—Unverified
2UniVS(Swin-L)mIoU59.8—Unverified
3Tube-Link(Swin-large)mIoU59.6—Unverified
4MRCFA(MiT-B5)mIoU49.9—Unverified
5CFFM(MiT-B5)mIoU49.3—Unverified
#ModelMetricClaimedVerifiedStatus
1WaSR-T (ResNet-101)Q60.1—Unverified
2TMANet (ResNet-50)Q57.5—Unverified
3CSANet (ResNet-101)Q49.1—Unverified
#ModelMetricClaimedVerifiedStatus
1MVNet(DeepLabV3)mIoU54.52—Unverified
2MVNet(PSPNet)mIoU54.36—Unverified
3MVNet(FCN)mIoU53.9—Unverified