SOTAVerified

Video Instance Segmentation

The goal of video instance segmentation is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain.

To facilitate research on this new task, a large-scale benchmark called YouTube-VIS, which consists of 2,883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks is built.

Papers

Showing 1–10 of 148 papers

TitleStatusHype
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation—0
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects—0
SAM2Auto: Auto Annotation Using FLASH—0
ThinkVideo: High-Quality Reasoning Video Segmentation with Chain of ThoughtsCode0
FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching—0
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection—0
RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety—0
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation—0
Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance Segmentation—0
Decoupled Motion Expression Video Segmentation—0
Show:102550
← PrevPage 1 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1DVIS-DAQ(VIT-L, Offline)mask AP57.1—Unverified
2CAVIS(VIT-L, Offline)mask AP57.1—Unverified
3DVIS++(VIT-L,Offline)mask AP53.4—Unverified
4GLEE-Promask AP50.4—Unverified
5DVIS(Swin-L, Offline)mask AP49.9—Unverified
6DVIS++(VIT-L, Online)mask AP49.6—Unverified
7UNINEXT (ViT-H, Online)mask AP49—Unverified
8DVIS(Swin-L, Online)mask AP47.1—Unverified
9CTVIS (Swin-L)mask AP46.9—Unverified
10RefineVIS (Swin-L, offline)mask AP46—Unverified