SOTAVerified

Video Instance Segmentation

The goal of video instance segmentation is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain.

To facilitate research on this new task, a large-scale benchmark called YouTube-VIS, which consists of 2,883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks is built.

Papers

Showing 1–10 of 148 papers

TitleStatusHype
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation—0
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects—0
SAM2Auto: Auto Annotation Using FLASH—0
ThinkVideo: High-Quality Reasoning Video Segmentation with Chain of ThoughtsCode0
FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching—0
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection—0
RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety—0
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation—0
Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance Segmentation—0
Decoupled Motion Expression Video Segmentation—0
Show:102550
← PrevPage 1 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1CAVIS(ViT-L, Online)mask AP68.9—Unverified
2DVIS++(ViT-L, Online)mask AP67.7—Unverified
3DVISmask AP64.9—Unverified
4Tube-Linkmask AP64.6—Unverified
5MinVIS (Swin-L)mask AP61.6—Unverified
6Mask2Former (Swin-L)mask AP60.4—Unverified
7UniVS(Swin-L)mask AP60—Unverified
8MDQE(Swin-L)mask AP59.9—Unverified
9SeqFormer (Swin-L)mask AP59.3—Unverified
10DeVIS (Swin-L)mask AP57.1—Unverified