SOTAVerified

Semi-Supervised Video Object Segmentation

The semi-supervised scenario assumes the user inputs a full mask of the object(s) of interest in the first frame of a video sequence. Methods have to produce the segmentation mask for that object(s) in the subsequent frames.

Papers

Showing 1–10 of 147 papers

TitleStatusHype
THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation—0
Exploring Enhanced Contextual Information for Video-Level Object TrackingCode2
A Distractor-Aware Memory for Visual Object Tracking with SAM2Code3
LiVOS: Light Video Object Segmentation with Gated Linear MatchingCode1
Memory Matching is not Enough: Jointly Improving Memory Matching and Decoding for Video Object Segmentation—0
SAM 2: Segment Anything in Images and VideosCode12
Global Motion Understanding in Large-Scale Video Object Segmentation—0
Spatial-Temporal Multi-level Association for Video Object Segmentation—0
Efficient Video Object Segmentation via Modulated Cross-Attention MemoryCode2
Video Object Segmentation with Dynamic Query ModulationCode1
Show:102550
← PrevPage 1 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SwinB-AOTv2-L (MS)J&F93—Unverified
2SwinB-AOST (L'=3, MS)J&F93—Unverified
3SwinB-DeAOT-LJ&F92.9—Unverified
4XMem (MS)J&F92.7—Unverified
5SwinB-AOTv2-LJ&F92.4—Unverified
6SwinB-AOST (L'=3)J&F92.4—Unverified
7R50-DeAOT-LJ&F92.3—Unverified
8R50-AOST (L'=3)J&F92.1—Unverified
9XMem (BL30K)J&F92—Unverified
10DeAOT-LJ&F92—Unverified