SOTAVerified

Instance Segmentation

Instance Segmentation is a computer vision task that involves identifying and separating individual objects within an image, including detecting the boundaries of each object and assigning a unique label to each object. The goal of instance segmentation is to produce a pixel-wise segmentation map of the image, where each pixel is assigned to a specific object instance.

Image Credit: Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers, CVPR'21

Papers

Showing 151–200 of 2262 papers

TitleStatusHype
Semantic and Sequential Alignment for Referring Video Object Segmentation—0
Insightful Instance Features for 3D Instance Segmentation—0
Decoupled Motion Expression Video Segmentation—0
PolarNeXt: Rethink Instance Segmentation with Polar Representation—0
DefMamba: Deformable Visual State Space Model—0
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAMCode0
A Novel Shape Guided Transformer Network for Instance Segmentation in Remote Sensing Images—0
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers—0
RelationField: Relate Anything in Radiance FieldsCode2
ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance SegmentationCode1
PyPotteryLens: An Open-Source Deep Learning Framework for Automated Digitisation of Archaeological Pottery DocumentationCode0
SAM-IF: Leveraging SAM for Incremental Few-Shot Instance Segmentation—0
Classification Drives Geographic Bias in Street Scene Segmentation—0
RapidNet: Multi-Level Dilated Convolution Based Mobile BackboneCode1
STEAM: Squeeze and Transform Enhanced Attention Module—0
MaskTerial: A Foundation Model for Automated 2D Material Flake DetectionCode2
Open-Vocabulary High-Resolution 3D (OVHR3D) Data Segmentation and Annotation Framework—0
Integrating YOLO11 and Convolution Block Attention Module for Multi-Season Segmentation of Tree Trunks and Branches in Commercial Apple Orchards—0
DreamColour: Controllable Video Colour Editing without TrainingCode2
Towards Real-Time Open-Vocabulary Video Instance SegmentationCode0
Vision Transformers for Weakly-Supervised Microorganism EnumerationCode0
A2VIS: Amodal-Aware Approach to Video Instance Segmentation—0
Holistic Understanding of 3D Scenes as Universal Scene Description—0
3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting—0
SyncVIS: Synchronized Video Instance SegmentationCode1
Token Cropr: Faster ViTs for Quite a Few TasksCode1
LDA-AQU: Adaptive Query-guided Upsampling via Local Deformable AttentionCode0
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception—0
A Bilayer Segmentation-Recombination Network for Accurate Segmentation of Overlapping C. elegans—0
TinyViM: Frequency Decoupling for Tiny Hybrid Vision MambaCode2
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps—0
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking—0
CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance Segmentation—0
Learn from Foundation Model: Fruit Detection Model without Manual AnnotationCode1
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks—0
CompetitorFormer: Competitor Transformer for 3D Instance Segmentation—0
Entropy Bootstrapping for Weakly Supervised Nuclei Detection—0
DIS-Mine: Instance Segmentation for Disaster-Awareness in Poor-Light Condition in Underground Mines—0
Zero-Shot Automatic Annotation and Instance Segmentation using LLM-Generated Datasets: Eliminating Field Imaging and Manual Annotation for Deep Learning Model Development—0
RETR: Multi-View Radar Detection Transformer for Indoor PerceptionCode1
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing ImageryCode0
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation—0
Horticultural Temporal Fruit Monitoring via 3D Instance Segmentation and Re-Identification using Point CloudsCode0
Data-Centric Learning Framework for Real-Time Detection of Aiming Beam in Fluorescence Lifetime Imaging Guided Surgery—0
Fast and Efficient Transformer-based Method for Bird's Eye View Instance PredictionCode1
SA3DIP: Segment Any 3D Instance with Potential 3D PriorsCode0
Tree level change detection over Ahmedabad city using very high resolution satellite images and Deep Learning—0
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation—0
Automated Classification of Cell Shapes: A Comparative Evaluation of Shape DescriptorsCode1
MV-Adapter: Enhancing Underwater Instance Segmentation via Adaptive Channel Attention—0
Show:102550
← PrevPage 4 of 46Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1InternImage-HAP5080.8—Unverified
2ResNeSt-200 (multi-scale)AP5070.2—Unverified
3CenterMask + VoVNetV2-99 (multi-scale)AP5066.2—Unverified
4CenterMask + VoVNetV2-57 (single-scale)AP5060.8—Unverified
5Co-DETRmask AP57.1—Unverified
6CBNetV2 (EVA02, single-scale)mask AP56.1—Unverified
7ISDA (ResNet-50)APL55.7—Unverified
8EVAmask AP55.5—Unverified
9FD-SwinV2-Gmask AP55.4—Unverified
10Mask Frozen-DETRmask AP55.3—Unverified
#ModelMetricClaimedVerifiedStatus
1InternImage-BGFLOPs501—Unverified
2Co-DETRmask AP56.6—Unverified
3ViT-CoMer-L (Mask RCNN, DINOv2)mask AP55.9—Unverified
4InternImage-Hmask AP55.4—Unverified
5EVAmask AP55—Unverified
6Mask Frozen-DETRmask AP54.9—Unverified
7MasK DINO (SwinL, multi-scale)mask AP54.5—Unverified
8GLEE-Promask AP54.2—Unverified
9ViT-Adapter-L (HTC++, BEiTv2, O365, multi-scale)mask AP54.2—Unverified
10SwinV2-G (HTC++)mask AP53.7—Unverified