SOTAVerified

Open Vocabulary Object Detection

Open-vocabulary detection (OVD) aims to generalize beyond the limited number of base classes labeled during the training phase. The goal is to detect novel classes defined by an unbounded (open) vocabulary at inference.

Papers

Showing 101–145 of 145 papers

TitleStatusHype
OpenIns3D: Snap and Lookup for 3D Open-vocabulary Instance SegmentationCode2
Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection—0
Taming Self-Training for Open-Vocabulary Object DetectionCode1
Described Object Detection: Liberating Object Detection with Flexible ExpressionsCode1
Open-Vocabulary Object Detection via Scene Graph Discovery—0
Scaling Open-Vocabulary Object DetectionCode0
Multi-Modal Classifiers for Open-Vocabulary Object DetectionCode1
Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision TransformersCode1
DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment—0
V3Det: Vast Vocabulary Visual Detection DatasetCode1
MaMMUT: A Simple Architecture for Joint Learning for MultiModal TasksCode0
Prompt-Guided Transformers for End-to-End Open-Vocabulary Object Detection—0
CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingCode1
Open-Vocabulary Object Detection using Pseudo Caption Labels—0
Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection—0
Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionCode1
Aligning Bag of Regions for Open-Vocabulary Object DetectionCode1
OvarNet: Towards Open-vocabulary Object Attribute RecognitionCode1
Open-Vocabulary Object Detection With an Open Corpus—0
Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionCode1
Learning To Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceCode1
Learning to Detect and Segment for Open Vocabulary Object Detection—0
X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionCode1
Learning Object-Language Alignments for Open-Vocabulary Object DetectionCode1
Open-vocabulary Attribute DetectionCode1
PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningCode2
Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language ModelsCode1
Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection—0
F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsCode0
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection networkCode3
Exploiting Unlabeled Data with Vision and Language Models for Object DetectionCode1
Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionCode2
Open Vocabulary Object Detection with Proposal Mining and Prediction EqualizationCode1
GLIPv2: Unifying Localization and Vision-Language UnderstandingCode4
Simple Open-Vocabulary Object Detection with Vision TransformersCode0
Localized Vision-Language Matching for Open-vocabulary Object DetectionCode1
Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelCode1
Open-Vocabulary DETR with Conditional MatchingCode2
Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationCode1
Detecting Twenty-thousand Classes using Image-level SupervisionCode3
RegionCLIP: Region-based Language-Image PretrainingCode1
PointCLIP: Point Cloud Understanding by CLIPCode1
Open Vocabulary Object Detection with Pseudo Bounding-Box LabelsCode1
Open-vocabulary Object Detection via Vision and Language Knowledge DistillationCode1
Open-Vocabulary Object Detection Using CaptionsCode1
Show:102550
← PrevPage 3 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Cooperative Foundational ModelsAP 0.550.3—Unverified
2DE-ViTAP 0.550—Unverified
3Yolov8-nanoAP 0.547.2—Unverified
4DITOAP 0.546.1—Unverified
5OV-DQUO(RN50x4)AP 0.545.6—Unverified
6LP-OVOD (OWL-ViT Proposals)AP 0.544.9—Unverified
7CLIPSelfAP 0.544.3—Unverified
8CORA+AP 0.543.1—Unverified
9BARONAP 0.542.7—Unverified
10SIA-OVD (RN50x4)AP 0.541.9—Unverified
#ModelMetricClaimedVerifiedStatus
1LaMI-DETRAP novel-LVIS base training43.4—Unverified
2DITOAP novel-LVIS base training40.4—Unverified
3OV-DQUO(ViT-L/14)AP novel-LVIS base training39.3—Unverified
4CoDet (EVA02-L)AP novel-LVIS base training37—Unverified
5CLIPSelfAP novel-LVIS base training34.9—Unverified
6OVMRAP novel-LVIS base training34.4—Unverified
7DE-ViTAP novel-LVIS base training34.3—Unverified
8CFM-ViTAP novel-LVIS base training33.9—Unverified
9CLIM (RN50x64)AP novel-LVIS base training32.3—Unverified
10RO-ViTAP novel-LVIS base training32.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5022.3—Unverified
2ViLDmask AP5018.2—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5042.9—Unverified
2Deticmask AP5042.2—Unverified