SOTAVerified

Open Vocabulary Object Detection

Open-vocabulary detection (OVD) aims to generalize beyond the limited number of base classes labeled during the training phase. The goal is to detect novel classes defined by an unbounded (open) vocabulary at inference.

Papers

Showing 76–100 of 145 papers

TitleStatusHype
Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object Detection—0
Generating Enhanced Negatives for Training Language-Based Object DetectorsCode0
GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object DetectionCode1
Weakly Supervised Open-Vocabulary Object Detection—0
CLIM: Contrastive Language-Image Mosaic for Region RepresentationCode1
Simple Image-level Classification Improves Open-vocabulary Object DetectionCode1
ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object DetectionCode1
Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection—0
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understandingCode1
Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher LearningCode1
Enhancing Novel Object Detection via Cooperative Foundational ModelsCode1
Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and RetentionCode1
Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCode1
Spuriosity Rankings for Free: A Simple Framework for Last Layer Retraining Based on Object Detection—0
YOLOv8-Based Visual Detection of Road Hazards: Potholes, Sewer Covers, and Manholes—0
LP-OVOD: Open-Vocabulary Object Detection by Linear ProbingCode1
CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object DetectionCode1
OV-VG: A Benchmark for Open-Vocabulary Visual GroundingCode1
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionCode2
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object DetectionCode1
Region-centric Image-Language Pretraining for Open-Vocabulary DetectionCode0
MoCaE: Mixture of Calibrated Experts Significantly Improves Object DetectionCode1
Detect Everything with Few ExamplesCode2
EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment—0
Contrastive Feature Masking Open-Vocabulary Vision Transformer—0
Show:102550
← PrevPage 4 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Cooperative Foundational ModelsAP 0.550.3—Unverified
2DE-ViTAP 0.550—Unverified
3Yolov8-nanoAP 0.547.2—Unverified
4DITOAP 0.546.1—Unverified
5OV-DQUO(RN50x4)AP 0.545.6—Unverified
6LP-OVOD (OWL-ViT Proposals)AP 0.544.9—Unverified
7CLIPSelfAP 0.544.3—Unverified
8CORA+AP 0.543.1—Unverified
9BARONAP 0.542.7—Unverified
10SIA-OVD (RN50x4)AP 0.541.9—Unverified
#ModelMetricClaimedVerifiedStatus
1LaMI-DETRAP novel-LVIS base training43.4—Unverified
2DITOAP novel-LVIS base training40.4—Unverified
3OV-DQUO(ViT-L/14)AP novel-LVIS base training39.3—Unverified
4CoDet (EVA02-L)AP novel-LVIS base training37—Unverified
5CLIPSelfAP novel-LVIS base training34.9—Unverified
6OVMRAP novel-LVIS base training34.4—Unverified
7DE-ViTAP novel-LVIS base training34.3—Unverified
8CFM-ViTAP novel-LVIS base training33.9—Unverified
9CLIM (RN50x64)AP novel-LVIS base training32.3—Unverified
10RO-ViTAP novel-LVIS base training32.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5022.3—Unverified
2ViLDmask AP5018.2—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5042.9—Unverified
2Deticmask AP5042.2—Unverified