SOTAVerified

Open Vocabulary Object Detection

Open-vocabulary detection (OVD) aims to generalize beyond the limited number of base classes labeled during the training phase. The goal is to detect novel classes defined by an unbounded (open) vocabulary at inference.

Papers

Showing 51–100 of 145 papers

TitleStatusHype
Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCode1
MoCaE: Mixture of Calibrated Experts Significantly Improves Object DetectionCode1
Enhancing Novel Object Detection via Cooperative Foundational ModelsCode1
CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingCode1
Multi-Modal Classifiers for Open-Vocabulary Object DetectionCode1
Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionCode1
Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and RetentionCode1
Exploiting Unlabeled Data with Vision and Language Models for Object DetectionCode1
Open-vocabulary Attribute DetectionCode1
Open-Vocabulary Object Detection Using CaptionsCode1
Open-vocabulary Object Detection via Vision and Language Knowledge DistillationCode1
DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model TrainingCode1
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object DetectionCode1
Open Vocabulary Object Detection with Proposal Mining and Prediction EqualizationCode1
Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationCode1
V3Det: Vast Vocabulary Visual Detection DatasetCode1
Described Object Detection: Liberating Object Detection with Flexible ExpressionsCode1
Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher LearningCode1
OvarNet: Towards Open-vocabulary Object Attribute RecognitionCode1
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects SupervisionCode1
OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesCode1
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object TrackingCode1
OV-VG: A Benchmark for Open-Vocabulary Visual GroundingCode1
OW-OVD: Unified Open World and Open Vocabulary Object DetectionCode1
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel ObjectsCode1
PointCLIP: Point Cloud Understanding by CLIPCode1
ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object DetectionCode1
Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision TransformersCode1
RegionCLIP: Region-based Language-Image PretrainingCode1
Retrieval-Augmented Open-Vocabulary Object DetectionCode1
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object DetectionCode1
Open-Vocabulary Object Detection via Scene Graph Discovery—0
An Application-Agnostic Automatic Target Recognition System Using Vision Language Models—0
An Iterative Feedback Mechanism for Improving Natural Language Class Descriptions in Open-Vocabulary Object Detection—0
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction—0
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs—0
Boosting Open-Vocabulary Object Detection by Handling Background Samples—0
Contrastive Feature Masking Open-Vocabulary Vision Transformer—0
DenseVLM: A Retrieval and Decoupled Alignment Framework for Open-Vocabulary Dense Prediction—0
DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment—0
DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection—0
DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection—0
EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment—0
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting—0
Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024—0
Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection—0
Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection—0
Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object Detection—0
Few-shot target-driven instance detection based on open-vocabulary object detection models—0
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation—0
Show:102550
← PrevPage 2 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Cooperative Foundational ModelsAP 0.550.3—Unverified
2DE-ViTAP 0.550—Unverified
3Yolov8-nanoAP 0.547.2—Unverified
4DITOAP 0.546.1—Unverified
5OV-DQUO(RN50x4)AP 0.545.6—Unverified
6LP-OVOD (OWL-ViT Proposals)AP 0.544.9—Unverified
7CLIPSelfAP 0.544.3—Unverified
8CORA+AP 0.543.1—Unverified
9BARONAP 0.542.7—Unverified
10SIA-OVD (RN50x4)AP 0.541.9—Unverified
#ModelMetricClaimedVerifiedStatus
1LaMI-DETRAP novel-LVIS base training43.4—Unverified
2DITOAP novel-LVIS base training40.4—Unverified
3OV-DQUO(ViT-L/14)AP novel-LVIS base training39.3—Unverified
4CoDet (EVA02-L)AP novel-LVIS base training37—Unverified
5CLIPSelfAP novel-LVIS base training34.9—Unverified
6OVMRAP novel-LVIS base training34.4—Unverified
7DE-ViTAP novel-LVIS base training34.3—Unverified
8CFM-ViTAP novel-LVIS base training33.9—Unverified
9CLIM (RN50x64)AP novel-LVIS base training32.3—Unverified
10RO-ViTAP novel-LVIS base training32.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5022.3—Unverified
2ViLDmask AP5018.2—Unverified
#ModelMetricClaimedVerifiedStatus
1Object-Centric-OVDmask AP5042.9—Unverified
2Deticmask AP5042.2—Unverified