SOTAVerified

Zero-Shot Image Classification

Zero-shot image classification is a technique in computer vision where a model can classify images into categories that were not present during training. This is achieved by leveraging semantic information about the categories, such as textual descriptions or relationships between classes.

Papers

Showing 51–100 of 111 papers

TitleStatusHype
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP InversionCode0
Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsCode0
KPL: Training-Free Medical Knowledge Mining of Vision-Language ModelsCode0
Learning from Children: Improving Image-Caption Pretraining via CurriculumCode0
Open-vocabulary vs. Closed-set: Best Practice for Few-shot Object Detection Considering Text DescribabilityCode0
Multilingual Vision-Language Pre-training for the Remote Sensing DomainCode0
What Do You See? Enhancing Zero-Shot Image Classification with Multimodal Large Language ModelsCode0
Language-Driven Anchors for Zero-Shot Adversarial RobustnessCode0
Who's in and who's out? A case study of multimodal CLIP-filtering in DataCompCode0
Semantically-Prompted Language Models Improve Visual Descriptions—0
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations—0
A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision—0
Altogether: Image Captioning via Re-aligning Alt-text—0
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene—0
BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models—0
Bayesian Test-Time Adaptation for Vision-Language Models—0
Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation—0
Bridge the Modality and Capability Gaps in Vision-Language Model Selection—0
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization—0
CLAMP: Contrastive LAnguage Model Prompt-tuning—0
Class Knowledge Overlay to Visual Feature Learning for Zero-Shot Image Classification—0
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance—0
CoAPT: Context Attribute words for Prompt Tuning—0
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features—0
DiRaC-I: Identifying Diverse and Rare Training Classes for Zero-Shot Learning—0
Efficient Model-Agnostic Multi-Group Equivariant Networks—0
Efficient Multilingual Multi-modal Pre-training through Triple Contrastive Loss—0
Exploring Low-Resource Medical Image Classification with Weakly Supervised Prompt Learning—0
Gaze Embeddings for Zero-Shot Image Classification—0
Generative Negative Text Replay for Continual Vision-Language Pretraining—0
GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training—0
I2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification—0
I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification—0
CLIPPO: Image-and-Language Understanding from Pixels Only—0
Improving Semantic Embedding Consistency by Metric Learning for Zero-Shot Classification—0
Integrating Propositional and Relational Label Side Information for Hierarchical Zero-Shot Image Classification—0
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap—0
Language to Network: Conditional Parameter Adaptation with Natural Language Descriptions—0
Large-Scale Zero-Shot Image Classification from Rich and Diverse Textual Descriptions—0
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships—0
LightCLIP: Learning Multi-Level Interaction for Lightweight Vision-Language Models—0
LoGra-Med: Long Context Multi-Graph Alignment for Medical Vision-Language Model—0
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification—0
MoDE: CLIP Data Experts via Clustering—0
Multi-method Integration with Confidence-based Weighting for Zero-shot Image Classification—0
Noise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models—0
PaLI: A Jointly-Scaled Multilingual Language-Image Model—0
PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining—0
RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training—0
Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation—0
Show:102550
← PrevPage 2 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1OpenClip H/14 (34B)(Laion2B)Top-1 accuracy30.01—Unverified
#ModelMetricClaimedVerifiedStatus
1CLIP (ViT B-32)Average Score56.64—Unverified
#ModelMetricClaimedVerifiedStatus
1GLIP (Tiny A)Average Score11.4—Unverified