| Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples | May 25, 2024 | Active LearningImage-text Retrieval | —Unverified | 0 |
| Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training | Jan 1, 2025 | Image-text RetrievalImage to text | —Unverified | 0 |
| AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection | Apr 28, 2025 | Adversarial AttackAnomaly Detection | —Unverified | 0 |
| Anatomy-Aware Conditional Image-Text Retrieval | Mar 10, 2025 | AnatomyContrastive Learning | —Unverified | 0 |
| AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models | Oct 7, 2024 | Image CaptioningImage-text Retrieval | —Unverified | 0 |
| Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment | Nov 30, 2024 | Image-text RetrievalRepresentation Learning | —Unverified | 0 |
| Assessing Brittleness of Image-Text Retrieval Benchmarks from Vision-Language Models Perspective | Jul 21, 2024 | Image-text RetrievalInformation Retrieval | —Unverified | 0 |
| Asymmetrically Weighted CCA And Hierarchical Kernel Sentence Embedding For Image & Text Retrieval | Nov 19, 2015 | Image-text RetrievalModel Selection | —Unverified | 0 |
| Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP | Jul 18, 2023 | AttributeImage-text Retrieval | —Unverified | 0 |
| Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses | Dec 11, 2024 | Image-text RetrievalQuestion Answering | —Unverified | 0 |
| Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval | Jun 9, 2024 | Image-text RetrievalPerson Retrieval | —Unverified | 0 |
| Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models | May 20, 2025 | Image-text RetrievalText Retrieval | —Unverified | 0 |
| Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs | Apr 24, 2025 | Image-text RetrievalInstruction Following | —Unverified | 0 |
| CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval | Aug 21, 2022 | ClusteringContrastive Learning | —Unverified | 0 |
| CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval | Feb 15, 2022 | Image-text RetrievalRepresentation Learning | —Unverified | 0 |
| Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data | Dec 11, 2023 | Image CaptioningImage-text Retrieval | —Unverified | 0 |
| Constructing Image-Text Pair Dataset from Books | Oct 3, 2023 | Image-text RetrievalOptical Character Recognition (OCR) | —Unverified | 0 |
| Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text Retrieval | Sep 12, 2021 | FormImage-text Retrieval | —Unverified | 0 |
| Constructing Phrase-level Semantic Labels to Form Multi-GrainedSupervision for Image-Text Retrieval | Nov 16, 2021 | FormImage-text Retrieval | —Unverified | 0 |
| Context-Aware Attention Network for Image-Text Retrieval | Jun 1, 2020 | Image-text RetrievalRetrieval | —Unverified | 0 |
| Continual learning in cross-modal retrieval | Apr 14, 2021 | Continual Learningcross-modal alignment | —Unverified | 0 |
| Contrastive Feature Masking Open-Vocabulary Vision Transformer | Sep 2, 2023 | Contrastive LearningImage-text Retrieval | —Unverified | 0 |
| CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging | Jul 10, 2024 | Contrastive LearningImage-text Retrieval | —Unverified | 0 |
| COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval | Apr 15, 2022 | Contrastive LearningCross-Modal Retrieval | —Unverified | 0 |
| CPL: Counterfactual Prompt Learning for Vision and Language Models | Oct 19, 2022 | counterfactualimage-classification | —Unverified | 0 |