SOTAVerified

Image to text

Papers

Showing 101150 of 246 papers

TitleStatusHype
Zero-shot Nuclei Detection via Visual-Language Pre-trained ModelsCode0
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language ModelsCode0
Align before Search: Aligning Ads Image to Text for Accurate Cross-Modal Sponsored SearchCode0
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face DescriptionsCode0
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags0
Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation0
Retrieval-Augmented Multimodal Language Modeling0
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning0
Revisiting DETR Pre-training for Object Detection0
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization0
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization0
Robustifying Vision-Language Models via Dynamic Token Reweighting0
See then Tell: Enhancing Key Information Extraction with Vision Grounding0
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs0
Sequential Semantic Generative Communication for Progressive Text-to-Image Generation0
SingleInsert: Inserting New Concepts from a Single Image into Text-to-Image Models for Flexible Editing0
SLAN: Self-Locator Aided Network for Cross-Modal Understanding0
SLAN: Self-Locator Aided Network for Vision-Language Understanding0
SRCB at SemEval-2022 Task 5: Pretraining Based Image to Text Late Sequential Fusion System for Multimodal Misogynous Meme Identification0
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution0
Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval0
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment0
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image0
Synthesizing Novel Pairs of Image and Text0
Task-Oriented Multi-Modal Mutual Leaning for Vision-Language Models0
TMCIR: Token Merge Benefits Composed Image Retrieval0
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP0
Towards a Visual-Language Foundation Model for Computational Pathology0
Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering0
TrojVLM: Backdoor Attack Against Vision Language Models0
Turbo Learning for Captionbot and Drawingbot0
Two-stream Hierarchical Similarity Reasoning for Image-text Matching0
Uncertainty-based Cross-Modal Retrieval with Probabilistic Representations0
Understanding the Effect of using Semantically Meaningful Tokens for Visual Representation Learning0
UNITE-FND: Reframing Multimodal Fake News Detection through Unimodal Scene Translation0
Using Inter-Sentence Diverse Beam Search to Reduce Redundancy in Visual Storytelling0
Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor Languages0
Vision-Braille: An End-to-End Tool for Chinese Braille Image-to-Text Translation0
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation0
When are Lemons Purple? The Concept Association Bias of Vision-Language Models0
X-Fusion: Introducing New Modality to Frozen Large Language Models0
15M Multimodal Facial Image-Text Dataset0
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning0
Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution0
ABC: Achieving Better Control of Multimodal Embeddings using VLMs0
Accept the Modality Gap: An Exploration in the Hyperbolic Space0
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training0
AICoderEval: Improving AI Domain Code Generation of Large Language Models0
AI Recommendation System for Enhanced Customer Experience: A Novel Image-to-Text Method0
An End-to-End Neural Network for Image-to-Audio Transformation0
Show:102550
← PrevPage 3 of 5Next →

No leaderboard results yet.