SOTAVerified

Image to text

Papers

Showing 101–125 of 246 papers

TitleStatusHype
Zero-shot Nuclei Detection via Visual-Language Pre-trained ModelsCode0
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language ModelsCode0
Align before Search: Aligning Ads Image to Text for Accurate Cross-Modal Sponsored SearchCode0
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face DescriptionsCode0
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags—0
Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation—0
Retrieval-Augmented Multimodal Language Modeling—0
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning—0
Revisiting DETR Pre-training for Object Detection—0
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization—0
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization—0
Robustifying Vision-Language Models via Dynamic Token Reweighting—0
See then Tell: Enhancing Key Information Extraction with Vision Grounding—0
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs—0
Sequential Semantic Generative Communication for Progressive Text-to-Image Generation—0
SingleInsert: Inserting New Concepts from a Single Image into Text-to-Image Models for Flexible Editing—0
SLAN: Self-Locator Aided Network for Cross-Modal Understanding—0
SLAN: Self-Locator Aided Network for Vision-Language Understanding—0
SRCB at SemEval-2022 Task 5: Pretraining Based Image to Text Late Sequential Fusion System for Multimodal Misogynous Meme Identification—0
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution—0
Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval—0
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment—0
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image—0
Synthesizing Novel Pairs of Image and Text—0
Task-Oriented Multi-Modal Mutual Leaning for Vision-Language Models—0
Show:102550
← PrevPage 5 of 10Next →

No leaderboard results yet.