SOTAVerified

Image to text

Papers

Showing 201–225 of 246 papers

TitleStatusHype
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing—0
GPC: Generative and General Pathology Image Classifier—0
GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks—0
GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training—0
Hierarchical Gumbel Attention Network for Text-based Person Search—0
HyCIR: Boosting Zero-Shot Composed Image Retrieval with Synthetic Labels—0
I2T2I: Learning Text to Image Synthesis with Textual Data Augmentation—0
Illegible Text to Readable Text: An Image-to-Image Transformation using Conditional Sliced Wasserstein Adversarial Networks—0
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models—0
Image Captioners Sometimes Tell More Than Images They See—0
Image Semantic Relation Generation—0
Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module—0
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything—0
Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation—0
Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration—0
Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling—0
Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards—0
Instruction Tuning-free Visual Token Complement for Multimodal LLMs—0
Interpreting Vision and Language Generative Models with Semantic Visual Priors—0
Is Cross-modal Information Retrieval Possible without Training?—0
I See Dead People: Gray-Box Adversarial Attack on Image-To-Text Models—0
Knowledge Aware Semantic Concept Expansion for Image-Text Matching—0
Knowledge driven Description Synthesis for Floor Plan Interpretation—0
Semantically Grounded QFormer for Efficient Vision Language Understanding—0
Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision—0
Show:102550
← PrevPage 9 of 10Next →

No leaderboard results yet.