SOTAVerified

Image to text

Papers

Showing 201–246 of 246 papers

TitleStatusHype
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing—0
GPC: Generative and General Pathology Image Classifier—0
GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks—0
GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training—0
Hierarchical Gumbel Attention Network for Text-based Person Search—0
HyCIR: Boosting Zero-Shot Composed Image Retrieval with Synthetic Labels—0
I2T2I: Learning Text to Image Synthesis with Textual Data Augmentation—0
Illegible Text to Readable Text: An Image-to-Image Transformation using Conditional Sliced Wasserstein Adversarial Networks—0
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models—0
Image Captioners Sometimes Tell More Than Images They See—0
Image Semantic Relation Generation—0
Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module—0
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything—0
Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation—0
Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration—0
Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling—0
Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards—0
Instruction Tuning-free Visual Token Complement for Multimodal LLMs—0
Interpreting Vision and Language Generative Models with Semantic Visual Priors—0
Is Cross-modal Information Retrieval Possible without Training?—0
I See Dead People: Gray-Box Adversarial Attack on Image-To-Text Models—0
Knowledge Aware Semantic Concept Expansion for Image-Text Matching—0
Knowledge driven Description Synthesis for Floor Plan Interpretation—0
Semantically Grounded QFormer for Efficient Vision Language Understanding—0
Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision—0
Learning Deep Structure-Preserving Image-Text Embeddings—0
Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection—0
Leveraging AI to Generate Audio for User-generated Content in Video Games—0
Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency—0
MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering—0
MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant—0
MFP-CLIP: Exploring the Efficacy of Multi-Form Prompts for Zero-Shot Industrial Anomaly Detection—0
Category-Oriented Representation Learning for Image to Multi-Modal Retrieval—0
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset—0
Multimodal Intelligence: Representation Learning, Information Fusion, and Applications—0
Multimodal Neurons in Pretrained Text-Only Transformers—0
Natural Language Generation—0
Natural Language Generation from Visual Sequences: Challenges and Future Directions—0
Offline Detection of Misspelled Handwritten Words by Convolving Recognition Model Features with Text Labels—0
On the Importance of Text Preprocessing for Multimodal Representation Learning and Pathology Report Generation—0
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation—0
Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text Retrieval—0
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models—0
PiTL: Cross-modal Retrieval with Weakly-supervised Vision-language Pre-training via Prompting—0
RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement—0
Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API—0
Show:102550
← PrevPage 5 of 5Next →

No leaderboard results yet.