SOTAVerified

Caption Generation

Papers

Showing 126–150 of 310 papers

TitleStatusHype
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning—0
LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images—0
Automated Audio Captioning: An Overview of Recent Progress and New Challenges—0
Knowledge driven Description Synthesis for Floor Plan Interpretation—0
Efficient Audio Captioning Transformer with Patchout and Text Guidance—0
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits—0
Common Subspace for Model and Similarity: Phrase Learning for Caption Generation From Images—0
Language Production Dynamics with Recurrent Neural Networks—0
LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation—0
Clue: Cross-modal Coherence Modeling for Caption Generation—0
DS@BioMed at ImageCLEFmedical Caption 2024: Enhanced Attention Mechanisms in Medical Caption Generation through Concept Detection Integration—0
Domain Adaptation for Neural Networks by Parameter Augmentation—0
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SCICAP Challenge 2023—0
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?—0
Image Captioning using Facial Expression and Attention—0
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation—0
Image Caption Generation Framework for Assamese News using Attention Mechanism—0
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning—0
Image Caption Generation for Low-Resource Assamese Language—0
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers—0
Chittron: An Automatic Bangla Image Captioning System—0
Image to Bengali Caption Generation Using Deep CNN and Bidirectional Gated Recurrent Unit—0
Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space—0
Image Captioning with Integrated Bottom-Up and Multi-level Residual Top-Down Attention for Game Scene Understanding—0
Improving Image Captioning with Better Use of Caption—0
Show:102550
← PrevPage 6 of 13Next →

No leaderboard results yet.