SOTAVerified

Caption Generation

Papers

Showing 101150 of 310 papers

TitleStatusHype
Bi-directional Contextual Attention for 3D Dense Captioning0
Dual-path Collaborative Generation Network for Emotional Video CaptioningCode0
XMeCap: Meme Caption Generation with Sub-Image Adaptability0
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing ImagesCode0
Explainable Image Captioning using CNN- CNN architecture and Hierarchical Attention0
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?0
Enhancing Cross-Prompt Transferability in Vision-Language Models through Contextual Injection of Target TokensCode0
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon CaptioningCode0
DS@BioMed at ImageCLEFmedical Caption 2024: Enhanced Attention Mechanisms in Medical Caption Generation through Concept Detection Integration0
Multi-Modal Generative Embedding Model0
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation0
MICap: A Unified Model for Identity-aware Movie Descriptions0
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation0
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge0
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival0
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning0
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback0
LLMs in Political Science: Heralding a New Era of Visual Analysis0
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation0
Social Media Ready Caption Generation for Brands0
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving0
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning0
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERTCode0
Enhancing Image Captioning with Neural Models0
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers0
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism0
Dense Video Captioning: A Survey of Techniques, Datasets and Evaluation Protocols0
Visual Analytics for Efficient Image Exploration and User-Guided Image Captioning0
LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation0
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools0
ViPE: Visualise Pretty-much EverythingCode0
A Comparative Study of Pre-trained CNNs and GRU-Based Attention for Image Caption Generation0
FaceGemma: Enhancing Image Captioning with Facial Attributes for Portrait Images0
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning0
ViCo: Engaging Video Comment Generation with Human Preference Rewards0
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human FeedbackCode0
AIC-AB NET: A Neural Network for Image Captioning with Spatial Attention and Text Attributes0
Multi-Similarity Contrastive Learning0
Knowledge Distillation for Efficient Audio-Visual Video Captioning0
SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure CaptioningCode0
CapText: Large Language Model-based Caption Generation From Image Context and Description0
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment0
HAAV: Hierarchical Aggregation of Augmented Views for Image Captioning0
DiffCap: Exploring Continuous Diffusion on Image Captioning0
Efficient Audio Captioning Transformer with Patchout and Text Guidance0
Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models0
Multi-modal reward for visual relationships-based image captioning0
GNNFormer: A Graph-based Framework for Cytopathology Report Generation0
Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text SummarizationCode0
Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning0
Show:102550
← PrevPage 3 of 7Next →

No leaderboard results yet.