SOTAVerified

Caption Generation

Papers

Showing 2650 of 310 papers

TitleStatusHype
Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning0
FE-LWS: Refined Image-Text Representations via Decoder Stacking and Fused Encodings for Remote Sensing Image Captioning0
Expertized Caption Auto-Enhancement for Video-Text RetrievalCode0
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SCICAP Challenge 20230
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsCode4
MAMS: Model-Agnostic Module Selection Framework for Video Captioning0
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing0
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing0
Multi-LLM Collaborative Caption Generation in Scientific DocumentsCode0
Time Series Language Model for Descriptive Caption Generation0
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning0
Multimodal Preference Data Synthetic Alignment with Reward ModelCode0
Learning from Massive Human Videos for Universal Humanoid Pose Control0
From Simple to Professional: A Combinatorial Controllable Image Captioning AgentCode0
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding0
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language ModelsCode2
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains0
Everything is a Video: Unifying Modalities through Next-Frame Prediction0
Grounded Video Caption Generation0
PPLLaVA: Varied Video Sequence Understanding With Prompt GuidanceCode2
Croc: Pretraining Large Multimodal Models with Cross-Modal ComprehensionCode1
MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based AnnotationsCode1
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMsCode0
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning0
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and TrainingCode2
Show:102550
← PrevPage 2 of 13Next →

No leaderboard results yet.