SOTAVerified

Caption Generation

Papers

Showing 125 of 310 papers

TitleStatusHype
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning0
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real WorldCode2
SonicVerse: Multi-Task Learning for Music Feature-Informed CaptioningCode2
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits0
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation0
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual FusionCode2
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationCode1
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID0
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance0
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks0
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives0
LoVR: A Benchmark for Long Video Retrieval in Multimodal ContextsCode1
VideoMultiAgents: A Multi-Agent Framework for Video Question AnsweringCode1
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation0
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training0
3D CoCa: Contrastive Learners are 3D CaptionersCode0
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention0
Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering0
LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images0
Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic CognitionCode1
Large-scale Pre-training for Grounded Video Caption GenerationCode1
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification0
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models0
Fine-Grained Video Captioning through Scene Graph Consolidation0
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models0
Show:102550
← PrevPage 1 of 13Next →

No leaderboard results yet.