SOTAVerified

Image Comprehension

Papers

Showing 2649 of 49 papers

TitleStatusHype
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web0
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models0
Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models0
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models0
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models0
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition0
Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges0
Teach Multimodal LLMs to Comprehend Electrocardiographic Images0
Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens0
Unveiling Glitches: A Deep Dive into Image Encoding Bugs within CLIP0
What Large Language Models Bring to Text-rich VQA?0
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA0
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation0
On the Performance of Multimodal Language Models0
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving0
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and CompositionCode0
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and OutputCode0
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image InputsCode0
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-TuningCode0
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image InsertionCode0
RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human FeedbackCode0
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal RetrievalCode0
CLIC: Contrastive Learning Framework for Unsupervised Image Complexity RepresentationCode0
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained ClassificationCode0
Show:102550
← PrevPage 2 of 2Next →

No leaderboard results yet.