SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 21512177 of 2177 papers

TitleStatusHype
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering0
Fine-tuning vs From Scratch: Do Vision & Language Models Have Similar Capabilities on Out-of-Distribution Visual Question Answering?0
An experimental study of the vision-bottleneck in VQA0
Learning to Disambiguate by Asking Discriminative Questions0
Correlation Information Bottleneck: Towards Adapting Pretrained Multimodal Models for Robust Visual Question Answering0
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents0
V-Doc : Visual questions answers with Documents0
V-Doc: Visual Questions Answers With Documents0
An Evaluation of GPT-4V and Gemini in Online VQA0
Neural Reasoning, Fast and Slow, for Video Question Answering0
Learning to Recognize the Unseen Visual Predicates0
Learning to Select Question-Relevant Relations for Visual Question Answering0
Learning to Specialize with Knowledge Distillation for Visual Question Answering0
Fine-tuning Large Language Models with Sequential Instructions0
Learning Visual Knowledge Memory Networks for Visual Question Answering0
An Empirical Study on the Language Modal in Visual Question Answering0
Learning What Makes a Difference from Counterfactual Examples and Gradient Supervision0
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models0
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?0
Fine-Grained Retrieval-Augmented Generation for Visual Question Answering0
Less Is More: Linear Layers on CLIP Features as Powerful VizWiz Model0
Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation Evaluation0
Leveraging Medical Visual Question Answering with Supporting Facts0
Leveraging Visual Question Answering for Image-Caption Ranking0
Leveraging Visual Question Answering to Improve Text-to-Image Synthesis0
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models0
Lightweight In-Context Tuning for Multimodal Unified Models0
Show:102550
← PrevPage 44 of 44Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified