SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 17411750 of 2177 papers

TitleStatusHype
Learning Compositional Representation for Few-shot Visual Question Answering0
Answer Questions with Right Image Regions: A Visual Attention Regularization ApproachCode0
An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games0
Unanswerable Questions about Images and Texts0
Visual Question Answering based on Local-Scene-Aware Referring Expression Generation0
Understanding in Artificial Intelligence0
Latent Variable Models for Visual Question Answering0
Understanding the Role of Scene Graphs in Visual Question Answering0
Predicting Relative Depth between Objects from Semantic Features0
Self Supervision for Attention NetworksCode0
Show:102550
← PrevPage 175 of 218Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified