SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 20012010 of 2177 papers

TitleStatusHype
Question Relevance in Visual Question Answering0
Latent Alignment and Variational AttentionCode0
NMT-Keras: a Very Flexible Toolkit with a Focus on Interactive NMT and Online Learning0
Semantically Equivalent Adversarial Rules for Debugging NLP modelsCode0
Connecting Language and Vision to Actions0
Pushing the Limits of Radiology with Joint Modeling of Visual and Textual Information0
End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video FeaturesCode0
Learning Conditioned Graph Structures for Interpretable Visual Question AnsweringCode0
Learning Visual Knowledge Memory Networks for Visual Question Answering0
iParaphrasing: Extracting Visually Grounded Paraphrases via an ImageCode0
Show:102550
← PrevPage 201 of 218Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified