SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 501510 of 2177 papers

TitleStatusHype
Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question AnsweringCode1
Change Detection Meets Visual Question AnsweringCode1
Debiased Visual Question Answering from Feature and Sample PerspectivesCode1
Searching the Search Space of Vision TransformerCode1
UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language ModelingCode1
Florence: A New Foundation Model for Computer VisionCode1
Many Heads but One Brain: Fusion Brain -- a Competition and a Single Multimodal Multitask ArchitectureCode1
ViVQA: Vietnamese Visual Question AnsweringCode1
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language ReasoningCode1
Label-Descriptive Patterns and Their Application to Characterizing Classification ErrorsCode1
Show:102550
← PrevPage 51 of 218Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified