SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 10011050 of 2177 papers

TitleStatusHype
Cascaded Mutual Modulation for Visual ReasoningCode0
MIRTT: Learning Multimodal Interaction Representations from Trilinear Transformers for Visual Question AnsweringCode0
End-to-End Instance Segmentation with Recurrent AttentionCode0
End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video FeaturesCode0
MHSAN: Multi-Head Self-Attention Network for Visual Semantic EmbeddingCode0
Mixture-of-Subspaces in Low-Rank AdaptationCode0
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language ModelsCode0
Measuring Faithful and Plausible Visual Grounding in VQACode0
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next ParadigmCode0
Learning to Count Objects in Natural Images for Visual Question AnsweringCode0
Effective Approaches to Batch Parallelization for Dynamic Neural Network ArchitecturesCode0
MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation ModelsCode0
Learning to Follow Object-Centric Image Editing Instructions FaithfullyCode0
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsCode0
Marten: Visual Question Answering with Mask Generation for Multi-modal Document UnderstandingCode0
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question AnsweringCode0
A Question-Centric Model for Visual Question Answering in Medical ImagingCode0
EaSe: A Diagnostic Tool for VQA based on Answer DiversityCode0
MaMMUT: A Simple Architecture for Joint Learning for MultiModal TasksCode0
LXMERT Model Compression for Visual Question AnsweringCode0
Applying recent advances in Visual Question Answering to Record LinkageCode0
LININ: Logic Integrated Neural Inference Network for Explanatory Visual Question AnsweringCode0
LPF: A Language-Prior Feedback Objective Function for De-biased Visual Question AnsweringCode0
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language ResamplersCode0
Dynamic Task and Weight Prioritization Curriculum Learning for Multimodal ImageryCode0
Logical Implications for Visual Question Answering ConsistencyCode0
Locally Smoothed Neural NetworksCode0
Dynamic Memory Networks for Visual and Textual Question AnsweringCode0
Dynamic Key-value Memory Enhanced Multi-step Graph Reasoning for Knowledge-based Visual Question AnsweringCode0
LLaVA-OneVision: Easy Visual Task TransferCode0
Learning Visual Question Answering by Bootstrapping Hard AttentionCode0
LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic SurgeryCode0
Loss re-scaling VQA: Revisiting the LanguagePrior Problem from a Class-imbalance ViewCode0
Overcoming Language Priors in Visual Question Answering via Distinguishing Superficially Similar InstancesCode0
Siamese Tracking with Lingual Object ConstraintsCode0
Learning Visual Knowledge Memory Networks for Visual Question Answering0
Dynamic Fusion With Intra- and Inter-Modality Attention Flow for Visual Question Answering0
Learning to Specialize with Knowledge Distillation for Visual Question Answering0
Learning to Select Question-Relevant Relations for Visual Question Answering0
Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering0
Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks0
Learning to Recognize the Unseen Visual Predicates0
Neural Reasoning, Fast and Slow, for Video Question Answering0
DUBLIN -- Document Understanding By Language-Image Network0
BuDDIE: A Business Document Dataset for Multi-task Information Extraction0
Learning to Disambiguate by Asking Discriminative Questions0
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering0
Learning to Compose Diversified Prompts for Image Emotion Classification0
DualNet: Domain-Invariant Network for Visual Question Answering0
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets0
Show:102550
← PrevPage 21 of 44Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified