SOTAVerified

Visual Question Answering

MLLM Leaderboard

Papers

Showing 501550 of 2177 papers

TitleStatusHype
Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question AnsweringCode1
Change Detection Meets Visual Question AnsweringCode1
Debiased Visual Question Answering from Feature and Sample PerspectivesCode1
Searching the Search Space of Vision TransformerCode1
UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language ModelingCode1
Many Heads but One Brain: Fusion Brain -- a Competition and a Single Multimodal Multitask ArchitectureCode1
Florence: A New Foundation Model for Computer VisionCode1
ViVQA: Vietnamese Visual Question AnsweringCode1
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language ReasoningCode1
Label-Descriptive Patterns and Their Application to Characterizing Classification ErrorsCode1
Pano-AVQA: Grounded Audio-Visual Question Answering on 360^ VideosCode1
Coarse-to-Fine Reasoning for Visual Question AnsweringCode1
Counterfactual Samples Synthesizing and Training for Robust Visual Question AnsweringCode1
The Spoon Is in the Sink: Assisting Visually Impaired People in the KitchenCode1
Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesCode1
Does Vision-and-Language Pretraining Improve Lexical Grounding?Code1
xGQA: Cross-Lingual Visual Question AnsweringCode1
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQACode1
Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question AnsweringCode1
WebQA: Multihop and Multimodal QACode1
SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionCode1
X-modaler: A Versatile and High-performance Codebase for Cross-modal AnalyticsCode1
Task-Oriented Multi-User Semantic Communications for VQA TaskCode1
Sparse Continuous Distributions and Fenchel-Young LossesCode1
Check It Again:Progressive Visual Question Answering via Visual EntailmentCode1
Greedy Gradient Ensemble for Robust Visual Question AnsweringCode1
Separating Skills and Concepts for Novel Visual Question AnsweringCode1
How Much Can CLIP Benefit Vision-and-Language Tasks?Code1
Graphhopper: Multi-Hop Scene Graph Reasoning for Visual Question AnsweringCode1
Zero-shot Visual Question Answering using Knowledge GraphCode1
Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question AnsweringCode1
RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsCode1
Predicting Human Scanpaths in Visual Question AnsweringCode1
Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic TestingCode1
Probing Image-Language Transformers for Verb UnderstandingCode1
Check It Again: Progressive Visual Question Answering via Visual EntailmentCode1
Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-TrainingCode1
Multiple Meta-model Quantifying for Medical Visual Question AnsweringCode1
Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using CapsulesCode1
Passage Retrieval for Outside-Knowledge Visual Question AnsweringCode1
MDETR -- Modulated Detection for End-to-End Multi-Modal UnderstandingCode1
GraghVQA: Language-Guided Graph Neural Networks for Graph-based Visual Question AnsweringCode1
Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question AnsweringCode1
MMBERT: Multimodal BERT Pretraining for Improved Medical VQACode1
VisQA: X-raying Vision and Language Reasoning in TransformersCode1
Are Bias Mitigation Techniques for Deep Learning Effective?Code1
Towards General Purpose Vision SystemsCode1
Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersCode1
Multi-Modal Answer Validation for Knowledge-Based VQACode1
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout TransformerCode1
Show:102550
← PrevPage 11 of 44Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MMCTAgent (GPT-4 + GPT-4V)GPT-4 score74.24Unverified
2Qwen2-VL-72BGPT-4 score74Unverified
3InternVL2.5-78BGPT-4 score72.3Unverified
4GPT-4o +text rationale +IoTGPT-4 score72.2Unverified
5Lyra-ProGPT-4 score71.4Unverified
6GLM-4V-PlusGPT-4 score71.1Unverified
7Phantom-7BGPT-4 score70.8Unverified
8InternVL2.5-38BGPT-4 score68.8Unverified
9InternVL2-26B (SGP, token ratio 64%)GPT-4 score65.6Unverified
10Baichuan-Omni (7B)GPT-4 score65.4Unverified