SOTAVerified

Visual Question Answering (VQA)

Visual Question Answering (VQA) is a task in computer vision that involves answering questions about an image. The goal of VQA is to teach machines to understand the content of an image and answer questions about it in natural language.

Image Source: visualqa.org

Papers

Showing 1–10 of 2167 papers

TitleStatusHype
VisionThink: Smart and Efficient Vision Language Model via Reinforcement LearningCode0
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM—0
Describe Anything Model for Visual Question Answering on Text-rich ImagesCode1
Evaluating Attribute Confusion in Fashion Text-to-Image Generation—0
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation—0
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and GrounderCode1
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning—0
Bridging Video Quality Scoring and Justification via Large Multimodal Models—0
DrishtiKon: Multi-Granular Visual Grounding for Text-Rich Document ImagesCode0
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering—0
Show:102550
← PrevPage 1 of 217Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Gemini Ultra (pixel only)ANLS80.3—Unverified
2SMoLA-PaLI-X SpecialistANLS66.2—Unverified
3ScreenAI 5B (4.62 B params, w/ OCR)ANLS65.9—Unverified
4SMoLA-PaLI-X GeneralistANLS65.6—Unverified
5UDOP (aux)ANLS63—Unverified
6PaLI-3 (w/ OCR)ANLS62.4—Unverified
7TILT-LargeANLS61.2—Unverified
8PaLI-3ANLS57.8—Unverified
9ChatGPT 3.5 with LAPDoc Prompt (SpatialFormat)ANLS54.9—Unverified
10PaLI-X (Single-task FT w/ OCR)ANLS54.8—Unverified