SOTAVerified

Visual Reasoning

Ability to understand actions and reasoning associated with any visual images

Papers

Showing 651698 of 698 papers

TitleStatusHype
Making History Matter: History-Advantage Sequence Training for Visual Dialog0
Can We Automate Diagrammatic Reasoning?0
When Causal Intervention Meets Adversarial Examples and Image Masking for Deep Neural NetworksCode0
Visual Entailment: A Novel Task for Fine-Grained Image Understanding0
Visual Reasoning of Feature Attribution with Deep Recurrent Neural Networks0
CLEVR-Ref+: Diagnosing Visual Reasoning with Referring ExpressionsCode0
Spatial Knowledge Distillation to aid Visual Reasoning0
Learning to Assemble Neural Module Tree Networks for Visual Grounding0
Explainable and Explicit Visual Reasoning over Scene GraphsCode0
Learning to Compose Dynamic Tree Structures for Visual ContextsCode2
A Corpus for Reasoning About Natural Language Grounded in PhotographsCode0
Cascaded Mutual Modulation for Visual ReasoningCode0
Mapping Natural Language Commands to Web ElementsCode0
Visual Reasoning with Multi-hop Feature ModulationCode0
Weakly Supervised Semantic Parsing with Abstract Examples0
Modularity Matters: Learning Invariant Relational Reasoning Tasks0
Object Level Visual Reasoning in VideosCode0
Visual Reasoning by Progressive Module NetworksCode0
Lexical Conceptual Structure of Literal and Metaphorical Spatial Language: A Case Study of ``Push''0
Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal ReasoningCode0
Object Ordering with Bidirectional Matchings for Visual Reasoning0
Iterative Visual Reasoning Beyond Convolutions0
A Dataset and Architecture for Visual Reasoning with a Working MemoryCode0
Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual ReasoningCode0
Compositional Attention Networks for Machine ReasoningCode1
Same-different problems strain convolutional neural networks0
Benchmark Visual Question Answer Models by using Focus Map0
Not-So-CLEVR: Visual Relations Strain Feedforward Neural Networks0
Learning to Act Properly: Predicting and Explaining Affordances from Images0
Multi-Label Zero-Shot Learning with Structured Knowledge GraphsCode0
Weakly-supervised Semantic Parsing with Abstract ExamplesCode0
Complete 3D Scene Parsing from an RGBD ImageCode0
FigureQA: An Annotated Figure Dataset for Visual ReasoningCode0
Visual Reasoning with Natural Language0
FiLM: Visual Reasoning with a General Conditioning LayerCode1
VSE++: Improving Visual-Semantic Embeddings with Hard NegativesCode1
Learning Visual Reasoning Without Strong PriorsCode0
End-to-End Learning of Semantic Grasping0
A Corpus of Natural Language for Visual Reasoning0
How a General-Purpose Commonsense Ontology can Improve Performance of Learning-Based Image RetrievalCode0
Inferring and Executing Programs for Visual ReasoningCode0
EgoReID: Cross-view Self-Identification and Human Re-identification in Egocentric and Surveillance Videos0
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual ReasoningCode1
Dual Local-Global Contextual Pathways for Recognition in Aerial Imagery0
Filling in the details: Perceiving from low fidelity images0
Are Elephants Bigger than Butterflies? Reasoning about Sizes of Objects0
Predicting Complete 3D Models of Indoor ScenesCode0
Factorization of View-Object Manifolds for Joint Object Recognition and Pose Estimation0
Show:102550
← PrevPage 14 of 14Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4o + CAText Score75.5Unverified
2GPT-4V (CoT, pick b/w two options)Text Score75.25Unverified
3GPT-4V (pick b/w two options)Text Score69.25Unverified
4MMICL + CoCoTText Score64.25Unverified
5GPT-4V + CoCoTText Score58.5Unverified
6OpenFlamingo + CoCoTText Score58.25Unverified
7GPT-4VText Score54.5Unverified
8FIBER (EqSim)Text Score51.5Unverified
9FIBER (finetuned, Flickr30k)Text Score51.25Unverified
10MMICL + CCoTText Score51Unverified
#ModelMetricClaimedVerifiedStatus
1BEiT-3Accuracy91.51Unverified
2X2-VLM (large)Accuracy88.7Unverified
3XFM (base)Accuracy87.6Unverified
4X2-VLM (base)Accuracy86.2Unverified
5CoCaAccuracy86.1Unverified
6VLMoAccuracy85.64Unverified
7VK-OODAccuracy84.6Unverified
8SimVLMAccuracy84.53Unverified
9X-VLM (base)Accuracy84.41Unverified
10VK-OODAccuracy83.9Unverified
#ModelMetricClaimedVerifiedStatus
1BEiT-3Accuracy92.58Unverified
2X2-VLM (large)Accuracy89.4Unverified
3XFM (base)Accuracy88.4Unverified
4X2-VLM (base)Accuracy87Unverified
5CoCaAccuracy87Unverified
6VLMoAccuracy86.86Unverified
7SimVLMAccuracy85.15Unverified
8X-VLM (base)Accuracy84.76Unverified
9BLIP-129MAccuracy83.09Unverified
10ALBEF (14M)Accuracy82.55Unverified
#ModelMetricClaimedVerifiedStatus
1AI CoreAverage-per ques.95.24Unverified
2redherringAverage-per ques.91.14Unverified
3VRDPAverage-per ques.90.24Unverified
4FightttttAverage-per ques.88.71Unverified
5neuralAverage-per ques.88.27Unverified
6NERVAverage-per ques.88.05Unverified
7DCLAverage-per ques.75.52Unverified
8troublesolverAverage-per ques.73.3Unverified
9v0.1Average-per ques.73.1Unverified
10First_testAverage-per ques.69.65Unverified
#ModelMetricClaimedVerifiedStatus
1Gemini-2.0 + CA2-Class Accuracy93.6Unverified
2GPT-4o + CA2-Class Accuracy92.8Unverified
3Human2-Class Accuracy91Unverified
4SNAIL2-Class Accuracy64Unverified
5InstructBLIP + GPT-42-Class Accuracy63.8Unverified
6BLIP-2 + ChatGPT (Fine-tuned)2-Class Accuracy63.3Unverified
7InstructBLIP + ChatGPT + Neuro-Symbolic2-Class Accuracy55.5Unverified
8ChatCaptioner + ChatGPT2-Class Accuracy49.3Unverified
9Otter2-Class Accuracy49.3Unverified
#ModelMetricClaimedVerifiedStatus
1HumansJaccard Index90Unverified
2ViLT (Zero-Shot)Jaccard Index52Unverified
3X-VLM (Zero-Shot)Jaccard Index46Unverified
4CLIP-ViT-B/32 (Zero-Shot)Jaccard Index41Unverified
5CLIP-ViT-L/14 (Zero-Shot)Jaccard Index40Unverified
6CLIP-RN50x64/14 (Zero-Shot)Jaccard Index38Unverified
7CLIP-RN50 (Zero-Shot)Jaccard Index35Unverified
8CLIP-ViL (Zero-Shot)Jaccard Index15Unverified
#ModelMetricClaimedVerifiedStatus
1LXMERTaccuracy70.1Unverified
2ViLTaccuracy69.3Unverified
3CLIP (finetuned)accuracy65.1Unverified
4CLIP (frozen)accuracy56Unverified
5VisualBERTaccuracy55.2Unverified
#ModelMetricClaimedVerifiedStatus
1RPINAUCCESS42.2Unverified
2Dec[Joint]1fAUCCESS40.3Unverified
3Dynamics-Aware DQNAUCCESS39.9Unverified
4DQNAUCCESS36.8Unverified
#ModelMetricClaimedVerifiedStatus
1RPINAUCCESS85.2Unverified
2Dynamics-Aware DQNAUCCESS85.2Unverified
3Dec[Joint]1fAUCCESS80Unverified
4DQNAUCCESS77.6Unverified
#ModelMetricClaimedVerifiedStatus
1Swin1:1 Accuracy52.9Unverified
2ConvNeXt1:1 Accuracy51.2Unverified
3ViT1:1 Accuracy50.3Unverified
4DEiT1:1 Accuracy47.2Unverified
#ModelMetricClaimedVerifiedStatus
1Humans1-of-100 Accuracy100Unverified
#ModelMetricClaimedVerifiedStatus
1VisualBERTAccuracy (Dev)67.4Unverified