SOTAVerified

Visual Dialog

Visual Dialog requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a follow-up question about the image, the task is to answer the question.

Papers

Showing 76100 of 118 papers

TitleStatusHype
What Should I Ask? Using Conversationally Informative Rewards for Goal-oriented Visual Dialog.0
What Should I Ask? Using Conversationally Informative Rewards for Goal-Oriented Visual Dialog0
What's to know? Uncertainty as a Guide to Asking Goal-oriented Questions0
How to Fool Systems and Humans in Visually Grounded Interaction: A Case Study on Adversarial Attacks on Visual Dialog0
ICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report0
Image-Question-Answer Synergistic Network for Visual Dialog0
Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning0
Knowledge Transfer with Visual Prompt in multi-modal Dialogue Understanding and Generation0
Learning Goal-Oriented Visual Dialog Agents: Imitating and Surpassing Analytic Experts0
Factor Graph AttentionCode0
Visual Dialogue without Vision or DialogueCode0
Recursive Visual Attention in Visual DialogCode0
Collecting Visually-Grounded Dialogue with A Game Of SortsCode0
CLEVR-Dialog: A Diagnostic Dataset for Multi-Round Reasoning in Visual DialogCode0
Examining Cooperation in Visual Dialog ModelsCode0
SeqDialN: Sequential Visual Dialog Network in Joint Visual-Linguistic Representation SpaceCode0
Enhancing Visual Dialog Questioner with Entity-based Strategy Learning and Augmented GuesserCode0
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple InputsCode0
LAVIS: A Library for Language-Vision IntelligenceCode0
Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic RepresentationCode0
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueCode0
SeqDialN: Sequential Visual Dialog Networks in Joint Visual-Linguistic Representation SpaceCode0
Learning Goal-Oriented Visual Dialog via Tempered Policy GradientCode0
Spot the Difference: A Cooperative Object-Referring Game in Non-Perfectly Co-Observable SceneCode0
Learning to Reason: End-to-End Module Networks for Visual Question AnsweringCode0
Show:102550
← PrevPage 4 of 5Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SingleNDCG (x 100)78.7Unverified
2P1P2+Distill+EnsembleNDCG (x 100)77.92Unverified
3Ensemble + Fine-tuningNDCG (x 100)76.43Unverified
4ensemble, finetuneNDCG (x 100)76.17Unverified
5VD-PCRNDCG (x 100)76.14Unverified
6EnsembleNDCG (x 100)75.35Unverified
7Ensemble + FinetuneNDCG (x 100)74.88Unverified
8bert-double-stream-finetuningNDCG (x 100)74.62Unverified
9CE-finetuned, single modelNDCG (x 100)74.47Unverified
102NDCG (x 100)73.36Unverified
#ModelMetricClaimedVerifiedStatus
19xFGA (VGG)MRR68.92Unverified
2DANMRR66.38Unverified
3CorefNMN (ResNet-152)MRR64.1Unverified
4CoAttMRR63.98Unverified
5CorefNMNMRR63.6Unverified
6DualVDMRR62.94Unverified
7SF-QIH-se-2MRR62.42Unverified
8HCIAE-NP-ATTMRR62.22Unverified
9HieCoAtt-QIMRR57.88Unverified
10AMEMR@148.53Unverified
#ModelMetricClaimedVerifiedStatus
15xFGA + LSNDCG64.04Unverified
25xFGA + LS*+MRR0.71Unverified
3Two-StepMRR0.7Unverified
#ModelMetricClaimedVerifiedStatus
1Multi-Modal BlenderBotBLEU-41Unverified
#ModelMetricClaimedVerifiedStatus
1Multi-Modal BlenderBotBLEU-41.1Unverified
#ModelMetricClaimedVerifiedStatus
1Multi-Modal BlenderBotBLEU-41.5Unverified
#ModelMetricClaimedVerifiedStatus
1Multi-Modal BlenderBotBLEU-440Unverified
#ModelMetricClaimedVerifiedStatus
1Multi-Modal BlenderBotBLEU-42.2Unverified