| A Study of Autoregressive Decoders for Multi-Tasking in Computer Vision | Mar 30, 2023 | DecoderMulti-Task Learning | —Unverified | 0 |
| MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks | Mar 29, 2023 | Cross-Modal RetrievalDecoder | CodeCode Available | 0 |
| Curriculum Learning for Compositional Visual Reasoning | Mar 27, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering | Mar 22, 2023 | Question AnsweringVisual Question Answering | CodeCode Available | 0 |
| 3D Concept Learning and Reasoning from Multi-View Images | Mar 20, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| FVQA 2.0: Introducing Adversarial Samples into Fact-based Visual Question Answering | Mar 19, 2023 | Common Sense ReasoningInformation Retrieval | —Unverified | 0 |
| Logical Implications for Visual Question Answering Consistency | Mar 16, 2023 | Language ModelingLanguage Modelling | CodeCode Available | 0 |
| Polar-VQA: Visual Question Answering on Remote Sensed Ice sheet Imagery from Polar Region | Mar 13, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images | Mar 13, 2023 | Common Sense ReasoningExplanation Generation | —Unverified | 0 |
| Vision-Language Models as Success Detectors | Mar 13, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning | Mar 10, 2023 | Few-Shot Image Classificationimage-classification | —Unverified | 0 |
| Toward Unsupervised Realistic Visual Question Answering | Mar 9, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Interpretable Visual Question Answering Referring to Outside Knowledge | Mar 8, 2023 | DiversityImage Captioning | —Unverified | 0 |
| Graph Neural Networks in Vision-Language Image Understanding: A Survey | Mar 7, 2023 | Image CaptioningImage Retrieval | —Unverified | 0 |
| Knowledge-Based Counterfactual Queries for Visual Question Answering | Mar 5, 2023 | counterfactualDecision Making | —Unverified | 0 |
| VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning | Mar 5, 2023 | Answer GenerationEntity Alignment | CodeCode Available | 0 |
| VQA with Cascade of Self- and Co-Attention Blocks | Feb 28, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Language Is Not All You Need: Aligning Perception with Language Models | Feb 27, 2023 | AllImage Captioning | —Unverified | 0 |
| Medical visual question answering using joint self-supervised learning | Feb 25, 2023 | DecoderDiversity | —Unverified | 0 |
| EVJVQA Challenge: Multilingual Visual Question Answering | Feb 23, 2023 | Language ModelingLanguage Modelling | —Unverified | 0 |
| VinVL+L: Enriching Visual Representation with Location Context in VQA | Feb 22, 2023 | Question AnsweringTAG | CodeCode Available | 0 |
| Reusable Slotwise Mechanisms | Feb 21, 2023 | Future predictionObject | —Unverified | 0 |
| Interpretable Medical Image Visual Question Answering via Multi-Modal Relationship Graph Learning | Feb 19, 2023 | Graph LearningMedical Visual Question Answering | —Unverified | 0 |
| Few-shot Multimodal Multitask Multilingual Learning | Feb 19, 2023 | Few-Shot LearningIn-Context Learning | —Unverified | 0 |
| Bridge Damage Cause Estimation Using Multiple Images Based on Visual Question Answering | Feb 18, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |