| Vision-Language Models as Success Detectors | Mar 13, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images | Mar 13, 2023 | Common Sense ReasoningExplanation Generation | —Unverified | 0 |
| Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models | Mar 10, 2023 | Language ModelingLanguage Modelling | CodeCode Available | 1 |
| Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning | Mar 10, 2023 | Few-Shot Image Classificationimage-classification | —Unverified | 0 |
| Toward Unsupervised Realistic Visual Question Answering | Mar 9, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Interpretable Visual Question Answering Referring to Outside Knowledge | Mar 8, 2023 | DiversityImage Captioning | —Unverified | 0 |
| Graph Neural Networks in Vision-Language Image Understanding: A Survey | Mar 7, 2023 | Image CaptioningImage Retrieval | —Unverified | 0 |
| PaLM-E: An Embodied Multimodal Language Model | Mar 6, 2023 | Language ModelingLanguage Modelling | CodeCode Available | 2 |
| Knowledge-Based Counterfactual Queries for Visual Question Answering | Mar 5, 2023 | counterfactualDecision Making | —Unverified | 0 |
| VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning | Mar 5, 2023 | Answer GenerationEntity Alignment | CodeCode Available | 0 |
| Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering | Mar 3, 2023 | Language ModellingLarge Language Model | CodeCode Available | 2 |
| ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax | Mar 2, 2023 | DescriptiveImage Captioning | CodeCode Available | 1 |
| BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs | Mar 2, 2023 | ArticlesMedical Visual Question Answering | CodeCode Available | 1 |
| MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering | Mar 2, 2023 | Mixture-of-ExpertsQuestion Answering | CodeCode Available | 1 |
| RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training | Mar 1, 2023 | Question AnsweringRetrieval | CodeCode Available | 1 |
| VQA with Cascade of Self- and Co-Attention Blocks | Feb 28, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Language Is Not All You Need: Aligning Perception with Language Models | Feb 27, 2023 | AllImage Captioning | —Unverified | 0 |
| Medical visual question answering using joint self-supervised learning | Feb 25, 2023 | DecoderDiversity | —Unverified | 0 |
| Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? | Feb 23, 2023 | Open-Domain Question AnsweringQuestion Answering | CodeCode Available | 1 |
| EVJVQA Challenge: Multilingual Visual Question Answering | Feb 23, 2023 | Language ModelingLanguage Modelling | —Unverified | 0 |
| VinVL+L: Enriching Visual Representation with Location Context in VQA | Feb 22, 2023 | Question AnsweringTAG | CodeCode Available | 0 |
| Reusable Slotwise Mechanisms | Feb 21, 2023 | Future predictionObject | —Unverified | 0 |
| Few-shot Multimodal Multitask Multilingual Learning | Feb 19, 2023 | Few-Shot LearningIn-Context Learning | —Unverified | 0 |
| Interpretable Medical Image Visual Question Answering via Multi-Modal Relationship Graph Learning | Feb 19, 2023 | Graph LearningMedical Visual Question Answering | —Unverified | 0 |
| Bridge Damage Cause Estimation Using Multiple Images Based on Visual Question Answering | Feb 18, 2023 | Question AnsweringVisual Question Answering | —Unverified | 0 |