SOTAVerified

Visual Entailment

Visual Entailment (VE) - is a task consisting of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks. The goal is to predict whether the image semantically entails the text.

Papers

Showing 3140 of 56 papers

TitleStatusHype
Visual Entailment Task for Visually-Grounded Language Learning0
Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment0
AlignVE: Visual Entailment Recognition Based on Alignment Relations0
Answer-Me: Multi-Task Open-Vocabulary Visual Question Answering0
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks0
Prompt Tuning for Generative Multimodal Pretrained Models0
Visual Entailment: A Novel Task for Fine-Grained Image Understanding0
Segment-Phrase Table for Semantic Segmentation, Visual Entailment and Paraphrasing0
How Much Can CLIP Benefit Vision-and-Language Tasks?0
Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning0
Show:102550
← PrevPage 4 of 6Next →

No leaderboard results yet.