SOTAVerified

Visual Entailment

Visual Entailment (VE) - is a task consisting of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks. The goal is to predict whether the image semantically entails the text.

Papers

Showing 41–50 of 56 papers

TitleStatusHype
Compound Tokens: Channel Fusion for Vision-Language Representation Learning—0
Playing Lottery Tickets with Vision and Language—0
Pre-training image-language transformers for open-vocabulary tasks—0
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training—0
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-training—0
Few-shot Multimodal Multitask Multilingual Learning—0
Prompt Tuning for Generative Multimodal Pretrained Models—0
Visual Entailment: A Novel Task for Fine-Grained Image Understanding—0
Segment-Phrase Table for Semantic Segmentation, Visual Entailment and Paraphrasing—0
How Much Can CLIP Benefit Vision-and-Language Tasks?—0
Show:102550
← PrevPage 5 of 6Next →

No leaderboard results yet.