Soft Prompt Generation for Domain Generalization

2024-04-30Code Available1· sign in to hype

Shuanghao Bai, Yuedi Zhang, Wanqi Zhou, Zhirong Luan, Badong Chen

Code Available — Be the first to reproduce this paper.

Code

github.com/renytek13/soft-prompt-generation-with-cgan
OfficialIn paperpytorch★ 32

Abstract

Large pre-trained vision language models (VLMs) have shown impressive zero-shot ability on downstream tasks with manually designed prompt. To further adapt VLMs to downstream tasks, soft prompt is proposed to replace manually designed prompt, which undergoes fine-tuning based on specific domain data. Prior prompt learning methods primarily learn a fixed prompt or residuled prompt from training samples. However, the learned prompts lack diversity and ignore information about unseen domains. In this paper, we reframe the prompt learning framework from a generative perspective and propose a simple yet efficient method for the Domain Generalization (DG) task, namely Soft Prompt Generation (SPG). Specifically, SPG consists of a two-stage training phase and an inference phase. During the training phase, we introduce soft prompt label for each domain, aiming to incorporate the generative model domain knowledge. During the inference phase, the generator of the generative model is employed to obtain instance-specific soft prompts for the unseen target domain. Extensive experiments on five domain generalization benchmarks of three DG tasks demonstrate that SPG achieves state-of-the-art performance. The code is available at https://github.com/renytek13/Soft-Prompt-Generation-with-CGAN.

Tasks

Diversity Domain Generalization Prompt Learning

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
DomainNet	SPG (CLIP, ViT-B/16)	Average Accuracy	60.1	—	Unverified
DomainNet	SPG (CLIP, ResNet-50)	Average Accuracy	50.1	—	Unverified
Office-Home	SPG (CLIP, ViT-B/16)	Average Accuracy	83.6	—	Unverified
Office-Home	SPG (CLIP, ResNet-50)	Average Accuracy	73.8	—	Unverified
PACS	SPG (CLIP, ViT-B/16)	Average Accuracy	97	—	Unverified
PACS	SPG (CLIP, ResNet-50)	Average Accuracy	92.8	—	Unverified
TerraIncognita	SPG (CLIP, ViT-B/16)	Average Accuracy	50.2	—	Unverified
VLCS	SPG (CLIP, ViT-B/16)	Average Accuracy	82.4	—	Unverified
VLCS	SPG (CLIP, ResNet-50)	Average Accuracy	84	—	Unverified

Soft Prompt Generation for Domain Generalization

Code

Abstract

Tasks

Benchmark Results

Reproductions