Optimizing Relevance Maps of Vision Transformers Improves Robustness

2022-06-02Code Available1· sign in to hype

Hila Chefer, Idan Schwartz, Lior Wolf

Code Available — Be the first to reproduce this paper.

Code

github.com/hila-chefer/robustvit
OfficialIn paperpytorch★ 134

Abstract

It has been observed that visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and manipulate it such that the model is focused on the foreground object. This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required.

Tasks

Image Classification Out-of-Distribution Generalization

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
ObjectNet	AR-L (Opt Relevance)	Top-1 Accuracy	52	—	Unverified
ObjectNet	AR-B (Opt Relevance)	Top-1 Accuracy	47.1	—	Unverified
ObjectNet	AR-L	Top-1 Accuracy	46.5	—	Unverified
ObjectNet	ViT-L (Opt Relevance)	Top-1 Accuracy	43.2	—	Unverified
ObjectNet	ViT-B (Opt Relevance)	Top-1 Accuracy	42.2	—	Unverified
ObjectNet	AR-B	Top-1 Accuracy	41.4	—	Unverified
ObjectNet	AR-S (Opt Relevance)	Top-1 Accuracy	39.3	—	Unverified
ObjectNet	ViT-L	Top-1 Accuracy	37.4	—	Unverified
ObjectNet	DeiT-L (Opt Relevance)	Top-1 Accuracy	36.3	—	Unverified
ObjectNet	ViT-B	Top-1 Accuracy	35.1	—	Unverified
ObjectNet	AR-S	Top-1 Accuracy	34.3	—	Unverified
ObjectNet	DeiT-S (Opt Relevance)	Top-1 Accuracy	31.6	—	Unverified
ObjectNet	DeiT-L	Top-1 Accuracy	31.4	—	Unverified
ObjectNet	DeiT-S	Top-1 Accuracy	28.3	—	Unverified

Optimizing Relevance Maps of Vision Transformers Improves Robustness

Code

Abstract

Tasks

Benchmark Results

Reproductions