Activating More Pixels in Image Super-Resolution Transformer

2022-05-09CVPR 2023Code Available3· sign in to hype

Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, Chao Dong

Code Available — Be the first to reproduce this paper.

Code

github.com/xpixelgroup/hat
OfficialIn paperpytorch★ 1,499
github.com/chxy95/hat
OfficialIn paperpytorch★ 1,497

Abstract

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is still not fully exploited in existing networks. In order to activate more input pixels for better reconstruction, we propose a novel Hybrid Attention Transformer (HAT). It combines both channel attention and window-based self-attention schemes, thus making use of their complementary advantages of being able to utilize global statistics and strong local fitting capability. Moreover, to better aggregate the cross-window information, we introduce an overlapping cross-attention module to enhance the interaction between neighboring window features. In the training stage, we additionally adopt a same-task pre-training strategy to exploit the potential of the model for further improvement. Extensive experiments show the effectiveness of the proposed modules, and we further scale up the model to demonstrate that the performance of this task can be greatly improved. Our overall method significantly outperforms the state-of-the-art methods by more than 1dB. Codes and models are available at https://github.com/XPixelGroup/HAT.

Tasks

Image Super-Resolution Super-Resolution

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
BSD100 - 2x upscaling	HAT-L	PSNR	32.74	—	Unverified
BSD100 - 2x upscaling	HAT	PSNR	32.69	—	Unverified
BSD100 - 3x upscaling	HAT	PSNR	29.59	—	Unverified
BSD100 - 3x upscaling	HAT-L	PSNR	29.63	—	Unverified
BSD100 - 4x upscaling	HAT-L	PSNR	28.09	—	Unverified
BSD100 - 4x upscaling	HAT	PSNR	28.05	—	Unverified
Manga109 - 2x upscaling	HAT-L	PSNR	41.01	—	Unverified
Manga109 - 2x upscaling	HAT	PSNR	40.71	—	Unverified
Manga109 - 3x upscaling	HAT	PSNR	35.84	—	Unverified
Manga109 - 3x upscaling	HAT-L	PSNR	36.02	—	Unverified
Manga109 - 4x upscaling	HAT-L	SSIM	0.93	—	Unverified
Manga109 - 4x upscaling	HAT	SSIM	0.93	—	Unverified
Set14 - 2x upscaling	HAT-L	PSNR	35.29	—	Unverified
Set14 - 2x upscaling	HAT	PSNR	35.13	—	Unverified
Set14 - 3x upscaling	HAT	PSNR	31.33	—	Unverified
Set14 - 3x upscaling	HAT-L	PSNR	31.47	—	Unverified
Set14 - 4x upscaling	HAT-L	PSNR	29.47	—	Unverified
Set14 - 4x upscaling	HAT	PSNR	29.38	—	Unverified
Set5 - 2x upscaling	HAT-L	PSNR	38.91	—	Unverified
Set5 - 2x upscaling	HAT	PSNR	38.73	—	Unverified
Set5 - 3x upscaling	HAT	PSNR	35.16	—	Unverified
Set5 - 3x upscaling	HAT-L	PSNR	35.28	—	Unverified
Set5 - 4x upscaling	HAT-L	PSNR	33.3	—	Unverified
Urban100 - 2x upscaling	HAT-L	PSNR	35.09	—	Unverified
Urban100 - 2x upscaling	HAT	PSNR	34.81	—	Unverified
Urban100 - 3x upscaling	HAT-L	PSNR	30.92	—	Unverified
Urban100 - 3x upscaling	HAT	PSNR	30.7	—	Unverified
Urban100 - 4x upscaling	HAT	PSNR	28.37	—	Unverified
Urban100 - 4x upscaling	HAT-L	PSNR	28.6	—	Unverified

Activating More Pixels in Image Super-Resolution Transformer

Code

Abstract

Tasks

Benchmark Results

Reproductions