AutoFocusFormer: Image Segmentation off the Grid

2023-04-24CVPR 2023Code Available1· sign in to hype

Chen Ziwen, Kaushik Patnaik, Shuangfei Zhai, Alvin Wan, Zhile Ren, Alex Schwing, Alex Colburn, Li Fuxin

Code Available — Be the first to reproduce this paper.

Code

github.com/apple/ml-autofocusformer
Officialpytorch★ 125

Abstract

Real world images often have highly imbalanced content density. Some areas are very uniform, e.g., large patches of blue sky, while other areas are scattered with many small objects. Yet, the commonly used successive grid downsampling strategy in convolutional deep networks treats all areas equally. Hence, small objects are represented in very few spatial locations, leading to worse results in tasks such as segmentation. Intuitively, retaining more pixels representing small objects during downsampling helps to preserve important information. To achieve this, we propose AutoFocusFormer (AFF), a local-attention transformer image recognition backbone, which performs adaptive downsampling by learning to retain the most important pixels for the task. Since adaptive downsampling generates a set of pixels irregularly distributed on the image plane, we abandon the classic grid structure. Instead, we develop a novel point-based local attention block, facilitated by a balanced clustering module and a learnable neighborhood merging module, which yields representations for our point-based versions of state-of-the-art segmentation heads. Experiments show that our AutoFocusFormer (AFF) improves significantly over baseline models of similar sizes.

Tasks

Image Segmentation Instance Segmentation Panoptic Segmentation Segmentation Semantic Segmentation

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
Cityscapes val	AFF-Base (single-scale, point-based Mask2Former)	mask AP	46.2	—	Unverified
Cityscapes val	AFF-Small (single-scale, point-based Mask2Former)	mask AP	44	—	Unverified

AutoFocusFormer: Image Segmentation off the Grid

Code

Abstract

Tasks

Benchmark Results

Reproductions