SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

2022-03-19CVPR 2022Code Available2· sign in to hype

Mingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu, Dahua Lin, Shenggao Zhu, Nicholas Yuan, Kai Ding, Lianwen Jin

Code Available — Be the first to reproduce this paper.

Code

github.com/mxin262/swintextspotter
OfficialIn paperpytorch★ 288
github.com/jacobtyo/swintextspotter
pytorch★ 0

Abstract

End-to-end scene text spotting has attracted great attention in recent years due to the success of excavating the intrinsic synergy of the scene text detection and recognition. However, recent state-of-the-art methods usually incorporate detection and recognition simply by sharing the backbone, which does not directly take advantage of the feature interaction between the two tasks. In this paper, we propose a new end-to-end scene text spotting framework termed SwinTextSpotter. Using a transformer encoder with dynamic head as the detector, we unify the two tasks with a novel Recognition Conversion mechanism to explicitly guide text localization through recognition loss. The straightforward design results in a concise framework that requires neither additional rectification module nor character-level annotation for the arbitrarily-shaped text. Qualitative and quantitative experiments on multi-oriented datasets RoIC13 and ICDAR 2015, arbitrarily-shaped datasets Total-Text and CTW1500, and multi-lingual datasets ReCTS (Chinese) and VinText (Vietnamese) demonstrate SwinTextSpotter significantly outperforms existing methods. Code is available at https://github.com/mxin262/SwinTextSpotter.

Tasks

Scene Text Detection Text Detection Text Spotting

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
ICDAR 2015	SwinTextSpotter	F-measure (%) - Strong Lexicon	83.9	—	Unverified
Inverse-Text	SwinTextSpotter	F-measure (%) - No Lexicon	55.4	—	Unverified
SCUT-CTW1500	SwinTextSpotter	F-measure (%) - No Lexicon	51.8	—	Unverified
Total-Text	SwinTextSpotter	F-measure (%) - No Lexicon	74.3	—	Unverified

SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

Code

Abstract

Tasks

Benchmark Results

Reproductions