HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

Mingxin Huang Yuliang Liu Zhenghao Peng Chongyu Liu Dahua Lin Shenggao Zhu Nicholas Yuan Kai Ding Lianwen Jin

SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

Abstract

End-to-end scene text spotting has attracted great attention in recent years due to the success of excavating the intrinsic synergy of the scene text detection and recognition. However, recent state-of-the-art methods usually incorporate detection and recognition simply by sharing the backbone, which does not directly take advantage of the feature interaction between the two tasks. In this paper, we propose a new end-to-end scene text spotting framework termed SwinTextSpotter. Using a transformer encoder with dynamic head as the detector, we unify the two tasks with a novel Recognition Conversion mechanism to explicitly guide text localization through recognition loss. The straightforward design results in a concise framework that requires neither additional rectification module nor character-level annotation for the arbitrarily-shaped text. Qualitative and quantitative experiments on multi-oriented datasets RoIC13 and ICDAR 2015, arbitrarily-shaped datasets Total-Text and CTW1500, and multi-lingual datasets ReCTS (Chinese) and VinText (Vietnamese) demonstrate SwinTextSpotter significantly outperforms existing methods. Code is available at https://github.com/mxin262/SwinTextSpotter.

Code Repositories

jacobtyo/swintextspotter
pytorch
Mentioned in GitHub
mxin262/swintextspotter
Official
pytorch
Mentioned in GitHub

Benchmarks

BenchmarkMethodologyMetrics
text-spotting-on-icdar-2015SwinTextSpotter
F-measure (%) - Generic Lexicon: 70.5
F-measure (%) - Strong Lexicon: 83.9
F-measure (%) - Weak Lexicon: 77.3
text-spotting-on-inverse-textSwinTextSpotter
F-measure (%) - Full Lexicon: 67.9
F-measure (%) - No Lexicon: 55.4
text-spotting-on-scut-ctw1500SwinTextSpotter
F-Measure (%) - Full Lexicon: 77.0
F-measure (%) - No Lexicon: 51.8
text-spotting-on-total-textSwinTextSpotter
F-measure (%) - Full Lexicon: 84.1
F-measure (%) - No Lexicon: 74.3

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp