Multi-granularity Deep Local Representations for Irregular Scene Text Recognition

Recognizing irregular text from natural scene images is challenging due to the unconstrained appearance of text, such as curvature, orientation, and distortion. Recent recognition networks regard this task as a text sequence labeling problem and most networks capture the sequence only from a single-granularity visual representation, which to some extent limits the performance of recognition. In this article, we propose a hierarchical attention network to capture multi-granularity deep local representations for recognizing irregular scene text. It consists of several hierarchical attention blocks, and each block contains a Local Visual Representation Module (LVRM) and a Decoder Module (DM). Based on the hierarchical attention network, we propose a scene text recognition network. The extensive experiments show that our proposed network achieves the state-of-the-art performance on several benchmark datasets including IIIT-5K, SVT, CUTE, SVT-Perspective, and ICDAR datasets under shorter training time.

Download Full-text

Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v33i01.33018610 ◽

2019 ◽

Vol 33 ◽

pp. 8610-8617 ◽

Cited By ~ 18

Author(s):

Hui Li ◽

Peng Wang ◽

Chunhua Shen ◽

Guyu Zhang

Keyword(s):

State Of The Art ◽

Text Recognition ◽

Natural Scene ◽

Fine Grained ◽

Word Level ◽

Scene Text ◽

Sophisticated Model ◽

Regular Text ◽

Algorithm Implementation ◽

Scene Text Recognition

Recognizing irregular text in natural scene images is challenging due to the large variance in text appearance, such as curvature, orientation and distortion. Most existing approaches rely heavily on sophisticated model designs and/or extra fine-grained annotations, which, to some extent, increase the difficulty in algorithm implementation and data collection. In this work, we propose an easy-to-implement strong baseline for irregular scene text recognition, using offthe-shelf neural network components and only word-level annotations. It is composed of a 31-layer ResNet, an LSTMbased encoder-decoder framework and a 2-dimensional attention module. Despite its simplicity, the proposed method is robust. It achieves state-of-the-art performance on irregular text recognition benchmarks and comparable results on regular text datasets. The code will be released.

Download Full-text

TextScanner: Reading Characters in Order for Robust Scene Text Recognition

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v34i07.6891 ◽

2020 ◽

Vol 34 (07) ◽

pp. 12120-12127 ◽

Cited By ~ 1

Author(s):

Zhaoyi Wan ◽

Minghang He ◽

Haoran Chen ◽

Xiang Bai ◽

Cong Yao

Keyword(s):

State Of The Art ◽

Semantic Segmentation ◽

Text Recognition ◽

Context Modeling ◽

Alternative Approach ◽

Scene Text ◽

Benchmark Datasets ◽

Character Position ◽

Scene Text Recognition ◽

Character Class

Driven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters.

Download Full-text