DSRN: A Deep Scale Relationship Network for Scene Text Detection

Nowadays, scene text detection has become increasingly important and popular. However, the large variance of text scale remains the main challenge and limits the detection performance in most previous methods. To address this problem, we propose an end-to-end architecture called Deep Scale Relationship Network (DSRN) to map multi-scale convolution features onto a scale invariant space to obtain uniform activation of multi-size text instances. Firstly, we develop a Scale-transfer module to transfer the multi-scale feature maps to a unified dimension. Due to the heterogeneity of features, simply concatenating feature maps with multi-scale information would limit the detection performance. Thus we propose a Scale Relationship module to aggregate the multi-scale information through bi-directional convolution operations. Finally, to further reduce the miss-detected instances, a novel Recall Loss is proposed to force the network to concern more about miss-detected text instances by up-weighting poor-classified examples. Compared with previous approaches, DSRN efficiently handles the large-variance scale problem without complex hand-crafted hyperparameter settings (e.g. scale of default boxes) and complicated post processing. On standard datasets including ICDAR2015 and MSRA-TD500, the proposed algorithm achieves the state-of-art performance with impressive speed (8.8 FPS on ICDAR2015 and 13.3 FPS on MSRA-TD500).

Download Full-text

Multi-Scale Scene Text Detection Based on Convolutional Neural Network

2019 Chinese Automation Congress (CAC) ◽

10.1109/cac48633.2019.8996635 ◽

2019 ◽

Author(s):

Yan-Feng Lu ◽

Ai-Xuan Zhang ◽

Yi Li ◽

Qian-Hui Yu ◽

Hong Qiao

Keyword(s):

Neural Network ◽

Convolutional Neural Network ◽

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

Multi-scale Scene Text Detection via Resolution Transform

2019 IEEE International Conference on Multimedia and Expo (ICME) ◽

10.1109/icme.2019.00174 ◽

2019 ◽

Cited By ~ 1

Author(s):

Peirui Cheng ◽

Weiqiang Wang ◽

Yuanqiang Cai

Keyword(s):

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

R-Net: A Relationship Network for Efficient and Accurate Scene Text Detection

IEEE Transactions on Multimedia ◽

10.1109/tmm.2020.2995290 ◽

2020 ◽

pp. 1-1 ◽

Cited By ~ 2

Author(s):

Yuxin Wang ◽

Hongtao Xie ◽

Zheng-Jun Zha ◽

Youliang Tian ◽

Zilong Fu ◽

...

Keyword(s):

Text Detection ◽

Scene Text Detection ◽

Scene Text ◽

Relationship Network

Download Full-text

Realtime multi-scale scene text detection with scale-based region proposal network

Pattern Recognition ◽

10.1016/j.patcog.2019.107026 ◽

2020 ◽

Vol 98 ◽

pp. 107026 ◽

Cited By ~ 4

Author(s):

Wenhao He ◽

Xu-Yao Zhang ◽

Fei Yin ◽

Zhenbo Luo ◽

Jean-Marc Ogier ◽

...

Keyword(s):

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

Scene text detection based on multi-scale SWT and edge filtering

2016 23rd International Conference on Pattern Recognition (ICPR) ◽

10.1109/icpr.2016.7899707 ◽

2016 ◽

Cited By ~ 2

Author(s):

Yuanyuan Feng ◽

Yonghong Song ◽

Yuanlin Zhang

Keyword(s):

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

Adaptive Multi-Scale HyperNet with Bi-Direction Residual Attention Module for Scene Text Detection

Journal of Information Hiding and Privacy Protection ◽

10.32604/jihpp.2021.017181 ◽

2021 ◽

Vol 3 (2) ◽

pp. 83-89

Author(s):

Junjie Qu ◽

Jin Liu ◽

Chao Yu

Keyword(s):

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

Natural scene text detection by multi-scale adaptive color clustering and non-text filtering

Neurocomputing ◽

10.1016/j.neucom.2016.07.016 ◽

2016 ◽

Vol 214 ◽

pp. 1011-1025 ◽

Cited By ~ 13

Author(s):

Hui Wu ◽

Beiji Zou ◽

Yu-qian Zhao ◽

Zailiang Chen ◽

Chengzhang Zhu ◽

...

Keyword(s):

Text Detection ◽

Natural Scene ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text ◽

Text Filtering ◽

Color Clustering

Download Full-text

Deep Multi-Scale Context Aware Feature Aggregation for Curved Scene Text Detection

IEEE Transactions on Multimedia ◽

10.1109/tmm.2019.2952978 ◽

2020 ◽

Vol 22 (8) ◽

pp. 1969-1984 ◽

Cited By ~ 1

Author(s):

Pengwen Dai ◽

Hua Zhang ◽

Xiaochun Cao

Keyword(s):

Text Detection ◽

Context Aware ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text ◽

Feature Aggregation

Download Full-text

Attention Guided Multi-Scale Regression for Scene Text Detection

Proceedings of the 2020 European Symposium on Software Engineering ◽

10.1145/3393822.3432316 ◽

2020 ◽

Author(s):

Zhiwei Zheng

Keyword(s):

Text Detection ◽

Multi Scale ◽

Scene Text Detection ◽

Scene Text

Download Full-text

Omnidirectional Scene Text Detection with Sequential-free Box Discretization

Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence ◽

10.24963/ijcai.2019/423 ◽

2019 ◽

Cited By ~ 9

Author(s):

Yuliang Liu ◽

Sheng Zhang ◽

Lianwen Jin ◽

Lele Xie ◽

Yaqiang Wu ◽

...

Keyword(s):

State Of The Art ◽

Detection Performance ◽

Text Detection ◽

Detection Methods ◽

Bounding Box ◽

Scene Text Detection ◽

Scene Text ◽

Art Methods ◽

In The Wild ◽

Ablation Study

Scene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection methods. However, recent researches reveal that introducing quadrilateral bounding box for scene text detection will bring a label confusion issue which is easily overlooked, and this issue may significantly undermine the detection performance. To address this issue, in this paper, we propose a novel method called Sequential-free Box Discretization (SBD) by discretizing the bounding box into key edges (KE) which can further derive more effective methods to improve detection performance. Experiments showed that the proposed method can outperform state-of-the-art methods in many popular scene text benchmarks, including ICDAR 2015, MLT, and MSRA-TD500. Ablation study also showed that simply integrating the SBD into Mask R-CNN framework, the detection performance can be substantially improved. Furthermore, an experiment on the general object dataset HRSC2016 (multi-oriented ships) showed that our method can outperform recent state-of-the-art methods by a large margin, demonstrating its powerful generalization ability.

Download Full-text