Automated Text line Segmentation and Table detection for Pre-Printed Document Image Analysis Systems

The paper presents the algorithm for text line segmentation based on the oriented anisotropic Gaussian kernel. Initially, the document image is split into connected components achieved by bounding boxes. These connected components are cleared from redundant fragments. Furthermore, the binary moments are applied to each of these connected components evaluating local text skewing. According to this information the orientation of the anisotropic Gaussian kernel is set. After the algorithm application the boundary growing areas around connected components are established. These areas are of major importance for the evaluation of text line segmentation. For testing purposes, the algorithm is evaluated under different text samples. Comparative analysis between algorithm with and without orientation based on the anisotropic Gaussian kernel is made. The results show the improvement in the domain of text line segmentation.

Download Full-text

Text Line Segmentation With Water Flow Algorithm Based on Power Function

Journal of Electrical Engineering ◽

10.2478/jee-2015-0021 ◽

2015 ◽

Vol 66 (3) ◽

pp. 132-141 ◽

Cited By ~ 2

Author(s):

Darko Brodić

Keyword(s):

Power Function ◽

Water Flow ◽

Document Image ◽

Text Line ◽

Flow Function ◽

Basic Algorithm ◽

Process Stage ◽

Flow Algorithm ◽

Text Line Segmentation ◽

Line Segmentation

Abstarct This manuscript proposes an extension to the water flow algorithm for text line segmentation. Basic algorithm assumes hypothetical water flows under few specified angles of the document image frame from left to right and vice versa. As a result, unwetted image regions that incorporate text are extracted. These regions are of the major importance for text line segmentation. The extension of the basic algorithm means modification of water flow function that creates the unwetted region. Hence, the linear water flow function used in the basic algorithm is changed with its power function counterpart. Extended method was tested, examined and evaluated under different text samples. Results are encouraging due to improving text line segmentation which is a key process stage.

Download Full-text

Experimental Application of a Japanese Historical Document Image Synthesis Method to Text Line Segmentation

Proceedings of the 10th International Conference on Pattern Recognition Applications and Methods ◽

10.5220/0010330206280634 ◽

2021 ◽

Author(s):

Naoto Inuzuka ◽

Tetsuya Suzuki

Keyword(s):

Synthesis Method ◽

Image Synthesis ◽

Document Image ◽

Text Line ◽

Historical Document ◽

Experimental Application ◽

Text Line Segmentation ◽

Line Segmentation

Download Full-text

Text Line Detection in Multicolumn for Indian Scripts Using Histogram: A Document Image Analysis Application

International Conference on Advanced Computer Theory and Engineering (ICACTE 2009) ◽

10.1115/1.802977.paper18 ◽

2009 ◽

pp. 161-168

Keyword(s):

Image Analysis ◽

Document Image ◽

Line Detection ◽

Text Line ◽

Document Image Analysis ◽

Analysis Application

Download Full-text

A novel method of text line segmentation for historical document image of the uchen Tibetan

Journal of Visual Communication and Image Representation ◽

10.1016/j.jvcir.2019.01.021 ◽

2019 ◽

Vol 61 ◽

pp. 23-32 ◽

Cited By ~ 2

Author(s):

Zhenjiang Li ◽

Weilan Wang ◽

Yang Chen ◽

Yusheng Hao

Keyword(s):

Document Image ◽

Text Line ◽

Historical Document ◽

Novel Method ◽

Text Line Segmentation ◽

Line Segmentation

Download Full-text

Enabling Text-Line Segmentation in Run-Length Encoded Handwritten Document Image Using Entropy-Driven Incremental Learning

Proceedings of 3rd International Conference on Computer Vision and Image Processing - Advances in Intelligent Systems and Computing ◽

10.1007/978-981-32-9088-4_20 ◽

2019 ◽

pp. 233-245

Author(s):

R. Amarnath ◽

P. Nagabhushan ◽

Mohammed Javed

Keyword(s):

Incremental Learning ◽

Document Image ◽

Text Line ◽

Run Length ◽

Handwritten Document ◽

Text Line Segmentation ◽

Line Segmentation

Download Full-text

A fast multiresolution text line and non text-line structures extraction and discrimination scheme for document image analysis

Proceedings of 1st International Conference on Image Processing ◽

10.1109/icip.1994.413290 ◽

2002 ◽

Cited By ~ 9

Author(s):

O. Deforges ◽

D. Barba

Keyword(s):

Image Analysis ◽

Document Image ◽

Text Line ◽

Document Image Analysis

Download Full-text

DATASET AND GROUND TRUTH FOR HANDWRITTEN TEXT IN FOUR DIFFERENT SCRIPTS

International Journal of Pattern Recognition and Artificial Intelligence ◽

10.1142/s0218001412530011 ◽

2012 ◽

Vol 26 (04) ◽

pp. 1253001 ◽

Cited By ~ 19

Author(s):

ALIREZA ALAEI ◽

UMAPADA PAL ◽

P. NAGABHUSHAN

Keyword(s):

Ground Truth ◽

Document Image ◽

Text Line ◽

Document Recognition ◽

Handwritten Text ◽

Latin Script ◽

Handwritten Document ◽

Text Line Segmentation ◽

Line Segmentation ◽

Content Information

In document image analysis (DIA) especially in handwritten document recognition, standard databases play significant roles for evaluating performances of algorithms and comparing results obtained by different groups of researchers. The field of DIA regard to Indo-Persian documents is still at its infancy compared to Latin script-based documents; as such standard datasets are not still available in literature. This paper is an effort towards alleviating this gap. In this paper, an unconstrained handwritten dataset containing documents of Persian, Bangla, Oriya and Kannada (PBOK) is introduced. The PBOK contains 707 text-pages written in four different languages (Persian, Bangla, Oriya and Kannada) by 436 individuals. Total number of text-lines, words/subwords and characters are 12,565, 104,541 and 423,980, respectively. In most documents of PBOK dataset contain either an overlapping or a touching text-lines. The average number of text-lines in text-pages of the PBOK dataset is 18. Two types of ground truths, based on pixels information and content information, are generated for the dataset. Because of such ground truths, the PBOK dataset can be utilized in many areas of document image processing e.g. text-line segmentation, word segmentation and word recognition. To provide an insight for other researches, recent text-line segmentation results on this dataset are also reported.

Download Full-text

A Novel Text Line Segmentation Method Based on Contour Curve Tracking for Tibetan Historical Documents

International Journal of Pattern Recognition and Artificial Intelligence ◽

10.1142/s0218001418540253 ◽

2018 ◽

Vol 32 (10) ◽

pp. 1854025 ◽

Cited By ~ 3

Author(s):

Fengming Zhou ◽

Weilan Wang ◽

Qiang Lin

Keyword(s):

Image Data ◽

Document Image ◽

Connected Components ◽

Data Sets ◽

Contour Tracking ◽

Text Line ◽

Segmentation Method ◽

Connected Component ◽

Text Line Segmentation ◽

Line Segmentation

In this paper, we proposed a novel method for text line segmentation of Tibetan historical document image with uchen script based on contour tracking. Our method is mainly to segment the text lines from the image documents using the contour curve of the text lines, which consists of three parts: First, we calculate the barycentre coordinates of the connected components for the text regions, and then the barycentre of each text line is connected in order, so that the main part of each text line is connected and a new connected component is formed; then the contour curve of the connected component is obtained using the contour tracing algorithm; Second, the contour curve and the barycentre gravity are used to assign key elements (such as the syllable point, the upper vowel, the lower vowel, and the broken strokes and so on) of the text lines, and next the candidate text lines are obtained based on these connected components; Finally, the contour tracking algorithm is used to calculate the contour curve of the candidate text lines and segment the text lines. We evaluated our text line segmentation method on the 200 document image data sets. Experimental results show that the proposed method based on contour curve tracing can accurately segment the text lines of image documents and achieve the encouraging results.

Download Full-text

A two-step framework for text line segmentation in historical Arabic and Latin document images

International Journal on Document Analysis and Recognition (IJDAR) ◽

10.1007/s10032-021-00377-1 ◽

2021 ◽

Author(s):

Olfa Mechi ◽

Maroua Mehri ◽

Rolf Ingold ◽

Najoua Essoukri Ben Amara

Keyword(s):

Text Line ◽

Document Images ◽

Text Line Segmentation ◽

Line Segmentation

Download Full-text