A Novel Text Line Segmentation Method Based on Contour Curve Tracking for Tibetan Historical Documents

Author(s):  
Fengming Zhou ◽  
Weilan Wang ◽  
Qiang Lin

In this paper, we proposed a novel method for text line segmentation of Tibetan historical document image with uchen script based on contour tracking. Our method is mainly to segment the text lines from the image documents using the contour curve of the text lines, which consists of three parts: First, we calculate the barycentre coordinates of the connected components for the text regions, and then the barycentre of each text line is connected in order, so that the main part of each text line is connected and a new connected component is formed; then the contour curve of the connected component is obtained using the contour tracing algorithm; Second, the contour curve and the barycentre gravity are used to assign key elements (such as the syllable point, the upper vowel, the lower vowel, and the broken strokes and so on) of the text lines, and next the candidate text lines are obtained based on these connected components; Finally, the contour tracking algorithm is used to calculate the contour curve of the candidate text lines and segment the text lines. We evaluated our text line segmentation method on the 200 document image data sets. Experimental results show that the proposed method based on contour curve tracing can accurately segment the text lines of image documents and achieve the encouraging results.

2013 ◽  
Vol 64 (4) ◽  
pp. 238-243 ◽  
Author(s):  
Darko Brodić ◽  
Zoran N. Milivojević

The paper presents the algorithm for text line segmentation based on the oriented anisotropic Gaussian kernel. Initially, the document image is split into connected components achieved by bounding boxes. These connected components are cleared from redundant fragments. Furthermore, the binary moments are applied to each of these connected components evaluating local text skewing. According to this information the orientation of the anisotropic Gaussian kernel is set. After the algorithm application the boundary growing areas around connected components are established. These areas are of major importance for the evaluation of text line segmentation. For testing purposes, the algorithm is evaluated under different text samples. Comparative analysis between algorithm with and without orientation based on the anisotropic Gaussian kernel is made. The results show the improvement in the domain of text line segmentation.


2015 ◽  
Vol 66 (3) ◽  
pp. 132-141 ◽  
Author(s):  
Darko Brodić

Abstarct This manuscript proposes an extension to the water flow algorithm for text line segmentation. Basic algorithm assumes hypothetical water flows under few specified angles of the document image frame from left to right and vice versa. As a result, unwetted image regions that incorporate text are extracted. These regions are of the major importance for text line segmentation. The extension of the basic algorithm means modification of water flow function that creates the unwetted region. Hence, the linear water flow function used in the basic algorithm is changed with its power function counterpart. Extended method was tested, examined and evaluated under different text samples. Results are encouraging due to improving text line segmentation which is a key process stage.


Sign in / Sign up

Export Citation Format

Share Document