Title
An Adaptive Text-Line Extraction Algorithm For Printed Arabic Documents With Diacritics
Abstract
The performance of document text recognition depends on text line segmentation algorithms, which heavily relies on the type of language, author's writing style, pen type, and document quality. In this paper, we present a novel unsupervised text-line segmentation algorithm for printed Arabic documents with and without diacritics. The presented approach employs a projection profile along with connected components in an iterative manner to detect text-lines. The primary benefits of the presented algorithm are (i) it is not threshold dependent, (ii) it is not required a training phase for threshold selection, and (iii) it is robust towards page rotation, font type, size, and style variation for both with and without diacritics documents. The extensive computational simulations on manually collected dataset prove the efficiency of the proposed scheme compared with several baseline and states of the art methods, including, Voronoi, X-Y Cut, Docstrum, Smearing and Seam-carving methods. Computational time analysis also presented.
Year
DOI
Venue
2021
10.1007/s11042-020-09737-1
MULTIMEDIA TOOLS AND APPLICATIONS
Keywords
DocType
Volume
Arabic character recognition, Line segmentation, Baseline, Diacritics
Journal
80
Issue
ISSN
Citations 
2
1380-7501
0
PageRank 
References 
Authors
0.34
0
5
Name
Order
Citations
PageRank
Khader Mohammad1135.22
Aziz Qaroush2106.86
Mahdi Washha300.34
Sos S. Agaian474483.01
Iyad Tumar5224.77