Text Line Extraction in Historical Documents Using Mask R-CNN

Ahmad Droby; Berat Kurar Barakat; Reem Alaasam; Boraq Madi; Irina Rabaev; Jihad El-Sana

doi:10.3390/signals3030032

Signals (Aug 2022)

Text Line Extraction in Historical Documents Using Mask R-CNN

Ahmad Droby,
Berat Kurar Barakat,
Reem Alaasam,
Boraq Madi,
Irina Rabaev,
Jihad El-Sana

Affiliations

Ahmad Droby: Department of Computer Science, Ben-Gurion University of the Negev, Be’er Sheva 8410501, Israel
Berat Kurar Barakat: Department of Computer Science, Ben-Gurion University of the Negev, Be’er Sheva 8410501, Israel
Reem Alaasam: Department of Computer Science, Ben-Gurion University of the Negev, Be’er Sheva 8410501, Israel
Boraq Madi: Department of Computer Science, Ben-Gurion University of the Negev, Be’er Sheva 8410501, Israel
Irina Rabaev: Department of Software Engineering, Shamoon College of Engineering, Be’er Sheva 8410802, Israel
Jihad El-Sana: Department of Computer Science, Ben-Gurion University of the Negev, Be’er Sheva 8410501, Israel

DOI: https://doi.org/10.3390/signals3030032
Journal volume & issue: Vol. 3, no. 3
pp. 535 – 549

Abstract

Read online

Text line extraction is an essential preprocessing step in many handwritten document image analysis tasks. It includes detecting text lines in a document image and segmenting the regions of each detected line. Deep learning-based methods are frequently used for text line detection. However, only a limited number of methods tackle the problems of detection and segmentation together. This paper proposes a holistic method that applies Mask R-CNN for text line extraction. A Mask R-CNN model is trained to extract text lines fractions from document patches, which are further merged to form the text lines of an entire page. The presented method was evaluated on the two well-known datasets of historical documents, DIVA-HisDB and ICDAR 2015-HTR, and achieved state-of-the-art results. In addition, we introduce a new challenging dataset of Arabic historical manuscripts, VML-AHTE, where numerous diacritics are present. We show that the presented Mask R-CNN-based method can successfully segment text lines, even in such a challenging scenario.

Published in Signals

ISSN: 2624-6120 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Applied mathematics. Quantitative methods
Website: https://www.mdpi.com/journal/signals

About the journal

Abstract

Keywords