Adapting multilingual vision language transformers for low-resource Urdu optical character recognition (OCR)

Musa Dildar Ahmed Cheema; Mohammad Daniyal Shaiq; Farhaan Mirza; Ali Kamal; M. Asif Naeem

doi:10.7717/peerj-cs.1964

PeerJ Computer Science (Apr 2024)

Adapting multilingual vision language transformers for low-resource Urdu optical character recognition (OCR)

Musa Dildar Ahmed Cheema,
Mohammad Daniyal Shaiq,
Farhaan Mirza,
Ali Kamal,
M. Asif Naeem

Affiliations

Musa Dildar Ahmed Cheema: Department of Artificial Intelligence and Data Science, National University of Computer and Emerging Sciences, Islamabad, Pakistan
Mohammad Daniyal Shaiq: Department of Artificial Intelligence and Data Science, National University of Computer and Emerging Sciences, Islamabad, Pakistan
Farhaan Mirza: School of Computer, Engineering and Mathematical Sciences, Auckland University of Technology, Auckland, New Zealand
Ali Kamal: Department of Artificial Intelligence and Data Science, National University of Computer and Emerging Sciences, Islamabad, Pakistan
M. Asif Naeem: Department of Artificial Intelligence and Data Science, National University of Computer and Emerging Sciences, Islamabad, Pakistan

DOI: https://doi.org/10.7717/peerj-cs.1964
Journal volume & issue: Vol. 10
p. e1964

Abstract

Read online Read online

In the realm of digitizing written content, the challenges posed by low-resource languages are noteworthy. These languages, often lacking in comprehensive linguistic resources, require specialized attention to develop robust systems for accurate optical character recognition (OCR). This article addresses the significance of focusing on such languages and introduces ViLanOCR, an innovative bilingual OCR system tailored for Urdu and English. Unlike existing systems, which struggle with the intricacies of low-resource languages, ViLanOCR leverages advanced multilingual transformer-based language models to achieve superior performances. The proposed approach is evaluated using the character error rate (CER) metric and achieves state-of-the-art results on the Urdu UHWR dataset, with a CER of 1.1%. The experimental results demonstrate the effectiveness of the proposed approach, surpassing state of the-art baselines in Urdu handwriting digitization.

Published in PeerJ Computer Science

ISSN: 2376-5992 (Online)
Publisher: PeerJ Inc.
Country of publisher: United States
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://peerj.com/computer-science/

About the journal

Abstract

Keywords