The Design of a Script Identification Algorithm and Its Application in Constructing a Text Language Identification Dataset

Mamtimin Qasim; Wushour Silamu; Minghui Qiu

doi:10.3390/data9110134

Data (Nov 2024)

The Design of a Script Identification Algorithm and Its Application in Constructing a Text Language Identification Dataset

Mamtimin Qasim,
Wushour Silamu,
Minghui Qiu

Affiliations

Mamtimin Qasim: School of Information Technology and Engineering, Guangzhou College of Commerce, Guangzhou 511363, China
Wushour Silamu: School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China
Minghui Qiu: School of Information Technology and Engineering, Guangzhou College of Commerce, Guangzhou 511363, China

DOI: https://doi.org/10.3390/data9110134
Journal volume & issue: Vol. 9, no. 11
p. 134

Abstract

Read online

Script identification is easier to implement than language identification, and its identification rate is very high. The fewer languages are identified when using a language identification algorithm, the higher the identification rate is. However, no systematic study on SI involving multiple languages and determining how to construct relevant language identification datasets has been conducted. Therefore, in this paper, we discuss and design a script identification algorithm and the construction of a language identification dataset based on script groups. The data sources in this paper comprise 261 different languages’ text corpora from the Leipzig Corpora Collection, which are grouped into 23 different script groups. In the Unicode encoding scheme, different scripts are arranged into different code regions. Based on this feature, we propose a written script identification algorithm based on regular expression matching, the micro F-score of which reaches 0.9929 in sentence-level script identification experiments. To reduce noise when constructing the language identification dataset for each script, a script identification algorithm is used to filter out other-script content in each text.

Published in Data

ISSN: 2306-5729 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Bibliography. Library science. Information resources
Website: http://www.mdpi.com/journal/data

About the journal

Abstract

Keywords