Developing Core Technologies for Resource-Scarce Nguni Languages

Jakobus S. du Toit; Martin J. Puttkammer

doi:10.3390/info12120520

Information (Dec 2021)

Developing Core Technologies for Resource-Scarce Nguni Languages

Jakobus S. du Toit,
Martin J. Puttkammer

Affiliations

Jakobus S. du Toit: Centre for Text Technology, North-West University, Potchefstroom 2520, South Africa
Martin J. Puttkammer: Centre for Text Technology, North-West University, Potchefstroom 2520, South Africa

DOI: https://doi.org/10.3390/info12120520
Journal volume & issue: Vol. 12, no. 12
p. 520

Abstract

Read online

The creation of linguistic resources is crucial to the continued growth of research and development efforts in the field of natural language processing, especially for resource-scarce languages. In this paper, we describe the curation and annotation of corpora and the development of multiple linguistic technologies for four official South African languages, namely isiNdebele, Siswati, isiXhosa, and isiZulu. Development efforts included sourcing parallel data for these languages and annotating each on token, orthographic, morphological, and morphosyntactic levels. These sets were in turn used to create and evaluate three core technologies, viz. a lemmatizer, part-of-speech tagger, morphological analyzer for each of the languages. We report on the quality of these technologies which improve on previously developed rule-based technologies as part of a similar initiative in 2013. These resources are made publicly accessible through a local resource agency with the intention of fostering further development of both resources and technologies that may benefit the NLP industry in South Africa.

Published in Information

ISSN: 2078-2489 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://www.mdpi.com/journal/information/

About the journal

Abstract

Keywords