Amharic spoken digits recognition using convolutional neural network

Tewodros Alemu Ayall; Changjun Zhou; Huawen Liu; Getnet Mezgebu Brhanemeskel; Solomon Teferra Abate; Michael Adjeisah

doi:10.1186/s40537-024-00910-z

Journal of Big Data (May 2024)

Amharic spoken digits recognition using convolutional neural network

Tewodros Alemu Ayall,
Changjun Zhou,
Huawen Liu,
Getnet Mezgebu Brhanemeskel,
Solomon Teferra Abate,
Michael Adjeisah

Affiliations

Tewodros Alemu Ayall: School of Computer Science and Technology, Zhejiang Normal University
Changjun Zhou: School of Computer Science and Technology, Zhejiang Normal University
Huawen Liu: Department of Computer Science, Shaoxing University
Getnet Mezgebu Brhanemeskel: School of Information Science, Addis Ababa University
Solomon Teferra Abate: School of Information Science, Addis Ababa University
Michael Adjeisah: School of Computer Science and Technology, Zhejiang Normal University

DOI: https://doi.org/10.1186/s40537-024-00910-z
Journal volume & issue: Vol. 11, no. 1
pp. 1 – 23

Abstract

Read online

Abstract Spoken digits recognition (SDR) is a type of supervised automatic speech recognition, which is required in various human–machine interaction applications. It is utilized in phone-based services like dialing systems, certain bank operations, airline reservation systems, and price extraction. However, the design of SDR is a challenging task that requires the development of labeled audio data, the proper choice of feature extraction method, and the development of the best performing model. Even if several works have been done for various languages, such as English, Arabic, Urdu, etc., there is no developed Amharic spoken digits dataset (AmSDD) to build Amharic spoken digits recognition (AmSDR) model for the Amharic language, which is the official working language of the government of Ethiopia. Therefore, in this study, we developed a new AmSDD that contains 12,000 utterances of 0 (Zaero) to 9 (zet’enyi) digits which were recorded from 120 volunteer speakers of different age groups, genders, and dialects who repeated each digit ten times. Mel frequency cepstral coefficients (MFCCs) and Mel-Spectrogram feature extraction methods were used to extract trainable features from the speech signal. We conducted different experiments on the development of the AmSDR model using the AmSDD and classical supervised learning algorithms such as Linear Discriminant Analysis (LDA), K-Nearest Neighbors (KNN), Support Vector Machine (SVM), and Random Forest (RF) as the baseline. To further improve the performance recognition of AmSDR, we propose a three layers Convolutional Neural Network (CNN) architecture with Batch normalization. The results of our experiments show that the proposed CNN model outperforms the baseline algorithms and scores an accuracy of 99% and 98% using MFCCs and Mel-Spectrogram features, respectively.

Published in Journal of Big Data

ISSN: 2196-1115 (Online)
Publisher: SpringerOpen
Country of publisher: United Kingdom
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Electronics: Computer engineering. Computer hardware; Technology: Technology (General): Industrial engineering. Management engineering: Information technology; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://journalofbigdata.springeropen.com

About the journal

Abstract

Keywords