Continuous Speech Recognition of Kazakh Language

Mamyrbayev Оrken; Turdalyuly Mussa; Mekebayev Nurbapa; Mukhsina Kuralay; Keylan Alimukhan; BabaAli Bagher; Nabieva Gulnaz; Duisenbayeva Aigerim; Akhmetov Bekturgan

doi:10.1051/itmconf/20192401012

ITM Web of Conferences (Jan 2019)

Continuous Speech Recognition of Kazakh Language

Mamyrbayev Оrken,
Turdalyuly Mussa,
Mekebayev Nurbapa,
Mukhsina Kuralay,
Keylan Alimukhan,
BabaAli Bagher,
Nabieva Gulnaz,
Duisenbayeva Aigerim,
Akhmetov Bekturgan

Affiliations

Mamyrbayev Оrken: Institut of Information and Computational Technology
Turdalyuly Mussa: Institut of Information and Computational Technology
Mekebayev Nurbapa: Information Technology Department
Mukhsina Kuralay: Information Technology Department
Keylan Alimukhan: Institut of Information and Computational Technology
BabaAli Bagher: Institut of Information and Computational Technology
Nabieva Gulnaz: Institut of Information and Computational Technology
Duisenbayeva Aigerim: Information Technology Department
Akhmetov Bekturgan: Institut of Information and Computational Technology

DOI: https://doi.org/10.1051/itmconf/20192401012
Journal volume & issue: Vol. 24
p. 01012

Abstract

Read online

This article describes the methods of creating a system of recognizing the continuous speech of Kazakh language. Studies on recognition of Kazakh speech in comparison with other languages began relatively recently, that is after obtaining independence of the country, and belongs to low resource languages. A large amount of data is required to create a reliable system and evaluate it accurately. A database has been created for the Kazakh language, consisting of a speech signal and corresponding transcriptions. The continuous speech has been composed of 200 speakers of different genders and ages, and the pronunciation vocabulary of the selected language. Traditional models and deep neural networks have been used to train the system. As a result, a word error rate (WER) of 30.01% has been obtained.

Published in ITM Web of Conferences

ISSN: 2271-2097 (Online)
Publisher: EDP Sciences
Country of publisher: France
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://www.itm-conferences.org/

About the journal