ITM Web of Conferences (Jan 2019)

Continuous Speech Recognition of Kazakh Language

  • Mamyrbayev Оrken,
  • Turdalyuly Mussa,
  • Mekebayev Nurbapa,
  • Mukhsina Kuralay,
  • Keylan Alimukhan,
  • BabaAli Bagher,
  • Nabieva Gulnaz,
  • Duisenbayeva Aigerim,
  • Akhmetov Bekturgan

DOI
https://doi.org/10.1051/itmconf/20192401012
Journal volume & issue
Vol. 24
p. 01012

Abstract

Read online

This article describes the methods of creating a system of recognizing the continuous speech of Kazakh language. Studies on recognition of Kazakh speech in comparison with other languages began relatively recently, that is after obtaining independence of the country, and belongs to low resource languages. A large amount of data is required to create a reliable system and evaluate it accurately. A database has been created for the Kazakh language, consisting of a speech signal and corresponding transcriptions. The continuous speech has been composed of 200 speakers of different genders and ages, and the pronunciation vocabulary of the selected language. Traditional models and deep neural networks have been used to train the system. As a result, a word error rate (WER) of 30.01% has been obtained.