A Method of Speech Coding for Speech Recognition Using a Convolutional Neural Network

Mariusz Kubanek; Janusz Bobulski; Joanna Kulawik

doi:10.3390/sym11091185

Symmetry (Sep 2019)

A Method of Speech Coding for Speech Recognition Using a Convolutional Neural Network

Mariusz Kubanek,
Janusz Bobulski,
Joanna Kulawik

Affiliations

Mariusz Kubanek: Faculty of Mechanical Engineering and Computer Science, Institute of Computer and Information Sciences, Czestochowa University of Technology, Dabrowskiego 73, 42-201 Czestochowa, Poland
Janusz Bobulski: Faculty of Mechanical Engineering and Computer Science, Institute of Computer and Information Sciences, Czestochowa University of Technology, Dabrowskiego 73, 42-201 Czestochowa, Poland
Joanna Kulawik: Faculty of Mechanical Engineering and Computer Science, Institute of Computer and Information Sciences, Czestochowa University of Technology, Dabrowskiego 73, 42-201 Czestochowa, Poland

DOI: https://doi.org/10.3390/sym11091185
Journal volume & issue: Vol. 11, no. 9
p. 1185

Abstract

Read online

This work presents a new approach to speech recognition, based on the specific coding of time and frequency characteristics of speech. The research proposed the use of convolutional neural networks because, as we know, they show high resistance to cross-spectral distortions and differences in the length of the vocal tract. Until now, two layers of time convolution and frequency convolution were used. A novel idea is to weave three separate convolution layers: traditional time convolution and the introduction of two different frequency convolutions (mel-frequency cepstral coefficients (MFCC) convolution and spectrum convolution). This application takes into account more details contained in the tested signal. Our idea assumes creating patterns for sounds in the form of RGB (Red, Green, Blue) images. The work carried out research for isolated words and continuous speech, for neural network structure. A method for dividing continuous speech into syllables has been proposed. This method can be used for symmetrical stereo sound.

Published in Symmetry

ISSN: 2073-8994 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science: Mathematics
Website: http://www.mdpi.com/journal/symmetry/

About the journal

Abstract

Keywords