DeepVoCoder: A CNN Model for Compression and Coding of Narrow Band Speech

Hacer Yalim Keles; Jan Rozhon; H. Gokhan Ilk; Miroslav Voznak

doi:10.1109/ACCESS.2019.2920663

IEEE Access (Jan 2019)

DeepVoCoder: A CNN Model for Compression and Coding of Narrow Band Speech

Hacer Yalim Keles,
Jan Rozhon,
H. Gokhan Ilk,
Miroslav Voznak

Affiliations

Hacer Yalim Keles: Computer Engineering Department, Ankara University, Ankara, Turkey
Jan Rozhon: Department of Telecommunications, Faculty of Electrical Engineering and Computer Science, VSB–Technical University of Ostrava, Ostrava, Czech Republic
H. Gokhan Ilk: ORCiD; IT4 Innovations, VSB–Technical University of Ostrava, Ostrava, Czech Republic
Miroslav Voznak: Department of Telecommunications, Faculty of Electrical Engineering and Computer Science, VSB–Technical University of Ostrava, Ostrava, Czech Republic

DOI: https://doi.org/10.1109/ACCESS.2019.2920663
Journal volume & issue: Vol. 7
pp. 75081 – 75089

Abstract

Read online

This paper proposes a convolutional neural network (CNN)-based encoder model to compress and code speech signal directly from raw input speech. Although the model can synthesize wideband speech by implicit bandwidth extension, narrowband is preferred for IP telephony and telecommunications purposes. The model takes time domain speech samples as inputs and encodes them using a cascade of convolutional filters in multiple layers, where pooling is applied after some layers to downsample the encoded speech by half. The final bottleneck layer of the CNN encoder provides an abstract and compact representation of the speech signal. In this paper, it is demonstrated that this compact representation is sufficient to reconstruct the original speech signal in high quality using the CNN decoder. This paper also discusses the theoretical background of why and how CNN may be used for end-to-end speech compression and coding. The complexity, delay, memory requirements, and bit rate versus quality are discussed in the experimental results.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords