On Training Targets and Activation Functions for Deep Representation Learning in Text-Dependent Speaker Verification

Achintya Kumar Sarkar; Zheng-Hua Tan

doi:10.3390/acoustics5030042

Acoustics (Jul 2023)

On Training Targets and Activation Functions for Deep Representation Learning in Text-Dependent Speaker Verification

Achintya Kumar Sarkar,
Zheng-Hua Tan

Affiliations

Achintya Kumar Sarkar: Indian Institute of Information Technology, Sri City/Chittoor 517646, India
Zheng-Hua Tan: Department of Electronic Systems, Aalborg University, 9220 Aalborg, Denmark

DOI: https://doi.org/10.3390/acoustics5030042
Journal volume & issue: Vol. 5, no. 3
pp. 693 – 713

Abstract

Read online

Deep representation learning has gained significant momentum in advancing text-dependent speaker verification (TD-SV) systems. When designing deep neural networks (DNN) for extracting bottleneck (BN) features, the key considerations include training targets, activation functions, and loss functions. In this paper, we systematically study the impact of these choices on the performance of TD-SV. For training targets, we consider speaker identity, time-contrastive learning (TCL), and auto-regressive prediction coding, with the first being supervised and the last two being self-supervised. Furthermore, we study a range of loss functions when speaker identity is used as the training target. With regard to activation functions, we study the widely used sigmoid function, rectified linear unit (ReLU), and Gaussian error linear unit (GELU). We experimentally show that GELU is able to reduce the error rates of TD-SV significantly compared to sigmoid, irrespective of the training target. Among the three training targets, TCL performs the best. Among the various loss functions, cross-entropy, joint-softmax, and focal loss functions outperform the others. Finally, the score-level fusion of different systems is also able to reduce the error rates. To evaluate the representation learning methods, experiments are conducted on the RedDots 2016 challenge database consisting of short utterances for TD-SV systems based on classic Gaussian mixture model-universal background model (GMM-UBM) and i-vector methods.

Published in Acoustics

ISSN: 2624-599X (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science: Physics
Website: https://www.mdpi.com/journal/acoustics

About the journal

Abstract

Keywords