Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech

Yeunju Choi; Youngmoon Jung; Youngjoo Suh; Hoirin Kim

doi:10.1109/ACCESS.2022.3175810

IEEE Access (Jan 2022)

Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech

Yeunju Choi,
Youngmoon Jung,
Youngjoo Suh,
Hoirin Kim

Affiliations

Yeunju Choi: ORCiD; School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
Youngmoon Jung: ORCiD; School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
Youngjoo Suh: Voice Group, Konan Technology Inc., Seoul, South Korea
Hoirin Kim: ORCiD; School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea

DOI: https://doi.org/10.1109/ACCESS.2022.3175810
Journal volume & issue: Vol. 10
pp. 52621 – 52629

Abstract

Read online

Although recent neural text-to-speech (TTS) systems have achieved high-quality speech synthesis, there are cases where a TTS system generates low-quality speech, mainly caused by limited training data or information loss during knowledge distillation. Therefore, we propose a novel method to improve speech quality by training a TTS model under the supervision of perceptual loss, which measures the distance between the maximum possible speech quality score and the predicted one. We first pre-train a mean opinion score (MOS) prediction model and then train a TTS model to maximize the MOS of synthesized speech using the pre-trained MOS prediction model. The proposed method can be applied independently regardless of the TTS model architecture or the cause of speech quality degradation and efficiently without increasing the inference time or model complexity. The evaluation results for the MOS and phone error rate demonstrate that our proposed approach improves previous models in terms of both naturalness and intelligibility.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords