Pronunciation augmentation for Mandarin-English code-switching speech recognition

Yanhua Long; Shuang Wei; Jie Lian; Yijie Li

doi:10.1186/s13636-021-00222-7

EURASIP Journal on Audio, Speech, and Music Processing (Aug 2021)

Pronunciation augmentation for Mandarin-English code-switching speech recognition

Yanhua Long,
Shuang Wei,
Jie Lian,
Yijie Li

Affiliations

Yanhua Long: SHNU-Unisound Joint Laboratory of Natural Human-Computer Interaction, Shanghai Engineering Research Center of Intelligent Education and Bigdata, Shanghai Normal University
Shuang Wei: SHNU-Unisound Joint Laboratory of Natural Human-Computer Interaction, Shanghai Engineering Research Center of Intelligent Education and Bigdata, Shanghai Normal University
Jie Lian: SHNU-Unisound Joint Laboratory of Natural Human-Computer Interaction, Shanghai Engineering Research Center of Intelligent Education and Bigdata, Shanghai Normal University
Yijie Li: Unisound AI Technology Co., Ltd.

DOI: https://doi.org/10.1186/s13636-021-00222-7
Journal volume & issue: Vol. 2021, no. 1
pp. 1 – 14

Abstract

Read online

Abstract Code-switching (CS) refers to the phenomenon of using more than one language in an utterance, and it presents great challenge to automatic speech recognition (ASR) due to the code-switching property in one utterance, the pronunciation variation phenomenon of the embedding language words and the heavy training data sparse problem. This paper focuses on the Mandarin-English CS ASR task. We aim at dealing with the pronunciation variation and alleviating the sparse problem of code-switches by using pronunciation augmentation methods. An English-to-Mandarin mix-language phone mapping approach is first proposed to obtain a language-universal CS lexicon. Based on this lexicon, an acoustic data-driven lexicon learning framework is further proposed to learn new pronunciations to cover the accents, mis-pronunciations, or pronunciation variations of those embedding English words. Experiments are performed on real CS ASR tasks. Effectiveness of the proposed methods are examined on all of the conventional, hybrid, and the recent end-to-end speech recognition systems. Experimental results show that both the learned phone mapping and augmented pronunciations can significantly improve the performance of code-switching speech recognition.

Published in EURASIP Journal on Audio, Speech, and Music Processing

ISSN: 1687-4722 (Online)
Publisher: SpringerOpen
Country of publisher: United Kingdom
LCC subjects: Science: Physics: Acoustics. Sound; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://asmp-eurasipjournals.springeropen.com

About the journal

Abstract

Keywords