On‐device audio‐visual multi‐person wake word spotting

Yidi Li; Guoquan Wang; Zhan Chen; Hao Tang; Hong Liu

doi:10.1049/cit2.12189

CAAI Transactions on Intelligence Technology (Dec 2023)

On‐device audio‐visual multi‐person wake word spotting

Yidi Li,
Guoquan Wang,
Zhan Chen,
Hao Tang,
Hong Liu

Affiliations

Yidi Li: Key Laboratory of Machine Perception Peking University Shenzhen Graduate School Shenzhen China
Guoquan Wang: Key Laboratory of Machine Perception Peking University Shenzhen Graduate School Shenzhen China
Zhan Chen: Key Laboratory of Machine Perception Peking University Shenzhen Graduate School Shenzhen China
Hao Tang: Computer Vision Lab ETH Zurich Zurich Switzerland
Hong Liu: Key Laboratory of Machine Perception Peking University Shenzhen Graduate School Shenzhen China

DOI: https://doi.org/10.1049/cit2.12189
Journal volume & issue: Vol. 8, no. 4
pp. 1578 – 1589

Abstract

Read online

Abstract Audio‐visual wake word spotting is a challenging multi‐modal task that exploits visual information of lip motion patterns to supplement acoustic speech to improve overall detection performance. However, most audio‐visual wake word spotting models are only suitable for simple single‐speaker scenarios and require high computational complexity. Further development is hindered by complex multi‐person scenarios and computational limitations in mobile environments. In this paper, a novel audio‐visual model is proposed for on‐device multi‐person wake word spotting. Firstly, an attention‐based audio‐visual voice activity detection module is presented, which generates an attention score matrix of audio and visual representations to derive active speaker representation. Secondly, the knowledge distillation method is introduced to transfer knowledge from the large model to the on‐device model to control the size of our model. Moreover, a new audio‐visual dataset, PKU‐KWS, is collected for sentence‐level multi‐person wake word spotting. Experimental results on the PKU‐KWS dataset show that this approach outperforms the previous state‐of‐the‐art methods.

Published in CAAI Transactions on Intelligence Technology

ISSN: 2468-2322 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Language and Literature: Philology. Linguistics: Computational linguistics. Natural language processing; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/24682322

About the journal

Abstract

Keywords