On‐device audio‐visual multi‐person wake word spotting-Reference-Cited by-同舟云学术

On‐device audio‐visual multi‐person wake word spotting

Published:2023-03 Issue:4 Volume:8 Page:1578-1589
ISSN:2468-2322
Container-title:CAAI Transactions on Intelligence Technology
language:en
Short-container-title:CAAI Trans on Intel Tech

Author:

Li Yidi¹^ORCID,Wang Guoquan¹²,Chen Zhan¹,Tang Hao³,Liu Hong¹

Affiliation:

1. Key Laboratory of Machine Perception Peking University Shenzhen Graduate School Shenzhen China

2. College of Computer and Information Hefei University of Technology Hefei China

3. Computer Vision Lab ETH Zurich Zurich Switzerland

Abstract

AbstractAudio‐visual wake word spotting is a challenging multi‐modal task that exploits visual information of lip motion patterns to supplement acoustic speech to improve overall detection performance. However, most audio‐visual wake word spotting models are only suitable for simple single‐speaker scenarios and require high computational complexity. Further development is hindered by complex multi‐person scenarios and computational limitations in mobile environments. In this paper, a novel audio‐visual model is proposed for on‐device multi‐person wake word spotting. Firstly, an attention‐based audio‐visual voice activity detection module is presented, which generates an attention score matrix of audio and visual representations to derive active speaker representation. Secondly, the knowledge distillation method is introduced to transfer knowledge from the large model to the on‐device model to control the size of our model. Moreover, a new audio‐visual dataset, PKU‐KWS, is collected for sentence‐level multi‐person wake word spotting. Experimental results on the PKU‐KWS dataset show that this approach outperforms the previous state‐of‐the‐art methods.

Publisher

Institution of Engineering and Technology (IET)

Subject

Artificial Intelligence,Computer Networks and Communications,Computer Vision and Pattern Recognition,Human-Computer Interaction,Information Systems

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1049/cit2.12189

Reference51 articles.

1. Customized Wake-Up Word with Key Word Spotting using Convolutional Neural Network

2. Wake-up-word spotting using end-to-end deep neural network system

3. Towards Data-Efficient Modeling for Wake Word Spotting

4. Gao Y. et al.:On front‐end gain invariant modeling for wake word spotting. arXiv preprint arXiv:2010.06676 (2020)

5. Deep spoken keyword spotting: an overview;López‐Espejo I.;IEEE Access,2021