Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers-Reference-Cited by-同舟云学术

Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers

Published:2020-04-03 Issue:04 Volume:34 Page:6917-6924
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Zhao Ya,Xu Rui,Wang Xinchao,Hou Peng,Tang Haihong,Song Mingli

Abstract

Lip reading has witnessed unparalleled development in recent years thanks to deep learning and the availability of large-scale datasets. Despite the encouraging results achieved, the performance of lip reading, unfortunately, remains inferior to the one of its counterpart speech recognition, due to the ambiguous nature of its actuations that makes it challenging to extract discriminant features from the lip movement videos. In this paper, we propose a new method, termed as Lip by Speech (LIBS), of which the goal is to strengthen lip reading by learning from speech recognizers. The rationale behind our approach is that the features extracted from speech recognizers may provide complementary and discriminant clues, which are formidable to be obtained from the subtle movements of the lips, and consequently facilitate the training of lip readers. This is achieved, specifically, by distilling multi-granularity knowledge from speech recognizers to lip readers. To conduct this cross-modal knowledge distillation, we utilize an efficacious alignment scheme to handle the inconsistent lengths of the audios and videos, as well as an innovative filtering strategy to refine the speech recognizer's prediction. The proposed method achieves the new state-of-the-art performance on the CMLR and LRS2 datasets, outperforming the baseline by a margin of 7.66% and 2.75% in character error rate, respectively.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 39 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Cross-modal knowledge distillation for continuous sign language recognition;Neural Networks;2024-11

2. Deep Learning for Visual Speech Analysis: A Survey;IEEE Transactions on Pattern Analysis and Machine Intelligence;2024-09

3. Multi-modal co-learning for silent speech recognition based on ultrasound tongue images;Speech Communication;2024-09

4. Cantonese sentence dataset for lip‐reading;IET Image Processing;2024-06-18

5. HNet: A deep learning based hybrid network for speaker dependent visual speech recognition;International Journal of Hybrid Intelligent Systems;2024-06-03