Application of deep learning in Mandarin Chinese lip-reading recognition-Reference-Cited by-同舟云学术

Application of deep learning in Mandarin Chinese lip-reading recognition

Published:2023-09-05 Issue:1 Volume:2023 Page:
ISSN:1687-1499
Container-title:EURASIP Journal on Wireless Communications and Networking
language:en
Short-container-title:J Wireless Com Network

Author:

Xing Guangxin^ORCID,Han Lingkun,Zheng Yelong,Zhao Meirong

Abstract

AbstractLip-reading is an emerging technology in recent years, and it can be applied to the field of language recovery, criminal investigation, identity authentication, etc. We aim to recognize what the speaker is saying without audio but only video. Because of the different mouth shapes and the influence of homophones, the current Mandarin Chinese lip-reading network is proposed, an end-to-end model based on long short-term memory (LSTM) encoder-decoder architecture. The model incorporates the LSTM encoder-decode architecture, the spatiotemporal convolutional neural network (STCNN), Word2Vec, and the Attention model. The STCNN captures continuously encoded motion information, Word2Vec converts words into word vectors for feature encoding, and the Attention model assigns weights to the target words. Based on the video dataset we built, we completed training and testing. Experiments have proved that the accuracy of the Mandarin Chinese lip-reading model is about 72%. Therefore, MCLRN can be used to identify the words spoken by the speaker.

Funder

Joint Fund of the Ministry of Education for Equipment Pre research

National Key Research and Development Program of China

Publisher

Springer Science and Business Media LLC

Subject

Computer Networks and Communications,Computer Science Applications,Signal Processing

Link

https://link.springer.com/content/pdf/10.1186/s13638-023-02283-y.pdf

Reference31 articles.

1. X. Chen, J. Du, H. Zhang, Lipreading with DenseNet and resBi-LSTM. Signal image video 14(5), 981–989 (2020). https://doi.org/10.1007/s11760-019-01630-1

2. X. Zhao, S. Yang, S. Shan, X. Chen, Mutual information maximization for effective lip reading, in Proceedings of IEEE International Conference on Automatic Face and Gesture Recognition (2020), pp. 420–427. https://doi.org/10.1109/FG47880.2020.00133

3. S. Jiang, H. Ruan, Z. Wang, H. Zhang, H. Zhao, L. Li, Microwave lip reading of chinese mandarin based on programmable metasurface, in Proceedings of IEEE MTT-S International Microwave Workshop Series on Advanced Materials and Processes for RF and THz Applications (2021), pp. 376–378. https://doi.org/10.1109/IMWS-AMP53428.2021.9643862

4. Ü. Atila, F. Sabaz, Turkish lip-reading using Bi-LSTM and deep learning models. Eng. Sci. Technol. Int. J. 35, 101206 (2022). https://doi.org/10.1016/j.jestch.2022.101206

5. J. Xiao, S. Yang, Y. Zhang, S. Shan, X. Chen, Deformation flow based two-stream network for lip reading, in International Conference on Automatic Face and Gesture Recognition. (2020), pp. 364–370. https://doi.org/10.1109/FG47880.2020.00132

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Script Generation for Silent Speech in E-Learning;Advances in Educational Technologies and Instructional Design;2024-06-03

2. Retraction Note: Application of deep learning in Mandarin Chinese lip-reading recognition;EURASIP Journal on Wireless Communications and Networking;2024-05-21

3. AI LipReader-Transcribing Speech from Lip Movements;2024 International Conference on Emerging Smart Computing and Informatics (ESCI);2024-03-05