Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices-Reference-Cited by-同舟云学术

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices

Published:2023-02-17 Issue:4 Volume:23 Page:2284
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Ryumin Dmitry¹^ORCID,Ivanko Denis¹^ORCID,Ryumina Elena¹^ORCID

Affiliation:

1. St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia

Abstract

Audio-visual speech recognition (AVSR) is one of the most promising solutions for reliable speech recognition, particularly when audio is corrupted by noise. Additional visual information can be used for both automatic lip-reading and gesture recognition. Hand gestures are a form of non-verbal communication and can be used as a very important part of modern human–computer interaction systems. Currently, audio and video modalities are easily accessible by sensors of mobile devices. However, there is no out-of-the-box solution for automatic audio-visual speech and gesture recognition. This study introduces two deep neural network-based model architectures: one for AVSR and one for gesture recognition. The main novelty regarding audio-visual speech recognition lies in fine-tuning strategies for both visual and acoustic features and in the proposed end-to-end model, which considers three modality fusion approaches: prediction-level, feature-level, and model-level. The main novelty in gesture recognition lies in a unique set of spatio-temporal features, including those that consider lip articulation information. As there are no available datasets for the combined task, we evaluated our methods on two different large-scale corpora—LRW and AUTSL—and outperformed existing methods on both audio-visual speech recognition and gesture recognition tasks. We achieved AVSR accuracy for the LRW dataset equal to 98.76% and gesture recognition rate for the AUTSL dataset equal to 98.56%. The results obtained demonstrate not only the high performance of the proposed methodology, but also the fundamental possibility of recognizing audio-visual speech and gestures by sensors of mobile devices.

Funder

Russian Science Foundation

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/23/4/2284/pdf

Reference146 articles.

1. Miao, Z., Liu, H., and Yang, B. (2020, January 11–14). Part-based Lipreading for Audio-Visual Speech Recognition. Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics (SMC), IEEE, Toronto, ON, Canada.

2. Bayesian Feature Enhancement using Independent Vector Analysis and Reverberation Parameter Re-Estimation for Noisy Reverberant Speech Recognition;Cho;Comput. Speech Lang.,2017

3. Yu, W., Zeiler, S., and Kolossa, D. (2021, January 6–11). Fusing Information Streams in End-to-End Audio-Visual Speech Recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Toronto, ON, Canada.