A Review of Recent Advances on Deep Learning Methods for Audio-Visual Speech Recognition-Reference-Cited by-同舟云学术

A Review of Recent Advances on Deep Learning Methods for Audio-Visual Speech Recognition

Published:2023-06-12 Issue:12 Volume:11 Page:2665
ISSN:2227-7390
Container-title:Mathematics
language:en
Short-container-title:Mathematics

Author:

Ivanko Denis¹^ORCID,Ryumin Dmitry¹^ORCID,Karpov Alexey¹^ORCID

Affiliation:

1. St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia

Abstract

This article provides a detailed review of recent advances in audio-visual speech recognition (AVSR) methods that have been developed over the last decade (2013–2023). Despite the recent success of audio speech recognition systems, the problem of audio-visual (AV) speech decoding remains challenging. In comparison to the previous surveys, we mainly focus on the important progress brought with the introduction of deep learning (DL) to the field and skip the description of long-known traditional “hand-crafted” methods. In addition, we also discuss the recent application of DL toward AV speech fusion and recognition. We first discuss the main AV datasets used in the literature for AVSR experiments since we consider it a data-driven machine learning (ML) task. We then consider the methodology used for visual speech recognition (VSR). Subsequently, we also consider recent AV methodology advances. We then separately discuss the evolution of the core AVSR methods, pre-processing and augmentation techniques, and modality fusion strategies. We conclude the article with a discussion on the current state of AVSR and provide our vision for future research.

Funder

RFBR

Grant

Leading scientific school

State research grant

Publisher

MDPI AG

Subject

General Mathematics,Engineering (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2227-7390/11/12/2665/pdf

Reference148 articles.

1. Ryumin, D., Kagirov, I., Axyonov, A., Pavlyuk, N., Saveliev, A., Kipyatkova, I., Zelezny, M., Mporas, I., and Karpov, A. (2020). A Multimodal User Interface for an Assistive Robotic Shopping Cart. Electronics, 9.

2. Medical Exoskeleton “Remotion” with an Intelligent Control System: Modeling, implementation, and Testing;Kagirov;Simul. Model. Pract. Theory,2021

3. A Novel Human-Vehicle Interaction Assistive Device for Arab Drivers Using Speech Recognition;Jaradat;IEEE Access,2022

4. Ivanko, D. (2022). Audio-Visual Russian Speech Recognition. [Ph.D. Thesis, Universität Ulm].

5. Potamianos, G., Neti, C., Luettin, J., and Matthews, I. (2004). Issues in Visual and Audio-Visual Speech Processing, MIT Press.

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Audio-Visual Speech Recognition In-The-Wild: Multi-Angle Vehicle Cabin Corpus and Attention-Based Method;ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP);2024-04-14

2. Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain;ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP);2024-04-14

3. Data Protection Issues in Automated Decision-Making Systems Based on Machine Learning: Research Challenges;Network;2024-03-01

4. EMOLIPS: Towards Reliable Emotional Speech Lip-Reading;Mathematics;2023-11-27

5. Deep Models for Low-Resourced Speech Recognition: Livvi-Karelian Case;Mathematics;2023-09-05