Effect of Various Visual Speech Units on Language Identification Using Visual Speech Recognition-Reference-Cited by-同舟云学术

Effect of Various Visual Speech Units on Language Identification Using Visual Speech Recognition

Published:2020-10 Issue:04 Volume:20 Page:2050029
ISSN:0219-4678
Container-title:International Journal of Image and Graphics
language:en
Short-container-title:Int. J. Image Grap.

Author:

Brahme Aparna¹,Bhadade Umesh²

Affiliation:

1. MET’s Institute of Engineering, Adgaon, Nashik 422003, India

2. SSBT’s College of Engineering and Technology, Bhambori, Jalgaon, India

Abstract

In this paper, we describe our work in Spoken language Identification using Visual Speech Recognition (VSR) and analyze the effect of various visual speech units used to transcribe the visual speech on language recognition. We have proposed a new approach of word recognition followed by the word N-gram language model (WRWLM), which uses high-level syntactic features and the word bigram language model for language discrimination. Also, as opposed to the traditional visemic approach, we propose a holistic approach of using the signature of a whole word, referred to as a “Visual Word” as visual speech unit for transcribing visual speech. The result shows Word Recognition Rate (WRR) of 88% and Language Recognition Rate (LRR) of 94% in speaker dependent cases and 58% WRR and 77% LRR in speaker independent cases for English and Marathi digit classification task. The proposed approach is also evaluated for continuous speech input. The result shows that the Spoken Language Identification rate of 50% is possible even though the WRR using Visual Speech Recognition is below 10%, using only 1[Formula: see text]s of speech. Also, there is an improvement of about 5% in language discrimination as compared to traditional visemic approaches.

Publisher

World Scientific Pub Co Pte Lt

Subject

Computer Graphics and Computer-Aided Design,Computer Science Applications,Computer Vision and Pattern Recognition

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0219467820500291

Reference45 articles.

1. Visual units and confusion modelling for automatic lip-reading

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Shuffle Attention U-Net for Speech Enhancement in Time Domain;International Journal of Image and Graphics;2023-03-31

2. A highly stretchable and sensitive strain sensor for lip-reading extraction and speech recognition;Journal of Materials Chemistry C;2023

3. A Review of Machine Learning-Based Recognition of Sign Language;International Journal of Image and Graphics;2022-10-05

4. Data science and AI in FinTech: an overview;International Journal of Data Science and Analytics;2021-08

5. Data science and AI in FinTech: An overview;SSRN Electronic Journal;2021