Sign2Pose: A Pose-Based Approach for Gloss Prediction Using a Transformer Model-Reference-Cited by-同舟云学术

Sign2Pose: A Pose-Based Approach for Gloss Prediction Using a Transformer Model

Published:2023-03-06 Issue:5 Volume:23 Page:2853
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Eunice Jennifer¹,J Andrew²^ORCID,Sei Yuichi³^ORCID,Hemanth D. Jude¹^ORCID

Affiliation:

1. Department of Electronics and Communication Engineering, Karunya Institute of Technology and Sciences, Coimbatore 641114, India

2. Computer Science and Engineering, Manipal Institute of Technology, Manipal Academy of Higher Education, Manipal 576104, India

3. Department of Informatics, The University of Electro-Communications, Tokyo 182-8585, Japan

Abstract

Word-level sign language recognition (WSLR) is the backbone for continuous sign language recognition (CSLR) that infers glosses from sign videos. Finding the relevant gloss from the sign sequence and detecting explicit boundaries of the glosses from sign videos is a persistent challenge. In this paper, we propose a systematic approach for gloss prediction in WLSR using the Sign2Pose Gloss prediction transformer model. The primary goal of this work is to enhance WLSR’s gloss prediction accuracy with reduced time and computational overhead. The proposed approach uses hand-crafted features rather than automated feature extraction, which is computationally expensive and less accurate. A modified key frame extraction technique is proposed that uses histogram difference and Euclidean distance metrics to select and drop redundant frames. To enhance the model’s generalization ability, pose vector augmentation using perspective transformation along with joint angle rotation is performed. Further, for normalization, we employed YOLOv3 (You Only Look Once) to detect the signing space and track the hand gestures of the signers in the frames. The proposed model experiments on WLASL datasets achieved the top 1% recognition accuracy of 80.9% in WLASL100 and 64.21% in WLASL300. The performance of the proposed model surpasses state-of-the-art approaches. The integration of key frame extraction, augmentation, and pose estimation improved the performance of the proposed gloss prediction model by increasing the model’s precision in locating minor variations in their body posture. We observed that introducing YOLOv3 improved gloss prediction accuracy and helped prevent model overfitting. Overall, the proposed model showed 17% improved performance in the WLASL 100 dataset.

Funder

JSPS KAKENHI

JST, PRESTO

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/23/5/2853/pdf

Reference71 articles.

1. Automatic Sign Language Finger Spelling Using Convolution Neural Network: Analysis;Dept;Int. J. Pure Appl. Math.,2017

2. Deep CNN for Static Indian Sign Language Digits Recognition;Frontiers in Artificial Intelligence and Applications,2022

3. Handwritten mathematical symbols dataset;Chajri;Data Br.,2016

4. Huang, J., Zhou, W., Zhang, Q., Li, H., and Li, W. (2018, January 2–7). Video-based sign language recognition without temporal segmentation. Proceedings of the 32nd Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.

5. Static sign language recognition using deep learning;Tolentino;Int. J. Mach. Learn. Comput.,2019

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Techniques for Generating Sign Language a Comprehensive Review;Journal of The Institution of Engineers (India): Series B;2024-07-13

2. A machine learning-driven web application for sign language learning;Frontiers in Artificial Intelligence;2024-06-18

3. Real-Time Arabic Sign Language Recognition Using a Hybrid Deep Learning Model;Sensors;2024-06-06

4. Synthetic Corpus Generation for Deep Learning-Based Translation of Spanish Sign Language;Sensors;2024-02-24

5. Recent Advances on Deep Learning for Sign Language Recognition;Computer Modeling in Engineering & Sciences;2024