Vision Transformers and Transfer Learning Approaches for Arabic Sign Language Recognition-Reference-Cited by-同舟云学术

Vision Transformers and Transfer Learning Approaches for Arabic Sign Language Recognition

Published:2023-10-24 Issue:21 Volume:13 Page:11625
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Alharthi Nojood M.¹,Alzahrani Salha M.¹^ORCID

Affiliation:

1. Department of Computer Science, College of Computers and Information Technology, Taif University, P.O. Box 11099, Taif 21944, Saudi Arabia

Abstract

Sign languages are complex, but there are ongoing research efforts in engineering and data science to recognize, understand, and utilize them in real-time applications. Arabic sign language recognition (ArSL) has been examined and applied using various traditional and intelligent methods. However, there have been limited attempts to enhance this process by utilizing pretrained models and large-sized vision transformers designed for image classification tasks. This study aimed to create robust transfer learning models trained on a dataset of 54,049 images depicting 32 alphabets from an ArSL dataset. The goal was to accurately classify these images into their corresponding Arabic alphabets. This study included two methodological parts. The first one was the transfer learning approach, wherein we utilized various pretrained models namely MobileNet, Xception, Inception, InceptionResNet, DenseNet, and BiT, and two vision transformers namely ViT, and Swin. We evaluated different variants from base-sized to large-sized pretrained models and vision transformers with weights initialized from the ImageNet dataset or otherwise randomly. The second part was the deep learning approach using convolutional neural networks (CNNs), wherein several CNN architectures were trained from scratch to be compared with the transfer learning approach. The proposed methods were evaluated using the accuracy, AUC, precision, recall, F1 and loss metrics. The transfer learning approach consistently performed well on the ArSL dataset and outperformed other CNN models. ResNet and InceptionResNet obtained a comparably high performance of 98%. By combining the concepts of transformer-based architecture and pretraining, ViT and Swin leveraged the strengths of both architectures and reduced the number of parameters required for training, making them more efficient and stable than other models and existing studies for ArSL classification. This demonstrates the effectiveness and robustness of using transfer learning with vision transformers for sign language recognition for other low-resourced languages.

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/13/21/11625/pdf

Reference77 articles.

1. Occupational hearing loss;May;Am. J. Ind. Med.,2000

2. Helping Hearing-Impaired in Emergency Situations: A Deep Learning-Based Approach;Areeb;IEEE Access,2022

3. Arabic Sign Language Recognition System for Alphabets Using Machine Learning Techniques;Tharwat;J. Electr. Comput. Eng.,2021

4. Pan, T.Y., Lo, L.Y., Yeh, C.W., Li, J.W., Liu, H.T., and Hu, M.C. (2016, January 20–22). Real-Time Sign Language Recognition in Complex Background Scene Based on a Hierarchical Clustering Classification Method. Proceedings of the IEEE Second International Conference on Multimedia Big Data (BigMM), Taipei, Taiwan.

5. A review on Arabic sign language translator systems;Mohammed;J. Phys. Conf. Ser.,2021

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Intelligent real-life key-pixel image detection system for early Arabic sign language learners;PeerJ Computer Science;2024-06-14

2. Convolutional Neural Networks for Indian Sign Language Recognition;International Journal of Innovative Science and Research Technology (IJISRT);2024-06-11

3. Efhamni: A Deep Learning-Based Saudi Sign Language Recognition Application;Sensors;2024-05-14

4. Applying Swin Architecture to Diverse Sign Language Datasets;Electronics;2024-04-16

5. A Brief Review of Sign Language Recognition Methods and Cutting-edge Technologies;2024 5th International Conference on Computer Engineering and Application (ICCEA);2024-04-12