Egocentric-View Fingertip Detection for Air Writing Based on Convolutional Neural Networks-Reference-Cited by-同舟云学术

Egocentric-View Fingertip Detection for Air Writing Based on Convolutional Neural Networks

Published:2021-06-26 Issue:13 Volume:21 Page:4382
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Chen Yung-Han,Huang Chi-Hsuan,Syu Sin-Wun,Kuo Tien-Ying^ORCID,Su Po-Chyi^ORCID

Abstract

This research investigated real-time fingertip detection in frames captured from the increasingly popular wearable device, smart glasses. The egocentric-view fingertip detection and character recognition can be used to create a novel way of inputting texts. We first employed Unity3D to build a synthetic dataset with pointing gestures from the first-person perspective. The obvious benefits of using synthetic data are that they eliminate the need for time-consuming and error-prone manual labeling and they provide a large and high-quality dataset for a wide range of purposes. Following that, a modified Mask Regional Convolutional Neural Network (Mask R-CNN) is proposed, consisting of a region-based CNN for finger detection and a three-layer CNN for fingertip location. The process can be completed in 25 ms per frame for 640×480 RGB images, with an average error of 8.3 pixels. The speed is high enough to enable real-time “air-writing”, where users are able to write characters in the air to input texts or commands while wearing smart glasses. The characters can be recognized by a ResNet-based CNN from the fingertip trajectories. Experimental results demonstrate the feasibility of this novel methodology.

Funder

Ministry of Science and Technology

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/21/13/4382/pdf

Reference31 articles.

1. A Human Body Analysis System

2. Skin color-based video segmentation under time-varying illumination

3. Model-Based 3D Hand Pose Estimation from Monocular Video

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Next Generation Computing and Communication Hub for First Responders in Smart Cities;Sensors;2024-04-08

2. Hover-Based Japanese Input Method without Restricting Arm Position for XR;Proceedings of the 35th Australian Computer-Human Interaction Conference;2023-12-02

3. The Virtual Air Canvas Using Image Processing;2023 International Conference on New Frontiers in Communication, Automation, Management and Security (ICCAMS);2023-10-27

4. Open Scene Understanding: Grounded Situation Recognition Meets Segment Anything for Helping People with Visual Impairments;2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW);2023-10-02

5. Fingertip Detection Algorithm Based on Maximum Discrimination HOG Feature in Complex Background;IEEE Access;2023