What we see is what we do: a practical Peripheral Vision-Based HMM framework for gaze-enhanced recognition of actions in a medical procedural task-Reference-Cited by-同舟云学术

What we see is what we do: a practical Peripheral Vision-Based HMM framework for gaze-enhanced recognition of actions in a medical procedural task

Published:2023-01-04 Issue:4 Volume:33 Page:939-965
ISSN:0924-1868
Container-title:User Modeling and User-Adapted Interaction
language:en
Short-container-title:User Model User-Adap Inter

Author:

Wang Felix S.^ORCID,Kreiner Thomas^ORCID,Lutz Alexander^ORCID,Lohmeyer Quentin^ORCID,Meboldt Mirko^ORCID

Abstract

AbstractDeep learning models have shown remarkable performances in egocentric video-based action recognition (EAR), but rely heavily on a large quantity of training data. In specific applications with only limited data available, eye movement data may provide additional valuable sensory information to achieve accurate classification performances. However, little is known about the effectiveness of gaze data as a modality for egocentric action recognition. We, therefore, propose the new Peripheral Vision-Based HMM (PVHMM) classification framework, which utilizes context-rich and object-related gaze features for the detection of human action sequences. Gaze information is quantified using two features, the object-of-interest hit and the object–gaze distance, and human action recognition is achieved by employing a hidden Markov model. The classification performance of the framework is tested and validated on a safety-critical medical device handling task sequence involving seven distinct action classes, using 43 mobile eye tracking recordings. The robustness of the approach is evaluated using the addition of Gaussian noise. Finally, the results are then compared to the performance of a VGG-16 model. The gaze-enhanced PVHMM achieves high classification performances in the investigated medical procedure task, surpassing the purely image-based classification model. Consequently, this gaze-enhanced EAR approach shows the potential for the implementation in action sequence-dependent real-world applications, such as surgical training, performance assessment, or medical procedural tasks.

Funder

Innosuisse - Schweizerische Agentur für Innovationsförderung

Swiss Federal Institute of Technology Zurich

Publisher

Springer Science and Business Media LLC

Subject

Computer Science Applications,Human-Computer Interaction,Education

Link

https://link.springer.com/content/pdf/10.1007/s11257-022-09352-9.pdf

Reference69 articles.

1. Allahverdyan, A., Galstyan, A.: Comparative analysis of viterbi training and maximum likelihood estimation for HMMs. In: Advances in neural information processing systems 24: 25th annual conference on neural information processing systems 2011, NIPS 2011. https://arxiv.org/abs/1312.4551v1. (2011, December 16)

2. Almaadeed, N., Elharrouss, O., Al-Maadeed, S., Bouridane, A., Beghdadi, A.: A novel approach for robust multi human action recognition and summarization based on 3D convolutional neural networks. https://www.researchgate.net/publication/334735494. (2019)

3. Arabacı, M.A., Özkan, F., Surer, E., Jančovič, P., Temizel, A.: Multi-modal egocentric activity recognition using audio-visual features. Multimed. Tools Appl 80(11), 16299–16328 (2018). https://doi.org/10.1007/s11042-020-08789-7

4. Bandini, A., Zariffa, J.: Analysis of the hands in egocentric vision: a survey. IEEE Trans. Pattern Anal. Mach. Intell (2020). https://doi.org/10.1109/tpami.2020.2986648

5. Basha, S.H.S., Dubey, S.R., Pulabaigari, V., Mukherjee, S.: Impact of fully connected layers on performance of convolutional neural networks for image classification. Neurocomputing 378, 112–119 (2020a). https://doi.org/10.1016/J.NEUCOM.2019.10.008

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Joint pyramidal perceptual attention and hierarchical consistency constraint for gaze estimation;Computer Vision and Image Understanding;2024-11

2. FreeGaze: A Framework for 3D Gaze Estimation Using Appearance Cues from a Facial Video;Sensors;2023-12-04

3. Use of eye-tracking technology for appreciation-based information in design decisions related to product details: Furniture example;Multimedia Tools and Applications;2023-06-13

4. Gaze is more than just a point: Rethinking visual attention analysis using peripheral vision-based gaze mapping;2023 Symposium on Eye Tracking Research and Applications;2023-05-30