YOLOv8-PoseBoost: Advancements in Multimodal Robot Pose Keypoint Detection-Reference-Cited by-同舟云学术

YOLOv8-PoseBoost: Advancements in Multimodal Robot Pose Keypoint Detection

Published:2024-03-11 Issue:6 Volume:13 Page:1046
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Wang Feng¹^ORCID,Wang Gang²³,Lu Baoli⁴⁵

Affiliation:

1. Engineering and Technology College, Hubei University of Technology, Wuhan 430068, China

2. School of Computing and Data Engineering, NingboTech University, Ningbo 315100, China

3. Department of Bioengineering, Imperial College London, London SW7 2AZ, UK

4. School of Computing, University of Portsmouth, Portsmouth PO1 3HE, UK

5. Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China

Abstract

In the field of multimodal robotics, achieving comprehensive and accurate perception of the surrounding environment is a highly sought-after objective. However, current methods still have limitations in motion keypoint detection, especially in scenarios involving small target detection and complex scenes. To address these challenges, we propose an innovative approach known as YOLOv8-PoseBoost. This method introduces the Channel Attention Module (CBAM) to enhance the network’s focus on small targets, thereby increasing sensitivity to small target individuals. Additionally, we employ multiple scale detection heads, enabling the algorithm to comprehensively detect individuals of varying sizes in images. The incorporation of cross-level connectivity channels further enhances the fusion of features between shallow and deep networks, reducing the rate of missed detections for small target individuals. We also introduce a Scale Invariant Intersection over Union (SIoU) redefined bounding box regression localization loss function, which accelerates model training convergence and improves detection accuracy. Through a series of experiments, we validate YOLOv8-PoseBoost’s outstanding performance in motion keypoint detection for small targets and complex scenes. This innovative approach provides an effective solution for enhancing the perception and execution capabilities of multimodal robots. It has the potential to drive the development of multimodal robots across various application domains, holding both theoretical and practical significance.

Funder

Ningbo Key R&D Program

Zhejiang Province Postdoctoral Research Funding Project

Ningbo Natural Science Foundation

Publisher

MDPI AG

Link

https://www.mdpi.com/2079-9292/13/6/1046/pdf

Reference31 articles.

1. Cheng, B., Xiao, B., Wang, J., Shi, H., Huang, T.S., and Zhang, L. (2020, January 13–19). Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.

2. Moon, G., Yu, S.I., Wen, H., Shiratori, T., and Lee, K.M. (2020, January 23–28). Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XX 16.

3. Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 15–20). Deep high-resolution representation learning for human pose estimation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.

4. Sattler, T., Zhou, Q., Pollefeys, M., and Leal-Taixe, L. (2019, January 15–20). Understanding the limitations of cnn-based absolute camera pose regression. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.

5. Iskakov, K., Burkov, E., Lempitsky, V., and Malkov, Y. (November, January 27). Learnable triangulation of human pose. Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea.

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. CBA-YOLOv5s: A hip dysplasia detection algorithm based on YOLOv5s using angle consistency and bi-level routing attention;Biomedical Signal Processing and Control;2024-09

2. A Robust Pointer Meter Reading Recognition Method Based on TransUNet and Perspective Transformation Correction;Electronics;2024-06-21

3. Leveraging Advanced Computer Vision for Hazardous Behavior Monitoring on Campus Safety Maintenance;2024 IEEE 4th International Conference on Electronic Communications, Internet of Things and Big Data (ICEIB);2024-04-19