Multi-Angle Models and Lightweight Unbiased Decoding-Based Algorithm for Human Pose Estimation-Reference-Cited by-同舟云学术

Multi-Angle Models and Lightweight Unbiased Decoding-Based Algorithm for Human Pose Estimation

Published:2023-06-30 Issue:08 Volume:37 Page:
ISSN:0218-0014
Container-title:International Journal of Pattern Recognition and Artificial Intelligence
language:en
Short-container-title:Int. J. Patt. Recogn. Artif. Intell.

Author:

He Jianghai¹^ORCID,Zhang Weitong¹,Shang Ronghua¹,Feng Jie¹,Jiao Licheng¹

Affiliation:

1. Key Laboratory of Intelligent Perception and Image Understanding, Ministry of Education, School of Artificial Intelligence, Xidian University, Xi’an, Shaanxi Province 710071, P. R. China

Abstract

When a top-down method is taken to the task of human pose estimation, the accuracy of joint point localization is often limited by the accuracy of human detection. In addition, conventional algorithms commonly encode the image to generate a heat map before processing, but the systematic error in decoding the heat map back to the original image has an impact on the positioning. Therefore, to address the two problems, we propose an algorithm that uses multiple angle models to generate the human boxes and then performs lightweight decoding to recover the image. The new boxes can better fit humans and the recovery error can be reduced. First, we split the backbone network into three sub-networks, the first sub-network is responsible for generating the original human box, the second sub-network is responsible for generating a coarse pose estimation in the boxes, and the third sub-network is responsible for a high-precision pose estimation. In order to make the human box fit the human body better, with only a small number of interfering pixels inside the box, models of the human boxes with multiple rotation angles are generated. The results from the second sub-network are used to select the best human box. Using this human box as input to the third sub-network can significantly improve the accuracy of the pose estimation. Then to reduce the errors arising from image decoding, we propose a lightweight unbiased decoding strategy that differs from traditional methods by combining multiple possible offsets to select the direction and size of the final offset. On the MPII dataset and the COCO dataset, we compare the proposed algorithm with 11 state-of-the-art algorithms. The experimental results show that the algorithm achieves a large improvement in accuracy for a wide range of image sizes and different metrics.

Funder

Innovative Research Group Project of the National Natural Science Foundation of China

Natural Science Basic Research Program of Shaanxi Province

Open Research Projects of Zhejiang Lab

Basic and Applied Basic Research Foundation of Guangdong Province

Research Project of SongShan Laboratory

Fundamental Research Funds for the Central Universities

Publisher

World Scientific Pub Co Pte Ltd

Subject

Artificial Intelligence,Computer Vision and Pattern Recognition,Software

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0218001423560141

Reference38 articles.

1. Learning Delicate Local Representations for Multi-person Pose Estimation

2. Human action recognition in videos based on spatiotemporal features and bag-of-poses

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Identification of Coffee Leaf Pests and Diseases based on Transfer Learning and Knowledge Distillation;Frontiers in Computing and Intelligent Systems;2023-09-12