Bidirectional Temporal Pose Matching for Tracking-Reference-Cited by-同舟云学术

Bidirectional Temporal Pose Matching for Tracking

Published:2024-01-21 Issue:2 Volume:13 Page:442
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Fang Yichuan¹,Shi Qingxuan¹,Yang Zhen¹

Affiliation:

1. Hebei Machine Vision Engineering Research Center, Hebei University, Baoding 071002, China

Abstract

Multi-person pose tracking is a challenging task. It requires identifying the human poses in each frame and matching them across time. This task still faces two main challenges. Firstly, sudden camera zooming and drastic pose changes between adjacent frames may result in mismatched poses between them. Secondly, the time relationships modeled by most existing methods provide insufficient information in scenarios with long-term occlusion. In this paper, to address the first challenge, we propagate the bounding boxes of the current frame to the previous frame for pose estimation, and match the estimated results with the previous ones, which we call the Backward Temporal Pose-Matching (BTPM) module. To solve the second challenge, we design an Association Across Multiple Frames (AAMF) module that utilizes long-term temporal relationships to supplement tracking information lost in the previous frames as a Re-identification (Re-id) technique. Specifically, we select keyframes with a fixed step size in the videos and label other frames as general frames. In the keyframes, we use the BTPM module and the AAMF module to perform tracking. In the general frames, we propagate poses in the previous frame to the current frame for pose estimation and association, which we call the Forward Temporal Pose-Matching (FTPM) module. If the pose association fails, the current frame will be set as a keyframe, and tracking will be re-performed. In the PoseTrack 2018 benchmark tests, our method shows significant improvements over the baseline methods, with improvements of 2.1 and 1.1 in mean Average Precision (mAP) and Multi-Object Tracking Accuracy (MOTA), respectively.

Funder

Natural Science Foundation of Hebei Province

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/13/2/442/pdf

Reference53 articles.

1. Wei, S.E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (2016, January 27–30). Convolutional pose machines. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.

2. Newell, A., Yang, K., and Deng, J. (2016, January 11–14). Stacked hourglass networks for human pose estimation. Proceedings of the Computer Vision—ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Part VIII 14.

3. Yang, W., Li, S., Ouyang, W., Li, H., and Wang, X. (2017, January 22–29). Learning feature pyramids for human pose estimation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.

4. Ke, L., Chang, M.C., Qi, H., and Lyu, S. (2018, January 8–14). Multi-scale structure-aware network for human pose estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.

5. Andriluka, M., Pishchulin, L., Gehler, P., and Schiele, B. (2014, January 23–28). 2d human pose estimation: New benchmark and state of the art analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.