Video2mesh: 3D human pose and shape recovery by a temporal convolutional transformer network-Reference-Cited by-同舟云学术

Video2mesh: 3D human pose and shape recovery by a temporal convolutional transformer network

Published:2023-02-27 Issue:4 Volume:17 Page:379-388
ISSN:1751-9632
Container-title:IET Computer Vision
language:en
Short-container-title:IET Computer Vision

Author:

Chao Xianjin¹^ORCID,Ge Zhipeng²,Leung Howard¹

Affiliation:

1. City University of Hong Kong Hong Kong China

2. Nan Jing University Nan Jing China

Abstract

AbstractFrom a 2D video of a person in action, human mesh recovery aims to infer the 3D human pose and shape frame by frame. Despite progress on video‐based human pose and shape estimation, it is still challenging to guarantee high accuracy and smoothness simultaneously. To tackle this problem, we propose a Video2mesh, a temporal convolutional transformer (TConvTransformer) based temporal network which is able to recover accurate and smooth human mesh from 2D video. The temporal convolution block achieves the sequence‐level smoothness by aggregating image features from adjacent frames. The subsequent multi‐attention transformer improves the accuracy due to its multi‐subspace for better middle‐frame feature representation. Meanwhile, we add a TConvTransformer discriminator which is trained together with our 3D human mesh temporal encoder. This TConvTransformer discriminator further improves the accuracy and smoothness by restricting the pose and shape in a more reliable space based on the AMASS dataset. We conduct extensive experiments on three standard benchmark datasets and show that our proposed Video2mesh outperforms other state‐of‐the‐art methods in both accuracy and smoothness.

Funder

City University of Hong Kong

Publisher

Institution of Engineering and Technology (IET)

Subject

Computer Vision and Pattern Recognition,Software

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1049/cvi2.12172

Reference52 articles.

1. Kanazawa A. et al.:End‐to‐end recovery of human shape and pose. In:Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp.7122–7131(2018)

2. Elhayek A. et al.:Fully automatic multi‐person human motion capture for vr applications. In:International Conference on Virtual Reality and Augmented Reality(2018)

3. Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image

4. Lassner C. et al.:Unite the people: closing the loop between 3d and 2d human representations. In:Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp.6050–6059(2017)