research-article

Open access

PoseVocab: Learning Joint-structured Pose Embeddings for Human Avatar Modeling

Authors:

Zhe Li,

Zerong Zheng,

Yuxiao Liu,

Boyao Zhou,

Yebin LiuAuthors Info & Claims

SIGGRAPH '23: ACM SIGGRAPH 2023 Conference Proceedings

Article No.: 8, Pages 1 - 11

https://doi.org/10.1145/3588432.3591490

Published: 23 July 2023 Publication History

All formats PDF

Abstract

Creating pose-driven human avatars is about modeling the mapping from the low-frequency driving pose to high-frequency dynamic human appearances, so an effective pose encoding method that can encode high-fidelity human details is essential to human avatar modeling. To this end, we present PoseVocab, a novel pose encoding method that encourages the network to discover the optimal pose embeddings for learning the dynamic human appearance. Given multi-view RGB videos of a character, PoseVocab constructs key poses and latent embeddings based on the training poses. To achieve pose generalization and temporal consistency, we sample key rotations in so(3) of each joint rather than the global pose vectors, and assign a pose embedding to each sampled key rotation. These joint-structured pose embeddings not only encode the dynamic appearances under different key poses, but also factorize the global pose embedding into joint-structured ones to better learn the appearance variation related to the motion of each joint. To improve the representation ability of the pose embedding while maintaining memory efficiency, we introduce feature lines, a compact yet effective 3D representation, to model more fine-grained details of human appearances. Furthermore, given a query pose and a spatial position, a hierarchical query strategy is introduced to interpolate pose embeddings and acquire the conditional pose feature for dynamic human synthesis. Overall, PoseVocab effectively encodes the dynamic details of human appearance and enables realistic and generalized animation under novel poses. Experiments show that our method outperforms other state-of-the-art baselines both qualitatively and quantitatively in terms of synthesis quality. Code is available at https://github.com/lizhe00/PoseVocab.

Supplemental Material

MP4 File

presentation

Download
139.23 MB

ZIP File

Supplementary document and video

Download
222.21 MB

References

[1]

Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. 2018. Video based reconstruction of 3d people models. In CVPR. 8387–8397.

Abstract

Supplemental Material

References

Cited By

Index Terms

Recommendations

Tracking Human Motion in Structured Environments Using a Distributed-Camera System

Single-View Dressed Human Modeling via Morphable Template

Joint Reasoning for Camera and 3D Human Pose Estimation

Comments

Information

Published In

Sponsors

Publisher

Publication History

Check for updates

Author Tags

Qualifiers

Funding Sources

Conference

Acceptance Rates

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Cited By

View options

PDF

eReader

HTML Format

Get Access

Login options

Full Access

Figures

Other

Share

Share this Publication link

Share on social media

Affiliations