Affiliation:
1. George Mason University, Fairfax, United States of America
Abstract
We present an approach to generate virtual activity snippets, which comprise sequenced keyframes of multi-character, multi-object interaction scenarios in 3D environments, by learning from recordings of human-scene interactions. The generation consists of two stages. First, we use a sequential deep graph generative model with a temporal module to iteratively generate keyframe descriptions, which represent abstract interactions using graphs, while preserving spatial-temporal relations through the activities. Second, we devise an optimization framework to instantiate the activity snippets in virtual 3D environments guided by the generated keyframe descriptions. Our approach optimizes the poses of character and object instances encoded by the graph nodes to satisfy the relations and constraints encoded by the graph edges. The instantiation process includes a coarse 2D optimization followed by a fine 3D optimization to effectively explore the complex solution space for placing and posing the instances. Through experiments and a perceptual study, we applied our approach to generate plausible activity snippets under different settings.
Publisher
Association for Computing Machinery (ACM)
Subject
Computer Graphics and Computer-Aided Design
Reference61 articles.
1. Text2SceneVR
2. Task-based locomotion
3. Nikos Athanasiou , Mathis Petrovich , Michael J Black , and Gül Varol . 2022 . TEACH: Temporal Action Composition for 3D Humans. arXiv preprint arXiv:2209.04066 (2022). Nikos Athanasiou, Mathis Petrovich, Michael J Black, and Gül Varol. 2022. TEACH: Temporal Action Composition for 3D Humans. arXiv preprint arXiv:2209.04066 (2022).
4. Synthesis of concurrent object manipulation tasks;Bai Yunfei;ACM Transactions on Graphics,2012
5. ActivityNet: A large-scale video benchmark for human activity understanding
Cited by
5 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献