Iterative Text-Based Editing of Talking-Heads Using Neural Retargeting-Reference-Cited by-同舟云学术

Iterative Text-Based Editing of Talking-Heads Using Neural Retargeting

Published:2021-06-30 Issue:3 Volume:40 Page:1-14
ISSN:0730-0301
Container-title:ACM Transactions on Graphics
language:en
Short-container-title:ACM Trans. Graph.

Author:

Yao Xinwei¹^ORCID,Fried Ohad²,Fatahalian Kayvon¹,Agrawala Maneesh¹

Affiliation:

1. Stanford University, Stanford, CA, USA

2. The Interdisciplinary Center Herzliya, Herzliya, Israel

Abstract

We present a text-based tool for editing talking-head video that enables an iterative editing workflow. On each iteration users can edit the wording of the speech, further refine mouth motions if necessary to reduce artifacts, and manipulate non-verbal aspects of the performance by inserting mouth gestures (e.g., a smile) or changing the overall performance style (e.g., energetic, mumble). Our tool requires only 2 to 3 minutes of the target actor video and it synthesizes the video for each iteration in about 40 seconds, allowing users to quickly explore many editing possibilities as they iterate. Our approach is based on two key ideas. (1) We develop a fast phoneme search algorithm that can quickly identify phoneme-level subsequences of the source repository video that best match a desired edit. This enables our fast iteration loop. (2) We leverage a large repository of video of a source actor and develop a new self-supervised neural retargeting technique for transferring the mouth motions of the source actor to the target actor. This allows us to work with relatively short target actor videos, making our approach applicable in many real-world editing scenarios. Finally, our, refinement and performance controls give users the ability to further fine-tune the synthesized results.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Graphics and Computer-Aided Design

Link

https://dl.acm.org/doi/pdf/10.1145/3449063

Reference52 articles.

1. Google LLC. 2020a. Google Cloud Speech to Text API. https://cloud.google.com/speech-to-text Google LLC. 2020a. Google Cloud Speech to Text API. https://cloud.google.com/speech-to-text

2. Google LLC. 2020b. Google Cloud Text to Speech API. https://cloud.google.com/text-to-speech Google LLC. 2020b. Google Cloud Text to Speech API. https://cloud.google.com/text-to-speech

3. Descript Inc. 2020. Lyrebird AI. https://www.descript.com/lyrebird-ai Descript Inc. 2020. Lyrebird AI. https://www.descript.com/lyrebird-ai

4. Rev.com Inc. 2020. Rev. https://rev.com Rev.com Inc. 2020. Rev. https://rev.com

5. Martín Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Mané Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Viégas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https://www.tensorflow.org/ Software available from tensorflow.org. Martín Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Mané Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Viégas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https://www.tensorflow.org/ Software available from tensorflow.org.

Cited by 14 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Deep Learning for Visual Speech Analysis: A Survey;IEEE Transactions on Pattern Analysis and Machine Intelligence;2024-09

2. Audio-to-Deep-Lip: Speaking lip synthesis based on 3D landmarks;Computers & Graphics;2024-05

3. Facial Parameter Splicing: A Novel Approach to Efficient Talking Face Generation;ACM Multimedia Asia 2023;2023-12-06

4. MusicFace: Music-driven expressive singing face synthesis;Computational Visual Media;2023-11-30

5. StableFace: Analyzing and Improving Motion Stability for Talking Face Generation;IEEE Journal of Selected Topics in Signal Processing;2023-11