articleTop 1% cited
Synthesizing Obama
ACM Transactions on Graphics · 2017 · Vol. 36(4) · pp. 1–13
Supasorn Suwajanakorn✉(University of Washington)Steven M. Seitz(University of Washington)Ira Kemelmacher-Shlizerman(University of Washington)
Abstract
Given audio of President Barack Obama, we synthesize a high quality video of him speaking with accurate lip sync, composited into a target video clip. Trained on many hours of his weekly address footage, a recurrent neural network learns the mapping from raw audio features to mouth shapes. Given the mouth shape at each time instant, we synthesize high quality mouth texture, and composite it with proper 3D pose matching to change what he appears to be saying in a target video to match the input audio track. Our approach produces photorealistic results.
Face recognition and analysisSpeech and Audio ProcessingGenerative Adversarial Networks and Image SynthesisComputer sciencesyncArtificial intelligenceTexture (cosmology)Computer visionComputer graphics (images)Track (disk drive)Image (mathematics)Telecommunications
Funding
- Samsung
Citations
1,051
FWCI
30.89
field-weighted impact
References
51
Percentile
100%
vs. same field & year
Citations per year
Cited by
Deep video portraits
ACM Transactions on Graphics · 2018 · 649 citations
References
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
Neural Networks · 2005 · 5,322 citations
Active appearance models
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2001 · 5,440 citations
A multiresolution spline with application to image mosaics
ACM Transactions on Graphics · 1983 · 1,099 citations
Citation Network
How this paper connects to the literature. Drag to explore, click any node to open that paper.
