Talking Head
An animated face or avatar that moves its mouth and head to match speech - creating the illusion of someone talking.
Think of it like
A puppet master animating a doll to lip-sync.
Example
Given a portrait photo and speech audio, generate a video of the face talking naturally.
How it actually works
Input: single image + audio. Output: video of that person talking. Requires: (1) lip-sync, (2) head pose/gaze matching content, (3) realistic motion. Methods: keypoint-based (detect facial landmarks, animate), flow-based (predict motion fields), or diffusion (generate frames autoregressively or in latent space). Challenges: convincingness (uncanny valley), preserving identity, handling out-of-domain audio (accents, singing). Models like Wav2Lip, Metaface, or diffusion-based approaches.
For product teams
Enables virtual presenters, customer service bots, personalized video messages.
For engineers
Input: still image + audio. Output: video. Lip-sync critical. Benchmarks: LRS3 (audio-video pairs). Metrics: sync accuracy, visual quality (LPIPS, FID).
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome