Lip Sync
Aligning a video of mouth movements with audio - or generating video that matches audio.
Think of it like
Dubbing: the actors lips must match the audio dialogue.
Example
A talking head avatar syncs its mouth to generated speech. A video with mismatched audio is corrected.
How it actually works
Can mean: (1) verifying alignment (is this video dubbed?), or (2) generating video frames of a mouth matched to audio. Deterministic: use detected lip keypoints + audio to measure sync loss. Generative: audio -> visual features -> video frames (requires autoregressive or diffusion model). Challenges: accents, realistic teeth/tongue, avoiding uncanny valley.
For product teams
Powers deepfakes, talking-head avatars, audiobook video, and dubbed content.
For engineers
Audio-video alignment loss: sync between phoneme onsets and mouth movements. Generative: condition on audio embeddings, generate frame sequences.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome