Decoder. plain-English AI glossary

Talking Head

▲ Rising

An animated face or avatar that moves its mouth and head to match speech - creating the illusion of someone talking.

Think of it like

A puppet master animating a doll to lip-sync.

Example

Given a portrait photo and speech audio, generate a video of the face talking naturally.

How it actually works

Input: single image + audio. Output: video of that person talking. Requires: (1) lip-sync, (2) head pose/gaze matching content, (3) realistic motion. Methods: keypoint-based (detect facial landmarks, animate), flow-based (predict motion fields), or diffusion (generate frames autoregressively or in latent space). Challenges: convincingness (uncanny valley), preserving identity, handling out-of-domain audio (accents, singing). Models like Wav2Lip, Metaface, or diffusion-based approaches.

For product teams

Enables virtual presenters, customer service bots, personalized video messages.

For engineers

Input: still image + audio. Output: video. Lip-sync critical. Benchmarks: LRS3 (audio-video pairs). Metrics: sync accuracy, visual quality (LPIPS, FID).

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome