Decoder. plain-English AI glossary

Music Generation

▲ Rising

Creating original musical compositions - melodies, harmonies, or full arrangements - often from text or other constraints.

Think of it like

A composer jamming: they start with a mood or theme and improvise the rest.

Example

Text: "Upbeat electronic dance music in minor key" -> model generates a full track. Or: MIDI prompt -> extend or remix it.

How it actually works

Can operate on symbolic level (MIDI, musical notation) or audio (waveforms, spectrograms). Symbolic: discrete tokens, easier to learn long-form structure, less data needed, but limited expressiveness. Audio: richer, but harder to train and slower. Transformer-based models (e.g., Jukebox), diffusion models (Riffusion), and VAE-based approaches. Challenges: long-range structure (songs have verse-chorus), coherence, diversity, and domain specificity (different genres).

For product teams

Accelerates music production, enables personalized content, supports creators.

For engineers

Datasets: Spotify, MAESTRO (MIDI), FSD50K. Metrics: musicality, adherence to prompt, diversity. Models: MusicGen, Jukebox.

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome