Decoder. plain-English AI glossary

Autoscaling

● Core

Automatically spinning up more instances when traffic spikes and tearing them down when it drops.

Think of it like

A restaurant that hires temp workers when the lunch rush hits and sends them home at 2pm.

Example

Your inference service normally runs 2 replicas; at 9am traffic jumps and the system automatically brings 5 more online; by noon it scales back to 2.

How it actually works

Autoscaling uses metrics like CPU, memory, or request queue depth to decide when to add or remove instances. Scaling *up* takes time (cold start overhead); scaling *down* means losing capacity mid-spike if you're not careful. Metrics lag behind actual load, so good autoscaling tuning matters.

For product teams

Autoscaling lets you handle spikes without overpaying for idle capacity.

For engineers

Set scaling thresholds carefully — too aggressive and you'll waste money on temporary instances; too conservative and you'll serve slow responses during spikes.

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome