Autoscaling
Automatically spinning up more instances when traffic spikes and tearing them down when it drops.
Think of it like
A restaurant that hires temp workers when the lunch rush hits and sends them home at 2pm.
Example
Your inference service normally runs 2 replicas; at 9am traffic jumps and the system automatically brings 5 more online; by noon it scales back to 2.
How it actually works
Autoscaling uses metrics like CPU, memory, or request queue depth to decide when to add or remove instances. Scaling *up* takes time (cold start overhead); scaling *down* means losing capacity mid-spike if you're not careful. Metrics lag behind actual load, so good autoscaling tuning matters.
For product teams
Autoscaling lets you handle spikes without overpaying for idle capacity.
For engineers
Set scaling thresholds carefully — too aggressive and you'll waste money on temporary instances; too conservative and you'll serve slow responses during spikes.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome