Spot instances for the workloads you were told not to
The Lizrd team · · 3 min read
Spot instances are the deepest discount in the cloud — routinely 60–90% off on-demand — and most teams use them for exactly one thing: stateless batch jobs. The received wisdom is that anything stateful, latency-sensitive, or user-facing is off-limits, because Spot can be reclaimed with two minutes’ notice.
That wisdom is a decade out of date. The reclamation is real, but the tooling around it has matured to the point where a lot of the workloads you were told to keep on-demand can run on Spot safely — if you design for interruption instead of pretending it won’t happen.
Interruption is a signal, not a disaster
The two-minute reclamation warning is the whole game. A workload that listens for it and reacts gracefully — drain connections, checkpoint state, deregister from the load balancer, hand off in-flight work — treats an interruption as a routine event, not an outage. The difference between “Spot is too risky” and “Spot is fine” is almost entirely whether your workload handles that signal.
Kubernetes makes this almost free: a node-termination handler catches the warning and cordons-and-drains the node, and your pods reschedule the same way they would during any node rotation. If your cluster survives a node going away — and it should — it survives Spot.
Diversify so you don’t lose the whole pool at once
The failure mode that scares people is a correlated reclamation — the entire fleet pulled simultaneously because it was all the same instance type in the same AZ. The defense is diversification: draw from many instance types across many AZs, so the loss of any one Spot pool costs you a slice, not the service.
# Don't bet the fleet on one pool
- instance_types = ["m6i.2xlarge"]
+ instance_types = ["m6i.2xlarge", "m6a.2xlarge", "m5.2xlarge", "m5a.2xlarge"]
+ # spread across AZs; let capacity-optimized allocation pick the safe pools
Capacity-optimized allocation strategies then place you in the pools least likely to be interrupted. Breadth plus a smart allocation strategy turns Spot from a gamble into a statistical near-certainty.
What can actually move
With interruption handling and diversification in place, the eligible set is much wider than batch:
- Stateless web and API tiers behind a load balancer — drain on the warning, reschedule elsewhere.
- Kubernetes worker nodes for interruption-tolerant workloads — the ecosystem is built for this.
- CI/CD runners — a killed build retries; the savings are enormous.
- Queue consumers and stream processors — re-deliver on interruption, which they already handle.
The genuinely off-limits list shrinks to a small core: singleton stateful services with no failover, and hard real-time paths where a two-minute drain is too slow.
Know your baseline before you go elastic
Spot is the elastic layer on top of a committed baseline — so the prerequisite is knowing which of your workloads are genuinely interruption-tolerant and how big that baseline really is. Move the wrong thing and you learn the hard way; leave the right things on-demand and you’re paying 3–5x for no reason. The answer lives in your workloads’ behavior, which most teams have never actually characterized.
That’s the read Lizrd gives you: it looks at your fleet and utilization, identifies which workloads are safe Spot candidates versus which belong on committed capacity, and proposes the diversified configuration as a concrete change. Spot isn’t just for batch — it’s for everything you’ve taught to survive an interruption.