ExtraSysWork out the decision first. The vocabulary follows.
Spot Instances — What You Are Actually Buying

FORK 03/Pay for what?/Spot Instances — What You Are Actually Buying

Spot Instances — What You Are Actually Buying

By the ExtraSys desk · Pay for what? · 2 min read

You are not buying compute. You are buying an option on spare capacity, and the provider holds the other side of that trade.

The Preemption Model, Honestly

A spot instance — called a spot instance on AWS, a preemptible VM on Google Cloud, a Spot VM on Azure — is idle capacity in a provider's fleet that would otherwise sit warm and unused. The provider rents it to you at a significant discount. In exchange, you accept one condition that changes everything: the provider can reclaim that capacity, typically with very little warning, whenever demand from higher-priority customers requires it.

The warning window varies by provider and is not guaranteed beyond what each provider documents for their specific product. On AWS the interruption notice is two minutes; Google Cloud gives only thirty seconds. The number matters less than the implication: your workload must be able to stop, cleanly, on extremely short notice — or be willing to die without ceremony.

That is the whole trade. The discount is not a loyalty reward or a volume incentive. It is compensation for the interruption risk you absorb.

What Survives Preemption, and What Doesn't

A bundle of network cables with printed adhesive labels at the termination point, tied with velcro straps

The workload characteristics that make spot viable are specific. Batch processing is the canonical fit: rendering, genomic sequencing, log analysis, machine learning training jobs that checkpoint state to object storage frequently enough that a restart loses minutes rather than hours. The key property is idempotency — running the same chunk of work twice produces the same result, and no downstream system notices the retry.

Stateless web tier nodes behind a load balancer can tolerate spot, provided the pool is large enough that losing one or two instances degrades capacity rather than ending service. Auto-scaling groups that mix spot and on-demand instances give you the discount on a fraction of your fleet while on-demand nodes hold the floor.

Auto-scaling groups that mix spot and on-demand instances give you the discount on a fraction of your fleet while on-demand nodes hold the floor.

What cannot tolerate preemption: anything with a session, anything holding a lock, anything mid-transaction, any primary database node, any job where partial progress is either lost or corrupting. If your workload has state that lives only in memory and cannot be reconstructed from a durable store, spot will eventually ruin your day — and eventually is not a long wait.

The discipline spot instances impose is actually useful: they force architects to treat compute as disposable, to externalise state, and to design for failure explicitly. Those are good properties regardless of billing model. But the forcing function is real. If your architecture cannot survive a preemption cleanly today, buying spot does not create that resilience — it just makes the gap visible, expensively and at the worst possible moment.