Blog / Guides

Spot vs on-demand GPUs: is the discount worth the interruption?

Published ยท AWS interruption and pricing documentation checked 6 September 2026; cost example is illustrative

Spot capacity can make an evaluation run much cheaper. The same capacity can make a coding assistant disappear in the middle of a task. The decision depends on what an interruption does to your work and how reliably you can recover.

Start with the workload: can it wait, can it retry, and is completed work saved somewhere durable? Then compare the cost of finishing it. A large hourly discount is useful only if enough of the billed time produces results you can keep.

Know which provider's spot product you are buying

AWS Spot Instances use spare EC2 capacity at a rate that can change. Capacity may be reclaimed. On-demand removes that particular spare-capacity interruption mechanism, but it does not guarantee an instance will be available for every new launch.

Do not assume another GPU provider has identical pricing, notice periods or recovery behavior. Some offerings use different terms or do not offer a spot equivalent for the hardware you want. Record the product, region, instance type and price basis in your comparison.

Two minutes of warning is not two minutes of downtime

AWS documents a two-minute warning before a Spot Instance is stopped or terminated, delivered on a best-effort basis. Hibernation begins immediately and does not provide that two-minute lead time. Design recovery so it also works when a notice is missed.

An LLM service must find replacement capacity, start the runtime, load weights, warm up and pass a real inference check before it can serve again. Requests already generating may fail. Saved weights shorten one part of that process; they do not provide a replacement GPU or preserve an in-flight answer.

For batch jobs, persist completed results and acknowledge a queue item only when its result is durable. For an agent that can write files or call external tools, give retried operations stable identifiers so a retry does not repeat an action that already succeeded.

Compare the cost per useful hour

Here is an illustrative calculation, not a price quote. Suppose the on-demand instance costs $2/hour and spot costs $0.80/hour. If 15% of billed spot time is spent loading or repeating lost work, the effective compute cost is:

$0.80 / (1 - 0.15) = $0.94 per useful hour

Under these assumptions, spot remains cheaper on compute until wasted billed time reaches 60%. That does not mean a service with 60% waste is acceptable. Capacity waits may cost little compute while still missing a deadline, and operator time or a second replica can outweigh the saving.

WorkloadWhen spot is worth testingWhat must work first
Batch inference or evaluationsThe queue can wait and completed results are retainedDurable results, safe retries and a completion deadline
Development endpointUsers accept occasional unavailabilityClear status, restart procedure and a spending limit
Interactive production serviceOther capacity can carry traffic during recoveryTested failover, enough surviving capacity and controlled fallback costs

Use interruption history carefully

AWS Spot Instance Advisor provides historical interruption-frequency bands and savings information. Treat them as planning inputs, not the probability that your next request will fail. A price discount alone does not establish an interruption rate.

Broader placement choices can improve your chances of finding capacity, but alternatives still need enough memory, the right engine support and acceptable data residency. Two replicas also need a routing and failure-handling design; merely launching two instances does not create a dependable service.

Run a recovery drill before making spot the default

  1. Measure a normal launch and a launch using saved weights.
  2. Stop the worker during a representative request and verify how the client or queue responds.
  3. Test a replacement-capacity failure and confirm the maximum retry time.
  4. If on-demand fallback is allowed, verify the price ceiling and the return policy.
  5. Reconcile completed work, repeated work, unavailable minutes and the provider bill.

LLM Hangar's spot documentation describes its recovery and fallback controls. Match those settings to the budget and outage tolerance you have actually tested. For a deadline-bound batch, paying more for the final hours can be a sensible planned fallback.

Questions people ask

Are spot GPUs always cheaper overall?

No. Add startup, repeated work, storage, fallback capacity and operating effort. Compare the cost of completed work and whether it finishes on time.

Does an interruption always give two minutes of warning?

No. AWS describes stop and termination notices as best effort, and hibernation begins immediately. Other providers have their own rules. Recovery should also handle an abrupt loss.

Do saved model weights prevent spot downtime?

No. They can reduce loading work, but replacement capacity, engine startup and health checks still take time. Keep in-flight request recovery separate from weight storage.