🦎 Lizrd is in private beta. Request early access →

Use case

Move batch and fine-tuning jobs to Spot GPUs

Find interruptible GPU workloads paying the full on-demand rate, and get the exact change to run them on managed Spot capacity — with automatic resume.

Cost Optimization $6,932/mo found
1
Rightsize Downsize the prod EKS node group

eks · prod-workers

$4,800/mo

High

2
Idle Stop the idle staging database

rds · analytics-staging

$1,180/mo

Medium

3
Rightsize Trim the over-provisioned checkout API

ecs · checkout-api

$640/mo

High

4
Orphaned Delete 14 unattached EBS volumes

ec2 · 14 × gp3

$312/mo

High

The problem

Training, fine-tuning, and batch-scoring jobs run on-demand at the premium GPU rate even when they don’t need to. These jobs are often already interruptible — they checkpoint, they run off-hours, they can retry — which makes them ideal for Spot capacity at up to ~70% off. But spotting which jobs qualify, and doing it safely, is rarely anyone’s job.

How Lizrd fixes it

Lizrd identifies GPU workloads that look batch-shaped and interruptible read-only, and recommends moving them to managed Spot training (SageMaker, Vertex, Azure ML) with automatic resume — with the evidence, the confidence level, and the exact configuration change. On-demand stays the default for anything latency-sensitive; only genuinely interruptible work is flagged.

The outcome

The same jobs finish for a fraction of the GPU cost, with checkpoint-and-resume handling interruptions, and Lizrd validates the realized savings against your bill.

“Our nightly fine-tuning already checkpointed — it just never occurred to anyone it could run on Spot for a fraction of the price.”
— Engineering lead, AI startup ~$1.8k/mo

More ways to save

Find this in your own cloud

Connect read-only and Lizrd surfaces the highest-impact fixes — with the exact change to make.