What we optimize
Where Lizrd finds cloud waste
The areas Lizrd continuously explores across AWS, GCP, and Azure. Every finding comes with the evidence behind it, a confidence level, effort & risk, and a copy-paste fix — and the saving is validated against your real bill after you apply it.
Click any area to expand — what it is, what we detect, and a real example.
Over-provisioning Paying for capacity you don't use Resources sized "to be safe" and never revisited — usually the biggest, lowest-risk saving.
Infrastructure gets sized once — generously — and then forgotten. Multiply a little headroom across a fleet and it becomes one of the largest lines on the bill, yet it’s invisible in a dashboard that only shows totals. It’s the biggest, lowest-risk saving there is, and it grows silently as you scale.
What Lizrd explores
- ◆ 30-day CPU, memory, and true daily-peak on compute, managed databases, caches, and container tasks.
- ◆ Serverless functions carrying more memory than they need.
- ◆ A genuinely-needed peak vs. a rare burst, so it never strips headroom a workload really uses.
How it helps → names the exact smaller instance/class/allocation, with the diff and an easy rollback.
Example
The problem: a staging API on m6i.2xlarge averages 8% CPU and 19% memory over 30 days, with a single 34% spike one day a month.
Lizrd proposes: downsize to m6i.large — keeps headroom for the monthly spike — with the exact Terraform diff and one-line rollback. ≈ $210/mo
Coverage: AWS · GCP · Azure — VMs, RDS/Cloud SQL/Azure DB (incl. DocumentDB, Neptune, Spanner, Bigtable, AlloyDB, Cosmos), caches, ECS tasks, Lambda/Cloud Functions.
Idle & orphaned Resources doing no work at all Experiments, leftovers, and unused resources still billing 24/7 for nothing.
The easiest waste to create and the hardest to notice: things left over from experiments, migrations, and one-off tests that keep billing around the clock. Nobody misses what nothing is using — which is why it lingers for months. Pure, near-zero-risk savings.
What Lizrd explores
- ◆ Compute, databases, and caches at near-zero utilization.
- ◆ Unattached disks, stale snapshots, and allocated-but-unused IPs.
- ◆ Load balancers and NAT gateways passing no meaningful traffic.
- ◆ Whether a resource is owned-but-idle (stop/downsize) or abandoned (safe to remove).
How it helps → the concrete action for each — stop, downsize, or delete — with a one-command fix.
Example
The problem: a db.r6g.large "analytics-staging" database shows 1% CPU and 0 connections for 3 weeks.
Lizrd proposes: stop it, or downsize to db.t4g.medium if it’s still needed — with the evidence and a reversible change. ≈ $1,180/mo
Coverage: AWS deepest; core idle detection across GCP · Azure.
Storage The wrong media, tier, and performance Older disk types, over-provisioned IOPS, and cold data on hot tiers — all in-place fixes.
Storage quietly overspends three ways: older/pricier disk types with a cheaper equivalent, performance provisioned far above what’s used, and cold data on hot tiers. Individually small, collectively large, almost never revisited — and every fix is in-place with no application risk.
What Lizrd explores
- ◆ Volumes on older/costlier media with a drop-in cheaper equivalent (gp2→gp3, pd-ssd→pd-balanced, Premium→Standard SSD).
- ◆ Provisioned IOPS/throughput far above actual usage.
- ◆ Buckets holding cold data with no lifecycle or auto-tiering (S3 Intelligent-Tiering, GCS Autoclass).
How it helps → the exact SKU or tier change, applied in place — no downtime.
Example
The problem: a 4 TB gp2 volume is on last-generation media with the same workload demand.
Lizrd proposes: switch to gp3 — ~20% cheaper per GB with a higher performance floor, changed in place with no downtime. ≈ $80/mo
Coverage: AWS · GCP · Azure.
Modernization Better price/performance, same workload The same workload runs ~20% cheaper on Arm silicon.
The same workload frequently runs ~20% cheaper on Arm silicon (Graviton, Cobalt, T2A), but the migration gets deprioritized because "it works today." A durable win on compute you’re already paying for.
What Lizrd explores
- ◆ Instances and functions whose shape has a drop-in Arm equivalent and no architecture blocker.
How it helps → flags the candidates and the migration path, with the price delta.
Example
The problem: a group of m5.xlarge web nodes runs a container image that already builds for Arm.
Lizrd proposes: move to m6g.xlarge (Graviton) — ~20% cheaper compute for the same performance. ≈ $640/mo
Coverage: AWS (Graviton) · GCP (T2A) · Azure (Cobalt).
Commitments & pricing Full on-demand rates on steady spend Predictable, always-on usage that a commitment would discount up to ~64%.
Predictable, always-on usage billed at full on-demand rates can be discounted up to ~64% with the right commitment — but commitments are easy to get wrong in both directions. It’s the biggest percentage lever on predictable spend.
What Lizrd explores
- ◆ Sustained run-rate and uptime to size Savings Plans / Reserved Instances (AWS), Committed Use Discounts (GCP), Reservations (Azure).
- ◆ Steady, well-utilized data warehouses suited to reserved nodes or slot commitments.
- ◆ Commitments only on genuinely steady spend — never spiky workloads.
How it helps → the right commitment type and coverage level, grounded in real utilization.
Example
The problem: ~$18k/mo of EC2 has run flat at 92% uptime for 60 days with zero commitment coverage.
Lizrd proposes: a 1-year compute Savings Plan sized to that steady baseline (not the spiky top) at up to ~40% off. ≈ $6,000/mo
Coverage: AWS · GCP · Azure.
Serverless & consumption billing When "pay for what you use" quietly breaks Always-warm capacity and the wrong billing mode — pure billing fixes, no code change.
Serverless promises you only pay for what you invoke — but always-warm capacity and the wrong billing mode quietly undo that. These are pure billing-model fixes: no code change, immediate effect.
What Lizrd explores
- ◆ Provisioned concurrency billed 24/7 but barely used.
- ◆ NoSQL tables on provisioned capacity with spiky or low traffic that belong on on-demand.
How it helps → reduce or remove the warm capacity, or switch the billing mode.
Example
The problem: a Lambda holds 50 units of provisioned (always-warm) concurrency billed 24/7 but peaks at 6.
Lizrd proposes: cut it to 8 — covers the peak with margin, drops the round-the-clock charge. ≈ $430/mo
Coverage: AWS (Lambda, DynamoDB).
Data warehouses High unit cost, hidden inefficiency Idle clusters and inefficient scans, hidden behind one "analytics" line.
Warehouses are expensive per hour and per terabyte scanned, so idle clusters and inefficient queries translate into large dollars fast — and both hide behind a single "analytics" line.
What Lizrd explores
- ◆ Clusters idle for a sustained window that could be paused (data preserved).
- ◆ Over-provisioned nodes or slots.
- ◆ Projects that scan far more than they store — a partition/cluster signal.
How it helps → pause, right-size (snapped to valid increments), commit, or restructure the scan pattern.
Example
The problem: a 4-node Redshift cluster sits below 5% utilization every night and weekend but bills 24/7.
Lizrd proposes: pause it off-hours (data preserved, resumes on demand) or drop to 2 nodes. ≈ $2,100/mo
Coverage: AWS (Redshift) · GCP (BigQuery).
Kubernetes Where over-provisioning multiplies Padded pod requests and under-packed nodes — waste that scales per replica.
Clusters are where padding compounds: every pod’s over-sized request is paid for on a node, and under-packed nodes bill for empty capacity. Requests are set once and rarely revisited, so the waste scales with every replica you run.
What Lizrd explores
- ◆ Deployments/StatefulSets requesting far more CPU/memory than they use, from real cluster metrics.
- ◆ Idle workloads that could scale to zero.
- ◆ Under-packed node pools where nodes can be consolidated and reclaimed.
How it helps → the exact resources.requests edit in your manifest — linked to the owning file, template-proof via Argo/Flux — plus node reclaim.
Example
The problem: a payments-api Deployment requests 2 vCPU / 4 GiB per pod but uses ~0.3 / 0.8 across 12 replicas.
Lizrd proposes: lower requests to 0.5 vCPU / 1 GiB — the exact edit in payments/deployment.yaml. ≈ $1,300/mo
Coverage: AWS (EKS) · GCP (GKE).
Scheduling & elasticity Flat capacity for demand that isn't flat Non-prod idling overnight, or spiky workloads paying peak-sized capacity 24/7.
A lot of capacity runs 24/7 for demand that clearly isn’t — non-prod idling overnight, or spiky workloads paying peak-sized capacity around the clock. Matching capacity to the actual demand curve is a large, safe saving most teams never get to.
What Lizrd explores
- ◆ Non-prod compute running outside business hours.
- ◆ Workloads with a low baseline and recurring, predictable spikes that fit autoscaling.
- ◆ The inverse, for reliability: resources saturating their ceiling that should scale up before customers feel it.
How it helps → off-hours schedules, a move to dynamic scaling (ASG / MIG / VM Scale Set / ECS), or a timely scale-up.
Example
The problem: 20 non-prod EC2 instances run 24/7 but are only used ~9–6 on weekdays.
Lizrd proposes: a stop/start schedule for evenings + weekends — ~65% fewer running hours. ≈ $3,400/mo
Coverage: AWS · GCP · Azure.
AI & GPU infrastructure The fastest-growing, easiest-to-overpay line Idle GPUs and inference endpoints — the highest-dollar waste you can surface.
GPU and inference is now a top-growing cost center — and the easiest to overpay on. Average GPU utilization runs shockingly low and idle endpoints bill around the clock; because a GPU is the most expensive unit of compute you buy, a low-utilization one is the highest-dollar waste you can surface.
What Lizrd explores
- ◆ Real GPU + VRAM utilization and endpoint traffic across SageMaker & Bedrock, Vertex AI, Azure ML & OpenAI.
- ◆ Idle GPU instances, and inference endpoints and notebooks billing 24/7 with no traffic.
- ◆ Training-class accelerators (A100/H100) doing inference work a cheaper one would handle.
- ◆ Bursty endpoints that should scale to zero; interruptible jobs that belong on Spot.
- ◆ For managed LLMs: token traffic that’s uncached, un-batched, or over-powered for the task.
How it helps → the specific move — stop, right-accelerator, scale-to-zero, Spot, commitment, or the token-economics change — each backed by the measured signal.
Example
The problem: a SageMaker endpoint fraud-scoring-v2 on ml.g5.xlarge shows 1% GPU and 0 invocations for 14 days.
Lizrd proposes: delete it (model + config retained) or move it behind a scale-to-zero config — savings start immediately. ≈ $2,190/mo
Coverage: AWS · GCP · Azure.
Safe by design
How every finding stays trustworthy
- Read-only & you approve everything — Lizrd never changes your infrastructure; every fix is a proposal you apply.
- Evidence + earned confidence — each finding shows the metrics behind it and a confidence level — no black-box guesses.
- Downsize-safety vetoes — a resource that looks idle on CPU but is pushing heavy disk I/O or network is held back from a downsize.
- Burst-aware — a rare monthly spike won’t distort a right-size; recurring peaks keep headroom, paired with autoscaling.
- Proven savings — after you apply a change, the realized saving is validated against your actual bill.
Work with the founder
Want an expert pass over your infrastructure?
Book a founder-led infrastructure review — a prioritized, explained plan to cut your cloud bill. Starts with a free 30-minute consultation.
See what Lizrd finds in your cloud
Connect read-only and get your first prioritized optimizations — the exact fix included.