🦎 Lizrd is in private beta. Request early access →

Rightsizing databases without a 3 a.m. page

The Lizrd team · · 3 min read

Nobody wants to shrink the database. It’s stateful, it’s on the critical path, and the failure mode isn’t a slow endpoint — it’s an outage with your name on it. So databases get provisioned for the worst day imaginable and then left there, permanently, at a size the workload has never actually needed.

Which makes them both the scariest thing to rightsize and often the most over-provisioned line on the bill. You can do it safely — you just have to read the right signals instead of trusting the instance class someone picked at launch.

CPU is the least interesting number

For stateless compute, CPU tells the story. For a database, CPU is often the least useful signal — a healthy database can sit at 10% CPU and still be correctly sized, because it’s constrained by something else. Look at the full picture before you touch anything:

  • Connections — peak and sustained. A pool maxing out points to a client-side problem, not an undersized DB.
  • IOPS and throughput — are you actually using the provisioned I/O, or paying for headroom you never hit?
  • Freeable memory — the working set has to fit. This is the real guardrail on shrinking; cross it and you thrash.
  • CPU credits (on burstable classes) — a balance that never depletes means the instance is oversized for its burst pattern.

Two levers, very different risk

There are two distinct moves, and treating them the same is how the page happens:

- instance_class = "db.r6g.2xlarge"
+ instance_class = "db.r6g.xlarge"

Scaling the instance class down is the big saving and the higher-risk one — you’re reducing memory and CPU, so the working-set check matters. Do it with a read replica promoted in a maintenance window, or on a Multi-AZ standby, so there’s a rollback that isn’t “restore from backup.”

Moving storage from io1/io2 to gp3, or dropping over-provisioned IOPS to match real usage, is often a large saving at near-zero risk — I/O is decoupled from the compute your queries run on.

Verify against latency, not vibes

After the change, the question isn’t “is it up?” — it’s “did P99 query latency move?” Watch it across a full business cycle, including the daily peak and any batch jobs. An estimated saving that quietly added 40ms to every checkout isn’t a saving. Closing that loop — comparing realized latency and realized cost against the estimate — is what makes the next rightsizing an easy yes instead of a fight.

Make the evidence do the arguing

The reason database rightsizing stalls is that it’s an argument, and the person defending the current size has fear on their side and no data on the other. Flip that: bring 30 days of connections, memory, and IOPS to the table and the oversize is no longer a judgment call.

That’s what Lizrd assembles for you. It reads database utilization and cost continuously, identifies the instances whose real signals — not their CPU — say they’re oversized, and proposes the class or storage change as an exact diff with the evidence and rollback attached. Rightsizing the database stops being a 3 a.m. risk when the data makes the call.

Keep reading

Stop reading about savings — find yours

Connect read-only and Lizrd surfaces your highest-impact fixes, with the exact change to make.