Lizrd
Infrastructure Cost Review

Meridian Fleet — Cloud Infrastructure Review

Prepared for Meridian Fleet, Inc.
Reviewer Georg Hennerbichler
Scope AWS + GCP · 3 regions
Period Trailing 30 days
Confidential · Sample
Current spend
$48,200/mo
AWS $43,900 · GCP $4,300
Identified savings
$15,410/mo
32% of the bill · $185K/yr
Utilization rating
2.4/5
Underutilized
Resources analyzed
437
across 2 clouds, 3 regions
Bottom line. Meridian is paying full on-demand rates for a workload that is remarkably steady, and running non-production around the clock at production size. Nearly a third of the monthly bill is recoverable without touching the customer-facing SLA — and roughly $4.4K/mo of it is spend that wasn't attributed to any resource until this review.
01

Business context

Meridian Fleet is a B2B SaaS platform for real-time fleet management, used by mid-market logistics operators running 200–2,000-vehicle fleets. The product ingests GPS and telemetry from in-vehicle devices, powers a live operations dashboard, scores driver behaviour, and offers route optimization on its premium tier. Customers are billed per active vehicle per month. ~45 engineers.

What the tech stack is for. Three things drive the architecture and, with it, the bill: (1) low-latency telemetry ingest — a vehicle's position must appear on the dashboard in under five seconds, which is both the core SLA and the top churn driver; (2) dashboard availability for 24/7 logistics operations; and (3) the route-optimization ML that differentiates the premium tier. Infrastructure is ~22% of cost of goods sold, so cost efficiency is gross margin — every $10K/mo recovered is roughly a point of margin.

02

Infrastructure summary

Resources analyzed
437
Clouds
2
AWS · GCP
Regions
3
us-east-1 · eu-west-1 · GCP
EKS clusters
2
prod + staging
EC2 / managed nodes
38
Databases & caches
14

Architecture, as it stands.

  • A modular-monolith API runs on an EKS cluster in us-east-1 (plus a full staging cluster), behind an ALB and CloudFront. The prod node group is a fixed set of six m6i.2xlarge — no Cluster Autoscaler or Karpenter.
  • The telemetry pipeline is device → API Gateway + Lambda → MSK (Kafka) → consumer deployments → TimescaleDB on RDS Postgres for positions, with Redis (ElastiCache) holding live vehicle state.
  • Route-ML is served from a SageMaker real-time endpoint (ml.g5.xlarge) for premium customers.
  • Analytics is exported nightly to BigQuery, which backs customer-facing reporting dashboards.
  • A second region, eu-west-1, serves EU customers — but also carries a full staging environment that appears to have been left running after a migration.
03

What’s working well

Before the findings below — the parts of this setup that are in good shape, and shouldn’t change.

  • Resource tagging is broadly consistent — ~85% of spend carries an owner and environment tag, which is exactly what made this attribution possible in the first place.
  • S3 lifecycle policies are in place on the primary buckets — cold telemetry already tiers to Infrequent Access and Glacier, so storage isn’t a runaway line.
  • No security red flags surfaced incidentally — no public buckets, EBS/RDS encryption is on, and no obviously over-broad IAM wildcards in the paths reviewed. (A dedicated security review was out of scope for this cost engagement.)
  • The core telemetry path is well-isolated on its own cluster and brokers, so the recommended compute and database changes can proceed without risking the 5-second SLA.
04

Risks & bottlenecks

High
RDS primary is a single writer serving both ingest and analytics
The Postgres/Timescale primary peaks at 82% of max connections and there is no read replica — month-end customer reporting queries run against the same writer that ingests telemetry. A reporting spike during peak ingest risks connection saturation and dashboard latency breaching the 5-second SLA.
High
Kafka ingest topic is under-partitioned
The positions topic has 6 partitions against 9 consumer replicas; consumer lag spikes to ~90 seconds at peak ingest, directly threatening the "position in under 5s" SLA. The cluster is simultaneously over-provisioned on brokers (see optimization 07) — a capacity/partitioning mismatch, not a capacity shortage.
Medium
No autoscaling on the production API node group
Fixed six-node group sized for peak. Off-peak CPU sits at ~19%; there is also no headroom controller for an unexpected spike — the worst of both worlds (paying for peak, no elasticity for above-peak).
Medium
Single NAT gateway per region + chatty cross-AZ consumers
Egress is a per-region single NAT gateway (a SPOF for outbound), and the telemetry consumers talk across AZs to the brokers and database, driving both cross-AZ transfer and NAT processing charges that are invisible in the top-line service breakdown.
Medium
Zero commitment coverage on a very steady baseline
Compute and database load is strikingly flat week-to-week, yet 100% of it is billed on-demand. The baseline is predictable enough to commit against with low risk.
Low
CloudWatch logs never expire; high-cardinality custom metrics
Several log groups are set to never-expire, and per-vehicle custom metrics create high cardinality. Neither is a reliability risk, but both quietly compound.
05

Optimizations

Nine changes, ranked by monthly saving. Each lists the reason, the evidence, and the effort/risk. Expand for detail. None of these touches the customer-facing SLA; the first three are low-risk and capture most of the value.

▸ 01 · Adopt a 1-year Compute Savings Plan on the steady baseline
EKS + Fargate + RDS baseline · commitment coverage 0%
$4,000/mo

Why. Compute and database usage is flat week-over-week, but everything is billed on-demand. A 1-year, no-upfront Compute Savings Plan sized to the observed p5 floor (not the peak) converts the predictable base load to a ~30% lower rate while leaving spikes on-demand. Start at ~70% coverage of the floor to stay conservative.

0%
current coverage
~30%
rate reduction
p5
commit floor
1 yr
term, no upfront
Effort Low Risk Low Reversible commitment, not architectural
▸ 02 · Rightsize the production API node group
eks/prod-api · m6i.2xlarge ×6 → m6i.xlarge ×6
$3,100/mo

Why. 30-day CPU averages 19% and memory 34% on the prod API nodes, with peaks well within a node half the size. Halving the instance size keeps 2× headroom over observed peak. Pair with Karpenter (optimization 06 context) so the group also scales for genuine spikes instead of standing at peak 24/7.

19%
CPU 30-day avg
34%
mem 30-day avg
41%
CPU p95
6
nodes (unchanged)
Effort Low Risk Low Change node group instance type
▸ 03 · Schedule non-production to business hours
eks/staging + eu-west-1 staging + 5 dev RDS · 24/7 → ~50 h/week
$2,400/mo

Why. Staging (prod-sized) and the dev databases run around the clock but see no traffic outside working hours. A scheduled scale-to-zero / stop overnight and on weekends cuts their runtime by ~70% with zero production impact. The forgotten eu-west-1 staging environment (see hidden costs) is included here.

168 h
current runtime/wk
~50 h
needed/wk
~70%
runtime cut
0
prod impact
Effort Low Risk Low Change scheduled stop/start
▸ 04 · Move RDS to Graviton, add a read replica, downsize the primary
rds/meridian-core · db.r6i.4xlarge → db.r6g.2xlarge + 1 replica
$2,200/mo

Why. This both saves money and removes the §04 High risk. Graviton (r6g) is ~20% cheaper at equal performance; a read replica offloads analytics/reporting reads off the writer, which then safely downsizes two sizes. Net: lower cost and the writer is no longer a single point of contention at month-end.

82%
peak conn. (writer)
~20%
Graviton saving
2
sizes down
+1
read replica
Effort Medium Risk Medium — failover window Also fixes §04 writer contention
▸ 05 · Decommission the idle route-ML endpoint
sagemaker/route-opt-v2 · ml.g5.xlarge · 0 invocations in 21 days
$1,350/mo

Why. A GPU inference endpoint has served zero requests in 21 days — the premium route-optimization feature was moved to a batch job and the real-time endpoint was never torn down. Delete it (or convert to a serverless/async endpoint that scales to zero if real-time is ever needed again).

0
calls / 21 days
1%
GPU util
24/7
billed uptime
g5.xlarge
instance
Effort Low Risk Low — confirm batch path Change delete endpoint
▸ 06 · Cut NAT egress with VPC endpoints
vpc · S3 + ECR + CloudWatch via gateway/interface endpoints
$1,000/mo

Why. A large share of NAT-gateway processing is S3, ECR image pulls, and CloudWatch traffic routed out through NAT. Gateway endpoints (S3) and interface endpoints (ECR, CloudWatch, STS) keep that traffic on the AWS backbone — removing both the NAT data-processing charge and part of the cross-AZ transfer surfaced below.

61%
NAT traffic to S3/ECR
4
endpoints to add
Low
blast radius
Effort Low Risk Low Change add VPC endpoints
▸ 07 · Right-size the MSK (Kafka) cluster
msk/telemetry · 6 × m5.large → 3 × m5.large, repartition topic
$820/mo

Why. Brokers run at 28% CPU and well under half their storage throughput — the cluster is over-provisioned on brokers while the hot topic is under-partitioned (the §04 lag). Halving brokers and raising the positions topic from 6 → 18 partitions both cuts cost and resolves the consumer-lag bottleneck. Do the repartition first, validate lag, then remove brokers.

28%
broker CPU
6 → 3
brokers
6 → 18
partitions
90s → <5s
target lag
Effort Medium Risk Medium — sequence carefully Also fixes §04 consumer lag
▸ 08 · Tighten CloudWatch logs retention & metric cardinality
cloudwatch · 23 never-expire log groups + per-vehicle metrics
$360/mo

Why. 23 log groups never expire (years of debug logs retained), and per-vehicle custom metrics multiply cardinality. Set 30–90 day retention by group and aggregate the per-vehicle metrics to per-fleet dimensions. No observability lost that anyone uses.

23
never-expire groups
30–90d
target retention
per-vehicle
metric cardinality
Effort Low Risk Low
▸ 09 · Reclaim orphaned storage
ebs · 1.2 TB unattached volumes + snapshots with no source
$180/mo

Why. 1.2 TB of EBS volumes are unattached and a set of snapshots reference volumes that no longer exist — residue from past instance churn. Snapshot-then-delete the volumes; age out the orphan snapshots. Small, but free.

1.2 TB
unattached
31
orphan snapshots
Effort Low Risk Low
06

Hidden costs surfaced

Spend that doesn't show up under a recognizable service line — it hides in "EC2-Other," duplicate environments, and usage-based services. ~$4,380/mo was not attributed to any resource before this review.

Hidden cost How it hides Per month
Cross-AZ data transfer telemetry consumers ↔ brokers/DB across AZs Billed as "EC2-Other," no resource tag $2,300
Forgotten eu-west-1 staging full env left after an EU migration Spread across EC2/RDS/EKS in a second region $1,100
BigQuery full-table daily scan reporting job re-scans history nightly On-demand bytes-scanned, no partition pruning $900
Idle load balancers 2 ALBs with zero healthy targets Small per-ALB hourly charge, long forgotten $44
Unassociated Elastic IPs 4 EIPs not attached to anything Charged only because they're unused $36

Note: the cross-AZ and eu-west-1 items partly overlap with optimizations 03 and 06 — they are listed here to show where the money was hiding, not double-counted in the headline savings.

07

Utilization rating

2.4/ 5 Underutilized
AI / ML 1.2
Compute 2.1
Databases 2.6
Networking 2.8
Analytics 3.0
Storage 3.4

The rating weights each resource by spend, so the big, underused line items (prod compute, the idle GPU endpoint, the oversized writer) pull the overall score down hardest — which is exactly where the top three optimizations are aimed.

08

Prioritized action roadmap

Where to start, and in what order. The numbers below map to the optimizations in §05. Four-fifths of the savings are low-risk quick wins you can capture this week without touching the customer-facing SLA; the structural work that follows pays for itself and clears the two High risks on the way.

● quick wins (low effort, no SLA impact) · ● structural (also fixes a High risk). The spend clusters top-left: the biggest savings are the easiest to capture.

This week
Quick wins
≈ $12,390/mo · ~80%
  • 01Adopt the Compute Savings Plan
  • 02Rightsize the prod API node group
  • 03Schedule non-production to business hours
  • 05Decommission the idle route-ML endpoint
  • 06Add VPC endpoints to cut NAT egress
  • 08Tighten CloudWatch retention
  • 09Reclaim orphaned storage
This month
Structural
≈ $3,020/mo + fixes 2 High risks
  • 04RDS → Graviton, add a read replica, downsize the writer (removes the writer-contention risk)
  • 07Repartition then right-size MSK (removes the consumer-lag risk)
Ongoing
Keep it tight
sustain the savings
  • ·Review Savings Plan coverage quarterly as the baseline grows
  • ·Keep tagging + log-retention hygiene in CI so drift doesn't creep back
  • ·Re-measure realized vs. identified savings after each phase

Method. Meridian's AWS and GCP accounts were connected to Lizrd read-only. Lizrd built a knowledge graph of 437 resources, reconciled it against 30 days of billed spend and utilization telemetry, and surfaced the candidates above; each was then reviewed by hand for business context, risk, and sequencing. No changes were made to any environment — every recommendation is Meridian's to approve and implement, and Lizrd tracks realized savings as they land.

This is a sample deliverable for a fictional company. Want one for your infrastructure? Book a review at calendly.com/georg-lizrd/30min.

Produced with Lizrd · lizrd.ai · Confidential — Sample