Why this matters
- StreamHub's monthly AWS bill exceeded $2.4M before FinOps practices — 30% was waste from over-provisioned instances, orphaned resources, and un-tiered storage.
- Without cost visibility per team, no one owns the bill and optimisation becomes nobody's job.
- Spot instances, reserved capacity, and storage lifecycle policies can cut compute costs 40–70% with minimal reliability trade-offs.
FinOps practices
- Cost allocation tags — tag every resource with
team,service,environmentfor chargeback. - Right-sizing — match instance type and count to actual CPU/memory utilisation (target 60–70%).
- Reserved / savings plans — commit to 1-year usage for 30–40% discount on predictable workloads.
- Spot / preemptible instances — 60–90% cheaper for fault-tolerant batch jobs (transcoding, ML training).
- Storage tiering — lifecycle policies move cold data to cheaper tiers automatically.
FinOps observability loop
Cost visibility dashboard
| Service | Monthly cost | % of total | Optimisation opportunity |
|---|---|---|---|
| EKS compute | $980K | 41% | Right-size pods; move batch to spot |
| RDS / Aurora | $420K | 17% | Reserved instances; read replica audit |
| S3 storage | $310K | 13% | Lifecycle to Glacier; delete incomplete uploads |
| CloudFront CDN | $280K | 12% | Already optimised; cache hit ratio 94% |
| Kafka (MSK) | $190K | 8% | Reduce retention from 7d to 3d for non-critical topics |
| Other | $220K | 9% | Orphaned EBS volumes, unused Elastic IPs |
EKS compute
Monthly cost$980K% of total41%Optimisation opportunityRight-size pods; move batch to spotRDS / Aurora
Monthly cost$420K% of total17%Optimisation opportunityReserved instances; read replica auditS3 storage
Monthly cost$310K% of total13%Optimisation opportunityLifecycle to Glacier; delete incomplete uploadsCloudFront CDN
Monthly cost$280K% of total12%Optimisation opportunityAlready optimised; cache hit ratio 94%Kafka (MSK)
Monthly cost$190K% of total8%Optimisation opportunityReduce retention from 7d to 3d for non-critical topicsOther
Monthly cost$220K% of total9%Optimisation opportunityOrphaned EBS volumes, unused Elastic IPs
Example StreamHub cost breakdown — EKS compute is the largest optimisation target.
Right-sizing Kubernetes workloads
Most pods are over-provisioned. Use Vertical Pod Autoscaler (VPA) recommendations or Grafana metrics.
# Before: over-provisioned
resources:
requests:
cpu: 2000m
memory: 4Gi
limits:
cpu: 4000m
memory: 8Gi
# After: right-sized based on p95 utilisation
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: 1000m
memory: 2Gi
Spot instances for batch workloads
StreamHub's video transcoding pipeline runs on spot instances with fallback.
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: transcoder-spot
spec:
template:
spec:
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot"]
- key: node.kubernetes.io/instance-type
operator: In
values: ["c6i.4xlarge", "c6i.8xlarge"]
taints:
- key: workload
value: transcoding
effect: NoSchedule
limits:
cpu: 500
memory: 1000Gi
disruption:
consolidationPolicy: WhenUnderutilized
| Aspect | Strategy | Savings |
|---|---|---|
| Reserved instances (1yr) | 30–40% off on-demand | Predictable baseline workloads |
| Spot instances | 60–90% off on-demand | Fault-tolerant batch, stateless workers |
| Savings Plans | Flexible commitment across instance families | Mixed workload environments |
| Graviton (ARM) | 20% better price-performance | Java/Go services with ARM builds |
| S3 Intelligent-Tiering | Auto-moves to cheaper tier | Unpredictable access patterns |
Reserved instances (1yr)
Strategy30–40% off on-demandSavingsPredictable baseline workloadsSpot instances
Strategy60–90% off on-demandSavingsFault-tolerant batch, stateless workersSavings Plans
StrategyFlexible commitment across instance familiesSavingsMixed workload environmentsGraviton (ARM)
Strategy20% better price-performanceSavingsJava/Go services with ARM buildsS3 Intelligent-Tiering
StrategyAuto-moves to cheaper tierSavingsUnpredictable access patterns
Combine strategies: reserved for baseline, spot for burst, Graviton where compatible.
Storage lifecycle automation
{
"Rules": [
{
"ID": "tier-streamhub-logs",
"Filter": { "Prefix": "logs/" },
"Status": "Enabled",
"Transitions": [
{ "Days": 30, "StorageClass": "STANDARD_IA" },
{ "Days": 90, "StorageClass": "GLACIER_IR" }
],
"Expiration": { "Days": 365 }
},
{
"ID": "cleanup-temp-uploads",
"Filter": { "Prefix": "uploads/temp/" },
"Status": "Enabled",
"Expiration": { "Days": 7 }
}
]
}
Eliminating waste
Common waste sources
- Orphaned EBS volumes — detached volumes from terminated instances; automate cleanup with Lambda.
- Over-retained logs — 7-day Kafka retention for debug topics that nobody reads after 24 hours.
- Idle load balancers — NLBs/ALBs with zero traffic still cost $20/month each.
- Uncompressed data — enable S3 compression for JSON logs; 5x storage reduction.
- NAT gateway charges — $0.045/GB processed; use VPC endpoints for S3/DynamoDB to bypass NAT.
Quick recall
Everything you need if you only revisit this box.
- FinOps = cost visibility + accountability + continuous optimisation, not one-time audits.
- Tag every resource for chargeback; right-size pods to 60–70% utilisation.
- Reserved instances for baseline; spot for fault-tolerant batch; Graviton for compatible workloads.
- S3 lifecycle policies auto-tier cold data; delete temp uploads after 7 days.
- Weekly cost reviews and waste audits (orphaned volumes, idle LBs, over-retained logs).
Test yourself
Answer these before moving on — recall is what makes it stick.