PrepZone Logo
PrepZone

Cost Optimization and FinOps

Right-size instances, spot capacity, storage tiers and the metrics that keep cloud bills predictable.

Why this matters

  • StreamHub's monthly AWS bill exceeded $2.4M before FinOps practices — 30% was waste from over-provisioned instances, orphaned resources, and un-tiered storage.
  • Without cost visibility per team, no one owns the bill and optimisation becomes nobody's job.
  • Spot instances, reserved capacity, and storage lifecycle policies can cut compute costs 40–70% with minimal reliability trade-offs.

FinOps practices

  • Cost allocation tags — tag every resource with team, service, environment for chargeback.
  • Right-sizing — match instance type and count to actual CPU/memory utilisation (target 60–70%).
  • Reserved / savings plans — commit to 1-year usage for 30–40% discount on predictable workloads.
  • Spot / preemptible instances — 60–90% cheaper for fault-tolerant batch jobs (transcoding, ML training).
  • Storage tiering — lifecycle policies move cold data to cheaper tiers automatically.

FinOps observability loop

COMPUTE
AWS resourcestagged by team
OPS
Cost Explorerdaily spend
OPS
CloudWatchutilization
OPS
Compute Optimiz…right-size
EXTERNAL
ActionsSpot · RI · GP3
Tagged resources → Cost Explorer + Compute Optimizer → right-sizing actions.

Cost visibility dashboard

ServiceMonthly cost% of totalOptimisation opportunity
EKS compute$980K41%Right-size pods; move batch to spot
RDS / Aurora$420K17%Reserved instances; read replica audit
S3 storage$310K13%Lifecycle to Glacier; delete incomplete uploads
CloudFront CDN$280K12%Already optimised; cache hit ratio 94%
Kafka (MSK)$190K8%Reduce retention from 7d to 3d for non-critical topics
Other$220K9%Orphaned EBS volumes, unused Elastic IPs
  • EKS compute

    Monthly cost$980K
    % of total41%
    Optimisation opportunityRight-size pods; move batch to spot
  • RDS / Aurora

    Monthly cost$420K
    % of total17%
    Optimisation opportunityReserved instances; read replica audit
  • S3 storage

    Monthly cost$310K
    % of total13%
    Optimisation opportunityLifecycle to Glacier; delete incomplete uploads
  • CloudFront CDN

    Monthly cost$280K
    % of total12%
    Optimisation opportunityAlready optimised; cache hit ratio 94%
  • Kafka (MSK)

    Monthly cost$190K
    % of total8%
    Optimisation opportunityReduce retention from 7d to 3d for non-critical topics
  • Other

    Monthly cost$220K
    % of total9%
    Optimisation opportunityOrphaned EBS volumes, unused Elastic IPs

Example StreamHub cost breakdown — EKS compute is the largest optimisation target.

Right-sizing Kubernetes workloads

Most pods are over-provisioned. Use Vertical Pod Autoscaler (VPA) recommendations or Grafana metrics.

Java
# Before: over-provisioned
resources:
  requests:
    cpu: 2000m
    memory: 4Gi
  limits:
    cpu: 4000m
    memory: 8Gi

# After: right-sized based on p95 utilisation
resources:
  requests:
    cpu: 500m
    memory: 1Gi
  limits:
    cpu: 1000m
    memory: 2Gi

Spot instances for batch workloads

StreamHub's video transcoding pipeline runs on spot instances with fallback.

Java
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
  name: transcoder-spot
spec:
  template:
    spec:
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot"]
        - key: node.kubernetes.io/instance-type
          operator: In
          values: ["c6i.4xlarge", "c6i.8xlarge"]
      taints:
        - key: workload
          value: transcoding
          effect: NoSchedule
  limits:
    cpu: 500
    memory: 1000Gi
  disruption:
    consolidationPolicy: WhenUnderutilized
AspectStrategySavings
Reserved instances (1yr)30–40% off on-demandPredictable baseline workloads
Spot instances60–90% off on-demandFault-tolerant batch, stateless workers
Savings PlansFlexible commitment across instance familiesMixed workload environments
Graviton (ARM)20% better price-performanceJava/Go services with ARM builds
S3 Intelligent-TieringAuto-moves to cheaper tierUnpredictable access patterns
  • Reserved instances (1yr)

    Strategy30–40% off on-demand
    SavingsPredictable baseline workloads
  • Spot instances

    Strategy60–90% off on-demand
    SavingsFault-tolerant batch, stateless workers
  • Savings Plans

    StrategyFlexible commitment across instance families
    SavingsMixed workload environments
  • Graviton (ARM)

    Strategy20% better price-performance
    SavingsJava/Go services with ARM builds
  • S3 Intelligent-Tiering

    StrategyAuto-moves to cheaper tier
    SavingsUnpredictable access patterns

Combine strategies: reserved for baseline, spot for burst, Graviton where compatible.

Storage lifecycle automation

Java
{
  "Rules": [
    {
      "ID": "tier-streamhub-logs",
      "Filter": { "Prefix": "logs/" },
      "Status": "Enabled",
      "Transitions": [
        { "Days": 30, "StorageClass": "STANDARD_IA" },
        { "Days": 90, "StorageClass": "GLACIER_IR" }
      ],
      "Expiration": { "Days": 365 }
    },
    {
      "ID": "cleanup-temp-uploads",
      "Filter": { "Prefix": "uploads/temp/" },
      "Status": "Enabled",
      "Expiration": { "Days": 7 }
    }
  ]
}

Eliminating waste

Common waste sources

  • Orphaned EBS volumes — detached volumes from terminated instances; automate cleanup with Lambda.
  • Over-retained logs — 7-day Kafka retention for debug topics that nobody reads after 24 hours.
  • Idle load balancers — NLBs/ALBs with zero traffic still cost $20/month each.
  • Uncompressed data — enable S3 compression for JSON logs; 5x storage reduction.
  • NAT gateway charges — $0.045/GB processed; use VPC endpoints for S3/DynamoDB to bypass NAT.

Quick recall

Everything you need if you only revisit this box.

  • FinOps = cost visibility + accountability + continuous optimisation, not one-time audits.
  • Tag every resource for chargeback; right-size pods to 60–70% utilisation.
  • Reserved instances for baseline; spot for fault-tolerant batch; Graviton for compatible workloads.
  • S3 lifecycle policies auto-tier cold data; delete temp uploads after 7 days.
  • Weekly cost reviews and waste audits (orphaned volumes, idle LBs, over-retained logs).

Test yourself

Answer these before moving on — recall is what makes it stick.