PrepZone Logo
PrepZone

Observability Stack

Metrics, logs, traces, SLOs and error budgets with OpenTelemetry and modern dashboards.

Why this matters

  • StreamHub's checkout path touches 6 services — without distributed tracing, a 3-second latency spike is impossible to diagnose from logs alone.
  • SLOs and error budgets turn observability data into engineering priorities: burn rate alerts fire before users notice.
  • OpenTelemetry (OTel) is the vendor-neutral standard in 2024–2026; instrument once, export to Prometheus, Jaeger, Grafana, or Datadog.

Observability stack components

  • Instrumentation — OTel SDK in each service; auto-instrument HTTP, gRPC, DB, Kafka.
  • Collector — receives OTLP data; processes, batches, and routes to backends.
  • Metrics backend — Prometheus scrapes or receives remote-write; Grafana visualises.
  • Logs backend — Loki, Elasticsearch, or CloudWatch Logs; correlated by trace ID.
  • Traces backend — Jaeger, Tempo, or Honeycomb; distributed request timelines.
  • Alerting — Alertmanager or Grafana Alerting → PagerDuty/Slack on SLO breach.

Pipeline architecture

Observability on AWS

OTLPalertCOMPUTE
EKS podsOTel SDK
OPS
ADOT collectorOTLP
OPS
CloudWatchmetrics · logs
OPS
X-Raytraces
OPS
Managed Grafanadashboards
EXTERNAL
PagerDutySLO alerts
OpenTelemetry from EKS → ADOT → CloudWatch + X-Ray + Grafana → PagerDuty.

OpenTelemetry configuration

Java
# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 5s
    send_batch_size: 1024
  memory_limiter:
    check_interval: 1s
    limit_mib: 512

exporters:
  prometheus:
    endpoint: 0.0.0.0:8889
  otlp/jaeger:
    endpoint: jaeger:4317
    tls:
      insecure: true
  loki:
    endpoint: http://loki:3100/loki/api/v1/push

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [prometheus]
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp/jaeger]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [loki]

Application instrumentation

Java
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(
    endpoint="otel-collector:4317"
)))
trace.set_tracer_provider(provider)
FastAPIInstrumentor.instrument_app(app)

SLOs and error budgets

Define what "good" means, then measure against it.

ServiceSLISLO targetError budget (30d)
API gatewayAvailability (non-5xx)99.95%21.6 min downtime
Video playbackStartup latency p99< 2s0.1% of requests
Live chatMessage delivery p99< 500ms0.1% of messages
SearchQuery latency p95< 200ms0.5% of queries
  • API gateway

    SLIAvailability (non-5xx)
    SLO target99.95%
    Error budget (30d)21.6 min downtime
  • Video playback

    SLIStartup latency p99
    SLO target< 2s
    Error budget (30d)0.1% of requests
  • Live chat

    SLIMessage delivery p99
    SLO target< 500ms
    Error budget (30d)0.1% of messages
  • Search

    SLIQuery latency p95
    SLO target< 200ms
    Error budget (30d)0.5% of queries

Error budget = (1 - SLO) × time window. Burn it on risky deploys; protect it during peak events.

Java
# Grafana alerting rule
groups:
  - name: streamhub-slos
    rules:
      - alert: APIErrorBudgetBurn
        expr: |
          1 - (
            sum(rate(http_requests_total{status!~"5.."}[1h]))
            / sum(rate(http_requests_total[1h]))
          ) > 0.0005
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "API error budget burning at >2x normal rate"

Structured logging with trace correlation

Java
{
  "timestamp": "2026-04-01T14:32:10.123Z",
  "level": "ERROR",
  "service": "streamhub-payment",
  "trace_id": "abc123def456",
  "span_id": "789ghi",
  "message": "Payment gateway timeout",
  "user_id": "usr_42",
  "amount_cents": 1999,
  "duration_ms": 5023
}

Inject trace_id into every log line via MDC (Mapped Diagnostic Context). Click a trace in Jaeger → jump to correlated logs in Loki.

AspectSignalBest for
MetricsAggregated counters, gauges, histogramsDashboards, alerting, capacity planning
LogsDiscrete events with contextDebugging specific failures, audit trails
TracesRequest flow across servicesLatency analysis, dependency mapping
ProfilesCPU/memory flame graphsPerformance optimisation, hot path analysis
  • Metrics

    SignalAggregated counters, gauges, histograms
    Best forDashboards, alerting, capacity planning
  • Logs

    SignalDiscrete events with context
    Best forDebugging specific failures, audit trails
  • Traces

    SignalRequest flow across services
    Best forLatency analysis, dependency mapping
  • Profiles

    SignalCPU/memory flame graphs
    Best forPerformance optimisation, hot path analysis

Use all three together — metrics alert, traces narrow, logs confirm.

Quick recall

Everything you need if you only revisit this box.

  • Three pillars: metrics (aggregated), logs (events), traces (request flow) — unified by OpenTelemetry.
  • OTel Collector receives OTLP, processes, and routes to Prometheus/Jaeger/Loki backends.
  • SLOs define reliability targets; error budgets quantify allowed downtime.
  • Correlate logs with traces via trace_id in every log line.
  • Sample traces in production (10%); always capture errors at 100%.

Test yourself

Answer these before moving on — recall is what makes it stick.