Why this matters
- StreamHub's checkout path touches 6 services — without distributed tracing, a 3-second latency spike is impossible to diagnose from logs alone.
- SLOs and error budgets turn observability data into engineering priorities: burn rate alerts fire before users notice.
- OpenTelemetry (OTel) is the vendor-neutral standard in 2024–2026; instrument once, export to Prometheus, Jaeger, Grafana, or Datadog.
Observability stack components
- Instrumentation — OTel SDK in each service; auto-instrument HTTP, gRPC, DB, Kafka.
- Collector — receives OTLP data; processes, batches, and routes to backends.
- Metrics backend — Prometheus scrapes or receives remote-write; Grafana visualises.
- Logs backend — Loki, Elasticsearch, or CloudWatch Logs; correlated by trace ID.
- Traces backend — Jaeger, Tempo, or Honeycomb; distributed request timelines.
- Alerting — Alertmanager or Grafana Alerting → PagerDuty/Slack on SLO breach.
Pipeline architecture
Observability on AWS
OpenTelemetry configuration
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 1024
memory_limiter:
check_interval: 1s
limit_mib: 512
exporters:
prometheus:
endpoint: 0.0.0.0:8889
otlp/jaeger:
endpoint: jaeger:4317
tls:
insecure: true
loki:
endpoint: http://loki:3100/loki/api/v1/push
service:
pipelines:
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [prometheus]
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp/jaeger]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]
Application instrumentation
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(
endpoint="otel-collector:4317"
)))
trace.set_tracer_provider(provider)
FastAPIInstrumentor.instrument_app(app)
SLOs and error budgets
Define what "good" means, then measure against it.
| Service | SLI | SLO target | Error budget (30d) |
|---|---|---|---|
| API gateway | Availability (non-5xx) | 99.95% | 21.6 min downtime |
| Video playback | Startup latency p99 | < 2s | 0.1% of requests |
| Live chat | Message delivery p99 | < 500ms | 0.1% of messages |
| Search | Query latency p95 | < 200ms | 0.5% of queries |
API gateway
SLIAvailability (non-5xx)SLO target99.95%Error budget (30d)21.6 min downtimeVideo playback
SLIStartup latency p99SLO target< 2sError budget (30d)0.1% of requestsLive chat
SLIMessage delivery p99SLO target< 500msError budget (30d)0.1% of messagesSearch
SLIQuery latency p95SLO target< 200msError budget (30d)0.5% of queries
Error budget = (1 - SLO) × time window. Burn it on risky deploys; protect it during peak events.
# Grafana alerting rule
groups:
- name: streamhub-slos
rules:
- alert: APIErrorBudgetBurn
expr: |
1 - (
sum(rate(http_requests_total{status!~"5.."}[1h]))
/ sum(rate(http_requests_total[1h]))
) > 0.0005
for: 5m
labels:
severity: critical
annotations:
summary: "API error budget burning at >2x normal rate"
Structured logging with trace correlation
{
"timestamp": "2026-04-01T14:32:10.123Z",
"level": "ERROR",
"service": "streamhub-payment",
"trace_id": "abc123def456",
"span_id": "789ghi",
"message": "Payment gateway timeout",
"user_id": "usr_42",
"amount_cents": 1999,
"duration_ms": 5023
}
Inject trace_id into every log line via MDC (Mapped Diagnostic Context). Click a trace in Jaeger → jump to correlated logs in Loki.
| Aspect | Signal | Best for |
|---|---|---|
| Metrics | Aggregated counters, gauges, histograms | Dashboards, alerting, capacity planning |
| Logs | Discrete events with context | Debugging specific failures, audit trails |
| Traces | Request flow across services | Latency analysis, dependency mapping |
| Profiles | CPU/memory flame graphs | Performance optimisation, hot path analysis |
Metrics
SignalAggregated counters, gauges, histogramsBest forDashboards, alerting, capacity planningLogs
SignalDiscrete events with contextBest forDebugging specific failures, audit trailsTraces
SignalRequest flow across servicesBest forLatency analysis, dependency mappingProfiles
SignalCPU/memory flame graphsBest forPerformance optimisation, hot path analysis
Use all three together — metrics alert, traces narrow, logs confirm.
Quick recall
Everything you need if you only revisit this box.
- Three pillars: metrics (aggregated), logs (events), traces (request flow) — unified by OpenTelemetry.
- OTel Collector receives OTLP, processes, and routes to Prometheus/Jaeger/Loki backends.
- SLOs define reliability targets; error budgets quantify allowed downtime.
- Correlate logs with traces via trace_id in every log line.
- Sample traces in production (10%); always capture errors at 100%.
Test yourself
Answer these before moving on — recall is what makes it stick.