Why this matters
- At StreamHub scale (40+ microservices), implementing retries, timeouts, and mutual TLS in every language and framework is unmaintainable.
- Istio or Linkerd sidecars give platform engineers uniform observability and security without asking each team to reinvent the wheel.
- Canary deployments and fault injection become configuration changes, not code changes.
Service mesh layers
- Data plane — Envoy sidecar proxies on every pod; handle actual traffic routing, encryption, metrics.
- Control plane — Istiod configures all sidecars; distributes certificates, routing rules, policies.
- mTLS — automatic mutual TLS between services; no application code changes required.
- Traffic management — weighted routing, retries, timeouts, circuit breakers via VirtualService.
- Observability — automatic trace propagation, request metrics, access logs from every hop.
Sidecar architecture
Istio service mesh on EKS
Traffic never flows directly from Service A to Service B. Both pods communicate through their local Envoy sidecars, which the control plane configures.
apiVersion: apps/v1
kind: Deployment
metadata:
name: streamhub-catalog
spec:
template:
metadata:
annotations:
sidecar.istio.io/inject: "true"
spec:
containers:
- name: catalog
image: registry.streamhub.io/catalog:3.1.0
ports:
- containerPort: 8080
# Envoy sidecar injected automatically by Istio
mTLS without code changes
Istio issues short-lived certificates to each sidecar. All inter-service traffic encrypts automatically.
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: streamhub
spec:
mtls:
mode: STRICT
Traffic management and canary deploys
Route 95% of traffic to the stable version and 5% to the canary.
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: streamhub-api
spec:
hosts:
- streamhub-api
http:
- route:
- destination:
host: streamhub-api
subset: stable
weight: 95
- destination:
host: streamhub-api
subset: canary
weight: 5
retries:
attempts: 3
perTryTimeout: 2s
retryOn: 5xx,reset,connect-failure
timeout: 10s
---
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: streamhub-api
spec:
host: streamhub-api
subsets:
- name: stable
labels:
version: v2.8.0
- name: canary
labels:
version: v2.8.1
Circuit breaking
Prevent cascading failures when a downstream service is unhealthy.
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: streamhub-payment
spec:
host: streamhub-payment
trafficPolicy:
connectionPool:
tcp:
maxConnections: 100
http:
h2UpgradePolicy: DO_NOT_UPGRADE
http1MaxPendingRequests: 50
maxRequestsPerConnection: 10
outlierDetection:
consecutive5xxErrors: 5
interval: 30s
baseEjectionTime: 60s
maxEjectionPercent: 50
| Aspect | Concern | Without mesh |
|---|---|---|
| mTLS | Manual cert management per service | Automatic via sidecar |
| Retries | Duplicated in every HTTP client | VirtualService policy |
| Canary deploy | Custom load balancer rules | Weight-based routing |
| Distributed tracing | Manual header propagation | Auto-injected trace context |
| Cost | Zero sidecar overhead | ~50-100 MB RAM per sidecar |
mTLS
ConcernManual cert management per serviceWithout meshAutomatic via sidecarRetries
ConcernDuplicated in every HTTP clientWithout meshVirtualService policyCanary deploy
ConcernCustom load balancer rulesWithout meshWeight-based routingDistributed tracing
ConcernManual header propagationWithout meshAuto-injected trace contextCost
ConcernZero sidecar overheadWithout mesh~50-100 MB RAM per sidecar
The mesh tax (CPU/memory per sidecar) is justified above ~15 services or when compliance requires uniform mTLS.
When not to use a mesh
Skip the mesh when
- You have fewer than 10 services and a mature API gateway handles L7 routing.
- Latency-sensitive paths (sub-1 ms internal calls) cannot tolerate sidecar hop overhead.
- Team lacks operational maturity to debug proxy-level issues.
- Edge-only security (API gateway mTLS to clients) is sufficient for your threat model.
Quick recall
Everything you need if you only revisit this box.
- Service mesh = sidecar proxies (data plane) + control plane (Istiod) for uniform networking policy.
- mTLS encrypts all inter-service traffic without application changes.
- VirtualService controls routing weights, retries, timeouts — enables canary deploys.
- Outlier detection ejects unhealthy endpoints to prevent cascading failures.
- Mesh overhead (~50-100 MB RAM per pod) is justified at 15+ services or compliance-driven mTLS.
Test yourself
Answer these before moving on — recall is what makes it stick.