PrepZone Logo
PrepZone

DNS, Service Discovery, and Health-Aware Routing

How ShardPay instances find each other and how health checks prevent routing to dead nodes.

Why this matters

  • Hardcoded IPs break the moment autoscaling replaces a pod.
  • DNS TTL vs real-time discovery is a classic trade-off.
  • ShardPay uses discovery with health checks, not static lists.
  • Stale endpoints cause silent traffic loss during deploys and failures.

DNS lookup path (Route 53)

querydelegationaliasCLIENT
Browser / appapi.streamhub.com
NETWORK
Route 53 Resolv…recursive lookup
NETWORK
Hosted ZoneA/AAAA alias
NETWORK
ALB / CloudFronttarget IP
Every request starts here — api.streamhub.com resolves to ALB or CloudFront.

Finding healthy instances

In a dynamic fleet, the question "where is the ledger service?" has a different answer every few minutes. Pods start, stop, crash, and migrate. Hardcoded IP lists break on the first deploy. DNS provides name resolution but caches answers for seconds to minutes. Service discovery bridges the gap: a live registry that tracks which instances are healthy right now.

ShardPay ledger pods register on startup with Consul and deregister on graceful shutdown. The API tier resolves ledger.shardpay.internal to current healthy endpoints via a discovery client that refreshes every 5 seconds — faster than DNS TTL, slower than per-request lookup.

Discovery mechanisms

  • DNS — resolves a hostname to one or more IP addresses, cached by clients and resolvers with a TTL (time to live). ShardPay's Kubernetes internal DNS has a 30-second TTL; during a rolling deploy, up to 30 seconds of traffic may hit terminated pods.
  • Service discovery — dynamic registry (Consul, etcd, Kubernetes Endpoints) updated in real time as instances register and deregister. ShardPay's Consul checks update within 2 seconds of a pod failing health checks.
  • Health check — periodic probe (HTTP /health, TCP connect, gRPC health protocol) that determines whether an instance is ready for traffic. ShardPay's ledger readiness probe checks Raft leader status and replica lag, not just process alive.
  • Client-side load balancing — the client fetches the instance list from discovery and picks one (round-robin, least-connections, consistent hash). ShardPay's gRPC client-side LB routes by accountId hash to the correct ledger shard.

DNS TTL vs real-time discovery

DNS is simple and universal but stale. A TTL of 30 seconds means clients may send traffic to dead pods for up to 30 seconds after termination. For stateless API pods, retries mask this. For stateful ledger leaders, 30 seconds of writes to a dead leader causes data loss risk.

ShardPay uses a layered approach: Kubernetes DNS for coarse resolution (which service exists), Consul for fine-grained health-aware routing (which instances are ready), and client-side load balancing for shard-aware routing (which shard owns this account).

Walkthrough: deploy with stale DNS

ShardPay deploys a new ledger version. Kubernetes terminates old pods and starts new ones over 60 seconds. With DNS TTL = 30s, API clients cache old pod IPs for up to 30 seconds. During this window, ~4% of transfer requests fail with connection refused.

Fix: readiness-gated registration. New pods register with Consul only after passing health checks. Old pods deregister before termination (preStop hook). Consul removes unhealthy endpoints within 2 seconds. API clients refresh every 5 seconds. Failure rate during deploy drops to under 0.1%.

Java
// ShardPay — discovery-backed gRPC client with health-aware refresh
public class LedgerDiscoveryClient {
    private final ConsulClient consul;
    private volatile List<Endpoint> healthyEndpoints = List.of();
    private final ScheduledExecutorService refresher = Executors.newSingleThreadScheduledExecutor();

    public LedgerDiscoveryClient(ConsulClient consul) {
        this.consul = consul;
        refresher.scheduleAtFixedRate(this::refresh, 0, 5, TimeUnit.SECONDS);
    }

    private void refresh() {
        healthyEndpoints = consul.getHealthyInstances("ledger", "v1");
    }

    public LedgerShardClient route(AccountId accountId) {
        List<Endpoint> endpoints = healthyEndpoints;
        if (endpoints.isEmpty()) throw new UnavailableException("no healthy ledger instances");

        Endpoint target = consistentHash.pick(endpoints, accountId);
        return grpcClient.forEndpoint(target);
    }
}

Health checks that matter

A health check that only verifies "process is running" is insufficient. ShardPay's ledger health check validates: (1) Raft node is part of a quorum, (2) replica lag is under 2 seconds, (3) disk write latency is under 10ms. A pod that is alive but cannot serve correct data is removed from the pool before it causes errors.

Liveness vs readiness distinction: liveness means the process should be restarted (it is stuck). Readiness means the instance should receive traffic (it is caught up and healthy). ShardPay restarts on liveness failure; removes from load balancer on readiness failure without restart.

Health check design

  • Liveness probe — "should this pod be killed and restarted?" ShardPay's liveness checks for deadlock (thread pool exhausted for 60s) or OOM imminent.
  • Readiness probe — "should this pod receive traffic?" ShardPay's readiness checks Raft state, replica lag, and disk health. Failing readiness removes the pod from discovery without killing it.
  • Graceful shutdown — deregister from discovery, drain in-flight requests, then terminate. ShardPay's preStop hook waits 15 seconds for active transfers to complete before pod deletion.
  • Cascading failure prevention — if health checks are too aggressive, a brief slowdown marks all pods unhealthy, leaving zero endpoints. ShardPay uses hysteresis: two consecutive failures before removal, one success before re-addition.

Service mesh and discovery at scale

At ShardPay's scale (40 API pods, 18 ledger pods, 6 shards × 3 replicas), manual endpoint management is impossible. Kubernetes Endpoints controller watches pod labels and updates DNS automatically. Consul adds health-aware filtering on top. Istio service mesh (adopted in module 11) adds mTLS, retry policies, and circuit breaking at the discovery layer.

The principle remains: never hardcode endpoints. Every service name resolves through discovery. Every resolution is health-aware. Every deploy is graceful.

Quick recall

Everything you need if you only revisit this box.

  1. Discovery beats hardcoded endpoints; health checks must gate routing.
  2. DNS TTL trades freshness for resolver load — too high causes stale traffic during deploys.
  3. Readiness ≠ liveness: remove unhealthy pods from routing without restarting them.

Test yourself

Answer these before moving on — recall is what makes it stick.