PrepZone Logo
PrepZone

Stateless vs Stateful Tiers

Which ShardPay services can scale freely and which need sticky routing, replication, or partitioning.

Why this matters

  • Misclassifying stateful services causes data loss on scale-in.
  • Session stickiness fights even load distribution.
  • ShardPay keeps transfer state in the ledger, not API pods.
  • The stateless/stateful boundary determines how every tier scales.
Stateless API
Any instanceNo local session state
Scale out freely
Stateful shard
Owns partitionLedger rows A–M
Sticky or hashed routing
Stateless instances scale horizontally; stateful nodes need sticky routing or partitioning.

Where state lives

The fundamental scaling question is: where does the data live? If every API pod holds account balances in local memory, scaling out means copying or partitioning that data. If API pods hold no account data and delegate to a ledger service, any pod can handle any request — true statelessness.

ShardPay's architecture enforces a strict boundary: API pods are stateless (no account data, no session state beyond a JWT). Ledger shards are stateful (authoritative balances, write-ahead logs, Raft state). Fraud pods are stateless (model in memory, loaded from S3 on startup). Notification workers are stateless (consume from Kafka, no local durable state).

State classification

  • Stateless — any instance can serve any request; no data pinned to a specific node. ShardPay's API pods validate a JWT, route to the correct ledger shard, and return the response — no pod-specific state survives the request.
  • Stateful — data or processing is pinned to specific nodes or shards. ShardPay's ledger shard 2 owns accounts hashing to shard 2; moving an account requires explicit rebalancing, not just adding a pod.
  • Sticky sessions — route the same client to the same instance to preserve in-memory session state. ShardPay avoids stickiness; it caused hot instances when a few merchants generated 40% of traffic on one pod.
  • Externalized state — session or cache data stored in Redis, a database, or object store so compute pods remain stateless. ShardPay stores rate-limit counters in Redis, not in API pod memory, so pods scale in and out freely.

Designing for stateless compute

Stateless tiers are the easy scaling path. ShardPay's API deployment scales from 10 to 40 pods during Black Friday with no data migration, no rebalancing, and no split-brain risk. Kubernetes adds pods; the load balancer distributes traffic; done.

The discipline is pushing state down to dedicated tiers. When ShardPay engineers want to cache a balance in the API layer for speed, the architecture review asks: "What happens when this pod dies mid-request? What happens on scale-in?" If the answer is "stale data" or "lost update," the cache belongs in Redis with TTL, not in pod memory.

Walkthrough: scale-in data loss

ShardPay's early prototype stored in-flight transfer status in a ConcurrentHashMap on the API pod. During a deploy, Kubernetes terminated a pod with 200 in-flight transfers. Those status lookups returned 404 until the client retried. The fix: transfer status lives in the ledger (durable) or Redis (externalized), never in pod memory. Scale-in became safe.

Java
// ShardPay API — stateless handler; all durable state in ledger
@RestController
public class TransferController {
    private final LedgerClient ledger;       // state lives here, not in this pod
    private final ShardRouter router;        // routing logic, no account data

    @PostMapping("/v1/transfers")
    public TransferResponse transfer(@RequestBody TransferRequest req,
                                     @RequestHeader("Idempotency-Key") String key) {
        AccountId from = req.fromAccount();
        LedgerShard shard = router.route(from);  // deterministic routing, no local state
        return ledger.transfer(shard, req.withIdempotencyKey(key));
    }
}

// Anti-pattern ShardPay removed — stateful API pod
// private final Map<String, TransferStatus> inFlight = new ConcurrentHashMap<>();
// ^ lost on pod termination, breaks horizontal scaling

Stateful tier patterns

Stateful tiers need one of three patterns: replication (copies of the same data), partitioning (each node owns a slice), or single-writer (one leader, many followers). ShardPay's ledger uses all three: each shard is a partition, each partition has three Raft replicas, and only the leader accepts writes.

Scaling stateful tiers is harder. Adding a ledger shard requires rebalancing account data — moving keys between nodes while serving traffic. ShardPay uses consistent hashing with virtual nodes so adding a shard moves only ~1/N of accounts, not a full reshuffle.

Stateful scaling patterns

  • Replication — multiple copies of the same data for read scaling and failover. ShardPay's ledger followers serve bounded-staleness reads, reducing load on the leader.
  • Partitioning (sharding) — each node owns a key range; writes go to one partition. ShardPay hashes accountId to determine shard ownership.
  • Single-writer (leader) — one node accepts writes; followers replicate. ShardPay's Raft leader per shard ensures linearizable writes without write conflicts.
  • Rebalancing — moving data between partitions when adding or removing nodes. ShardPay rebalances during low-traffic windows with dual-write to old and new shards during migration.

When stickiness is a last resort

Sticky sessions are a code smell in distributed systems. They fight the load balancer's goal of even distribution and break when the sticky instance dies. ShardPay uses stickiness only for WebSocket connections (merchant dashboard live feed), where reconnecting to a different pod would drop the stream. Even then, the pod stores no business state — only the socket handle.

For HTTP request handling, externalize everything. JWTs carry auth state. Idempotency keys live in the ledger. Rate limits live in Redis. The API pod is a pure function: request in, response out, no memory of previous requests.

Quick recall

Everything you need if you only revisit this box.

  1. Push state to dedicated tiers; keep compute pods stateless.
  2. Stateful tiers need replication, partitioning, or single-writer — not just more pods.
  3. Sticky sessions are a last resort; externalize session data to Redis or the database.

Test yourself

Answer these before moving on — recall is what makes it stick.