PrepZone Logo
PrepZone

The 8 Fallacies of Distributed Computing

Why assuming reliable networks and zero latency leads to production outages — and how ShardPay designs around each fallacy.

Why this matters

  • Interviewers use fallacies to test whether you design for reality or the happy path.
  • Timeouts, retries, and idempotency exist because networks are not reliable.
  • ShardPay's first cross-region deployment failed on three fallacies at once.
  • Naming the fallacy behind a bug speeds root-cause analysis and prevents repeat incidents.
Network is reliablePackets drop
Latency is zeroRTT adds up
Bandwidth is infinitePipes saturate
Network is secureTLS everywhere
Topology is staticNodes churn
One administratorMany teams
Transport cost is zeroEgress bills
Network is homogeneousMixed hardware
Each fallacy is a design assumption that fails in production.

The fallacies in production

Peter Deutsch and James Gosling named eight assumptions that bite every distributed team. They are not academic — they are the reason ShardPay has timeouts on every outbound call, mTLS between services, and a FinOps line item for cross-region egress. Before shipping a feature, ShardPay engineers ask: which fallacy does this design accidentally assume is false?

Local development hides most of them. Your laptop runs all services on localhost with sub-millisecond RTT, one operator (you), and no packet loss. Production introduces the real world: autoscaling replaces pods, TLS certificates expire, and a merchant in Singapore talks to a ledger in Virginia.

Network assumptions

  • Network is reliable — packets drop, connections reset, and load balancers silently drain unhealthy backends. ShardPay wraps every ledger RPC in a retry policy with idempotency keys so a transient TCP reset does not create a duplicate $500 debit.
  • Latency is zero — round-trip time, queuing, and GC pauses add up; a "fast" service becomes slow under load. ShardPay allocates a 200ms end-to-end transfer budget and fails fast when fraud scoring exceeds its 50ms slice.
  • Bandwidth is infinite — large payloads saturate links and inflate bills; replication of full account snapshots overwhelmed ShardPay's cross-AZ links until they switched to change-data-capture events.
  • Network is secure — unencrypted internal traffic is a lateral movement path for attackers. ShardPay runs mTLS between every service hop, not just the public API edge, after a penetration test flagged plain-text gRPC.

Operations and topology fallacies

The second group of fallacies concerns who runs the system and how often it changes. Kubernetes deploys replace pods every few minutes. Three teams own ledger, fraud, and notifications with different on-call rotations. Hardware in the same "instance family" can have different disk throughput after a silent platform migration.

ShardPay's staging environment once passed all tests because it assumed static topology — three fixed ledger pods with stable IPs. Production autoscaling changed the pod set during a load test, DNS TTL cached stale endpoints for 30 seconds, and 4% of transfers failed until health-aware service discovery was wired in.

Operations assumptions

  • Topology is static — nodes join, leave, and migrate during deploys, autoscaling, and failures; routing must adapt in seconds, not hours. ShardPay ledger pods register with Consul on startup and deregister on graceful shutdown so the API never routes to a draining instance.
  • There is only one administrator — ownership is split across teams with different runbooks, SLOs, and release cadences. A ledger deploy must not require simultaneous fraud-team approval unless the contract between services changed.
  • Transport cost is zero — cross-region replication, log shipping, and backup sync have real dollar costs. ShardPay's EU expansion added a regional ledger cell partly to avoid shipping every read request across the Atlantic.
  • The network is homogeneous — ARM and x86, NVMe and network-attached storage, and different instance generations coexist in one fleet. ShardPay load tests on the slowest instance type in the pool so p99 does not surprise production.

Mapping fallacies to ShardPay defenses

Every resilience pattern in this track counters at least one fallacy. Retries and circuit breakers counter unreliable networks. Latency budgets counter zero-latency assumptions. CDNs and compression counter infinite bandwidth. Service mesh and mTLS counter insecure networks. Autoscaling and discovery counter static topology.

Example: Black Friday cross-fallacy stress

During ShardPay's first Black Friday, four fallacies collided. Traffic spiked (latency is not zero — queues formed). Autoscaling added pods (topology is not static — discovery lagged). Cross-shard replication bandwidth saturated (bandwidth is not infinite). A misconfigured security group dropped internal traffic (network is not reliable). The fix was not one knob — it was headroom, faster discovery, compressed replication, and verified mTLS paths.

Java
// ShardPay API client — every fallacy countered in one config block
LedgerClient client = LedgerClient.builder()
    .endpoint(discovery.resolve("ledger.shardpay.internal"))  // topology changes
    .connectTimeout(Duration.ofMillis(100))                   // latency is not zero
    .requestTimeout(Duration.ofMillis(80))                    // budget per hop
    .retryPolicy(RetryPolicy.exponentialBackoff(3, Duration.ofMillis(50))
        .withIdempotencyKey(request.idempotencyKey()))         // network is not reliable
    .tls(TlsConfig.mtls(clientCert, caCert))                    // network is not secure
    .build();

Building a fallacy checklist

ShardPay's architecture review template includes a fallacy column. New features must document which assumptions they violate and which defense applies. "We'll call the ledger synchronously" triggers latency and reliability review. "We'll cache DNS for an hour" triggers topology review. "We'll replicate everything to EU" triggers transport cost review.

This discipline prevents designing for the laptop and deploying to the WAN. It also gives interview answers structure: name the fallacy, name the production symptom, name the pattern that fixes it.

Quick recall

Everything you need if you only revisit this box.

  1. Eight assumptions that fail in real networks and operations — every outage maps to at least one.
  2. Retries, budgets, discovery, and mTLS exist to counter specific fallacies.
  3. ShardPay reviews new features against a fallacy checklist before production deploy.

Test yourself

Answer these before moving on — recall is what makes it stick.