PrepZone Logo
PrepZone

When to Distribute (and When Not)

The operational tax of distribution, when a modulith wins, and how ShardPay knew it was time to split.

Why this matters

  • Staff engineers are judged on whether they distribute too early or too late.
  • A modular monolith often outperforms microservices for teams under 20 engineers.
  • ShardPay split ledger, fraud, and notifications only after clear scaling bottlenecks.
  • Premature distribution adds latency, failure modes, and on-call burden before you need them.
Phase 1
Single JVMAll modules in-process
Ledger
Fraud
Notify
Shard AAccounts 0–N
Shard B
Shard C
Split when proven limits appear — not when microservices sound fashionable.

The distribution tax

Every network hop adds latency, failure modes, and versioning problems. ShardPay's monolith handled 2,000 transfers/sec on one beefy machine with a single Postgres instance and one deploy pipeline. After splitting into three services, the team needed Kubernetes, distributed tracing, contract tests, and a three-team on-call rotation before matching that throughput — and p99 latency was higher for six months.

The distribution tax is real and recurring. It is not paid once at the split; it is paid on every feature that crosses a service boundary, every incident that spans three log systems, and every deploy that must coordinate schema versions across teams.

Costs of distribution

  • Operational tax — deploy pipelines per service, service mesh config, distributed debugging, and correlated tracing across hops. ShardPay's mean time to diagnose incidents doubled until OpenTelemetry was mandatory on every RPC.
  • Data tax — a single-database transaction becomes a saga with compensating actions, outbox tables, and eventual consistency windows. ShardPay's cross-shard transfer went from one BEGIN…COMMIT block to a choreographed saga with explicit failure states.
  • Team tax — Conway's law says system boundaries mirror communication patterns; bad boundaries create coordination overhead. ShardPay split along team ownership (ledger, fraud, notifications) only when those teams already worked mostly independently.
  • Modulith — clear module boundaries inside one deployable unit, with enforced package-level APIs and separate schemas. ShardPay kept reporting and admin inside the modulith for two years because they shared the same release cadence and data model.

When ShardPay stayed monolithic

For its first 18 months, ShardPay was a modular monolith: separate Java packages for ledger, API, and admin, one WAR file, one database. The team was eight engineers shipping weekly. Cross-cutting changes (add a field to transfers) touched one repo, one migration, one deploy. Latency was predictable because there were no network hops between ledger logic and fraud rules — just in-process calls.

The monolith was not "legacy." It was the correct architecture until proven constraints appeared. Premature microservices would have slowed feature delivery without solving a real bottleneck.

When ShardPay split

Three independent scaling profiles justified three services. Fraud scoring needed GPU nodes and a Python runtime with a different release cadence than the Java ledger. The ledger needed CP consistency and dedicated I/O tuning. Notifications were AP, bursty, and tolerable at seconds of delay. None of these profiles could share one machine sizing or one deploy schedule without waste or risk.

Walkthrough: the ledger split decision

ShardPay's ledger hit 32 cores and disk I/O saturation at 4,500 transfers/sec. Vertical scaling to a larger instance bought three months. Horizontal sharding required extracting the ledger into its own service with a stable RPC contract, so API pods could scale independently. The split took six weeks; the decision took six months of metrics proving the monolith could not shard in place.

Java
// Before split — in-process call inside the monolith
@Service
public class TransferService {
    private final LedgerRepository ledger;
    private final FraudRulesEngine fraud;  // same JVM, same transaction boundary

    public TransferResult transfer(TransferRequest req) {
        fraud.check(req);                          // ~5ms, no network
        return ledger.debitAndCredit(req);         // single DB transaction
    }
}

// After split — network hop, separate failure domain, saga for cross-shard
@Service
public class TransferOrchestrator {
    private final LedgerClient ledger;           // gRPC, 80ms budget
    private final FraudClient fraud;             // separate deploy, GPU pool

    public TransferResult transfer(TransferRequest req) {
        fraud.check(req);                          // RPC, may timeout independently
        return ledger.transfer(req);               // saga if accounts cross shards
    }
}

Signs you are ready to distribute

ShardPay used a checklist before each split: (1) a measured bottleneck that cannot be fixed in the monolith, (2) a team boundary that matches the service boundary, (3) a stable domain API that will not change weekly, and (4) operational readiness for independent deploys and monitoring. Missing any item meant staying in the modulith longer.

Conversely, signs you are not ready: "we want microservices," "other companies do it," or "the monolith feels messy." Messy code splits into messy services. Clean module boundaries in a modulith are cheaper to maintain than distributed spaghetti.

Split triggers at ShardPay

  • Independent scaling — fraud needed 4× GPU capacity during model retraining while ledger traffic was flat. Separate services let each tier scale on its own metrics.
  • Different availability targets — notifications at 99.9% vs ledger at 99.99% with durability guarantees. Coupling them in one deploy risked unnecessary ledger downtime for notification bugs.
  • Regulatory isolation — PCI scope reduction by moving card-token handling to a dedicated service with a smaller audit surface.
  • Technology mismatch — Python ML stack for fraud vs Java ledger; forcing one runtime would slow both teams.

Quick recall

Everything you need if you only revisit this box.

  1. Distribution trades simplicity for scale and autonomy — the tax is ongoing, not one-time.
  2. Split on proven bottlenecks and team boundaries, not fashion.
  3. A well-structured modulith is a valid and often optimal end state.

Test yourself

Answer these before moving on — recall is what makes it stick.