PrepZone Logo
PrepZone

Ordering Guarantees Across Services

Per-partition ordering, total order costs, and when ShardPay accepts out-of-order settlement.

Why this matters

  • Global ordering requires a single Kafka partition — throughput ceiling unusable for ShardPay's millions of events per minute.
  • Key-based partitioning gives order where the business needs it: all events for accountId X on one partition.
  • Cross-partition operations (A pays B across two account shards) still need saga or application coordination — partition order does not solve everything.
  • Interviewers ask how you guarantee debit-before-credit visibility for one account while scaling horizontally.
No orderIndependent events
Per-partitionSame key ordered
Global totalSingle sequencer
Global order is expensive; partition-key order is the Kafka sweet spot.

Ordering scopes

Total order: one sequence for all events cluster-wide — simple mental model, single bottleneck. Partition order: Kafka preserves order within one partition; no guarantee across partitions. Per-key ordering: hash accountId (or transferId) to partition so all events for that key share one ordered stream.

ShardPay publishes TransferCompleted with key = merchantAccountId so settlement, notification, and analytics consumers see consistent per-merchant ordering without a global sequencer.

Walkthrough: same account, two transfers

Two transfers debit merchant account M-42 in quick succession. Both events hash to partition 17. Consumers process T1 before T2 matching commit order on the ledger shard. Merchant running balance projection never shows T2 before T1 for the same account.

Key points

  • Total order — every consumer sees identical global sequence; impractical at ShardPay scale except for low-volume control topics (Raft metadata changelog with single partition by design).
  • Partition order — broker guarantees order of records in one partition; reordering never happens inside partition. ShardPay relies on this for per-account settlement streams keyed by merchantAccountId.
  • Per-key ordering — choose partition key to match consistency boundary. ShardPay keys by creditor account for settlement; keys by transferId for saga orchestrator steps so all lifecycle events for one transfer co-locate.
  • Cross-partition coordination — operations touching two keys (payer on partition 3, payee on partition 9) have independent order; saga state machine reconciles with idempotent steps and explicit state transitions.
  • Reorder buffer — consumer-side window sorting when upstream cannot guarantee key discipline (legacy producers). ShardPay maintains 500ms reorder buffer on analytics only; financial path rejects out-of-order keys at schema validation.

Cross-partition transfers

Account A on shard 1 pays account B on shard 2: ledger emits two events with different partition keys. Nothing in Kafka guarantees global debit-before-credit order across partitions. ShardPay's saga orchestrator consumes a transferId-keyed topic that sequences orchestration commands — RESERVE → COMMIT → NOTIFY — while per-account projections consume account-keyed topics independently.

Dispute investigations use transferId correlation, not global offset comparison.

Walkthrough: out-of-order across merchants

Transfer batch job publishes 10k events in random merchant order across 64 partitions. Merchant M-42's events stay ordered on partition 17; merchant M-99 on partition 3. A dashboard aggregating all merchants may interleave arbitrarily — correct for global KPIs. M-42's statement page consumes only partition 17 via keyed stream — correct for merchant-facing order.

Java
// ShardPay event publisher — partition key matches ordering scope (Java 17)
public void publishTransferCompleted(TransferCompleted event) {
    ProducerRecord<String, TransferCompleted> record = new ProducerRecord<>(
        "shardpay.transfer.completed.v1",
        event.merchantAccountId(),  // key → per-merchant order
        event
    );
    record.headers().add("eventId", event.eventId().getBytes(StandardCharsets.UTF_8));
    record.headers().add("transferId", event.transferId().getBytes(StandardCharsets.UTF_8));
    kafkaTemplate.send(record);
}

// Consumer assumes order only within same merchantAccountId partition
public void onTransferCompleted(TransferCompleted event) {
    balanceProjector.applyInOrder(event.merchantAccountId(), event);
}

When to widen or narrow the key

Wrong partition key causes false concurrency or hot partitions. ShardPay almost keyed by merchantId until hot merchants saturated single partitions — moved to merchantAccountId sub-accounts for top 50 merchants. Ordering scope narrowed to the entity that needs serial apply.

Changing partition key is a breaking migration — dual-write or new topic version required.

Quick recall

Everything you need if you only revisit this box.

  1. Order is per partition — choose partition key to match your consistency boundary.
  2. Cross-partition operations need saga or orchestrator coordination, not broker global order.
  3. Hot keys require sub-key splitting; changing keys is a migration, not a config tweak.

Test yourself

Answer these before moving on — recall is what makes it stick.