PrepZone Logo
PrepZone

Ordering, Causality, and Partition Strategy

Design partition keys so causally related VaultCommerce events always land on the same partition.

Why this matters

  • PaymentCaptured must not be processed before OrderPlaced — same orderId key ensures order.
  • Cross-aggregate causality (parent order → child returns) needs careful key design.
  • Headers can carry correlationId for tracing but not for ordering guarantees.
  • Expert topic tying partitioning, sagas, and stream joins together.
At-most-once
Fire and forgetNo retry
Risk: loss
At-least-once
Retry on failureKafka default
Risk: duplicates
Exactly-once
Transactions + idempotence
Cost: complexity
At-most-once may lose messages. At-least-once may duplicate. Exactly-once needs broker and application cooperation.

Partition key as causality boundary

All events for order-42 use key order-42 on all topics in the saga. VaultCommerce enforces via shared EventPublisher wrapper that rejects publishes without key. Different orders interleave freely across partitions.

Key points

  • Causality — event A must be observed before dependent event B
  • Correlation ID — tracing across services; does not enforce order
  • Partition key scope — ordering guarantee boundary
  • Cross-topic ordering — not guaranteed; design per-saga keys
  • Event versioning — ignore stale events with lower version numbers

Cross-topic ordering limits

Kafka does not order across topics — only within one partition of one topic. VaultCommerce puts saga steps on separate topics but same key per order. Consumers processing multiple topics must not assume global order — use event versioning or state machine.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Same aggregate key → same partition → causal order. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
// VaultCommerce EventPublisher — enforces partition key = aggregateId
public void publishDomainEvent(String topic, String aggregateId, Object event) {
    if (aggregateId == null || aggregateId.isBlank())
        throw new IllegalArgumentException("aggregateId required for ordering");
    ProducerRecord<String, Object> record = new ProducerRecord<>(topic, aggregateId, event);
    record.headers().add("causation-id", currentEventId.getBytes(UTF_8));
    kafkaTemplate.send(record);
}

Quick recall

Everything you need if you only revisit this box.

  1. Same aggregate key → same partition → causal order.
  2. No cross-topic global ordering.
  3. Enforce keys at publisher boundary.

Test yourself

Answer these before moving on — recall is what makes it stick.