PrepZone Logo
PrepZone

Ordering Guarantees and Idempotent Handlers

Why duplicate messages happen and how idempotent consumers keep VaultCommerce inventory correct.

Why this matters

  • Global order across all VaultCommerce events is impossible at scale; partition keys scope ordering to what matters.
  • Duplicate OrderPlaced events without idempotency keys caused double shipment requests in an early bug.
  • Idempotency keys belong in the consumer, not assumed away by the broker.
  • Foundation for every later producer and consumer tuning article.
At-most-once
Fire and forgetNo retry
Risk: loss
At-least-once
Retry on failureKafka default
Risk: duplicates
Exactly-once
Transactions + idempotence
Cost: complexity
At-most-once may lose messages. At-least-once may duplicate. Exactly-once needs broker and application cooperation.

Partition-local ordering

Kafka guarantees order within a single partition. VaultCommerce keys by orderId so OrderPlaced, PaymentCaptured, and OrderShipped for the same checkout arrive in publish order. Events for different orders on different partitions may interleave — that's fine.

Key points

  • Partition key — determines which partition receives the record; same key → same partition
  • Idempotency key — unique business identifier (eventId, orderId+type) stored before side effect
  • Natural idempotency — SQL UPDATE ... WHERE status = 'PENDING' makes retries no-ops
  • Out-of-order — possible across partitions; design aggregates to tolerate interleaving
  • Deduplication window — TTL on idempotency store; VaultCommerce keeps 30 days

Idempotent handler patterns

Store processed event IDs in Postgres with a unique constraint. Use natural keys: reserveInventory(orderId, sku, qty) upserts rather than blind inserts. VaultCommerce's payment consumer checks payment_events(event_id) before calling the gateway capture API.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Order is per-partition, not global. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
// VaultCommerce — idempotent reserve with natural key
public void reserveInventory(String orderId, List<LineItem> items) {
    for (LineItem item : items) {
        int updated = jdbcTemplate.update("""
            UPDATE inventory SET reserved = reserved + ?
            WHERE sku = ? AND available - reserved >= ?
              AND NOT EXISTS (
                SELECT 1 FROM reservations WHERE order_id = ? AND sku = ?
              )
            """, item.qty(), item.sku(), item.qty(), orderId, item.sku());
        if (updated == 0) throw new InsufficientStockException(item.sku());
        jdbcTemplate.update(
            "INSERT INTO reservations (order_id, sku, qty) VALUES (?, ?, ?)",
            orderId, item.sku(), item.qty());
    }
}

Quick recall

Everything you need if you only revisit this box.

  1. Order is per-partition, not global.
  2. Idempotent handlers make at-least-once safe.
  3. Key by the entity whose lifecycle must stay ordered (orderId).

Test yourself

Answer these before moving on — recall is what makes it stick.