PrepZone Logo
PrepZone

acks, Retries, and Durability Tuning

acks=all, min.insync.replicas, and retry policy — the producer settings that prevent silent data loss.

Why this matters

  • VaultCommerce production uses acks=all with min.insync.replicas=2 — no silent loss on single broker failure.
  • acks=1 saved 2ms but lost orders during a follower lag incident — never again on financial topics.
  • Retries without idempotence create duplicates — pair retries with enable.idempotence=true.
  • Most common producer tuning question in senior interviews.
acks=0
No wait
May loseFire and forget
acks=1
Leader ack
Leader crash risk
acks=all
ISR ack
DurableProduction
acks=0 is fastest but may lose data. acks=all waits for all ISR replicas — production default.

acks levels

acks=0: fire-and-forget, no broker ack. acks=1: leader wrote to local log. acks=all: all ISR replicas acked. VaultCommerce uses all for order and payment topics; 1 only for debug telemetry.

Key points

  • acks=all — wait for ISR replication before success
  • min.insync.replicas — broker-side gate; writes fail if ISR too small
  • retries — producer resends on retriable exceptions
  • delivery.timeout.ms — upper bound for total send attempt duration
  • NotLeaderOrFollowerException — transient; retries after metadata refresh

Retry and delivery timeout

retries=Integer.MAX_VALUE with delivery.timeout.ms=120000 retries transient errors (NOT_LEADER, timeout). request.timeout.ms bounds single request wait. Combined with idempotence, retries are safe from duplicate sequence numbers.

VaultCommerce rollout checklist

Before promoting changes that touch vaultcommerce.orders.placed.v1 and vaultcommerce.payments.captured.v1, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. acks=all + min.insync.replicas=2 for financial events. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
// VaultCommerce payment-producer — durability-first settings
props.put(ProducerConfig.ACKS_CONFIG, "all");
props.put(ProducerConfig.RETRIES_CONFIG, Integer.MAX_VALUE);
props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true); // implies acks=all, retries>0
props.put(ProducerConfig.MAX_IN_FLIGHT_REQUESTS_PER_CONNECTION, 5);
props.put(ProducerConfig.DELIVERY_TIMEOUT_MS_CONFIG, 120_000);
props.put(ProducerConfig.REQUEST_TIMEOUT_MS_CONFIG, 30_000);
// Cluster: default.replication.factor=3, min.insync.replicas=2

Quick recall

Everything you need if you only revisit this box.

  1. acks=all + min.insync.replicas=2 for financial events.
  2. Enable idempotence before max retries.
  3. delivery.timeout.ms caps total retry window.

Test yourself

Answer these before moving on — recall is what makes it stick.