PrepZone Logo
PrepZone

DLT, Retry Topics, and Poison Messages

Route poison VaultCommerce payloads to a dead-letter topic after bounded retries — never block the partition.

Why this matters

  • VaultCommerce routes failed inventory messages to vaultcommerce.orders.placed.v1.DLT after 3 retries with exponential backoff.
  • Poison pill (malformed Avro) would infinite-loop without DLT escape hatch.
  • DLT consumers alert on-call — human triage before replay or discard.
  • Spring Kafka DefaultErrorHandler implements this pattern natively.
At-most-once
Fire and forgetNo retry
Risk: loss
At-least-once
Retry on failureKafka default
Risk: duplicates
Exactly-once
Transactions + idempotence
Cost: complexity
At-most-once may lose messages. At-least-once may duplicate. Exactly-once needs broker and application cooperation.

Retry topic pattern

Main consumer fails → publish to topic-retry-30000 with header retry-count. Retry consumer waits 30s (or uses delay topic + separate consumer) then republishes to main topic. VaultCommerce uses Spring's FixedBackOff + DeadLetterPublishingRecoverer.

Key points

  • DLT (dead-letter topic) — quarantine for permanently failed messages
  • Retry topic — intermediate topic with delay before reattempt
  • Poison message — record that always fails processing
  • Backoff — exponential delay between retries
  • Non-blocking error handling — skip and continue partition processing

DLT operations

DLT retention 30 days. Dashboard shows DLT rate. Fix code, then kafka-console-consumer inspect payload, replay via admin tool with fixed schema. VaultCommerce never auto-replays DLT without human ack.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. DLT prevents poison messages blocking partitions. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
// VaultCommerce Spring Kafka — retry + DLT
@Bean
DefaultErrorHandler errorHandler(KafkaTemplate<String, Object> template) {
    var recoverer = new DeadLetterPublishingRecoverer(template,
        (record, ex) -> new TopicPartition(record.topic() + ".DLT", record.partition()));
    var handler = new DefaultErrorHandler(recoverer, new FixedBackOff(2000L, 3L));
    handler.addNotRetryableExceptions(AvroTypeException.class);
    return handler;
}

Quick recall

Everything you need if you only revisit this box.

  1. DLT prevents poison messages blocking partitions.
  2. Bounded retries with backoff before DLT.
  3. Human triage before DLT replay.

Test yourself

Answer these before moving on — recall is what makes it stick.