PrepZone Logo
PrepZone

Error Handlers, Retry, and DLT in Spring

DefaultErrorHandler, exponential backoff, and DeadLetterPublishingRecoverer for VaultCommerce consumers.

Why this matters

  • VaultCommerce inventory listener retries 3× with 2s backoff then DLT.
  • Classifier exceptions: AvroTypeException → not retryable, straight to DLT.
  • SeekToCurrentErrorHandler legacy — use DefaultErrorHandler in Spring Kafka 3.x.
  • Observability: metric tags on retry count and DLT publish.
OrderServiceVaultCommerce
vaultcommerce.orders.placed.v1
InventoryConsumer
KafkaTemplate publishes to a topic. @KafkaListener consumes from a consumer group.

Error handler setup

DefaultErrorHandler + DeadLetterPublishingRecoverer + FixedBackOff. Custom ConsumerRecordRecoverer for logging PII-scrubbed payloads. VaultCommerce registers handler on ConcurrentKafkaListenerContainerFactory bean.

Key points

  • DefaultErrorHandler — Spring Kafka 3 unified error handling
  • DeadLetterPublishingRecoverer — publishes failed record to DLT
  • FixedBackOff / ExponentialBackOff — delay between retries
  • @RetryableTopic — declarative retry topic generation
  • notRetryableExceptions — skip retry for deterministic failures

RetryableTopic alternative

@RetryableTopic(attempts=4, backoff=@Backoff(delay=1000, multiplier=2)) creates retry topics automatically. Non-blocking — other partitions continue. VaultCommerce uses this for payment listeners; manual DLT for inventory.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. DefaultErrorHandler + DeadLetterPublishingRecoverer. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
@Configuration
public class KafkaErrorConfig {
    @Bean
    ConcurrentKafkaListenerContainerFactory<String, OrderPlaced> kafkaListenerContainerFactory(
            ConsumerFactory<String, OrderPlaced> consumerFactory,
            KafkaTemplate<String, Object> template) {
        var factory = new ConcurrentKafkaListenerContainerFactory<String, OrderPlaced>();
        factory.setConsumerFactory(consumerFactory);
        factory.setCommonErrorHandler(new DefaultErrorHandler(
            new DeadLetterPublishingRecoverer(template,
                (r, e) -> new TopicPartition(r.topic() + ".DLT", r.partition())),
            new FixedBackOff(2000L, 3)));
        return factory;
    }
}

Quick recall

Everything you need if you only revisit this box.

  1. DefaultErrorHandler + DeadLetterPublishingRecoverer.
  2. Classify non-retryable exceptions.
  3. @RetryableTopic for declarative retry chains.

Test yourself

Answer these before moving on — recall is what makes it stick.