PrepZone Logo
PrepZone

Saga Pattern — Choreography

Services react to events without a central orchestrator — trade-offs for ShardPay refunds.

Why this matters

  • Not every multi-step flow needs a central orchestrator; simple refund chains suit choreography well.
  • ShardPay uses choreography for merchant refund flows (three services, clear event contracts) and orchestration for cross-shard transfers (five steps, strict ordering).
  • Debugging choreography requires excellent tracing — the flow lives in event chains, not one state table.
  • Interviewers ask you to compare both styles and pick based on complexity, team boundaries, and observability budget.
OrderPlaced
InventoryReserved
PaymentCaptured
PaymentFailed → ReleaseInventoryCompensate
Each service listens and publishes domain events. Failure triggers compensating events on the same log.

Event-driven saga without a coordinator

In choreography, there is no central saga DB. The debit service publishes TransferDebited; the credit service consumes it and publishes TransferCredited or CreditFailed; the debit service listens for CreditFailed and publishes DebitCompensated.

ShardPay's refund flow works this way: RefundRequested → billing validates → RefundApproved → ledger credits merchant → RefundCompleted → notification emails the merchant. Each handler is idempotent on refundId.

Key points

  • Choreography — decentralized saga where services react to domain events without a central coordinator. ShardPay's refund pipeline has no orchestrator DB — state is implied by the latest event in the chain.
  • Event chain — ordered causality via published events, not a central state machine. A failed credit emits CreditFailed; downstream handlers subscribe to compensation triggers.
  • Event contract — schema, versioning, and idempotency rules per topic. ShardPay uses Avro schemas in Schema Registry with refundId as the idempotency key on every handler.
  • Visibility cost — reconstructing "where is this refund?" requires tracing across topics. ShardPay compensates with OpenTelemetry links from each event handler span to the originating refundId.

When ShardPay chooses choreography

Choreography fits when the flow is short (≤4 steps), teams own services independently, and event contracts are stable. ShardPay's refund saga has three services with infrequent schema changes — a good fit.

Refund walkthrough

  1. Merchant API publishes RefundRequested{refundId, merchantId, amount} to shardpay.refunds.v1.
  2. Billing service consumes, validates merchant balance, publishes RefundApproved{refundId}.
  3. Ledger service consumes, credits merchant account idempotently, publishes RefundCompleted{refundId}.
  4. Notification service consumes, sends email — no compensation needed (at-most-once email is acceptable).

If ledger credit fails, it publishes RefundFailed{refundId, reason}. Billing listens and publishes RefundApprovalRevoked — no central orchestrator coordinates the rollback.

Java
// ShardPay refund choreography — ledger handler
@KafkaListener(topics = "shardpay.refunds.approved.v1")
public void onRefundApproved(ConsumerRecord<String, RefundApproved> record) {
    RefundApproved event = record.value();
    if (processedEvents.exists(event.refundId())) return;

    try {
        ledgerService.creditMerchant(event.merchantId(), event.amountCents());
        kafkaTemplate.send("shardpay.refunds.completed.v1",
            event.refundId(), RefundCompleted.from(event));
        processedEvents.mark(event.refundId());
    } catch (InsufficientFundsException e) {
        kafkaTemplate.send("shardpay.refunds.failed.v1",
            event.refundId(), RefundFailed.from(event, e.getMessage()));
    }
}

Choreography vs orchestration

ShardPay uses both deliberately. The decision matrix below is documented in their architecture decision records (ADRs).

CriterionChoreographyOrchestration
Step count≤4, linear5+, branching
Failure visibilityNeeds strong tracingSaga DB query
Team ownershipOne team per servicePlatform owns orchestrator
Schema churnLowHigh (central contract)
ShardPay exampleRefundsCross-shard transfers
  • Step count

    Choreography≤4, linear
    Orchestration5+, branching
  • Failure visibility

    ChoreographyNeeds strong tracing
    OrchestrationSaga DB query
  • Team ownership

    ChoreographyOne team per service
    OrchestrationPlatform owns orchestrator
  • Schema churn

    ChoreographyLow
    OrchestrationHigh (central contract)
  • ShardPay example

    ChoreographyRefunds
    OrchestrationCross-shard transfers

When ShardPay chooses choreography vs orchestration

Pitfalls in choreography

Missing compensation listener — if debit service does not subscribe to CreditFailed, a failed credit leaves a orphaned debit. ShardPay's CI checks that every forward event type has a documented compensation path in the event catalog.

Cyclic dependencies — service A waits for B's event while B waits for A's. ShardPay's refund chain is strictly acyclic: Requested → Approved → Completed.

Key points

  • Decentralized control — no single point of failure for saga logic, but also no single source of truth. ShardPay runs a nightly job that scans for refunds with Approved but no Completed after 1 hour.
  • Eventual visibility — lag between steps equals consumer processing time. ShardPay's refund p99 completes in 4 seconds; support tools join events by refundId in the data warehouse.
  • Schema evolution — adding a step means new topic + new consumer; orchestration adds a case in one switch statement. ShardPay migrated one refund validation step from choreography to orchestration when fraud rules became too complex.
  • Testing — contract tests per event pair; ShardPay uses Pact to verify RefundApproved producers match ledger consumer expectations.

Quick recall

Everything you need if you only revisit this box.

  1. Choreography distributes saga logic via event chains — no central coordinator.
  2. Best for short flows with stable contracts and strong tracing.
  3. ShardPay uses choreography for refunds, orchestration for cross-shard transfers.

Test yourself

Answer these before moving on — recall is what makes it stick.