Why this matters
- Not every multi-step flow needs a central orchestrator; simple refund chains suit choreography well.
- ShardPay uses choreography for merchant refund flows (three services, clear event contracts) and orchestration for cross-shard transfers (five steps, strict ordering).
- Debugging choreography requires excellent tracing — the flow lives in event chains, not one state table.
- Interviewers ask you to compare both styles and pick based on complexity, team boundaries, and observability budget.
Event-driven saga without a coordinator
In choreography, there is no central saga DB. The debit service publishes TransferDebited; the credit service consumes it and publishes TransferCredited or CreditFailed; the debit service listens for CreditFailed and publishes DebitCompensated.
ShardPay's refund flow works this way: RefundRequested → billing validates → RefundApproved → ledger credits merchant → RefundCompleted → notification emails the merchant. Each handler is idempotent on refundId.
Key points
- Choreography — decentralized saga where services react to domain events without a central coordinator. ShardPay's refund pipeline has no orchestrator DB — state is implied by the latest event in the chain.
- Event chain — ordered causality via published events, not a central state machine. A failed credit emits
CreditFailed; downstream handlers subscribe to compensation triggers. - Event contract — schema, versioning, and idempotency rules per topic. ShardPay uses Avro schemas in Schema Registry with
refundIdas the idempotency key on every handler. - Visibility cost — reconstructing "where is this refund?" requires tracing across topics. ShardPay compensates with OpenTelemetry links from each event handler span to the originating
refundId.
When ShardPay chooses choreography
Choreography fits when the flow is short (≤4 steps), teams own services independently, and event contracts are stable. ShardPay's refund saga has three services with infrequent schema changes — a good fit.
Refund walkthrough
- Merchant API publishes
RefundRequested{refundId, merchantId, amount}toshardpay.refunds.v1. - Billing service consumes, validates merchant balance, publishes
RefundApproved{refundId}. - Ledger service consumes, credits merchant account idempotently, publishes
RefundCompleted{refundId}. - Notification service consumes, sends email — no compensation needed (at-most-once email is acceptable).
If ledger credit fails, it publishes RefundFailed{refundId, reason}. Billing listens and publishes RefundApprovalRevoked — no central orchestrator coordinates the rollback.
// ShardPay refund choreography — ledger handler
@KafkaListener(topics = "shardpay.refunds.approved.v1")
public void onRefundApproved(ConsumerRecord<String, RefundApproved> record) {
RefundApproved event = record.value();
if (processedEvents.exists(event.refundId())) return;
try {
ledgerService.creditMerchant(event.merchantId(), event.amountCents());
kafkaTemplate.send("shardpay.refunds.completed.v1",
event.refundId(), RefundCompleted.from(event));
processedEvents.mark(event.refundId());
} catch (InsufficientFundsException e) {
kafkaTemplate.send("shardpay.refunds.failed.v1",
event.refundId(), RefundFailed.from(event, e.getMessage()));
}
}
Choreography vs orchestration
ShardPay uses both deliberately. The decision matrix below is documented in their architecture decision records (ADRs).
| Criterion | Choreography | Orchestration |
|---|---|---|
| Step count | ≤4, linear | 5+, branching |
| Failure visibility | Needs strong tracing | Saga DB query |
| Team ownership | One team per service | Platform owns orchestrator |
| Schema churn | Low | High (central contract) |
| ShardPay example | Refunds | Cross-shard transfers |
Step count
Choreography≤4, linearOrchestration5+, branchingFailure visibility
ChoreographyNeeds strong tracingOrchestrationSaga DB queryTeam ownership
ChoreographyOne team per serviceOrchestrationPlatform owns orchestratorSchema churn
ChoreographyLowOrchestrationHigh (central contract)ShardPay example
ChoreographyRefundsOrchestrationCross-shard transfers
When ShardPay chooses choreography vs orchestration
Pitfalls in choreography
Missing compensation listener — if debit service does not subscribe to CreditFailed, a failed credit leaves a orphaned debit. ShardPay's CI checks that every forward event type has a documented compensation path in the event catalog.
Cyclic dependencies — service A waits for B's event while B waits for A's. ShardPay's refund chain is strictly acyclic: Requested → Approved → Completed.
Key points
- Decentralized control — no single point of failure for saga logic, but also no single source of truth. ShardPay runs a nightly job that scans for refunds with
Approvedbut noCompletedafter 1 hour. - Eventual visibility — lag between steps equals consumer processing time. ShardPay's refund p99 completes in 4 seconds; support tools join events by
refundIdin the data warehouse. - Schema evolution — adding a step means new topic + new consumer; orchestration adds a case in one switch statement. ShardPay migrated one refund validation step from choreography to orchestration when fraud rules became too complex.
- Testing — contract tests per event pair; ShardPay uses Pact to verify
RefundApprovedproducers match ledger consumer expectations.
Kafka trackSee saga choreography in Kafka track
Quick recall
Everything you need if you only revisit this box.
- Choreography distributes saga logic via event chains — no central coordinator.
- Best for short flows with stable contracts and strong tracing.
- ShardPay uses choreography for refunds, orchestration for cross-shard transfers.
Test yourself
Answer these before moving on — recall is what makes it stick.