Why this matters
- Heisenbugs in distributed systems reproduce only under partition, load, or specific shard routing — printf debugging across 40 services fails.
- Structured logs with trace-id are mandatory; ShardPay rejects PRs that add unstructured log lines to the transfer path.
- On-call runbooks start with "find the trace" — everything else follows from that anchor.
- Senior interviews expect a systematic workflow, not "I'd check the logs."
The debugging workflow
ShardPay's on-call runbook for transfer failures follows five steps: anchor on business ID, pull trace, correlate logs, check metrics at incident time, reproduce or mitigate.
Step 1: Anchor on a single request
Support tickets include transferId. If missing, search by merchant ID + timestamp in the API access log. Everything downstream keys off this one request — resist the urge to grep globally first.
Step 2: Find the trace
Jaeger search by baggage transferId=TXN-c8e2. If no trace exists, check sampling (was it a fast success trace discarded?) or broken propagation (orphan subtree in a different trace ID).
Step 3: Correlate logs
Click trace ID in Grafana → log panel filtered by trace_id. ShardPay's JSON logs include trace_id, span_id, service, shard_id, and level. Follow the error log line backward to the first WARN.
Step 4: Check metrics at timestamp
Overlay trace timestamp on golden signals dashboard. Was shard 7 saturated? Did deploy happen 5 minutes before? Metrics explain why the span was slow; the trace shows where.
Step 5: Mitigate, then reproduce
Mitigate first (scale pool, rollback deploy, route around bad shard). Reproduce in staging with traffic shadowing if root cause is unclear.
Key points
- Correlation ID — business identifier (transferId) plus technical identifier (trace_id) linking all observability signals. ShardPay sets both at the gateway; internal services must not generate new trace IDs mid-request.
- Structured logging — JSON with fixed fields, not string concatenation.
{"level":"ERROR","trace_id":"abc","shard_id":7,"msg":"debit timeout"}is searchable;"Error on shard 7"is not. - Three pillars correlation — traces show path and latency; logs show decision context; metrics show system-wide state at that moment. ShardPay's Grafana dashboards join all three by trace_id.
- Time alignment — ensure clock sync (NTP) across services. A 30-second skew between trace timestamps and metric timestamps sends on-call chasing the wrong minute.
Walkthrough: missing transfer report
Merchant M-3301 reports a $2,400 transfer " disappeared" — debited from their account but not credited to the recipient 20 minutes ago.
- Trace lookup:
transferId=TXN-f1a9→ trace shows saga reachedDEBITEDon shard 2, credit call to shard 6 returned HTTP 504 after 30s. - Logs: orchestrator log at 18:14:02 —
CreditException: gateway timeout shard-6; saga state persisted asCOMPENSATING. - Metrics: at 18:14, shard 6 pool utilization was 99% — saturation caused timeouts, not logic bug.
- Resolution: compensation completed at 18:14:35 — debit reversed. Merchant saw funds return but no explanation.
- Root cause: shard 6 pool undersized for end-of-day batch overlap with transfers. Fix: separate pool for batch vs API, increase max connections.
// ShardPay structured logging — MDC populated from OTel context
@Slf4j
public class TransferOrchestrator {
public void creditShard(SagaInstance saga) {
MDC.put("transfer_id", saga.transferId());
MDC.put("saga_id", saga.sagaId());
MDC.put("shard_id", String.valueOf(saga.toShard()));
try {
log.info("credit attempt started amount_cents={}", saga.amountCents());
shardClient.credit(saga.toShard(), saga.request());
log.info("credit succeeded");
} catch (TimeoutException e) {
log.error("credit timeout after_ms={}", 30_000, e);
compensateDebit(saga);
} finally {
MDC.clear();
}
}
}
Advanced techniques
When the standard workflow fails — intermittent bugs, race conditions, partition-specific failures — ShardPay escalates to traffic shadowing and chaos reproduction.
Traffic shadowing
Production traffic is mirrored to staging with synthetic account IDs. Staging runs the same code path without affecting real balances. ShardPay caught a race condition in saga recovery that only appeared at >500 transfers/sec — invisible in unit tests, visible in shadow traffic.
Partition and failure injection
After mitigation, reproduce in a controlled environment. ShardPay's staging cluster runs Toxiproxy between orchestrator and shards to simulate 504 timeouts. Confirm compensation fires correctly before declaring fix verified.
Postmortem discipline
Every Sev-2+ incident produces a blameless postmortem with timeline, root cause, action items, and runbook updates. ShardPay's missing-transfer report above updated the runbook: check saga state in orchestrator DB before assuming data loss.
Key points
- Traffic shadowing — replay prod requests in staging safely. ShardPay scrubs PII and replaces account IDs with test fixtures; timing and payload shapes match prod.
- Exemplars — link metric data points to example traces. ShardPay's p99 latency graph click-through jumps to the slowest trace in that bucket — saves manual Jaeger search.
- Runbook-driven on-call — decision trees, not tribal knowledge. "504 on credit → check shard saturation → check saga state → compensate if stuck" is written, not remembered.
- Blameless postmortem — focus on systems and process, not individuals. Action items tracked to completion; ShardPay reviews open postmortem actions in weekly reliability meeting.
Quick recall
Everything you need if you only revisit this box.
- Anchor on one transferId → trace → logs → metrics at that timestamp.
- Structured JSON logs with trace_id are non-negotiable for distributed debugging.
- Mitigate first, reproduce with shadow traffic second, postmortem always.
Test yourself
Answer these before moving on — recall is what makes it stick.