PrepZone Logo
PrepZone

Distributed Tracing — Context Propagation

Trace IDs across ShardPay services — follow one transfer through gateway, ledger, and fraud.

Why this matters

  • Logs alone cannot reconstruct cross-service latency or determine which hop added 150ms to a transfer.
  • OpenTelemetry (OTel) is the industry standard; ShardPay exports traces to Jaeger with 7-day retention.
  • Every on-call incident at ShardPay starts with "find the trace" — not grep across 40 microservice log streams.
  • Interviewers expect you to explain context propagation, sampling trade-offs, and how tracing complements metrics and logs.
Gateway spantrace-id: abc
Ledger spanparent: gateway
Fraud spanparent: ledger
traceparent header links spans across service boundaries.

Context propagation

When a merchant hits POST /transfers, the API gateway creates a root span and injects W3C traceparent into outbound headers. Every downstream service — fraud, ledger shards, notification — extracts the context, creates a child span, and passes context to its own outbound calls.

ShardPay's transfer TXN-a4f2 shows this chain: gateway (12ms) → fraud scoring (94ms) → orchestrator (8ms) → shard-3 debit (41ms) → shard-9 credit (38ms) → outbox publish (6ms). Total 199ms — fraud is the bottleneck, visible immediately in Jaeger.

Key points

  • Trace — the end-to-end journey of one request, identified by a 128-bit trace ID. ShardPay support looks up traces by transferId via a baggage field set at the gateway.
  • Span — one unit of work in one service: name, start/end timestamps, status, attributes. The fraud span includes attributes merchantId, riskScore, and modelVersion.
  • traceparent — W3C header (00-{trace-id}-{span-id}-{flags}) propagated on HTTP and gRPC. ShardPay's gRPC interceptors auto-inject; missing propagation breaks the trace chain.
  • Baggage — optional key-value pairs propagated alongside trace context. ShardPay sets baggage: transferId=TXN-a4f2 at the gateway so every span is searchable by business ID.

Instrumentation in ShardPay

ShardPay uses OTel Java agent for automatic HTTP/gRPC instrumentation and manual spans for business-critical sections like saga state transitions and shard routing decisions.

Automatic vs manual spans

Automatic instrumentation covers ingress, egress, and JDBC calls — sufficient for 80% of latency analysis. Manual spans cover orchestrator state transitions (saga.advance), shard selection (route.toShard), and outbox publish confirmation — where automatic spans lack business context.

Java
// ShardPay manual span around saga step
public void debitShard(SagaInstance saga) {
    Span span = tracer.spanBuilder("saga.debit")
        .setAttribute("transfer.id", saga.transferId())
        .setAttribute("shard.id", saga.fromShard())
        .startSpan();
    try (Scope scope = span.makeCurrent()) {
        DebitResult result = shardClient.debit(saga.fromShard(), saga.request());
        span.setAttribute("debit.amount_cents", saga.amountCents());
        saga.recordDebit(result);
    } catch (Exception e) {
        span.recordException(e);
        span.setStatus(StatusCode.ERROR);
        throw e;
    } finally {
        span.end();
    }
}

Sampling and storage

Tracing every request at 1M transfers/day generates terabytes. ShardPay uses tail-based sampling in the OTel collector: always keep errors and slow traces (>500ms), sample 1% of successful fast traces.

Walkthrough: debugging a slow transfer

  1. Merchant reports transfer took 8 seconds; support retrieves transferId=TXN-b7c1.
  2. Jaeger search by baggage transferId=TXN-b7c1 returns one trace (slow traces are always sampled).
  3. Flame graph shows 7.2 seconds in shardClient.credit on shard 9 — not fraud, not gateway.
  4. Drill into shard 9 span attributes: db.connection_wait_ms=7100 — connection pool exhausted.
  5. Metrics confirm pool saturation on shard 9; scale connection pool and shard replicas.

Key points

  • Head-based sampling — decide at trace start whether to keep (e.g., 1% random). Simple but may discard the slow trace you need. ShardPay moved away from pure head-based sampling for this reason.
  • Tail-based sampling — decide after trace completes based on latency, errors, or attributes. ShardPay's collector keeps all errors, all p99+ latency, and 1% of the rest.
  • Span events and links — span events mark point-in-time occurrences (e.g., retry attempt 2); span links connect async work (Kafka consumer) to the producer span. ShardPay links outbox publish spans to settlement consumer spans.
  • Trace correlation in logs — every log line includes trace_id and span_id from MDC. ShardPay's structured JSON logs join to Jaeger in Grafana with one click.

Quick recall

Everything you need if you only revisit this box.

  1. Trace ID + traceparent link spans across every service hop.
  2. Tail-based sampling keeps errors and slow traces without storing everything.
  3. Baggage fields like transferId make traces searchable by business ID.

Test yourself

Answer these before moving on — recall is what makes it stick.