StreamHub generates IDs for streams, clips, chat messages, and payments — billions per month. Auto-increment IDs are simple but leak business metrics and fail across shards. Snowflake IDs give global uniqueness, time ordering, and no central coordinator.
Requirements
What a good ID generator provides
- Uniqueness: No collisions across all services and regions.
- Sortability (optional): Time-ordered IDs improve DB index locality and pagination.
- Generation speed: 10K+ IDs/sec per worker without DB round trip.
- No single point of failure: Decentralized generation preferred.
- Compact: 64-bit integer fits in BIGINT; UUID is 128-bit (16 bytes).
Approach comparison
| Approach | Pros | Cons |
|---|---|---|
| Auto-increment (DB) | Simple, sequential, index-friendly | Single-writer bottleneck; leaks count; sharding breaks it |
| UUID v4 (random) | No coordination; 122 bits of randomness | Not sortable; index fragmentation; 16 bytes |
| UUID v7 (time-ordered) | Sortable; no coordination | Still 16 bytes; newer standard |
| Snowflake (64-bit) | Sortable, compact, decentralized | Clock skew breaks ordering; needs worker ID assignment |
| DB sequence per shard | Simple within shard | Global uniqueness needs shard prefix; coordination on shard assign |
Auto-increment (DB)
ProsSimple, sequential, index-friendlyConsSingle-writer bottleneck; leaks count; sharding breaks itUUID v4 (random)
ProsNo coordination; 122 bits of randomnessConsNot sortable; index fragmentation; 16 bytesUUID v7 (time-ordered)
ProsSortable; no coordinationConsStill 16 bytes; newer standardSnowflake (64-bit)
ProsSortable, compact, decentralizedConsClock skew breaks ordering; needs worker ID assignmentDB sequence per shard
ProsSimple within shardConsGlobal uniqueness needs shard prefix; coordination on shard assign
64-bit Snowflake ID
41 bitstimestamp (ms)
5 bitsdatacenter ID
5 bitsworker ID
12 bitssequence
Snowflake layout
64-bit Snowflake ID structure
- 41 bits timestamp: Milliseconds since custom epoch — ~69 years of IDs.
- 5 bits datacenter ID: Up to 32 datacenters (US-East, EU-West, APAC).
- 5 bits worker ID: Up to 32 workers per datacenter.
- 12 bits sequence: Up to 4096 IDs/ms per worker — 4M IDs/sec theoretical max.
EPOCH = 1_700_000_000_000 # custom epoch ms
def generate_id(datacenter_id: int, worker_id: int, seq: int, now_ms: int) -> int:
return ((now_ms - EPOCH) << 22) | (datacenter_id << 17) | (worker_id << 12) | seq
Worker ID assignment
# Worker ID via K8s StatefulSet ordinal (pseudo)
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: id-generator
spec:
replicas: 32
# Pod ordinal 0–31 maps directly to worker_id
StreamHub ID strategy
| Aspect | Snowflake (internal) | UUID v7 (external API) |
|---|---|---|
| Used for | Internal DB primary keys, Kafka message IDs | Public API resource IDs (stream_id, clip_id) |
| Reason | Compact, sortable, fast index inserts | Opaque to clients; no business metric leakage |
| Format | 64-bit integer stored as BIGINT | String like clip_01HXYZ in API layer |
Used for
Snowflake (internal)Internal DB primary keys, Kafka message IDsUUID v7 (external API)Public API resource IDs (stream_id, clip_id)Reason
Snowflake (internal)Compact, sortable, fast index insertsUUID v7 (external API)Opaque to clients; no business metric leakageFormat
Snowflake (internal)64-bit integer stored as BIGINTUUID v7 (external API)String like clip_01HXYZ in API layer
CREATE TABLE clips (
id BIGINT PRIMARY KEY, -- Snowflake internally
public_id VARCHAR(32) UNIQUE, -- opaque external ID
stream_id BIGINT NOT NULL,
created_at TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX idx_clips_stream ON clips (stream_id, id DESC);
Quick recall
Everything you need if you only revisit this box.
- Auto-increment fails at scale — single writer, leaks metrics, breaks sharding.
- Snowflake: 64-bit, time-sortable, 4096 IDs/ms per worker, needs unique worker IDs.
- UUID v4: no coordination but not sortable; UUID v7 adds time ordering.
- Assign worker IDs via ZooKeeper, etcd, or K8s StatefulSet ordinals.
- Guard against clock skew — NTP sync and monotonic timestamp enforcement.
- Expose opaque public IDs; use Snowflake internally for index performance.
Test yourself
Answer these before moving on — recall is what makes it stick.