PrepZone Logo
PrepZone

Column-Family Stores with Cassandra

Write-heavy event logs, partition keys, and tunable consistency for VaultCommerce analytics.

Why this matters

  • VaultCommerce ingests millions of clickstream and order-status events per day; writing every event to Postgres would saturate the primary.
  • Cassandra's tunable consistency and append-friendly storage fit time-series and audit logs.
  • Interviews test whether you can explain partition keys, clustering columns, and why SELECT * without a partition key is a cluster-wide scan.
Vertical
usersid, name, email
user_profilesbio, avatar
Horizontal
Shard Auser_id 0–999
Shard Buser_id 1000–1999
Vertical splits tables by column group. Horizontal splits rows across nodes by key.

Rows, partitions, and clustering

Cassandra tables are sorted maps: a partition key routes a row to a node; clustering columns order rows inside that partition.

Cassandra vocabulary

  • Partition key — determines which node stores the data; keep partitions under ~100 MB.
  • Clustering columns — sort rows within a partition (e.g. event time descending).
  • Column family — logical grouping of columns; think wide rows, not skinny SQL tables.
  • Tunable consistency — per-query ONE, QUORUM, or ALL depending on freshness needs.
Java
CREATE TABLE order_events (
  order_id    UUID,
  event_time  TIMESTAMP,
  event_type  TEXT,
  payload     TEXT,
  PRIMARY KEY (order_id, event_time)
) WITH CLUSTERING ORDER BY (event_time DESC);

VaultCommerce's order timeline query — "show status history for order X" — hits one partition keyed by order_id. No cross-partition join required.

Query-first table design

In Cassandra you create one table per query pattern. VaultCommerce maintains separate tables for "events by order" and "events by customer per day" because a single flexible table cannot serve both efficiently.

Java
-- Access pattern: daily funnel per customer
CREATE TABLE customer_daily_events (
  customer_id UUID,
  day         DATE,
  event_time  TIMESTAMP,
  event_type  TEXT,
  PRIMARY KEY ((customer_id, day), event_time)
);

The double parentheses ((customer_id, day), event_time) make (customer_id, day) the partition key — all events for one customer on one day live together.

Write path and compaction

Cassandra appends writes to a commit log and memtable, then flushes immutable SSTables. Compaction merges files in the background. This design favours sequential writes over in-place updates — ideal for append-only event logs, painful for frequently updated counters unless you batch them.

AspectPostgres (VaultCommerce orders)Cassandra (VaultCommerce events)
Sweet spotACID checkout, joinsHigh-volume append-only writes
Query modelAd hoc SQLPartition-key lookups only
ScalingVertical + read replicasLinear horizontal write scale
ConsistencyStrong per transactionTunable per read/write
  • Sweet spot

    Postgres (VaultCommerce orders)ACID checkout, joins
    Cassandra (VaultCommerce events)High-volume append-only writes
  • Query model

    Postgres (VaultCommerce orders)Ad hoc SQL
    Cassandra (VaultCommerce events)Partition-key lookups only
  • Scaling

    Postgres (VaultCommerce orders)Vertical + read replicas
    Cassandra (VaultCommerce events)Linear horizontal write scale
  • Consistency

    Postgres (VaultCommerce orders)Strong per transaction
    Cassandra (VaultCommerce events)Tunable per read/write

Operational trade-offs

Repairs, tombstones, and hot partitions are real production concerns. A celebrity product launch that funnels all traffic through one product_id partition will bottleneck a single node. VaultCommerce salts hot keys or buckets high-cardinality IDs across synthetic suffixes.

Quick recall

Everything you need if you only revisit this box.

  1. Cassandra partitions data by a key you choose — design tables around queries, not entities.
  2. Clustering columns sort rows inside a partition for range reads.
  3. VaultCommerce uses Cassandra for append-only order and clickstream events.
  4. Avoid hot partitions; salt or bucket keys that concentrate traffic.
  5. Tunable consistency trades freshness for availability on every read and write.

Test yourself

Answer these before moving on — recall is what makes it stick.