PrepZone Logo
PrepZone

Graph Databases with Neo4j

Nodes, edges, and traversal queries for recommendations and fraud detection.

Why this matters

  • VaultCommerce recommendation and fraud teams ask relationship questions that explode in SQL: multi-hop paths, shared devices, circular referral chains.
  • Neo4j's Cypher query language expresses traversals in lines, not nested subqueries.
  • Graph stores are specialised — knowing when they earn their ops cost separates thoughtful architects from resume-driven design.
SQL (Postgres)
ACID transactions
Complex joins
Fixed schema
NoSQL
Flexible schema
Horizontal scale
Specialised access
Structured relations with joins favour SQL. Flexible schema and horizontal scale favour NoSQL.

Nodes, relationships, and properties

Graph primitives

  • Node — an entity with labels and properties (:User {id, email}).
  • Relationship — directed, typed edge (:PURCHASED, :REFERRED, :SHARES_DEVICE).
  • Traversal — walk edges to arbitrary depth with indexed start points.
  • Pattern matching — Cypher describes shapes, not join order.
Java
// Products co-purchased with items in the current cart
MATCH (u:User {id: $userId})-[:PURCHASED]->(p:Product)
MATCH (p)-[:BOUGHT_WITH]->(rec:Product)
WHERE NOT rec.id IN $cartProductIds
RETURN rec.id, rec.name, count(*) AS strength
ORDER BY strength DESC
LIMIT 8

VaultCommerce precomputes :BOUGHT_WITH edges nightly from order history so product pages serve recommendations in under 20 ms.

Recommendation vs relational joins

A three-hop SQL query — users who bought A also bought B through mutual categories — devolves into temporary tables and minutes of tuning. The same question in Neo4j is a declarative pattern:

Java
MATCH (target:Product {id: $productId})<-[:PURCHASED]-(u:User)
MATCH (u)-[:PURCHASED]->(related:Product)
WHERE related.id <> $productId
RETURN related.name, count(u) AS buyers
ORDER BY buyers DESC
LIMIT 5

Fraud detection patterns

VaultCommerce's risk engine flags accounts that share a device, shipping address, or payment fingerprint within two hops of a known fraudster. Graph traversals with depth limits (*1..3) replace brittle rule engines.

Java
MATCH (suspect:User {id: $userId})
MATCH path = (suspect)-[:SHARES_DEVICE|SHARES_ADDRESS*1..3]-(flagged:User)
WHERE flagged.risk_score > 80
RETURN path LIMIT 10

Index node properties used as traversal entry points (User.id, Device.fingerprint). Relationship-only queries without an indexed anchor scan the graph.

Syncing with the system of record

Neo4j is not VaultCommerce's ledger. Orders and payments stay in Postgres; a CDC pipeline projects PURCHASED and REFERRED edges into the graph. Eventual consistency is acceptable for recommendations; fraud scoring tolerates seconds of lag with compensating rules on checkout.

Use caseGraph fitVaultCommerce store
Checkout totalsPoor — needs ACIDPostgres
Co-purchase recommendationsExcellent — 2-hop traversalNeo4j
Product search by keywordPoor — no inverted indexElasticsearch
Referral chain payoutsGood — path queriesNeo4j + Postgres ledger
  • Checkout totals

    Graph fitPoor — needs ACID
    VaultCommerce storePostgres
  • Co-purchase recommendations

    Graph fitExcellent — 2-hop traversal
    VaultCommerce storeNeo4j
  • Product search by keyword

    Graph fitPoor — no inverted index
    VaultCommerce storeElasticsearch
  • Referral chain payouts

    Graph fitGood — path queries
    VaultCommerce storeNeo4j + Postgres ledger

Quick recall

Everything you need if you only revisit this box.

  1. Graph databases optimise relationship traversals, not tabular reporting.
  2. Model nodes, typed edges, and properties; write Cypher as pattern matches.
  3. VaultCommerce uses Neo4j for recommendations and fraud rings, not checkout.
  4. Index entry-point properties; always anchor traversals from a selective start.
  5. Sync from the system of record via CDC — the graph is a derived view.

Test yourself

Answer these before moving on — recall is what makes it stick.