PrepZone Logo
PrepZone

Elasticsearch Architecture

Clusters, nodes, shards, and replicas — how VaultCommerce indexes millions of products.

Why this matters

  • VaultCommerce product search cannot rely on SQL LIKE at catalog scale — Elasticsearch indexes titles, descriptions, and facets for sub-second fuzzy search.
  • Shard and replica counts are fixed at index creation (with split/shrink limits) — wrong sizing causes reindex pain later.
  • Understanding cluster topology separates "search feels slow" triage from application bugs.

Cluster building blocks

Cluster
Node 1Shard 0 primary
Node 2Shard 1 primary
Node 3Shard 0 replica
An index is split into shards across nodes. Replicas provide failover and read scaling.

Core vocabulary

  • Cluster — one or more nodes sharing a cluster name; a node belongs to exactly one cluster.
  • Index — logical namespace (like a database table) for similar documents — e.g., vaultcommerce-products.
  • Shard — horizontal slice of an index; primary shards handle writes; replica shards copy primaries.
  • Document — JSON object with _id, _source, and analyzed fields for search.

VaultCommerce runs a three-node cluster: two primary shards plus one replica each — survives single-node loss while serving browse traffic from replicas.

Indexing flow

A write routes to the primary shard for the document ID hash. The primary indexes in Lucene, then replicates to replica shards asynchronously.

Java
curl -X PUT "localhost:9200/vaultcommerce-products" -H 'Content-Type: application/json' -d'
{
  "settings": {
    "number_of_shards": 2,
    "number_of_replicas": 1
  }
}'
Java
curl -X POST "localhost:9200/vaultcommerce-products/_doc" -H 'Content-Type: application/json' -d'
{
  "sku": "SKU-8842",
  "title": "Trail Pack 40L",
  "category": "outdoor",
  "price": 89.00,
  "tags": ["hiking", "waterproof"]
}'

Node roles (Elasticsearch 8.x)

Modern clusters assign roles: master, data, ingest, ml. Small VaultCommerce staging clusters combine roles; production separates dedicated master-eligible nodes (three for quorum) from data nodes.

Java
curl -s localhost:9200/_cat/nodes?v&h=name,node.role,heap.percent,disk.used_percent
curl -s localhost:9200/_cluster/health?pretty

Near-real-time search

Documents are searchable after refresh (default ~1 s). Bulk indexing during catalog imports should temporarily raise refresh_interval to reduce segment churn.

Quick recall

Everything you need if you only revisit this box.

  • Cluster → indices → shards (primary + replica) on nodes.
  • Writes hit primary shards; replicas copy for HA and read scale.
  • VaultCommerce indexes products as JSON documents in vaultcommerce-products.
  • Size shards at creation — plan nodes before shard explosion.
  • Separate master-eligible nodes in production for quorum stability.
  • Refresh interval trades indexing throughput against search visibility latency.

Test yourself

Answer these before moving on — recall is what makes it stick.