Why this matters
- VaultCommerce product search cannot rely on SQL
LIKEat catalog scale — Elasticsearch indexes titles, descriptions, and facets for sub-second fuzzy search. - Shard and replica counts are fixed at index creation (with split/shrink limits) — wrong sizing causes reindex pain later.
- Understanding cluster topology separates "search feels slow" triage from application bugs.
Cluster building blocks
Core vocabulary
- Cluster — one or more nodes sharing a cluster name; a node belongs to exactly one cluster.
- Index — logical namespace (like a database table) for similar documents — e.g.,
vaultcommerce-products. - Shard — horizontal slice of an index; primary shards handle writes; replica shards copy primaries.
- Document — JSON object with
_id,_source, and analyzed fields for search.
VaultCommerce runs a three-node cluster: two primary shards plus one replica each — survives single-node loss while serving browse traffic from replicas.
Indexing flow
A write routes to the primary shard for the document ID hash. The primary indexes in Lucene, then replicates to replica shards asynchronously.
curl -X PUT "localhost:9200/vaultcommerce-products" -H 'Content-Type: application/json' -d'
{
"settings": {
"number_of_shards": 2,
"number_of_replicas": 1
}
}'
curl -X POST "localhost:9200/vaultcommerce-products/_doc" -H 'Content-Type: application/json' -d'
{
"sku": "SKU-8842",
"title": "Trail Pack 40L",
"category": "outdoor",
"price": 89.00,
"tags": ["hiking", "waterproof"]
}'
Node roles (Elasticsearch 8.x)
Modern clusters assign roles: master, data, ingest, ml. Small VaultCommerce staging clusters combine roles; production separates dedicated master-eligible nodes (three for quorum) from data nodes.
curl -s localhost:9200/_cat/nodes?v&h=name,node.role,heap.percent,disk.used_percent
curl -s localhost:9200/_cluster/health?pretty
Near-real-time search
Documents are searchable after refresh (default ~1 s). Bulk indexing during catalog imports should temporarily raise refresh_interval to reduce segment churn.
Quick recall
Everything you need if you only revisit this box.
- Cluster → indices → shards (primary + replica) on nodes.
- Writes hit primary shards; replicas copy for HA and read scale.
- VaultCommerce indexes products as JSON documents in
vaultcommerce-products. - Size shards at creation — plan nodes before shard explosion.
- Separate master-eligible nodes in production for quorum stability.
- Refresh interval trades indexing throughput against search visibility latency.
Test yourself
Answer these before moving on — recall is what makes it stick.