PrepZone Logo
PrepZone

Elasticsearch Production Ops

Analyzers, index lifecycle management, reindexing, and the OpenSearch fork note.

Why this matters

  • VaultCommerce autocomplete needs edge n-grams at index time; applying n-grams to the main title field bloats index size — use a sub-field or dedicated suggest index.
  • Search analytics indices grow daily — ILM rolls vaultcommerce-search-logs-* to warm nodes, then deletes after 90 days.
  • Elasticsearch 8.x licensing shifted; AWS OpenSearch is a fork — know which engine your org runs and test client compatibility.
Cluster
Node 1Shard 0 primary
Node 2Shard 1 primary
Node 3Shard 0 replica
An index is split into shards across nodes. Replicas provide failover and read scaling.

Custom analyzers

Java
PUT /vaultcommerce-products
{
  "settings": {
    "analysis": {
      "analyzer": {
        "title_search": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding", "english_stemmer"]
        },
        "title_suggest": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": ["lowercase", "edge_ngram_filter"]
        }
      },
      "filter": {
        "english_stemmer": { "type": "stemmer", "language": "english" },
        "edge_ngram_filter": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "title_search",
        "fields": {
          "suggest": { "type": "text", "analyzer": "title_suggest" }
        }
      }
    }
  }
}

Search queries target title; typeahead targets title.suggest — different analyzers, different purposes.

Analyzer components

  • Character filters — strip HTML, map chars (& → and).
  • Tokenizer — split stream into tokens (standard, whitespace, keyword).
  • Token filters — lowercase, stem, n-gram, synonym expansion.

Index Lifecycle Management

Java
PUT _ilm/policy/vaultcommerce-logs-policy
{
  "policy": {
    "phases": {
      "hot": {
        "actions": {
          "rollover": { "max_size": "50gb", "max_age": "7d" }
        }
      },
      "warm": {
        "min_age": "14d",
        "actions": { "shrink": { "number_of_shards": 1 } }
      },
      "delete": { "min_age": "90d", "actions": { "delete": {} } }
    }
  }
}

VaultCommerce attaches this policy to vaultcommerce-search-logs-000001 with "index.lifecycle.rollover_alias": "vaultcommerce-search-logs".

Reindex without downtime

Mapping changes require reindex to a new index, then alias swap:

Java
POST _reindex
{ "source": { "index": "vaultcommerce-products-v1" }, "dest": { "index": "vaultcommerce-products-v2" } }

POST _aliases
{
  "actions": [
    { "remove": { "index": "vaultcommerce-products-v1", "alias": "vaultcommerce-products" } },
    { "add":    { "index": "vaultcommerce-products-v2", "alias": "vaultcommerce-products" } }
  ]
}

Quick recall

Everything you need if you only revisit this box.

  • Analyzers = character filters + tokenizer + token filters — applied at index and query time.
  • Use separate sub-fields for suggest n-grams vs main search analyzer.
  • ILM automates rollover, warm shrink, and delete for growing log indices.
  • Reindex + alias swap is the safe path for mapping changes at scale.
  • Test analyzer output with _analyze API before bulk indexing.
  • Confirm Elasticsearch vs OpenSearch compatibility in your deployment.

Test yourself

Answer these before moving on — recall is what makes it stick.