Migrating your index from Elasticsearch or OpenSearch
September 2026 · 9 min read
The Index Server speaks the OpenSearch 3.5 REST API, so a migration from Elasticsearch or OpenSearch isn't a re-platforming project — it's a data copy, a relevance check, and an endpoint swap. This is the end-to-end playbook we use, including the two built-in console tools that do most of the work: Migrate (zero-touch import) and Evaluate (prove the results match before you cut over).
Step 0 — stand up the server
One line on a Linux host installs the server (and the inference engine for semantic search) as a systemd service:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Set server.api-key in conf/server.properties before
exposing it, then open the admin console at http://<host>:9200/console
and log in with the key. Everything below can be done from the console UI or with
curl.
Step 1 — copy the data: two paths
Path A: zero-touch migration (console Migrate tab)
The console's Migrate tab connects to your running Elasticsearch or OpenSearch, discovers its indexes, and imports them — mappings pass through as-is and documents stream over the scroll API. No exporter to write, no intermediate files:
- Discover — point it at the source (
http://source-host:9200, with the same basic-auth or Bearer credentials your source accepts). It lists every index with doc counts. - Start — select the indexes and go. One job runs at a time with per-index progress you can watch from the tab.
- Verify — each finished index records source vs target document counts, so the completeness check is built into the job.
The same tool imports from Qdrant, Pinecone, Chroma, Weaviate and Solr if
you're consolidating vector stores at the same time. Two current limits to plan
around: sources are plain http:// (for TLS-only cloud endpoints, run
a local tunnel and point the migrator at that), and a job doesn't resume across a
server restart — re-run it instead.
Path B: client-side reindex (when you want to transform)
Because both sides speak the same API, the standard scan-and-bulk loop works
with the client you already use — and it's the right path when you want to
reshape documents in flight (rename fields, drop legacy ones, add a
knn_vector field for semantic search):
from opensearchpy import OpenSearch, helpers
src = OpenSearch("http://old-cluster:9200", http_auth=("user", "pass"))
dst = OpenSearch("http://index-server:9200", http_auth=("admin", "YOUR_API_KEY"))
# create the target with the source mapping (tweak here if reshaping)
mapping = src.indices.get(index="products")["products"]
dst.indices.create(index="products", body={
"settings": {"index": {"knn": True}},
"mappings": mapping["mappings"],
})
def stream():
for hit in helpers.scan(src, index="products", size=1000):
yield {"_index": "products", "_id": hit["_id"], "_source": hit["_source"]}
helpers.bulk(dst, stream(), chunk_size=500)
Writing to a missing index also just works — 1.1.0 auto-creates it with a mapping inferred from the first document (index templates win over inference), which makes quick trial imports frictionless. For production indexes, create the real mapping explicitly as above.
Step 2 — prove relevance before you switch (console Evaluate tab)
The part most migrations skip — and then regret. The console's Evaluate tab holds golden query sets (real queries from your application, optionally with graded judgments). A run replays every query against the index server and, with your old cluster configured as the baseline, sends the same body to both engines and compares:
- overlap@10 (Jaccard) and top-1 agreement between the two engines, per query and averaged,
- NDCG@10 where you've supplied judgments,
- latency on both sides, from the same machine.
Run history is kept, and a drop of more than 0.05 in average overlap or NDCG vs the previous run is flagged as a regression — so post-migration config changes stay honest too. When the numbers say "same results, at these latencies," you have your cutover evidence.
Step 3 — point your clients at it
For OpenSearch clients (opensearch-py, opensearch-java,
the REST high-level client) the migration is the endpoint string. Auth is
standard basic auth (any username, your API key as the password) or a Bearer
token. For Elasticsearch clients, set compat.flavor=elasticsearch
in the server config — the handshake, product headers and
dense_vector/top-level-knn surface then present as
Elasticsearch 8.x, so elasticsearch-py/-java 8.x connect
unchanged.
This is exactly how SearchBlox SearchAI 12.2 runs it in production: the
platform points at the index server like any OpenSearch endpoint on
:9200 — console, crawlers, facets, RAG — with no application
changes.
Step 4 — cutover checklist
- API key set, port fronted by your gateway/LB; health check is
unauthenticated
GET /. - Hit counting: set
index.track-total-hits=1000(the same cap Elasticsearch defaults to) — exact totals on large matched sets are the one place latency differs materially. - Templates & analyzers: index templates are supported; custom analyzers cover standard/whitespace/ngram/edge_ngram tokenizers, stopwords (applied server-side), stemming, ascii-folding, and query-time synonyms — you can edit synonym lists without reindexing. Char filters and per-field search analyzers aren't implemented; check the compatibility docs if you rely on exotic analysis chains.
- Semantic upgrade (optional but why not): add a
knn_vectorfield with anembedblock to any mapping and the bundled inference server embeds at index time; queries gainneuralandhybrid(RRF) clauses your old cluster never had. - Backups: register a filesystem
_snapshotrepository and snapshot on your usual schedule. (Note: Lucene snapshots from your old cluster can't be restored here — migrate via Migrate/reindex, then snapshot natively.) - HA later, config not migration: a single node can become a 3-voter raft cluster with four config lines per node when traffic demands it — see the scaling article.
What to expect
On our published Graviton benchmarks, the index server ingests text 17–44% faster than OpenSearch 3.8 in roughly 1/15th of the memory, serves interactive-range query latencies, and adds vector + hybrid search that leads the vector-database field on ingest. The honest trade-offs are published on the same page — phrase-heavy query mixes and doc-values-style aggregations remain OpenSearch strengths today.
Start with the Getting Started guide, or grab a build from the downloads page and have the Migrate tab pulling your first index over in the next ten minutes.