Introducing the Index Server: a private, OpenSearch-compatible search engine
September 2026 · 6 min read
Search is infrastructure. If your product searches documents, powers a RAG pipeline, or ranks results, you depend on a search engine — and today that usually means either running OpenSearch/Elasticsearch yourself, or renting a hosted vector database and paying by the document and the query. We wanted a third option: a single, private engine you own, that speaks the API your code already knows, and that scales at a fixed cost on hardware you control. That's the Index Server.
What it is
The Index Server is a self-hosted search + vector engine that implements
the OpenSearch 3.5 REST API. It's a single self-contained binary on an
embedded high-performance vector-engine core with HNSW/IVF/DiskANN dense + sparse
vectors, inverted-index scalar filters, BM25 full-text search, native hybrid RRF
fusion, and WAL durability. On top of that engine, the server translates the
OpenSearch query and aggregation DSL, so existing OpenSearch clients —
opensearch-py, opensearch-java, the REST high-level client — connect
unchanged. It's the search backend SearchBlox points at on :9200 over TLS, and
it's yours to run the same way.
A genuine drop-in
"Compatible" is easy to claim and easy to get wrong, so we hold it to a generated compatibility matrix that records exactly what works. The supported surface includes:
- Full-text search: the
match/match_phrase/match_bool_prefix/match_phrase_prefix/multi_match/query_string/simple_query_stringfamily, on BM25. - Structured filters:
term,terms,range,exists,prefix,wildcard,regexp,ids,fuzzy,bool,constant_score,function_score,nested,geo_distance,span_near— pushed down to the engine. - Vector & hybrid:
knn,script_scoreknn_scorein the SearchBlox RAG shape, and nativehybridqueries that fuse full-text and vector results with Reciprocal Rank Fusion. - Aggregations: terms, the metric family (avg/sum/min/max/stats/
extended_stats/percentiles/percentile_ranks/cardinality/value_count),
date_histogram,auto_date_histogram,range/date_range,filter/filters,nested,top_hits,multi_terms,missing— the facets SearchBlox renders. - The rest of the plumbing: highlight,
_sourcefiltering, sort, from/size,search_after, scroll, suggest; index/mapping CRUD,_bulk,mget,msearch,update,delete_by_query, aliases, index templates; and ops routes like_cat/indices,_stats,_cluster/health,_count,_refresh,_forcemerge.
Just as important is what it cleanly rejects instead of faking:
significant_terms and composite aggregations, point-in-time, and reindex
all return honest 4xx/501 responses. A green row in the matrix means the contract is
accepted and the response shape is right — not that ranking is byte-identical to
Lucene. We validate ranking fidelity separately with a differential suite against a
live OpenSearch.
One engine, many dialects
Teams rarely standardize on one client. So beyond OpenSearch, the server offers
optional data-plane adapters — the same engine and data behind each — for
Qdrant, Pinecone, Chroma and Algolia, plus an
Elasticsearch 8.x presentation flavor (set compat.flavor=elasticsearch)
that satisfies elasticsearch-py/java 8.x clients. These are deliberately a
subset — indexing and query, not each vendor's full control-plane or advanced
features — and they reject what they don't implement rather than pretend. It means
you can point a Pinecone or Qdrant client at your own box for the operations that
matter, and consolidate onto one engine. See the
dialect matrix for the exact coverage.
Backups that are actually backups
The server implements a native OpenSearch-compatible _snapshot repository
API for filesystem repositories. You register a repository, snapshot all or selected
indices point-in-time-consistently — the engine briefly quiesces writers to each
index, flushes, copies and validates — and the server stays live throughout.
Restore rewrites the index to a fresh id and opens it immediately. A cold copy of the
self-contained data directory works too. We're honest about the edges: filesystem
repositories only (no object-store types yet), full copies (not incremental), and
node-local (not cluster-coordinated). And a replica is never a backup. The details
are in Backup & restore.
Private, and fixed-cost
Everything runs inside your network — on-prem, air-gapped, regulated or edge. No documents, queries or vectors leave the box; there's no per-document or per-query metered billing. You size a machine and you're done — cost scales with hardware, not usage. Security is built in: API-key auth with an admin key plus scoped, per-index keys (read / write / admin roles), in-process TLS on all listeners, and encrypted inter-node cluster traffic.
Where the edges are
We'd rather you deploy this knowing its limits. Two are worth stating up front:
- Clustering is asynchronous primary/replica, not consensus. One primary takes writes; followers replay the op-log and serve reads locally. That gives you read scale-out and read replicas — but not quorum HA. Split-brain is possible under a partition, there's no automatic failover yet, and extra nodes add read replicas, not shard capacity (a dataset must fit one node). The supported production posture today is single-node, or an externally-fenced single primary + read replicas. See Clustering — contract & limits.
- Engine memory under heavy, wide ingest. Full-text memory can grow with segment count on very wide (many-text-field) mappings; the mitigations are to narrow which fields get full-text, keep auto-optimize aggressive, and provision RAM to measured peak. Full detail in Production readiness.
These are the same notes we hold ourselves to internally — we publish them because an accurate boundary is more useful than a marketing claim.
Try it
One line on a Linux host installs the self-contained bundle as a systemd service:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
curl http://<host>:9200/ # OpenSearch 3.5 handshake
curl http://<host>:9200/health # {"status":"green"}
Then create an index, index a document, and search — the Getting Started guide walks it end to end with curl and opensearch-py. Downloads for Linux arm64/amd64 are on the downloads page; the full documentation hub is here.
It's free to use under the Elastic License 2.0. If you want enterprise search, RAG and commercial support on top of it, that's the SearchAI platform — talk to us.