Documentation
What the Index Server implements, how it behaves, and how to run it safely. The pages below are the product's own docs, summarized. Where a section describes a boundary or limitation, that boundary is stated honestly — accuracy over marketing.
start here
Getting Started
Install the server, verify the OpenSearch 3.5 handshake, create an index, index a document, and run full-text, kNN and hybrid searches — with copy-paste curl and opensearch-py examples.
OpenSearch 3.5query + agg DSL
Compatibility matrix
The generated compatibility matrix records exactly which OpenSearch 3.5
capabilities work, which are cleanly rejected with an honest 4xx/501, and which
(if any) are gaps. It is generated against compat.version=3.5.0.
- Queries (supported): match_all, term, terms, range, exists, prefix, wildcard, regexp, ids, match, match_phrase, match_bool_prefix, match_phrase_prefix, multi_match, query_string, simple_query_string, bool, fuzzy, constant_score, function_score, knn, nested, geo_distance, span_near.
- Aggregations (supported): terms, avg, sum, min, max, stats, extended_stats, value_count, cardinality, percentiles, percentile_ranks, date_histogram, auto_date_histogram, range, date_range, filter, filters, nested, top_hits, geo_distance, multi_terms, missing.
- Search features (supported): highlight, _source filtering, sort, from/size, search_after, scroll, suggest.
- Docs & index (supported): create index, get mapping, bulk, mget, msearch, update, delete_by_query, alias, index template.
- Ops (supported): _cat/indices, _stats, _cluster/health, _count, _refresh, _forcemerge, snapshot.
- Cleanly rejected (by design): significant_terms & composite aggregations (501), point-in-time (400), reindex (400).
A supported row means the contract is accepted and returns the right
response shape — not that ranking is byte-identical to Lucene/OpenSearch. Ranking
fidelity (BM25 vs Lucene, RRF), aggregation numeric correctness at scale, and
analyzer token-stream fidelity are validated separately by the differential suite.
See docs/COMPAT-MATRIX.md in the release bundle for the full generated table.
QdrantPineconeChromaAlgoliaElasticsearch 8
Multi-dialect adapters
Optional listeners let clients built for other vector databases index and search the same engine and data. The adapters are versioned compatibility adapters — a deliberate data-plane subset (CRUD + query) of each vendor's surface, not the full control-plane / advanced feature set.
- Qdrant: create/get collection, upsert / query / filtered query / scroll / delete points. (Snapshots, recommend/discover — not implemented.)
- Pinecone: upsert, query, query + metadata filter, fetch, update, delete, describe_index_stats, list. (Control-plane index create/list — not implemented.)
- Chroma (v2): heartbeat, version, create collection, add, query, get.
- Algolia: batch (save objects), query, facets, filters, set settings. (Query Rules — not implemented.)
- Elasticsearch 8.x flavor: product header + handshake, dense_vector mapping, top-level knn — for elasticsearch-py/java 8.x clients (
compat.flavor=elasticsearch).
Each adapter cleanly rejects the activities it does not implement (honest 4xx).
Vendor ranking/relevance is served with BM25/RRF, not each vendor's native model.
See docs/DIALECT-MATRIX.md for the full generated table.
MCPClaudeAI agentsStreamable HTTP
MCP for AI agents
An optional Model Context Protocol listener presents the server to Claude and other agents as a set of tools — create indices, load documents, and run full-text, vector, and hybrid search directly, with no glue code. One engine and dataset, an extra protocol on its own port; disabled by default.
Enable it in server.properties (or with the admin-console toggle):
dialects.mcp-port=9400
It speaks MCP over Streamable HTTP at POST /mcp, authenticated
with the same API key (or a scoped key) as every other listener — a scoped key keeps
its per-index limits through the agent. Connect Claude Code:
claude mcp add searchai --transport http http://<host>:9400/mcp \
--header "Authorization: Bearer <api-key>"
Thirteen tools cover the workflow: list_indices, create_index,
index_document, bulk_index, get_document,
delete_document, delete_index, get_mapping,
search (full query-DSL passthrough), count,
hybrid_search (BM25 + vector, RRF-fused), cluster_health,
and server_stats. Write tools refresh by default, so a follow-up search
sees the change immediately.
Auto-embed vector fields embed at ingest through the bundled inference engine and default to 1024 dimensions — near-full quality at ~half the storage and search cost of a full-width model, with the query embedded at the same width. The admin console also shows a live resource panel: CPU load, cores, memory, disk and uptime.
_snapshotfilesystem repos
Backup & restore
The server implements a native OpenSearch-compatible _snapshot repository
API for filesystem repositories. Register a repository, snapshot all or a
subset of indices, and restore — the server stays live throughout.
PUT /_snapshot/my_repo
{ "type": "fs", "settings": { "location": "/backups/searchai" } }
PUT /_snapshot/my_repo/snap_2026_09_09?wait_for_completion=true
{ "indices": "idx001,idx002" } # omit for every index
POST /_snapshot/my_repo/snap_2026_09_09/_restore
{ "indices": "idx001",
"rename_pattern": "(.+)", "rename_replacement": "$1_restored" }
Each index is copied point-in-time consistently (the engine briefly quiesces
writers to that index, flushes, copies and validates); all other indices and reads
stay live. A cold copy of the self-contained data.dir is also supported for
whole-node backups.
Limitations: filesystem repositories only (no S3/object-store types);
snapshots are full copies (not incremental/deduplicated); snapshot/restore is
node-local (not cluster-coordinated); restore requires the target index to be
absent. A replica is not a backup. See docs/BACKUP.md.
primary/replicaread scale-out
Clustering — contract & limits
Multi-node mode is asynchronous primary/replica replication (full replication — every node stores every index), not a quorum-consensus cluster. One primary accepts writes; followers forward writes to it and replay the op-log; reads are served locally on any node. Inter-node traffic is authenticated with a cluster secret and can be TLS-encrypted.
What it is not — do not rely on these: no quorum/consensus, split-brain is possible under a partition, acknowledged writes can be lost if the primary dies before a replica catches up, extra nodes add read replicas (not shard capacity — a dataset must fit one node), and snapshot/restore is node-local.
Supported production posture: single node (recommended; pair with regular
backups), or an externally-fenced single primary + read replicas where an
external mechanism guarantees exactly one primary. Do not deploy multi-node as an
HA system-of-record until durable consensus with fencing lands. See
docs/CLUSTERING.md.
honest assessmentread before public exposure
Production readiness
The product ships an honest production-readiness assessment. It works today as a SearchBlox drop-in and in demos; the notes below are what to plan for before a public, high-ingest deployment.
- Engine memory under real ingest: per-segment FTS memory can grow with segment count on very wide (many-text-field) mappings. Mitigations: narrow which fields get full-text, tune auto-optimize to keep segment count low, and provision RAM to measured peak, not average.
- Cluster: single-primary, asynchronous, no automatic failover (see Clustering above). Run single-node or externally-fenced until consensus lands.
- Security for public exposure: always set API keys (never run with auth
off), use CA-issued TLS certs, do not expose
:9200directly — front it with a gateway/WAF in a private subnet, and add rate limiting. - Observability:
/health,_console/statsand optimize logs exist; a metrics endpoint, slow-query log and alerting are recommended additions.
Foundations already in place: API-key auth (admin + scoped keys), TLS listeners,
op-log fsync durability with torn-tail recovery, native _snapshot/restore, and
compatibility + fidelity matrices run in CI. See docs/PRODUCTION-READINESS.md.