Changelog
Every release of the Index Server, newest first. All binaries are built natively per platform and published on the downloads page with sha256 checksums.
1.3.1October 5, 2026 current
1.3.1 — OpenSearch drop-in compatibility fixes
A drop-in patch over 1.3.0 — same config, same on-disk format, same wire compatibility. It sharpens behavior for OpenSearch clients that probe the server across multiple URL schemes and HTTP verbs during startup. Raw notes: RELEASE-NOTES-1.3.1.md.
Fixed
- A client that opened a TLS connection to the plaintext port could stall the connection's read loop. The server now detects a TLS handshake on a plaintext listener and closes it immediately, so the client fails fast and falls back — matching stock OpenSearch behavior.
HEADon an index or document now reports the sameContent-LengthaGETwould (with the body omitted), instead ofContent-Length: 0.POST /{index}/_openand/_closenow return200 acknowledged(indices are always open), so client startup routines that toggle index state no longer error.
1.3.0September 23, 2026
1.3.0 — Agent-ready: MCP tools + a live admin console
Drive the index server from Claude and other AI agents, and watch the node in real time from the console. Drop-in over 1.2.0 — same config, same on-disk format, same wire compatibility; every new surface is opt-in. Raw notes: RELEASE-NOTES-1.3.0.md.
MCP for AI agents
- Model Context Protocol listener — a new opt-in listener presents the server to Claude and other agents as a set of tools: create indices, load documents, and run full-text, vector, and hybrid search directly, with no glue code. One engine and dataset, an extra protocol on its own port.
- Streamable HTTP at
POST /mcp, authenticated with the same API key (or a scoped key) as everything else — a scoped key keeps its per-index limits through the agent. Enable withdialects.mcp-portor the console toggle; disabled by default. - Thirteen tools including
search(full query DSL passthrough),hybrid_search(BM25 + vector, RRF-fused), bulk loading, and index management. Writes refresh by default for read-after-write.
Admin console
- Live resource panel — CPU load (1/5/15 min) and process CPU usage, core count, resident memory, disk, and uptime, refreshing on an interval so the node's health is visible at a glance.
- Dialect toggles include MCP — enable or disable the MCP listener from the console, with the exact endpoint URL and the one-line command to connect Claude Code shown inline.
Existing deployments are unaffected until they opt in: the MCP listener stays off unless a port is set.
1.2.0September 17, 2026
1.2.0 — Throughput release
Higher clustered-write throughput, lower replicated-write latency, faster reads, and reduced per-request memory. Drop-in over 1.1.0 — same config, same on-disk format, same wire compatibility. Raw notes: RELEASE-NOTES-1.2.0.md.
Write path
- Group-commit replication log — a bulk (and each replicated batch) fsyncs the op-log once per batch instead of once per document; clustered ingest is no longer bounded by per-document fsync rate.
- Event-driven replication — replicas receive fresh writes as they are appended rather than on a fixed tick; acked clustered writes commit in single-digit milliseconds.
- Keep-alive inter-node connections — replication and consensus RPCs reuse a persistent connection per peer (no new socket or TLS handshake per message).
- Batched bulk
update/delete; writes retry across a background merge window instead of failing an item.
Read path
- Compiled filter scans — cached-column filters resolve fields, bounds and terms once per query rather than per row.
- Verbatim
_sourcepassthrough for full-document hits (no parse/re-serialize round trip); lower per-request memory on large scored + aggregation queries; reader-writer locking on index lookups. - Exact
_countabove 10,000 matches for filter queries. - Reused inference connections for auto-embedding and reranking.
Operations
- Non-blocking snapshots —
wait_for_completion=falsereturns immediately and runs in the background; status viaGET _snapshot. - New
/_metricscountersearchai_fallback_iterator_scans_total; tuned HTTP connection reuse, keep-alive probing, and large-body handling.
1.1.0September 15, 2026
1.1.0 — Raft clustering, production guardrails, drop-in hardening
Validated end-to-end against a production SearchBlox 12.2.2 stack. Raw notes: RELEASE-NOTES-1.1.0.md.
Added
- Raft consensus clustering (
cluster.mode=raft) — leader election, quorum-committed writes (no acknowledged-write loss), automatic failover, snapshot bootstrap for late or restarted nodes, and live membership changes. Validated as a 3-node cluster serving a live crawl.single_primaryremains the default clustered mode; see clustering docs. - Auto-create-on-write (
index.auto-create, default on) — writing to a missing index creates it with a mapping inferred from the document, matching OpenSearchaction.auto_create_index. Index templates take precedence over inference. - Installer bundles the SearchAI Inference Server (embedding + reranker models),
so vector, hybrid and semantic search work natively after one install
(
--no-inferenceopts out). - Prometheus
/_metricsendpoint — per-index segments, FTS columns, RSS, circuit-breaker counters. - Aggregation bucket budget (
search.max-buckets) and per-request memory budget (search.max-request-mb) breakers; streamedsize:0aggregations.
Improved
- Memory guardrails proven on a real crawl — the emergency segment merge (default ceiling 16) bounds the per-document segment storm that previously drove 19 GB RSS at 114 docs; the same crawl now holds 2.5–4.5 GB. Replicas enforce the same guardrails as the primary.
- query_string — fielded search (
field:term,field:"phrase",field:(group)),+/-,AND/OR/NOT, cross-field AND, and textmust_not.
Fixed
- HTTP Basic auth accepted alongside Bearer (password checked against the API key) — required by clients that authenticate with basic auth.
_msearchheaderindexas an array (Java high-level client).- Request bodies on GET/DELETE (e.g.
GET _analyze/_searchwith a body) no longer break keep-alive connections.
Changed
- Background optimize sweep interval default 300 → 20 s (cheap on idle indices; run the same guardrail config on every cluster node).
- Known limits:
knn_vector/geo/date can't be inferred by auto-create (use a mapping or template); query_string boost/slop are parsed but not yet ranked.
1.0.0September 3, 2026
1.0.0 — First release
A private, OpenSearch-compatible search + vector engine on an embedded high-performance vector engine — one small native binary per platform. Raw notes: RELEASE-NOTES-1.0.0.md.
- OpenSearch-compatible REST (verified with opensearch-py and the Java client); Elasticsearch 8.x presentation mode; Qdrant/Pinecone/Chroma/Algolia wire-protocol listeners.
- Vector search 1.8× faster than OpenSearch at equal recall on Graviton; fastest vector ingest of the five engines benchmarked; filtered kNN ahead of Qdrant.
- Facets and sorts at or above OpenSearch speed; ~92 MB RSS after a 100k-doc query storm (vs ~1.4 GB for OpenSearch).
- Admin console (
/console): Dev Tools, Documents, Search tracing, zero-touch Migrate from OpenSearch/Elasticsearch/Qdrant/Pinecone/Chroma/Weaviate/Solr, and Evaluate (golden queries, side-by-side, regression gates). - Custom analyzers (stopwords, query-time synonyms, ngram, jieba), array and rank_features fields, neural search + rerank via the SearchAI Inference Server.
- Scoped API keys, native TLS everywhere, full-replica clustering with op-log replication.