← Home · Downloads · Docs

Changelog

Every release of the Index Server, newest first. All binaries are built natively per platform and published on the downloads page with sha256 checksums.

1.3.1October 5, 2026 current

1.3.1 — OpenSearch drop-in compatibility fixes

A drop-in patch over 1.3.0 — same config, same on-disk format, same wire compatibility. It sharpens behavior for OpenSearch clients that probe the server across multiple URL schemes and HTTP verbs during startup. Raw notes: RELEASE-NOTES-1.3.1.md.

Fixed

  • A client that opened a TLS connection to the plaintext port could stall the connection's read loop. The server now detects a TLS handshake on a plaintext listener and closes it immediately, so the client fails fast and falls back — matching stock OpenSearch behavior.
  • HEAD on an index or document now reports the same Content-Length a GET would (with the body omitted), instead of Content-Length: 0.
  • POST /{index}/_open and /_close now return 200 acknowledged (indices are always open), so client startup routines that toggle index state no longer error.

1.3.0September 23, 2026

1.3.0 — Agent-ready: MCP tools + a live admin console

Drive the index server from Claude and other AI agents, and watch the node in real time from the console. Drop-in over 1.2.0 — same config, same on-disk format, same wire compatibility; every new surface is opt-in. Raw notes: RELEASE-NOTES-1.3.0.md.

MCP for AI agents

  • Model Context Protocol listener — a new opt-in listener presents the server to Claude and other agents as a set of tools: create indices, load documents, and run full-text, vector, and hybrid search directly, with no glue code. One engine and dataset, an extra protocol on its own port.
  • Streamable HTTP at POST /mcp, authenticated with the same API key (or a scoped key) as everything else — a scoped key keeps its per-index limits through the agent. Enable with dialects.mcp-port or the console toggle; disabled by default.
  • Thirteen tools including search (full query DSL passthrough), hybrid_search (BM25 + vector, RRF-fused), bulk loading, and index management. Writes refresh by default for read-after-write.

Admin console

  • Live resource panel — CPU load (1/5/15 min) and process CPU usage, core count, resident memory, disk, and uptime, refreshing on an interval so the node's health is visible at a glance.
  • Dialect toggles include MCP — enable or disable the MCP listener from the console, with the exact endpoint URL and the one-line command to connect Claude Code shown inline.

Existing deployments are unaffected until they opt in: the MCP listener stays off unless a port is set.

1.2.0September 17, 2026

1.2.0 — Throughput release

Higher clustered-write throughput, lower replicated-write latency, faster reads, and reduced per-request memory. Drop-in over 1.1.0 — same config, same on-disk format, same wire compatibility. Raw notes: RELEASE-NOTES-1.2.0.md.

Write path

  • Group-commit replication log — a bulk (and each replicated batch) fsyncs the op-log once per batch instead of once per document; clustered ingest is no longer bounded by per-document fsync rate.
  • Event-driven replication — replicas receive fresh writes as they are appended rather than on a fixed tick; acked clustered writes commit in single-digit milliseconds.
  • Keep-alive inter-node connections — replication and consensus RPCs reuse a persistent connection per peer (no new socket or TLS handshake per message).
  • Batched bulk update/delete; writes retry across a background merge window instead of failing an item.

Read path

  • Compiled filter scans — cached-column filters resolve fields, bounds and terms once per query rather than per row.
  • Verbatim _source passthrough for full-document hits (no parse/re-serialize round trip); lower per-request memory on large scored + aggregation queries; reader-writer locking on index lookups.
  • Exact _count above 10,000 matches for filter queries.
  • Reused inference connections for auto-embedding and reranking.

Operations

  • Non-blocking snapshots — wait_for_completion=false returns immediately and runs in the background; status via GET _snapshot.
  • New /_metrics counter searchai_fallback_iterator_scans_total; tuned HTTP connection reuse, keep-alive probing, and large-body handling.

1.1.0September 15, 2026

1.1.0 — Raft clustering, production guardrails, drop-in hardening

Validated end-to-end against a production SearchBlox 12.2.2 stack. Raw notes: RELEASE-NOTES-1.1.0.md.

Added

  • Raft consensus clustering (cluster.mode=raft) — leader election, quorum-committed writes (no acknowledged-write loss), automatic failover, snapshot bootstrap for late or restarted nodes, and live membership changes. Validated as a 3-node cluster serving a live crawl. single_primary remains the default clustered mode; see clustering docs.
  • Auto-create-on-write (index.auto-create, default on) — writing to a missing index creates it with a mapping inferred from the document, matching OpenSearch action.auto_create_index. Index templates take precedence over inference.
  • Installer bundles the SearchAI Inference Server (embedding + reranker models), so vector, hybrid and semantic search work natively after one install (--no-inference opts out).
  • Prometheus /_metrics endpoint — per-index segments, FTS columns, RSS, circuit-breaker counters.
  • Aggregation bucket budget (search.max-buckets) and per-request memory budget (search.max-request-mb) breakers; streamed size:0 aggregations.

Improved

  • Memory guardrails proven on a real crawl — the emergency segment merge (default ceiling 16) bounds the per-document segment storm that previously drove 19 GB RSS at 114 docs; the same crawl now holds 2.5–4.5 GB. Replicas enforce the same guardrails as the primary.
  • query_string — fielded search (field:term, field:"phrase", field:(group)), +/-, AND/OR/NOT, cross-field AND, and text must_not.

Fixed

  • HTTP Basic auth accepted alongside Bearer (password checked against the API key) — required by clients that authenticate with basic auth.
  • _msearch header index as an array (Java high-level client).
  • Request bodies on GET/DELETE (e.g. GET _analyze/_search with a body) no longer break keep-alive connections.

Changed

  • Background optimize sweep interval default 300 → 20 s (cheap on idle indices; run the same guardrail config on every cluster node).
  • Known limits: knn_vector/geo/date can't be inferred by auto-create (use a mapping or template); query_string boost/slop are parsed but not yet ranked.

1.0.0September 3, 2026

1.0.0 — First release

A private, OpenSearch-compatible search + vector engine on an embedded high-performance vector engine — one small native binary per platform. Raw notes: RELEASE-NOTES-1.0.0.md.

  • OpenSearch-compatible REST (verified with opensearch-py and the Java client); Elasticsearch 8.x presentation mode; Qdrant/Pinecone/Chroma/Algolia wire-protocol listeners.
  • Vector search 1.8× faster than OpenSearch at equal recall on Graviton; fastest vector ingest of the five engines benchmarked; filtered kNN ahead of Qdrant.
  • Facets and sorts at or above OpenSearch speed; ~92 MB RSS after a 100k-doc query storm (vs ~1.4 GB for OpenSearch).
  • Admin console (/console): Dev Tools, Documents, Search tracing, zero-touch Migrate from OpenSearch/Elasticsearch/Qdrant/Pinecone/Chroma/Weaviate/Solr, and Evaluate (golden queries, side-by-side, regression gates).
  • Custom analyzers (stopwords, query-time synonyms, ngram, jieba), array and rank_features fields, neural search + rerank via the SearchAI Inference Server.
  • Scoped API keys, native TLS everywhere, full-replica clustering with op-log replication.