# SearchAI Index Server 1.3.0 (2026-09-23)

Agent-ready release: drive the server from Claude and other AI agents over the
Model Context Protocol, and watch node health live from the admin console. A
drop-in upgrade over 1.2.0 — no configuration or on-disk format changes, and
every new surface is opt-in.

## Platforms (all built natively)
- searchai-index-server-1.3.0-linux-arm64.tar.gz (Graviton/ARM servers)
- searchai-index-server-1.3.0-linux-amd64.tar.gz (x86_64 glibc servers)
- searchai-index-server-1.3.0-osx-arm64.tar.gz (Apple Silicon dev)

## MCP for AI agents
- **New MCP listener** presents the server to Claude and other agents as a set
  of tools — create indices, load documents, and run full-text, vector, and
  hybrid search directly, with no glue code. One engine and dataset, an extra
  protocol on its own port.
- **Streamable HTTP** at `POST /mcp`, authenticated with the same API key (or a
  scoped key) as every other listener. The caller's principal is propagated
  into each operation, so a scoped key keeps its per-index limits through the
  agent.
- **Thirteen tools**: `list_indices`, `cluster_health`, `server_stats`,
  `get_mapping`, `create_index`, `delete_index`, `index_document`,
  `bulk_index`, `get_document`, `delete_document`, `search` (full query-DSL
  passthrough or convenience fields), `count`, and `hybrid_search` (BM25 +
  vector, RRF-fused). Write tools refresh by default for read-after-write.
- **Opt-in**: enable with `dialects.mcp-port=<port>` in `server.properties`, or
  the console toggle, or `POST /_console/dialects {"mcp": <port>}`. Off by
  default — existing deployments are unaffected until they opt in.

Connect Claude Code:
```
claude mcp add searchai --transport http http://<host>:<port>/mcp \
  --header "Authorization: Bearer <api-key>"
```

## Admin console
- **Live resource panel**: CPU load average (1/5/15 min) and process CPU usage,
  logical core count, resident memory, disk free/total, and uptime, refreshing
  on an interval so a node's health is visible at a glance.
- **MCP dialect toggle**: enable or disable the MCP listener from the console
  alongside the other dialects, with the endpoint URL and the one-line command
  to connect Claude Code shown inline.

## Embeddings: Matryoshka dimensions + a 1024 default
- A `knn_vector` field's declared `dimension` is now passed to the inference
  server (ingest and neural query alike), so a Matryoshka-capable embed model
  produces vectors natively at that width — smaller storage and faster search,
  with the query embedded at the same width as the index.
- An auto-embed field may omit `dimension`; it defaults to **1024** — near-full
  quality at ~half the storage/search cost of a full-width model.

## Reliability fixes
- **HNSW `ef_search` clamp:** the dimension-scaled default could exceed the
  engine's ef_search limit at high dimensions (e.g. a 2048-wide model), failing
  every vector query; ef_search is now clamped to the engine maximum on all
  paths.
- **Open-file limit:** the server raises its own `RLIMIT_NOFILE` at startup so
  the per-column engine isn't starved by a launcher's low default (which
  surfaced as "Too many open files" and failed full-text materialization).

## New in `/_console/stats` and `/_metrics`
- `cpu_cores`, `load_avg` (1/5/15-minute), and cumulative process `cpu_seconds`
  are now reported by `/_console/stats`.
- `/_metrics` gains `searchai_load1` and `searchai_cpu_seconds_total` gauges.

## Upgrade notes
Drop-in over 1.2.0 — same configuration, same on-disk format, same wire
compatibility. No action required beyond replacing the binary/bundle. In a
cluster, upgrade nodes one at a time as usual. The MCP listener remains disabled
until a port is configured.

## Verification
Each bundle was built and smoke-tested natively on its platform with the
patched embedded engine, including the new MCP endpoint (handshake, tools/list,
and tool calls) and the console resource panel. Full unit/probe suite plus the
complete pytest suite (cluster, raft failover/membership/snapshot/chaos,
crash-recovery, TLS, security, dialects via real client libraries, bulk,
aggregations, scroll, snapshots), plus live auto-embedding and neural-query
validation against the inference server.
