Index straight from Claude
September 2026 · 7 min read
Most RAG problems aren't retrieval problems — they're corpus problems. The wrong documents got indexed, duplicates drowned the good ones, or nobody ever measured whether a source actually improved answers. The Model Context Protocol listener changes who does that work: Claude (or any MCP-capable agent) indexes content directly into the Index Server as tools, so the agent that understands your documents is the same one curating and testing the corpus — with no glue code between them.
The agent has the tools already
Point Claude Code at the server's MCP endpoint and it gets a working
toolset — create_index, index_document,
bulk_index, search, hybrid_search,
get_mapping, delete_document and more — over the same
API and auth as everything else:
claude mcp add searchai --transport http http://<host>:9400/mcp \
--header "Authorization: Bearer <api-key>"
Now "read these docs and index the useful parts" is a single agent turn. Vectors are embedded at ingest by the bundled inference engine (a 1024-dim default), so the agent sends text — not embeddings — and writes refresh by default, so the next search sees them immediately.
Curation happens at ingest, not after
Because the agent is in the loop as it writes, it can do the janitorial work that usually never happens:
- Normalize and chunk — split long documents on real boundaries, strip boilerplate, and index clean, self-contained passages instead of raw dumps.
- Tag and enrich — add
source,section,date,doc_typefields as keywords so retrieval can filter by them later. The agent knows what each document is; capture that at write time. - Dedupe and skip — check with a quick
searchbefore indexing, and drop near-duplicates or low-value pages rather than letting them dilute recall. - Keep provenance — store the URL / file path / author alongside the text so every retrieved chunk can cite where it came from.
{ "name": "bulk_index", "arguments": {
"index": "kb",
"documents": [
{ "id": "handbook#leave", "doc": {
"body": "Annual leave carries over up to five days …",
"source": "employee-handbook", "section": "leave",
"doc_type": "policy", "url": "https://…/handbook#leave" } }
] } }
Pick sources by measuring, not guessing
The part teams usually skip: proving a source helps. Because the same index answers keyword, semantic, hybrid and reranked queries (see One engine, four ways to search), the agent can run your real questions against the freshly-indexed corpus and read the results back:
- Does this source ever win? Run the golden questions with
hybrid_search; if a source never appears in the top results, it's cost without benefit — drop it. - Is it adding noise? Index a candidate source, re-run the queries, and see whether good answers get pushed down. The agent can compare before/after and decide.
- Which retrieval mode fits? Exact-term catalogs favor keyword; natural-language knowledge bases favor hybrid; answer-critical RAG adds a rerank pass on the top window. Test all four on your data — the console's Compare search strategies tab and the Evaluate harness score overlap and relevance for you.
The loop is tight: bulk_index a candidate source →
hybrid_search the golden set → judge → keep, re-tag, or
delete_document. Corpus curation becomes an agent workflow you can
re-run whenever the sources change, instead of a one-time guess.
Why one engine matters here
Indexing, keyword search, vector search, hybrid fusion and reranking are all the same server over one dataset and one API. The agent never juggles a vector store plus a search system plus an embedding service — it writes text, searches it four ways, and curates in place. Scoped API keys mean an agent can be limited to exactly the indexes it should touch, so "let the agent manage the corpus" is safe to actually do.
Enable the listener in server.properties (or the console toggle)
and connect an agent:
dialects.mcp-port=9400
One line installs the server and inference engine:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Full tool reference is in the MCP docs; start with the Getting Started guide.