← Blog · Home · Getting Started

From Chroma to the Index Server

September 2026 · 8 min read

Chroma is one of the fastest ways to stand up a RAG prototype: a pip install, an embedded collection, a few add and query calls, and you're retrieving. That simplicity is exactly why it's everywhere in notebooks and demos. The questions arrive when the prototype has to become a product — and they're the same questions every time: How do I add keyword precision to semantic recall? How fast are filtered queries at real selectivity? Where do auth, TLS, backups and failover come from? What happens when one process isn't enough?

The Index Server is built to answer those without re-platforming. It keeps your Chroma client working through a compatible dialect, and it adds the production layer a search engine is supposed to have: server-side BM25 + hybrid search, fast filtered vectors, aggregations, API-key auth, TLS, snapshots, a Prometheus endpoint, and full-replica or Raft clustering.

What changes when RAG goes to production

needChromaIndex Server
Keyword + semantic in one queryvector similarity only — no ranked BM25 fusionserver-side hybrid (BM25 + vector, RRF) as one query clause
Filtered vector search at real selectivitypost-filtering; ~33 ms p50 at 50% selectivity on our bench3 ms p50 (filter pushdown into the index)
Recall vs costneeds ef≈1024 to reach ~0.99 recall@10recall 1.0 at a low ef
Securitynone built in (front it yourself)API-key auth (admin + scoped keys), in-process TLS
Durability / backupslocal persistence; no snapshot APIfilesystem & S3 _snapshot / restore
High availability / scale-outsingle processfull-replica reads or Raft quorum HA with failover
Beyond vectorsvector storefull-text, aggregations, facets, geo — a complete search server

The one that usually decides it: hybrid search

Pure vector retrieval is what makes RAG feel magic — and also what feeds the classic failure mode: it returns plausible-but-wrong context because it matched vibe, not tokens. Production RAG wants both legs:

Chroma has no ranked keyword fusion, so hybrid means bolting on a second system and fusing client-side. Here it's one clause:

POST /docs/_search
{ "query": { "hybrid": { "queries": [
    { "match": { "body": "export invoices E4021" } },
    { "neural": { "embedding": { "query_text": "invoice export error", "k": 20 } } }
] } } }

RRF fuses the two rank lists with no score calibration between BM25 and cosine scales. And because the vector leg can be a neural clause, the server embeds the query text itself — your app sends one plain-text hybrid query and never calls an embedding API.

Embed at ingest — no separate pipeline

In a Chroma stack you typically embed documents yourself (call an embedding API, then add the vectors). Here you map a knn_vector field with an embed block once and just index text; the bundled inference engine embeds inline and writes the vector into the document, so replicas and snapshots never re-embed:

PUT /docs
{ "mappings": { "properties": {
  "body": { "type": "text" },
  "embedding": { "type": "knn_vector", "dimension": 1024,
                 "embed": { "source_field": "body", "model": "your-embed-model" } }
}}}

POST /docs/_bulk
{"index":{"_id":"1"}}
{"body":"annual leave carries over up to five days"}

Keep your Chroma client — the dialect is built in

You don't have to rewrite your retrieval layer to evaluate it. The server exposes a Chroma-compatible dialect on its own listener, so existing collection add/query/get code points at the new host and runs:

# enable the Chroma-dialect listener (or set dialects.chroma-port in config)
POST /_console/dialects   {"kind":"chroma","port":8000}

# your existing client, new base URL
collection.add(ids=["1"], documents=["..."], embeddings=[[...]])
collection.query(query_embeddings=[[...]], n_results=10, where={"team":"support"})

When you're ready, the same documents are reachable over the OpenSearch REST API — with full-text, filters, aggregations and hybrid — so you can consolidate two systems (a vector store plus whatever you bolted on for keyword search) into one.

Migrating what you already have

Two hands-off paths:

What the numbers look like

On our published Graviton benchmarks (20k × 384-dim, engines one at a time on a c8g.4xlarge):

The honest trade-off: Chroma is lighter to start — an in-process library with almost no ceremony, which is genuinely the right tool for a notebook. This is the engine you move to when that prototype needs keyword precision, filtered speed, and the operational surface of a real server.

Try it in ten minutes

One line on a Linux host installs the server and the inference engine as a service:

curl -fsSL https://index-server.searchblox.com/install | sudo bash

Set server.api-key, open http://<host>:9200/console, map one embed field, bulk-load a sample of your text, and run the same top-k queries you run against Chroma — then add a keyword leg and watch the grounding improve.

Start with the Getting Started guide or grab a build from the downloads page. The full comparison and honest edges are on the benchmarks page.