From Chroma to the Index Server
September 2026 · 8 min read
Chroma is one of the fastest ways to stand up a RAG prototype: a
pip install, an embedded collection, a few add and
query calls, and you're retrieving. That simplicity is exactly
why it's everywhere in notebooks and demos. The questions arrive when the
prototype has to become a product — and they're the same questions every
time: How do I add keyword precision to semantic recall? How fast are
filtered queries at real selectivity? Where do auth, TLS, backups and
failover come from? What happens when one process isn't enough?
The Index Server is built to answer those without re-platforming. It keeps your Chroma client working through a compatible dialect, and it adds the production layer a search engine is supposed to have: server-side BM25 + hybrid search, fast filtered vectors, aggregations, API-key auth, TLS, snapshots, a Prometheus endpoint, and full-replica or Raft clustering.
What changes when RAG goes to production
| need | Chroma | Index Server |
|---|---|---|
| Keyword + semantic in one query | vector similarity only — no ranked BM25 fusion | server-side hybrid (BM25 + vector, RRF) as one query clause |
| Filtered vector search at real selectivity | post-filtering; ~33 ms p50 at 50% selectivity on our bench | 3 ms p50 (filter pushdown into the index) |
| Recall vs cost | needs ef≈1024 to reach ~0.99 recall@10 | recall 1.0 at a low ef |
| Security | none built in (front it yourself) | API-key auth (admin + scoped keys), in-process TLS |
| Durability / backups | local persistence; no snapshot API | filesystem & S3 _snapshot / restore |
| High availability / scale-out | single process | full-replica reads or Raft quorum HA with failover |
| Beyond vectors | vector store | full-text, aggregations, facets, geo — a complete search server |
The one that usually decides it: hybrid search
Pure vector retrieval is what makes RAG feel magic — and also what feeds the classic failure mode: it returns plausible-but-wrong context because it matched vibe, not tokens. Production RAG wants both legs:
- Exact identifiers inside natural questions — "error E4021 when
exporting invoices". The embedding blurs
E4021; BM25 nails it while the vector leg ranks the conceptually-similar docs around it. Same for SKUs, part numbers, API names, legal citations. - Grounding — the keyword leg anchors retrieval to documents that literally contain the entities in the question, so the model is less likely to be handed convincing-but-irrelevant passages.
- Out-of-vocabulary terms — new product names, internal codenames, fresh jargon the embedding model has never seen but the inverted index matches trivially.
Chroma has no ranked keyword fusion, so hybrid means bolting on a second system and fusing client-side. Here it's one clause:
POST /docs/_search
{ "query": { "hybrid": { "queries": [
{ "match": { "body": "export invoices E4021" } },
{ "neural": { "embedding": { "query_text": "invoice export error", "k": 20 } } }
] } } }
RRF fuses the two rank lists with no score calibration between BM25 and
cosine scales. And because the vector leg can be a neural
clause, the server embeds the query text itself — your app sends one
plain-text hybrid query and never calls an embedding API.
Embed at ingest — no separate pipeline
In a Chroma stack you typically embed documents yourself (call an
embedding API, then add the vectors). Here you map a
knn_vector field with an embed block once and just
index text; the bundled inference engine embeds inline and writes the vector
into the document, so replicas and snapshots never re-embed:
PUT /docs
{ "mappings": { "properties": {
"body": { "type": "text" },
"embedding": { "type": "knn_vector", "dimension": 1024,
"embed": { "source_field": "body", "model": "your-embed-model" } }
}}}
POST /docs/_bulk
{"index":{"_id":"1"}}
{"body":"annual leave carries over up to five days"}
Keep your Chroma client — the dialect is built in
You don't have to rewrite your retrieval layer to evaluate it. The server
exposes a Chroma-compatible dialect on its own listener, so existing
collection add/query/get code points at
the new host and runs:
# enable the Chroma-dialect listener (or set dialects.chroma-port in config)
POST /_console/dialects {"kind":"chroma","port":8000}
# your existing client, new base URL
collection.add(ids=["1"], documents=["..."], embeddings=[[...]])
collection.query(query_embeddings=[[...]], n_results=10, where={"team":"support"})
When you're ready, the same documents are reachable over the OpenSearch REST API — with full-text, filters, aggregations and hybrid — so you can consolidate two systems (a vector store plus whatever you bolted on for keyword search) into one.
Migrating what you already have
Two hands-off paths:
- Console Migrate tab — connects to your Chroma server, lists the collections, and streams vectors + metadata over; it records source-vs-target counts so completeness is part of the job. (It also imports from Qdrant, Pinecone, Weaviate and OpenSearch.)
- Re-embed from source text (recommended when you can) — index the original text and let the server embed. You drop the old vectors, never pay to move them, and get a clean mapping with hybrid enabled.
What the numbers look like
On our published Graviton benchmarks (20k × 384-dim, engines one at a time on a c8g.4xlarge):
- Filtered kNN: ~3 ms p50 vs Chroma's ~33 ms — filter pushdown into the index instead of post-filtering.
- Recall: 1.0 at a low ef, where Chroma needs ef≈1024 to reach ~0.99.
- Vector ingest: ≈4,380 docs/s (with inline embedding) vs Chroma's ≈2,260 — and it's the fastest ingest of the five engines tested.
- Footprint: a full search server (BM25, aggregations, snapshots, clustering) in ~527 MB RSS.
The honest trade-off: Chroma is lighter to start — an in-process library with almost no ceremony, which is genuinely the right tool for a notebook. This is the engine you move to when that prototype needs keyword precision, filtered speed, and the operational surface of a real server.
Try it in ten minutes
One line on a Linux host installs the server and the inference engine as a service:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Set server.api-key, open http://<host>:9200/console,
map one embed field, bulk-load a sample of your text, and run the
same top-k queries you run against Chroma — then add a keyword leg and watch
the grounding improve.
Start with the Getting Started guide or grab a build from the downloads page. The full comparison and honest edges are on the benchmarks page.