From Qdrant to the Index Server
September 2026 · 7 min read
Let's lead with the honest part: Qdrant is an excellent dedicated vector database, and in our own benchmark it is the fastest raw ANN query engine of the five — 2.40 ms p50 kNN at recall 1.0, and the fastest filtered kNN too. If raw vector latency is the only thing you need, Qdrant is a great answer and this article isn't trying to talk you out of it.
The reason teams move to the Index Server is consolidation: they don't just need nearest-neighbors, they need keyword precision, hybrid ranking, aggregations and facets over the same documents — and they'd rather run one engine on the familiar OpenSearch API than a vector DB plus a separate search system stitched together in the application.
Where a dedicated vector DB stops
Qdrant does dense + sparse vectors and payload filtering superbly. What it doesn't do server-side is the rest of a search stack:
| need | Qdrant | Index Server |
|---|---|---|
| Raw ANN latency | fastest here (2.40 ms p50) | 3.19 ms p50 (recall 1.0) |
| Ranked keyword (BM25) | sparse vectors — BM25 computed client-side | native BM25 full-text |
| Hybrid | fuses dense + sparse vectors | BM25 + vector, RRF-fused, one query clause |
| Aggregations / facets | — | terms/metric/histogram/range/nested/geo |
| API | Qdrant REST/gRPC | OpenSearch REST — and a Qdrant-compatible dialect |
| Ingest (docs/s) | 3,486 | 4,383 (with inline embedding) |
The practical version of "hybrid" matters here. Qdrant fuses dense and sparse vectors, which means real BM25 has to be computed by your application before the query ever reaches the database. On the index server, lexical BM25 and semantic kNN are fused server-side by RRF in a single clause — no client-side sparse encoding, no score calibration.
Keep your Qdrant client — the dialect is built in
You can evaluate it without touching your data layer. The server exposes a Qdrant-compatible dialect on its own listener, so existing upsert/query/scroll code points at the new host and runs:
# enable the Qdrant-dialect listener (or set dialects.qdrant-port in config)
POST /_console/dialects {"kind":"qdrant","port":6333}
# your existing client, new base URL
client.upsert(collection_name="docs", points=[...])
client.search(collection_name="docs", query_vector=[...], limit=10,
query_filter={"must":[{"key":"team","match":{"value":"support"}}]})
When you're ready to use the extras — BM25, hybrid, aggregations — the same documents are right there over the OpenSearch REST API. That's the whole point: one engine, two ways in, no second system to run.
Migrating the collection
- Console Migrate tab — connects to Qdrant, reads the collection
config (vector params →
knn_vectormapping), and streams points + payloads over scroll, with source-vs-target counts. - Re-embed from source text (recommended) — index the original text
with an
embedmapping and let the server embed inline; you get hybrid for free and never move the old vectors.
Honest scorecard
From the Graviton benchmarks (c8g.4xlarge, engines one at a time): Qdrant wins raw kNN (2.40 vs 3.19 ms p50) and filtered kNN (2.69 vs 3.03 ms) — a real, if small, edge. SearchAI wins vector ingest (4,383 vs 3,486 docs/s) and runs lighter here (527 vs 681 MB RSS), and it adds the entire full-text + hybrid + aggregation surface that a dedicated vector engine doesn't have. Choose Qdrant if the last millisecond of raw ANN is the product; choose the index server if "vector + keyword + hybrid + facets over one API" is.
One line installs the server and inference engine:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Start with the Getting Started guide or a build from the downloads page; the full numbers and trade-offs are on the benchmarks page.