← Blog · Home · Getting Started

From Qdrant to the Index Server

September 2026 · 7 min read

Let's lead with the honest part: Qdrant is an excellent dedicated vector database, and in our own benchmark it is the fastest raw ANN query engine of the five — 2.40 ms p50 kNN at recall 1.0, and the fastest filtered kNN too. If raw vector latency is the only thing you need, Qdrant is a great answer and this article isn't trying to talk you out of it.

The reason teams move to the Index Server is consolidation: they don't just need nearest-neighbors, they need keyword precision, hybrid ranking, aggregations and facets over the same documents — and they'd rather run one engine on the familiar OpenSearch API than a vector DB plus a separate search system stitched together in the application.

Where a dedicated vector DB stops

Qdrant does dense + sparse vectors and payload filtering superbly. What it doesn't do server-side is the rest of a search stack:

needQdrantIndex Server
Raw ANN latencyfastest here (2.40 ms p50)3.19 ms p50 (recall 1.0)
Ranked keyword (BM25)sparse vectors — BM25 computed client-sidenative BM25 full-text
Hybridfuses dense + sparse vectorsBM25 + vector, RRF-fused, one query clause
Aggregations / facets—terms/metric/histogram/range/nested/geo
APIQdrant REST/gRPCOpenSearch REST — and a Qdrant-compatible dialect
Ingest (docs/s)3,4864,383 (with inline embedding)

The practical version of "hybrid" matters here. Qdrant fuses dense and sparse vectors, which means real BM25 has to be computed by your application before the query ever reaches the database. On the index server, lexical BM25 and semantic kNN are fused server-side by RRF in a single clause — no client-side sparse encoding, no score calibration.

Keep your Qdrant client — the dialect is built in

You can evaluate it without touching your data layer. The server exposes a Qdrant-compatible dialect on its own listener, so existing upsert/query/scroll code points at the new host and runs:

# enable the Qdrant-dialect listener (or set dialects.qdrant-port in config)
POST /_console/dialects   {"kind":"qdrant","port":6333}

# your existing client, new base URL
client.upsert(collection_name="docs", points=[...])
client.search(collection_name="docs", query_vector=[...], limit=10,
              query_filter={"must":[{"key":"team","match":{"value":"support"}}]})

When you're ready to use the extras — BM25, hybrid, aggregations — the same documents are right there over the OpenSearch REST API. That's the whole point: one engine, two ways in, no second system to run.

Migrating the collection

Honest scorecard

From the Graviton benchmarks (c8g.4xlarge, engines one at a time): Qdrant wins raw kNN (2.40 vs 3.19 ms p50) and filtered kNN (2.69 vs 3.03 ms) — a real, if small, edge. SearchAI wins vector ingest (4,383 vs 3,486 docs/s) and runs lighter here (527 vs 681 MB RSS), and it adds the entire full-text + hybrid + aggregation surface that a dedicated vector engine doesn't have. Choose Qdrant if the last millisecond of raw ANN is the product; choose the index server if "vector + keyword + hybrid + facets over one API" is.

One line installs the server and inference engine:

curl -fsSL https://index-server.searchblox.com/install | sudo bash

Start with the Getting Started guide or a build from the downloads page; the full numbers and trade-offs are on the benchmarks page.