One engine, four ways to search
September 2026 · 8 min read
"Search" isn't one thing. Keyword (BM25), semantic (vector), hybrid (BM25 + vector, RRF-fused) and reranking (a cross-encoder re-scoring the top results) are four different retrieval strategies with genuinely different strengths — and the wrong default quietly costs you relevance or latency. The useful part: the Index Server runs all four over the same index and the same OpenSearch REST API, so choosing isn't a re-platforming decision — it's a query.
The admin console even has a Compare search strategies tab that runs one query as all four side by side, with per-strategy latency and the overlap between them — so you can see, on your data, where they agree and where they don't. Everything below comes from running that on a 20,000-product e-commerce catalog.
The four strategies, and what each is actually for
1 · Keyword (BM25) — exact words, instantly
Classic inverted-index scoring. Unbeatable when the query is the vocabulary: SKUs, part numbers, error codes, product names, legal citations — anything where the exact token matters. It's also the cheapest by a wide margin (single-digit milliseconds).
Where it fails: it can't match what isn't literally there. A typo
(drees for "dress") returns nothing. And it's easily fooled by
term-matching without intent — on our catalog, "elegant dress for a summer
party" returned perfumes, because BM25 latched onto "summer" in
"Blue Spirit Summer EDT."
POST /products/_search
{ "query": { "multi_match": { "query": "linen shirt",
"fields": ["product_name^2", "summary", "description"] } } }
2 · Semantic (neural) — meaning via embeddings
The query is embedded into a vector and matched by nearest-neighbor, so it
retrieves by meaning. Robust to typos, synonyms and natural-language
intent: drees still finds dresses, "sneakers" finds trainers, and
"elegant dress for a summer party" returns actual dresses. Send text, not a
vector — the server embeds the query with the same model it used at ingest.
Where it costs: a query-time embedding round-trip (tens to a couple hundred milliseconds, model-dependent), and it can rank a vibe-similar document above one that literally contains the term you typed.
POST /products/_search
{ "query": { "neural": { "vec": {
"query_text": "something elegant for a summer party", "k": 20 } } } }
3 · Hybrid (RRF) — the safe default
Run both legs and fuse their rank lists with Reciprocal Rank Fusion — no score calibration between BM25 and cosine scales. You get semantic recall grounded by keyword precision: the exact-term matches anchor the results, the vector leg fills in the conceptually-related ones. For most product search, site search and RAG retrieval, this is the right default.
POST /products/_search
{ "query": { "hybrid": { "queries": [
{ "multi_match": { "query": "elegant summer party dress",
"fields": ["product_name^2", "summary"] } },
{ "neural": { "vec": { "query_text": "elegant summer party dress", "k": 20 } } }
] } } }
4 · Reranking — top-k precision for RAG
Retrieve a wider window (say top-50) with any of the above, then have a cross-encoder re-score each candidate against the query and reorder. Because it reads the query and document together, it's the most accurate at getting the top 3–5 exactly right — which is precisely what a RAG pipeline feeds an LLM. It's also the most expensive: the reranker scores every candidate, so latency scales with the window.
POST /products/_search
{ "query": { "hybrid": { "queries": [ /* … */ ] } },
"rerank": { "model": "your-reranker", "query": "elegant summer party dress",
"field": "product_name", "window": 20 } }
The numbers: latency vs. quality
One query — "short blue skirt" — across all four on the 20k catalog (dev box, CPU inference). Overlap is versus the BM25 result set:
| strategy | p50 latency | overlap w/ BM25 | best for |
|---|---|---|---|
| Keyword (BM25) | ~36 ms | — | exact terms, SKUs, codes, known vocabulary |
| Semantic | ~280 ms | ~10% | typos, synonyms, natural-language intent |
| Hybrid (RRF) | ~262 ms | ~50% | the default — recall + precision together |
| + Rerank | ~3,500 ms | 100% (re-ordered) | RAG / answer quality, where top-3 must be right |
The 10% overlap between keyword and semantic is the whole story: they surface largely different documents, which is exactly why fusing them helps — and why picking only one leaves relevance on the table. Reranking the hybrid set here cost ~3.5 s because it's a large cross-encoder scoring 50 candidates on a CPU; in production you shrink the window (e.g. 20), use a smaller/faster reranker, or put inference on a GPU. The keyword, semantic and hybrid latencies are dominated by the query-embedding call, not the index — the engine's own search is single-digit milliseconds.
How to choose
- Autocomplete, code/SKU lookup, "I know the term": keyword. It's 6× faster and precise when the vocabulary is fixed.
- Q&A, discovery, messy human queries: hybrid. It absorbs typos and intent without dropping exact-match precision — the safe default.
- RAG and anywhere the top 3 feed a model or a human decision: hybrid + rerank. Pay the reranker only for the last mile that matters.
- Pure recall research over a clean vocabulary: semantic alone can be enough — but measure it against hybrid before committing.
Don't guess — the console's Compare tab (and the Evaluate harness for golden query sets) let you A/B these on your own corpus with overlap and latency in front of you.
Why one engine matters
The reason all four are one query away is that they share one index. Text is stored once with a BM25 inverted index; vectors are embedded at ingest (the bundled inference engine — no separate embedding pipeline, and the query is embedded at the same width, a 1024-dim default that halves storage and search cost versus a full-width model); RRF fusion and the reranker run server-side. No second datastore to keep in sync, no client-side score juggling, and the same OpenSearch API the rest of your stack already speaks.
One line installs the server and inference engine:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Then map one embed field, load your text, and open
/console → Search to compare strategies on your own data.
Start with the Getting Started guide;
the query-DSL details are in the docs.