← Blog · Home · Docs

One engine, four ways to search

September 2026 · 8 min read

"Search" isn't one thing. Keyword (BM25), semantic (vector), hybrid (BM25 + vector, RRF-fused) and reranking (a cross-encoder re-scoring the top results) are four different retrieval strategies with genuinely different strengths — and the wrong default quietly costs you relevance or latency. The useful part: the Index Server runs all four over the same index and the same OpenSearch REST API, so choosing isn't a re-platforming decision — it's a query.

The admin console even has a Compare search strategies tab that runs one query as all four side by side, with per-strategy latency and the overlap between them — so you can see, on your data, where they agree and where they don't. Everything below comes from running that on a 20,000-product e-commerce catalog.

The four strategies, and what each is actually for

1 · Keyword (BM25) — exact words, instantly

Classic inverted-index scoring. Unbeatable when the query is the vocabulary: SKUs, part numbers, error codes, product names, legal citations — anything where the exact token matters. It's also the cheapest by a wide margin (single-digit milliseconds).

Where it fails: it can't match what isn't literally there. A typo (drees for "dress") returns nothing. And it's easily fooled by term-matching without intent — on our catalog, "elegant dress for a summer party" returned perfumes, because BM25 latched onto "summer" in "Blue Spirit Summer EDT."

POST /products/_search
{ "query": { "multi_match": { "query": "linen shirt",
    "fields": ["product_name^2", "summary", "description"] } } }

2 · Semantic (neural) — meaning via embeddings

The query is embedded into a vector and matched by nearest-neighbor, so it retrieves by meaning. Robust to typos, synonyms and natural-language intent: drees still finds dresses, "sneakers" finds trainers, and "elegant dress for a summer party" returns actual dresses. Send text, not a vector — the server embeds the query with the same model it used at ingest.

Where it costs: a query-time embedding round-trip (tens to a couple hundred milliseconds, model-dependent), and it can rank a vibe-similar document above one that literally contains the term you typed.

POST /products/_search
{ "query": { "neural": { "vec": {
    "query_text": "something elegant for a summer party", "k": 20 } } } }

3 · Hybrid (RRF) — the safe default

Run both legs and fuse their rank lists with Reciprocal Rank Fusion — no score calibration between BM25 and cosine scales. You get semantic recall grounded by keyword precision: the exact-term matches anchor the results, the vector leg fills in the conceptually-related ones. For most product search, site search and RAG retrieval, this is the right default.

POST /products/_search
{ "query": { "hybrid": { "queries": [
    { "multi_match": { "query": "elegant summer party dress",
        "fields": ["product_name^2", "summary"] } },
    { "neural": { "vec": { "query_text": "elegant summer party dress", "k": 20 } } }
] } } }

4 · Reranking — top-k precision for RAG

Retrieve a wider window (say top-50) with any of the above, then have a cross-encoder re-score each candidate against the query and reorder. Because it reads the query and document together, it's the most accurate at getting the top 3–5 exactly right — which is precisely what a RAG pipeline feeds an LLM. It's also the most expensive: the reranker scores every candidate, so latency scales with the window.

POST /products/_search
{ "query": { "hybrid": { "queries": [ /* … */ ] } },
  "rerank": { "model": "your-reranker", "query": "elegant summer party dress",
              "field": "product_name", "window": 20 } }

The numbers: latency vs. quality

One query — "short blue skirt" — across all four on the 20k catalog (dev box, CPU inference). Overlap is versus the BM25 result set:

strategyp50 latencyoverlap w/ BM25best for
Keyword (BM25)~36 ms—exact terms, SKUs, codes, known vocabulary
Semantic~280 ms~10%typos, synonyms, natural-language intent
Hybrid (RRF)~262 ms~50%the default — recall + precision together
+ Rerank~3,500 ms100% (re-ordered)RAG / answer quality, where top-3 must be right

The 10% overlap between keyword and semantic is the whole story: they surface largely different documents, which is exactly why fusing them helps — and why picking only one leaves relevance on the table. Reranking the hybrid set here cost ~3.5 s because it's a large cross-encoder scoring 50 candidates on a CPU; in production you shrink the window (e.g. 20), use a smaller/faster reranker, or put inference on a GPU. The keyword, semantic and hybrid latencies are dominated by the query-embedding call, not the index — the engine's own search is single-digit milliseconds.

How to choose

Don't guess — the console's Compare tab (and the Evaluate harness for golden query sets) let you A/B these on your own corpus with overlap and latency in front of you.

Why one engine matters

The reason all four are one query away is that they share one index. Text is stored once with a BM25 inverted index; vectors are embedded at ingest (the bundled inference engine — no separate embedding pipeline, and the query is embedded at the same width, a 1024-dim default that halves storage and search cost versus a full-width model); RRF fusion and the reranker run server-side. No second datastore to keep in sync, no client-side score juggling, and the same OpenSearch API the rest of your stack already speaks.

One line installs the server and inference engine:

curl -fsSL https://index-server.searchblox.com/install | sudo bash

Then map one embed field, load your text, and open /console → Search to compare strategies on your own data. Start with the Getting Started guide; the query-DSL details are in the docs.