Benchmarks on a Graviton server
Index Server 1.3.0 vs OpenSearch 3.8.0, Qdrant 1.19, Weaviate and Chroma 1.5.9 — ingestion and search rates measured on one AWS c8g.4xlarge (Graviton4, 16 vCPU, 32 GB), engines run one at a time on the same box, September 2026. Every ANN index was verified actually built before timing, and recall is measured against exact brute-force ground truth.
Vector search — 20k × 384-dim, cosine, HNSW m=24 / ef_construction=128
Latency is p50 over 200 iterations after 20 warmups; recall@10 vs exact search. Each engine's search-time ef runs at its default or its calibrated equal-recall setting.
| metric | SearchAI 1.3.0 | OpenSearch 3.8.0 | Qdrant 1.19 | Weaviate | Chroma 1.5.9 |
|---|---|---|---|---|---|
| vector ingest (docs/s) | 4,383 | 1,669 | 3,486 | 2,063 | 2,259 |
| kNN k=10 p50 (ms) | 3.19 | 3.28 | 2.40 | 8.30 | 8.18 |
| kNN k=10 p95 (ms) | 3.24 | 3.83 | 5.23 | 8.67 | 8.59 |
| filtered kNN p50 (ms) | 3.03 | 4.54 | 2.69 | 4.86 | 32.80 |
| recall@10 | 1.00 | 0.73 | 1.00 | 0.99 | 0.99 |
| RSS after workload (MB) | 527* | 5,085 | 681 | 282 | 507 |
*SearchAI's RSS is after the full workload — the 100k-doc full-text phase plus vectors — and includes a complete search server (BM25, aggregations, snapshots, clustering), not a vector-only store. OpenSearch ran with a 4 GB heap. Chroma's ~33 ms filtered kNN is its documented post-filtering behavior at this selectivity.
Read: SearchAI has the fastest vector ingest of the five engines (2.6× OpenSearch, 1.3× Qdrant) and groups with Qdrant and OpenSearch in the 3 ms query tier — with the tightest p95 of the group — while Weaviate and Chroma sit 2–3× behind. Qdrant is the fastest raw ANN query engine here; it is also vector-only, speaks its own API, and (like Weaviate and Chroma) has no BM25/text leg in this comparison because it doesn't do full-text search of this kind.
Hybrid search — keyword BM25 + vector, fused with RRF
The query that AI applications actually run: a lexical leg and a semantic leg over the same documents, fused by reciprocal-rank fusion. Corpus: 20k docs each carrying a text body and a 384-dim vector; k=10; 200 iterations after 20 warmups. Only engines with server-side BM25 + fusion can run this leg at all.
| metric | SearchAI 1.3.0 | OpenSearch 3.8.0 | Weaviate | Qdrant | Chroma |
|---|---|---|---|---|---|
| text+vector ingest (docs/s) | 4,310 | 1,582 | 2,175 | no server-side BM25 — hybrid needs client-side sparse encoding | no ranked keyword fusion |
| hybrid k=10 p50 (ms) | 3.70 | 7.86 | 7.58 | ||
| hybrid k=10 p95 (ms) | 3.88 | 9.41 | 8.15 | ||
| RRF fusion, how invoked | one query clause | search pipeline you create + reference | GraphQL fusionType |
Read: SearchAI answers hybrid queries ~2× faster than both
engines that can run them (3.70 vs 7.86/7.58 ms) — hybrid costs it only
+0.5 ms over its pure-kNN latency — and ingests the combined text+vector
documents 2–2.7× faster. The invocation difference matters in practice:
in SearchAI, RRF hybrid is a plain query-DSL clause
({"hybrid":{"queries":[{"match":…},{"knn":…}]}}); OpenSearch requires
creating a search pipeline with a score-ranker processor and referencing it on every
request; Weaviate requires its GraphQL dialect. Qdrant fuses dense+sparse
vectors — BM25 must be computed client-side before it ever sees the query —
and Chroma has no ranked keyword fusion, which is why running one engine instead of
two is the actual differentiator here.
Why hybrid search — the cases where one leg fails
- Exact identifiers inside natural questions. "error E4021 when
exporting invoices" — embeddings blur
E4021into generic error-space; BM25 nails the token while the vector leg ranks the conceptually-similar docs around it. Same story for SKUs, part numbers, API names, and legal citations. - RAG grounding. Pure-vector retrieval happily returns plausible-but-wrong context (the classic hallucination feeder). The keyword leg anchors retrieval to documents that literally contain the entities in the question; the vector leg still catches paraphrase and synonyms. RRF needs no score calibration between BM25 and cosine scales — rank positions fuse directly.
- Vocabulary the model has never seen. New product names, internal project codenames, fresh jargon — out-of-vocabulary for the embedding model, trivial for the inverted index. Hybrid degrades gracefully to the leg that knows.
- E-commerce and catalog search. "nike pegasus 40 trail" needs brand and model matched exactly and "similar running shoes" ranked semantically beneath the exact hits.
- Compliance / e-discovery. Must-match phrases and clause language (lexical, auditable) combined with conceptual recall of paraphrased variants.
With the SearchAI Inference Server bundled by the installer, the vector leg can
be a neural clause — the server embeds the query text itself, so the
application sends one plain-text hybrid query and never touches an embedding.
Full-text — 100k docs, single bulk client (SearchAI vs OpenSearch)
Only the two OpenSearch-API engines can run
this phase. SearchAI configured with the recommended
index.track-total-hits=1000 cap (same knob Elasticsearch defaults to);
OpenSearch 3.8.0 single-node, security off, 4 GB heap.
| metric | SearchAI 1.3.0 | OpenSearch 3.8.0 |
|---|---|---|
| text bulk ingest (docs/s) | 20,385 | 17,332 |
| realistic-doc ingest, ~4.5 KB bodies (docs/s · MB/s) | 3,308 · 15.1 | 2,299 · 10.5 |
| term / match / bool / range p50 (ms) | 3.0 – 3.8 | 1.2 – 1.8 |
| terms_agg p50 (ms) | 0.87 | 1.03 |
| sort-by-field p50 (ms) | 4.3 | 1.2 |
| match_phrase p50 (ms) | 103.3 | 4.3 |
| RSS after workload (MB) | 527 | 5,085 |
Read: SearchAI ingests text 18% faster at this scale — and
44% faster on realistic ~4.5 KB documents — in roughly
1/10th of the memory. It now edges OpenSearch on terms_agg
(0.87 vs 1.03 ms — the columnar aggregation path landed in 1.2.0), while
OpenSearch answers most other text queries 1.5–3 ms faster on this 16-core box;
both engines are comfortably in interactive range. The honest exception is ours:
phrase queries remain an OpenSearch strength (a tracked upstream-engine work item).
With default exact hit-counting instead of the capped setting, SearchAI's filter/text
latencies rise to 30–100 ms on large matched sets — use the cap.
Workload, server & memory
Corpus sizes
| phase | docs | per-document | corpus |
|---|---|---|---|
| full-text | 100,000 | ~270 B — 40-word body + title, keyword, numeric, date fields | ~26 MB |
| realistic documents | 20,000 | ~4.7 KB — 700-word bodies (crawled-page shape) | ~90 MB |
| vectors | 20,000 | 384-dim float32 (1,536 B raw) + keyword tag | ~31 MB raw vectors |
| hybrid | 20,000 | 30-word body (~200 B) + 384-dim vector | ~35 MB |
Deliberately a small-to-mid corpus: it fits every engine's comfortable zone, so the comparison measures engine efficiency rather than who hits a memory wall first. All ingest is a single bulk client, chunk size 500.
Server
One AWS c8g.4xlarge — Graviton4 (arm64), 16 vCPU, 32 GB RAM (~30 GiB usable), gp3 EBS, Amazon Linux 2023, us-east-1. Engines ran strictly one at a time on this box with wiped data directories; clients ran on the same host over loopback (no network variance, no auth on any engine).
Memory — observed and what you should plan for
| engine | RSS after its workload | configured |
|---|---|---|
| Index Server | 323 MB | defaults (no memory limit set) |
| OpenSearch 3.8.0 | 5,095 MB | 4 GB JVM heap (-Xms4g/-Xmx4g), security off |
| Qdrant 1.19.0 | 537 MB | defaults |
| Chroma 1.5.9 | 472 MB | defaults |
| Weaviate | 281 MB | defaults, anonymous access |
SearchAI's 323 MB is after the entire suite — 100k text docs + 90 MB realistic docs + 20k vectors + query storms — on a full search server. Practical sizing: this workload class runs comfortably on a 2 GB instance for SearchAI; OpenSearch's own minimum for the same data is effectively an 8 GB box once heap + off-heap are provisioned. One honest caveat for wide mappings: SearchAI's full-text memory scales with the number of distinct text columns queried (engine-level cost, ~113 MB per materialized FTS column), so schema width — not document count — is what to watch on small boxes.
How the vectors ("embeddings") were produced
The vectors are synthetic: seeded, L2-normalized 384-dimensional random
vectors (NumPy RandomState(7)), identical bytes for every engine,
cosine distance. No embedding model ran during any timed phase — ingest and query
numbers measure the engines, not a model, and random vectors are the
worst case for HNSW (no cluster structure to exploit), so real embeddings
should perform at least this well. 384 dimensions was chosen to match the
small-embedder class used in production RAG (e.g. bge-small-en-v1.5,
all-MiniLM-L6-v2). In a real deployment the bundled SearchAI
Inference Server produces the embeddings server-side (embed mapping
block at index time, neural clause at query time) — model inference
cost is then a property of your chosen model, identical whichever search engine
consumes the vectors.
Methodology
- Box: AWS c8g.4xlarge (Graviton4, 16 vCPU, 32 GB, gp3), Amazon Linux 2023 arm64, us-east-1 — full details in "Workload, server & memory" above. Engines run strictly one at a time; data dirs wiped between runs.
- Builds: Index Server 1.3.0 release bundle — the SearchAI leg was re-run on 1.3.0 and reproduced the 1.2.0 numbers within run-to-run noise (1.3.0's changes are additive; the engine is unchanged), so the five-engine table stands (auth disabled to match the other engines' unauthenticated loopback protocol); OpenSearch 3.8.0 tarball; Qdrant 1.19.0 official arm64 binary; Weaviate official arm64 binary; Chroma 1.5.9 (pip).
- Vector workload: seeded 20k × 384 unit vectors (NumPy RandomState(7) — random vectors are HNSW's worst case), identical for every engine; 256 query vectors; 200 timed iterations after 20 warmups; recall@10 against exact brute-force top-10. Filtered kNN = same query with a 50%-selectivity tag filter.
- ANN honesty: we verified each engine's index was actually built before
timing. Two engines needed intervention: Qdrant's default
indexing_threshold=20000never builds HNSW at exactly 20k points (it silently serves exact scan), and SearchAI 1.3.0's crawl-tuned vector-index build threshold (100k staged docs) similarly left vectors on the staged scan path. Both were configured to build (both thresholds lowered to 1,000) and timing started only after latency stabilized. - Text workload: 100k docs (title/body/keyword/long/date) plus a 20k × ~4.5 KB realistic-document phase; single bulk client, chunk 500.
- Hybrid workload: same 20k vectors, each doc adding a seeded 30-word
body from a shared vocabulary; 256 (word, vector) query pairs; k=10; RRF
fusion on every engine (OpenSearch via its score-ranker search pipeline,
Weaviate via
fusionType: rankedFusion); timing starts after latency stabilizes. - Reproduce:
tests/bench/in the source tree —bench.py(OpenSearch-API engines),bench_qdrant.py,bench_vecdb.py(Weaviate/Chroma),bench_hybrid.py(hybrid RRF). Raw JSON results ship in the same directory.