← Blog · Home · Docs

Scaling your AI search with nodes — and how it differs from OpenSearch and Qdrant

September 2026 · 8 min read

Most AI search workloads have a particular shape: the corpus is gigabytes, not petabytes — product catalogs, documentation, knowledge bases, RAG chunk stores — while the query traffic is what grows. Every chat turn fires retrieval; every agent step searches again. You need more queries per second and higher availability far more often than you need more terabytes.

The Index Server's clustering is built for exactly that shape, and it makes a deliberately different trade than OpenSearch or Qdrant: every node holds a full copy of every index. No shards, no placement, no rebalancing — nodes are interchangeable replicas of the whole engine.

What scaling looks like here

A node is one small native binary. Scaling out is configuration, not architecture:

Read scaling — one line. Point a new node at the cluster and it replays the op-log and serves reads locally:

# on the new node
cluster.seed=10.0.0.1:9200
cluster.secret=<SHARED_SECRET>

High availability — four lines. Switch on Raft consensus (cluster.mode=raft) on 3 or 5 voters and you get leader election, quorum-committed writes, and automatic failover:

cluster.mode=raft
cluster.secret=<SHARED_SECRET>
cluster.advertise=10.0.0.1:9200
cluster.peers=10.0.0.2:9200,10.0.0.3:9200

In raft mode a write is acknowledged only after a majority has committed it — an acked write survives any single-node failure, and a write that can't reach a quorum returns a 503 instead of a false success. Kill the leader and the survivors elect a new one; writes resume with no operator action. A node that was down catches up automatically (incremental log replay, or a full snapshot install if it's far behind). Nodes can be added, promoted and removed live, without a restart.

Because every node has all the data, any node answers any query entirely locally. Behind a load balancer, N nodes ≈ N× the query throughput, and writes can be sent to any node — followers forward them to the leader transparently. The docs include a production load-balancer reference (nginx/HAProxy/ALB): health-check GET /, balance everything across all nodes.

How OpenSearch scales

OpenSearch (like Elasticsearch) is a sharded system, and that's its superpower and its tax. Indexes are split into shards; shards and their replicas are allocated across data nodes; a cluster manager (with its own quorum of manager-eligible nodes) tracks cluster state; shards rebalance when topology changes. This is genuinely the right design for data that cannot fit on one machine — it scales storage horizontally, which the index server's full-replica model deliberately does not.

But you pay for that capability whether or not you use it: shard-count planning (hard to change later), heap sizing and GC behavior, rebalancing storms after node loss, hot shards, scatter-gather queries whose tail latency is set by the slowest shard, and the operational discipline of running dedicated manager nodes. For a corpus that fits on one box — which is most AI retrieval corpora — a shard cluster is complexity rented but not used. And each OpenSearch node is a JVM: our Graviton benchmark measured ~1.5 GB RSS for a single idle-ish node versus ~300 MB for the index server with the same data loaded.

How Qdrant scales

Qdrant's distributed mode shards collections across nodes too, using Raft for cluster metadata while point data moves through shard transfers, with a per-collection replication factor and tunable read/write consistency. It's a capable design for very large vector corpora, with the same category of responsibilities: choosing shard and replication numbers, re-sharding as data grows, and reasoning about consistency settings per request.

The bigger difference is scope. Qdrant nodes replicate vectors and payloads; the index server replicates a complete search engine — BM25 full-text with analyzers, filters, aggregations, highlighting, hybrid RRF ranking, snapshots, and the OpenSearch API surface, alongside HNSW vectors. Scaling a RAG stack on a vector-only store usually still means running a second system for keyword/faceted search; here the node you replicate is both systems.

The honest trade

Picking a topology

Start with one node from the downloads page (or curl -fsSL https://index-server.searchblox.com/install | sudo bash), and turn it into a cluster the day traffic demands it — it's a config change, not a migration. The full contract, including exactly what each mode does and does not guarantee, is in the clustering documentation.