Replacing Pinecone with the Index Server
September 2026 · 8 min read
A managed vector database like Pinecone bills you three ways that all grow with your data: managed storage per vector, read/write units for query and upsert traffic, and — because the index doesn't embed for you — a separate per-token bill from whatever embedding API you call before every upsert and every query. The Index Server collapses all three into one line item you already pay for: a server you run. It embeds at ingest with a bundled inference engine (no per-token API), stores vectors on your own disk, and speaks a Pinecone-compatible dialect so your existing client keeps working.
This is the playbook for moving over — and the cost math that usually makes the decision.
Where the money and time actually go
| cost / time driver | Managed vector DB (Pinecone) | Index Server |
|---|---|---|
| Embeddings | separate embedding API, billed per token, on every upsert and every query | bundled inference engine embeds at index time — $0 per token, runs on your box |
| Vector storage | managed, priced per vector / per pod-hour with a provider markup | your own EBS or local NVMe at raw cloud disk price |
| Query / upsert | metered read/write units | unmetered — bounded only by your instance |
| Ingestion pipeline | embed-then-upsert: two network hops per batch, two services to run | one bulk call; the server embeds inline as it indexes |
| Data egress | vectors + queries leave your VPC to a third-party cloud | stays inside your account / VPC |
The ingestion win: one hop instead of two
The usual Pinecone ingestion loop is: read your documents → call an embedding API for each batch (pay per token, wait on the network) → upsert the vectors to Pinecone (pay write units, wait again). Two services, two bills, two round trips per batch.
Here you map a knn_vector field with an embed block
once, then just index text. The server calls its bundled inference engine inline,
writes the vector into the document, and the embedding travels with the doc — so
replicas and snapshots never re-embed:
PUT /docs
{ "mappings": { "properties": {
"body": { "type": "text" },
"embedding": { "type": "knn_vector", "dimension": 1024,
"embed": { "source_field": "body", "model": "your-embed-model" } }
}}}
# then just index text — no embedding step in your pipeline
POST /docs/_bulk
{"index":{"_id":"1"}}
{"body":"annual leave carries over up to five days"}
At query time you send text, not a vector — the server embeds the query the same way:
POST /docs/_search
{ "query": { "neural": { "embedding": { "query_text": "how much PTO rolls over", "k": 10 } } } }
Your application never touches an embedding API, never holds an embedding-API key, and never pays per token. On our published Graviton benchmarks the server posts the fastest vector ingest of five engines (≈4,380 docs/s including the inline embed step), filtered kNN at 3 ms p50 at recall 1.0, in about 527 MB of RAM for a full search server — not a vector-only store.
Keep your Pinecone client — the dialect is built in
You don't have to rewrite your data layer to try it. The server exposes a Pinecone-compatible dialect on its own listener, so existing upsert/query/fetch code points at the new host and runs:
# start the Pinecone-dialect listener (or set dialects.pinecone-port in config)
POST /_console/dialects {"kind":"pinecone","port":5081}
# your existing client, new base URL
index.upsert(vectors=[{"id":"1","values":[...],"metadata":{"body":"..."}}])
index.query(vector=[...], top_k=10, filter={"team":"support"})
Prefer to consolidate onto one API? The same documents are reachable over the OpenSearch REST API with full-text, filters, aggregations and hybrid (BM25 + vector, RRF-fused) search — things a pure vector database can't do server-side at all.
Migrating the vectors you already have
Two paths, both hands-off:
Console Migrate tab
The built-in Migrate tool connects to your Pinecone index, lists it, and streams the vectors + metadata over — no exporter to write. (It imports from Qdrant, Chroma, Weaviate and OpenSearch too, if you're retiring more than one store.) It records source-vs-target counts so completeness is part of the job.
Client-side re-embed (recommended when you can)
If you still have the source text, the cleanest migration re-indexes the
text and lets the server embed — you drop the old vectors entirely and
never pay to move them. Point a bulk loader at the text, map the
embed field as above, and you're done; the vectors are regenerated
locally at no per-token cost.
What this looks like on AWS (and any provider)
Because it's a single self-contained binary, the deployment is just an instance and a disk:
- Compute: one right-sized instance. Our numbers are from an AWS c8g.4xlarge (Graviton4) — Graviton gives the best price/performance here, but any x86_64 or arm64 host works. You pay for the instance you'd run anyway, not per-vector fees on top.
- Storage: vectors and full-text live on EBS gp3 (or local NVMe) at raw disk price. HNSW + scalar quantization keep the footprint small, and RSS stays modest because the engine scans native columns rather than paging a JVM heap.
- Embeddings: computed on the same box (CPU works; add a GPU instance only if your ingest volume needs it). No embedding-API line item, no tokens leaving your account.
- Network: vectors and queries never leave your VPC — which also removes the data-egress and data-residency conversation entirely.
- HA when you need it: a single node becomes a 3-voter Raft cluster with four config lines per node (scaling article) — still your instances, still no per-vector fee.
The shape of the savings: a managed vector DB's monthly cost scales with your vector count and traffic and carries a parallel embedding-API bill that scales with tokens. Self-hosting replaces both curves with a flat instance + disk cost you control — the crossover comes fast once you're past a few million vectors or any steady query load.
Install and try it in ten minutes
One line on a Linux host installs the server and the inference engine as a service:
curl -fsSL https://index-server.searchblox.com/install | sudo bash
Set server.api-key, open http://<host>:9200/console,
map one embed field, and bulk-load a sample of your text. Then run
the same top-k queries you run against Pinecone today and compare the results —
and the bill.
Start with the Getting Started guide or grab a build from the downloads page. The honest trade-offs — where a specialized vector engine still edges raw ANN latency, and where full-text phrase queries favor a Lucene engine — are laid out on the benchmarks page.