← Blog · Home · Getting Started

Replacing Pinecone with the Index Server

September 2026 · 8 min read

A managed vector database like Pinecone bills you three ways that all grow with your data: managed storage per vector, read/write units for query and upsert traffic, and — because the index doesn't embed for you — a separate per-token bill from whatever embedding API you call before every upsert and every query. The Index Server collapses all three into one line item you already pay for: a server you run. It embeds at ingest with a bundled inference engine (no per-token API), stores vectors on your own disk, and speaks a Pinecone-compatible dialect so your existing client keeps working.

This is the playbook for moving over — and the cost math that usually makes the decision.

Where the money and time actually go

cost / time driverManaged vector DB (Pinecone)Index Server
Embeddingsseparate embedding API, billed per token, on every upsert and every querybundled inference engine embeds at index time — $0 per token, runs on your box
Vector storagemanaged, priced per vector / per pod-hour with a provider markupyour own EBS or local NVMe at raw cloud disk price
Query / upsertmetered read/write unitsunmetered — bounded only by your instance
Ingestion pipelineembed-then-upsert: two network hops per batch, two services to runone bulk call; the server embeds inline as it indexes
Data egressvectors + queries leave your VPC to a third-party cloudstays inside your account / VPC

The ingestion win: one hop instead of two

The usual Pinecone ingestion loop is: read your documents → call an embedding API for each batch (pay per token, wait on the network) → upsert the vectors to Pinecone (pay write units, wait again). Two services, two bills, two round trips per batch.

Here you map a knn_vector field with an embed block once, then just index text. The server calls its bundled inference engine inline, writes the vector into the document, and the embedding travels with the doc — so replicas and snapshots never re-embed:

PUT /docs
{ "mappings": { "properties": {
  "body": { "type": "text" },
  "embedding": { "type": "knn_vector", "dimension": 1024,
                 "embed": { "source_field": "body", "model": "your-embed-model" } }
}}}

# then just index text — no embedding step in your pipeline
POST /docs/_bulk
{"index":{"_id":"1"}}
{"body":"annual leave carries over up to five days"}

At query time you send text, not a vector — the server embeds the query the same way:

POST /docs/_search
{ "query": { "neural": { "embedding": { "query_text": "how much PTO rolls over", "k": 10 } } } }

Your application never touches an embedding API, never holds an embedding-API key, and never pays per token. On our published Graviton benchmarks the server posts the fastest vector ingest of five engines (≈4,380 docs/s including the inline embed step), filtered kNN at 3 ms p50 at recall 1.0, in about 527 MB of RAM for a full search server — not a vector-only store.

Keep your Pinecone client — the dialect is built in

You don't have to rewrite your data layer to try it. The server exposes a Pinecone-compatible dialect on its own listener, so existing upsert/query/fetch code points at the new host and runs:

# start the Pinecone-dialect listener (or set dialects.pinecone-port in config)
POST /_console/dialects   {"kind":"pinecone","port":5081}

# your existing client, new base URL
index.upsert(vectors=[{"id":"1","values":[...],"metadata":{"body":"..."}}])
index.query(vector=[...], top_k=10, filter={"team":"support"})

Prefer to consolidate onto one API? The same documents are reachable over the OpenSearch REST API with full-text, filters, aggregations and hybrid (BM25 + vector, RRF-fused) search — things a pure vector database can't do server-side at all.

Migrating the vectors you already have

Two paths, both hands-off:

Console Migrate tab

The built-in Migrate tool connects to your Pinecone index, lists it, and streams the vectors + metadata over — no exporter to write. (It imports from Qdrant, Chroma, Weaviate and OpenSearch too, if you're retiring more than one store.) It records source-vs-target counts so completeness is part of the job.

Client-side re-embed (recommended when you can)

If you still have the source text, the cleanest migration re-indexes the text and lets the server embed — you drop the old vectors entirely and never pay to move them. Point a bulk loader at the text, map the embed field as above, and you're done; the vectors are regenerated locally at no per-token cost.

What this looks like on AWS (and any provider)

Because it's a single self-contained binary, the deployment is just an instance and a disk:

The shape of the savings: a managed vector DB's monthly cost scales with your vector count and traffic and carries a parallel embedding-API bill that scales with tokens. Self-hosting replaces both curves with a flat instance + disk cost you control — the crossover comes fast once you're past a few million vectors or any steady query load.

Install and try it in ten minutes

One line on a Linux host installs the server and the inference engine as a service:

curl -fsSL https://index-server.searchblox.com/install | sudo bash

Set server.api-key, open http://<host>:9200/console, map one embed field, and bulk-load a sample of your text. Then run the same top-k queries you run against Pinecone today and compare the results — and the bill.

Start with the Getting Started guide or grab a build from the downloads page. The honest trade-offs — where a specialized vector engine still edges raw ANN latency, and where full-text phrase queries favor a Lucene engine — are laid out on the benchmarks page.