← Blog · Home · Getting Started

Add semantic search to your AI app without changing your stack

September 2026 · 7 min read

Adding semantic search to an existing application usually means three changes at once: new infrastructure (a vector database), new client code (its SDK and query language), and a new embedding pipeline (something has to turn text into vectors, at index time and at query time). Each one is a project. Together they're the reason "let's add semantic search" sits in so many backlogs.

The Index Server collapses all three. It's one self-hosted engine that speaks the API your code already uses — OpenSearch/Elasticsearch, Qdrant, Pinecone, Chroma, Algolia — and it embeds text server-side, so your app keeps sending plain documents and plain queries. Here's what that looks like from each starting point.

One install, embeddings included

Since 1.1.0 the installer sets up the SearchAI Inference Server alongside the index server, with embedding and reranker models loaded and wired in. One line on a Linux host and semantic search is a mapping away:

curl -fsSL https://index-server.searchblox.com/install | sudo bash

Already on OpenSearch or Elasticsearch? Change an endpoint.

The server implements the OpenSearch 3.5 REST API (with an Elasticsearch 8.x presentation mode), so opensearch-py, the Java clients, and your existing query DSL work unchanged — _bulk, _search, _msearch, aggregations, highlighting. To make it semantic, attach an embed block to a vector field in the mapping:

PUT /articles
{"mappings": {"properties": {
  "title": {"type": "text"},
  "body":  {"type": "text"},
  "vec":   {"type": "knn_vector", "dimension": 384, "space_type": "cosinesimil",
            "embed": {"source_field": "body", "model": "bge-small-en-v1.5"}}
}}}

That's the whole pipeline. Your app keeps bulk-indexing the same JSON it always has — no vectors in sight — and the server batches one embeddings call per field and fills them in. Querying is just as plain: a neural clause embeds the query text server-side and runs kNN, and the native hybrid query fuses a lexical leg and a semantic leg with reciprocal-rank fusion:

POST /articles/_search
{"query": {"hybrid": {"queries": [
  {"match":  {"body": "how do I rotate api keys"}},
  {"neural": {"vec": {"query_text": "how do I rotate api keys", "k": 20}}}
]}}}

In Python, nothing about the client changes:

from opensearchpy import OpenSearch
os = OpenSearch("http://search.internal:9200")   # was: your OpenSearch cluster
os.index(index="articles", body={"title": "Key rotation", "body": "..."})
hits = os.search(index="articles", body={
    "query": {"neural": {"vec": {"query_text": "rotating credentials", "k": 5}}}})

On Qdrant, Pinecone or Chroma? Keep your client too.

Each vector-database dialect is a separate listener you switch on in conf/server.properties — same engine, same data, a different wire protocol on its own port:

dialects.qdrant-port=6333     # Qdrant REST data-plane
dialects.pinecone-port=6334   # Pinecone data-plane
dialects.chroma-port=8000     # Chroma v2 API
dialects.algolia-port=7700    # Algolia search API

Then your existing code points at the new host and keeps running:

# Qdrant (REST client) — upsert, query, filtered query, scroll
from qdrant_client import QdrantClient
q = QdrantClient(url="http://search.internal:6333")

# Pinecone — upsert, query with metadata filter, fetch, update, stats
index.upsert(vectors=[...])           # pointed at http://search.internal:6334

# Chroma — HttpClient against the v2 API
import chromadb
c = chromadb.HttpClient(host="search.internal", port=8000)
col = c.get_or_create_collection("docs")
col.add(ids=[...], embeddings=[...], documents=[...])
col.query(query_embeddings=[...], n_results=10)

A Qdrant collection, a Pinecone index and a Chroma collection are all the same thing underneath: an index with knn_vector fields. Which means data written through one dialect is searchable through the others — including the full OpenSearch surface. Index from your existing Chroma ingestion job, and on the same data run BM25 full-text queries, filters, aggregations, highlighting, and hybrid ranking that a vector-only store doesn't have.

The adapters are honest, versioned data-plane subsets — the CRUD and query surface a client library actually uses day to day (29 supported activities across the dialects, verified by live probes against the real client libraries). Control-plane and vendor-specific extras (Qdrant recommend/discover and gRPC, Pinecone index management, Algolia Rules) return a clean 4xx rather than a wrong answer. The full matrix is in the docs.

Migrating instead of coexisting

If you'd rather move the data once and retire the old system, the built-in admin console (/console) has a zero-touch Migrate tool that imports from OpenSearch, Elasticsearch, Qdrant, Pinecone, Chroma, Weaviate and Solr — and an Evaluate tab that replays golden queries side-by-side against your current engine so you can verify relevance before switching traffic.

Why server-side embedding matters

Client-side embedding couples every writer and every reader to a model: each service that indexes or queries needs the model dependency, the same version of it, and GPU/CPU budget for it. When the model changes, everything redeploys. Putting embedding behind the index server turns all of that into a mapping detail: writers send text, readers send text, and the model is configured in exactly one place. Reranking works the same way — a rerank block on any query widens the candidate window, scores it with a cross-encoder on the inference server, and returns the reordered top-k.

The short version

Get started with the Getting Started guide, grab a build from the downloads page, or read the changelog to see what shipped in 1.1.0. For enterprise search, RAG and commercial support on top, that's the SearchAI platform.