Orchestrate

Vector databases explained: indexes, trade-offs and how to choose

What a vector index does, how HNSW trades a little accuracy for speed, and how to compare local stores, database extensions, managed services and SAP HANA Cloud.

Updated Oct 4, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An AI assistant that answers from your documents first has to find the right passages. Each passage is stored as an embedding: a long list of numbers that captures its meaning. A question becomes an embedding too, and the search looks for the stored embeddings closest to it.

With a few thousand passages, a computer can simply compare the question with every one of them. That is exact search, and it is fast enough.

With millions, comparing everything gets slow and expensive. A vector database keeps an index, a map that jumps to the right neighbourhood and checks only a small part of the data. It is much faster, but it can miss a few true matches. This trade is called approximate nearest neighbour (ANN) search.

So the real questions are not "which vector database is best?" but: how many passages do we have, how much accuracy can we give up for speed, how do we filter by company code or process, and where should the data live?

Why it matters to the business

The vector store sits on the critical path of every answer. Its choices show up in four places.

  • Answer quality. If the index misses the passage that holds the answer, the assistant answers from the next best one. The miss is silent: nobody sees the passage that wasn't returned.
  • Cost. Indexes are usually held in memory. Memory grows with the number of passages and with the length of each embedding. Two million passages from a model with 1,536 numbers per embedding need about 11 GiB for the raw numbers alone, before text, metadata and copies.
  • Control. A sales assistant for company code 2000 must not return passages for company code 3000. Filters by company code, plant or document type are a precision tool and an access tool. Some indexes handle filters poorly, especially very selective ones.
  • Operations. A separate vector database is one more system to secure, back up, monitor and keep in sync with SAP. A vector feature inside a database you already run avoids that.

Take the running order-to-cash example. A credit team asks, "Why is this customer's order blocked?" The assistant searches credit policies, past case notes and help notes. With 20,000 notes, exact search answers in about a millisecond on an ordinary computer, as the deep layer shows. With several million case notes across all company codes, you need an index, a tested accuracy level, and filters that respect who is asking.

How SAP does it

As of October 2026, SAP offers two main places to keep vectors.

  • SAP HANA Cloud vector engine. Generally available since the April 2024 release. Vectors live in the same database as relational, graph, spatial and JSON data, so one query can combine similarity with ordinary business filters. It supports an HNSW vector index (the most common ANN method), and since early 2025 SAP lets you change an index's settings after creation. Vector data can also sit in HANA's disk-based storage extension to cut memory cost. The course set it up in Set up for Unit 7.
  • SAP AI Core grounding (generative AI hub), Vector API. A managed store: you create a collection, name the embedding model it uses, upload documents as chunks with metadata, and query it through the retrieval API. You don't run or tune an index yourself.

SAP also maintains langchain-hana, an open-source package that connects the popular LangChain framework to the HANA Cloud vector engine, including creating an HNSW index.

The options side by side

Kind Examples Good for Watch out for
Plain exact search A few lines of numpy code Up to tens of thousands of passages; tests; ground truth Time grows in step with the data
Store inside your program Chroma (local), Faiss Learning, prototypes, single-user tools One machine; you build backup and access control
Vector feature in a database you run pgvector for PostgreSQL, SAP HANA Cloud vector engine Vectors next to business data; joins and filters in SQL Index tuning and memory are now database concerns
Dedicated vector database Qdrant Very large collections; advanced filtering Another system to run, secure and sync
Managed retrieval service SAP AI Core Vector API Teams that want no index to operate Less control over the index; check metrics and limits

What to compare, in this order:

  1. Where the data and its permissions already live. Keeping vectors next to SAP data often beats a faster engine somewhere else.
  2. Filtering. Can it filter by company code, plant, date and language at search time, and does accuracy hold for rare values?
  3. Scale. Number of vectors, numbers per vector, growth per year, and the memory that implies.
  4. Accuracy. What share of true matches does the index return (its recall), at what speed? Ask for a measurement on your data.
  5. Updates and deletes. How fast do new or corrected documents become searchable, and can old ones be removed cleanly?
  6. Cost and operations. Memory, licences, backups, monitoring, high availability, and who is on call.

Questions to ask

  • How many passages will we index on day one and in three years, and how long is each embedding?
  • Do we need an index at all, or is exact search fast enough at our size?
  • What recall does the index reach on our own test questions, and at what response time?
  • How are filters by company code, sales organization or plant applied? What happens with a filter that matches very few documents?
  • How are deleted or superseded documents removed from the index, and how fast?
  • Where do the vectors live, and does that location meet our data residency and security rules?
  • If we use SAP HANA Cloud, what does the vector index add to our memory sizing and cost?
  • If we use SAP AI Core's Vector API, which usage metrics are billed, and what are the limits per collection?

Common misconceptions

  • "Every RAG project needs a vector database." Tens of thousands of passages can be searched exactly in milliseconds. Start simple; add an index when measurements say so.
  • "The index returns the true nearest matches." ANN indexes are approximate. They can miss matches, and the rate depends on settings and data. Measure it.
  • "Filters are free." A filter can slow searches down or, with some engines, return fewer results than asked. Very selective filters are the hard case.
  • "A filter is the security model." Filters only help if they come from the user's real authorizations, applied on the server. Unit 7's topic on grounding with SAP authorizations covers this.
  • "The fastest engine on a benchmark is the best choice." Data location, filtering, operations and cost usually decide. Benchmarks don't use your data or your filters.

Key terms

  • Embedding (vector): a list of numbers that represents the meaning of a text.
  • Nearest neighbours: the stored vectors closest to a question's vector.
  • Exact search: comparing the question with every stored vector. Always correct, slower at scale.
  • ANN (approximate nearest neighbour) search: using an index to check only part of the data; faster, may miss matches.
  • HNSW: a common ANN index that links vectors into layered graphs of neighbours.
  • Recall at k: the share of the true top k matches that the index actually returns.
  • Distance metric: the rule for "close", for example cosine similarity or straight-line (Euclidean) distance.
  • Metadata filter: a condition on labels such as company code or process, applied during the search.
  • Collection: a named set of vectors that share one embedding model and settings.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Your team has 15,000 help notes to search. A partner proposes a dedicated vector database. What is the best first response?

    Answer: B. Exact search compares the question with every note, and at tens of thousands of notes it takes milliseconds. An index adds speed at scale, but also cost and another system, so measure first.
  2. 2What does an approximate (ANN) index trade away to gain speed?

    Answer: C. An ANN index checks only part of the data, so it can skip a true match. The miss is silent, which is why you measure recall on your own test questions.
  3. 3Which question best tests whether a vector store can serve a sales assistant used across company codes?

    Answer: D. Filtering by company code is both a precision and an access need. Very selective filters are where some indexes lose accuracy or return fewer results than asked.
  4. 4Why might an SAP customer prefer the SAP HANA Cloud vector engine over a separate vector database?

    Answer: A. SAP positions the vector engine as part of the same database as relational and other data, so similarity and business filters meet in one query and one system to run. It still needs sizing and authorization checks.
  5. 5A team wants retrieval on SAP BTP without running or tuning any index. Which SAP option fits?

    Answer: C. The Vector API is a managed store: you create a collection with an embedding model, upload chunks with metadata, and query it. SAP runs the store, so you trade some control for less operations work.
  6. 6Two million passages, each embedded as 1,536 numbers, are planned. What should a leader expect?

    Answer: B. Each number takes 4 bytes, so 2 million times 1,536 numbers is about 11 GiB before text, metadata, index links and copies. Memory is a real cost driver, which is why scale belongs in the comparison.
  7. 7A vendor says its filters make retrieval secure. What is the right concern?

    Answer: D. A filter is a condition in the query. It enforces access only when the server derives it from who the user is, not from what the client sends.
Deep layer · 40 min read

Mental model

A vector index is a shortcut through the data that you pay for twice: once in memory and build time, and again in recall. Exact search is the truth. Every index is judged against it.

Three consequences follow:

  • Measure recall, not just speed. Run the same questions through exact search and through the index, and count how many true neighbours the index returned. A fast index with 70 percent recall silently drops answers.
  • Every index has dials. For HNSW they are the number of links per vector (M, called max_neighbors in Chroma), the effort spent building (ef_construction) and the effort spent per search (ef_search). More effort means better recall and more time or memory.
  • Filters change the problem. A filter shrinks the set of valid answers. An index built for the whole collection may not find them efficiently, so filtered search needs its own measurement.

In RAG fundamentals the "index" box was a black box. This topic opens it.

How it works

Distance: what "close" means

A distance metric turns two vectors into one number. Three are common:

Metric Meaning Chroma space pgvector operator SAP HANA Cloud function
Cosine Angle between vectors; ignores length cosine <=> (cosine distance) COSINE_SIMILARITY
Euclidean (L2) Straight-line distance l2 (the default) <-> L2DISTANCE
Inner product Dot product; larger is closer ip <#> (negative inner product) Not covered here

If every vector is scaled to length 1 (normalised), all three put results in the same order. The course's embedding code normalises, and uses cosine everywhere. Two traps: Chroma's default is l2, so set cosine explicitly; and a distance goes down as similarity goes up. Chroma reports cosine distance as 1 minus cosine similarity.

Exact (or flat) search computes the distance to every stored vector and keeps the best k. The Faiss guidelines call the flat index the baseline for every other index, and recommend it when you run only a few searches, because an index's build time is never paid back. pgvector, too, does exact search unless you create an index, with perfect recall.

Its cost grows in step with the data. In this topic's lab, 20,000 vectors take about 1.2 ms per search in numpy; 100,000 take about 9 ms.

IVF: cluster first, then search a few clusters

An inverted file (IVF) index groups vectors into clusters. A search looks only in the clusters nearest the question. pgvector's IVFFlat calls the clusters lists and the clusters searched per query probes (default 1). Faiss gives rules of thumb for the number of clusters by collection size. IVF builds fast and uses little extra memory, but recall depends heavily on probes.

HNSW: a layered graph of neighbours

HNSW (Hierarchical Navigable Small World) links each vector to a handful of near neighbours, forming a graph. Malkov and Yashunin's paper adds layers: every vector is in the bottom layer, and each one is given a random top layer, with higher layers exponentially rarer. The top layers are sparse "highways"; the bottom layer is the full "street map".

flowchart TD
  Q[Question vector] --> L2[Top layer: few vectors, long jumps]
  L2 -->|closest found becomes start point| L1[Middle layer: more vectors]
  L1 -->|closest found becomes start point| L0[Bottom layer: every vector]
  L0 --> R[Best ef_search candidates kept, top k returned]

A search enters at the top, walks greedily towards the question, then drops a layer and repeats. On the bottom layer it keeps a list of the best ef_search candidates seen so far and returns the top k of them. The paper reports logarithmic scaling of search cost, which is why HNSW stays fast as data grows.

The three dials:

Dial Chroma name and default pgvector name and default Effect
Links per vector max_neighbors, 16 m, 16 More links: better recall, more memory, slower build
Build effort ef_construction, 100 ef_construction, 64 Higher: better graph, slower build
Search effort ef_search, 100 hnsw.ef_search, 40 Higher: better recall, slower search

Chroma lets you change ef_search after the collection exists, but not max_neighbors or ef_construction. Changing those means rebuilding the index.

Memory

The Faiss guidelines estimate HNSW memory as about d * 4 + M * 2 * 4 bytes per vector, where d is the number of dimensions: 4 bytes per number, plus about 2 × M links of 4 bytes. The raw numbers dominate. For 1 million vectors of 384 numbers at M = 16 that is about 1.55 GiB; for 2 million of 1,536 numbers, about 11.7 GiB. Compression such as product quantization (PQ) stores short codes instead of full vectors, trading recall for memory. Smaller numbers help too: pgvector's halfvec type uses half-precision numbers and can index up to 4,000 dimensions, against 2,000 for vector.

Filtering

Real queries carry conditions: company code, process, language, valid-from date. There are three ways to combine them with an index:

  • Post-filter. Search the index, then drop results that fail the filter. pgvector documents this behaviour: filtering is applied after the index scan, so a query can return fewer rows than asked. Since version 0.8.0 it can automatically scan more of the index (iterative index scans).
  • Pre-filter. Find all rows that pass the filter, then search only those. Exact, but Qdrant's write-up notes that a broad filter turns this into brute force.
  • Filter during traversal. Qdrant's approach: extra graph links between points that share an indexed value, a repair step for strict filters, and a planner that estimates how many points pass and picks a path, including a full scan for very selective filters.

The lesson is not "use engine X". It is: test filtered queries separately, with common and rare values. The lab does exactly that.

Updates and deletes

Indexes are cheapest to build once. Real content changes: new case notes daily, superseded manuals, documents that must be deleted. Check how your store handles it. The Faiss guidelines note that its HNSW index does not support removing vectors. Database-backed stores handle deletes as row deletes, but the index still has to absorb them.

Build it yourself: an index lab with recall and filters

Before you start: complete Set up your computer for this course and Set up for Unit 7. They create your orchestrate-course folder with its .venv, install Chroma (numpy comes with it) and create the unit07 folder.

You will build vector_index_lab.py. It stores 20,000 made-up vectors twice: as a plain array for exact search, and in a Chroma collection with an HNSW index. Each vector gets SAP-style labels: a company code and a process. Then you measure, on 200 test questions, how often the index finds the true top 10 (recall) and how long it takes, with and without a company code filter.

The vectors are synthetic: random points scattered around 200 topic centres, 384 numbers each, the same shape as the Unit 3 embedding model's output. That way there is no model download, and you can make 100,000 of them in seconds. The index can't tell the difference.

flowchart LR
  G[build: 20,000 made-up vectors<br/>with company code and process] --> A[numpy array<br/>exact search]
  G --> C[(Chroma collection<br/>HNSW index)]
  A --> B[bench: recall and ms<br/>per ef_search]
  C --> B
  A --> F[filter: recall and ms<br/>per company code]
  C --> F

What you need

  • Your course folder with the Unit 7 setup done. Cost: free.
  • No new library: Chroma and numpy are already installed.
  • About 45 minutes. No account and no network; everything runs on your computer.
  • About 400 MB of free disk space for the optional 100,000-vector run.

Step 1: Open your course folder

  1. Open a terminal (on Windows, PowerShell; on macOS, Terminal) and turn on your environment.

    Windows (PowerShell):

    cd $HOME\orchestrate-course
    .\.venv\Scripts\Activate.ps1

    macOS/Linux:

    cd ~/orchestrate-course
    source .venv/bin/activate
  2. Check that Chroma and numpy are there:

    pip show chromadb numpy

    You should see Name: chromadb and Name: numpy, each with a Version: line. If one says WARNING: Package(s) not found, add chromadb and numpy to requirements.txt and run pip install -r requirements.txt.

Step 2: Create the script

  1. In VS Code, right-click the unit07 folder, choose New File, name it vector_index_lab.py, paste the code below and save.
"""Unit 7: a vector index lab. Compare exact search with an HNSW index in Chroma.

It makes made-up vectors shaped like the Unit 3 model's output (384 numbers each), labels them
with SAP-style metadata (company code, process), and stores them twice: as a plain array for
exact search, and in a Chroma collection with an HNSW index. Then it measures how often the
index finds the true nearest neighbours (recall) and how long each search takes.

How to run (from your course folder, with .venv turned on):
    python unit07/vector_index_lab.py build                      # make the data and the index
    python unit07/vector_index_lab.py bench                      # exact vs HNSW at several ef_search values
    python unit07/vector_index_lab.py filter                     # search inside one company code
    python unit07/vector_index_lab.py size --n 2000000 --dim 1536  # memory estimate for a bigger index
Options for build: --n (how many vectors, default 20000), --dim (default 384),
--m (HNSW max_neighbors, default 16), --efc (HNSW ef_construction, default 100).
"""
import argparse
import json
import shutil
import subprocess
import sys
import time
from pathlib import Path

try:
    import numpy as np
except ImportError:
    sys.exit("numpy is not installed. Run: pip install -r requirements.txt")

HERE = Path(__file__).resolve().parent
STORE = HERE / "index_lab_store"      # the Chroma database folder
DATA = HERE / "index_lab_data.npz"    # the same vectors, for exact search
INFO = HERE / "index_lab_info.json"   # the settings used by build
COLLECTION = "index_lab"
K = 10                                # how many neighbours each search returns
QUERIES = 200                         # how many test searches to run
EF_VALUES = [10, 20, 50, 100, 200]    # ef_search settings to compare

# Made-up metadata: company code 3000 is a small subsidiary with only 0.5% of the documents.
COMPANY_CODES = ["1000", "2000", "3000"]
COMPANY_SHARE = [0.60, 0.395, 0.005]
PROCESSES = ["order-to-cash", "procure-to-pay", "plan-to-produce"]


def open_chroma():
    try:
        import chromadb
        from chromadb.config import Settings
    except ImportError:
        sys.exit("chromadb is not installed. Run: pip install -r requirements.txt (see Set up for Unit 7)")
    return chromadb.PersistentClient(path=str(STORE), settings=Settings(anonymized_telemetry=False))


def normalise(vectors):
    """Scale every vector to length 1, so a dot product equals cosine similarity."""
    return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)


def make_vectors(rng, n, dim, centres):
    """Points scattered around topic centres, like embeddings of documents about similar things."""
    topic = rng.integers(0, len(centres), size=n)
    noise = rng.normal(size=(n, dim)).astype(np.float32)
    return normalise(centres[topic] + 0.9 * normalise(noise))


def exact_top_k(vectors, query, k):
    """Exact search: score every stored vector, keep the k best. This is the ground truth."""
    scores = vectors @ query
    best = np.argpartition(-scores, k)[:k]
    return best[np.argsort(-scores[best])]


def load_all():
    if not (DATA.exists() and INFO.exists() and STORE.exists()):
        sys.exit("No index yet. Run first: python unit07/vector_index_lab.py build")
    data = np.load(DATA)
    info = json.loads(INFO.read_text())
    collection = open_chroma().get_collection(COLLECTION)
    return data["vectors"], data["queries"], data["company"], info, collection


def hnsw_ids(collection, queries, where=None):
    """Ask Chroma for the top K of each query; return the ids found and the average milliseconds."""
    found, start = [], time.perf_counter()
    for query in queries:
        result = collection.query(query_embeddings=[query.tolist()], n_results=K, where=where, include=[])
        found.append([int(i) for i in result["ids"][0]])
    return found, 1000 * (time.perf_counter() - start) / len(queries)


def recall(found, truth):
    """Share of the true top K that the index also returned, averaged over all queries."""
    hits = [len(set(f) & set(t.tolist())) / max(len(t), 1) for f, t in zip(found, truth)]
    return sum(hits) / len(hits)


def cmd_build(args):
    if STORE.exists():
        shutil.rmtree(STORE)
    rng = np.random.default_rng(7)
    centres = normalise(rng.normal(size=(200, args.dim)).astype(np.float32))
    vectors = make_vectors(rng, args.n, args.dim, centres)
    queries = make_vectors(rng, QUERIES, args.dim, centres)
    company = rng.choice(len(COMPANY_CODES), size=args.n, p=COMPANY_SHARE)
    process = rng.integers(0, len(PROCESSES), size=args.n)
    np.savez(DATA, vectors=vectors, queries=queries, company=company)

    client = open_chroma()
    collection = client.create_collection(
        name=COLLECTION, embedding_function=None,
        configuration={"hnsw": {"space": "cosine", "max_neighbors": args.m, "ef_construction": args.efc}})
    start = time.perf_counter()
    batch = 5000
    for first in range(0, args.n, batch):
        rows = range(first, min(first + batch, args.n))
        collection.add(
            ids=[str(i) for i in rows],
            embeddings=vectors[first:first + len(rows)].tolist(),
            metadatas=[{"company_code": COMPANY_CODES[company[i]], "process": PROCESSES[process[i]]} for i in rows])
        print(f"  added {rows[-1] + 1:,} of {args.n:,}")
    seconds = time.perf_counter() - start
    INFO.write_text(json.dumps({"n": args.n, "dim": args.dim, "m": args.m, "efc": args.efc}))
    print(f"Built an HNSW index of {args.n:,} vectors x {args.dim} numbers in {seconds:.1f} s "
          f"(max_neighbors={args.m}, ef_construction={args.efc})")
    counts = {code: int((company == i).sum()) for i, code in enumerate(COMPANY_CODES)}
    print("Documents per company code:", counts)


def probe(ef, code=None):
    """Run the test searches at one ef_search value in a fresh Python process.

    Chroma keeps an index in memory once it is loaded, so a new ef_search only takes effect
    when the index is loaded again. A fresh process for each setting is the simplest way."""
    command = [sys.executable, str(Path(__file__).resolve()), "probe", "--ef", str(ef)]
    if code:
        command += ["--company", code]
    done = subprocess.run(command, capture_output=True, text=True)
    if done.returncode != 0:
        sys.exit(done.stderr.strip() or done.stdout.strip())
    return json.loads(done.stdout.strip().splitlines()[-1])


def reset_ef():
    """Put ef_search back to Chroma's default (100) so the saved collection is left as it was."""
    open_chroma().get_collection(COLLECTION).modify(configuration={"hnsw": {"ef_search": 100}})


def cmd_probe(args):
    """Hidden helper used by bench and filter: one ef_search value, one optional filter."""
    vectors, queries, company, _, collection = load_all()
    collection.modify(configuration={"hnsw": {"ef_search": args.ef}})
    collection = open_chroma().get_collection(COLLECTION)
    rows = np.arange(len(vectors))
    where = None
    if args.company:
        rows = np.flatnonzero(company == COMPANY_CODES.index(args.company))
        where = {"company_code": args.company}
    truth = [rows[exact_top_k(vectors[rows], q, min(K, len(rows)))] for q in queries]
    found, ms = hnsw_ids(collection, queries, where)
    print(json.dumps({"recall": recall(found, truth), "ms": ms}))


def cmd_bench(args):
    vectors, queries, _, info, _ = load_all()
    start = time.perf_counter()
    for q in queries:
        exact_top_k(vectors, q, K)
    exact_ms = 1000 * (time.perf_counter() - start) / len(queries)
    print(f"{info['n']:,} vectors, {len(queries)} test searches, top {K}\n")
    print(f"{'method':<26}{'recall@10':>10}{'ms/search':>11}")
    print(f"{'exact (numpy, every row)':<26}{1.0:>10.3f}{exact_ms:>11.2f}")
    for ef in EF_VALUES:
        result = probe(ef)
        print(f"{'HNSW ef_search=' + str(ef):<26}{result['recall']:>10.3f}{result['ms']:>11.2f}")
    reset_ef()
    print("\nRecall 1.000 means the index found every true neighbour. Times include Chroma's own overhead.")


def cmd_filter(args):
    _, _, company, info, _ = load_all()
    print(f"Filtered search, top {K}, ef_search={args.ef}\n")
    print(f"{'company code':<14}{'share':>8}{'recall@10':>11}{'ms/search':>11}")
    for i, code in enumerate(COMPANY_CODES):
        result = probe(args.ef, code)
        share = f"{100 * (company == i).sum() / info['n']:.1f}%"
        print(f"{code:<14}{share:>8}{result['recall']:>11.3f}{result['ms']:>11.2f}")
    reset_ef()
    print("\nTruth is computed by exact search inside the same company code.")


def cmd_size(args):
    raw = args.n * args.dim * 4                  # 4 bytes per number (32-bit float)
    graph = args.n * args.m * 2 * 4              # about 2*M links of 4 bytes per vector (FAISS rule of thumb)
    gib = 1024 ** 3
    print(f"{args.n:,} vectors x {args.dim} numbers, M={args.m}")
    print(f"  raw vectors : {raw / gib:6.2f} GiB")
    print(f"  HNSW links  : {graph / gib:6.2f} GiB")
    print(f"  total       : {(raw + graph) / gib:6.2f} GiB  (before metadata, text and copies)")


def main():
    parser = argparse.ArgumentParser(description="Compare exact search with an HNSW index in Chroma.")
    sub = parser.add_subparsers(dest="command", required=True)
    build = sub.add_parser("build", help="make the data and build the index")
    build.add_argument("--n", type=int, default=20000, help="number of vectors (default 20000)")
    build.add_argument("--dim", type=int, default=384, help="numbers per vector (default 384)")
    build.add_argument("--m", type=int, default=16, help="HNSW max_neighbors (default 16)")
    build.add_argument("--efc", type=int, default=100, help="HNSW ef_construction (default 100)")
    sub.add_parser("bench", help="recall and speed: exact vs HNSW")
    flt = sub.add_parser("filter", help="search inside one company code")
    flt.add_argument("--ef", type=int, default=100, help="ef_search to use (default 100)")
    prb = sub.add_parser("probe", help=argparse.SUPPRESS)
    prb.add_argument("--ef", type=int, required=True)
    prb.add_argument("--company", choices=COMPANY_CODES)
    size = sub.add_parser("size", help="estimate memory for an index")
    size.add_argument("--n", type=int, default=1_000_000)
    size.add_argument("--dim", type=int, default=384)
    size.add_argument("--m", type=int, default=16)
    args = parser.parse_args()
    {"build": cmd_build, "bench": cmd_bench, "filter": cmd_filter, "probe": cmd_probe,
     "size": cmd_size}[args.command](args)


if __name__ == "__main__":
    main()
  1. The lab writes three things you can always rebuild, so keep them out of Git. Open .gitignore, add these lines at the end and save:

    unit07/index_lab_store/
    unit07/index_lab_data.npz
    unit07/index_lab_info.json

Step 3: Build the data and the index

  1. Run:

    python unit07/vector_index_lab.py build
  2. You should see something like this (times differ by computer):

      added 5,000 of 20,000
      added 10,000 of 20,000
      added 15,000 of 20,000
      added 20,000 of 20,000
    Built an HNSW index of 20,000 vectors x 384 numbers in 7.0 s (max_neighbors=16, ef_construction=100)
    Documents per company code: {'1000': 12065, '2000': 7842, '3000': 93}

Company code 3000 is a small subsidiary with only 93 documents. You'll need it in Step 5.

Step 4: Compare exact search with HNSW

  1. Run:

    python unit07/vector_index_lab.py bench
  2. You should see a table like this. It takes about a minute, because each ef_search value runs in a fresh Python process:

    20,000 vectors, 200 test searches, top 10
    
    method                     recall@10  ms/search
    exact (numpy, every row)       1.000       1.15
    HNSW ef_search=10              0.892       0.72
    HNSW ef_search=20              0.983       0.86
    HNSW ef_search=50              1.000       0.84
    HNSW ef_search=100             1.000       0.94
    HNSW ef_search=200             1.000       1.09
    
    Recall 1.000 means the index found every true neighbour. Times include Chroma's own overhead.
  3. Read it row by row:

    • At ef_search=10 the index missed about one true neighbour in ten. That is the "approximate" in ANN.
    • From ef_search=50 on, it found them all.
    • At 20,000 vectors, exact search is about as fast as the index. At this size you don't need an index.
  4. Now make the collection five times bigger and run the comparison again:

    python unit07/vector_index_lab.py build --n 100000
    python unit07/vector_index_lab.py bench

    The build takes about 40 seconds. You should see something like:

    100,000 vectors, 200 test searches, top 10
    
    method                     recall@10  ms/search
    exact (numpy, every row)       1.000       8.74
    HNSW ef_search=10              0.646       1.57
    HNSW ef_search=20              0.807       1.63
    HNSW ef_search=50              0.983       1.58
    HNSW ef_search=100             0.999       1.77
    HNSW ef_search=200             1.000       1.76

    Exact search got about eight times slower; HNSW barely moved. But the same ef_search=10 now finds only two thirds of the true neighbours. A setting that was fine at one size is not fine at another. That is why you re-measure when the data grows.

  5. Go back to the smaller collection for the next step:

    python unit07/vector_index_lab.py build

Step 5: Search inside one company code

  1. Run:

    python unit07/vector_index_lab.py filter
  2. You should see:

    Filtered search, top 10, ef_search=100
    
    company code     share  recall@10  ms/search
    1000             60.3%      1.000      36.29
    2000             39.2%      1.000      29.62
    3000              0.5%      1.000      10.97
    
    Truth is computed by exact search inside the same company code.
  3. Now lower the search effort:

    python unit07/vector_index_lab.py filter --ef 10
    company code     share  recall@10  ms/search
    1000             60.3%      0.959      36.23
    2000             39.2%      0.983      27.33
    3000              0.5%      0.964       5.24
  4. Two lessons:

    • In this local Chroma setup, a filtered search took up to about 40 times longer than an unfiltered one (compare with Step 4). Filters are not free; measure them.
    • Recall held up even for company code 3000, with 93 documents. That is a property of this engine at this size, not a law. pgvector's documentation, for example, warns that filtering after an approximate index scan can return fewer rows than asked. Run this test on whichever store you choose.

Step 6: Estimate memory before you build

  1. Run the estimate for a large collection with a long embedding:

    python unit07/vector_index_lab.py size --n 2000000 --dim 1536
    2,000,000 vectors x 1536 numbers, M=16
      raw vectors :  11.44 GiB
      HNSW links  :   0.24 GiB
      total       :  11.68 GiB  (before metadata, text and copies)
  2. Try --dim 384 with the same --n. A shorter embedding cuts memory by three quarters. Embedding length is a cost decision, not only a quality one.

Step 7: Save your work

  1. Check what Git will save:

    git status

    You should see unit07/vector_index_lab.py and .gitignore. You must not see index_lab_store or index_lab_data.npz. If you do, check the lines you added to .gitignore in Step 2.

  2. Save it:

    git add .gitignore unit07/vector_index_lab.py
    git commit -m "Add vector index lab"

What the code does

Part What it does
COMPANY_CODES, COMPANY_SHARE Three made-up company codes; 3000 holds only 0.5 percent of documents
normalise, make_vectors Points scattered around 200 topic centres, scaled to length 1 so a dot product equals cosine similarity
exact_top_k Scores every vector with one matrix product and keeps the best k: the ground truth
cmd_build Saves vectors to index_lab_data.npz and adds them, with metadata, to a Chroma collection with space: cosine and your max_neighbors and ef_construction
hnsw_ids Runs each test question through Chroma, optionally with where={"company_code": ...}, and times it
recall Share of the true top 10 the index returned, averaged over the 200 questions
probe Starts a fresh Python process per ef_search value, because Chroma keeps a loaded index in memory with the setting it was loaded with
cmd_bench, cmd_filter Print the comparison tables, then set ef_search back to Chroma's default of 100
cmd_size The Faiss rule of thumb: 4 bytes per number plus about 2 × M links of 4 bytes per vector

If something goes wrong

What you see What it means What to do
python is not recognized / command not found Python isn't on your PATH, or the terminal was open before you installed it Close and reopen the terminal; see Set up your computer for this course
chromadb is not installed or numpy is not installed The library is missing in the Python you're using Check for (.venv) in the prompt, then pip install -r requirements.txt
pip install fails with a proxy or SSL error The company network blocks the package index Ask IT to allow pypi.org, or try on another network
No index yet. Run first: ... bench or filter ran before build Run python unit07/vector_index_lab.py build
error: argument command: invalid choice The subcommand is misspelled Use build, bench, filter or size
The build is very slow or the computer runs out of memory --n is too large for your machine Use the default 20,000, or close other programs
bench shows recall 1.000 on every row Your data is easy for the index at this size Build with --n 100000, or --m 8, and run it again
Times jump around between runs Other programs are using the computer Run again; compare patterns, not single numbers

The SAP way

As of October 2026, two SAP services keep vectors for you.

SAP HANA Cloud vector engine

SAP made the vector engine generally available with the April 2024 release of SAP HANA Cloud. It stores vectors in a column of type REAL_VECTOR, in the same database as your relational data, and searches them with SQL functions such as COSINE_SIMILARITY and L2DISTANCE. You ran such a query in Set up for Unit 7.

Exact search first. Without a vector index, HANA compares the question with every row. An SAP product expert measured this in April 2024: 300,000 vectors of 1,024 numbers, about 97 ms per similarity search, and about 7 ms when the same query also filtered on an ordinary column. The pattern matches the lab: exact search is a sound start, and a selective business filter in SQL cuts the work.

HNSW index. SAP's own open-source langchain-hana package (version 1.2.0, August 2026) creates an HNSW index with SQL of this shape. The dials have the same names as in this topic:

-- SKETCH: an HNSW index on an example table HELP_NOTES with a REAL_VECTOR column EMBEDDING.
-- Ranges checked by SAP's langchain-hana package: M 4 to 1000, efConstruction 1 to 100000, efSearch 1 to 100000.
CREATE HNSW VECTOR INDEX HELP_NOTES_COSINE_IDX
  ON HELP_NOTES ("EMBEDDING")
  SIMILARITY FUNCTION COSINE_SIMILARITY
  BUILD CONFIGURATION '{"M": 16, "efConstruction": 100}'
  SEARCH CONFIGURATION '{"efSearch": 100}'
  ONLINE;

The same thing from Python with SAP's package (also a sketch; it needs pip install langchain-hana, a HANA connection as in Unit 7 setup, and an embedding object):

# SKETCH: needs an SAP HANA Cloud connection and a LangChain embeddings object.
from langchain_hana import HanaDB

store = HanaDB(connection=connection, embedding=embeddings, table_name="HELP_NOTES")
store.create_hnsw_index(m=16, ef_construction=100, ef_search=100)   # leave out to use database defaults

The index is tied to the similarity function: create it for COSINE_SIMILARITY if your queries rank by cosine.

Changing and sizing. SAP's March 2025 release notes added ALTER VECTOR INDEX, to change an existing index's settings, and the option to keep vector data in the native storage extension, HANA's disk-based store, to reduce memory cost for large collections. The same release allowed the VECTOR_EMBEDDING function, which creates embeddings inside the database, in calculation view expressions. SAP's langchain-hana package also accepts a HALF_VECTOR column type, which stores each number in 2 bytes instead of 4; check that your HANA Cloud release supports it before you plan on the saving.

Filtering. In HANA, a filter is an ordinary WHERE clause on ordinary columns: company code, plant, document type, valid-from date. That is the main argument for keeping vectors next to business data. How the optimizer combines a WHERE clause with an HNSW index is not something this course could confirm from SAP documentation; test filtered queries with common and rare values, as in Step 5.

SAP AI Core grounding: the Vector API

The managed path. The SAP Cloud SDK for AI documents it as: create a collection with an embeddingConfig naming the embedding model (for example text-embedding-3-small); add documents made of chunks, each with metadata; and search through the retrieval API with dataRepositoryType: 'vector' and a maxChunkCount. You don't choose an index type or tune ef_search. The chunks.jsonl file from Chunking and document preparation is the input this path expects.

Licensing notes

  • Chroma, Faiss, pgvector and langchain-hana are open source.
  • In SAP HANA Cloud, a vector index adds memory, so include it in your instance sizing. The trial is free but stops nightly.
  • SAP AI Core grounding usage is billed against your SAP BTP account; check the metrics in your contract.

Build vs. SAP

Situation Chroma or Faiss (this topic) pgvector or a dedicated vector DB SAP HANA Cloud vector engine SAP AI Core Vector API
Learning, offline tests Best Needs a server Trial stops nightly Needs SAP AI Core
Tens of thousands of chunks Exact search is fine Exact search is fine Exact search is fine Fine
Millions of chunks One machine's memory Built for it HNSW index, disk storage option SAP runs it; check limits
Filters on SAP fields Metadata only Columns or payload SQL WHERE on business columns Key/value metadata
Data already in SAP Copy out Copy out Often already there Upload chunks
Index tuning Full control Full control M, efConstruction, efSearch None
Operations Yours Yours or the vendor's Your HANA Cloud team SAP

A practical path: prototype with exact search or Chroma, build an evaluation set, then choose. If the business data and its permissions live in SAP HANA Cloud, start there. If you want no index to operate, use the Vector API. Reach for a dedicated vector database when scale or filtering needs exceed both, and you can staff its operations.

Production concerns

  • Authorizations. Store access labels (company code, sales organization, confidentiality) with every vector, and build the filter on the server from the user's identity. Never trust a filter sent by the client. Unit 7's topic on grounding with SAP authorizations goes further.
  • Evaluation. Keep a recall test like the lab's: a fixed set of questions with their exact top k. Run it after every rebuild, setting change and data load. Track filtered and unfiltered recall separately.
  • Sizing. Estimate memory from count × dimensions × 4 bytes plus links, then add text, metadata, replicas and growth. Shorter embeddings and compression cut memory; test their effect on recall.
  • Freshness and deletes. Decide how fast new documents must be searchable. Test deletes: a superseded manual that is still retrievable gives contradictory answers.
  • One model per collection. Vectors from different embedding models, or different versions, are not comparable. Changing the model means re-embedding everything; record the model name with the collection.
  • Latency budget. Retrieval shares the response time with the model call. Measure p95 latency, not averages, with filters on.
  • Backups and rebuilds. Keep the source chunks and metadata; an index can always be rebuilt from them, but not the other way round.
  • Clean core. Vector stores sit side by side on SAP BTP or in SAP HANA Cloud. Read SAP data through released APIs and never write vectors into S/4HANA tables.

Pitfalls

  • Adding an index too early. At tens of thousands of chunks, exact search is fast and always correct.
  • Never measuring recall. Without exact-search ground truth, a misconfigured index looks fine.
  • Copying settings across sizes. The lab's ef_search=10 gave 0.89 recall at 20,000 vectors and 0.65 at 100,000.
  • Mixing distance settings. Building with l2 and thinking in cosine, or ranking by distance in the wrong direction.
  • Testing only unfiltered queries. Filters change both speed and recall; rare values are the hard case.
  • Treating the vector store as the source of truth. Keep the chunks and metadata elsewhere so you can rebuild.
  • Ignoring memory until go-live. Long embeddings at millions of rows are a sizing line item.

Exercise

Find the cheapest index settings that keep recall at 0.95 or more on 100,000 vectors, and write them down. The report, unit07/index_lab_report.txt, is your baseline for the SAP HANA Cloud vector engine topic later in Unit 7, where you will set the same M, efConstruction and efSearch in SQL and compare.

  1. Build the large collection with fewer links:

    python unit07/vector_index_lab.py build --n 100000 --m 8
  2. Save the comparison to the report.

    Windows (PowerShell):

    python unit07/vector_index_lab.py bench | Out-File -Encoding utf8 unit07/index_lab_report.txt

    macOS/Linux:

    python unit07/vector_index_lab.py bench > unit07/index_lab_report.txt
  3. Rebuild with --m 16, then with --m 32, each followed by a bench run appended to the same file.

    Windows (PowerShell):

    python unit07/vector_index_lab.py build --n 100000 --m 32
    python unit07/vector_index_lab.py bench | Out-File -Append -Encoding utf8 unit07/index_lab_report.txt

    macOS/Linux:

    python unit07/vector_index_lab.py build --n 100000 --m 32
    python unit07/vector_index_lab.py bench >> unit07/index_lab_report.txt

    Note the build time printed by each build run.

  4. Open unit07/index_lab_report.txt. For each M, find the lowest ef_search with recall@10 of 0.95 or more.

  5. Add three lines at the end of the file by hand: the build time for each M; the setting you would choose; and one sentence on why (recall, search time, build time, memory from size --n 100000 --m <M>).

  6. Rebuild the default collection (python unit07/vector_index_lab.py build) and commit the report:

    git add unit07/index_lab_report.txt
    git commit -m "Add vector index report"

Done when unit07/index_lab_report.txt holds three bench tables (M = 8, 16 and 32 at 100,000 vectors) and your three hand-written lines, including the chosen M and ef_search and the reason.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the lab compute exact search alongside the HNSW index?

    Answer: B. Recall is the share of the true top k that the index returns, and only exact search knows the true top k. Without it, a misconfigured index looks the same as a good one.
  2. 2In HNSW, what does raising ef_search do?

    Answer: C. ef_search is the size of the candidate list on the bottom layer during a search. More candidates find more true neighbours at some cost in time. Links per vector are M (max_neighbors), and build effort is ef_construction.
  3. 3The lab shows recall 0.89 at ef_search=10 with 20,000 vectors and 0.65 with 100,000. What should you conclude?

    Answer: D. The same setting gave very different recall at two sizes. Settings are tied to the data, so re-run the recall test after loads and rebuilds rather than copying a value that once worked.
  4. 4Why does vector_index_lab.py run each ef_search value in a fresh Python process?

    Answer: A. In the lab's testing, a changed ef_search took effect only when the index was loaded again. A fresh process per setting is the simplest reliable way to measure each value.
  5. 5A sales assistant filters by company code, and one subsidiary holds 0.5 percent of documents. What is the right test before go-live?

    Answer: C. Filters change both speed and recall, and very selective values are the hard case. pgvector documents that post-filtering can return fewer rows, so test the store you choose with real filter values.
  6. 6You plan 2 million chunks with 1,536-number embeddings in an HNSW index with M = 16. Roughly what memory do the vectors and links need?

    Answer: B. Each number is 4 bytes, so 2 million × 1,536 × 4 is about 11.4 GiB, and links add about 2 × M × 4 bytes per vector, about 0.24 GiB. The raw vectors dominate, so embedding length is a cost lever.
  7. 7In SAP HANA Cloud, which statement shape creates an HNSW index for cosine ranking, as SAP's langchain-hana package builds it?

    Answer: D. The package generates CREATE HNSW VECTOR INDEX with a SIMILARITY FUNCTION, plus optional build and search configuration as JSON. The index is tied to the similarity function, so it must match how queries rank.
  8. 8A team wants retrieval for SAP data with company code filters and no new system to operate, and the data is already in SAP HANA Cloud. What would you propose first?

    Answer: A. Keeping vectors next to the business data lets similarity and business filters meet in one SQL query and one system. A separate store adds syncing and operations, and a low ef_search trades away recall without a measured need.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in