Orchestrate

Set up for Unit 7: vector search: HANA Cloud and a local store

Install a local vector store and SAP's HANA Python driver, add SAP HANA Cloud to your BTP trial, and prove both can store and search vectors.

Updated Oct 2, 2026Foundational 7 minDeep 40 min
Foundational layer · 7 min read

The 60-second version

Unit 7 teaches retrieval-augmented generation (RAG): before a model answers, the system looks up the most relevant passages from the company's own documents and data, and hands them to the model. The lookup needs a vector store, a database that keeps each passage as a list of numbers (an embedding) and finds the passages closest in meaning to a question.

This setup gives each learner two vector stores:

  • A local store on the laptop, using the free, open-source library Chroma. It needs no account and works offline.
  • SAP HANA Cloud in the learner's SAP BTP trial. Its vector engine stores embeddings in normal database tables and searches them with SQL. It is where SAP-shaped RAG lives later in the unit.

Setup takes 45 to 75 minutes, a good part of it waiting for the database to be created. Money: nothing, as long as the learner uses the BTP trial.

Why it matters to the business

A model alone doesn't know your credit policy, your tolerance rules for invoices or last week's change to MRP settings. RAG is the usual way to give it that knowledge without retraining. The vector store is the part that decides which passages the model sees, so its quality and its security drive the quality and security of the answer.

Three points matter to a leader:

  • Local first, then the platform. A local store lets a team try ideas in an afternoon, on made-up data, with no procurement. The platform store is where real data, real users and authorizations come in.
  • Data next to data. SAP HANA Cloud keeps embeddings beside business tables. That makes filters such as company code or sales organization ordinary SQL, which matters when answers must respect who may see what. Unit 7 has a topic on exactly that.
  • Trials have rules. The BTP trial database stops every night and must be started again each day. The trial ends after 90 days at most. It is for learning, never for customer data.

Picture the running example. A credit clerk asks why order 4711 is blocked. A RAG system finds the two most relevant notes from the credit policy and passes them to the model, which then explains the block in plain words and cites the policy. This setup builds the "find the notes" part, twice.

How SAP does it

As of October 2026:

  • SAP HANA Cloud is SAP's cloud database on SAP BTP. Its vector engine adds a data type, REAL_VECTOR, and SQL functions such as COSINE_SIMILARITY and L2DISTANCE. SAP Learning describes vectors of 1 to 65,000 numbers.
  • In a BTP trial, SAP's tutorials add SAP HANA Cloud through the subaccount's entitlements and the SAP HANA Cloud Central administration tool. No payment details are needed.
  • The free tier is a different offer for productive (Pay-As-You-Go or CPEA) accounts. SAP's tutorial describes a free tier instance with 30 GB of memory, 2 vCPUs and 120 GB of storage. A company team that wants a shared sandbox would look here, not at personal trials.
  • SAP also offers a 30-day "basic trial" of SAP HANA Cloud on sap.com. It runs on a shared tenant with guided tours and a subset of functions. It is not what this course uses: the course needs a database of your own that your Python code connects to.
  • SAP's Python driver, hdbcli, connects Python programs to SAP HANA. It is free to install, under SAP's own developer licence.

Some vector features, such as computing embeddings inside the database, need extra options on the instance. This course computes embeddings in Python, so the setup doesn't depend on them.

What this unit adds

Tool What it is Cost Used in
Chroma (chromadb) Open-source vector store that runs inside a Python program Free RAG fundamentals; Advanced retrieval
hdbcli SAP's Python driver for SAP HANA Free SAP HANA Cloud vector engine; Grounding on SAP data
SAP HANA Cloud in your BTP trial SAP's cloud database with a vector engine Free, time-limited SAP HANA Cloud vector engine; later units
Sentence Transformers (from Unit 3) Turns text into embeddings on the laptop Free Every Unit 7 topic

Model calls later in the unit use the SAP generative AI hub access from Set up for Unit 5. That trial is short, so each topic also offers a path without a model.

Time and money

  • Time: 45 to 75 minutes. Creating the database takes a good part of it; the rest is installs and two short tests.
  • Money: nothing on the BTP trial. A team instance on a productive account uses the free tier or a paid plan; check your contract.
  • Daily habit: start the trial database each day you use it, and wait until it is running before you connect.
  • Trial clock: the BTP trial from Set up for Unit 3 lasts up to 90 days and is suspended after 30 days without a sign-in.

Questions to ask IT

  • May learners install chromadb and SAP's hdbcli driver from the Python package index?
  • Does the company network allow outgoing database connections on port 443 to SAP HANA Cloud addresses?
  • May a learner's trial database accept connections from any IP address, as SAP's tutorials suggest for external tools? If not, which addresses should it allow?
  • Does the company already run SAP HANA Cloud, and is a free tier or sandbox instance available for team experiments?
  • Which data may be loaded into a vector store at all, and who approves that? (For the setup, the answer is: made-up data only.)

Common misconceptions

  • "A vector store is a new kind of system to buy." Often it isn't. SAP HANA Cloud has a vector engine built in, and Chroma is a free library. The decision is where embeddings should live, not which new product to add.
  • "The trial database runs all the time." It stops every night. A morning "connection failed" usually means it needs starting, not that something broke.
  • "The sap.com HANA trial and the BTP trial are the same." They aren't. The course needs the database inside your BTP trial, which your own code can reach.
  • "Vectors replace permissions." A vector search finds similar text. It knows nothing about who may read it. Filters on fields such as company code must be added on purpose.

Key terms

  • RAG (retrieval-augmented generation): looking up relevant passages first, then giving them to the model with the question.
  • Embedding: a list of numbers that stands for the meaning of a text.
  • Vector store: a database that saves embeddings and finds the closest ones to a question.
  • Chroma: an open-source vector store that runs on a laptop inside Python.
  • SAP HANA Cloud: SAP's cloud database on SAP BTP.
  • Vector engine: the part of SAP HANA Cloud that stores and compares embeddings (REAL_VECTOR, COSINE_SIMILARITY).
  • SAP HANA Cloud Central: the web tool for creating, starting and stopping SAP HANA Cloud instances.
  • Instance: one running database you created.
  • hdbcli: SAP's Python driver for connecting to SAP HANA.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does a vector store do in a RAG system?

    Answer: B. A vector store keeps passages as embeddings and returns the closest ones to a question. The model then answers from those passages; nothing is retrained.
  2. 2Why does the course set up a local store as well as SAP HANA Cloud?

    Answer: D. Chroma runs on the laptop with no account or procurement. SAP HANA Cloud is where real data, users and authorizations come in later.
  3. 3A learner says the database "broke" this morning. What is the most likely cause?

    Answer: C. SAP's tutorials say the trial database shuts down at the end of each day. Start it again in SAP HANA Cloud Central and wait until it is running.
  4. 4Your team wants a shared SAP HANA Cloud sandbox for a pilot. Which offer should you look at?

    Answer: A. Free tier plans exist for Pay-As-You-Go or CPEA accounts and are meant for this kind of start. Personal trials are for individual learning, and the basic trial is a shared, guided tenant.
  5. 5Why does keeping embeddings in SAP HANA Cloud help with security?

    Answer: D. Embeddings sit beside business columns, so the system can filter on them in the same query. The filter still has to be designed; vectors don't carry permissions on their own.
  6. 6Which question to IT matters most before connecting from a laptop?

    Answer: B. SAP's Python driver reaches SAP HANA Cloud on port 443. If the network blocks that, the database can't be reached even when everything else is right.
Deep layer · 40 min read

Mental model: the same search, in two places

Unit 7 does one thing over and over: turn text into vectors, store them, and find the nearest ones to a question. This setup proves you can do that in two places. Chroma runs inside your Python program and saves to a folder. SAP HANA Cloud runs in SAP's cloud, and you talk to it with SQL through SAP's hdbcli driver.

flowchart LR
  T[Help notes] --> E[Embeddings<br/>Sentence Transformers]
  E --> C[(Chroma<br/>unit07/chroma_store)]
  E --> H[(SAP HANA Cloud<br/>REAL_VECTOR column)]
  Q[Question] --> E
  C -->|nearest notes| R[Results]
  H -->|COSINE_SIMILARITY| R

The embedding step is the same for both. Only the storage and the search syntax differ. That is why later topics can switch between them with little change.

How it works

Chroma, a vector store inside your program

Chroma is an open-source vector store published under the Apache 2.0 licence. You install it with pip, and it runs inside your Python program; there is no server to start. Three ideas are enough for now:

Idea What it means
PersistentClient(path=...) Opens a store saved in a folder, so data survives between runs
Collection A named set of items: an ID, the text, optional metadata (labels such as process) and an embedding
query(...) Returns the nearest items to a question vector, optionally filtered with where on metadata

Chroma can compute embeddings for you, but this course passes its own from Sentence Transformers. That keeps one embedding model across Chroma and SAP HANA Cloud.

Chroma's documentation lists three distance measures for a collection: l2 (the default), cosine and ip. The course uses cosine, set with configuration={"hnsw": {"space": "cosine"}}. Chroma reports a distance of 1 minus the cosine similarity, so the script prints 1 - distance to show similarity, the measure you used in Embeddings and semantic similarity.

SAP HANA Cloud and its vector engine

SAP HANA Cloud is SAP's database on SAP BTP. Its vector engine adds a column type and functions to normal SQL, as SAP Learning describes them:

SQL What it does
REAL_VECTOR(384) A column that holds a vector of 384 numbers (1 to 65,000 are allowed)
TO_REAL_VECTOR('[0.1, 0.8]') Turns text in square brackets into a vector
COSINE_SIMILARITY(a, b) How alike two vectors point, from -1 to 1
L2DISTANCE(a, b) The straight-line distance between two vectors

A search is an ordinary SELECT that sorts by similarity and keeps the top rows. Because the vectors sit in a normal table, you can add WHERE conditions on other columns in the same query.

SAP HANA Cloud can also compute embeddings inside the database. The LangChain integration notes that this needs the NLP feature enabled on the instance. This setup doesn't rely on it: you compute embeddings in Python.

Reaching SAP HANA Cloud from Python

SAP's driver hdbcli follows Python's standard database interface. Its PyPI page says three things you need: SAP HANA Cloud uses port 443, encryption is on by default, and autocommit (each statement is saved at once) is on by default. As of 18 September 2026 the current version is 2.30.27, with builds for Windows, macOS and Linux and Python 3.10 to 3.14. It is published under the SAP Developer License Agreement, not an open-source licence.

The course stores the connection details in .env with the four names the LangChain integration uses: HANA_DB_ADDRESS, HANA_DB_PORT, HANA_DB_USER and HANA_DB_PASSWORD. Using the same names now means later topics can reuse them unchanged.

The trial instance's daily rhythm

SAP's trial tutorials describe two rules that shape how you work:

  • It stops every night. Start it in SAP HANA Cloud Central each day you use it, and wait until it shows as running.
  • Unused instances get removed. SAP's tutorial says a free tier instance that isn't restarted within 30 days is deleted. Assume the same risk for your trial instance, and keep your table definitions in code so you can recreate them.

The instance also has to accept connections from your laptop. SAP's tutorial for external tools chooses Allow all IP addresses in the connection settings, which you will do for your learning instance.

Build it yourself: install the Unit 7 tools and search vectors twice

You will install Chroma and hdbcli, build a small local vector store over made-up help notes, add SAP HANA Cloud to your BTP trial, run one vector search in it, and run a check. Steps 1 to 3 need no account.

Before you start: complete Set up your computer for this course, Set up for Unit 2 and Set up for Unit 3. They install Python, VS Code and Git, create your orchestrate-course folder with its .venv, .env and .gitignore, install Sentence Transformers and create your SAP BTP trial. This walkthrough doesn't repeat those steps.

flowchart LR
  S1[Steps 1-3<br/>Chroma, no account] --> S4[Step 4<br/>add HANA Cloud to trial]
  S4 --> S5[Step 5<br/>create instance, .env]
  S5 --> S6[Step 6<br/>hana_hello.py]
  S6 --> S7[Step 7<br/>check_unit07.py]

What you need

  • Your course folder from earlier units.
  • About 45 to 75 minutes.
  • Your SAP BTP trial from Unit 3, for Steps 4 to 6. If it has ended, create a new one as in Set up for Unit 3.
  • Cost: free.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate

Run every command in this topic from the course folder.

Step 2: Add Chroma and SAP's HANA driver

  1. Open requirements.txt and add these two lines at the end, then save:

    chromadb
    hdbcli
  2. Install (the same on every system):

    pip install -r requirements.txt

    Chroma brings several helper libraries, so this takes a minute or two.

  3. Check both:

    pip show chromadb hdbcli

What success looks like (trimmed; from our test, your versions may be newer):

Name: chromadb
Version: 1.5.9
...
Name: hdbcli
Version: 2.30.27

Step 3: Build a local vector store with Chroma

The script stores eight made-up help notes about blocked orders, invoice blocks and MRP exceptions. Each note has a process label. Then it finds the notes closest to your question, optionally only within one process.

  1. Make the Unit 7 folder:

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit07
    • macOS / Linux:

      mkdir -p unit07
  2. In VS Code, right-click unit07, choose New File, name it local_vector_store.py, paste the code below and save.

"""Unit 7 smoke test: a local vector store with Chroma, no account needed.

It stores eight short, made-up help notes about SAP process exceptions, then finds the notes
closest in meaning to your question. The store is saved in unit07/chroma_store, so a second
run reuses it instead of rebuilding.

How to run (from your course folder, with .venv turned on):
    python unit07/local_vector_store.py                                  # default question
    python unit07/local_vector_store.py "invoice quantity does not match the goods receipt"
    python unit07/local_vector_store.py "why is this stuck?" --process order-to-cash
    python unit07/local_vector_store.py --offline                        # no model download; toy embeddings
    python unit07/local_vector_store.py --reset                          # delete and rebuild the store
"""
import argparse
import hashlib
import math
import re
import shutil
import sys
from pathlib import Path

STORE = Path(__file__).resolve().parent / "chroma_store"
MODEL = "sentence-transformers/all-MiniLM-L6-v2"   # the Unit 3 model; 384 numbers per text

NOTES = [  # made-up help notes, shaped like the course's running examples
    ("n1", "order-to-cash", "A sales order is blocked for delivery when the customer's open items exceed the credit limit. Credit management must release it."),
    ("n2", "order-to-cash", "Orders for export customers stop at delivery when customs or export documents are missing."),
    ("n3", "order-to-cash", "A delivery block set by the sales team holds an order until pricing or terms are confirmed with the customer."),
    ("n4", "procure-to-pay", "An invoice is blocked for payment when the invoiced quantity is higher than the quantity received in goods receipt."),
    ("n5", "procure-to-pay", "A price variance between the purchase order and the supplier invoice beyond tolerance blocks the invoice."),
    ("n6", "procure-to-pay", "Three-way match compares purchase order, goods receipt and invoice before the invoice can be paid."),
    ("n7", "plan-to-produce", "MRP raises an exception message when a planned receipt arrives after the date the material is needed."),
    ("n8", "plan-to-produce", "A reschedule-in exception means MRP suggests moving a planned order earlier to cover demand."),
]


def toy_embed(text: str, size: int = 256) -> list:
    """A tiny stand-in for a real model: hash each word into one of 256 slots. Matches words, not meaning."""
    vector = [0.0] * size
    for word in re.findall(r"[a-z]+", text.lower()):
        slot = int(hashlib.md5(word.encode()).hexdigest(), 16) % size
        vector[slot] += 1.0
    length = math.sqrt(sum(v * v for v in vector)) or 1.0
    return [v / length for v in vector]


def get_embedder(offline: bool):
    """Return (name, function that turns a list of texts into a list of vectors)."""
    if offline:
        return "toy-hash-256", lambda texts: [toy_embed(t) for t in texts]
    try:
        from sentence_transformers import SentenceTransformer
    except ImportError:
        sys.exit("sentence-transformers is not installed. See Set up for Unit 3, or add --offline.")
    print(f"Loading {MODEL} (from disk if you ran Unit 3; otherwise it downloads once)...")
    try:
        model = SentenceTransformer(MODEL)
    except Exception as error:   # usually a blocked download on a company network
        sys.exit(f"Could not load the model ({type(error).__name__}). Check your network, or add --offline.")
    return "minilm-384", lambda texts: model.encode(texts, normalize_embeddings=True).tolist()


def main() -> None:
    parser = argparse.ArgumentParser(description="Search made-up SAP help notes with a local Chroma store.")
    parser.add_argument("question", nargs="?", default="Why can't this customer's order be delivered?")
    parser.add_argument("--process", choices=["order-to-cash", "procure-to-pay", "plan-to-produce"],
                        help="only search notes for this process (a metadata filter)")
    parser.add_argument("--top", type=int, default=3, help="how many notes to return (default 3)")
    parser.add_argument("--offline", action="store_true", help="use toy embeddings; no model needed")
    parser.add_argument("--reset", action="store_true", help="delete the saved store and rebuild it")
    args = parser.parse_args()

    try:
        import chromadb
        from chromadb.config import Settings
    except ImportError:
        sys.exit("chromadb is not installed. Run: pip install -r requirements.txt")

    if args.reset and STORE.exists():
        shutil.rmtree(STORE)
        print(f"Deleted {STORE.name}")

    name, embed = get_embedder(args.offline)
    client = chromadb.PersistentClient(path=str(STORE), settings=Settings(anonymized_telemetry=False))
    # One collection per embedding type: vectors of different sizes can't share a collection.
    collection = client.get_or_create_collection(
        name=f"help_notes_{name}", embedding_function=None, configuration={"hnsw": {"space": "cosine"}})

    if collection.count() == 0:
        collection.add(ids=[n[0] for n in NOTES], documents=[n[2] for n in NOTES],
                       metadatas=[{"process": n[1]} for n in NOTES], embeddings=embed([n[2] for n in NOTES]))
        print(f"Stored {collection.count()} notes in {STORE.name}/ (collection {collection.name})")
    else:
        print(f"Reusing {collection.count()} notes from {STORE.name}/ (collection {collection.name})")

    where = {"process": args.process} if args.process else None
    found = collection.query(query_embeddings=embed([args.question]), n_results=args.top, where=where)
    print(f'\nQuestion: "{args.question}"' + (f"  (only {args.process})" if args.process else ""))
    if not found["ids"][0]:
        print("No notes matched the filter. That is a valid empty result, not an error.")
    for rank, (nid, text, meta, dist) in enumerate(zip(found["ids"][0], found["documents"][0],
                                                       found["metadatas"][0], found["distances"][0]), 1):
        print(f"{rank}. [{nid}, {meta['process']}] similarity {1 - dist:.2f}\n   {text}")


if __name__ == "__main__":
    main()
  1. Run it with the real embedding model from Unit 3:

    python unit07/local_vector_store.py

What success looks like (shape only; our test environment couldn't download the model, so your ranking and numbers will differ):

Loading sentence-transformers/all-MiniLM-L6-v2 (from disk if you ran Unit 3; otherwise it downloads once)...
Stored 8 notes in chroma_store/ (collection help_notes_minilm-384)

Question: "Why can't this customer's order be delivered?"
1. [n1, order-to-cash] similarity 0.xx
   A sales order is blocked for delivery when the customer's open items exceed the credit limit. Credit management must release it.
2. ...

You should see order-to-cash notes near the top, because the model compares meaning.

  1. Run it again with --offline. This uses toy embeddings that only count shared words, so you can see the difference:

    python unit07/local_vector_store.py --offline

What success looks like (from our test):

Stored 8 notes in chroma_store/ (collection help_notes_toy-hash-256)

Question: "Why can't this customer's order be delivered?"
1. [n6, procure-to-pay] similarity 0.22
   Three-way match compares purchase order, goods receipt and invoice before the invoice can be paid.
2. [n3, order-to-cash] similarity 0.21
   A delivery block set by the sales team holds an order until pricing or terms are confirmed with the customer.
3. [n1, order-to-cash] similarity 0.20
   A sales order is blocked for delivery when the customer's open items exceed the credit limit. Credit management must release it.

The toy method ranks a procure-to-pay note first, because it shares words such as "order" and "be". It matches words, not meaning. That is the gap a real embedding model closes.

  1. Try a filter. Only plan-to-produce notes are searched:

    python unit07/local_vector_store.py "why is this stuck" --process plan-to-produce --top 2

    Add --offline if the model isn't available. Both results are MRP notes. If you ever see "No notes matched the filter", that is a valid empty result, not an error.

  2. The script saved the store in unit07/chroma_store. A second run says Reusing 8 notes. It is a rebuildable folder, so keep it out of Git: open .gitignore, add this line at the end and save:

    unit07/chroma_store/

What each part of the script does:

Part What it does
NOTES Eight made-up notes, each with an ID, a process label and text
toy_embed The --offline stand-in: hashes words into 256 slots; no model needed
get_embedder Loads all-MiniLM-L6-v2 from Unit 3, or the toy method with --offline
PersistentClient(path=...) Opens the store saved in unit07/chroma_store
get_or_create_collection(...) One collection per embedding type, with cosine distance and no built-in embedding function
collection.add(...) Stores IDs, texts, labels and your embeddings the first time
collection.query(..., where=...) Finds the nearest notes; where filters on the process label
--reset Deletes the folder and rebuilds, for when you change the notes

Step 4: Add SAP HANA Cloud to your BTP trial

These steps follow SAP's trial tutorial. Labels can change slightly; look for the closest match.

  1. Open the trial cockpit, https://cockpit.hanatrial.ondemand.com/trial/, and sign in. In your trial global account, click the tile named trial to open your subaccount, as in Unit 3.
  2. In the left menu, click Entitlements. Search for SAP HANA. You should see SAP HANA Cloud with plans including tools (Application) and hana.
  3. If they are missing, click Edit, then Add Service Plans, search for SAP HANA Cloud, tick the plans above, and click Save. If SAP HANA Cloud doesn't appear at all, your trial region may not offer it; see the troubleshooting table.
  4. In the left menu, click Services, then Service Marketplace. Search for SAP HANA Cloud and click it, then click Create at the top right.
  5. Choose SAP HANA Cloud as the service and tools as the plan, and click Create. This subscribes you to SAP HANA Cloud Central.
  6. Give yourself the admin role. In the left menu, click Security, then Users. Click your user, click Assign Role Collection, tick SAP HANA Cloud Administrator, and click Assign Role Collection.
  7. In the left menu, click Services, then Instances and Subscriptions. Under subscriptions, click SAP HANA Cloud. SAP HANA Cloud Central opens in a new tab. If it says you lack authorization, sign out and in again so the new role takes effect.

Step 5: Create the database and save its details in .env

  1. In SAP HANA Cloud Central, click Create Instance.

  2. Choose the SAP HANA Database type and click Next Step.

  3. Fill in the basics:

    • Instance Name: orchestrate-hana. Spaces aren't allowed.
    • Administrator Password and Confirm Administrator Password: a strong password. Write it in your password manager now.
  4. Click Next Step through the following pages, keeping the defaults, until you reach the connections setting. Choose Allow all IP addresses.

  5. Click Review and Create, check the summary, and click Create Instance.

  6. Wait. The status shows creation in progress, then changes to show the instance is running (SAP's tutorial describes a green Created status). Creating a database takes a while, so use the time to read Step 6.

  7. Open the instance's details and find the SQL Endpoint. It is a long host name followed by :443. Copy it.

  8. Open .env in your course folder and add these four lines at the end. Paste the host name without :443 in the first line, and use your own password:

    HANA_DB_ADDRESS=paste-the-host-name-here
    HANA_DB_PORT=443
    HANA_DB_USER=DBADMIN
    HANA_DB_PASSWORD=your-administrator-password
  9. Save .env.

Each day you use it: open SAP HANA Cloud Central, find orchestrate-hana, open its actions menu (the three dots), and click Start. Wait until it is running again.

Step 6: Run a vector search in SAP HANA Cloud

The script below connects with hdbcli, creates a tiny table with a REAL_VECTOR(3) column, stores three notes with hand-made 3-number vectors, ranks them against a question vector with COSINE_SIMILARITY, and drops the table. Real embeddings have 384 numbers; three keep the output readable.

  1. In VS Code, in unit07, create hana_hello.py, paste the code below and save.
"""Unit 7 smoke test: connect to SAP HANA Cloud from Python and run one tiny vector search.

It reads four HANA_DB_ lines from your .env file, connects with SAP's hdbcli driver, prints the
user it connected as, then creates a small table with a REAL_VECTOR column, ranks three rows by
COSINE_SIMILARITY, and drops the table again.

How to run (from your course folder, with .venv turned on):
    python unit07/hana_hello.py               # connect and run the test (needs a running instance)
    python unit07/hana_hello.py --dry-run     # no account: show the settings it found and the SQL it would run
    python unit07/hana_hello.py --keep        # leave the test table in place so you can look at it
"""
import argparse
import os
import sys

from dotenv import load_dotenv

TABLE = "COURSE_HELLO_VECTORS"
ROWS = [  # three made-up notes with hand-made 3-number vectors: [credit, documents, invoice]
    (1, "Order blocked: credit limit exceeded", "[0.9, 0.1, 0.0]"),
    (2, "Order blocked: export documents missing", "[0.1, 0.9, 0.0]"),
    (3, "Invoice blocked: quantity variance", "[0.0, 0.1, 0.9]"),
]
QUESTION = "[0.8, 0.2, 0.0]"   # "a question mostly about credit"
SQL = [
    f'CREATE COLUMN TABLE {TABLE} (ID INTEGER PRIMARY KEY, NOTE NVARCHAR(200), EMBEDDING REAL_VECTOR(3))',
    f'INSERT INTO {TABLE} VALUES (?, ?, TO_REAL_VECTOR(?))',
    f'SELECT TOP 2 ID, NOTE, COSINE_SIMILARITY(EMBEDDING, TO_REAL_VECTOR(?)) AS SCORE '
    f'FROM {TABLE} ORDER BY SCORE DESC',
    f'DROP TABLE {TABLE}',
]
NAMES = ["HANA_DB_ADDRESS", "HANA_DB_PORT", "HANA_DB_USER", "HANA_DB_PASSWORD"]


def main() -> None:
    parser = argparse.ArgumentParser(description="Connect to SAP HANA Cloud and run a tiny vector search.")
    parser.add_argument("--dry-run", action="store_true", help="don't connect; show settings and SQL")
    parser.add_argument("--keep", action="store_true", help="don't drop the test table at the end")
    args = parser.parse_args()

    load_dotenv()
    settings = {name: os.getenv(name, "") for name in NAMES}
    for name in NAMES:
        shown = "(set, hidden)" if name == "HANA_DB_PASSWORD" and settings[name] else settings[name] or "(missing)"
        print(f"{name:17} {shown}")

    if args.dry_run:
        print("\nDry run. With a running instance, the script would run:")
        for statement in SQL[:-1] + ([] if args.keep else SQL[-1:]):
            print("  " + statement)
        return

    missing = [name for name in NAMES if not settings[name]]
    if missing:
        sys.exit(f"\nMissing in .env: {', '.join(missing)}. See Step 5 of Set up for Unit 7, or use --dry-run.")
    try:
        from hdbcli import dbapi
    except ImportError:
        sys.exit("hdbcli is not installed. Run: pip install -r requirements.txt")

    print("\nConnecting (an instance that was stopped overnight refuses connections until you start it)...")
    try:
        conn = dbapi.connect(address=settings["HANA_DB_ADDRESS"], port=int(settings["HANA_DB_PORT"]),
                             user=settings["HANA_DB_USER"], password=settings["HANA_DB_PASSWORD"])
    except dbapi.Error as error:
        sys.exit(f"Could not connect: {error}\nCheck that the instance is running, the address and port, "
                 "the user and password, and that the instance allows connections from your IP address.")

    cursor = conn.cursor()
    cursor.execute("SELECT CURRENT_USER FROM DUMMY")
    print(f"Connected as {cursor.fetchone()[0]}.")
    try:
        try:
            cursor.execute(SQL[3])   # remove a table left over from an earlier --keep run
        except dbapi.Error:
            pass                     # normal: there was no table to remove
        cursor.execute(SQL[0])
        cursor.executemany(SQL[1], ROWS)
        cursor.execute(SQL[2], (QUESTION,))
        print(f"\nTop 2 notes for the question vector {QUESTION}:")
        for row_id, note, score in cursor.fetchall():
            print(f"  {row_id}. {note}  (cosine similarity {score:.3f})")
    except dbapi.Error as error:
        print(f"\nThe vector test failed: {error}")
    finally:
        if not args.keep:
            try:
                cursor.execute(SQL[3])
                print(f"\nDropped the test table {TABLE}.")
            except dbapi.Error:
                pass
        conn.close()


if __name__ == "__main__":
    main()
  1. No account yet? Run the dry run. It shows what it found in .env and the SQL it would send:

    python unit07/hana_hello.py --dry-run

What success looks like (from our test, before .env had the lines):

HANA_DB_ADDRESS   (missing)
HANA_DB_PORT      (missing)
HANA_DB_USER      (missing)
HANA_DB_PASSWORD  (missing)

Dry run. With a running instance, the script would run:
  CREATE COLUMN TABLE COURSE_HELLO_VECTORS (ID INTEGER PRIMARY KEY, NOTE NVARCHAR(200), EMBEDDING REAL_VECTOR(3))
  INSERT INTO COURSE_HELLO_VECTORS VALUES (?, ?, TO_REAL_VECTOR(?))
  SELECT TOP 2 ID, NOTE, COSINE_SIMILARITY(EMBEDDING, TO_REAL_VECTOR(?)) AS SCORE FROM COURSE_HELLO_VECTORS ORDER BY SCORE DESC
  DROP TABLE COURSE_HELLO_VECTORS
  1. With your instance running and .env filled in, run the real test:

    python unit07/hana_hello.py

What success looks like (expected shape; we couldn't reach SAP HANA Cloud from our test environment, so we didn't run this part):

HANA_DB_ADDRESS   xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx.hana....
HANA_DB_PORT      443
HANA_DB_USER      DBADMIN
HANA_DB_PASSWORD  (set, hidden)

Connecting (an instance that was stopped overnight refuses connections until you start it)...
Connected as DBADMIN.

Top 2 notes for the question vector [0.8, 0.2, 0.0]:
  1. Order blocked: credit limit exceeded  (cosine similarity 0.991)
  2. Order blocked: export documents missing  (cosine similarity 0.348)

Dropped the test table COURSE_HELLO_VECTORS.

The question vector leans towards "credit", so the credit note ranks first. The scores are what cosine similarity gives for these vectors; small rounding differences are fine.

What each part of the script does:

Part What it does
load_dotenv() Reads the four HANA_DB_ lines from .env
--dry-run Shows the settings (password hidden) and the SQL, without connecting
dbapi.connect(...) Opens an encrypted connection on port 443
SELECT CURRENT_USER FROM DUMMY A one-row test query: proves you are connected and as whom
CREATE COLUMN TABLE ... REAL_VECTOR(3) A table with a vector column of three numbers
TO_REAL_VECTOR(?) Turns the text "[0.9, 0.1, 0.0]" into a vector as it is inserted
COSINE_SIMILARITY(...) ... ORDER BY SCORE DESC Ranks rows by how alike their vector is to the question
DROP TABLE Cleans up; --keep leaves the table so you can look at it in SAP HANA Database Explorer

Step 7: Run the Unit 7 check

  1. In the course folder (not in unit07), create check_unit07.py, paste the code below and save. It uses only built-in Python, like the earlier checks.
"""Check that your computer is ready for Unit 7 (retrieval-augmented generation: vector stores).

Run it from your course folder:  python check_unit07.py
It uses built-in Python only. It looks for the Unit 7 libraries and files, reads the names (not the
values) of the HANA_DB_ lines in .env, and tries to reach your SAP HANA Cloud address on the network.
It changes nothing and never prints your password.
"""
import importlib.metadata
import importlib.util
import os
import socket
import ssl
import sys

problems = 0


def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
    """Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
    global problems
    if ok:
        print(f"  OK       {label}")
    elif optional:
        print(f"  LATER    {label}  ->  {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label}  ->  {fix}")


def library(module: str, package: str):
    """Return the installed version of a library, or '' if it isn't installed."""
    if importlib.util.find_spec(module) is None:
        return ""
    try:
        return importlib.metadata.version(package)
    except importlib.metadata.PackageNotFoundError:
        return "installed"


def read_env(path: str = ".env") -> dict:
    """Read NAME=value lines from .env with built-in Python (python-dotenv does this in the scripts)."""
    values = {}
    if os.path.exists(path):
        with open(path, encoding="utf-8") as handle:
            for line in handle:
                line = line.strip()
                if line and not line.startswith("#") and "=" in line:
                    name, value = line.split("=", 1)
                    values[name.strip()] = value.strip().strip('"').strip("'")
    return values


def reachable(host: str, port: int) -> str:
    """Open an encrypted connection to host:port and close it. Returns '' on success, else the reason."""
    try:
        with socket.create_connection((host, port), timeout=10) as raw:
            with ssl.create_default_context().wrap_socket(raw, server_hostname=host):
                return ""
    except (OSError, ValueError) as error:
        return type(error).__name__


print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
       "the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")

print("\n2. Python libraries")
for module, package in [("chromadb", "chromadb"), ("hdbcli", "hdbcli"), ("dotenv", "python-dotenv")]:
    version = library(module, package)
    report(bool(version), f"{package} {version}".strip(), "pip install -r requirements.txt (Step 2)")
version = library("sentence_transformers", "sentence-transformers")
report(bool(version), f"sentence-transformers {version}".strip(),
       "install it as in Set up for Unit 3; until then use --offline", optional=True)

print("\n3. Course folder")
for path, step in [(os.path.join("unit07", "local_vector_store.py"), "3"),
                   (os.path.join("unit07", "hana_hello.py"), "6")]:
    report(os.path.exists(path), path, f"create it (Step {step})")
report(os.path.isdir(os.path.join("unit07", "chroma_store")), "unit07/chroma_store (your local vector store)",
       "run python unit07/local_vector_store.py (Step 3)", optional=True)
ignored = os.path.exists(".gitignore") and "chroma_store" in open(".gitignore", encoding="utf-8").read()
report(ignored, ".gitignore leaves out chroma_store", "add the line unit07/chroma_store/ (Step 3)", optional=True)

print("\n4. SAP HANA Cloud (needed from the HANA Cloud vector engine topic)")
env = read_env()
names = ["HANA_DB_ADDRESS", "HANA_DB_PORT", "HANA_DB_USER", "HANA_DB_PASSWORD"]
missing = [name for name in names if not env.get(name)]
report(not missing, "HANA_DB_ lines in .env", f"add {', '.join(missing)} (Step 5)", optional=True)
if env.get("HANA_DB_ADDRESS"):
    try:
        port = int(env.get("HANA_DB_PORT") or 443)
    except ValueError:
        port = 443
    host = env["HANA_DB_ADDRESS"]
    reason = reachable(host, port)
    report(not reason, f"can reach {host if len(host) <= 40 else host[:40] + '...'} on port {port}",
           f"{reason}: check the address, your network or proxy; then run unit07/hana_hello.py", optional=True)

print()
if problems:
    print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
    sys.exit(1)
print("All set. Your computer is ready for Unit 7. Start your SAP HANA Cloud instance on the days you use it.")
  1. Run it:

    python check_unit07.py

What success looks like (from our test, before the SAP HANA Cloud steps; your versions will differ):

1. Python
  OK       Python 3.13.16
  OK       virtual environment is active

2. Python libraries
  OK       chromadb 1.5.9
  OK       hdbcli 2.30.27
  OK       python-dotenv 1.2.4
  OK       sentence-transformers 6.1.0

3. Course folder
  OK       unit07/local_vector_store.py
  OK       unit07/hana_hello.py
  OK       unit07/chroma_store (your local vector store)
  OK       .gitignore leaves out chroma_store

4. SAP HANA Cloud (needed from the HANA Cloud vector engine topic)
  LATER    HANA_DB_ lines in .env  ->  add HANA_DB_ADDRESS, HANA_DB_PORT, HANA_DB_USER, HANA_DB_PASSWORD (Step 5)

All set. Your computer is ready for Unit 7. Start your SAP HANA Cloud instance on the days you use it.

LATER lines don't stop you. RAG fundamentals and Advanced retrieval use only the local store. With .env filled in, section 4 also tries to reach your SAP HANA Cloud address on port 443. That proves the network path only; hana_hello.py proves the login and the vector engine.

Step 8: Save your work in Git

  1. Check what Git sees:

    git status

    You should see requirements.txt, .gitignore, check_unit07.py and unit07/. You must not see .env or unit07/chroma_store.

  2. Save:

    git add requirements.txt .gitignore check_unit07.py unit07
    git commit -m "Set up Unit 7: Chroma, hdbcli and SAP HANA Cloud"

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
chromadb is not installed or ModuleNotFoundError: No module named 'hdbcli' The library isn't in the Python you are using Check for (.venv) in the prompt, then pip install -r requirements.txt
pip fails while installing chromadb on an older Python Chroma's helper libraries need a recent Python Check python --version; the course needs 3.11 or newer
sentence-transformers is not installed Unit 3's library is missing from this .venv Follow Step 3 of Set up for Unit 3. Meanwhile use --offline
Could not load the model (...) The model download was blocked, often by a company proxy Try another network, or ask IT to allow Hugging Face downloads. Meanwhile use --offline
Chroma error about embedding dimensions You changed the notes or the model and reused an old store Run with --reset
SAP HANA Cloud is missing from Entitlements Your trial region may not offer it Check the region of your trial subaccount; SAP's tutorials can tell you which regions offer it. If none, use the local store for now
HANA Cloud Central says you aren't authorized The role collection isn't active yet Repeat Step 4.6, then sign out and in again
Missing in .env: ... The HANA_DB_ lines aren't in .env, or not in the course folder Repeat Step 5.8 and save. Run the script from the course folder
Could not connect: ... Cannot resolve host name The address is wrong or has :443 on the end Paste the host name only into HANA_DB_ADDRESS; put 443 in HANA_DB_PORT
Could not connect: ... authentication failed Wrong user or password Check HANA_DB_USER=DBADMIN and the password; reset it in HANA Cloud Central if needed
Could not connect with a timeout, the morning after it worked The instance stopped overnight Start it in HANA Cloud Central and wait until it is running
Could not connect with a timeout on a new instance Your IP address isn't allowed, or the network blocks port 443 Set Allow all IP addresses in the instance's connection settings; try another network or ask IT
The vector test failed: ... Connected, but the SQL was refused Read the message. If it names REAL_VECTOR, note the instance version and ask in the course questions
Windows: Activate.ps1 cannot be loaded PowerShell blocks scripts Run Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser, answer Y, and try again

Where this shows up in SAP

This section is short on purpose. The SAP HANA Cloud vector engine topic, later in Unit 7, covers SAP's offering in depth.

  • Vector engine in SQL. As of October 2026, SAP Learning documents the REAL_VECTOR type, TO_REAL_VECTOR, COSINE_SIMILARITY and L2DISTANCE. Your hana_hello.py uses all but the last.
  • In-database embeddings. SAP HANA Cloud can compute embeddings itself. The LangChain integration says this needs NLP enabled on the instance. Check before you plan on it; this course doesn't need it.
  • LangChain integration. SAP maintains langchain-hana, a vector store class for SAP HANA Cloud. It reads the same four HANA_DB_ settings you just saved. RAG topics later in the unit may use it.
  • Trial vs. free tier vs. paid. Your BTP trial is for one learner. The free tier lives in productive (Pay-As-You-Go or CPEA) accounts; SAP's tutorial lists 30 GB of memory, 2 vCPUs and 120 GB of storage for it. Production uses paid plans; check your contract.
Need Use Why
Try an idea on made-up data today Chroma on your laptop No account, no network, rebuilt in seconds
Learn SAP's vector SQL SAP HANA Cloud in your BTP trial Same engine as production, free for learning
A shared team sandbox SAP HANA Cloud free tier in a productive account Team access, no personal trials
Real data, real users Paid SAP HANA Cloud, with authorizations designed in Security, operations and support

Production concerns

  • Credentials. Keep database passwords in .env locally and in a secret store or service binding in the cloud. Never commit them. Rotate any password that leaks.
  • Least privilege. Applications connect as a dedicated user with only the rights they need, never as DBADMIN. Unit 11 covers agent permissions and SAP authorizations.
  • Network. "Allow all IP addresses" is for trials. Company instances allow only known addresses or connect through SAP BTP.
  • Data. Load only approved data into any vector store. A vector store copies text out of its source system, and with it the duty to protect it.
  • Operations. Trials stop nightly and can be removed. Keep table definitions and loading code in Git so any store can be rebuilt.

Pitfalls

  • Mixing embedding models. Vectors from two models can't be compared. The script keeps one Chroma collection per model for that reason.
  • Forgetting the daily start. Most "it worked yesterday" problems with the trial are a stopped instance.
  • Pasting the port into the address. HANA_DB_ADDRESS takes the host name only.
  • Committing the store or the password. Check git status before every commit.
  • Trusting the toy embeddings. --offline is for testing the plumbing. It matches words, not meaning.

Exercise

Extend the local store with your own notes. The result is the seed of the knowledge base that RAG fundamentals builds on.

  1. Open unit07/local_vector_store.py.

  2. Add two notes at the end of NOTES, for example ("n9", "procure-to-pay", "Goods received but no invoice yet means the GR/IR account shows an open balance.") and one order-to-cash note of your own. Use made-up wording, not text from a real system.

  3. Save, then rebuild the store:

    python unit07/local_vector_store.py "goods arrived but the supplier has not billed us" --reset

    Add --offline if the model isn't available.

  4. Run the same question with --process order-to-cash. Note how the filter changes the results.

  5. Commit: git add unit07/local_vector_store.py and git commit -m "Unit 7: add my own help notes".

Done when your new procure-to-pay note appears in the top 3 for the question in step 3 without a filter, and doesn't appear with --process order-to-cash.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the course pass its own embeddings to Chroma instead of letting Chroma compute them?

    Answer: B. The course uses Sentence Transformers for both stores, so results stay comparable and later topics can swap stores. Chroma can embed for you, but then the two stores could differ.
  2. 2Chroma returns a distance of 0.12 for a note in a cosine collection. What is the cosine similarity?

    Answer: C. Chroma's documentation defines cosine distance as 1 minus cosine similarity. The script prints 1 - distance for that reason.
  3. 3Which SQL ranks rows in SAP HANA Cloud by how alike their vectors are to a question?

    Answer: D. Higher cosine similarity means more alike, so you sort it descending. A larger L2 distance means less alike, so sorting it descending would put the worst matches first.
  4. 4Yesterday hana_hello.py worked. This morning it times out. What do you check first?

    Answer: A. SAP's tutorials say trial instances shut down at the end of each day. Start it in SAP HANA Cloud Central and wait until it is running.
  5. 5Your connection fails with "Cannot resolve host name ...:443". What is the likely fix?

    Answer: C. The SQL endpoint is shown as host plus port, but the driver takes them separately. SAP HANA Cloud uses port 443 with encryption on by default, so neither needs changing.
  6. 6A colleague wants the course's DBADMIN settings for a shared demo app. What do you say?

    Answer: B. DBADMIN controls the whole database, which is acceptable only for a personal trial with made-up data. Applications connect as a dedicated user with only the rights they need.
  7. 7Why do the toy --offline embeddings rank a procure-to-pay note first for a delivery question?

    Answer: D. The toy method hashes words, so notes that share common words score high whatever they mean. A real embedding model compares meaning, which is the point of the exercise.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in