Orchestrate

Embeddings and semantic similarity

How embeddings turn text into numbers that capture meaning, how similarity between them is measured, and how to compare SAP-style texts by meaning yourself.

Updated Sep 30, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An embedding is a list of numbers that stands for the meaning of a piece of text. A small AI model reads a sentence and returns, say, 384 numbers. Texts that mean similar things get similar lists of numbers.

Think of it as a map. Every text gets a location. "Customer over credit limit, order on hold" and "credit check failed for the new distributor" land close together. "Stock runs out before the next delivery" lands somewhere else, near other planning problems.

Semantic similarity is how close two texts are on that map. The computer measures the angle between two locations and turns it into a score. A score near 1 means "about the same thing". A score near 0 means "unrelated".

This beats keyword matching in one important way. Keyword search needs the same words. Embeddings can match "vendor bill" to "supplier invoice" even though the two share no word at all.

Embeddings don't answer questions and don't write text. They find, group and compare. They are the "retrieval" step in many enterprise AI assistants, including ones that look things up in SAP data before answering.

Why it matters to the business

A lot of SAP work is text that people type in their own words: notes on sales orders, reasons in tickets, material descriptions, supplier emails. Keyword search and exact matching miss most of it.

Take the running examples of this course. A service desk gets messages about blocked sales orders, three-way match exceptions on supplier invoices and MRP exceptions in planning. People write "finance has to release it", "vendor bill parked because of a variance" or "stock runs out before the next delivery". None of these uses the words a routing rule expects. With embeddings, each message can be compared with a short description of each exception type and sent to the right team.

The same idea covers other everyday problems:

  • Duplicate master data. "Hex bolt M8x40 zinc plated" and "Bolt, hexagon head, M8 x 40, galvanised" are the same part. Duplicates split stock and spend across two material numbers.
  • Finding the right document. A buyer asks about a contract clause; embeddings find the passages that talk about it, whatever the wording.
  • Grounding an AI assistant. Before a language model answers, the system finds the most relevant records and passages by embedding similarity. Unit 7 builds this.

The value is less manual sorting and fewer missed matches. The risk is a confident match that is wrong. In the hands-on part, "Hex bolt M8x40" and "Hex bolt M10x40" come out as the most similar pair of all. They are different parts. Similar text is not the same item, so a person or a rule must check matches that trigger an action.

Cost is usually small. A small open model runs on a laptop for free. Cloud embedding models typically charge by the amount of text, and each text is embedded once and then stored.

How SAP does it

As of September 2026, SAP offers the pieces of an embedding workflow in three places. Unit 5 and Unit 7 set them up.

  • SAP HANA Cloud vector engine. Generally available since the April 2024 release, it stores embeddings in a database column next to ordinary business tables. SQL functions compare them, for example COSINE_SIMILARITY. The point, as SAP describes it, is to keep vectors in the same database as relational, graph, spatial and JSON data.
  • Embeddings computed inside SAP HANA Cloud. A SQL function, VECTOR_EMBEDDING, can create embeddings in the database with an SAP-provided model. The LangChain integration for SAP HANA Cloud documents this for instances with the natural language processing (NLP) feature enabled.
  • Generative AI hub in SAP AI Core. It gives access to embedding models from model providers, such as OpenAI's text-embedding-3-large, through SAP's orchestration service. That service can mask personal data before the text reaches the embedding model.

For where these services sit among SAP's other AI offerings, see the SAP Business AI landscape.

Keyword or embedding? A decision guide

Your situation Better fit Why
Exact identifiers: order numbers, material numbers, tax IDs Exact match or keyword search An ID either matches or it doesn't. Embeddings blur exact values
Free text in people's own words Embeddings Matches meaning, not spelling
Short query against long documents Embeddings, plus a re-ranking step Different lengths and wording; a second model checks the top hits
Sizes, amounts, dates must match exactly Embeddings to find candidates, rules to confirm Embeddings rate "M8" and "M10" as nearly the same
Many languages in one system A multilingual embedding model Some models map several languages onto one map; check the model's documentation
An action follows automatically, such as merging two materials Embeddings plus human review A wrong match changes business data

Most real systems combine both: keywords for exact terms, embeddings for meaning. Unit 7 calls this hybrid search.

Questions to ask

  1. Which embedding model is used, and where does it run? On our laptops, inside SAP HANA Cloud, or at an external provider through the generative AI hub?
  2. Does our text leave our systems to be embedded? If yes, is personal data masked first?
  3. How was the match quality tested? On our own texts, with a list of pairs someone checked by hand? What share of the top matches were right?
  4. What score counts as a match, and who chose it? Scores differ between models, so a threshold from one model means nothing for another.
  5. What happens when we change the model? Every stored embedding must be recomputed. Who owns that, and what does it cost?
  6. Who can search the embeddings? Do search results respect the same authorizations as the source records?
  7. Where does a person check before an action? For example, before merging two materials or routing an escalation.

Common misconceptions

  • "An embedding is a compressed copy of the text." It is a summary of meaning for comparison. You can't read the original back from it. Still, it can reveal what the text was about, so protect it like the source.
  • "A high score means the two items are the same." It means the texts are about similar things. Two bolts of different sizes score very high.
  • "Each of the 384 numbers stands for a feature, like 'urgency'." Individual dimensions rarely mean anything a person can name. Only positions relative to each other matter.
  • "Scores from different models are comparable." They aren't. Each model has its own map. Vectors from two models can't even be compared with each other.
  • "Embeddings need a large language model." Small, free models do this well. The model used in the hands-on part runs on a laptop.
  • "Embeddings replace keyword search." They complement it. Exact IDs and codes still need exact matching.

Key terms

  • Embedding: a list of numbers (a vector) that represents the meaning of a text, image or record.
  • Vector: an ordered list of numbers. An embedding is a vector.
  • Dimension: one position in the vector. A 384-dimensional embedding has 384 numbers.
  • Embedding model: the neural network that turns text into embeddings.
  • Semantic similarity: closeness in meaning, measured between two embeddings.
  • Cosine similarity: the usual score, based on the angle between two vectors. 1 means same direction; 0 means unrelated.
  • Semantic search: finding the stored texts whose embeddings are closest to the embedding of a question.
  • Vector engine / vector database: a store that keeps embeddings and searches them quickly.
  • Threshold: the score above which you treat two items as a match.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does an embedding model give you for a sentence?

    Answer: B. An embedding is a vector that stands for the meaning of the text. Its value lies in comparison: similar texts get nearby vectors. It doesn't answer anything, and you can't read the original text back from it.
  2. 2A service desk gets messages like "vendor bill parked because of a variance". Why do embeddings route these better than keyword rules?

    Answer: C. Keyword rules only fire on the words they expect. Embeddings place texts by meaning, so a message and a category description can match with no word in common.
  3. 3A duplicate-finder rates "Hex bolt M8x40" and "Hex bolt M10x40" as the closest pair. What should you conclude?

    Answer: D. Embeddings capture that both are zinc-plated hex bolts, and barely notice the size. Sizes, amounts and IDs must be checked exactly before any merge. The model is doing what it was trained for.
  4. 4Your team wants to switch to a newer embedding model. What must be planned?

    Answer: A. Each model has its own map, so vectors from two models can't be compared. Stored embeddings must be recomputed, and a threshold chosen for the old model says nothing about the new one.
  5. 5Which SAP offering stores embeddings next to business tables and compares them with SQL?

    Answer: C. The vector engine, generally available since April 2024, adds a vector column type and similarity functions such as COSINE_SIMILARITY to SAP HANA Cloud, next to relational, graph, spatial and JSON data.
  6. 6Text is sent to an external embedding model through SAP's generative AI hub. What should you ask first?

    Answer: B. Embedding means sending the text to wherever the model runs. SAP's orchestration service can mask personal data before embedding, so ask whether that is switched on and what else protects the data.
  7. 7Which task is the worst fit for embeddings on their own?

    Answer: D. An order number either matches or it doesn't. Embeddings blur exact values, so exact match or keyword search is the right tool. The other three tasks are about meaning in people's own words.
Deep layer · 40 min read

Mental model

An embedding model is a function that places any text at a point in a space with a few hundred dimensions. It was trained so that texts with the same meaning land in the same direction from the origin. Semantic similarity is then just geometry: the smaller the angle between two arrows, the closer the meaning.

Everything else in this topic follows from that picture:

  • Search is "find the arrows closest to the question's arrow".
  • Routing is "which category description points the same way as this message".
  • Deduplication is "which pairs of arrows almost overlap".
  • What "close" means depends entirely on what the model was trained to pull together. That is why two bolts of different sizes can overlap.

In Neural networks from scratch, the hidden layer built its own intermediate signals. An embedding is the same thing, taken from a much larger network: the network's internal description of a text, used on its own.

How it works

From text to one vector

A sentence embedding model such as all-MiniLM-L6-v2 does four things when you call it.

flowchart LR
  T[Text] --> K[Split into<br/>word pieces]
  K --> N[Transformer network<br/>one vector per piece]
  N --> P[Mean pooling<br/>average the pieces]
  P --> L[Scale to length 1]
  L --> E[Embedding<br/>384 numbers]
  1. Tokenize. The text is split into word pieces (tokens). Common words stay whole; rare words are split into parts.
  2. Encode. A transformer network turns each piece into a vector. Each vector depends on the words around it. This is a contextual embedding: "order" in "sales order" and "order" in "out of order" get different vectors. Google's crash course contrasts this with older static embeddings such as word2vec, which gave each word one fixed vector. Unit 4 opens up the transformer.
  3. Pool. The model card for all-MiniLM-L6-v2 says it uses mean pooling: it averages the piece vectors, skipping padding, to get one vector for the whole text.
  4. Normalize. The vector is scaled to length 1. The model card shows this step explicitly. Then only the direction carries meaning.

The model card states the limits: 384 dimensions, and "input text longer than 256 word pieces is truncated". Anything after that point is silently ignored. The card describes the model as an encoder for sentences and short paragraphs. Long documents must be cut into chunks first; Unit 7 covers chunking.

How the model learned what "similar" means

The model wasn't told what words mean. According to its model card, it was trained on over 1 billion sentence pairs with a contrastive objective. Given one sentence, the model had to pick its true partner from a batch of randomly chosen other sentences. The pairs were things like question and answer, or two phrasings of the same question.

Each training step pulls true pairs closer and pushes the random others apart. This is the embedding-layer training from Google's crash course, at scale: weights are adjusted "so that embedding vectors for similar examples are closer to each other".

This tells you what the model will and won't treat as close:

  • Close: paraphrases, a question and its answer, the same topic in different words.
  • Not reliably separated: texts on the same topic that differ in one number, one negation or one unit. "M8" and "M10" are rare word pieces in a long shared context.
  • Weak: in-house jargon and abbreviations that never appeared in training, and languages the training data barely covered.

Measuring closeness

Three measures are common. The SAP HANA Cloud vector engine offers the first and the third as SQL functions.

Measure Formula (in words) Range Closer means
Cosine similarity Dot product divided by both lengths -1 to 1 Higher
Dot product Multiply matching positions, add up Any number Higher
Euclidean (L2) distance Straight-line distance between the points 0 or more Lower

For vectors of length 1, the three agree. Cosine similarity equals the dot product, and OpenAI's embeddings guide notes that ranking by Euclidean distance gives identical results. That is why most pipelines normalize once, at embedding time, and then use the cheaper dot product. The Sentence Transformers semantic search guide makes the same point.

A worked example, the one SAP Learning uses for COSINE_SIMILARITY: the vectors [1, 0, 0] and [1, 0, 1].

  • Dot product: 1×1 + 0×0 + 0×1 = 1.
  • Lengths: 1 and √2 ≈ 1.414.
  • Cosine: 1 / 1.414 = 0.7071, the cosine of 45 degrees.

Multiply the second vector by 5 and the cosine stays 0.7071, because the direction didn't change. The Euclidean distance changes a lot. That is why cosine is the default for text.

Symmetric and asymmetric comparison

The Sentence Transformers documentation separates two cases:

  • Symmetric: both sides look alike. Two material descriptions, or two ticket texts. Deduplication and routing to short category descriptions are mostly symmetric.
  • Asymmetric: a short question against longer passages, such as "what is the penalty for late delivery?" against contract paragraphs.

Some models are tuned for one case. Some expect a prefix or prompt that says whether a text is a query or a document; Sentence Transformers supports this with a prompt_name option. all-MiniLM-L6-v2 needs no prompt. Always check the model card.

Bi-encoders and cross-encoders

The model in this topic is a bi-encoder: it embeds each text on its own, so you can embed a million descriptions once, store them, and compare new texts in milliseconds.

A cross-encoder reads two texts together and returns one relevance score. It is more accurate and much slower, because nothing can be computed ahead of time. The Sentence Transformers documentation recommends combining them: the bi-encoder retrieves the top candidates, and a cross-encoder re-ranks them. Unit 7 builds this.

Scale: from comparing everything to indexes

Comparing a question with every stored vector is called exhaustive or brute-force search. It is fine for thousands of vectors and exact. For millions, vector stores use approximate nearest neighbour (ANN) indexes that skip most comparisons and accept a small chance of missing a result. HNSW is a common one; the LangChain integration for SAP HANA Cloud creates HNSW indexes on the vector column. Unit 7 compares exact and approximate search on SAP HANA Cloud.

Build it yourself: compare SAP-style texts by meaning

You will write one script that computes cosine similarity by hand, then compares nine made-up service-desk notes with a question two ways: by shared words and by meaning. It then routes each note to one of three exception types: credit block, three-way match or MRP exception. You will see where word matching fails and embeddings don't.

Before you start: complete Set up your computer for this course and Set up for Unit 3. They give you the orchestrate-course folder with its .venv, numpy, PyTorch, Sentence Transformers and the unit03 folder. Step 6 of the Unit 3 setup also downloads the embedding model used here. This walkthrough doesn't repeat those steps.

flowchart LR
  Q[Question] --> W[Word counts]
  N[9 notes] --> W
  Q --> E[Embedding model<br/>on your laptop]
  N --> E
  C[3 category<br/>descriptions] --> W
  C --> E
  W --> R1[Ranking and routing<br/>by shared words]
  E --> R2[Ranking and routing<br/>by meaning]

What you need

  • The course folder, .venv and unit03 folder from the setup topics.
  • About 40 minutes. No accounts, no API keys, no cost.
  • The embedding model from the Unit 3 setup, downloaded once from Hugging Face. If your network blocks the download, the --offline option runs everything except the embedding part.
  • All texts are made up. The model runs on your computer, so nothing is sent anywhere.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 3 folder:

    cd unit03

Step 2: Save the script

  1. In VS Code's file list, right-click unit03, choose New File and name it similarity.py.
  2. Paste the code below and save with File > Save.
"""Unit 3: compare SAP-style texts by meaning (embeddings) and by shared words.

How to run (from the unit03 folder, with the course .venv turned on):
    python similarity.py                      # embedding model (downloads once) and word counts
    python similarity.py --offline            # no download, no internet: word counts only
    python similarity.py --query "Why is my invoice not paid?"   # your own question
    python similarity.py --save               # also write note_vectors.npz

All texts are made up. Nothing is sent anywhere; the model runs on your computer.
"""
import argparse
import re
from pathlib import Path

import numpy as np

HERE = Path(__file__).parent
MODEL = "sentence-transformers/all-MiniLM-L6-v2"

NOTES = [  # made-up notes, shaped like what people write about SAP documents
    "Customer is over the credit limit, order on hold",
    "Finance has to release it before we can deliver",
    "Credit check failed for the new distributor",
    "Invoice price is higher than the purchase order, payment blocked",
    "Supplier billed 120 units but we only received 100",
    "Vendor bill parked because of a variance",
    "Planned order is late because a component is missing",
    "Requirement date moved earlier, reschedule the receipt",
    "Stock runs out before the next delivery arrives",
]
CATEGORIES = {  # one plain description per exception type
    "credit block": "Sales order blocked by the customer credit check",
    "three-way match": "Supplier invoice blocked because it does not match the purchase order or goods receipt",
    "MRP exception": "Material planning exception: shortage, late order or date change",
}
STOP = {"the", "a", "an", "is", "are", "be", "to", "of", "and", "or", "for", "it", "we", "but",
        "by", "on", "in", "this", "that", "has", "have", "can", "because", "before", "than", "does", "not"}


def cosine(a: np.ndarray, b: np.ndarray) -> float:
    """Cosine similarity: the dot product of the two vectors after scaling each to length 1."""
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))


def words(text: str) -> list[str]:
    return [w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP]


def word_vectors(texts: list[str]) -> np.ndarray:
    """One column per word in the texts; each value counts how often the word appears."""
    vocab = sorted({w for t in texts for w in words(t)})
    index = {w: i for i, w in enumerate(vocab)}
    vectors = np.zeros((len(texts), len(vocab)))
    for row, text in enumerate(texts):
        for w in words(text):
            vectors[row, index[w]] += 1
    return vectors


def embed(texts: list[str], model_name: str) -> np.ndarray:
    """Turn each text into one embedding with a Sentence Transformers model, scaled to length 1."""
    try:
        from sentence_transformers import SentenceTransformer
    except ImportError:
        raise SystemExit("sentence-transformers is not installed. See Set up for Unit 3, or use --offline.")
    print(f"Loading {model_name} (downloads once, then uses the copy on disk)...")
    model = SentenceTransformer(model_name)
    return model.encode(texts, normalize_embeddings=True)


def by_hand() -> None:
    print("1. Cosine similarity and distance by hand")
    a, b = np.array([1.0, 0.0, 0.0]), np.array([1.0, 0.0, 1.0])
    print(f"   cosine([1,0,0], [1,0,1])   = {cosine(a, b):.4f}   (45 degrees apart)")
    print(f"   cosine([1,0,0], [5,0,5])   = {cosine(a, 5 * b):.4f}   (longer arrow, same direction)")
    c, d = np.array([6.0, 3.0, 5.0]), np.array([6.0, 3.0, -5.0])
    print(f"   distance([6,3,5], [6,3,-5]) = {np.linalg.norm(c - d):.4f}")


def rank(title: str, query_vec: np.ndarray, note_vecs: np.ndarray) -> None:
    scores = [cosine(query_vec, v) if np.any(v) and np.any(query_vec) else 0.0 for v in note_vecs]
    print(f"   {title}")
    for score, note in sorted(zip(scores, NOTES), reverse=True)[:4]:
        print(f"     {score:5.2f}  {note}")


def route(title: str, note_vecs: np.ndarray, cat_vecs: np.ndarray) -> None:
    names = list(CATEGORIES)
    print(f"   {title}")
    for note, v in zip(NOTES, note_vecs):
        scores = [cosine(v, c) if np.any(v) and np.any(c) else 0.0 for c in cat_vecs]
        best = int(np.argmax(scores))
        label = names[best] if scores[best] > 0 else "(no shared words)"
        print(f"     {label:17s} {scores[best]:4.2f}  {note}")


def main() -> None:
    parser = argparse.ArgumentParser(description="Compare SAP-style texts by meaning and by shared words.")
    parser.add_argument("--query", default="Why is the client's delivery stuck?", help="question to rank notes against")
    parser.add_argument("--offline", action="store_true", help="skip the embedding model; word counts only")
    parser.add_argument("--model", default=MODEL, help="Sentence Transformers model name or local folder")
    parser.add_argument("--save", action="store_true", help="write note_vectors.npz with the note embeddings")
    args = parser.parse_args()

    by_hand()

    all_texts = [args.query] + NOTES + list(CATEGORIES.values())
    wv = word_vectors(all_texts)                  # word counts share one vocabulary across all texts
    wq, wn, wc = wv[0], wv[1:1 + len(NOTES)], wv[1 + len(NOTES):]

    print(f"\n2. Notes closest to: '{args.query}'")
    rank("By shared words (word counts):", wq, wn)
    if not args.offline:
        ev = embed(all_texts, args.model)
        print(f"   Each text became {ev.shape[1]} numbers.")
        eq, en, ec = ev[0], ev[1:1 + len(NOTES)], ev[1 + len(NOTES):]
        rank("By meaning (embeddings):", eq, en)

    print("\n3. Route each note to the closest exception type")
    route("By shared words:", wn, wc)
    if not args.offline:
        route("By meaning:", en, ec)

    if args.save:
        if args.offline:
            print("\n--save needs embeddings. Run without --offline.")
        else:
            np.savez(HERE / "note_vectors.npz", texts=np.array(NOTES), vectors=en, model=args.model)
            print(f"\nSaved {len(NOTES)} note embeddings to note_vectors.npz")
    if args.offline:
        print("\nOffline run: skipped the embedding model. Run without --offline to compare by meaning.")


if __name__ == "__main__":
    main()

Step 3: Run it offline first

Start without the model, so you see the keyword approach on its own.

  1. In the terminal, inside unit03, run:

    python similarity.py --offline

What success looks like:

1. Cosine similarity and distance by hand
   cosine([1,0,0], [1,0,1])   = 0.7071   (45 degrees apart)
   cosine([1,0,0], [5,0,5])   = 0.7071   (longer arrow, same direction)
   distance([6,3,5], [6,3,-5]) = 10.0000

2. Notes closest to: 'Why is the client's delivery stuck?'
   By shared words (word counts):
      0.18  Stock runs out before the next delivery arrives
      0.00  Vendor bill parked because of a variance
      0.00  Supplier billed 120 units but we only received 100
      0.00  Requirement date moved earlier, reschedule the receipt

3. Route each note to the closest exception type
   By shared words:
     credit block      0.50  Customer is over the credit limit, order on hold
     (no shared words) 0.00  Finance has to release it before we can deliver
     credit block      0.37  Credit check failed for the new distributor
     three-way match   0.53  Invoice price is higher than the purchase order, payment blocked
     three-way match   0.13  Supplier billed 120 units but we only received 100
     (no shared words) 0.00  Vendor bill parked because of a variance
     MRP exception     0.32  Planned order is late because a component is missing
     three-way match   0.14  Requirement date moved earlier, reschedule the receipt
     (no shared words) 0.00  Stock runs out before the next delivery arrives

Offline run: skipped the embedding model. Run without --offline to compare by meaning.

Step 4: Read the offline results

  • Section 1 is the SAP Learning example for COSINE_SIMILARITY, computed by you: 0.7071. Stretching the second arrow to [5,0,5] doesn't change the cosine. The distance example is SAP Learning's L2DISTANCE example: the two points differ only in the last number, by 10.
  • Section 2 shows keyword matching failing. The question is about a customer whose order won't ship, a credit block. The only note sharing a word ("delivery") is a planning note. Every other note scores 0.00, so their order below the first line means nothing.
  • Section 3 shows the same weakness in routing. Three notes share no word with any category. "Requirement date moved earlier" goes to three-way match only because of the word "receipt". Keyword routing is right when the writer happens to use the expected words.

Step 5: Run it with the embedding model

  1. Run:

    python similarity.py
  2. The first line says Loading sentence-transformers/all-MiniLM-L6-v2 .... If you ran Step 6 of the Unit 3 setup, the model is already on disk and loads in a few seconds. Otherwise it downloads once. No account or key is needed.

What success looks like (the offline lines from Step 3 appear too; only the new lines are shown):

Loading sentence-transformers/all-MiniLM-L6-v2 (downloads once, then uses the copy on disk)...
   Each text became 384 numbers.
   By meaning (embeddings):
     0.xx  ...
     0.xx  ...
     0.xx  ...
     0.xx  ...
...
   By meaning:
     credit block      0.xx  Customer is over the credit limit, order on hold
     ...

Your lines show real scores in place of 0.xx. Now every note has a non-zero score, because every text has a position on the map. Look for three things:

  1. Do the credit-block notes rise to the top of the ranking for the delivery question, above the planning note that shares the word "delivery"?
  2. Do "Vendor bill parked because of a variance" and "Finance has to release it" now get a category, and is it the right one?
  3. Which notes are still routed wrongly, and do their scores look lower than the correct ones?

Write the answers down; the exercise uses them. A small model won't get every note right, and that is the point of Step 6.

Step 6: Try your own questions

  1. Ask about a supplier invoice:

    python similarity.py --query "Why hasn't the supplier been paid yet?"
  2. Ask something unrelated to all nine notes:

    python similarity.py --query "Where is the canteen menu for next week?"
  3. Compare the top score in each case. An unrelated question still returns four "closest" notes, only with lower scores. Search always returns something. Deciding when the best match is not good enough is your job, with a threshold you choose from examples.

  4. Try a negation:

    python similarity.py --query "The order is not blocked for credit"

    Check whether the credit-block notes still come first. Sentence embeddings often don't separate a statement from its negation well, because both are about the same topic.

Step 7: Save the note embeddings

  1. Run:

    python similarity.py --save
  2. The last line should read:

    Saved 9 note embeddings to note_vectors.npz

The file holds the nine texts, their 384-number vectors and the model name. Storing the model name with the vectors matters: vectors from another model can't be compared with these.

Step 8: Save your work in Git

  1. Run:

    git add unit03/similarity.py
    git commit -m "Compare SAP-style notes by meaning and by shared words"

    If Git says the path doesn't exist, you are inside unit03; run git add similarity.py instead.

note_vectors.npz is generated output, so leave it out of Git. Anyone can recreate it with --save.

What each part of the script does

Part What it does
NOTES, CATEGORIES Nine made-up service-desk notes and one plain description per exception type
cosine Dot product divided by both lengths: the angle-based score from -1 to 1
words, STOP Lowercase, split into words, drop very common words such as "the"
word_vectors Keyword baseline: one column per word, counting how often it appears
embed Loads the Sentence Transformers model and returns one length-1 vector per text
by_hand Recomputes SAP Learning's cosine and L2 distance examples with numpy
rank Scores the question against every note and prints the top four
route Picks the category whose description is closest to each note
--offline Skips the model: no download, no internet
--query, --model Your own question, and another model name or a local model folder
--save Writes note_vectors.npz with texts, vectors and model name

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of the Unit 1 setup, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'numpy' The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
sentence-transformers is not installed Unit 3's library isn't in this .venv Follow Step 3 of Set up for Unit 3. Meanwhile, --offline works
ProxyError, 403 Forbidden, SSLError or ConnectionError while loading the model Your network or company proxy blocks huggingface.co Try another network once; after one successful download, the model works offline. Ask IT to allow huggingface.co. Meanwhile, use --offline
OSError: ... is not a local folder and is not a valid model identifier The model name is misspelled, or the download never finished Run without --model to use the default name
Asked for an API key or a Hugging Face token This model is public and needs neither Check that you are running similarity.py and haven't changed --model
can't open file ... similarity.py The terminal isn't in unit03, or the file has another name Run cd unit03 from the course folder, and check the file name
Scores differ from a classmate's Different library or model versions Compare the order of results, not the decimals

The SAP way

This section maps the build onto SAP's services, as of September 2026. Setting them up needs access you don't have yet; Unit 5 covers the generative AI hub and Unit 7 covers SAP HANA Cloud and its vector engine.

Storing and comparing: the SAP HANA Cloud vector engine

SAP announced the vector engine as generally available with the SAP HANA Cloud release of April 2024. SAP Learning's course describes the building blocks:

Your script SAP HANA Cloud vector engine
A numpy array of 384 numbers A column of type REAL_VECTOR: single-precision numbers, 1 to 65,000 dimensions
cosine(a, b) COSINE_SIMILARITY(a, b), returning -1 to 1
np.linalg.norm(c - d) L2DISTANCE(c, d), returning 0 or more; lower means closer
np.array([1.0, 0.0, 0.0]) TO_REAL_VECTOR('[1, 0, 0]')
note_vectors.npz A table with a text column and a vector column

SAP Learning's own examples look like this (sketch; it needs an SAP HANA Cloud instance):

SELECT COSINE_SIMILARITY(TO_REAL_VECTOR('[1, 0, 0]'), TO_REAL_VECTOR('[1, 0, 1]')) FROM DUMMY;

CREATE TABLE "VECTORTAB" ("ID" BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
                          "TEXT" NCLOB, "VECTOR" REAL_VECTOR);

The first query returns the 0.7071 your script printed. SAP Learning also notes that vectors have no order, so you can't sort or group by a vector column. You sort by the similarity score instead.

The LangChain integration for SAP HANA Cloud (langchain-hana) adds two points worth knowing now. Its vector store can use REAL_VECTOR or a HALF_VECTOR column. It uses cosine similarity by default, with Euclidean distance as an option, and can build an HNSW index for fast approximate search.

Creating embeddings inside the database

SAP HANA Cloud has a SQL function, VECTOR_EMBEDDING, that computes embeddings in the database with an SAP-provided model. The LangChain integration's HanaInternalEmbeddings class calls it, with the model ID SAP_NEB.20240715 in its example, and notes that it needs the NLP feature enabled on the instance. Check SAP's documentation for the current model IDs, languages and limits before you design around it; Unit 7 does this in detail.

The advantage is that text never leaves the database to be embedded. The trade-off is that you use the models SAP provides there.

Creating embeddings through the generative AI hub

The generative AI hub in SAP AI Core offers embedding models from model providers. The SAP Cloud SDK for AI (Python) shows an embeddings endpoint in the orchestration service, with text-embedding-3-large in its example. Its options include the number of dimensions and normalization, and the data masking module can anonymize personal data before the text is embedded.

For scale: OpenAI's guide lists 3,072 dimensions by default for text-embedding-3-large and an input limit of 8,192 tokens, against 384 dimensions and 256 word pieces for the laptop model. Larger is not automatically better for your texts; test on your own examples.

Licensing notes

all-MiniLM-L6-v2 is published on Hugging Face; check the licence on its model card before commercial use. SAP HANA Cloud and SAP AI Core are paid SAP BTP services, and what you pay for them and for provider models in the generative AI hub depends on your contract. Unit 5 covers the commercial side.

Build vs. SAP

Situation Use Why
Learning, prototyping, a few thousand texts Sentence Transformers on a laptop Free, private, every number visible
The data already lives in SAP HANA Cloud and must stay there VECTOR_EMBEDDING plus the vector engine No data movement; SQL next to business tables
You need a large or multilingual provider model Embeddings through the generative AI hub Managed access, data masking in the orchestration service
Millions of vectors searched by many users SAP HANA Cloud vector engine with an HNSW index Built for approximate search at scale, next to the source data
Exact IDs, codes and numbers SQL and keyword search, not embeddings Exact values need exact matching
Production matching that triggers actions Embeddings to find candidates, rules and people to confirm Similar text is not the same record

Production concerns

  • One model per vector store. Store the model name and version with every vector, as --save does. When the model changes, re-embed everything; mixing models gives meaningless scores.
  • Evaluate on your own texts. Build a list of pairs or queries with the right answers marked by a business expert. Measure how often the right item is in the top results, the retrieval version of the precision and recall from Classification and the metrics that matter. Unit 8 builds an evaluation harness.
  • Choose thresholds from data. Plot scores for true matches and non-matches, then pick the cut-off by the cost of each mistake. Re-check it for every new model.
  • Data protection. Embedding with an external model sends the text out. Use in-database embedding or a local model for sensitive text, or mask personal data first. Treat stored vectors as sensitive as their source: they are built to reveal what the text was about.
  • Authorizations. A vector store that holds sales notes from all company codes must filter results by what the user may see in SAP. The access rules from Calling your first SAP API still apply; Unit 7 covers retrieval that respects SAP authorizations.
  • Cost and speed. Embed each text once, when it is created or changed, not at every search. Batch the calls. Normalize once, then use the dot product.
  • Clean core. Read texts through released APIs and keep the vector store side by side, in SAP HANA Cloud or your own service. Don't add embedding columns to standard S/4HANA tables.

Pitfalls

  • Trusting a match on numbers. "M8" and "M10", "100 pcs" and "1,000 pcs" look nearly the same to a sentence model. Confirm sizes, amounts and IDs with rules.
  • Ignoring truncation. Text beyond the model's limit, 256 word pieces for all-MiniLM-L6-v2, is dropped without a warning. Chunk long documents.
  • Comparing vectors from two models. It runs without an error and returns nonsense.
  • No threshold. Search always returns a "closest" result, even for an unrelated question, as Step 6 showed.
  • Forgetting normalization. If vectors aren't length 1, the dot product favours long vectors. Normalize, or use cosine.
  • Embedding identifiers. Order and material numbers carry no meaning for the model. Keep them in normal columns and filter on them.
  • Negation and small words. "Blocked" and "not blocked" often score as similar. Don't use similarity alone to decide a status.
  • Judging a model by its general benchmark. In-house abbreviations and product codes may be unknown to it. Test on your texts.

Exercise: find duplicate material descriptions

You will rank every pair of twelve made-up material descriptions by similarity, save the top pairs to a CSV file and mark each one as a true or false duplicate by hand. This is the first step of the master data work in the next topic of this unit, semantic search over SAP master data.

  1. In VS Code, create unit03/find_duplicates.py in the same folder as similarity.py, paste the code below and save. It reuses functions from similarity.py, so both files must be in unit03.

    """Find possible duplicate material descriptions with embeddings.
    
    How to run (from the unit03 folder, with the course .venv turned on):
        python find_duplicates.py              # embedding model (downloaded by similarity.py)
        python find_duplicates.py --offline    # no model: shared words only
        python find_duplicates.py --top 10     # show more pairs
    
    Writes duplicate_candidates.csv. All descriptions are made up.
    """
    import argparse
    import csv
    from itertools import combinations
    from pathlib import Path
    
    import numpy as np
    
    from similarity import MODEL, cosine, embed, word_vectors
    
    HERE = Path(__file__).parent
    
    MATERIALS = [  # (made-up material number, description)
        ("M-1001", "Hex bolt M8x40 zinc plated"),
        ("M-1002", "Bolt, hexagon head, M8 x 40, galvanised"),
        ("M-1003", "Hex bolt M10x40 zinc plated"),
        ("M-1004", "Nitrile gloves size L, box of 100"),
        ("M-1005", "Gloves nitrile large 100 pcs"),
        ("M-1006", "Deep groove ball bearing 6204 2RS"),
        ("M-1007", "Bearing 6204-2RS sealed"),
        ("M-1008", "Hydraulic oil ISO VG 46, 20 l can"),
        ("M-1009", "Hydraulic fluid VG46 20 litre"),
        ("M-1010", "Safety glasses, clear lens"),
        ("M-1011", "Copy paper A4 80 g, 500 sheets"),
        ("M-1012", "Printer paper DIN A4 white ream"),
    ]
    
    
    def main() -> None:
        parser = argparse.ArgumentParser(description="Rank pairs of material descriptions by similarity.")
        parser.add_argument("--offline", action="store_true", help="use shared words instead of the embedding model")
        parser.add_argument("--model", default=MODEL, help="Sentence Transformers model name or local folder")
        parser.add_argument("--top", type=int, default=8, help="how many of the most similar pairs to show")
        args = parser.parse_args()
    
        texts = [text for _, text in MATERIALS]
        vectors = word_vectors(texts) if args.offline else embed(texts, args.model)
        method = "shared words" if args.offline else "embeddings"
    
        pairs = []
        for i, j in combinations(range(len(MATERIALS)), 2):   # every pair once: 12 items give 66 pairs
            score = cosine(vectors[i], vectors[j]) if np.any(vectors[i]) and np.any(vectors[j]) else 0.0
            pairs.append((score, MATERIALS[i], MATERIALS[j]))
        pairs.sort(key=lambda p: p[0], reverse=True)
    
        print(f"Top {args.top} of {len(pairs)} pairs by {method}:")
        for score, (id_a, text_a), (id_b, text_b) in pairs[:args.top]:
            print(f"  {score:4.2f}  {id_a} {text_a:40s} | {id_b} {text_b}")
    
        out = HERE / "duplicate_candidates.csv"
        with out.open("w", newline="", encoding="utf-8") as f:
            writer = csv.writer(f)
            writer.writerow(["score", "material_a", "text_a", "material_b", "text_b", "method", "same_item"])
            for score, (id_a, text_a), (id_b, text_b) in pairs[:args.top]:
                writer.writerow([f"{score:.2f}", id_a, text_a, id_b, text_b, method, ""])
        print(f"\nWrote {out.name}. Fill in same_item with yes or no for each row.")
    
    
    if __name__ == "__main__":
        main()
  2. Run the keyword baseline first:

    python find_duplicates.py --offline

    You should see:

    Top 8 of 66 pairs by shared words:
      0.80  M-1001 Hex bolt M8x40 zinc plated               | M-1003 Hex bolt M10x40 zinc plated
      0.61  M-1006 Deep groove ball bearing 6204 2RS        | M-1007 Bearing 6204-2RS sealed
      0.55  M-1004 Nitrile gloves size L, box of 100        | M-1005 Gloves nitrile large 100 pcs
      0.34  M-1008 Hydraulic oil ISO VG 46, 20 l can        | M-1009 Hydraulic fluid VG46 20 litre
      0.31  M-1011 Copy paper A4 80 g, 500 sheets           | M-1012 Printer paper DIN A4 white ream
      0.17  M-1001 Hex bolt M8x40 zinc plated               | M-1002 Bolt, hexagon head, M8 x 40, galvanised
      0.17  M-1002 Bolt, hexagon head, M8 x 40, galvanised  | M-1003 Hex bolt M10x40 zinc plated
      0.15  M-1004 Nitrile gloves size L, box of 100        | M-1008 Hydraulic oil ISO VG 46, 20 l can
    
    Wrote duplicate_candidates.csv. Fill in same_item with yes or no for each row.

    The top pair is not a duplicate: M8 and M10 are different bolts. The real duplicate of M-1001, M-1002, is only sixth.

  3. Rename the offline result so it isn't overwritten:

    • Windows (PowerShell):

      Rename-Item duplicate_candidates.csv duplicate_candidates_words.csv
    • macOS / Linux:

      mv duplicate_candidates.csv duplicate_candidates_words.csv
  4. Run it with the embedding model:

    python find_duplicates.py

    You should see Top 8 of 66 pairs by embeddings: followed by eight pairs with scores, and a new duplicate_candidates.csv.

  5. Open both CSV files in VS Code. In the same_item column of each row, type yes if the two descriptions are the same item, or no. The five true duplicate pairs are M-1001/M-1002, M-1004/M-1005, M-1006/M-1007, M-1008/M-1009 and M-1011/M-1012. Save both files.

  6. Create unit03/notes_embeddings.md with three short sections:

    • Routing: from Step 5 of the walkthrough, which notes the embeddings routed correctly that word counts missed, and which ones they still got wrong.
    • Duplicates: for each method, how many of the top 5 pairs are true duplicates (the precision at 5), and where the M8/M10 pair landed.
    • Rule: one sentence describing a rule that would stop the M8/M10 pair from being merged, such as comparing the numbers in both descriptions.
  7. Save your work:

    git add find_duplicates.py duplicate_candidates.csv duplicate_candidates_words.csv notes_embeddings.md
    git commit -m "Rank duplicate material descriptions by words and by embeddings"

Done when: both CSV files have yes or no in every same_item cell; your notes give the precision at 5 for both methods and name where the M8/M10 pair ranked; and git log shows the commit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does all-MiniLM-L6-v2 put "credit check failed" close to "customer over the credit limit"?

    Answer: C. The model card describes contrastive training on over a billion pairs: true partners are pulled together and random others pushed apart. Closeness in its space reflects what it learned to pair, not shared words.
  2. 2In similarity.py, why does embed pass normalize_embeddings=True?

    Answer: A. Scaling to length 1 leaves only the direction. Then the dot product equals the cosine, and Euclidean distance ranks the same way. It does nothing to make two models' vectors comparable.
  3. 3cosine([1,0,0], [5,0,5]) prints 0.7071, the same as for [1,0,1]. What does that show?

    Answer: D. The angle between the arrows stays 45 degrees when you stretch one of them. Euclidean distance would change a lot, which is why cosine is the default for comparing text.
  4. 4In the offline run, "Vendor bill parked because of a variance" gets "(no shared words)". What does that illustrate?

    Answer: B. The note says "vendor bill" and "parked"; the three-way match description says "supplier invoice" and "blocked". With no word in common, the word-count cosine is 0. Embeddings place both by meaning.
  5. 5Your company switches from all-MiniLM-L6-v2 to a 3,072-dimension provider model. What must happen to the vectors already stored?

    Answer: C. Each model has its own space; padding or scaling doesn't translate between them. Mixed vectors return scores that run fine and mean nothing. Store the model name with the vectors, as --save does, so you know when to re-embed.
  6. 6A 30-page supplier contract is embedded with all-MiniLM-L6-v2 as one text. What is the problem?

    Answer: B. The model card says input longer than 256 word pieces is truncated, without an error. Long documents must be split into chunks and each chunk embedded, which Unit 7 covers.
  7. 7Which SAP HANA Cloud statement matches cosine(a, b) in the script?

    Answer: D. COSINE_SIMILARITY is the vector engine's cosine function. L2DISTANCE is lower for closer vectors, TO_REAL_VECTOR only builds a vector, and SAP Learning notes that vectors have no order to sort by.
  8. 8A client wants to embed HR case notes but must not send personal data to an external model. What would you propose?

    Answer: C. External embedding sends the text out. In-database embedding with VECTOR_EMBEDDING or a local model keeps it in, and the orchestration service can mask personal data before embedding. Vectors still reveal what the text was about, so protect them too.
  9. 9The duplicate-finder ranks "Hex bolt M8x40" and "Hex bolt M10x40" first. What would you do before merging materials automatically?

    Answer: D. Sentence embeddings barely separate texts that differ in one number. A rule on sizes and amounts catches this reliably, and human review protects master data. A bigger model or a different threshold may still miss it.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in