How embeddings turn text into numbers that capture meaning, how similarity between them is measured, and how to compare SAP-style texts by meaning yourself.
An embedding is a list of numbers that stands for the meaning of a piece of text. A small AI model reads a sentence and returns, say, 384 numbers. Texts that mean similar things get similar lists of numbers.
Think of it as a map. Every text gets a location. "Customer over credit limit, order on hold" and "credit check failed for the new distributor" land close together. "Stock runs out before the next delivery" lands somewhere else, near other planning problems.
Semantic similarity is how close two texts are on that map. The computer measures the angle between two locations and turns it into a score. A score near 1 means "about the same thing". A score near 0 means "unrelated".
This beats keyword matching in one important way. Keyword search needs the same words. Embeddings can match "vendor bill" to "supplier invoice" even though the two share no word at all.
Embeddings don't answer questions and don't write text. They find, group and compare. They are the "retrieval" step in many enterprise AI assistants, including ones that look things up in SAP data before answering.
A lot of SAP work is text that people type in their own words: notes on sales orders, reasons in tickets, material descriptions, supplier emails. Keyword search and exact matching miss most of it.
Take the running examples of this course. A service desk gets messages about blocked sales orders, three-way match exceptions on supplier invoices and MRP exceptions in planning. People write "finance has to release it", "vendor bill parked because of a variance" or "stock runs out before the next delivery". None of these uses the words a routing rule expects. With embeddings, each message can be compared with a short description of each exception type and sent to the right team.
The same idea covers other everyday problems:
Duplicate master data. "Hex bolt M8x40 zinc plated" and "Bolt, hexagon head, M8 x 40, galvanised" are the same part. Duplicates split stock and spend across two material numbers.
Finding the right document. A buyer asks about a contract clause; embeddings find the passages that talk about it, whatever the wording.
Grounding an AI assistant. Before a language model answers, the system finds the most relevant records and passages by embedding similarity. Unit 7 builds this.
The value is less manual sorting and fewer missed matches. The risk is a confident match that is wrong. In the hands-on part, "Hex bolt M8x40" and "Hex bolt M10x40" come out as the most similar pair of all. They are different parts. Similar text is not the same item, so a person or a rule must check matches that trigger an action.
Cost is usually small. A small open model runs on a laptop for free. Cloud embedding models typically charge by the amount of text, and each text is embedded once and then stored.
As of September 2026, SAP offers the pieces of an embedding workflow in three places. Unit 5 and Unit 7 set them up.
SAP HANA Cloud vector engine. Generally available since the April 2024 release, it stores embeddings in a database column next to ordinary business tables. SQL functions compare them, for example COSINE_SIMILARITY. The point, as SAP describes it, is to keep vectors in the same database as relational, graph, spatial and JSON data.
Embeddings computed inside SAP HANA Cloud. A SQL function, VECTOR_EMBEDDING, can create embeddings in the database with an SAP-provided model. The LangChain integration for SAP HANA Cloud documents this for instances with the natural language processing (NLP) feature enabled.
Generative AI hub in SAP AI Core. It gives access to embedding models from model providers, such as OpenAI's text-embedding-3-large, through SAP's orchestration service. That service can mask personal data before the text reaches the embedding model.
"An embedding is a compressed copy of the text." It is a summary of meaning for comparison. You can't read the original back from it. Still, it can reveal what the text was about, so protect it like the source.
"A high score means the two items are the same." It means the texts are about similar things. Two bolts of different sizes score very high.
"Each of the 384 numbers stands for a feature, like 'urgency'." Individual dimensions rarely mean anything a person can name. Only positions relative to each other matter.
"Scores from different models are comparable." They aren't. Each model has its own map. Vectors from two models can't even be compared with each other.
"Embeddings need a large language model." Small, free models do this well. The model used in the hands-on part runs on a laptop.
"Embeddings replace keyword search." They complement it. Exact IDs and codes still need exact matching.
Pick one answer for each question. The explanation appears after you choose.
1What does an embedding model give you for a sentence?
Answer: B. An embedding is a vector that stands for the meaning of the text. Its value lies in comparison: similar texts get nearby vectors. It doesn't answer anything, and you can't read the original text back from it.
2A service desk gets messages like "vendor bill parked because of a variance". Why do embeddings route these better than keyword rules?
Answer: C. Keyword rules only fire on the words they expect. Embeddings place texts by meaning, so a message and a category description can match with no word in common.
3A duplicate-finder rates "Hex bolt M8x40" and "Hex bolt M10x40" as the closest pair. What should you conclude?
Answer: D. Embeddings capture that both are zinc-plated hex bolts, and barely notice the size. Sizes, amounts and IDs must be checked exactly before any merge. The model is doing what it was trained for.
4Your team wants to switch to a newer embedding model. What must be planned?
Answer: A. Each model has its own map, so vectors from two models can't be compared. Stored embeddings must be recomputed, and a threshold chosen for the old model says nothing about the new one.
5Which SAP offering stores embeddings next to business tables and compares them with SQL?
Answer: C. The vector engine, generally available since April 2024, adds a vector column type and similarity functions such as COSINE_SIMILARITY to SAP HANA Cloud, next to relational, graph, spatial and JSON data.
6Text is sent to an external embedding model through SAP's generative AI hub. What should you ask first?
Answer: B. Embedding means sending the text to wherever the model runs. SAP's orchestration service can mask personal data before embedding, so ask whether that is switched on and what else protects the data.
7Which task is the worst fit for embeddings on their own?
Answer: D. An order number either matches or it doesn't. Embeddings blur exact values, so exact match or keyword search is the right tool. The other three tasks are about meaning in people's own words.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
An embedding model is a function that places any text at a point in a space with a few hundred dimensions. It was trained so that texts with the same meaning land in the same direction from the origin. Semantic similarity is then just geometry: the smaller the angle between two arrows, the closer the meaning.
Everything else in this topic follows from that picture:
Search is "find the arrows closest to the question's arrow".
Routing is "which category description points the same way as this message".
Deduplication is "which pairs of arrows almost overlap".
What "close" means depends entirely on what the model was trained to pull together. That is why two bolts of different sizes can overlap.
In Neural networks from scratch, the hidden layer built its own intermediate signals. An embedding is the same thing, taken from a much larger network: the network's internal description of a text, used on its own.
A sentence embedding model such as all-MiniLM-L6-v2 does four things when you call it.
flowchart LR
T[Text] --> K[Split into<br/>word pieces]
K --> N[Transformer network<br/>one vector per piece]
N --> P[Mean pooling<br/>average the pieces]
P --> L[Scale to length 1]
L --> E[Embedding<br/>384 numbers]
Tokenize. The text is split into word pieces (tokens). Common words stay whole; rare words are split into parts.
Encode. A transformer network turns each piece into a vector. Each vector depends on the words around it. This is a contextual embedding: "order" in "sales order" and "order" in "out of order" get different vectors. Google's crash course contrasts this with older static embeddings such as word2vec, which gave each word one fixed vector. Unit 4 opens up the transformer.
Pool. The model card for all-MiniLM-L6-v2 says it uses mean pooling: it averages the piece vectors, skipping padding, to get one vector for the whole text.
Normalize. The vector is scaled to length 1. The model card shows this step explicitly. Then only the direction carries meaning.
The model card states the limits: 384 dimensions, and "input text longer than 256 word pieces is truncated". Anything after that point is silently ignored. The card describes the model as an encoder for sentences and short paragraphs. Long documents must be cut into chunks first; Unit 7 covers chunking.
The model wasn't told what words mean. According to its model card, it was trained on over 1 billion sentence pairs with a contrastive objective. Given one sentence, the model had to pick its true partner from a batch of randomly chosen other sentences. The pairs were things like question and answer, or two phrasings of the same question.
Each training step pulls true pairs closer and pushes the random others apart. This is the embedding-layer training from Google's crash course, at scale: weights are adjusted "so that embedding vectors for similar examples are closer to each other".
This tells you what the model will and won't treat as close:
Close: paraphrases, a question and its answer, the same topic in different words.
Not reliably separated: texts on the same topic that differ in one number, one negation or one unit. "M8" and "M10" are rare word pieces in a long shared context.
Weak: in-house jargon and abbreviations that never appeared in training, and languages the training data barely covered.
Three measures are common. The SAP HANA Cloud vector engine offers the first and the third as SQL functions.
Measure
Formula (in words)
Range
Closer means
Cosine similarity
Dot product divided by both lengths
-1 to 1
Higher
Dot product
Multiply matching positions, add up
Any number
Higher
Euclidean (L2) distance
Straight-line distance between the points
0 or more
Lower
For vectors of length 1, the three agree. Cosine similarity equals the dot product, and OpenAI's embeddings guide notes that ranking by Euclidean distance gives identical results. That is why most pipelines normalize once, at embedding time, and then use the cheaper dot product. The Sentence Transformers semantic search guide makes the same point.
A worked example, the one SAP Learning uses for COSINE_SIMILARITY: the vectors [1, 0, 0] and [1, 0, 1].
Dot product: 1×1 + 0×0 + 0×1 = 1.
Lengths: 1 and √2 ≈ 1.414.
Cosine: 1 / 1.414 = 0.7071, the cosine of 45 degrees.
Multiply the second vector by 5 and the cosine stays 0.7071, because the direction didn't change. The Euclidean distance changes a lot. That is why cosine is the default for text.
The Sentence Transformers documentation separates two cases:
Symmetric: both sides look alike. Two material descriptions, or two ticket texts. Deduplication and routing to short category descriptions are mostly symmetric.
Asymmetric: a short question against longer passages, such as "what is the penalty for late delivery?" against contract paragraphs.
Some models are tuned for one case. Some expect a prefix or prompt that says whether a text is a query or a document; Sentence Transformers supports this with a prompt_name option. all-MiniLM-L6-v2 needs no prompt. Always check the model card.
The model in this topic is a bi-encoder: it embeds each text on its own, so you can embed a million descriptions once, store them, and compare new texts in milliseconds.
A cross-encoder reads two texts together and returns one relevance score. It is more accurate and much slower, because nothing can be computed ahead of time. The Sentence Transformers documentation recommends combining them: the bi-encoder retrieves the top candidates, and a cross-encoder re-ranks them. Unit 7 builds this.
Comparing a question with every stored vector is called exhaustive or brute-force search. It is fine for thousands of vectors and exact. For millions, vector stores use approximate nearest neighbour (ANN) indexes that skip most comparisons and accept a small chance of missing a result. HNSW is a common one; the LangChain integration for SAP HANA Cloud creates HNSW indexes on the vector column. Unit 7 compares exact and approximate search on SAP HANA Cloud.
#Build it yourself: compare SAP-style texts by meaning
You will write one script that computes cosine similarity by hand, then compares nine made-up service-desk notes with a question two ways: by shared words and by meaning. It then routes each note to one of three exception types: credit block, three-way match or MRP exception. You will see where word matching fails and embeddings don't.
Before you start: complete Set up your computer for this course and Set up for Unit 3. They give you the orchestrate-course folder with its .venv, numpy, PyTorch, Sentence Transformers and the unit03 folder. Step 6 of the Unit 3 setup also downloads the embedding model used here. This walkthrough doesn't repeat those steps.
flowchart LR
Q[Question] --> W[Word counts]
N[9 notes] --> W
Q --> E[Embedding model<br/>on your laptop]
N --> E
C[3 category<br/>descriptions] --> W
C --> E
W --> R1[Ranking and routing<br/>by shared words]
E --> R2[Ranking and routing<br/>by meaning]
The course folder, .venv and unit03 folder from the setup topics.
About 40 minutes. No accounts, no API keys, no cost.
The embedding model from the Unit 3 setup, downloaded once from Hugging Face. If your network blocks the download, the --offline option runs everything except the embedding part.
All texts are made up. The model runs on your computer, so nothing is sent anywhere.
#Step 1: Open the course folder and turn on the environment
Open VS Code, choose File > Open Folder and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
Turn on the virtual environment if the prompt doesn't start with (.venv):
In VS Code's file list, right-click unit03, choose New File and name it similarity.py.
Paste the code below and save with File > Save.
"""Unit 3: compare SAP-style texts by meaning (embeddings) and by shared words.
How to run (from the unit03 folder, with the course .venv turned on):
python similarity.py # embedding model (downloads once) and word counts
python similarity.py --offline # no download, no internet: word counts only
python similarity.py --query "Why is my invoice not paid?" # your own question
python similarity.py --save # also write note_vectors.npz
All texts are made up. Nothing is sent anywhere; the model runs on your computer.
"""
import argparse
import re
from pathlib import Path
import numpy as np
HERE = Path(__file__).parent
MODEL = "sentence-transformers/all-MiniLM-L6-v2"
NOTES = [ # made-up notes, shaped like what people write about SAP documents
"Customer is over the credit limit, order on hold",
"Finance has to release it before we can deliver",
"Credit check failed for the new distributor",
"Invoice price is higher than the purchase order, payment blocked",
"Supplier billed 120 units but we only received 100",
"Vendor bill parked because of a variance",
"Planned order is late because a component is missing",
"Requirement date moved earlier, reschedule the receipt",
"Stock runs out before the next delivery arrives",
]
CATEGORIES = { # one plain description per exception type
"credit block": "Sales order blocked by the customer credit check",
"three-way match": "Supplier invoice blocked because it does not match the purchase order or goods receipt",
"MRP exception": "Material planning exception: shortage, late order or date change",
}
STOP = {"the", "a", "an", "is", "are", "be", "to", "of", "and", "or", "for", "it", "we", "but",
"by", "on", "in", "this", "that", "has", "have", "can", "because", "before", "than", "does", "not"}
def cosine(a: np.ndarray, b: np.ndarray) -> float:
"""Cosine similarity: the dot product of the two vectors after scaling each to length 1."""
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
def words(text: str) -> list[str]:
return [w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP]
def word_vectors(texts: list[str]) -> np.ndarray:
"""One column per word in the texts; each value counts how often the word appears."""
vocab = sorted({w for t in texts for w in words(t)})
index = {w: i for i, w in enumerate(vocab)}
vectors = np.zeros((len(texts), len(vocab)))
for row, text in enumerate(texts):
for w in words(text):
vectors[row, index[w]] += 1
return vectors
def embed(texts: list[str], model_name: str) -> np.ndarray:
"""Turn each text into one embedding with a Sentence Transformers model, scaled to length 1."""
try:
from sentence_transformers import SentenceTransformer
except ImportError:
raise SystemExit("sentence-transformers is not installed. See Set up for Unit 3, or use --offline.")
print(f"Loading {model_name} (downloads once, then uses the copy on disk)...")
model = SentenceTransformer(model_name)
return model.encode(texts, normalize_embeddings=True)
def by_hand() -> None:
print("1. Cosine similarity and distance by hand")
a, b = np.array([1.0, 0.0, 0.0]), np.array([1.0, 0.0, 1.0])
print(f" cosine([1,0,0], [1,0,1]) = {cosine(a, b):.4f} (45 degrees apart)")
print(f" cosine([1,0,0], [5,0,5]) = {cosine(a, 5 * b):.4f} (longer arrow, same direction)")
c, d = np.array([6.0, 3.0, 5.0]), np.array([6.0, 3.0, -5.0])
print(f" distance([6,3,5], [6,3,-5]) = {np.linalg.norm(c - d):.4f}")
def rank(title: str, query_vec: np.ndarray, note_vecs: np.ndarray) -> None:
scores = [cosine(query_vec, v) if np.any(v) and np.any(query_vec) else 0.0 for v in note_vecs]
print(f" {title}")
for score, note in sorted(zip(scores, NOTES), reverse=True)[:4]:
print(f" {score:5.2f} {note}")
def route(title: str, note_vecs: np.ndarray, cat_vecs: np.ndarray) -> None:
names = list(CATEGORIES)
print(f" {title}")
for note, v in zip(NOTES, note_vecs):
scores = [cosine(v, c) if np.any(v) and np.any(c) else 0.0 for c in cat_vecs]
best = int(np.argmax(scores))
label = names[best] if scores[best] > 0 else "(no shared words)"
print(f" {label:17s} {scores[best]:4.2f} {note}")
def main() -> None:
parser = argparse.ArgumentParser(description="Compare SAP-style texts by meaning and by shared words.")
parser.add_argument("--query", default="Why is the client's delivery stuck?", help="question to rank notes against")
parser.add_argument("--offline", action="store_true", help="skip the embedding model; word counts only")
parser.add_argument("--model", default=MODEL, help="Sentence Transformers model name or local folder")
parser.add_argument("--save", action="store_true", help="write note_vectors.npz with the note embeddings")
args = parser.parse_args()
by_hand()
all_texts = [args.query] + NOTES + list(CATEGORIES.values())
wv = word_vectors(all_texts) # word counts share one vocabulary across all texts
wq, wn, wc = wv[0], wv[1:1 + len(NOTES)], wv[1 + len(NOTES):]
print(f"\n2. Notes closest to: '{args.query}'")
rank("By shared words (word counts):", wq, wn)
if not args.offline:
ev = embed(all_texts, args.model)
print(f" Each text became {ev.shape[1]} numbers.")
eq, en, ec = ev[0], ev[1:1 + len(NOTES)], ev[1 + len(NOTES):]
rank("By meaning (embeddings):", eq, en)
print("\n3. Route each note to the closest exception type")
route("By shared words:", wn, wc)
if not args.offline:
route("By meaning:", en, ec)
if args.save:
if args.offline:
print("\n--save needs embeddings. Run without --offline.")
else:
np.savez(HERE / "note_vectors.npz", texts=np.array(NOTES), vectors=en, model=args.model)
print(f"\nSaved {len(NOTES)} note embeddings to note_vectors.npz")
if args.offline:
print("\nOffline run: skipped the embedding model. Run without --offline to compare by meaning.")
if __name__ == "__main__":
main()
Start without the model, so you see the keyword approach on its own.
In the terminal, inside unit03, run:
python similarity.py --offline
What success looks like:
1. Cosine similarity and distance by hand
cosine([1,0,0], [1,0,1]) = 0.7071 (45 degrees apart)
cosine([1,0,0], [5,0,5]) = 0.7071 (longer arrow, same direction)
distance([6,3,5], [6,3,-5]) = 10.0000
2. Notes closest to: 'Why is the client's delivery stuck?'
By shared words (word counts):
0.18 Stock runs out before the next delivery arrives
0.00 Vendor bill parked because of a variance
0.00 Supplier billed 120 units but we only received 100
0.00 Requirement date moved earlier, reschedule the receipt
3. Route each note to the closest exception type
By shared words:
credit block 0.50 Customer is over the credit limit, order on hold
(no shared words) 0.00 Finance has to release it before we can deliver
credit block 0.37 Credit check failed for the new distributor
three-way match 0.53 Invoice price is higher than the purchase order, payment blocked
three-way match 0.13 Supplier billed 120 units but we only received 100
(no shared words) 0.00 Vendor bill parked because of a variance
MRP exception 0.32 Planned order is late because a component is missing
three-way match 0.14 Requirement date moved earlier, reschedule the receipt
(no shared words) 0.00 Stock runs out before the next delivery arrives
Offline run: skipped the embedding model. Run without --offline to compare by meaning.
Section 1 is the SAP Learning example for COSINE_SIMILARITY, computed by you: 0.7071. Stretching the second arrow to [5,0,5] doesn't change the cosine. The distance example is SAP Learning's L2DISTANCE example: the two points differ only in the last number, by 10.
Section 2 shows keyword matching failing. The question is about a customer whose order won't ship, a credit block. The only note sharing a word ("delivery") is a planning note. Every other note scores 0.00, so their order below the first line means nothing.
Section 3 shows the same weakness in routing. Three notes share no word with any category. "Requirement date moved earlier" goes to three-way match only because of the word "receipt". Keyword routing is right when the writer happens to use the expected words.
The first line says Loading sentence-transformers/all-MiniLM-L6-v2 .... If you ran Step 6 of the Unit 3 setup, the model is already on disk and loads in a few seconds. Otherwise it downloads once. No account or key is needed.
What success looks like (the offline lines from Step 3 appear too; only the new lines are shown):
Loading sentence-transformers/all-MiniLM-L6-v2 (downloads once, then uses the copy on disk)...
Each text became 384 numbers.
By meaning (embeddings):
0.xx ...
0.xx ...
0.xx ...
0.xx ...
...
By meaning:
credit block 0.xx Customer is over the credit limit, order on hold
...
Your lines show real scores in place of 0.xx. Now every note has a non-zero score, because every text has a position on the map. Look for three things:
Do the credit-block notes rise to the top of the ranking for the delivery question, above the planning note that shares the word "delivery"?
Do "Vendor bill parked because of a variance" and "Finance has to release it" now get a category, and is it the right one?
Which notes are still routed wrongly, and do their scores look lower than the correct ones?
Write the answers down; the exercise uses them. A small model won't get every note right, and that is the point of Step 6.
python similarity.py --query "Why hasn't the supplier been paid yet?"
Ask something unrelated to all nine notes:
python similarity.py --query "Where is the canteen menu for next week?"
Compare the top score in each case. An unrelated question still returns four "closest" notes, only with lower scores. Search always returns something. Deciding when the best match is not good enough is your job, with a threshold you choose from examples.
Try a negation:
python similarity.py --query "The order is not blocked for credit"
Check whether the credit-block notes still come first. Sentence embeddings often don't separate a statement from its negation well, because both are about the same topic.
The file holds the nine texts, their 384-number vectors and the model name. Storing the model name with the vectors matters: vectors from another model can't be compared with these.
This section maps the build onto SAP's services, as of September 2026. Setting them up needs access you don't have yet; Unit 5 covers the generative AI hub and Unit 7 covers SAP HANA Cloud and its vector engine.
#Storing and comparing: the SAP HANA Cloud vector engine
SAP announced the vector engine as generally available with the SAP HANA Cloud release of April 2024. SAP Learning's course describes the building blocks:
Your script
SAP HANA Cloud vector engine
A numpy array of 384 numbers
A column of type REAL_VECTOR: single-precision numbers, 1 to 65,000 dimensions
cosine(a, b)
COSINE_SIMILARITY(a, b), returning -1 to 1
np.linalg.norm(c - d)
L2DISTANCE(c, d), returning 0 or more; lower means closer
np.array([1.0, 0.0, 0.0])
TO_REAL_VECTOR('[1, 0, 0]')
note_vectors.npz
A table with a text column and a vector column
SAP Learning's own examples look like this (sketch; it needs an SAP HANA Cloud instance):
The first query returns the 0.7071 your script printed. SAP Learning also notes that vectors have no order, so you can't sort or group by a vector column. You sort by the similarity score instead.
The LangChain integration for SAP HANA Cloud (langchain-hana) adds two points worth knowing now. Its vector store can use REAL_VECTOR or a HALF_VECTOR column. It uses cosine similarity by default, with Euclidean distance as an option, and can build an HNSW index for fast approximate search.
SAP HANA Cloud has a SQL function, VECTOR_EMBEDDING, that computes embeddings in the database with an SAP-provided model. The LangChain integration's HanaInternalEmbeddings class calls it, with the model ID SAP_NEB.20240715 in its example, and notes that it needs the NLP feature enabled on the instance. Check SAP's documentation for the current model IDs, languages and limits before you design around it; Unit 7 does this in detail.
The advantage is that text never leaves the database to be embedded. The trade-off is that you use the models SAP provides there.
#Creating embeddings through the generative AI hub
The generative AI hub in SAP AI Core offers embedding models from model providers. The SAP Cloud SDK for AI (Python) shows an embeddings endpoint in the orchestration service, with text-embedding-3-large in its example. Its options include the number of dimensions and normalization, and the data masking module can anonymize personal data before the text is embedded.
For scale: OpenAI's guide lists 3,072 dimensions by default for text-embedding-3-large and an input limit of 8,192 tokens, against 384 dimensions and 256 word pieces for the laptop model. Larger is not automatically better for your texts; test on your own examples.
all-MiniLM-L6-v2 is published on Hugging Face; check the licence on its model card before commercial use. SAP HANA Cloud and SAP AI Core are paid SAP BTP services, and what you pay for them and for provider models in the generative AI hub depends on your contract. Unit 5 covers the commercial side.
One model per vector store. Store the model name and version with every vector, as --save does. When the model changes, re-embed everything; mixing models gives meaningless scores.
Evaluate on your own texts. Build a list of pairs or queries with the right answers marked by a business expert. Measure how often the right item is in the top results, the retrieval version of the precision and recall from Classification and the metrics that matter. Unit 8 builds an evaluation harness.
Choose thresholds from data. Plot scores for true matches and non-matches, then pick the cut-off by the cost of each mistake. Re-check it for every new model.
Data protection. Embedding with an external model sends the text out. Use in-database embedding or a local model for sensitive text, or mask personal data first. Treat stored vectors as sensitive as their source: they are built to reveal what the text was about.
Authorizations. A vector store that holds sales notes from all company codes must filter results by what the user may see in SAP. The access rules from Calling your first SAP API still apply; Unit 7 covers retrieval that respects SAP authorizations.
Cost and speed. Embed each text once, when it is created or changed, not at every search. Batch the calls. Normalize once, then use the dot product.
Clean core. Read texts through released APIs and keep the vector store side by side, in SAP HANA Cloud or your own service. Don't add embedding columns to standard S/4HANA tables.
Trusting a match on numbers. "M8" and "M10", "100 pcs" and "1,000 pcs" look nearly the same to a sentence model. Confirm sizes, amounts and IDs with rules.
Ignoring truncation. Text beyond the model's limit, 256 word pieces for all-MiniLM-L6-v2, is dropped without a warning. Chunk long documents.
Comparing vectors from two models. It runs without an error and returns nonsense.
No threshold. Search always returns a "closest" result, even for an unrelated question, as Step 6 showed.
Forgetting normalization. If vectors aren't length 1, the dot product favours long vectors. Normalize, or use cosine.
Embedding identifiers. Order and material numbers carry no meaning for the model. Keep them in normal columns and filter on them.
Negation and small words. "Blocked" and "not blocked" often score as similar. Don't use similarity alone to decide a status.
Judging a model by its general benchmark. In-house abbreviations and product codes may be unknown to it. Test on your texts.
You will rank every pair of twelve made-up material descriptions by similarity, save the top pairs to a CSV file and mark each one as a true or false duplicate by hand. This is the first step of the master data work in the next topic of this unit, semantic search over SAP master data.
In VS Code, create unit03/find_duplicates.py in the same folder as similarity.py, paste the code below and save. It reuses functions from similarity.py, so both files must be in unit03.
"""Find possible duplicate material descriptions with embeddings.
How to run (from the unit03 folder, with the course .venv turned on):
python find_duplicates.py # embedding model (downloaded by similarity.py)
python find_duplicates.py --offline # no model: shared words only
python find_duplicates.py --top 10 # show more pairs
Writes duplicate_candidates.csv. All descriptions are made up.
"""
import argparse
import csv
from itertools import combinations
from pathlib import Path
import numpy as np
from similarity import MODEL, cosine, embed, word_vectors
HERE = Path(__file__).parent
MATERIALS = [ # (made-up material number, description)
("M-1001", "Hex bolt M8x40 zinc plated"),
("M-1002", "Bolt, hexagon head, M8 x 40, galvanised"),
("M-1003", "Hex bolt M10x40 zinc plated"),
("M-1004", "Nitrile gloves size L, box of 100"),
("M-1005", "Gloves nitrile large 100 pcs"),
("M-1006", "Deep groove ball bearing 6204 2RS"),
("M-1007", "Bearing 6204-2RS sealed"),
("M-1008", "Hydraulic oil ISO VG 46, 20 l can"),
("M-1009", "Hydraulic fluid VG46 20 litre"),
("M-1010", "Safety glasses, clear lens"),
("M-1011", "Copy paper A4 80 g, 500 sheets"),
("M-1012", "Printer paper DIN A4 white ream"),
]
def main() -> None:
parser = argparse.ArgumentParser(description="Rank pairs of material descriptions by similarity.")
parser.add_argument("--offline", action="store_true", help="use shared words instead of the embedding model")
parser.add_argument("--model", default=MODEL, help="Sentence Transformers model name or local folder")
parser.add_argument("--top", type=int, default=8, help="how many of the most similar pairs to show")
args = parser.parse_args()
texts = [text for _, text in MATERIALS]
vectors = word_vectors(texts) if args.offline else embed(texts, args.model)
method = "shared words" if args.offline else "embeddings"
pairs = []
for i, j in combinations(range(len(MATERIALS)), 2): # every pair once: 12 items give 66 pairs
score = cosine(vectors[i], vectors[j]) if np.any(vectors[i]) and np.any(vectors[j]) else 0.0
pairs.append((score, MATERIALS[i], MATERIALS[j]))
pairs.sort(key=lambda p: p[0], reverse=True)
print(f"Top {args.top} of {len(pairs)} pairs by {method}:")
for score, (id_a, text_a), (id_b, text_b) in pairs[:args.top]:
print(f" {score:4.2f} {id_a} {text_a:40s} | {id_b} {text_b}")
out = HERE / "duplicate_candidates.csv"
with out.open("w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerow(["score", "material_a", "text_a", "material_b", "text_b", "method", "same_item"])
for score, (id_a, text_a), (id_b, text_b) in pairs[:args.top]:
writer.writerow([f"{score:.2f}", id_a, text_a, id_b, text_b, method, ""])
print(f"\nWrote {out.name}. Fill in same_item with yes or no for each row.")
if __name__ == "__main__":
main()
Run the keyword baseline first:
python find_duplicates.py --offline
You should see:
Top 8 of 66 pairs by shared words:
0.80 M-1001 Hex bolt M8x40 zinc plated | M-1003 Hex bolt M10x40 zinc plated
0.61 M-1006 Deep groove ball bearing 6204 2RS | M-1007 Bearing 6204-2RS sealed
0.55 M-1004 Nitrile gloves size L, box of 100 | M-1005 Gloves nitrile large 100 pcs
0.34 M-1008 Hydraulic oil ISO VG 46, 20 l can | M-1009 Hydraulic fluid VG46 20 litre
0.31 M-1011 Copy paper A4 80 g, 500 sheets | M-1012 Printer paper DIN A4 white ream
0.17 M-1001 Hex bolt M8x40 zinc plated | M-1002 Bolt, hexagon head, M8 x 40, galvanised
0.17 M-1002 Bolt, hexagon head, M8 x 40, galvanised | M-1003 Hex bolt M10x40 zinc plated
0.15 M-1004 Nitrile gloves size L, box of 100 | M-1008 Hydraulic oil ISO VG 46, 20 l can
Wrote duplicate_candidates.csv. Fill in same_item with yes or no for each row.
The top pair is not a duplicate: M8 and M10 are different bolts. The real duplicate of M-1001, M-1002, is only sixth.
Rename the offline result so it isn't overwritten:
You should see Top 8 of 66 pairs by embeddings: followed by eight pairs with scores, and a new duplicate_candidates.csv.
Open both CSV files in VS Code. In the same_item column of each row, type yes if the two descriptions are the same item, or no. The five true duplicate pairs are M-1001/M-1002, M-1004/M-1005, M-1006/M-1007, M-1008/M-1009 and M-1011/M-1012. Save both files.
Create unit03/notes_embeddings.md with three short sections:
Routing: from Step 5 of the walkthrough, which notes the embeddings routed correctly that word counts missed, and which ones they still got wrong.
Duplicates: for each method, how many of the top 5 pairs are true duplicates (the precision at 5), and where the M8/M10 pair landed.
Rule: one sentence describing a rule that would stop the M8/M10 pair from being merged, such as comparing the numbers in both descriptions.
Save your work:
git add find_duplicates.py duplicate_candidates.csv duplicate_candidates_words.csv notes_embeddings.md
git commit -m "Rank duplicate material descriptions by words and by embeddings"
Done when: both CSV files have yes or no in every same_item cell; your notes give the precision at 5 for both methods and name where the M8/M10 pair ranked; and git log shows the commit.
Pick one answer for each question. The explanation appears after you choose.
1Why does all-MiniLM-L6-v2 put "credit check failed" close to "customer over the credit limit"?
Answer: C. The model card describes contrastive training on over a billion pairs: true partners are pulled together and random others pushed apart. Closeness in its space reflects what it learned to pair, not shared words.
2In similarity.py, why does embed pass normalize_embeddings=True?
Answer: A. Scaling to length 1 leaves only the direction. Then the dot product equals the cosine, and Euclidean distance ranks the same way. It does nothing to make two models' vectors comparable.
3cosine([1,0,0], [5,0,5]) prints 0.7071, the same as for [1,0,1]. What does that show?
Answer: D. The angle between the arrows stays 45 degrees when you stretch one of them. Euclidean distance would change a lot, which is why cosine is the default for comparing text.
4In the offline run, "Vendor bill parked because of a variance" gets "(no shared words)". What does that illustrate?
Answer: B. The note says "vendor bill" and "parked"; the three-way match description says "supplier invoice" and "blocked". With no word in common, the word-count cosine is 0. Embeddings place both by meaning.
5Your company switches from all-MiniLM-L6-v2 to a 3,072-dimension provider model. What must happen to the vectors already stored?
Answer: C. Each model has its own space; padding or scaling doesn't translate between them. Mixed vectors return scores that run fine and mean nothing. Store the model name with the vectors, as --save does, so you know when to re-embed.
6A 30-page supplier contract is embedded with all-MiniLM-L6-v2 as one text. What is the problem?
Answer: B. The model card says input longer than 256 word pieces is truncated, without an error. Long documents must be split into chunks and each chunk embedded, which Unit 7 covers.
7Which SAP HANA Cloud statement matches cosine(a, b) in the script?
Answer: D. COSINE_SIMILARITY is the vector engine's cosine function. L2DISTANCE is lower for closer vectors, TO_REAL_VECTOR only builds a vector, and SAP Learning notes that vectors have no order to sort by.
8A client wants to embed HR case notes but must not send personal data to an external model. What would you propose?
Answer: C. External embedding sends the text out. In-database embedding with VECTOR_EMBEDDING or a local model keeps it in, and the orchestration service can mask personal data before embedding. Vectors still reveal what the text was about, so protect them too.
9The duplicate-finder ranks "Hex bolt M8x40" and "Hex bolt M10x40" first. What would you do before merging materials automatically?
Answer: D. Sentence embeddings barely separate texts that differ in one number. A rule on sizes and amounts catches this reliably, and human review protects master data. A bigger model or a different threshold may still miss it.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Semantic Search (Sentence Transformers documentation)— symmetric vs asymmetric search; dot product is faster for normalized vectors; approximate nearest neighbor search and vector databases; retrieve and re-rank with a cross-encoder
Vector embeddings (OpenAI API guide)— text-embedding-3-small 1536 and -large 3072 dimensions by default; dimensions parameter to shorten; vectors normalized to length 1, so cosine equals the dot product and ranks like Euclidean distance; 8192-token input limit
SAP HANA Cloud Vector Engine (SAP Learning)— REAL_VECTOR type of single-precision floats, 1 to 65,000 dimensions, no ordering; COSINE_SIMILARITY returns -1 to 1; L2DISTANCE returns 0 or more; TO_REAL_VECTOR examples
SAP HANA Cloud vector engine integration (LangChain documentation)— langchain-hana package; HanaInternalEmbeddings runs VECTOR_EMBEDDING with model ID SAP_NEB.20240715 when NLP is enabled; REAL_VECTOR or HALF_VECTOR columns; cosine default, Euclidean option; HNSW index