Find products by what people mean, not only by the words they type, by combining keyword search, embeddings and master data filters, and prove it works.
Every SAP system holds thousands of records that describe things: products, suppliers, customers. People look them up all day. Usually they type a few words into a search box and hope the right record has the same words.
Semantic search looks records up by meaning. It turns each description into an embedding, the list of numbers from Embeddings and semantic similarity, and finds the records closest to the question. "Hearing protection" can find "Ear muffs, 30 dB" even though the two share no word.
Meaning alone isn't enough for master data, though. Part numbers, sizes and codes must match exactly. So good master data search combines three things:
Keyword search for exact words, numbers and codes.
Semantic search for meaning.
Filters on the structured fields SAP already has, such as product group or "marked for deletion".
Combining the first two is called hybrid search. The filters make sure nobody picks a blocked or deleted product. And a small set of test questions with known right answers tells you whether any of it works.
Master data search failures are quiet and expensive.
Take procure-to-pay. A buyer needs "galvanised screws, 8 mm". The system holds "Hex bolt M8x40 zinc plated". The keyword search finds nothing, so the buyer creates a new material. Now one part lives under two numbers. Stock is split, the spend analysis is wrong and MRP plans two items. The embeddings topic showed how to find such duplicates after the fact. Better search prevents many of them at the source.
The same pattern shows up elsewhere:
Order-to-cash. A service agent needs the material for "the blue pens we always order". A search by meaning finds the right product faster than scrolling a catalogue.
Service desk. Many requests start with "which product is this about?" before anyone can look at a blocked sales order or a three-way match exception.
AI assistants. An assistant that answers questions about products must first find the right product records. That retrieval step is this topic. Unit 7 builds the assistant on top of it.
The value is fewer duplicates, faster lookups and a solid base for AI assistants. The risks are specific. Semantic search always returns something, even for nonsense. It blurs numbers: M8 and M10 bolts look almost identical to it. And a search index copied out of SAP can show records a user may not see in SAP itself.
As of September 2026, SAP covers master data search in two ways: inside its master data applications, and as building blocks for your own search.
SAP Master Data Governance (MDG). SAP Learning describes duplicate checks built into the creation process, so a user sees likely duplicates before saving. The checks use either enterprise search, with free-text and optional fuzzy (typo-tolerant) search, or SAP HANA-based search, which ranks each hit.
AI-assisted central governance. SAP's Q4 2025 release highlights list this as generally available for MDG on SAP S/4HANA Cloud Private Edition. Users can search, display, create and change business partners in natural language through Joule.
SAP HANA Cloud vector engine. For your own search, SAP HANA Cloud stores embeddings next to business tables and compares them with SQL functions such as COSINE_SIMILARITY. A function called VECTOR_EMBEDDING can compute embeddings inside the database. SAP's CAP documentation shows it with an SAP-provided model and marks CAP's support for it as beta.
APIs to read the data. The product master is available through SAP APIs such as the Product Master API, API_PRODUCT_SRV, which you met in the style of Calling your first SAP API.
Check what your own licences include before you plan. MDG and SAP HANA Cloud are separate SAP products, and the AI features above depend on edition and release. For how these offerings fit together, see the SAP Business AI landscape.
Anything, when some products are deleted or blocked
Shows them unless filtered
Shows them unless filtered
Filters remove them before ranking
The pattern: keyword search is precise but literal, semantic search is flexible but vague, and filters carry the business rules neither of them knows. Most teams should aim for hybrid search with filters, and measure it.
Which master data objects and fields go into the search? Descriptions only, or also long texts, product groups and supplier part numbers?
How do we measure search quality? Is there a list of real searches with the right answers, checked by someone from the business? What share of searches find the right product in the top three?
What happens when nothing matches well? Does the system say "no good match", or does it always show something?
Which filters apply before the search? Deleted, blocked or wrong-plant products should never appear.
Does the search respect SAP authorizations? If the index is copied out of SAP, how are the user's permissions applied?
How fresh is the index? How long after a product changes in SAP does the search see it?
Which languages do our descriptions use, and does the embedding model support all of them?
Are we building this, or does MDG or another SAP application already cover our case?
"Semantic search replaces keyword search." Codes, sizes and part numbers still need exact matching. Hybrid search keeps both.
"If it returns a result, it found the right product." Semantic search always returns its closest records, even when none fits. You need a threshold and a "no good match" answer.
"Better AI will fix bad master data." Short, cryptic or duplicate descriptions hurt every search method. Search quality depends on data quality.
"We can judge search quality by trying a few queries." Opinions differ from one person to the next. A fixed test set with known right answers gives a number you can track.
"A search index outside SAP is just a copy, so it's harmless." It can expose records to people who can't see them in SAP. It needs the same access rules.
Pick one answer for each question. The explanation appears after you choose.
1A buyer searches "galvanised screws 8 mm" and finds nothing, so they create a new material. What is the main business cost?
Answer: B. A failed search often ends in a new material that duplicates an existing one. Stock, spend analysis and MRP then treat one part as two. Better search prevents this at the source.
2Why do most teams combine keyword and semantic search instead of choosing one?
Answer: D. Keyword search is precise but literal, and semantic search is flexible but blurs numbers such as M8 and M10. Hybrid search keeps the strengths of both.
3A user searches for "canteen menu" in the product search. What should a well-built system do?
Answer: A. Semantic search always has a "closest" result, even for unrelated questions. A threshold lets the system answer "no good match" instead of suggesting a wrong product.
4Your search shows products that are marked for deletion. What is the right fix?
Answer: C. The deletion flag is a structured field SAP already maintains. A filter applies that business rule reliably. Neither the model nor a threshold knows which products are deleted.
5A vendor says their search "works great" after a demo. What should you ask for?
Answer: D. A demo shows chosen examples. A fixed test set of your own searches, checked by the business, gives a number such as the hit rate at 3 that you can compare and track.
6The search index is copied out of SAP into a separate database. Which risk needs a plan?
Answer: B. A copy doesn't carry SAP's authorizations automatically. The search must filter results by what each user may see, and the index must be kept fresh.
7Your company needs duplicate checks when users create business partners. Where would you look first?
Answer: D. SAP Learning describes MDG's duplicate checks as part of the creation process, using enterprise search or SAP HANA-based search. Check what your licence covers before building your own.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
The gate is made of filters on structured fields: not marked for deletion, the right product group, what this user may see. It decides which records are allowed to appear at all.
The keyword ranker rewards records that share rare words with the query. It is exact and literal.
The semantic ranker rewards records whose embedding points the same way as the query's. It is flexible and vague.
Fusion merges the two rankings into one list. A threshold decides when even the best result isn't good enough. A test set tells you whether any of this beats what users have today.
Everything you build in this topic is one of those five parts. In Unit 7, the same parts become the retrieval step of a RAG system.
flowchart LR
S[SAP product API] --> T[Text per product<br/>plus fields]
T --> K[Keyword index<br/>BM25]
T --> V[Vector index<br/>embeddings]
Q[Query] --> G{Filters<br/>deleted, group}
G --> K
G --> V
K --> F[Fusion<br/>RRF]
V --> F
F --> R[Threshold<br/>and top results]
The Product Master API, API_PRODUCT_SRV, lives at the service path /sap/opu/odata/sap/API_PRODUCT_SRV. The SAP Cloud SDK's model of it, generated from the service metadata, lists these fields among many others:
Entity
Field
Type, length
Use in search
A_Product
Product
String, 40 (key)
Exact-match lookup; never embed it
A_Product
ProductGroup
String, 9
Filter
A_Product
IsMarkedForDeletion
Boolean
Filter: hide deleted products
A_Product
AuthorizationGroup
String, 4
Input to access filtering
A_Product
LastChangeDateTime
Date and time
Find changed products to re-index
A_ProductDescription
Product, Language
String 40, String 2 (key)
One description per language
A_ProductDescription
ProductDescription
String, 40
The text you search
A_Product links to its descriptions through the navigation property to_Description. The SDK describes the description as text of up to 40 characters, with one description per language.
Forty characters is short. Descriptions are full of abbreviations, sizes and codes: "Bearing 6204-2RS sealed". That shapes every design choice below. Product groups and other codes are configured per system, so the values in this topic are made up.
BM25 is the standard keyword ranking; Elasticsearch uses it by default. For each word the query and a record share, it adds up three effects:
Rare words count more. The inverse document frequency (IDF) is log(1 + (N - n + 0.5) / (n + 0.5)), where N is the number of records and n the number containing the word. "6204" in two of twenty descriptions outweighs "gloves" in three.
Repeats saturate. The parameter k1 limits how much a word repeated in one record can add. Elasticsearch's default is 1.2.
Short records get a boost. The parameter b scales scores by record length against the average. Elasticsearch's default is 0.75; b = 0 switches this off.
Two properties matter for master data. A record that shares no word with the query scores exactly 0, so keyword search can honestly say "nothing". And a word must match exactly as split: "M8x40" is one token, while "M8 x 40" is three.
This is the embeddings topic, applied to records. Embed every description once and store the vectors. At search time, embed the query and compute the cosine similarity with every stored vector. With length-1 vectors the dot product gives the same number, so one matrix product scores all records.
The Sentence Transformers documentation calls this asymmetric search when a short query meets longer documents, and symmetric when both sides look alike. Product search sits in between: short queries against 40-character descriptions. It also notes that comparing against every vector is fine up to about a million entries. Beyond that, approximate nearest neighbour indexes such as HNSW take over.
The two scores can't be added. BM25 scores run from 0 upwards with no fixed top; cosine scores sit between -1 and 1. Reciprocal rank fusion ignores the scores and uses only positions. As the Elasticsearch reference describes it, each record gets 1 / (k + rank) from every list it appears in, with rank starting at 1, and the totals decide the order. The default k is 60.
A worked example with k = 60:
Record
Keyword rank
Semantic rank
RRF score
Nitrile gloves, box of 100
1
3
1/61 + 1/63 = 0.0323
Ear muffs
not found
1
1/61 = 0.0164
Cut-resistant gloves
2
not found
1/62 = 0.0161
A record that both rankers like beats a record only one of them puts first. A large k flattens the difference between rank 1 and rank 5; a small k trusts the top positions more.
Filter first, then rank. If you rank first and filter the top 5 afterwards, a search whose top 5 are all deleted products returns nothing, although good matches sat at positions 6 to 10. The LangChain integration for SAP HANA Cloud supports this directly: its filter argument accepts operators such as $eq, $in and $like, and specific_metadata_columns stores frequently filtered fields as real columns for speed.
Keyword search can return nothing. Semantic search can't: every record has some cosine similarity with every query. A minimum score turns "the closest record" into "a record close enough". The right value depends on the model and your data, so you choose it from your test set, never from another project.
A test set is a list of real queries with the records a business user accepts as right. Two numbers summarize it:
Hit@3: how many queries have a right record in the top 3. This matches how users behave: they look at the first few results.
Mean reciprocal rank (MRR): for each query, 1 divided by the position of the first right record, or 0 if there is none in the top 10; then the average. Right at position 1 scores 1.0, position 2 scores 0.5, position 4 scores 0.25.
You will build one script, product_search.py, that searches twenty made-up product records three ways: by keyword (BM25), by meaning (embeddings) and hybrid (RRF). It filters on the deletion flag and product group first, and it scores each method on ten labelled queries. An optional mode reads real product descriptions from SAP's sandbox.
flowchart LR
D[20 made-up products<br/>or the sandbox] --> F{Filters}
Q[Your query] --> F
F --> K[BM25 keyword ranking]
F --> E[Embedding ranking]
K --> H[RRF hybrid ranking]
E --> H
L[10 labelled queries] --> M["hit@3 and MRR<br/>per method"]
K --> M
E --> M
H --> M
The course folder, .venv, unit03 folder and embedding model from the earlier Unit 3 topics.
About 40 minutes. No new accounts. No cost.
Optional, for Step 8: your SAP Business Accelerator Hub key in .env as SAP_API_KEY, from the Unit 1 setup, and a network that reaches sandbox.api.sap.com.
The twenty products are made up. They reuse the twelve materials from the embeddings exercise and add eight more, including one marked for deletion.
#Step 1: Open the course folder and turn on the environment
Open VS Code, choose File > Open Folder and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
Turn on the virtual environment if the prompt doesn't start with (.venv):
In VS Code's file list, right-click unit03, choose New File and name it product_search.py.
Paste the code below and save with File > Save.
"""Unit 3: search SAP-style product master data by keywords, by meaning, and both (hybrid).
How to run (from the unit03 folder, with the course .venv turned on):
python product_search.py --offline # keyword search only: no model, no internet
python product_search.py # keyword, meaning and hybrid side by side
python product_search.py --query "ear protection" # your own search
python product_search.py --group PPE --query "gloves" # filter on a master data field first
python product_search.py --evaluate # score each method on labelled queries
python product_search.py --evaluate --queries my_queries.csv # score on your own queries
python product_search.py --sandbox # real descriptions from SAP's sandbox (SAP_API_KEY)
python product_search.py --save # write product_index.npz with the vectors
The built-in products are made up. The model runs on your computer; only --sandbox goes online.
"""
import argparse
import csv
import math
import os
import re
import sys
from collections import Counter
from pathlib import Path
import numpy as np
HERE = Path(__file__).parent
MODEL = "sentence-transformers/all-MiniLM-L6-v2"
SERVICE = "/sap/opu/odata/sap/API_PRODUCT_SRV"
SANDBOX = "https://sandbox.api.sap.com/s4hanacloud" + SERVICE
# Made-up products shaped like A_Product plus its to_Description texts (A_ProductDescription).
# ProductGroup values are invented; every SAP system configures its own.
PRODUCTS = [
("M-1001", "FASTENERS", False, "Hex bolt M8x40 zinc plated", "Sechskantschraube M8x40 verzinkt"),
("M-1002", "FASTENERS", False, "Bolt, hexagon head, M8 x 40, galvanised", "Schraube Sechskant M8 x 40 feuerverzinkt"),
("M-1003", "FASTENERS", False, "Hex bolt M10x40 zinc plated", "Sechskantschraube M10x40 verzinkt"),
("M-1004", "PPE", False, "Nitrile gloves size L, box of 100", "Nitrilhandschuhe Gr. L, 100 Stück"),
("M-1005", "PPE", False, "Gloves nitrile large 100 pcs", "Handschuhe Nitril groß 100 St."),
("M-1006", "BEARINGS", False, "Deep groove ball bearing 6204 2RS", "Rillenkugellager 6204 2RS"),
("M-1007", "BEARINGS", False, "Bearing 6204-2RS sealed", "Lager 6204-2RS abgedichtet"),
("M-1008", "LUBES", False, "Hydraulic oil ISO VG 46, 20 l can", "Hydrauliköl ISO VG 46, 20-l-Kanister"),
("M-1009", "LUBES", False, "Hydraulic fluid VG46 20 litre", "Hydraulikflüssigkeit VG46 20 Liter"),
("M-1010", "PPE", False, "Safety glasses, clear lens", "Schutzbrille, klares Glas"),
("M-1011", "OFFICE", False, "Copy paper A4 80 g, 500 sheets", "Kopierpapier A4 80 g, 500 Blatt"),
("M-1012", "OFFICE", False, "Printer paper DIN A4 white ream", "Druckerpapier DIN A4 weiß, Paket"),
("M-1013", "FASTENERS", True, "Hex bolt M8x40 galv. OLD - do not use", "Sechskantschraube M8x40 ALT"),
("M-1014", "PPE", False, "Cut-resistant work gloves, level C", "Schnittschutzhandschuhe Stufe C"),
("M-1015", "PPE", False, "Goggles, indirect vent, anti-fog", "Vollsichtbrille, beschlagfrei"),
("M-1016", "PPE", False, "Ear muffs, 30 dB", "Kapselgehörschutz 30 dB"),
("M-1017", "ELECTRIC", False, "Cable tie 200 x 4.8 mm black, 100 pcs", "Kabelbinder 200 x 4,8 mm schwarz"),
("M-1018", "ELECTRIC", False, "Fuse 10 A slow blow 5x20 mm", "Feinsicherung 10 A träge 5x20"),
("M-1019", "LUBES", False, "Multipurpose grease cartridge 400 g", "Mehrzweckfett Kartusche 400 g"),
("M-1020", "OFFICE", False, "Ballpoint pen blue, box of 50", "Kugelschreiber blau, 50 Stück"),
]
# Labelled queries: what a buyer might type, and the products a buyer would accept as right.
QUERIES = [
("hex bolt M8", {"M-1001", "M-1002"}),
("galvanised screw 8 mm", {"M-1001", "M-1002"}),
("disposable gloves for the lab", {"M-1004", "M-1005"}),
("bearing 6204", {"M-1006", "M-1007"}),
("lubricant for the hydraulic press", {"M-1008", "M-1009"}),
("paper for the office printer", {"M-1011", "M-1012"}),
("eye protection", {"M-1010", "M-1015"}),
("hearing protection", {"M-1016"}),
("zip ties", {"M-1017"}),
("something to write with", {"M-1020"}),
]
# ------------------------------------------------------------------ getting the product data
def sample_products(language: str) -> list[dict]:
col = 3 if language == "EN" else 4
return [{"Product": p[0], "ProductGroup": p[1], "IsMarkedForDeletion": p[2], "text": p[col]}
for p in PRODUCTS]
def sandbox_products(language: str, max_rows: int) -> list[dict]:
"""Read products with their descriptions from the sandbox, page by page."""
try:
import requests
from dotenv import load_dotenv
load_dotenv()
except ImportError:
sys.exit("requests or python-dotenv is missing. Run pip install -r requirements.txt.")
key = os.environ.get("SAP_API_KEY")
if not key:
sys.exit("SAP_API_KEY is not set. Add it to .env, or run without --sandbox.")
rows, skip, page = [], 0, 50
while len(rows) < max_rows:
params = {"$select": "Product,ProductGroup,IsMarkedForDeletion,"
"to_Description/Language,to_Description/ProductDescription",
"$expand": "to_Description", "$top": str(min(page, max_rows - len(rows))),
"$skip": str(skip)}
try:
resp = requests.get(SANDBOX + "/A_Product", params=params, timeout=30,
headers={"APIKey": key, "Accept": "application/json"})
except requests.exceptions.RequestException as err:
sys.exit(f"Could not reach sandbox.api.sap.com ({type(err).__name__}). "
"Check your network or proxy, or run without --sandbox.")
if resp.status_code != 200:
sys.exit(f"The sandbox answered {resp.status_code}: {resp.text[:200]}")
batch = resp.json()["d"]["results"]
for r in batch:
texts = {d["Language"]: d["ProductDescription"] for d in r["to_Description"]["results"]}
if texts.get(language):
rows.append({"Product": r["Product"], "ProductGroup": r.get("ProductGroup") or "",
"IsMarkedForDeletion": bool(r["IsMarkedForDeletion"]),
"text": texts[language]})
if len(batch) < int(params["$top"]): # a short page means there is nothing more
break
skip += page
return rows
# ------------------------------------------------------------------ keyword search (BM25)
def tokens(text: str) -> list[str]:
return re.findall(r"[a-z0-9äöüß]+", text.lower())
class BM25:
"""Keyword ranking: rare words that appear in a short description score highest."""
def __init__(self, texts: list[str], k1: float = 1.2, b: float = 0.75):
self.docs = [Counter(tokens(t)) for t in texts]
self.lengths = [sum(d.values()) for d in self.docs]
self.avg = sum(self.lengths) / len(self.docs)
self.k1, self.b = k1, b
df = Counter(w for d in self.docs for w in d)
n = len(self.docs)
self.idf = {w: math.log(1 + (n - f + 0.5) / (f + 0.5)) for w, f in df.items()}
def scores(self, query: str) -> np.ndarray:
out = np.zeros(len(self.docs))
for i, (doc, length) in enumerate(zip(self.docs, self.lengths)):
for w in set(tokens(query)):
f = doc.get(w, 0)
if f:
norm = self.k1 * (1 - self.b + self.b * length / self.avg)
out[i] += self.idf[w] * f * (self.k1 + 1) / (f + norm)
return out
# ------------------------------------------------------------------ search by meaning
_MODELS = {} # models loaded in this run, so the model loads once
def embed(texts: list[str], model_name: str) -> np.ndarray:
"""One embedding per text, scaled to length 1, so a dot product is the cosine similarity."""
if model_name not in _MODELS:
try:
from sentence_transformers import SentenceTransformer
except ImportError:
raise SystemExit("sentence-transformers is not installed. See Set up for Unit 3, or use --offline.")
print(f"Loading {model_name} (downloads once, then uses the copy on disk)...")
_MODELS[model_name] = SentenceTransformer(model_name)
return _MODELS[model_name].encode(texts, normalize_embeddings=True)
# ------------------------------------------------------------------ ranking and fusion
def ranked(scores: np.ndarray, allowed: np.ndarray, top: int) -> list[int]:
"""Indexes of the best-scoring allowed products, best first."""
order = [i for i in np.argsort(-scores) if allowed[i]]
return order[:top]
def rrf(rankings: list[list[int]], k: int = 60) -> list[tuple[int, float]]:
"""Reciprocal rank fusion: add 1 / (k + rank) from each ranking, then sort by the total."""
fused = Counter()
for ranking in rankings:
for rank, i in enumerate(ranking, start=1):
fused[i] += 1 / (k + rank)
return fused.most_common()
class Searcher:
def __init__(self, products: list[dict], vectors: np.ndarray | None, embed_query):
self.products = products
self.bm25 = BM25([p["text"] for p in products])
self.vectors = vectors
self.embed_query = embed_query
def allowed(self, group: str | None, include_deleted: bool) -> np.ndarray:
return np.array([(include_deleted or not p["IsMarkedForDeletion"])
and (group is None or p["ProductGroup"] == group) for p in self.products])
def search(self, query: str, allowed: np.ndarray, top: int = 5, depth: int = 20,
min_score: float | None = None) -> dict:
kw = self.bm25.scores(query)
kw_ok = allowed & (kw > 0) # a keyword hit must share at least one word
results = {"keyword": [(i, kw[i]) for i in ranked(kw, kw_ok, top)]}
if self.vectors is not None:
sem = self.vectors @ self.embed_query(query) # dot product = cosine for length-1 vectors
sem_ok = allowed & (sem >= min_score) if min_score is not None else allowed
results["meaning"] = [(i, sem[i]) for i in ranked(sem, sem_ok, top)]
fused = rrf([ranked(kw, kw_ok, depth), ranked(sem, sem_ok, depth)])
results["hybrid"] = fused[:top]
return results
# ------------------------------------------------------------------ evaluation
def load_queries(path: str | None) -> list[tuple[str, set[str]]]:
if not path:
return QUERIES
with open(path, newline="", encoding="utf-8") as f:
return [(r["query"], {x.strip() for x in r["relevant"].split(";") if x.strip()})
for r in csv.DictReader(f)]
def evaluate(searcher: Searcher, queries: list[tuple[str, set[str]]], allowed: np.ndarray) -> None:
"""Hit@3: how many queries have a right product in the top 3.
MRR: average of 1 / position of the first right product (0 if none in the top 10)."""
methods = ["keyword", "meaning", "hybrid"] if searcher.vectors is not None else ["keyword"]
stats = {m: {"hits": 0, "rr": 0.0, "misses": []} for m in methods}
for query, relevant in queries:
results = searcher.search(query, allowed, top=10)
for m in methods:
ids = [searcher.products[i]["Product"] for i, _ in results[m]]
first = next((pos for pos, pid in enumerate(ids, start=1) if pid in relevant), None)
if first and first <= 3:
stats[m]["hits"] += 1
else:
stats[m]["misses"].append(query)
stats[m]["rr"] += 1 / first if first else 0.0
n = len(queries)
print(f"\nEvaluation on {n} labelled queries")
print(f" {'method':8s} hit@3 MRR")
for m in methods:
print(f" {m:8s} {stats[m]['hits']:2d}/{n} {stats[m]['rr'] / n:.2f}")
for m in methods:
if stats[m]["misses"]:
print(f" {m} missed: " + "; ".join(stats[m]["misses"]))
# ------------------------------------------------------------------ main
def show(searcher: Searcher, query: str, allowed: np.ndarray, min_score: float | None) -> None:
exact = [p for p in searcher.products if p["Product"].lower() == query.strip().lower()]
if exact:
print(f" Exact product number: {exact[0]['Product']} {exact[0]['text']}")
for method, hits in searcher.search(query, allowed, min_score=min_score).items():
print(f" By {method}:")
if not hits:
print({"keyword": " (no product shares a word with the search)",
"meaning": " (nothing scores above --min-score)",
"hybrid": " (no match from either method)"}[method])
for i, score in hits:
p = searcher.products[i]
print(f" {score:6.3f} {p['Product']:8s} {p['ProductGroup']:9s} {p['text']}")
def main() -> None:
parser = argparse.ArgumentParser(description="Search SAP-style product descriptions three ways.")
parser.add_argument("--query", default="gloves for handling chemicals", help="what to search for")
parser.add_argument("--offline", action="store_true", help="skip the embedding model; keyword search only")
parser.add_argument("--model", default=MODEL, help="Sentence Transformers model name or local folder")
parser.add_argument("--language", default="EN", choices=["EN", "DE"], help="description language")
parser.add_argument("--group", help="only products in this product group, e.g. PPE")
parser.add_argument("--include-deleted", action="store_true", help="also show products marked for deletion")
parser.add_argument("--min-score", type=float, help="hide meaning matches below this score")
parser.add_argument("--evaluate", action="store_true", help="score each method on labelled queries")
parser.add_argument("--queries", help="CSV file with columns query,relevant (IDs separated by ;)")
parser.add_argument("--sandbox", action="store_true", help="read products from SAP's sandbox (needs SAP_API_KEY)")
parser.add_argument("--max", type=int, default=200, help="with --sandbox: read at most this many products")
parser.add_argument("--save", action="store_true", help="write product_index.npz with the vectors")
args = parser.parse_args()
if args.sandbox:
products = sandbox_products(args.language, args.max)
print(f"Read {len(products)} products with a description in {args.language} from the sandbox")
if not products:
sys.exit("No descriptions in that language came back. Try --language DE or a larger --max.")
else:
products = sample_products(args.language)
print(f"Using {len(products)} made-up products, descriptions in {args.language}")
vectors, embed_query = None, None
if not args.offline:
vectors = embed([p["text"] for p in products], args.model)
embed_query = lambda q: embed([q], args.model)[0] # noqa: E731
print(f"Each description became {vectors.shape[1]} numbers.")
searcher = Searcher(products, vectors, embed_query)
allowed = searcher.allowed(args.group, args.include_deleted)
print(f"Searching {int(allowed.sum())} of {len(products)} products"
f"{' in group ' + args.group if args.group else ''}"
f"{'' if args.include_deleted else ', skipping those marked for deletion'}")
print(f"\nResults for: '{args.query}'")
show(searcher, args.query, allowed, args.min_score)
if args.evaluate:
if args.sandbox and not args.queries:
print("\n--evaluate uses labelled queries for the made-up products. "
"For sandbox data, write your own CSV and pass --queries.")
else:
evaluate(searcher, load_queries(args.queries), allowed)
if args.save:
if vectors is None:
print("\n--save needs embeddings. Run without --offline.")
else:
np.savez(HERE / "product_index.npz", products=np.array([p["Product"] for p in products]),
texts=np.array([p["text"] for p in products]), vectors=vectors,
model=args.model, language=args.language)
print(f"\nSaved {len(products)} product vectors to product_index.npz")
if args.offline:
print("\nOffline run: keyword search only. Run without --offline to search by meaning.")
if __name__ == "__main__":
main()
Using 20 made-up products, descriptions in EN
Searching 19 of 20 products, skipping those marked for deletion
Results for: 'gloves for handling chemicals'
By keyword:
1.923 M-1005 PPE Gloves nitrile large 100 pcs
1.792 M-1014 PPE Cut-resistant work gloves, level C
1.677 M-1004 PPE Nitrile gloves size L, box of 100
Evaluation on 10 labelled queries
method hit@3 MRR
keyword 6/10 0.55
keyword missed: eye protection; hearing protection; zip ties; something to write with
Offline run: keyword search only. Run without --offline to search by meaning.
What to notice:
The filter ran first. One of the twenty products, M-1013, is marked for deletion, so only 19 were searched.
Only three results. Just three descriptions share a word with the query, the word "gloves". Keyword search doesn't pad the list.
Cut-resistant gloves come second. BM25 sees "gloves" and nothing about chemicals. It can't tell a chemical glove from a cut glove.
Four of ten test queries fail. Each one uses words the descriptions don't: "protection", "zip ties", "write". That is the gap semantic search should close.
The exact token "m8x40" puts M-1001 first. The M10 bolt comes second because it shares "hex" and "bolt". The real duplicate, M-1002, comes last: it writes the size as "M8 x 40", three separate tokens.
Now include deleted products and filter on the product group:
The first status line now reads Searching 4 of 20 products in group FASTENERS, and M-1013 "Hex bolt M8x40 galv. OLD - do not use" appears in second place. In a real system that is the product a buyer must not order. Keep the deletion filter on.
Results for: 'M-1016'
Exact product number: M-1016 Ear muffs, 30 dB
By keyword:
(no product shares a word with the search)
The exact-match step catches the ID; the rankers don't, because IDs aren't in the searched text. That is by design: IDs belong in a lookup, not in an embedding.
The first lines say Loading sentence-transformers/all-MiniLM-L6-v2 ... and Each description became 384 numbers. The model loads from disk if you ran the embeddings topic; otherwise it downloads once. No account or key is needed.
What success looks like (your scores will differ):
Results for: 'gloves for handling chemicals'
By keyword:
1.923 M-1005 PPE Gloves nitrile large 100 pcs
...
By meaning:
0.xxx M-10xx PPE ...
... five lines ...
By hybrid:
0.0xx M-10xx PPE ...
... five lines ...
Evaluation on 10 labelled queries
method hit@3 MRR
keyword 6/10 0.55
meaning x/10 0.xx
hybrid x/10 0.xx
Your lines show real scores in place of x. The keyword lines stay exactly as in Step 3. Now look for four things and write them down; the exercise uses them:
Does "By meaning" fill the gaps? Check whether "hearing protection", "eye protection" and "zip ties" are still in the missed list for meaning.
Does meaning always show five results? It does, even when only two are relevant. Look at the scores of the last lines.
Which method has the best hit@3 and MRR? Hybrid often matches or beats the better single method, but not always. Twenty products and ten queries is a small test; one query changes the result by 10 points.
Where does the M10 bolt rank for "hex bolt M8" by meaning? Run python product_search.py --query "hex bolt M8" and compare with the keyword ranking.
Pick a number between the two top scores, for example halfway, and pass it as --min-score. For instance, if the canteen query topped out at 0.15 and hearing protection at 0.45:
python product_search.py --query "canteen menu for next week" --min-score 0.3
What success looks like: all three methods now answer with "no" lines:
By keyword:
(no product shares a word with the search)
By meaning:
(nothing scores above --min-score)
By hybrid:
(no match from either method)
Two queries are far too few to set a real threshold. The exercise uses your whole test set.
Every product also has a German description, as a real system would store under LanguageDE.
Run keyword search on them:
python product_search.py --offline --language DE --query "Handschuhe"
Results for: 'Handschuhe'
By keyword:
2.461 M-1005 PPE Handschuhe Nitril groß 100 St.
Only one of the three glove products appears. German joins words: "Nitrilhandschuhe" and "Schnittschutzhandschuhe" are single tokens, so "Handschuhe" doesn't match them.
Run the same search with the model:
python product_search.py --language DE --query "Handschuhe"
Check whether the other glove products appear by meaning. The model card lists all-MiniLM-L6-v2 as an English model, so don't expect much. The Sentence Transformers documentation points to multilingual models for this case; test one on your data before you choose it.
#Step 8 (optional): Search real descriptions from the sandbox
This step reads real products from SAP's sandbox with your key. Skip it if you have no key or your network blocks the sandbox; nothing later depends on it.
Check that .env in your course folder has a line SAP_API_KEY=.... The Unit 1 setup added it. If you need a key, follow the key steps in Set up your computer for this course: log on at https://api.sap.com, open an API's page and choose Show API Key.
The script asks the sandbox for A_Product with $expand=to_Description, 50 products per call, and keeps those with an English description.
What success looks like: a first line such as Read 100 products with a description in EN from the sandbox, then keyword results for the default query. The sandbox holds SAP's demo products, not gloves, so an empty keyword list is a valid result. Run it again with a word you saw in the results, using --query.
Drop --offline to add meaning and hybrid results. --evaluate won't score sandbox data without your own labelled queries, which the exercise shows how to write.
The last line reads Saved 20 product vectors to product_index.npz. Like note_vectors.npz in the embeddings topic, it stores the model name and language with the vectors, so nobody compares them with vectors from another model.
Commit the script:
git add product_search.py
git commit -m "Hybrid product search with filters and evaluation"
If Git says the path doesn't exist, you are in the course folder, not unit03; use git add unit03/product_search.py. Leave product_index.npz out of Git; --save recreates it.
This section maps the build onto SAP's services as of September 2026. Setting them up needs access you don't have yet. Unit 5 covers the generative AI hub, Unit 6 building apps on SAP BTP, and Unit 7 SAP HANA Cloud and its vector engine.
In a company system you read the same A_Product and to_Description data through the Product Master API with a communication user, as Calling your first SAP API described. Two fields from the API help keep the index honest:
LastChangeDateTime lets a nightly job fetch only products changed since the last run, instead of reloading everything.
IsMarkedForDeletion and AuthorizationGroup belong in the index as filter columns, not in the embedded text.
SAP Learning lists semantic search and similarity search among the vector engine's use cases. The pieces map onto your script like this:
Your script
SAP HANA Cloud
vectors array
A REAL_VECTOR column next to product number, group and deletion flag
embed(...) for descriptions
VECTOR_EMBEDDING(text, 'DOCUMENT', model) in the database
embed([query])
VECTOR_EMBEDDING(query, 'QUERY', model)
vectors @ query
COSINE_SIMILARITY(...)
Searcher.allowed
A WHERE clause on normal columns
--min-score
A condition on the similarity, as the CAP example uses with 0.75
The CAP documentation shows vector_embedding with 'DOCUMENT' for stored texts, QUERY for the question, and the SAP model ID SAP_GXY.20250407. A search could then look like this (sketch; it needs an SAP HANA Cloud instance with embedding enabled, and the table is yours to design):
SELECT TOP 5 "PRODUCT", "DESCRIPTION",
COSINE_SIMILARITY("EMBEDDING",
VECTOR_EMBEDDING(?, 'QUERY', 'SAP_GXY.20250407')) AS "SCORE"
FROM "PRODUCT_SEARCH"
WHERE "IS_MARKED_FOR_DELETION" = FALSE
AND "PRODUCT_GROUP" = ?
ORDER BY "SCORE" DESC;
Three cautions. CAP marks its vector_embedding support as beta and CAP Java only. The LangChain integration notes that VECTOR_EMBEDDING needs the NLP feature enabled on the instance. And model IDs, languages and limits change, so check SAP's current documentation first. Unit 7 does this in detail, including the keyword side of hybrid search in SAP HANA.
If you build in Python, the LangChain integration for SAP HANA Cloud gives you the same pattern: filter with $eq, $in or $like for the gate, specific_metadata_columns so filter fields are real columns, and an HNSW index for large catalogues.
SAP Master Data Governance builds duplicate checks into the creation of business partners, suppliers and customers. SAP Learning describes enterprise search with optional fuzzy search, or SAP HANA-based search with a ranking per hit.
AI-assisted central governance in MDG on SAP S/4HANA Cloud Private Edition, generally available per SAP's Q4 2025 release highlights, lets users search, display, create and change business partners in natural language through Joule.
Before building a custom product search, check whether one of these covers the need and is already licensed.
Authorizations. An index copied out of S/4HANA doesn't carry its authorization checks. Store fields such as AuthorizationGroup, plant or sales organization as filter columns, and apply the user's permissions in every query. Unit 11 covers agent permissions and SAP authorizations.
Freshness. Re-embed a product when its description changes, found through LastChangeDateTime. Remove or flag products marked for deletion at once, not at the next full reload.
One model per index. Store the model name and version with the vectors. Changing the model means re-embedding everything and choosing the threshold again.
Evaluation first. Build the test set from real search logs or users before tuning anything. Report hit@3 and MRR for the current search and the new one side by side. Re-run it after every change of model, text or k.
What to embed. Embed descriptions and meaningful long texts. Keep IDs, units and codes in normal columns. Adding a product group's name to the text can help short descriptions; measure it.
Languages. Search the description in the user's language, or use a multilingual model. Test each language separately; a model that works in English may fail in German.
Cost and speed. Embed each product once when it changes, not at search time. Query embedding is one short text per search.
Clean core. Read through released APIs and keep the index side by side, in SAP HANA Cloud or your own service. Don't add columns to standard S/4HANA tables for it.
Embedding product numbers. "M-1016" means nothing to the model. Use an exact lookup.
Filtering after ranking. Filter first, or deleted products can push every valid result out of the top 5.
No "no good match". Without a threshold, semantic search always proposes something.
Adding raw scores from both rankers. BM25 and cosine live on different scales. Fuse by rank with RRF, or normalize carefully.
Trusting sizes to embeddings. M8 and M10 bolts look alike to the model. Keep keyword search in the mix, or compare numbers with a rule.
Tokenization surprises. "M8x40" and "M8 x 40" are different tokens, and German compounds hide words. Test the queries your users actually type.
Tuning on the test set until it's perfect. With ten queries, you end up fitting those ten. Keep some queries aside to check the final choice.
Forgetting languages. One index in one language fails silently for users of another.
#Exercise: build a test set and choose a threshold
You will write your own labelled queries, score the three methods on them and choose a --min-score from the results. The test set is reused in Unit 7, where the same retrieval feeds a RAG assistant, and in Unit 8's evaluation harness.
In VS Code, create unit03/my_queries.csv. The first line must be exactly:
query,relevant
Add ten lines of your own queries against the twenty products in PRODUCTS. After the comma, list the product numbers that are right answers, separated by ;. For example:
gloves for oil and grease,M-1004;M-1005
something to seal a bearing,M-1007
Include at least: two queries in your own words that share no word with the right description, one with a size or number, one with a typo, and one query that no product answers. For that last one, leave the part after the comma empty, like canteen menu,.
Replace 0.3 with your value. Check whether the hit@3 dropped. If it did, your threshold hides right answers; lower it.
Create unit03/notes_search.md with three short sections:
Results: hit@3 and MRR for keyword, meaning and hybrid, with and without your threshold.
Misses: which of your queries each method missed, and why you think it did.
Decision: which method and threshold you would propose for a buyer's product search, in two sentences.
Save your work:
git add my_queries.csv notes_search.md
git commit -m "Test set and threshold for product search"
Done when:my_queries.csv has ten labelled queries including one without an answer; notes_search.md has hit@3 and MRR for all three methods with and without your threshold; and git log shows the commit.
Pick one answer for each question. The explanation appears after you choose.
1Why does the script apply the deletion and product group filters before ranking?
Answer: B. If the top 5 happen to be deleted products, filtering afterwards leaves an empty list while good matches sat at positions 6 to 10. Filtering first ranks only allowed products.
2The query "hearing protection" returns nothing by keyword but finds "Ear muffs, 30 dB" by meaning. Why?
Answer: C. A record that shares no token with the query gets a BM25 score of exactly 0. The embedding model places both texts by meaning, so they can be close with no word in common.
3Why does the hybrid ranking use RRF instead of adding BM25 and cosine scores?
Answer: A. BM25 runs from 0 upwards with no fixed top, while cosine sits between -1 and 1. RRF adds 1 / (k + rank) from each list, so only the positions matter and neither scale dominates.
4With k = 60, a product ranks 1st by keyword and 3rd by meaning. What is its RRF score?
Answer: D. Each list adds 1 divided by k plus the rank, with ranks starting at 1. So 1/61 from keyword and 1/63 from meaning, about 0.032.
5Why is "M-1002 Bolt, hexagon head, M8 x 40" last for "hex bolt M8x40" by keyword?
Answer: C. The tokenizer turns "M8x40" into one token and "M8 x 40" into three. Only M-1001 and the deleted M-1013 share "m8x40", so M-1002 matches on "bolt" alone.
6What does an MRR of 0.5 on a test set tell you?
Answer: B. MRR averages 1 divided by the position of the first right answer. A right answer always at position 2 gives 0.5; mixed positions can average to the same value.
7A client wants the product search index copied into SAP HANA Cloud. What must the design include?
Answer: D. A copied index doesn't carry S/4HANA's authorization checks. Storing fields like AuthorizationGroup as normal columns and filtering on them per user keeps results within what each user may see.
8You switch to a multilingual model and hit@3 rises from 6/10 to 8/10. What do you do before shipping?
Answer: B. Vectors from two models can't be mixed, and a threshold chosen for one model means nothing for another. Each language should be tested separately, and keyword search still keeps codes and sizes exact.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
@sap/cloud-sdk-vdm-product-service (npm, SAP Cloud SDK, version 2.1.0)— generated from the service metadata of API_PRODUCT_SRV; service path, A_Product fields (ProductGroup, IsMarkedForDeletion, AuthorizationGroup, LastChangeDateTime), to_Description, and A_ProductDescription with Product, Language (2) and ProductDescription (40 characters); package marked deprecated
Vector Embeddings (CAP documentation)— vector_embedding with 'DOCUMENT' for stored texts and TextType.QUERY for questions; model SAP_GXY.20250407; similarity threshold in the query; CAP support marked beta and CAP Java only
Semantic Search (Sentence Transformers documentation)— symmetric vs asymmetric search; util.semantic_search with top_k; exhaustive search fine up to about a million entries, approximate nearest neighbour beyond; multilingual models exist