Score a RAG search with recall, precision, MRR and nDCG on a labeled set of SAP-style questions, and compare chunking, hybrid search and re-ranking with numbers.
A RAG assistant answers in two steps. First a search finds a few passages. Then a model writes an answer from them. If the search brings back the wrong passages, the best model in the world writes a confident wrong answer.
So you test the search on its own. You write down realistic questions. For each one, an expert marks which documents actually answer it. Then you run the search and count:
Did the right document show up at all in the first few results? (hit rate, recall)
How much of what came back was useful? (precision)
How high up was the first right answer? (MRR)
Were the best documents at the very top? (nDCG)
Those numbers turn "the search feels better" into "the search found the right document for 17 of 20 questions, up from 15". With them, a team can pick a chunking method, a search method or a re-ranker on evidence instead of on a demo.
Most bad RAG answers start as bad search results. Take the running procure-to-pay example: an accounts payable clerk asks the assistant why an invoice for purchase order 4500017311 is blocked. If the search returns the general three-way match guide but misses the case note about the 10 missing pieces, the answer is vague. If it returns the price variance guide instead, the answer is wrong.
Measuring retrieval separately pays off in three ways:
Faster diagnosis. When an answer is wrong, you can tell whether the search or the model failed. They have different fixes and different owners.
Cheaper choices. Unit 7 offered many options: chunk sizes, keyword or vector search, hybrid fusion, re-ranking. Each costs build time, compute or licences. A labeled question set tells you which ones earn their cost on your documents.
Safe change. Re-indexing documents, switching embedding models or adding a new document source can quietly make search worse. Rerunning the same questions after each change catches it, the same regression idea as LLM evaluation fundamentals.
The main cost is expert time to label questions. A first set of a few dozen questions, labeled by someone who knows the process, is a modest piece of work. It is reused for every later decision.
As of the SAP AI Core service guide dated 4 September 2026, the document grounding part of the generative AI hub covers the retrieval side of RAG:
Pipelines take documents from a source and vectorize them. SAP's Python SDK documentation shows sources such as S3, and a vector API for feeding chunks directly.
A vector API manages collections and documents in the vector database. Each document carries chunks and metadata.
A retrieval API "searches data repositories" and returns the relevant chunks for a query. The guide also mentions metadata filtering and merging results from several repositories.
In the orchestration service, grounding has a setting for how many chunks to pass to the model (max_chunk_count in the Python SDK). That number is the k this topic measures.
The generative AI hub is part of SAP AI Core's extended service plan.
SAP AI Core also has Evaluations, which compares prompt and model configurations with system-defined and custom LLM-as-a-judge metrics. In the SAP documentation we could open, we found no built-in metric that scores the retrieval step itself with recall or nDCG. So plan to measure retrieval yourself: send your labeled questions to the retrieval API and score the returned chunks with the methods in this topic. The same method works for SAP HANA Cloud's vector engine (covered in Unit 7) or any other search.
A typical result, from the lab in this topic: 20 labeled questions about made-up SAP help notes, four search methods, looking at the top 3 results.
Method
Right note in top 3 (hit rate)
Share of top 3 that is useful (precision)
First right note, on average (MRR)
Best notes at the top (nDCG)
Keyword search
75%
27%
0.74
0.64
Vector search
80%
33%
0.74
0.70
Hybrid (both, fused)
85%
35%
0.80
0.73
Hybrid + re-ranker
100%
42%
1.00
0.93
How a leader should read it:
Hit rate and recall are about safety. If the right document isn't retrieved, the model can't use it. For RAG this is usually the first number to watch.
Precision is about noise and cost. Every useless passage costs tokens and can distract the model.
MRR and nDCG are about order. They matter when only the top one or two passages reach the model, or when a person reads the list.
Averages hide patterns. In the lab, keyword search was strong on exact codes like PRC-112 and weak on paraphrases; vector search was the opposite. Ask for results split by type of question.
A small set gives rough numbers. With 20 questions, a gap of a few points can be luck. Ask whether a difference was tested question by question (the deep layer shows how).
A perfect score is a warning. The lab's toy re-ranker was written while looking at these very questions, so its 100% is too good. New, unseen questions are the real test.
"If the answers look good, the search is fine." A strong model can paper over weak search on easy questions and fail on hard ones. Measure the search step on its own.
"Vector search is always better than keyword search." Embeddings struggle with exact codes, IDs and document numbers. In the lab, vector search found none of the identifier questions.
"Hybrid search always wins." Fusing a weak list with a strong one can drag the strong one down. In the lab, hybrid scored lower than vector search alone on paraphrased questions.
"One number is enough." Hit rate, precision and nDCG answer different questions. A change can raise one and lower another.
"More test questions can wait." Twenty questions are a good start, but they can't separate options that are close. Grow the set from real failures.
"Unlabeled means irrelevant." If a document was never shown to the labeler, it counts as wrong even when it is right. Label everything any method returns near the top.
Pick one answer for each question. The explanation appears after you choose.
1Why should a team measure the search step of a RAG assistant separately from the answers?
Answer: C. RAG answers depend on the passages the search returns. Measuring search on its own tells you whether to fix the search or the model, which are different jobs with different owners.
2A team wants to know if the right document reaches the model at all. Which number fits best?
Answer: A. Hit rate and recall ask whether the relevant documents are in the top k at all. If they aren't, the model can't use them. Precision and MRR matter too, but they answer questions about noise and order.
3Which test set gives the most trustworthy retrieval numbers?
Answer: C. Numbers only predict real use if the questions look like real use and the labels come from someone who knows the process. Questions used for tuning flatter the system, as the lab's perfect toy re-ranker shows.
4In the lab, keyword search scored well on codes like PRC-112 and badly on paraphrases. What is the leadership lesson?
Answer: D. An average can hide a method that fails a whole group of questions. Splitting by question type shows which method suits which need, and why hybrid search exists.
5Option B beats option A by two points on 20 test questions. What should you ask before switching?
Answer: B. With a small set, a small gap can come from a few lucky questions. A paired test, such as the bootstrap interval or the randomization test in the deep layer, shows whether the gap is likely to be real.
6What does SAP's document grounding offer for evaluating retrieval, as described in this topic?
Answer: C. The SAP AI Core guide describes a retrieval API that returns relevant chunks for a query. We found no built-in retrieval metric in the documentation we could open, so teams score the returned chunks themselves.
7A colleague says a document that was never labeled must be irrelevant. Why is that risky?
Answer: D. Metrics treat anything not marked relevant as irrelevant. Pooling, where a labeler judges the top results from every method, closes those gaps so no method is punished for finding something new.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Retrieval evaluation is a fixed exam with an answer key. The exam is a list of questions. The answer key, called qrels (query relevance judgments), says which documents are relevant to each question, and how much. A search method sits the exam and hands in a run: for each question, a ranked list of documents. A metric compares the run with the key.
Everything else is detail:
The answer key is written once, by people who know the process, and reused for every method and every change.
The metric decides what you care about: finding everything (recall), avoiding noise (precision), the first right answer (MRR), or the whole order (nDCG).
A comparison between two methods is only trustworthy if you look question by question, because the same questions were asked of both.
This is the information retrieval field's classic "test collection" method, applied to the search step of a RAG system. You met its building blocks in Set up for Unit 8, where ranx computed three metrics on a toy run.
flowchart LR
Q[Labeled questions<br/>qrels] --> S[Search method]
D[(Notes or chunks)] --> S
S --> R[Run: ranked<br/>note IDs]
R --> M[Metrics<br/>per question]
Q --> M
M --> A[Averages and<br/>splits by type]
M --> C[Paired comparison<br/>A vs B]
Each test question has an ID, the text, and graded labels. This lab uses two grades: 2 means the note answers the question, 1 means it helps. Notes not listed count as 0. Two more fields make the set useful for diagnosis: the process (order-to-cash, procure-to-pay, plan-to-produce, record-to-report) and the kind of question:
identifier: an exact code or number, such as PRC-112 or 4500017311
paraphrase: the meaning is there, the words are not ("client purchase frozen because they owe us money")
mixed: some shared words and some meaning
Where do good questions come from? In order of preference: real user questions from search logs or a pilot; ticket subjects and incident descriptions; questions process experts write from memory; and last, questions a model generates from the documents. Model-written questions tend to reuse the document's own words, which flatters keyword search.
You label documents, not chunks. Chunks change every time you change the chunking, and you don't want to relabel. When the search returns chunks, map each chunk back to its source document and keep the document's best position. That is what lets this lab compare three chunking choices on one answer key. The cost: a document-level label can't tell whether the right paragraph of a long document was returned. If your documents are long manuals, label at section level instead and keep section IDs stable.
For one question, let the ranked results be r1, r2, … and let rel(d) be the grade of document d (0 if unlabeled).
Metric
Formula in words
Range
Uses grades?
Sensitive to order?
Hit rate@k
1 if any relevant document is in the top k, else 0
0 or 1
No
No
Precision@k
relevant documents in the top k, divided by k
0–1
No
No
Recall@k
relevant documents in the top k, divided by all relevant documents
0–1
No
No
Reciprocal rank
1 divided by the position of the first relevant document
0–1
No
Yes
nDCG@k
sum of rel / log2(position + 1) over the top k, divided by the same sum for the ideal order
0–1
Yes
Yes
Each is computed per question and then averaged. The average reciprocal rank is MRR.
Some properties worth knowing:
Precision@k has a ceiling. If a question has one relevant note and k = 3, the best possible precision@3 is 0.33. Manning, Raghavan and Schütze point out that precision at k averages badly because the number of relevant documents varies so much between questions.
Recall@k needs the full answer key. If labels are missing, recall looks better or worse than it is.
nDCG is the field's usual headline metric. The BEIR benchmark picked nDCG@10 because it handles both yes/no and graded labels. Its authors note that precision and recall ignore order, and MRR cannot use grades.
There are two nDCG gain formulas. ranx's ndcg uses the grade as the gain; its ndcg_burges uses 2^grade − 1, which rewards highly relevant documents more. Say which you use when you report numbers.
A worked example. A question has labels {D07: 2, D06: 1}. The search returns D08, D07, D06.
Hit rate@3 = 1 (D07 is there). Recall@3 = 2/2 = 1.0. Precision@3 = 2/3 = 0.67.
Reciprocal rank = 1/2 (first relevant note is second).
Labeling every note for every question doesn't scale. The standard shortcut is pooling: run every method you want to compare, take the union of their top k results for each question, and label only those. Anything outside every pool is assumed irrelevant. The lab's pool command lists exactly the (question, note) pairs that some method put in its top k but nobody has labeled yet. When you add a new method later, pool again: otherwise it gets punished for finding relevant notes the old methods missed.
Two methods answer the same questions, so compare them pairwise: for each question, the score of B minus the score of A. Then:
Count wins, losses and ties. A method that wins 8 and loses 0 is different from one that wins 10 and loses 9.
Bootstrap: resample the questions with replacement 2,000 times and look at the spread of the mean difference. If the middle 95% of that spread is entirely above zero, B is very likely better on questions like these.
Randomization test: ranx's compare function offers Fisher's randomization test (stat_test="fisher") and a paired Student's t-test (the default, "student"). Its report marks a significant win with a superscript letter.
None of this fixes a biased question set. Statistics tell you the result is unlikely to be luck on these questions, not that the questions represent your users.
#Build it yourself: score four search methods and three chunking choices
Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder with its .venv, and install ranx. For the real embedding model and re-ranker you also need sentence-transformers from Set up for Unit 3; without it, use --offline on every command.
You will build retrieval_eval.py. It holds 18 made-up SAP help notes and 20 labeled questions. It searches the notes four ways, the methods from Unit 7: keyword (BM25), vector, hybrid with reciprocal rank fusion, and hybrid plus a re-ranker. It scores each with five metrics, compares chunking choices, tests whether a difference is luck, and lists results that still need labels. It carries its own small copy of the notes, so it runs even if you skipped a Unit 7 lab.
flowchart LR
I[init<br/>questions file] --> E[eval<br/>4 methods]
E --> C[chunking<br/>3 choices]
C --> P[compare<br/>A vs B]
P --> L[pool<br/>labels to add]
L --> E
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
If the prompt doesn't start with (.venv), turn it on:
Windows (PowerShell):
.venv\Scripts\Activate.ps1
macOS / Linux:
source .venv/bin/activate
Check that ranx is installed:
pip show ranx
You should see Name: ranx and a Version: line. If you see WARNING: Package(s) not found, repeat Step 2 of Set up for Unit 8. The script also works without ranx; only the --ranx options need it.
Run every command in this topic from the course folder.
In VS Code, right-click the unit08 folder, choose New File, name it retrieval_eval.py, paste the code below and save.
"""Unit 8: evaluate retrieval with a labeled query set, and compare search choices with numbers.
A small library of made-up SAP help notes is searched four ways (keyword BM25, vector, hybrid
fusion, hybrid plus a re-ranker) and chunked three ways. Every choice is scored on the same
labeled questions with hit rate, precision, recall, MRR and nDCG, and two choices can be
compared question by question to see whether a difference is real or luck.
How to run (from your course folder, with .venv turned on):
python unit08/retrieval_eval.py init # write the labeled questions to a file
python unit08/retrieval_eval.py eval # score every search method
python unit08/retrieval_eval.py chunking # score three chunking choices
python unit08/retrieval_eval.py compare hybrid rerank # is one method really better?
python unit08/retrieval_eval.py pool # results nobody has labeled yet
python unit08/retrieval_eval.py eval --report # also save unit08/retrieval_report.md
Add --offline to any command to use toy "concept" embeddings and a toy re-ranker (no download).
Add --ranx to eval or compare to check the numbers with the ranx library.
"""
import argparse
import json
import math
import random
import re
import sys
from collections import Counter
from datetime import date
from pathlib import Path
HERE = Path(__file__).resolve().parent
QUERIES_FILE = HERE / "retrieval_queries.jsonl"
REPORT_FILE = HERE / "retrieval_report.md"
EMBED_MODEL = "sentence-transformers/all-MiniLM-L6-v2" # the Unit 3 model
RERANK_MODEL = "cross-encoder/ms-marco-MiniLM-L6-v2" # the Unit 7 re-ranker
METHODS = ["bm25", "vector", "hybrid", "rerank"]
# ---------------------------------------------------------------------------
# 1. The library: made-up SAP help notes, each with a title and a few paragraphs.
# ---------------------------------------------------------------------------
DOCS = {
"D01": ("Releasing a sales order blocked by the credit check", "order-to-cash", [
"A sales order is blocked for delivery when the customer's credit exposure is above the "
"credit limit. Exposure includes open orders, deliveries and unpaid invoices.",
"The credit analyst reviews the open items. If the overrun is below 2 percent the analyst "
"releases the order; above that the credit manager approves the release."]),
"D02": ("Delivery blocks set by sales", "order-to-cash", [
"Sales can put a delivery block on an order, for example while export papers are missing "
"or because the customer asked to wait.",
"Once the reason is cleared, remove the delivery block in the order. The delivery is "
"created in the next delivery run."]),
"D03": ("Billing blocks and price checks", "order-to-cash", [
"A billing block stops the invoice for a delivered order, usually because a price or "
"discount has to be checked first.",
"After the pricing review, release the billing block so the order goes to billing."]),
"D04": ("Pricing error PRC-112", "order-to-cash", [
"Error PRC-112 means a mandatory price condition is missing on the order item.",
"Maintain the condition record for the customer and material, then redetermine prices "
"on the order."]),
"D05": ("Incomplete sales orders", "order-to-cash", [
"An order is incomplete when mandatory data is missing, such as the payment terms or the "
"ship-to party. Incomplete orders cannot be delivered or billed.",
"Open the incompletion log on the order to see which fields are missing and fill them."]),
"D06": ("Three-way match for supplier invoices", "procure-to-pay", [
"Each supplier invoice is checked against the purchase order and the goods receipt.",
"A quantity or price variance above the tolerance blocks the invoice for payment until "
"someone resolves it."]),
"D07": ("Case: quantity variance on purchase order 4500017311", "procure-to-pay", [
"The supplier delivered 90 of 100 pieces but invoiced all 100.",
"The invoice was blocked for payment. Accounts payable asked the supplier for a credit "
"memo for the 10 missing pieces."]),
"D08": ("Price variance on supplier invoices", "procure-to-pay", [
"The vendor billed a higher price than the one agreed on the purchase order.",
"If the price variance is above tolerance, the buyer contacts the vendor and asks for a "
"corrected invoice or a credit memo."]),
"D09": ("Releasing invoices blocked for payment", "procure-to-pay", [
"When the variance is resolved, the accounts payable clerk releases the blocked invoice.",
"Released invoices are picked up by the next payment run if they are due."]),
"D10": ("Payment terms and the payment run", "procure-to-pay", [
"The due date of an invoice is calculated from the baseline date and the payment terms of "
"the supplier.",
"The payment run pays every invoice that is due and not blocked."]),
"D11": ("Invoice arrives before the goods receipt", "procure-to-pay", [
"With goods-receipt-based invoice verification, an invoice cannot be posted for payment "
"until the goods receipt is posted.",
"Ask the warehouse to post the goods receipt, then the invoice can be matched."]),
"D12": ("Daily review of MRP exceptions", "plan-to-produce", [
"The planning run flags exception messages: missing parts, late receipts and orders to "
"reschedule.",
"Planners review the exceptions every morning, starting with critical components."]),
"D13": ("Case: shortage of component RM-4711", "plan-to-produce", [
"The planning run showed a shortage of component RM-4711 for production in week 38.",
"The planner moved the production order forward and expedited the supplier."]),
"D14": ("Reschedule in and reschedule out", "plan-to-produce", [
"Reschedule in means a receipt should arrive earlier than planned to cover demand.",
"Reschedule out means a receipt arrives too early and can be pushed later."]),
"D15": ("Safety stock in planning", "plan-to-produce", [
"Safety stock is a buffer quantity that planning keeps for uncertain demand.",
"When available stock falls below safety stock, the planning run creates a planned order."]),
"D16": ("Duplicate check when creating a supplier", "procure-to-pay", [
"Before creating a new supplier in master data, check for duplicates by tax number and "
"bank account.",
"Duplicate suppliers cause double payments and wrong spend reports."]),
"D17": ("Approving changes to customer bank details", "order-to-cash", [
"A change to a customer's bank details needs a second person to approve it.",
"This control reduces the risk of fraud in refunds and credit memos."]),
"D18": ("Month-end reconciliation of the GR/IR account", "record-to-report", [
"The GR/IR clearing account holds goods receipts that are not yet invoiced and invoices "
"not yet received.",
"At month end, accountants analyze open GR/IR items and clear or adjust them."]),
}
# ---------------------------------------------------------------------------
# 2. The labeled questions. Grades: 2 = answers it, 1 = helps. Unlisted notes count as 0.
# "kind" says what the question tests: an exact identifier, a paraphrase, or a mix.
# ---------------------------------------------------------------------------
BUILT_IN_QUERIES = [
{"id": "q01", "kind": "identifier", "process": "order-to-cash", "text": "PRC-112",
"relevant": {"D04": 2, "D03": 1}},
{"id": "q02", "kind": "identifier", "process": "procure-to-pay", "text": "4500017311",
"relevant": {"D07": 2}},
{"id": "q03", "kind": "identifier", "process": "plan-to-produce", "text": "RM-4711 status",
"relevant": {"D13": 2, "D12": 1}},
{"id": "q04", "kind": "identifier", "process": "record-to-report", "text": "GR/IR open items",
"relevant": {"D18": 2}},
{"id": "q05", "kind": "paraphrase", "process": "order-to-cash",
"text": "client purchase frozen because they owe us money",
"relevant": {"D01": 2}},
{"id": "q06", "kind": "paraphrase", "process": "procure-to-pay",
"text": "supplier sent fewer pieces than we were billed for",
"relevant": {"D07": 2, "D06": 1}},
{"id": "q07", "kind": "paraphrase", "process": "procure-to-pay",
"text": "vendor charged more than we agreed",
"relevant": {"D08": 2, "D06": 1}},
{"id": "q08", "kind": "paraphrase", "process": "plan-to-produce",
"text": "lacking parts to build the product",
"relevant": {"D13": 2, "D12": 1}},
{"id": "q09", "kind": "paraphrase", "process": "plan-to-produce",
"text": "what should a planner look at first thing each day",
"relevant": {"D12": 2}},
{"id": "q10", "kind": "paraphrase", "process": "order-to-cash",
"text": "why can we not send the goods to the client",
"relevant": {"D02": 2, "D01": 2, "D05": 1}},
{"id": "q11", "kind": "paraphrase", "process": "procure-to-pay",
"text": "when does the supplier get their money",
"relevant": {"D10": 2, "D09": 1}},
{"id": "q12", "kind": "paraphrase", "process": "procure-to-pay",
"text": "stop paying the same vendor twice",
"relevant": {"D16": 2}},
{"id": "q13", "kind": "paraphrase", "process": "order-to-cash",
"text": "who signs off when a client changes where refunds go",
"relevant": {"D17": 2}},
{"id": "q14", "kind": "paraphrase", "process": "plan-to-produce",
"text": "buffer quantity for uncertain demand",
"relevant": {"D15": 2}},
{"id": "q15", "kind": "mixed", "process": "order-to-cash",
"text": "who may release a credit block above 2 percent",
"relevant": {"D01": 2}},
{"id": "q16", "kind": "mixed", "process": "procure-to-pay",
"text": "release invoice blocked for payment after variance",
"relevant": {"D09": 2, "D06": 1, "D07": 1}},
{"id": "q17", "kind": "mixed", "process": "procure-to-pay",
"text": "invoice cannot be posted, no goods receipt yet",
"relevant": {"D11": 2, "D06": 1}},
{"id": "q18", "kind": "mixed", "process": "plan-to-produce",
"text": "meaning of reschedule out",
"relevant": {"D14": 2}},
{"id": "q19", "kind": "mixed", "process": "order-to-cash",
"text": "order missing payment terms cannot be delivered",
"relevant": {"D05": 2}},
{"id": "q20", "kind": "mixed", "process": "order-to-cash",
"text": "invoice for delivered order not created, price must be checked",
"relevant": {"D03": 2}},
]
# ---------------------------------------------------------------------------
# 3. Text helpers, keyword search (BM25) and toy "concept" embeddings
# ---------------------------------------------------------------------------
STOPWORDS = set("a an and are as at be because by can each for from has have if in is it its "
"of on or our so that the their them then they this to until us was we what "
"when where which who why with not no yet more than".split())
TOKEN = re.compile(r"[a-z0-9]+(?:[-/][a-z0-9]+)*")
def tokens(text: str) -> list:
return [t for t in TOKEN.findall(text.lower()) if t not in STOPWORDS]
class BM25:
"""Classic BM25 keyword ranking (k1 = 1.5, b = 0.75) over a list of passages."""
def __init__(self, passages: list, k1: float = 1.5, b: float = 0.75):
self.docs = [Counter(tokens(p)) for p in passages]
self.lengths = [sum(d.values()) for d in self.docs]
self.avg = sum(self.lengths) / len(self.lengths)
df = Counter(t for d in self.docs for t in d)
n = len(self.docs)
self.idf = {t: math.log(1 + (n - f + 0.5) / (f + 0.5)) for t, f in df.items()}
self.k1, self.b = k1, b
def scores(self, query: str) -> list:
out = []
for d, length in zip(self.docs, self.lengths):
s = 0.0
for t in tokens(query):
if t in d:
tf = d[t]
s += self.idf[t] * tf * (self.k1 + 1) / (
tf + self.k1 * (1 - self.b + self.b * length / self.avg))
out.append(s)
return out
# Toy embeddings: each word maps to one or more "concepts"; a text becomes a vector of concept
# counts. Different words with the same meaning share a concept, which is what real
# embeddings learn from data. Used only with --offline.
CONCEPTS = {
"money_owed": "credit exposure limit owe owes unpaid overrun debt",
"blocked": "blocked block blocks frozen stuck held hold stop stops stopped cannot",
"customer": "customer customer's client sales order purchase",
"supplier": "supplier vendor suppliers seller",
"invoice": "invoice invoices invoiced billed bill billing charged",
"price": "price prices pricing charged discount condition",
"quantity": "quantity pieces fewer delivered missing",
"payment": "payment pay paid paying money payments refunds",
"approve": "approve approves approval release releases releasing signs sign-off",
"parts": "parts component components shortage lacking material",
"plan": "planning planner planned mrp exceptions reschedule",
"daily": "daily morning day every first",
"build": "build production produce product assembly",
"ship": "delivery deliver delivered send goods ship ship-to",
"duplicate": "duplicate duplicates double twice same",
"bank": "bank refunds fraud",
"buffer": "buffer safety uncertain",
"receipt": "receipt receipts warehouse",
"early_late": "earlier later early late forward arrive",
"month_end": "month-end month end reconciliation clearing accountants",
}
WORD_TO_CONCEPTS = {}
for _concept, _words in CONCEPTS.items():
for _w in _words.split():
WORD_TO_CONCEPTS.setdefault(_w, []).append(_concept)
CONCEPT_INDEX = {c: i for i, c in enumerate(CONCEPTS)}
def toy_embed(texts: list) -> list:
vectors = []
for text in texts:
v = [0.0] * len(CONCEPTS)
for w in TOKEN.findall(text.lower()):
for c in WORD_TO_CONCEPTS.get(w, []):
v[CONCEPT_INDEX[c]] += 1.0
norm = math.sqrt(sum(x * x for x in v)) or 1.0
vectors.append([x / norm for x in v])
return vectors
def real_embed(texts: list) -> list:
try:
from sentence_transformers import SentenceTransformer
except ImportError:
sys.exit("sentence-transformers is not installed (Set up for Unit 3). Add --offline to use toy embeddings.")
try:
model = SentenceTransformer(EMBED_MODEL)
except Exception as e: # usually no network or a blocked download
sys.exit(f"Could not load {EMBED_MODEL} ({type(e).__name__}). Check your network, or add --offline.")
return [list(map(float, v)) for v in model.encode(texts, normalize_embeddings=True)]
def cosine(a: list, b: list) -> float:
return sum(x * y for x, y in zip(a, b))
# ---------------------------------------------------------------------------
# 4. Chunking: turn the notes into passages. Each passage remembers its note ID.
# ---------------------------------------------------------------------------
def make_chunks(strategy: str, size: int = 12) -> list:
"""Return (doc_id, text) pairs.
whole: one passage per note (title + all paragraphs)
sections: one passage per paragraph, with the note title in front
fixed: windows of `size` words cut across the note, no title (a blind cut)"""
chunks = []
for doc_id, (title, _process, paras) in DOCS.items():
if strategy == "whole":
chunks.append((doc_id, title + ". " + " ".join(paras)))
elif strategy == "sections":
chunks.extend((doc_id, f"{title}. {p}") for p in paras)
elif strategy == "fixed":
words = " ".join(paras).split()
chunks.extend((doc_id, " ".join(words[i:i + size])) for i in range(0, len(words), size))
else:
raise ValueError(strategy)
return chunks
# ---------------------------------------------------------------------------
# 5. The search pipelines. Each returns a ranked list of note IDs, best first.
# ---------------------------------------------------------------------------
def best_per_doc(chunks: list, scores: list) -> dict:
"""Several passages can come from one note: keep the note's best passage score."""
best = {}
for (doc_id, _text), s in zip(chunks, scores):
if doc_id not in best or s > best[doc_id]:
best[doc_id] = s
return best
def rank(scores: dict) -> list:
return sorted(scores, key=lambda d: (-scores[d], d))
class Searcher:
def __init__(self, chunks: list, offline: bool):
self.chunks = chunks
self.offline = offline
texts = [t for _d, t in chunks]
self.bm25 = BM25(texts)
self.embed = toy_embed if offline else real_embed
self.vectors = self.embed(texts)
self.cross = None
def bm25_rank(self, q: str) -> list:
scores = best_per_doc(self.chunks, self.bm25.scores(q))
return [d for d in rank(scores) if scores[d] > 0] # keyword search can return nothing
def vector_rank(self, q: str) -> list:
qv = self.embed([q])[0]
scores = best_per_doc(self.chunks, [cosine(qv, v) for v in self.vectors])
if self.offline: # a toy vector with no known concept matches nothing: return nothing
return [d for d in rank(scores) if scores[d] > 0]
return rank(scores)
def hybrid_rank(self, q: str, rrf_k: int = 60) -> list:
"""Reciprocal rank fusion: each list adds 1 / (rrf_k + position) to a note's score."""
fused = Counter()
for ranked in (self.bm25_rank(q), self.vector_rank(q)):
for pos, d in enumerate(ranked, 1):
fused[d] += 1 / (rrf_k + pos)
return rank(dict(fused))
def rerank(self, q: str, candidates: int = 8) -> list:
"""Stage 2: re-score the top hybrid candidates by reading query and passage together."""
shortlist = self.hybrid_rank(q)[:candidates]
passages = {d: " ".join([DOCS[d][0]] + DOCS[d][2]) for d in shortlist}
if self.offline:
qv = toy_embed([q])[0]
qt = set(tokens(q))
scored = {d: cosine(qv, toy_embed([p])[0]) + 0.3 * len(qt & set(tokens(p))) / max(len(qt), 1)
for d, p in passages.items()}
else:
if self.cross is None:
try:
from sentence_transformers import CrossEncoder
except ImportError:
sys.exit("sentence-transformers is not installed. Add --offline to use the toy re-ranker.")
try:
self.cross = CrossEncoder(RERANK_MODEL)
except Exception as e: # usually no network or a blocked download
sys.exit(f"Could not load {RERANK_MODEL} ({type(e).__name__}). Check your network, or add --offline.")
pairs = [(q, passages[d]) for d in shortlist]
scored = dict(zip(shortlist, map(float, self.cross.predict(pairs))))
reordered = rank(scored)
return reordered + [d for d in self.hybrid_rank(q) if d not in reordered]
def run(self, method: str, q: str) -> list:
return {"bm25": self.bm25_rank, "vector": self.vector_rank,
"hybrid": self.hybrid_rank, "rerank": self.rerank}[method](q)
# ---------------------------------------------------------------------------
# 6. The metrics, in plain Python. `ranked` = note IDs best first; `rel` = {note: grade}.
# ---------------------------------------------------------------------------
def hit_at_k(ranked, rel, k):
return 1.0 if any(d in rel for d in ranked[:k]) else 0.0
def precision_at_k(ranked, rel, k):
return sum(1 for d in ranked[:k] if d in rel) / k
def recall_at_k(ranked, rel, k):
return sum(1 for d in ranked[:k] if d in rel) / len(rel)
def reciprocal_rank(ranked, rel):
for pos, d in enumerate(ranked, 1):
if d in rel:
return 1 / pos
return 0.0
def ndcg_at_k(ranked, rel, k):
dcg = sum(rel.get(d, 0) / math.log2(i + 2) for i, d in enumerate(ranked[:k]))
ideal = sorted(rel.values(), reverse=True)[:k]
idcg = sum(g / math.log2(i + 2) for i, g in enumerate(ideal))
return dcg / idcg if idcg else 0.0
def per_query(ranked, rel, k):
return {f"hit@{k}": hit_at_k(ranked, rel, k), f"P@{k}": precision_at_k(ranked, rel, k),
f"R@{k}": recall_at_k(ranked, rel, k), "MRR": reciprocal_rank(ranked, rel),
f"nDCG@{k}": ndcg_at_k(ranked, rel, k)}
def mean(values):
values = list(values)
return sum(values) / len(values) if values else 0.0
# ---------------------------------------------------------------------------
# 7. Loading the labeled questions
# ---------------------------------------------------------------------------
def load_queries() -> list:
if not QUERIES_FILE.exists():
return BUILT_IN_QUERIES
queries = []
for n, line in enumerate(QUERIES_FILE.read_text(encoding="utf-8").splitlines(), 1):
if not line.strip():
continue
try:
q = json.loads(line)
except json.JSONDecodeError as e:
sys.exit(f"{QUERIES_FILE.name} line {n} is not valid JSON: {e}")
unknown = [d for d in q.get("relevant", {}) if d not in DOCS]
if not q.get("relevant") or unknown:
sys.exit(f"{QUERIES_FILE.name} line {n} ({q.get('id')}): needs 'relevant' with known note IDs; unknown: {unknown}")
queries.append(q)
return queries
def run_all(searcher, queries, methods):
return {m: {q["id"]: searcher.run(m, q["text"]) for q in queries} for m in methods}
def score_table(runs, queries, k):
table = {}
for method, results in runs.items():
rows = [per_query(results[q["id"]], q["relevant"], k) for q in queries]
table[method] = {metric: mean(r[metric] for r in rows) for metric in rows[0]}
return table
def print_table(table, title):
metrics = list(next(iter(table.values())))
print(title)
print(f"{'':10}" + "".join(f"{m:>9}" for m in metrics))
for name, row in table.items():
print(f"{name:10}" + "".join(f"{row[m]:9.3f}" for m in metrics))
def ranx_check(runs, queries, k):
try:
from ranx import Qrels, Run, evaluate
except ImportError:
print("\nranx is not installed; skipping the cross-check (see Set up for Unit 8).")
return
import warnings
warnings.filterwarnings("ignore", message="unsafe cast")
qrels = Qrels({q["id"]: q["relevant"] for q in queries})
print(f"\nranx cross-check (hit_rate@{k}, precision@{k}, recall@{k}, mrr, ndcg@{k}):")
for method, results in runs.items():
# ranx wants scores, not positions: give the first result the highest score.
run = Run({qid: {d: float(len(r) - i) for i, d in enumerate(r)} for qid, r in results.items()})
scores = evaluate(qrels, run, [f"hit_rate@{k}", f"precision@{k}", f"recall@{k}", "mrr", f"ndcg@{k}"])
print(f"{method:10}" + "".join(f"{v:9.3f}" for v in scores.values()))
# ---------------------------------------------------------------------------
# 8. Commands
# ---------------------------------------------------------------------------
def cmd_init(args):
if QUERIES_FILE.exists() and not args.force:
sys.exit(f"{QUERIES_FILE.name} already exists. Add --force to overwrite it with the built-in questions.")
with QUERIES_FILE.open("w", encoding="utf-8") as f:
for q in BUILT_IN_QUERIES:
f.write(json.dumps(q) + "\n")
kinds = Counter(q["kind"] for q in BUILT_IN_QUERIES)
print(f"Wrote {len(BUILT_IN_QUERIES)} labeled questions to unit08/{QUERIES_FILE.name}")
print(" by kind: " + ", ".join(f"{k} {n}" for k, n in kinds.items()))
print(f" notes in the library: {len(DOCS)}")
def cmd_eval(args):
queries = load_queries()
searcher = Searcher(make_chunks(args.chunking), args.offline)
runs = run_all(searcher, queries, METHODS)
table = score_table(runs, queries, args.k)
print(f"{len(queries)} questions, {len(DOCS)} notes, chunking '{args.chunking}', "
f"{'toy offline models' if args.offline else 'real models'}\n")
print_table(table, "Average over all questions:")
kinds = sorted({q.get("kind", "other") for q in queries})
print(f"\nnDCG@{args.k} by kind of question:")
print(f"{'':10}" + "".join(f"{k:>12}" for k in kinds))
for m in METHODS:
cells = []
for kind in kinds:
subset = [q for q in queries if q.get("kind", "other") == kind]
cells.append(mean(ndcg_at_k(runs[m][q["id"]], q["relevant"], args.k) for q in subset))
print(f"{m:10}" + "".join(f"{c:12.3f}" for c in cells))
if args.ranx:
ranx_check(runs, queries, args.k)
if args.report:
write_report(table, queries, args)
def cmd_chunking(args):
queries = load_queries()
table = {}
for strategy in ("whole", "sections", "fixed"):
searcher = Searcher(make_chunks(strategy, args.size), args.offline)
runs = run_all(searcher, queries, [args.method])
table[strategy] = score_table(runs, queries, args.k)[args.method]
print(f" {strategy:9} {len(searcher.chunks):3} passages")
print()
print_table(table, f"Method '{args.method}', {len(queries)} questions, by chunking choice:")
def paired_bootstrap(diffs, rounds=2000, seed=7):
"""95% interval for the mean difference, by resampling questions with replacement."""
rng = random.Random(seed)
n = len(diffs)
means = sorted(mean(rng.choice(diffs) for _ in range(n)) for _ in range(rounds))
return means[int(0.025 * rounds)], means[int(0.975 * rounds) - 1]
def cmd_compare(args):
queries = load_queries()
searcher = Searcher(make_chunks(args.chunking), args.offline)
runs = run_all(searcher, queries, [args.a, args.b])
metric = lambda r, q: ndcg_at_k(r, q["relevant"], args.k)
diffs = []
print(f"nDCG@{args.k} per question: {args.a} vs {args.b}\n")
for q in queries:
a, b = metric(runs[args.a][q["id"]], q), metric(runs[args.b][q["id"]], q)
diffs.append(b - a)
flag = "better" if b > a + 1e-9 else "worse" if b < a - 1e-9 else ""
print(f" {q['id']} {a:5.2f} {b:5.2f} {flag:6} {q['text'][:48]}")
wins = sum(d > 1e-9 for d in diffs)
losses = sum(d < -1e-9 for d in diffs)
low, high = paired_bootstrap(diffs)
print(f"\n{args.b} vs {args.a}: better on {wins}, worse on {losses}, tied on {len(diffs) - wins - losses}")
print(f"mean difference {mean(diffs):+.3f}, 95% bootstrap interval [{low:+.3f}, {high:+.3f}]")
if low > 0:
print(f"The whole interval is above zero: {args.b} is better on questions like these.")
elif high < 0:
print(f"The whole interval is below zero: {args.b} is worse on questions like these.")
else:
print("The interval includes zero: with this many questions, the difference could be luck.")
if args.ranx:
try:
from ranx import Qrels, Run, compare
except ImportError:
sys.exit("ranx is not installed; run without --ranx.")
import warnings
warnings.filterwarnings("ignore", message="unsafe cast")
qrels = Qrels({q["id"]: q["relevant"] for q in queries})
rx = []
for name in (args.a, args.b):
rx.append(Run({qid: {d: float(len(r) - i) for i, d in enumerate(r)}
for qid, r in runs[name].items()}, name=name))
print("\nranx compare (Fisher's randomization test, p < 0.05):")
print(compare(qrels, rx, [f"ndcg@{args.k}", "mrr"], stat_test="fisher", max_p=0.05))
def cmd_pool(args):
"""Pooling: collect the top-k notes from every method and list the ones nobody labeled."""
queries = load_queries()
searcher = Searcher(make_chunks(args.chunking), args.offline)
runs = run_all(searcher, queries, METHODS)
holes = 0
for q in queries:
pooled = []
for m in METHODS:
for d in runs[m][q["id"]][:args.k]:
if d not in pooled:
pooled.append(d)
unlabeled = [d for d in pooled if d not in q["relevant"]]
holes += len(unlabeled)
if unlabeled:
print(f"{q['id']} {q['text']}")
for d in unlabeled:
print(f" {d} {DOCS[d][0]}")
print(f"\n{holes} unlabeled (question, note) pairs in the top {args.k} of any method.")
print("Read each one. If it helps answer the question, add it to 'relevant' with grade 1 or 2;")
print("if not, leave it out (unlisted means 0).")
def write_report(table, queries, args):
lines = [f"# Retrieval evaluation report ({date.today().isoformat()})", "",
f"- Questions: {len(queries)} (labeled, graded 2 = answers, 1 = helps)",
f"- Library: {len(DOCS)} notes, chunking `{args.chunking}`",
f"- Models: {'toy offline' if args.offline else EMBED_MODEL + ' + ' + RERANK_MODEL}",
f"- Cut-off k = {args.k}", "",
"| Method | " + " | ".join(next(iter(table.values()))) + " |",
"| --- |" + " --- |" * len(next(iter(table.values())))]
for name, row in table.items():
lines.append(f"| {name} | " + " | ".join(f"{v:.3f}" for v in row.values()) + " |")
REPORT_FILE.write_text("\n".join(lines) + "\n", encoding="utf-8")
print(f"\nSaved the report to unit08/{REPORT_FILE.name}")
def main():
parser = argparse.ArgumentParser(description="Evaluate retrieval on labeled SAP-style questions.")
common = argparse.ArgumentParser(add_help=False)
common.add_argument("--offline", action="store_true", help="toy embeddings and re-ranker, no download")
common.add_argument("--k", type=int, default=3, help="cut-off for the @k metrics (default 3)")
sub = parser.add_subparsers(dest="command", required=True)
p = sub.add_parser("init", parents=[common], help="write the labeled questions to unit08/retrieval_queries.jsonl")
p.add_argument("--force", action="store_true")
p = sub.add_parser("eval", parents=[common], help="score every search method")
p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
p.add_argument("--ranx", action="store_true", help="cross-check with the ranx library")
p.add_argument("--report", action="store_true", help="save unit08/retrieval_report.md")
p = sub.add_parser("chunking", parents=[common], help="score three chunking choices with one method")
p.add_argument("--method", default="hybrid", choices=METHODS)
p.add_argument("--size", type=int, default=12, help="words per fixed-size chunk (default 12)")
p = sub.add_parser("compare", parents=[common], help="compare two methods question by question")
p.add_argument("a", choices=METHODS)
p.add_argument("b", choices=METHODS)
p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
p.add_argument("--ranx", action="store_true", help="also run ranx's significance test")
p = sub.add_parser("pool", parents=[common], help="list top results that no label covers yet")
p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
args = parser.parse_args()
if args.k < 1:
sys.exit("--k must be 1 or more")
{"init": cmd_init, "eval": cmd_eval, "chunking": cmd_chunking,
"compare": cmd_compare, "pool": cmd_pool}[args.command](args)
if __name__ == "__main__":
main()
What each part of the script does:
Part
What it does
DOCS
18 made-up help notes across four SAP processes, each a title plus paragraphs
BUILT_IN_QUERIES
20 labeled questions: text, kind, process and graded relevant notes
BM25
Keyword ranking; exact tokens such as PRC-112 and 4500017311 match strongly
CONCEPTS, toy_embed
Offline stand-in for embeddings: words with the same meaning share a "concept"
make_chunks
Cuts notes three ways: whole, sections (one paragraph plus title) and fixed (12-word windows)
best_per_doc
Maps chunks back to their note and keeps each note's best score, so labels stay at note level
Searcher
The four methods: bm25_rank, vector_rank, hybrid_rank (RRF) and rerank (stage 2 on 8 candidates)
hit_at_k … ndcg_at_k
The five metrics in plain Python, one question at a time
load_queries
Reads unit08/retrieval_queries.jsonl if it exists and checks every label names a real note
ranx_check
Turns each ranked list into scores and asks ranx for the same metrics
paired_bootstrap, cmd_compare
Per-question differences, win/loss counts and a 95% bootstrap interval
cmd_pool
Lists notes in any method's top k that have no label yet
write_report
Saves the score table as Markdown for your portfolio
Wrote 20 labeled questions to unit08/retrieval_queries.jsonl
by kind: identifier 4, paraphrase 10, mixed 6
notes in the library: 18
Open unit08/retrieval_queries.jsonl. Each line is one question:
{"id": "q06", "kind": "paraphrase", "process": "procure-to-pay", "text": "supplier sent fewer pieces than we were billed for", "relevant": {"D07": 2, "D06": 1}}
Note D07 (the case note on purchase order 4500017311) answers it: grade 2. D06 (the three-way match guide) helps: grade 1. From now on the script reads this file, so your edits count.
Running init again stops with already exists. That protects your edits. Add --force only if you want the original questions back.
Keyword search is strong on identifiers and mixed questions, and weak on paraphrases (0.36).
Vector search is the mirror image: 0.00 on identifiers. The toy embeddings don't know codes like RM-4711, just as real embeddings handle them poorly.
Hybrid fixes identifiers but scores lower than vector search on paraphrases (0.58 against 0.94). On those questions, keyword search matches words like "purchase", "goods" or "supplier" in the wrong notes. With RRF, a note that appears in both lists, even lower down, can overtake the note vector search put first. For q05, vector search ranks the credit note D01 first, but hybrid drops it out of the top 3. Fusion is a trade, not a free win.
Re-ranking scores best on every kind. Be suspicious: see Step 6.
Hybrid's hit rate rises to 1.000 and its recall to about 0.91, while precision falls to 0.27. Bigger k finds more and adds more noise. Choose k to match how many passages your prompt really receives.
If you have sentence-transformers and a network that allows Hugging Face downloads, run without --offline:
python unit08/retrieval_eval.py eval
The first run downloads all-MiniLM-L6-v2 (the Unit 3 embedding model) and ms-marco-MiniLM-L6-v2 (the Unit 7 re-ranker). Your numbers will differ from the toy ones. That is the point: you now measure real models on your labels.
Fixed 12-word windows lose about 0.09 nDCG. A blind cut separates a note's title from its body and splits sentences, the problem Chunking and document preparation showed by eye.
Whole notes and sections tie. These notes are only two paragraphs long, so splitting them gains nothing. With long manuals, the result would likely differ. Measure on your own documents instead of copying a rule.
Try the same comparison with keyword search, and with a different window size:
nDCG@3 per question: hybrid vs rerank
q01 0.76 0.76 PRC-112
q02 1.00 1.00 4500017311
q03 0.76 0.76 RM-4711 status
q04 1.00 1.00 GR/IR open items
q05 0.00 1.00 better client purchase frozen because they owe us money
...
q20 1.00 1.00 invoice for delivered order not created, price m
rerank vs hybrid: better on 8, worse on 0, tied on 12
mean difference +0.200, 95% bootstrap interval [+0.068, +0.349]
The whole interval is above zero: rerank is better on questions like these.
ranx compare (Fisher's randomization test, p < 0.05):
# Model NDCG@3 MRR
--- ------- -------- ------
a hybrid 0.731 0.804
b rerank 0.931ᵃ 1.000ᵃ
The superscript ᵃ means "significantly better than model a". Both tests agree.
hybrid vs bm25: better on 5, worse on 1, tied on 14
mean difference +0.092, 95% bootstrap interval [-0.001, +0.198]
The interval includes zero: with this many questions, the difference could be luck.
Hybrid's average is higher, but 20 questions can't rule out luck. More questions, especially paraphrases, would settle it.
One more lesson, and an honest one. The toy re-ranker (rerank with --offline) mixes concept overlap with word overlap, and we wrote it while looking at these 20 questions. A perfect MRR of 1.000 on the set you tuned on is a classic sign of overfitting the test set. The Exercise checks it on questions it has never seen.
What success looks like (first lines and last lines):
q04 GR/IR open items
D01 Releasing a sales order blocked by the credit check
D05 Incomplete sales orders
q05 client purchase frozen because they owe us money
D06 Three-way match for supplier invoices
D08 Price variance on supplier invoices
...
51 unlabeled (question, note) pairs in the top 3 of any method.
Read each one. If it helps answer the question, add it to 'relevant' with grade 1 or 2;
if not, leave it out (unlisted means 0).
Read a few. For q04 ("GR/IR open items"), D01 and D05 came up only because they contain the word "open": not relevant, leave them out. For q09 ("what should a planner look at first thing each day"), D14 (reschedule in and out) explains a kind of exception the planner reviews. You might judge it a 1.
If you decide a note helps, open unit08/retrieval_queries.jsonl, find the question's line, and add the note to relevant, for example "relevant": {"D12": 2, "D14": 1}. Save.
Run Step 4 again. Scores can go up or down: a newly labeled note raises the bar for recall and nDCG, and rewards methods that found it.
On SAP BTP, the search step of a RAG system usually runs in one of two places. Both can be evaluated with the same labeled questions and the same metric code.
As of the SAP AI Core service guide of 4 September 2026:
Pipelines vectorize documents from a source into the vector database. The Python SDK documentation shows S3 sources and polling the pipeline status until it completes.
The vector API manages collections and documents; each document holds chunks and metadata.
The retrieval API searches data repositories and returns the relevant chunks for a query, with metadata filtering and merging across repositories.
In orchestration, grounding is configured with a DocumentGroundingFilter that names the data repositories, and a search configuration with max_chunk_count. The retrieved context appears in the response under the grounding module result.
The generative AI hub requires the extended plan of SAP AI Core.
SAP's JavaScript SDK documents the retrieval search call like this (a sketch, quoted in shape from the SDK page):
// Sketch: needs SAP AI Core (extended plan) with a vectorized data repository.
const response: RetrievalSearchResults = await RetrievalApi.search(
{
query: 'supplier sent fewer pieces than we were billed for',
filters: [{
id: 'eval',
searchConfiguration: { maxChunkCount: 10 },
dataRepositories: ['*'],
dataRepositoryType: 'vector'
}]
},
{ 'AI-Resource-Group': 'default' }
).execute();
To evaluate it, you need one thing the API won't give you for free: a stable document ID on every chunk. When you create documents through the vector API, put your own ID (such as D07, or the SAP object key) in the chunk metadata. Then, for each labeled question:
Call the retrieval search with maxChunkCount at least as large as the k you measure.
Read the chunks in the order returned and map each to its document ID from the metadata.
Keep each document's first position, and score the list with the functions from this lab.
Two practical points. First, evaluate with the same filters and resource group the production assistant uses, or you measure a different search. Second, max_chunk_count in orchestration and the k in your metrics should match: if the prompt gets 3 chunks, recall@10 is the wrong headline.
If you built search directly on SAP HANA Cloud (SAP HANA Cloud vector engine), you control the SQL and get the ranked IDs directly. The same scoring applies. HANA Cloud can also use an approximate (HNSW) index, which trades a little recall for speed; Vector databases explained showed how to measure that against exact search. That index recall is a different question from this topic's: it asks "did the index find the true nearest vectors?", while retrieval evaluation asks "did the user get the right document?".
SAP AI Core Evaluations compares prompt templates and models as orchestration configurations, with system-defined metrics such as ROUGE, BLEU and COMET, tool-calling metrics, and custom LLM-as-a-judge metrics. It scores what the configuration outputs. In the documentation we could open, we found no system-defined metric for ranked retrieval such as recall@k or nDCG. A custom judge metric could rate whether retrieved context is relevant, but that is a judge's opinion, not a comparison with expert labels. Use Evaluations for answers (LLM evaluation fundamentals) and your own labeled set for retrieval.
ranx: the metric and comparison library used in this lab (MIT licence).
Ragas: its context precision rewards ranking relevant chunks above irrelevant ones, and context recall measures how much relevant information was retrieved. Each has LLM-based variants and non-LLM variants, including ones that compare document IDs. Every variant needs a reference: an answer, reference passages or reference IDs. The ID-based variants are close to this lab's precision and recall.
Retrieval API of document grounding, or SQL on HANA Cloud
Score retrieval against expert labels
Plain Python or ranx, full control of metrics and k
Not found as a built-in metric; call the API and score yourself
Score generated answers
Inspect, your own judges
SAP AI Core Evaluations with system-defined or custom judge metrics
Significance testing
Bootstrap or ranx compare
Not found in the documentation we opened
Labeling workflow
A JSONL file in Git, reviewed like code
No SAP labeling tool identified for this purpose
Cost
Free on a laptop
Extended plan of SAP AI Core for the generative AI hub
The sensible split for most SAP projects: production retrieval on SAP's managed services, retrieval evaluation in your own small harness that calls them. The next topics build that harness out.
Authorizations change the right answer. A question asked by a user who may only see company code 1000 has different relevant documents than the same question from a user who sees all. If production search applies access filters (Grounding on SAP data without breaking authorizations), store a test user with each question and run the evaluation as that user. Otherwise you measure a search no real user ever sees.
The test set is sensitive data. Real questions and ticket texts can contain customer names, prices and personal data. Store the set in a restricted repository, mask what you can, and treat it like the documents it points to.
Labels go stale. When a policy note is replaced by a new version, the old note's label becomes wrong. Give every label an owner, record the document version you judged, and re-pool after large document changes.
Keep a held-out set. Tune on one part of the questions and report on another that nobody looked at while tuning. Otherwise you get the toy re-ranker's perfect score.
Run it on every index change. Re-chunking, a new embedding model, a new source or a changed filter can all lower recall. Put the evaluation in the same pipeline as your other tests (Testing AI applications) and fail the build on a drop beyond an agreed margin.
Cost is small but real. Each test question is one search call per method, plus re-ranker calls. On SAP AI Core, those calls are billed like production calls, so a large set run on every commit adds up. Run the full set nightly and a small smoke set per change.
Clean core. Evaluation reads search results; it never needs to change SAP data or configuration. Keep it that way: test documents come from exports or the grounding repository, not from writes into S/4HANA.
Labeling only what one method returned. The others look worse for finding unlabeled relevant notes. Pool across methods.
Model-written questions. They copy the document's words and favour keyword search. Prefer real questions.
Precision@k with very different numbers of relevant documents. Its ceiling differs per question; read it alongside recall and nDCG.
Measuring at the wrong k. If the model gets 3 passages, recall@20 hides failures.
Chunk-level labels. They break every time chunking changes. Label documents or stable sections.
Comparing averages from different runs. Different models, filters or question files make numbers incomparable. Compare within one run.
Trusting a small gap. Look at wins and losses per question and a bootstrap interval before switching.
Forgetting questions with no answer. Questions that no document answers can't be scored by these metrics, because there is nothing to find. They test whether the assistant says "I don't know", covered in the next topic of Unit 8.
Test the methods on questions they have never seen, and save a report for the evaluation harness later in this unit.
Before you start: complete the Build it yourself steps above, so that unit08/retrieval_eval.py and unit08/retrieval_queries.jsonl exist.
Open unit08/retrieval_queries.jsonl in VS Code.
Read the titles of the 18 notes in the DOCS section of retrieval_eval.py.
Write six new questions, as a user would ask them, without looking at the note texts: two identifier questions, two paraphrases and two mixed. Use IDs q21 to q26.
For each, add a line at the end of the file in the same shape as the others, with your best labels, for example:
{"id": "q21", "kind": "paraphrase", "process": "procure-to-pay", "text": "the bill came before the delivery", "relevant": {"D11": 2}}
Save, then run the pool and add any labels you missed:
Open unit08/retrieval_report.md and add three sentences at the end: which method you would choose and at which k, whether the re-ranker's lead held up on your new questions, and which kind of question is weakest.
Done when your questions file has 26 questions that all load without errors, unit08/retrieval_report.md shows the score table for all four methods with your three sentences, and you can say whether the toy re-ranker still scores a perfect MRR on the larger set.
Pick one answer for each question. The explanation appears after you choose.
1Why does the lab label documents rather than chunks?
Answer: B. Chunking choices are exactly what you want to compare, so labels must not depend on them. The script maps each chunk to its note and keeps the note's best position. For long manuals, stable section IDs are the middle ground.
2A question has labels {D07: 2, D06: 1}. The search returns D08, D07, D06. What is the reciprocal rank?
Answer: C. Reciprocal rank is 1 divided by the position of the first relevant result. D07 is second, so it is 1/2. The 0.67 is precision@3, a different metric.
3In the lab, hybrid search scored lower than vector search on paraphrased questions. What explains it?
Answer: D. RRF adds credit from both lists, so a note that appears in both, even lower down, can overtake a note found first by vector search alone. Fusion helps on identifiers and costs something on paraphrases, which the split table shows and the average hides.
4The compare command reports "better on 5, worse on 1" with a 95% bootstrap interval of [-0.001, +0.198]. What do you conclude?
Answer: B. The interval includes zero, so the data is consistent with no real difference. More labeled questions, especially of the kind where the methods differ, would narrow it.
5The toy re-ranker scores a perfect MRR of 1.000. What is the right reaction?
Answer: C. A method tuned on the test questions learns those questions, not the task. A held-out set that nobody used for tuning shows whether the lead is real, which is the point of the Exercise.
6You add a new search method, and it scores lower than expected. The pool command lists several of its top results as unlabeled. What should you do?
Answer: B. Unlabeled notes count as irrelevant, so a method that finds relevant notes the old methods missed is punished. Pooling and judging the new results makes the comparison fair.
7You evaluate SAP document grounding through the retrieval API. What do you need to score its results against your labels?
Answer: D. Your labels name documents, and the retrieval API returns chunks. Putting your own ID in the chunk metadata when you create documents lets you map the returned order back to labeled documents and use the same metric code.
8Production search filters results by the user's company codes. How should the evaluation run?
Answer: C. The relevant documents depend on what the user is allowed to see. Running each question as a realistic test user measures the search people actually get, and checks that the filter doesn't hide what they need.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Metrics (ranx documentation)— hit_rate, precision, recall, mrr, ndcg and ndcg_burges with @k cut-offs; hit rate is the share of queries with at least one relevant result; ndcg uses rel / log2(i+1), ndcg_burges uses 2^rel - 1
Context Precision (Ragas documentation)— rewards ranking relevant chunks above irrelevant ones; LLM-based variants and non-LLM variants, including one that compares document IDs
Context Recall (Ragas documentation)— share of relevant documents or claims retrieved; every variant needs a reference (answer, contexts or IDs); ID-based variant needs no LLM
SAP AI Core service guide (PDF, 4 September 2026)— retrieval and vector APIs for document grounding; the retrieval API returns the relevant chunks for a query; metadata filtering; Evaluations with system-defined and custom LLM-as-a-judge metrics; generative AI hub in the extended plan
Document Grounding (SAP Cloud SDK for AI, Python)— pipelines vectorize documents; DocumentGroundingFilter with data repositories; GroundingFilterSearch(max_chunk_count=3); retrieved context read from module_results.grounding
Document Grounding (SAP Cloud SDK for AI, JavaScript)— RetrievalApi.search with a query and filters, searchConfiguration maxChunkCount, dataRepositories, dataRepositoryType vector; Vector API creates collections and documents with chunks and metadata