Orchestrate

Evaluating retrieval

Score a RAG search with recall, precision, MRR and nDCG on a labeled set of SAP-style questions, and compare chunking, hybrid search and re-ranking with numbers.

Updated Oct 5, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

A RAG assistant answers in two steps. First a search finds a few passages. Then a model writes an answer from them. If the search brings back the wrong passages, the best model in the world writes a confident wrong answer.

So you test the search on its own. You write down realistic questions. For each one, an expert marks which documents actually answer it. Then you run the search and count:

  • Did the right document show up at all in the first few results? (hit rate, recall)
  • How much of what came back was useful? (precision)
  • How high up was the first right answer? (MRR)
  • Were the best documents at the very top? (nDCG)

Those numbers turn "the search feels better" into "the search found the right document for 17 of 20 questions, up from 15". With them, a team can pick a chunking method, a search method or a re-ranker on evidence instead of on a demo.

Why it matters to the business

Most bad RAG answers start as bad search results. Take the running procure-to-pay example: an accounts payable clerk asks the assistant why an invoice for purchase order 4500017311 is blocked. If the search returns the general three-way match guide but misses the case note about the 10 missing pieces, the answer is vague. If it returns the price variance guide instead, the answer is wrong.

Measuring retrieval separately pays off in three ways:

  • Faster diagnosis. When an answer is wrong, you can tell whether the search or the model failed. They have different fixes and different owners.
  • Cheaper choices. Unit 7 offered many options: chunk sizes, keyword or vector search, hybrid fusion, re-ranking. Each costs build time, compute or licences. A labeled question set tells you which ones earn their cost on your documents.
  • Safe change. Re-indexing documents, switching embedding models or adding a new document source can quietly make search worse. Rerunning the same questions after each change catches it, the same regression idea as LLM evaluation fundamentals.

The main cost is expert time to label questions. A first set of a few dozen questions, labeled by someone who knows the process, is a modest piece of work. It is reused for every later decision.

How SAP does it

As of the SAP AI Core service guide dated 4 September 2026, the document grounding part of the generative AI hub covers the retrieval side of RAG:

  • Pipelines take documents from a source and vectorize them. SAP's Python SDK documentation shows sources such as S3, and a vector API for feeding chunks directly.
  • A vector API manages collections and documents in the vector database. Each document carries chunks and metadata.
  • A retrieval API "searches data repositories" and returns the relevant chunks for a query. The guide also mentions metadata filtering and merging results from several repositories.
  • In the orchestration service, grounding has a setting for how many chunks to pass to the model (max_chunk_count in the Python SDK). That number is the k this topic measures.
  • The generative AI hub is part of SAP AI Core's extended service plan.

SAP AI Core also has Evaluations, which compares prompt and model configurations with system-defined and custom LLM-as-a-judge metrics. In the SAP documentation we could open, we found no built-in metric that scores the retrieval step itself with recall or nDCG. So plan to measure retrieval yourself: send your labeled questions to the retrieval API and score the returned chunks with the methods in this topic. The same method works for SAP HANA Cloud's vector engine (covered in Unit 7) or any other search.

Reading a retrieval scorecard

A typical result, from the lab in this topic: 20 labeled questions about made-up SAP help notes, four search methods, looking at the top 3 results.

Method Right note in top 3 (hit rate) Share of top 3 that is useful (precision) First right note, on average (MRR) Best notes at the top (nDCG)
Keyword search 75% 27% 0.74 0.64
Vector search 80% 33% 0.74 0.70
Hybrid (both, fused) 85% 35% 0.80 0.73
Hybrid + re-ranker 100% 42% 1.00 0.93

How a leader should read it:

  • Hit rate and recall are about safety. If the right document isn't retrieved, the model can't use it. For RAG this is usually the first number to watch.
  • Precision is about noise and cost. Every useless passage costs tokens and can distract the model.
  • MRR and nDCG are about order. They matter when only the top one or two passages reach the model, or when a person reads the list.
  • Averages hide patterns. In the lab, keyword search was strong on exact codes like PRC-112 and weak on paraphrases; vector search was the opposite. Ask for results split by type of question.
  • A small set gives rough numbers. With 20 questions, a gap of a few points can be luck. Ask whether a difference was tested question by question (the deep layer shows how).
  • A perfect score is a warning. The lab's toy re-ranker was written while looking at these very questions, so its 100% is too good. New, unseen questions are the real test.

Questions to ask

  • Who wrote the test questions, and do they come from real users, tickets or search logs, not from the people who built the search?
  • Who decided which documents count as relevant, and was anyone checking a sample of those labels?
  • Which metric are we optimizing, and at what k? Does k match how many passages the model actually receives?
  • How do results split by process, language and type of question (codes, paraphrases, policy questions)?
  • Was the difference between the options tested for luck, or is it one average against another?
  • Are the test questions kept separate from the ones the team used to tune the search?
  • What happens to these numbers when documents are re-indexed, the embedding model changes or a new source is added?

Common misconceptions

  • "If the answers look good, the search is fine." A strong model can paper over weak search on easy questions and fail on hard ones. Measure the search step on its own.
  • "Vector search is always better than keyword search." Embeddings struggle with exact codes, IDs and document numbers. In the lab, vector search found none of the identifier questions.
  • "Hybrid search always wins." Fusing a weak list with a strong one can drag the strong one down. In the lab, hybrid scored lower than vector search alone on paraphrased questions.
  • "One number is enough." Hit rate, precision and nDCG answer different questions. A change can raise one and lower another.
  • "More test questions can wait." Twenty questions are a good start, but they can't separate options that are close. Grow the set from real failures.
  • "Unlabeled means irrelevant." If a document was never shown to the labeler, it counts as wrong even when it is right. Label everything any method returns near the top.

Key terms

  • Retrieval: the search step of RAG that picks passages for the model.
  • Labeled query set (qrels): test questions with the documents an expert marked as relevant, sometimes with grades such as 2 = answers it, 1 = helps.
  • Top k: the first k results; k is often the number of passages sent to the model.
  • Hit rate@k: share of questions with at least one relevant document in the top k.
  • Recall@k: share of all relevant documents that appear in the top k.
  • Precision@k: share of the top k results that are relevant.
  • MRR (mean reciprocal rank): average of 1 divided by the position of the first relevant result.
  • nDCG@k: a score from 0 to 1 that rewards putting the most relevant documents first.
  • Pooling: collecting the top results from several methods so a labeler judges all of them.
  • Statistical significance: evidence that a difference between two options is unlikely to be luck.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why should a team measure the search step of a RAG assistant separately from the answers?

    Answer: C. RAG answers depend on the passages the search returns. Measuring search on its own tells you whether to fix the search or the model, which are different jobs with different owners.
  2. 2A team wants to know if the right document reaches the model at all. Which number fits best?

    Answer: A. Hit rate and recall ask whether the relevant documents are in the top k at all. If they aren't, the model can't use them. Precision and MRR matter too, but they answer questions about noise and order.
  3. 3Which test set gives the most trustworthy retrieval numbers?

    Answer: C. Numbers only predict real use if the questions look like real use and the labels come from someone who knows the process. Questions used for tuning flatter the system, as the lab's perfect toy re-ranker shows.
  4. 4In the lab, keyword search scored well on codes like PRC-112 and badly on paraphrases. What is the leadership lesson?

    Answer: D. An average can hide a method that fails a whole group of questions. Splitting by question type shows which method suits which need, and why hybrid search exists.
  5. 5Option B beats option A by two points on 20 test questions. What should you ask before switching?

    Answer: B. With a small set, a small gap can come from a few lucky questions. A paired test, such as the bootstrap interval or the randomization test in the deep layer, shows whether the gap is likely to be real.
  6. 6What does SAP's document grounding offer for evaluating retrieval, as described in this topic?

    Answer: C. The SAP AI Core guide describes a retrieval API that returns relevant chunks for a query. We found no built-in retrieval metric in the documentation we could open, so teams score the returned chunks themselves.
  7. 7A colleague says a document that was never labeled must be irrelevant. Why is that risky?

    Answer: D. Metrics treat anything not marked relevant as irrelevant. Pooling, where a labeler judges the top results from every method, closes those gaps so no method is punished for finding something new.
Deep layer · 40 min read

Mental model

Retrieval evaluation is a fixed exam with an answer key. The exam is a list of questions. The answer key, called qrels (query relevance judgments), says which documents are relevant to each question, and how much. A search method sits the exam and hands in a run: for each question, a ranked list of documents. A metric compares the run with the key.

Everything else is detail:

  • The answer key is written once, by people who know the process, and reused for every method and every change.
  • The metric decides what you care about: finding everything (recall), avoiding noise (precision), the first right answer (MRR), or the whole order (nDCG).
  • A comparison between two methods is only trustworthy if you look question by question, because the same questions were asked of both.

This is the information retrieval field's classic "test collection" method, applied to the search step of a RAG system. You met its building blocks in Set up for Unit 8, where ranx computed three metrics on a toy run.

How it works

flowchart LR
  Q[Labeled questions<br/>qrels] --> S[Search method]
  D[(Notes or chunks)] --> S
  S --> R[Run: ranked<br/>note IDs]
  R --> M[Metrics<br/>per question]
  Q --> M
  M --> A[Averages and<br/>splits by type]
  M --> C[Paired comparison<br/>A vs B]

The labeled query set

Each test question has an ID, the text, and graded labels. This lab uses two grades: 2 means the note answers the question, 1 means it helps. Notes not listed count as 0. Two more fields make the set useful for diagnosis: the process (order-to-cash, procure-to-pay, plan-to-produce, record-to-report) and the kind of question:

  • identifier: an exact code or number, such as PRC-112 or 4500017311
  • paraphrase: the meaning is there, the words are not ("client purchase frozen because they owe us money")
  • mixed: some shared words and some meaning

Where do good questions come from? In order of preference: real user questions from search logs or a pilot; ticket subjects and incident descriptions; questions process experts write from memory; and last, questions a model generates from the documents. Model-written questions tend to reuse the document's own words, which flatters keyword search.

Document level or chunk level

You label documents, not chunks. Chunks change every time you change the chunking, and you don't want to relabel. When the search returns chunks, map each chunk back to its source document and keep the document's best position. That is what lets this lab compare three chunking choices on one answer key. The cost: a document-level label can't tell whether the right paragraph of a long document was returned. If your documents are long manuals, label at section level instead and keep section IDs stable.

The metrics, precisely

For one question, let the ranked results be r1, r2, … and let rel(d) be the grade of document d (0 if unlabeled).

Metric Formula in words Range Uses grades? Sensitive to order?
Hit rate@k 1 if any relevant document is in the top k, else 0 0 or 1 No No
Precision@k relevant documents in the top k, divided by k 0–1 No No
Recall@k relevant documents in the top k, divided by all relevant documents 0–1 No No
Reciprocal rank 1 divided by the position of the first relevant document 0–1 No Yes
nDCG@k sum of rel / log2(position + 1) over the top k, divided by the same sum for the ideal order 0–1 Yes Yes

Each is computed per question and then averaged. The average reciprocal rank is MRR.

Some properties worth knowing:

  • Precision@k has a ceiling. If a question has one relevant note and k = 3, the best possible precision@3 is 0.33. Manning, Raghavan and Schütze point out that precision at k averages badly because the number of relevant documents varies so much between questions.
  • Recall@k needs the full answer key. If labels are missing, recall looks better or worse than it is.
  • nDCG is the field's usual headline metric. The BEIR benchmark picked nDCG@10 because it handles both yes/no and graded labels. Its authors note that precision and recall ignore order, and MRR cannot use grades.
  • There are two nDCG gain formulas. ranx's ndcg uses the grade as the gain; its ndcg_burges uses 2^grade − 1, which rewards highly relevant documents more. Say which you use when you report numbers.

A worked example. A question has labels {D07: 2, D06: 1}. The search returns D08, D07, D06.

  • Hit rate@3 = 1 (D07 is there). Recall@3 = 2/2 = 1.0. Precision@3 = 2/3 = 0.67.
  • Reciprocal rank = 1/2 (first relevant note is second).
  • DCG@3 = 0/log2(2) + 2/log2(3) + 1/log2(4) = 0 + 1.26 + 0.5 = 1.76. Ideal order D07, D06: 2/1 + 1/1.58 = 2.63. nDCG@3 = 1.76 / 2.63 = 0.67.

Pooling: closing gaps in the answer key

Labeling every note for every question doesn't scale. The standard shortcut is pooling: run every method you want to compare, take the union of their top k results for each question, and label only those. Anything outside every pool is assumed irrelevant. The lab's pool command lists exactly the (question, note) pairs that some method put in its top k but nobody has labeled yet. When you add a new method later, pool again: otherwise it gets punished for finding relevant notes the old methods missed.

Is the difference real?

Two methods answer the same questions, so compare them pairwise: for each question, the score of B minus the score of A. Then:

  • Count wins, losses and ties. A method that wins 8 and loses 0 is different from one that wins 10 and loses 9.
  • Bootstrap: resample the questions with replacement 2,000 times and look at the spread of the mean difference. If the middle 95% of that spread is entirely above zero, B is very likely better on questions like these.
  • Randomization test: ranx's compare function offers Fisher's randomization test (stat_test="fisher") and a paired Student's t-test (the default, "student"). Its report marks a significant win with a superscript letter.

None of this fixes a biased question set. Statistics tell you the result is unlikely to be luck on these questions, not that the questions represent your users.

Build it yourself: score four search methods and three chunking choices

Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder with its .venv, and install ranx. For the real embedding model and re-ranker you also need sentence-transformers from Set up for Unit 3; without it, use --offline on every command.

You will build retrieval_eval.py. It holds 18 made-up SAP help notes and 20 labeled questions. It searches the notes four ways, the methods from Unit 7: keyword (BM25), vector, hybrid with reciprocal rank fusion, and hybrid plus a re-ranker. It scores each with five metrics, compares chunking choices, tests whether a difference is luck, and lists results that still need labels. It carries its own small copy of the notes, so it runs even if you skipped a Unit 7 lab.

flowchart LR
  I[init<br/>questions file] --> E[eval<br/>4 methods]
  E --> C[chunking<br/>3 choices]
  C --> P[compare<br/>A vs B]
  P --> L[pool<br/>labels to add]
  L --> E

What you need

  • Your course folder with the Unit 8 setup. Cost: free.
  • Optional: sentence-transformers from Unit 3. The first run without --offline downloads two small models from Hugging Face.
  • No accounts and no API keys.
  • Time: about 60 minutes.

Step 1: Open your course folder

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that ranx is installed:

    pip show ranx

    You should see Name: ranx and a Version: line. If you see WARNING: Package(s) not found, repeat Step 2 of Set up for Unit 8. The script also works without ranx; only the --ranx options need it.

Run every command in this topic from the course folder.

Step 2: Create the script

  1. In VS Code, right-click the unit08 folder, choose New File, name it retrieval_eval.py, paste the code below and save.
"""Unit 8: evaluate retrieval with a labeled query set, and compare search choices with numbers.

A small library of made-up SAP help notes is searched four ways (keyword BM25, vector, hybrid
fusion, hybrid plus a re-ranker) and chunked three ways. Every choice is scored on the same
labeled questions with hit rate, precision, recall, MRR and nDCG, and two choices can be
compared question by question to see whether a difference is real or luck.

How to run (from your course folder, with .venv turned on):
    python unit08/retrieval_eval.py init                          # write the labeled questions to a file
    python unit08/retrieval_eval.py eval                          # score every search method
    python unit08/retrieval_eval.py chunking                      # score three chunking choices
    python unit08/retrieval_eval.py compare hybrid rerank         # is one method really better?
    python unit08/retrieval_eval.py pool                          # results nobody has labeled yet
    python unit08/retrieval_eval.py eval --report                 # also save unit08/retrieval_report.md
Add --offline to any command to use toy "concept" embeddings and a toy re-ranker (no download).
Add --ranx to eval or compare to check the numbers with the ranx library.
"""
import argparse
import json
import math
import random
import re
import sys
from collections import Counter
from datetime import date
from pathlib import Path

HERE = Path(__file__).resolve().parent
QUERIES_FILE = HERE / "retrieval_queries.jsonl"
REPORT_FILE = HERE / "retrieval_report.md"
EMBED_MODEL = "sentence-transformers/all-MiniLM-L6-v2"   # the Unit 3 model
RERANK_MODEL = "cross-encoder/ms-marco-MiniLM-L6-v2"     # the Unit 7 re-ranker
METHODS = ["bm25", "vector", "hybrid", "rerank"]

# ---------------------------------------------------------------------------
# 1. The library: made-up SAP help notes, each with a title and a few paragraphs.
# ---------------------------------------------------------------------------
DOCS = {
    "D01": ("Releasing a sales order blocked by the credit check", "order-to-cash", [
        "A sales order is blocked for delivery when the customer's credit exposure is above the "
        "credit limit. Exposure includes open orders, deliveries and unpaid invoices.",
        "The credit analyst reviews the open items. If the overrun is below 2 percent the analyst "
        "releases the order; above that the credit manager approves the release."]),
    "D02": ("Delivery blocks set by sales", "order-to-cash", [
        "Sales can put a delivery block on an order, for example while export papers are missing "
        "or because the customer asked to wait.",
        "Once the reason is cleared, remove the delivery block in the order. The delivery is "
        "created in the next delivery run."]),
    "D03": ("Billing blocks and price checks", "order-to-cash", [
        "A billing block stops the invoice for a delivered order, usually because a price or "
        "discount has to be checked first.",
        "After the pricing review, release the billing block so the order goes to billing."]),
    "D04": ("Pricing error PRC-112", "order-to-cash", [
        "Error PRC-112 means a mandatory price condition is missing on the order item.",
        "Maintain the condition record for the customer and material, then redetermine prices "
        "on the order."]),
    "D05": ("Incomplete sales orders", "order-to-cash", [
        "An order is incomplete when mandatory data is missing, such as the payment terms or the "
        "ship-to party. Incomplete orders cannot be delivered or billed.",
        "Open the incompletion log on the order to see which fields are missing and fill them."]),
    "D06": ("Three-way match for supplier invoices", "procure-to-pay", [
        "Each supplier invoice is checked against the purchase order and the goods receipt.",
        "A quantity or price variance above the tolerance blocks the invoice for payment until "
        "someone resolves it."]),
    "D07": ("Case: quantity variance on purchase order 4500017311", "procure-to-pay", [
        "The supplier delivered 90 of 100 pieces but invoiced all 100.",
        "The invoice was blocked for payment. Accounts payable asked the supplier for a credit "
        "memo for the 10 missing pieces."]),
    "D08": ("Price variance on supplier invoices", "procure-to-pay", [
        "The vendor billed a higher price than the one agreed on the purchase order.",
        "If the price variance is above tolerance, the buyer contacts the vendor and asks for a "
        "corrected invoice or a credit memo."]),
    "D09": ("Releasing invoices blocked for payment", "procure-to-pay", [
        "When the variance is resolved, the accounts payable clerk releases the blocked invoice.",
        "Released invoices are picked up by the next payment run if they are due."]),
    "D10": ("Payment terms and the payment run", "procure-to-pay", [
        "The due date of an invoice is calculated from the baseline date and the payment terms of "
        "the supplier.",
        "The payment run pays every invoice that is due and not blocked."]),
    "D11": ("Invoice arrives before the goods receipt", "procure-to-pay", [
        "With goods-receipt-based invoice verification, an invoice cannot be posted for payment "
        "until the goods receipt is posted.",
        "Ask the warehouse to post the goods receipt, then the invoice can be matched."]),
    "D12": ("Daily review of MRP exceptions", "plan-to-produce", [
        "The planning run flags exception messages: missing parts, late receipts and orders to "
        "reschedule.",
        "Planners review the exceptions every morning, starting with critical components."]),
    "D13": ("Case: shortage of component RM-4711", "plan-to-produce", [
        "The planning run showed a shortage of component RM-4711 for production in week 38.",
        "The planner moved the production order forward and expedited the supplier."]),
    "D14": ("Reschedule in and reschedule out", "plan-to-produce", [
        "Reschedule in means a receipt should arrive earlier than planned to cover demand.",
        "Reschedule out means a receipt arrives too early and can be pushed later."]),
    "D15": ("Safety stock in planning", "plan-to-produce", [
        "Safety stock is a buffer quantity that planning keeps for uncertain demand.",
        "When available stock falls below safety stock, the planning run creates a planned order."]),
    "D16": ("Duplicate check when creating a supplier", "procure-to-pay", [
        "Before creating a new supplier in master data, check for duplicates by tax number and "
        "bank account.",
        "Duplicate suppliers cause double payments and wrong spend reports."]),
    "D17": ("Approving changes to customer bank details", "order-to-cash", [
        "A change to a customer's bank details needs a second person to approve it.",
        "This control reduces the risk of fraud in refunds and credit memos."]),
    "D18": ("Month-end reconciliation of the GR/IR account", "record-to-report", [
        "The GR/IR clearing account holds goods receipts that are not yet invoiced and invoices "
        "not yet received.",
        "At month end, accountants analyze open GR/IR items and clear or adjust them."]),
}

# ---------------------------------------------------------------------------
# 2. The labeled questions. Grades: 2 = answers it, 1 = helps. Unlisted notes count as 0.
#    "kind" says what the question tests: an exact identifier, a paraphrase, or a mix.
# ---------------------------------------------------------------------------
BUILT_IN_QUERIES = [
    {"id": "q01", "kind": "identifier", "process": "order-to-cash", "text": "PRC-112",
     "relevant": {"D04": 2, "D03": 1}},
    {"id": "q02", "kind": "identifier", "process": "procure-to-pay", "text": "4500017311",
     "relevant": {"D07": 2}},
    {"id": "q03", "kind": "identifier", "process": "plan-to-produce", "text": "RM-4711 status",
     "relevant": {"D13": 2, "D12": 1}},
    {"id": "q04", "kind": "identifier", "process": "record-to-report", "text": "GR/IR open items",
     "relevant": {"D18": 2}},
    {"id": "q05", "kind": "paraphrase", "process": "order-to-cash",
     "text": "client purchase frozen because they owe us money",
     "relevant": {"D01": 2}},
    {"id": "q06", "kind": "paraphrase", "process": "procure-to-pay",
     "text": "supplier sent fewer pieces than we were billed for",
     "relevant": {"D07": 2, "D06": 1}},
    {"id": "q07", "kind": "paraphrase", "process": "procure-to-pay",
     "text": "vendor charged more than we agreed",
     "relevant": {"D08": 2, "D06": 1}},
    {"id": "q08", "kind": "paraphrase", "process": "plan-to-produce",
     "text": "lacking parts to build the product",
     "relevant": {"D13": 2, "D12": 1}},
    {"id": "q09", "kind": "paraphrase", "process": "plan-to-produce",
     "text": "what should a planner look at first thing each day",
     "relevant": {"D12": 2}},
    {"id": "q10", "kind": "paraphrase", "process": "order-to-cash",
     "text": "why can we not send the goods to the client",
     "relevant": {"D02": 2, "D01": 2, "D05": 1}},
    {"id": "q11", "kind": "paraphrase", "process": "procure-to-pay",
     "text": "when does the supplier get their money",
     "relevant": {"D10": 2, "D09": 1}},
    {"id": "q12", "kind": "paraphrase", "process": "procure-to-pay",
     "text": "stop paying the same vendor twice",
     "relevant": {"D16": 2}},
    {"id": "q13", "kind": "paraphrase", "process": "order-to-cash",
     "text": "who signs off when a client changes where refunds go",
     "relevant": {"D17": 2}},
    {"id": "q14", "kind": "paraphrase", "process": "plan-to-produce",
     "text": "buffer quantity for uncertain demand",
     "relevant": {"D15": 2}},
    {"id": "q15", "kind": "mixed", "process": "order-to-cash",
     "text": "who may release a credit block above 2 percent",
     "relevant": {"D01": 2}},
    {"id": "q16", "kind": "mixed", "process": "procure-to-pay",
     "text": "release invoice blocked for payment after variance",
     "relevant": {"D09": 2, "D06": 1, "D07": 1}},
    {"id": "q17", "kind": "mixed", "process": "procure-to-pay",
     "text": "invoice cannot be posted, no goods receipt yet",
     "relevant": {"D11": 2, "D06": 1}},
    {"id": "q18", "kind": "mixed", "process": "plan-to-produce",
     "text": "meaning of reschedule out",
     "relevant": {"D14": 2}},
    {"id": "q19", "kind": "mixed", "process": "order-to-cash",
     "text": "order missing payment terms cannot be delivered",
     "relevant": {"D05": 2}},
    {"id": "q20", "kind": "mixed", "process": "order-to-cash",
     "text": "invoice for delivered order not created, price must be checked",
     "relevant": {"D03": 2}},
]

# ---------------------------------------------------------------------------
# 3. Text helpers, keyword search (BM25) and toy "concept" embeddings
# ---------------------------------------------------------------------------
STOPWORDS = set("a an and are as at be because by can each for from has have if in is it its "
                "of on or our so that the their them then they this to until us was we what "
                "when where which who why with not no yet more than".split())
TOKEN = re.compile(r"[a-z0-9]+(?:[-/][a-z0-9]+)*")


def tokens(text: str) -> list:
    return [t for t in TOKEN.findall(text.lower()) if t not in STOPWORDS]


class BM25:
    """Classic BM25 keyword ranking (k1 = 1.5, b = 0.75) over a list of passages."""

    def __init__(self, passages: list, k1: float = 1.5, b: float = 0.75):
        self.docs = [Counter(tokens(p)) for p in passages]
        self.lengths = [sum(d.values()) for d in self.docs]
        self.avg = sum(self.lengths) / len(self.lengths)
        df = Counter(t for d in self.docs for t in d)
        n = len(self.docs)
        self.idf = {t: math.log(1 + (n - f + 0.5) / (f + 0.5)) for t, f in df.items()}
        self.k1, self.b = k1, b

    def scores(self, query: str) -> list:
        out = []
        for d, length in zip(self.docs, self.lengths):
            s = 0.0
            for t in tokens(query):
                if t in d:
                    tf = d[t]
                    s += self.idf[t] * tf * (self.k1 + 1) / (
                        tf + self.k1 * (1 - self.b + self.b * length / self.avg))
            out.append(s)
        return out


# Toy embeddings: each word maps to one or more "concepts"; a text becomes a vector of concept
# counts. Different words with the same meaning share a concept, which is what real
# embeddings learn from data. Used only with --offline.
CONCEPTS = {
    "money_owed": "credit exposure limit owe owes unpaid overrun debt",
    "blocked": "blocked block blocks frozen stuck held hold stop stops stopped cannot",
    "customer": "customer customer's client sales order purchase",
    "supplier": "supplier vendor suppliers seller",
    "invoice": "invoice invoices invoiced billed bill billing charged",
    "price": "price prices pricing charged discount condition",
    "quantity": "quantity pieces fewer delivered missing",
    "payment": "payment pay paid paying money payments refunds",
    "approve": "approve approves approval release releases releasing signs sign-off",
    "parts": "parts component components shortage lacking material",
    "plan": "planning planner planned mrp exceptions reschedule",
    "daily": "daily morning day every first",
    "build": "build production produce product assembly",
    "ship": "delivery deliver delivered send goods ship ship-to",
    "duplicate": "duplicate duplicates double twice same",
    "bank": "bank refunds fraud",
    "buffer": "buffer safety uncertain",
    "receipt": "receipt receipts warehouse",
    "early_late": "earlier later early late forward arrive",
    "month_end": "month-end month end reconciliation clearing accountants",
}
WORD_TO_CONCEPTS = {}
for _concept, _words in CONCEPTS.items():
    for _w in _words.split():
        WORD_TO_CONCEPTS.setdefault(_w, []).append(_concept)
CONCEPT_INDEX = {c: i for i, c in enumerate(CONCEPTS)}


def toy_embed(texts: list) -> list:
    vectors = []
    for text in texts:
        v = [0.0] * len(CONCEPTS)
        for w in TOKEN.findall(text.lower()):
            for c in WORD_TO_CONCEPTS.get(w, []):
                v[CONCEPT_INDEX[c]] += 1.0
        norm = math.sqrt(sum(x * x for x in v)) or 1.0
        vectors.append([x / norm for x in v])
    return vectors


def real_embed(texts: list) -> list:
    try:
        from sentence_transformers import SentenceTransformer
    except ImportError:
        sys.exit("sentence-transformers is not installed (Set up for Unit 3). Add --offline to use toy embeddings.")
    try:
        model = SentenceTransformer(EMBED_MODEL)
    except Exception as e:  # usually no network or a blocked download
        sys.exit(f"Could not load {EMBED_MODEL} ({type(e).__name__}). Check your network, or add --offline.")
    return [list(map(float, v)) for v in model.encode(texts, normalize_embeddings=True)]


def cosine(a: list, b: list) -> float:
    return sum(x * y for x, y in zip(a, b))


# ---------------------------------------------------------------------------
# 4. Chunking: turn the notes into passages. Each passage remembers its note ID.
# ---------------------------------------------------------------------------
def make_chunks(strategy: str, size: int = 12) -> list:
    """Return (doc_id, text) pairs.
    whole:    one passage per note (title + all paragraphs)
    sections: one passage per paragraph, with the note title in front
    fixed:    windows of `size` words cut across the note, no title (a blind cut)"""
    chunks = []
    for doc_id, (title, _process, paras) in DOCS.items():
        if strategy == "whole":
            chunks.append((doc_id, title + ". " + " ".join(paras)))
        elif strategy == "sections":
            chunks.extend((doc_id, f"{title}. {p}") for p in paras)
        elif strategy == "fixed":
            words = " ".join(paras).split()
            chunks.extend((doc_id, " ".join(words[i:i + size])) for i in range(0, len(words), size))
        else:
            raise ValueError(strategy)
    return chunks


# ---------------------------------------------------------------------------
# 5. The search pipelines. Each returns a ranked list of note IDs, best first.
# ---------------------------------------------------------------------------
def best_per_doc(chunks: list, scores: list) -> dict:
    """Several passages can come from one note: keep the note's best passage score."""
    best = {}
    for (doc_id, _text), s in zip(chunks, scores):
        if doc_id not in best or s > best[doc_id]:
            best[doc_id] = s
    return best


def rank(scores: dict) -> list:
    return sorted(scores, key=lambda d: (-scores[d], d))


class Searcher:
    def __init__(self, chunks: list, offline: bool):
        self.chunks = chunks
        self.offline = offline
        texts = [t for _d, t in chunks]
        self.bm25 = BM25(texts)
        self.embed = toy_embed if offline else real_embed
        self.vectors = self.embed(texts)
        self.cross = None

    def bm25_rank(self, q: str) -> list:
        scores = best_per_doc(self.chunks, self.bm25.scores(q))
        return [d for d in rank(scores) if scores[d] > 0]     # keyword search can return nothing

    def vector_rank(self, q: str) -> list:
        qv = self.embed([q])[0]
        scores = best_per_doc(self.chunks, [cosine(qv, v) for v in self.vectors])
        if self.offline:   # a toy vector with no known concept matches nothing: return nothing
            return [d for d in rank(scores) if scores[d] > 0]
        return rank(scores)

    def hybrid_rank(self, q: str, rrf_k: int = 60) -> list:
        """Reciprocal rank fusion: each list adds 1 / (rrf_k + position) to a note's score."""
        fused = Counter()
        for ranked in (self.bm25_rank(q), self.vector_rank(q)):
            for pos, d in enumerate(ranked, 1):
                fused[d] += 1 / (rrf_k + pos)
        return rank(dict(fused))

    def rerank(self, q: str, candidates: int = 8) -> list:
        """Stage 2: re-score the top hybrid candidates by reading query and passage together."""
        shortlist = self.hybrid_rank(q)[:candidates]
        passages = {d: " ".join([DOCS[d][0]] + DOCS[d][2]) for d in shortlist}
        if self.offline:
            qv = toy_embed([q])[0]
            qt = set(tokens(q))
            scored = {d: cosine(qv, toy_embed([p])[0]) + 0.3 * len(qt & set(tokens(p))) / max(len(qt), 1)
                      for d, p in passages.items()}
        else:
            if self.cross is None:
                try:
                    from sentence_transformers import CrossEncoder
                except ImportError:
                    sys.exit("sentence-transformers is not installed. Add --offline to use the toy re-ranker.")
                try:
                    self.cross = CrossEncoder(RERANK_MODEL)
                except Exception as e:  # usually no network or a blocked download
                    sys.exit(f"Could not load {RERANK_MODEL} ({type(e).__name__}). Check your network, or add --offline.")
            pairs = [(q, passages[d]) for d in shortlist]
            scored = dict(zip(shortlist, map(float, self.cross.predict(pairs))))
        reordered = rank(scored)
        return reordered + [d for d in self.hybrid_rank(q) if d not in reordered]

    def run(self, method: str, q: str) -> list:
        return {"bm25": self.bm25_rank, "vector": self.vector_rank,
                "hybrid": self.hybrid_rank, "rerank": self.rerank}[method](q)


# ---------------------------------------------------------------------------
# 6. The metrics, in plain Python. `ranked` = note IDs best first; `rel` = {note: grade}.
# ---------------------------------------------------------------------------
def hit_at_k(ranked, rel, k):
    return 1.0 if any(d in rel for d in ranked[:k]) else 0.0


def precision_at_k(ranked, rel, k):
    return sum(1 for d in ranked[:k] if d in rel) / k


def recall_at_k(ranked, rel, k):
    return sum(1 for d in ranked[:k] if d in rel) / len(rel)


def reciprocal_rank(ranked, rel):
    for pos, d in enumerate(ranked, 1):
        if d in rel:
            return 1 / pos
    return 0.0


def ndcg_at_k(ranked, rel, k):
    dcg = sum(rel.get(d, 0) / math.log2(i + 2) for i, d in enumerate(ranked[:k]))
    ideal = sorted(rel.values(), reverse=True)[:k]
    idcg = sum(g / math.log2(i + 2) for i, g in enumerate(ideal))
    return dcg / idcg if idcg else 0.0


def per_query(ranked, rel, k):
    return {f"hit@{k}": hit_at_k(ranked, rel, k), f"P@{k}": precision_at_k(ranked, rel, k),
            f"R@{k}": recall_at_k(ranked, rel, k), "MRR": reciprocal_rank(ranked, rel),
            f"nDCG@{k}": ndcg_at_k(ranked, rel, k)}


def mean(values):
    values = list(values)
    return sum(values) / len(values) if values else 0.0


# ---------------------------------------------------------------------------
# 7. Loading the labeled questions
# ---------------------------------------------------------------------------
def load_queries() -> list:
    if not QUERIES_FILE.exists():
        return BUILT_IN_QUERIES
    queries = []
    for n, line in enumerate(QUERIES_FILE.read_text(encoding="utf-8").splitlines(), 1):
        if not line.strip():
            continue
        try:
            q = json.loads(line)
        except json.JSONDecodeError as e:
            sys.exit(f"{QUERIES_FILE.name} line {n} is not valid JSON: {e}")
        unknown = [d for d in q.get("relevant", {}) if d not in DOCS]
        if not q.get("relevant") or unknown:
            sys.exit(f"{QUERIES_FILE.name} line {n} ({q.get('id')}): needs 'relevant' with known note IDs; unknown: {unknown}")
        queries.append(q)
    return queries


def run_all(searcher, queries, methods):
    return {m: {q["id"]: searcher.run(m, q["text"]) for q in queries} for m in methods}


def score_table(runs, queries, k):
    table = {}
    for method, results in runs.items():
        rows = [per_query(results[q["id"]], q["relevant"], k) for q in queries]
        table[method] = {metric: mean(r[metric] for r in rows) for metric in rows[0]}
    return table


def print_table(table, title):
    metrics = list(next(iter(table.values())))
    print(title)
    print(f"{'':10}" + "".join(f"{m:>9}" for m in metrics))
    for name, row in table.items():
        print(f"{name:10}" + "".join(f"{row[m]:9.3f}" for m in metrics))


def ranx_check(runs, queries, k):
    try:
        from ranx import Qrels, Run, evaluate
    except ImportError:
        print("\nranx is not installed; skipping the cross-check (see Set up for Unit 8).")
        return
    import warnings
    warnings.filterwarnings("ignore", message="unsafe cast")
    qrels = Qrels({q["id"]: q["relevant"] for q in queries})
    print(f"\nranx cross-check (hit_rate@{k}, precision@{k}, recall@{k}, mrr, ndcg@{k}):")
    for method, results in runs.items():
        # ranx wants scores, not positions: give the first result the highest score.
        run = Run({qid: {d: float(len(r) - i) for i, d in enumerate(r)} for qid, r in results.items()})
        scores = evaluate(qrels, run, [f"hit_rate@{k}", f"precision@{k}", f"recall@{k}", "mrr", f"ndcg@{k}"])
        print(f"{method:10}" + "".join(f"{v:9.3f}" for v in scores.values()))


# ---------------------------------------------------------------------------
# 8. Commands
# ---------------------------------------------------------------------------
def cmd_init(args):
    if QUERIES_FILE.exists() and not args.force:
        sys.exit(f"{QUERIES_FILE.name} already exists. Add --force to overwrite it with the built-in questions.")
    with QUERIES_FILE.open("w", encoding="utf-8") as f:
        for q in BUILT_IN_QUERIES:
            f.write(json.dumps(q) + "\n")
    kinds = Counter(q["kind"] for q in BUILT_IN_QUERIES)
    print(f"Wrote {len(BUILT_IN_QUERIES)} labeled questions to unit08/{QUERIES_FILE.name}")
    print("  by kind: " + ", ".join(f"{k} {n}" for k, n in kinds.items()))
    print(f"  notes in the library: {len(DOCS)}")


def cmd_eval(args):
    queries = load_queries()
    searcher = Searcher(make_chunks(args.chunking), args.offline)
    runs = run_all(searcher, queries, METHODS)
    table = score_table(runs, queries, args.k)
    print(f"{len(queries)} questions, {len(DOCS)} notes, chunking '{args.chunking}', "
          f"{'toy offline models' if args.offline else 'real models'}\n")
    print_table(table, "Average over all questions:")
    kinds = sorted({q.get("kind", "other") for q in queries})
    print(f"\nnDCG@{args.k} by kind of question:")
    print(f"{'':10}" + "".join(f"{k:>12}" for k in kinds))
    for m in METHODS:
        cells = []
        for kind in kinds:
            subset = [q for q in queries if q.get("kind", "other") == kind]
            cells.append(mean(ndcg_at_k(runs[m][q["id"]], q["relevant"], args.k) for q in subset))
        print(f"{m:10}" + "".join(f"{c:12.3f}" for c in cells))
    if args.ranx:
        ranx_check(runs, queries, args.k)
    if args.report:
        write_report(table, queries, args)


def cmd_chunking(args):
    queries = load_queries()
    table = {}
    for strategy in ("whole", "sections", "fixed"):
        searcher = Searcher(make_chunks(strategy, args.size), args.offline)
        runs = run_all(searcher, queries, [args.method])
        table[strategy] = score_table(runs, queries, args.k)[args.method]
        print(f"  {strategy:9} {len(searcher.chunks):3} passages")
    print()
    print_table(table, f"Method '{args.method}', {len(queries)} questions, by chunking choice:")


def paired_bootstrap(diffs, rounds=2000, seed=7):
    """95% interval for the mean difference, by resampling questions with replacement."""
    rng = random.Random(seed)
    n = len(diffs)
    means = sorted(mean(rng.choice(diffs) for _ in range(n)) for _ in range(rounds))
    return means[int(0.025 * rounds)], means[int(0.975 * rounds) - 1]


def cmd_compare(args):
    queries = load_queries()
    searcher = Searcher(make_chunks(args.chunking), args.offline)
    runs = run_all(searcher, queries, [args.a, args.b])
    metric = lambda r, q: ndcg_at_k(r, q["relevant"], args.k)
    diffs = []
    print(f"nDCG@{args.k} per question: {args.a} vs {args.b}\n")
    for q in queries:
        a, b = metric(runs[args.a][q["id"]], q), metric(runs[args.b][q["id"]], q)
        diffs.append(b - a)
        flag = "better" if b > a + 1e-9 else "worse" if b < a - 1e-9 else ""
        print(f"  {q['id']}  {a:5.2f}  {b:5.2f}  {flag:6}  {q['text'][:48]}")
    wins = sum(d > 1e-9 for d in diffs)
    losses = sum(d < -1e-9 for d in diffs)
    low, high = paired_bootstrap(diffs)
    print(f"\n{args.b} vs {args.a}: better on {wins}, worse on {losses}, tied on {len(diffs) - wins - losses}")
    print(f"mean difference {mean(diffs):+.3f}, 95% bootstrap interval [{low:+.3f}, {high:+.3f}]")
    if low > 0:
        print(f"The whole interval is above zero: {args.b} is better on questions like these.")
    elif high < 0:
        print(f"The whole interval is below zero: {args.b} is worse on questions like these.")
    else:
        print("The interval includes zero: with this many questions, the difference could be luck.")
    if args.ranx:
        try:
            from ranx import Qrels, Run, compare
        except ImportError:
            sys.exit("ranx is not installed; run without --ranx.")
        import warnings
        warnings.filterwarnings("ignore", message="unsafe cast")
        qrels = Qrels({q["id"]: q["relevant"] for q in queries})
        rx = []
        for name in (args.a, args.b):
            rx.append(Run({qid: {d: float(len(r) - i) for i, d in enumerate(r)}
                           for qid, r in runs[name].items()}, name=name))
        print("\nranx compare (Fisher's randomization test, p < 0.05):")
        print(compare(qrels, rx, [f"ndcg@{args.k}", "mrr"], stat_test="fisher", max_p=0.05))


def cmd_pool(args):
    """Pooling: collect the top-k notes from every method and list the ones nobody labeled."""
    queries = load_queries()
    searcher = Searcher(make_chunks(args.chunking), args.offline)
    runs = run_all(searcher, queries, METHODS)
    holes = 0
    for q in queries:
        pooled = []
        for m in METHODS:
            for d in runs[m][q["id"]][:args.k]:
                if d not in pooled:
                    pooled.append(d)
        unlabeled = [d for d in pooled if d not in q["relevant"]]
        holes += len(unlabeled)
        if unlabeled:
            print(f"{q['id']}  {q['text']}")
            for d in unlabeled:
                print(f"      {d}  {DOCS[d][0]}")
    print(f"\n{holes} unlabeled (question, note) pairs in the top {args.k} of any method.")
    print("Read each one. If it helps answer the question, add it to 'relevant' with grade 1 or 2;")
    print("if not, leave it out (unlisted means 0).")


def write_report(table, queries, args):
    lines = [f"# Retrieval evaluation report ({date.today().isoformat()})", "",
             f"- Questions: {len(queries)} (labeled, graded 2 = answers, 1 = helps)",
             f"- Library: {len(DOCS)} notes, chunking `{args.chunking}`",
             f"- Models: {'toy offline' if args.offline else EMBED_MODEL + ' + ' + RERANK_MODEL}",
             f"- Cut-off k = {args.k}", "",
             "| Method | " + " | ".join(next(iter(table.values()))) + " |",
             "| --- |" + " --- |" * len(next(iter(table.values())))]
    for name, row in table.items():
        lines.append(f"| {name} | " + " | ".join(f"{v:.3f}" for v in row.values()) + " |")
    REPORT_FILE.write_text("\n".join(lines) + "\n", encoding="utf-8")
    print(f"\nSaved the report to unit08/{REPORT_FILE.name}")


def main():
    parser = argparse.ArgumentParser(description="Evaluate retrieval on labeled SAP-style questions.")
    common = argparse.ArgumentParser(add_help=False)
    common.add_argument("--offline", action="store_true", help="toy embeddings and re-ranker, no download")
    common.add_argument("--k", type=int, default=3, help="cut-off for the @k metrics (default 3)")
    sub = parser.add_subparsers(dest="command", required=True)
    p = sub.add_parser("init", parents=[common], help="write the labeled questions to unit08/retrieval_queries.jsonl")
    p.add_argument("--force", action="store_true")
    p = sub.add_parser("eval", parents=[common], help="score every search method")
    p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
    p.add_argument("--ranx", action="store_true", help="cross-check with the ranx library")
    p.add_argument("--report", action="store_true", help="save unit08/retrieval_report.md")
    p = sub.add_parser("chunking", parents=[common], help="score three chunking choices with one method")
    p.add_argument("--method", default="hybrid", choices=METHODS)
    p.add_argument("--size", type=int, default=12, help="words per fixed-size chunk (default 12)")
    p = sub.add_parser("compare", parents=[common], help="compare two methods question by question")
    p.add_argument("a", choices=METHODS)
    p.add_argument("b", choices=METHODS)
    p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
    p.add_argument("--ranx", action="store_true", help="also run ranx's significance test")
    p = sub.add_parser("pool", parents=[common], help="list top results that no label covers yet")
    p.add_argument("--chunking", default="sections", choices=["whole", "sections", "fixed"])
    args = parser.parse_args()
    if args.k < 1:
        sys.exit("--k must be 1 or more")
    {"init": cmd_init, "eval": cmd_eval, "chunking": cmd_chunking,
     "compare": cmd_compare, "pool": cmd_pool}[args.command](args)


if __name__ == "__main__":
    main()

What each part of the script does:

Part What it does
DOCS 18 made-up help notes across four SAP processes, each a title plus paragraphs
BUILT_IN_QUERIES 20 labeled questions: text, kind, process and graded relevant notes
BM25 Keyword ranking; exact tokens such as PRC-112 and 4500017311 match strongly
CONCEPTS, toy_embed Offline stand-in for embeddings: words with the same meaning share a "concept"
make_chunks Cuts notes three ways: whole, sections (one paragraph plus title) and fixed (12-word windows)
best_per_doc Maps chunks back to their note and keeps each note's best score, so labels stay at note level
Searcher The four methods: bm25_rank, vector_rank, hybrid_rank (RRF) and rerank (stage 2 on 8 candidates)
hit_at_k … ndcg_at_k The five metrics in plain Python, one question at a time
load_queries Reads unit08/retrieval_queries.jsonl if it exists and checks every label names a real note
ranx_check Turns each ranked list into scores and asks ranx for the same metrics
paired_bootstrap, cmd_compare Per-question differences, win/loss counts and a 95% bootstrap interval
cmd_pool Lists notes in any method's top k that have no label yet
write_report Saves the score table as Markdown for your portfolio

Step 3: Write the questions to a file

  1. Run:

    python unit08/retrieval_eval.py init

What success looks like:

Wrote 20 labeled questions to unit08/retrieval_queries.jsonl
  by kind: identifier 4, paraphrase 10, mixed 6
  notes in the library: 18
  1. Open unit08/retrieval_queries.jsonl. Each line is one question:

    {"id": "q06", "kind": "paraphrase", "process": "procure-to-pay", "text": "supplier sent fewer pieces than we were billed for", "relevant": {"D07": 2, "D06": 1}}

    Note D07 (the case note on purchase order 4500017311) answers it: grade 2. D06 (the three-way match guide) helps: grade 1. From now on the script reads this file, so your edits count.

  2. Running init again stops with already exists. That protects your edits. Add --force only if you want the original questions back.

Step 4: Score the four search methods

  1. Run with the toy models (no download):

    python unit08/retrieval_eval.py eval --offline

What success looks like (from our test):

20 questions, 18 notes, chunking 'sections', toy offline models

Average over all questions:
              hit@3      P@3      R@3      MRR   nDCG@3
bm25          0.750    0.267    0.592    0.735    0.639
vector        0.800    0.333    0.700    0.742    0.697
hybrid        0.850    0.350    0.742    0.804    0.731
rerank        1.000    0.417    0.875    1.000    0.931

nDCG@3 by kind of question:
            identifier       mixed  paraphrase
bm25             0.880       0.940       0.362
vector           0.000       0.755       0.941
hybrid           0.880       0.878       0.582
rerank           0.880       0.940       0.946
  1. Read the split table, not just the averages:

    • Keyword search is strong on identifiers and mixed questions, and weak on paraphrases (0.36).
    • Vector search is the mirror image: 0.00 on identifiers. The toy embeddings don't know codes like RM-4711, just as real embeddings handle them poorly.
    • Hybrid fixes identifiers but scores lower than vector search on paraphrases (0.58 against 0.94). On those questions, keyword search matches words like "purchase", "goods" or "supplier" in the wrong notes. With RRF, a note that appears in both lists, even lower down, can overtake the note vector search put first. For q05, vector search ranks the credit note D01 first, but hybrid drops it out of the top 3. Fusion is a trade, not a free win.
    • Re-ranking scores best on every kind. Be suspicious: see Step 6.
  2. Try a larger cut-off:

    python unit08/retrieval_eval.py eval --offline --k 5

    Hybrid's hit rate rises to 1.000 and its recall to about 0.91, while precision falls to 0.27. Bigger k finds more and adds more noise. Choose k to match how many passages your prompt really receives.

  3. If you have sentence-transformers and a network that allows Hugging Face downloads, run without --offline:

    python unit08/retrieval_eval.py eval

    The first run downloads all-MiniLM-L6-v2 (the Unit 3 embedding model) and ms-marco-MiniLM-L6-v2 (the Unit 7 re-ranker). Your numbers will differ from the toy ones. That is the point: you now measure real models on your labels.

Step 5: Compare chunking choices

  1. Run:

    python unit08/retrieval_eval.py chunking --offline

What success looks like:

  whole      18 passages
  sections   36 passages
  fixed      57 passages

Method 'hybrid', 20 questions, by chunking choice:
              hit@3      P@3      R@3      MRR   nDCG@3
whole         0.850    0.350    0.742    0.804    0.733
sections      0.850    0.350    0.742    0.804    0.731
fixed         0.850    0.317    0.692    0.733    0.646
  1. Read it:

    • Fixed 12-word windows lose about 0.09 nDCG. A blind cut separates a note's title from its body and splits sentences, the problem Chunking and document preparation showed by eye.
    • Whole notes and sections tie. These notes are only two paragraphs long, so splitting them gains nothing. With long manuals, the result would likely differ. Measure on your own documents instead of copying a rule.
  2. Try the same comparison with keyword search, and with a different window size:

    python unit08/retrieval_eval.py chunking --offline --method bm25
    python unit08/retrieval_eval.py chunking --offline --size 25

Step 6: Test whether a difference is real

  1. Compare hybrid with hybrid plus re-ranking. Add --ranx for the library's own test (the first ranx call takes up to a minute while it compiles):

    python unit08/retrieval_eval.py compare hybrid rerank --offline --ranx

What success looks like (trimmed in the middle):

nDCG@3 per question: hybrid vs rerank

  q01   0.76   0.76          PRC-112
  q02   1.00   1.00          4500017311
  q03   0.76   0.76          RM-4711 status
  q04   1.00   1.00          GR/IR open items
  q05   0.00   1.00  better  client purchase frozen because they owe us money
  ...
  q20   1.00   1.00          invoice for delivered order not created, price m

rerank vs hybrid: better on 8, worse on 0, tied on 12
mean difference +0.200, 95% bootstrap interval [+0.068, +0.349]
The whole interval is above zero: rerank is better on questions like these.

ranx compare (Fisher's randomization test, p < 0.05):
#    Model    NDCG@3    MRR
---  -------  --------  ------
a    hybrid   0.731     0.804
b    rerank   0.931ᵃ    1.000ᵃ

The superscript ᵃ means "significantly better than model a". Both tests agree.

  1. Now compare keyword search with hybrid:

    python unit08/retrieval_eval.py compare bm25 hybrid --offline

What success looks like (last lines):

hybrid vs bm25: better on 5, worse on 1, tied on 14
mean difference +0.092, 95% bootstrap interval [-0.001, +0.198]
The interval includes zero: with this many questions, the difference could be luck.

Hybrid's average is higher, but 20 questions can't rule out luck. More questions, especially paraphrases, would settle it.

  1. One more lesson, and an honest one. The toy re-ranker (rerank with --offline) mixes concept overlap with word overlap, and we wrote it while looking at these 20 questions. A perfect MRR of 1.000 on the set you tuned on is a classic sign of overfitting the test set. The Exercise checks it on questions it has never seen.

Step 7: Find labels that are missing

  1. Run:

    python unit08/retrieval_eval.py pool --offline

What success looks like (first lines and last lines):

q04  GR/IR open items
      D01  Releasing a sales order blocked by the credit check
      D05  Incomplete sales orders
q05  client purchase frozen because they owe us money
      D06  Three-way match for supplier invoices
      D08  Price variance on supplier invoices
...

51 unlabeled (question, note) pairs in the top 3 of any method.
Read each one. If it helps answer the question, add it to 'relevant' with grade 1 or 2;
if not, leave it out (unlisted means 0).
  1. Read a few. For q04 ("GR/IR open items"), D01 and D05 came up only because they contain the word "open": not relevant, leave them out. For q09 ("what should a planner look at first thing each day"), D14 (reschedule in and out) explains a kind of exception the planner reviews. You might judge it a 1.
  2. If you decide a note helps, open unit08/retrieval_queries.jsonl, find the question's line, and add the note to relevant, for example "relevant": {"D12": 2, "D14": 1}. Save.
  3. Run Step 4 again. Scores can go up or down: a newly labeled note raises the bar for recall and nDCG, and rewards methods that found it.

Step 8: Save a report and your work

  1. Save the score table:

    python unit08/retrieval_eval.py eval --offline --report

    The last line says Saved the report to unit08/retrieval_report.md. Open it: a Markdown table you can paste into a design document.

  2. Save in Git:

    git add unit08/retrieval_eval.py unit08/retrieval_queries.jsonl unit08/retrieval_report.md
    git commit -m "Evaluate retrieval on labeled SAP questions"

    The questions file is saved: it is test data, and its history shows how your labels changed.

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
can't open file ... retrieval_eval.py You aren't in the course folder, or the file has another name Run cd to orchestrate-course; check the file is unit08/retrieval_eval.py
sentence-transformers is not installed The Unit 3 library is missing Add --offline, or install it as in Set up for Unit 3
Could not load ... (ProxyError) or another download error Your network blocks Hugging Face downloads Use --offline; try another network, or ask IT to allow Hugging Face
ranx is not installed The --ranx option needs the Unit 8 library pip install -r requirements.txt with (.venv) on, or leave out --ranx
ranx takes 30 seconds or more It compiles code on first use Wait; later runs are faster
line N is not valid JSON An edit in the questions file broke a quote, comma or brace Fix that line; each line must be one complete {...} object
needs 'relevant' with known note IDs A label names a note that doesn't exist, or a question has no labels Use IDs D01 to D18, and give every question at least one relevant note
already exists. Add --force init protects your edited questions Do nothing, or add --force to start over
pip install fails with a proxy or SSL error The company network blocks the package index Ask IT to allow pypi.org, or try on another network
Windows: Activate.ps1 cannot be loaded PowerShell blocks scripts Run Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser, answer Y, and try again

The SAP way

On SAP BTP, the search step of a RAG system usually runs in one of two places. Both can be evaluated with the same labeled questions and the same metric code.

Document grounding in the generative AI hub

As of the SAP AI Core service guide of 4 September 2026:

  • Pipelines vectorize documents from a source into the vector database. The Python SDK documentation shows S3 sources and polling the pipeline status until it completes.
  • The vector API manages collections and documents; each document holds chunks and metadata.
  • The retrieval API searches data repositories and returns the relevant chunks for a query, with metadata filtering and merging across repositories.
  • In orchestration, grounding is configured with a DocumentGroundingFilter that names the data repositories, and a search configuration with max_chunk_count. The retrieved context appears in the response under the grounding module result.
  • The generative AI hub requires the extended plan of SAP AI Core.

SAP's JavaScript SDK documents the retrieval search call like this (a sketch, quoted in shape from the SDK page):

// Sketch: needs SAP AI Core (extended plan) with a vectorized data repository.
const response: RetrievalSearchResults = await RetrievalApi.search(
  {
    query: 'supplier sent fewer pieces than we were billed for',
    filters: [{
      id: 'eval',
      searchConfiguration: { maxChunkCount: 10 },
      dataRepositories: ['*'],
      dataRepositoryType: 'vector'
    }]
  },
  { 'AI-Resource-Group': 'default' }
).execute();

To evaluate it, you need one thing the API won't give you for free: a stable document ID on every chunk. When you create documents through the vector API, put your own ID (such as D07, or the SAP object key) in the chunk metadata. Then, for each labeled question:

  1. Call the retrieval search with maxChunkCount at least as large as the k you measure.
  2. Read the chunks in the order returned and map each to its document ID from the metadata.
  3. Keep each document's first position, and score the list with the functions from this lab.

Two practical points. First, evaluate with the same filters and resource group the production assistant uses, or you measure a different search. Second, max_chunk_count in orchestration and the k in your metrics should match: if the prompt gets 3 chunks, recall@10 is the wrong headline.

SAP HANA Cloud vector engine

If you built search directly on SAP HANA Cloud (SAP HANA Cloud vector engine), you control the SQL and get the ranked IDs directly. The same scoring applies. HANA Cloud can also use an approximate (HNSW) index, which trades a little recall for speed; Vector databases explained showed how to measure that against exact search. That index recall is a different question from this topic's: it asks "did the index find the true nearest vectors?", while retrieval evaluation asks "did the user get the right document?".

SAP AI Core Evaluations

SAP AI Core Evaluations compares prompt templates and models as orchestration configurations, with system-defined metrics such as ROUGE, BLEU and COMET, tool-calling metrics, and custom LLM-as-a-judge metrics. It scores what the configuration outputs. In the documentation we could open, we found no system-defined metric for ranked retrieval such as recall@k or nDCG. A custom judge metric could rate whether retrieved context is relevant, but that is a judge's opinion, not a comparison with expert labels. Use Evaluations for answers (LLM evaluation fundamentals) and your own labeled set for retrieval.

Open tools that do the same job

  • ranx: the metric and comparison library used in this lab (MIT licence).
  • Ragas: its context precision rewards ranking relevant chunks above irrelevant ones, and context recall measures how much relevant information was retrieved. Each has LLM-based variants and non-LLM variants, including ones that compare document IDs. Every variant needs a reference: an answer, reference passages or reference IDs. The ID-based variants are close to this lab's precision and recall.

Build vs. SAP

Need Build it yourself (this lab, ranx) SAP-managed services
Run retrieval for test questions Your own search code, or calls to any API Retrieval API of document grounding, or SQL on HANA Cloud
Score retrieval against expert labels Plain Python or ranx, full control of metrics and k Not found as a built-in metric; call the API and score yourself
Score generated answers Inspect, your own judges SAP AI Core Evaluations with system-defined or custom judge metrics
Significance testing Bootstrap or ranx compare Not found in the documentation we opened
Labeling workflow A JSONL file in Git, reviewed like code No SAP labeling tool identified for this purpose
Cost Free on a laptop Extended plan of SAP AI Core for the generative AI hub

The sensible split for most SAP projects: production retrieval on SAP's managed services, retrieval evaluation in your own small harness that calls them. The next topics build that harness out.

Production concerns

  • Authorizations change the right answer. A question asked by a user who may only see company code 1000 has different relevant documents than the same question from a user who sees all. If production search applies access filters (Grounding on SAP data without breaking authorizations), store a test user with each question and run the evaluation as that user. Otherwise you measure a search no real user ever sees.
  • The test set is sensitive data. Real questions and ticket texts can contain customer names, prices and personal data. Store the set in a restricted repository, mask what you can, and treat it like the documents it points to.
  • Labels go stale. When a policy note is replaced by a new version, the old note's label becomes wrong. Give every label an owner, record the document version you judged, and re-pool after large document changes.
  • Keep a held-out set. Tune on one part of the questions and report on another that nobody looked at while tuning. Otherwise you get the toy re-ranker's perfect score.
  • Run it on every index change. Re-chunking, a new embedding model, a new source or a changed filter can all lower recall. Put the evaluation in the same pipeline as your other tests (Testing AI applications) and fail the build on a drop beyond an agreed margin.
  • Cost is small but real. Each test question is one search call per method, plus re-ranker calls. On SAP AI Core, those calls are billed like production calls, so a large set run on every commit adds up. Run the full set nightly and a small smoke set per change.
  • Clean core. Evaluation reads search results; it never needs to change SAP data or configuration. Keep it that way: test documents come from exports or the grounding repository, not from writes into S/4HANA.

Pitfalls

  • Labeling only what one method returned. The others look worse for finding unlabeled relevant notes. Pool across methods.
  • Model-written questions. They copy the document's words and favour keyword search. Prefer real questions.
  • Precision@k with very different numbers of relevant documents. Its ceiling differs per question; read it alongside recall and nDCG.
  • Measuring at the wrong k. If the model gets 3 passages, recall@20 hides failures.
  • Chunk-level labels. They break every time chunking changes. Label documents or stable sections.
  • Comparing averages from different runs. Different models, filters or question files make numbers incomparable. Compare within one run.
  • Trusting a small gap. Look at wins and losses per question and a bootstrap interval before switching.
  • Forgetting questions with no answer. Questions that no document answers can't be scored by these metrics, because there is nothing to find. They test whether the assistant says "I don't know", covered in the next topic of Unit 8.

Exercise

Test the methods on questions they have never seen, and save a report for the evaluation harness later in this unit.

Before you start: complete the Build it yourself steps above, so that unit08/retrieval_eval.py and unit08/retrieval_queries.jsonl exist.

  1. Open unit08/retrieval_queries.jsonl in VS Code.

  2. Read the titles of the 18 notes in the DOCS section of retrieval_eval.py.

  3. Write six new questions, as a user would ask them, without looking at the note texts: two identifier questions, two paraphrases and two mixed. Use IDs q21 to q26.

  4. For each, add a line at the end of the file in the same shape as the others, with your best labels, for example:

    {"id": "q21", "kind": "paraphrase", "process": "procure-to-pay", "text": "the bill came before the delivery", "relevant": {"D11": 2}}
  5. Save, then run the pool and add any labels you missed:

    python unit08/retrieval_eval.py pool --offline
  6. Score all methods and save the report:

    python unit08/retrieval_eval.py eval --offline --report
  7. Compare the re-ranker with hybrid on the larger set:

    python unit08/retrieval_eval.py compare hybrid rerank --offline
  8. Open unit08/retrieval_report.md and add three sentences at the end: which method you would choose and at which k, whether the re-ranker's lead held up on your new questions, and which kind of question is weakest.

  9. Save your work:

    git add unit08/retrieval_queries.jsonl unit08/retrieval_report.md
    git commit -m "Add unseen retrieval questions and report"

Done when your questions file has 26 questions that all load without errors, unit08/retrieval_report.md shows the score table for all four methods with your three sentences, and you can say whether the toy re-ranker still scores a perfect MRR on the larger set.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the lab label documents rather than chunks?

    Answer: B. Chunking choices are exactly what you want to compare, so labels must not depend on them. The script maps each chunk to its note and keeps the note's best position. For long manuals, stable section IDs are the middle ground.
  2. 2A question has labels {D07: 2, D06: 1}. The search returns D08, D07, D06. What is the reciprocal rank?

    Answer: C. Reciprocal rank is 1 divided by the position of the first relevant result. D07 is second, so it is 1/2. The 0.67 is precision@3, a different metric.
  3. 3In the lab, hybrid search scored lower than vector search on paraphrased questions. What explains it?

    Answer: D. RRF adds credit from both lists, so a note that appears in both, even lower down, can overtake a note found first by vector search alone. Fusion helps on identifiers and costs something on paraphrases, which the split table shows and the average hides.
  4. 4The compare command reports "better on 5, worse on 1" with a 95% bootstrap interval of [-0.001, +0.198]. What do you conclude?

    Answer: B. The interval includes zero, so the data is consistent with no real difference. More labeled questions, especially of the kind where the methods differ, would narrow it.
  5. 5The toy re-ranker scores a perfect MRR of 1.000. What is the right reaction?

    Answer: C. A method tuned on the test questions learns those questions, not the task. A held-out set that nobody used for tuning shows whether the lead is real, which is the point of the Exercise.
  6. 6You add a new search method, and it scores lower than expected. The pool command lists several of its top results as unlabeled. What should you do?

    Answer: B. Unlabeled notes count as irrelevant, so a method that finds relevant notes the old methods missed is punished. Pooling and judging the new results makes the comparison fair.
  7. 7You evaluate SAP document grounding through the retrieval API. What do you need to score its results against your labels?

    Answer: D. Your labels name documents, and the retrieval API returns chunks. Putting your own ID in the chunk metadata when you create documents lets you map the returned order back to labeled documents and use the same metric code.
  8. 8Production search filters results by the user's company codes. How should the evaluation run?

    Answer: C. The relevant documents depend on what the user is allowed to see. Running each question as a realistic test user measures the search people actually get, and checks that the filter doesn't hide what they need.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in