Unit 8 is about evaluation: proving with numbers whether an AI system gives the right answers, and noticing when a change makes it worse. Without it, every demo is an opinion.
This setup adds two free, open-source tools and one file:
Inspect, an evaluation framework from the UK AI Security Institute. It runs a list of test cases through a model, scores each answer and keeps a log you can browse.
ranx, a small library that scores search results with the standard measures of retrieval quality.
A golden set: a short file of test cases with known right answers. You build it from the triage notes you judged at the end of What is an SAP FDE? in Unit 1.
No new accounts. Setup takes 30 to 50 minutes and costs nothing. A real model call later in the unit is a small per-request charge on the model key you already have.
Leaders hear two kinds of claims about AI: "it works great" and "it hallucinates". Neither helps a decision. Evaluation turns both into a number on a fixed set of cases, such as "on 50 past blocked orders, the draft note was right 41 times". That number can be compared before and after a change, and between vendors.
Take the running example. In Unit 1 an AI drafted notes for blocked sales orders, and you judged five of them as right, partly right or wrong. Those judgments are the start of an evaluation set. Each time someone changes the prompt, the model or the data the AI reads, the team reruns the same cases. If the share of right notes drops, the change doesn't ship.
Three points for a leader:
Evaluation starts with people, not tools. Someone who knows the process must judge real cases first. The tools only repeat that judgment at scale.
The first number is often unflattering. That is useful. It is the baseline that later work must beat.
Open-source tools cost nothing to try. The cost is in the expert time to build and maintain good test cases.
As of September 2026, SAP's service guide for SAP AI Core describes Evaluations in the generative AI hub:
It benchmarks models and prompts by running them as orchestration configurations, the same setup you met in Unit 5.
You can choose system-defined metrics. SAP names ROUGE, BLEU and COMET (scores that compare text with a reference answer) and metrics for tool calling.
You can define custom LLM-as-a-judge metrics, where a model grades answers against rating criteria you write.
The generative AI hub is available only in the extended service plan of SAP AI Core.
The SAP Cloud SDK for AI that you installed in Set up for Unit 5 already contains an evaluations module. SAP's reference shows it registering test data as datasets in SAP AI Core and reading files from an object store. That needs storage and access that a personal trial may not have, so this course teaches the ideas with open tools first. Later Unit 8 topics show the SAP path as a sketch.
Time: 30 to 50 minutes. Installing Inspect takes a few minutes because it brings many helper libraries. ranx is slow the first time it runs, then fast.
Money: nothing for the setup. The mock judge in the setup makes no model calls.
Later in the unit: running six cases through a real model is a small per-request charge. A golden set of 50 cases, rerun after each change, adds up; check the price of the model you pick.
May learners install inspect-ai and ranx from the Python package index?
Does the network allow a one-time download from openaipublic.blob.core.windows.net? Inspect may fetch a small tokenizer file from there to count tokens.
May evaluation logs, which contain every prompt and answer, be stored on laptops? Where should they live for a real project?
Which past cases may be used as test data, and who approves that? For the setup, the answer is: made-up data and your own judged notes only.
Does the company have SAP AI Core with the extended plan, and an object store connected to it, for evaluations on SAP's side?
"Evaluation needs a big labelled dataset before we can start." Five judged cases are enough to start. The set grows as the team finds new failure cases.
"A model can grade itself, so we don't need experts." A model judge only helps once its verdicts agree with an expert's on known cases. Measuring that agreement is the first job.
"An evaluation tool is a quality guarantee." The tool counts. The test cases decide what it counts. Weak cases give a confident but meaningless number.
"We need SAP's evaluation service to start." SAP's service is useful in production on SAP AI Core. The habits (golden set, metrics, reruns) start on a laptop.
Pick one answer for each question. The explanation appears after you choose.
1What does the Unit 8 setup turn your judged Unit 1 triage notes into?
Answer: B. The judged notes become test cases that every later change is scored against. Nothing is retrained; the cases measure the system.
2The first evaluation shows only half of the draft notes are right. What is the best reading?
Answer: C. An unflattering first number is still the baseline. Its value is that every later change can be compared with it on the same cases.
3A vendor says their model can grade its own answers, so no expert time is needed. What do you ask?
Answer: B. A model judge is only useful once its verdicts match an expert's on cases with known answers. Agreement is the number that earns trust.
4Which SAP offering covers evaluating models and prompts, as of September 2026?
Answer: C. SAP's service guide describes Evaluations in the generative AI hub, with system-defined and custom LLM-as-a-judge metrics. The generative AI hub is only in the extended service plan of SAP AI Core.
5Why does the course start with open-source tools instead of SAP's evaluation service?
Answer: B. SAP's evaluations module works with datasets in SAP AI Core and files in an object store. The habits are the same, so the course teaches them on a laptop first and shows SAP's path later.
6Which question to IT matters for evaluation logs specifically?
Answer: D. An evaluation log saves every input and output of a run. With real data, that is sensitive, so where logs live is a data protection decision.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 35 min read
#Mental model: a fixed exam, rerun after every change
Evaluation is an exam with fixed questions. The golden set holds the questions and the right answers. A solver (your prompt plus a model) writes the answers. A scorer marks them. You rerun the same exam after every change and compare the marks.
flowchart LR
G[(Golden set<br/>cases + right answers)] --> S[Solver<br/>prompt + model]
S --> A[Answers]
A --> SC[Scorer<br/>rule or judge model]
G --> SC
SC --> M[Metrics<br/>accuracy, recall, MRR]
M --> L[(Log<br/>every case, every score)]
This setup builds each box once, small. Inspect provides the solver, scorer and log. ranx provides retrieval metrics. You provide the golden set.
Inspect is an evaluation framework developed by the UK AI Security Institute and Meridian Labs, published on PyPI under the MIT licence. As of 2 October 2026 the current version is 0.3.276, for Python 3.10 or newer. Its documentation builds every evaluation from three parts:
Part
What it is
In this setup
Dataset
A list of Sample(input=..., target=...)
One sample per judged triage note; the target is your verdict
Solver
Steps that produce an answer, such as system_message(...) then generate()
The judge instructions, then one model call
Scorer
A function that marks the answer
pattern(...), which pulls VERDICT: right out of the reply and compares it with the target
Inspect's scorer list also includes match, includes, exact, f1, choice and model-graded scorers such as model_graded_qa. The tutorial says logs go to ./logs by default and that inspect view opens a viewer in your browser. Models are named by provider, such as anthropic/<model>, which reads ANTHROPIC_API_KEY.
For the no-account path, the setup uses Inspect's mock model, mockllm/model, with a fixed reply. It behaves like a lazy judge that says "partly" to everything. That is useful, not just a stand-in: a constant answer is the floor any real judge must beat.
Retrieval is scored differently from answers. You need two things per question: which documents are relevant (qrels, short for relevance judgments) and what the search returned, in order (run). ranx builds both from Python dictionaries and computes metrics by name, such as recall@3, mrr and ndcg@3. It is MIT-licensed and uses Numba, a compiler for numeric Python, so its first call is slow while it compiles and later calls are fast.
Metric
Plain meaning
Recall@k
Share of the relevant documents found in the top k
MRR
Average of 1 / position of the first relevant result
nDCG@k
Rewards relevant results near the top, and very relevant ones more
The Evaluating retrieval topic explains them in depth. Here you only prove ranx gives the same numbers as a few lines of plain Python.
The golden set is a JSONL file: one JSON object per line, easy to add to and easy to compare in Git. Each line holds one judged note:
{"id": "t2", "order": {"SalesOrder": "9000002", "HeaderBillingBlockReason": "02"}, "draft": {"likely_cause": "...", "next_check": "...", "confidence": "low"}, "verdict": "right", "reason": "Admits what the data can't tell."}
The order fields and draft keys match what triage.py printed in Unit 1. verdict is your judgment: right, partly or wrong. reason is one sentence on why. That sentence matters: it is what turns a verdict into a rule a judge model can follow later.
#Build it yourself: install the Unit 8 tools and run a first evaluation
You will install Inspect and ranx, create a golden set, score a small search result two ways, run an evaluation with a mock judge, and run a check. Every step works without an account.
The script writes six made-up judged notes, shaped like Unit 1's output, and checks the file. You will add your own notes in the Exercise.
Make the Unit 8 folder:
Windows (PowerShell):
New-Item -ItemType Directory -Force unit08
macOS / Linux:
mkdir -p unit08
In VS Code, right-click unit08, choose New File, name it golden_set.py, paste the code below and save.
"""Unit 8: turn judged triage notes into a golden set, and check that the file is well formed.
A golden set is a fixed list of cases with a known right answer. Here each case is one blocked
sales order, the draft note an AI wrote for it in Unit 1, and your verdict on that note.
The file is unit08/data/triage_golden.jsonl: one JSON object per line.
How to run (from your course folder, with .venv turned on):
python unit08/golden_set.py --sample # write six made-up judged notes to start from
python unit08/golden_set.py # check the file and print a summary
python unit08/golden_set.py --sample --force # overwrite the file with the samples again
"""
import argparse
import json
import sys
from collections import Counter
from pathlib import Path
GOLDEN = Path(__file__).resolve().parent / "data" / "triage_golden.jsonl"
VERDICTS = ("right", "partly", "wrong")
REQUIRED = ("id", "order", "draft", "verdict", "reason")
def order(number, customer, amount, delivery_block, billing_block, credit_status):
"""Build the header fields Unit 1's triage.py read. The codes are placeholders, not SAP meanings."""
return {"SalesOrder": number, "SoldToParty": customer, "TotalNetAmount": amount,
"TransactionCurrency": "USD", "TotalCreditCheckStatus": credit_status,
"DeliveryBlockReason": delivery_block, "HeaderBillingBlockReason": billing_block}
def draft(cause, check, confidence):
return {"likely_cause": cause, "next_check": check, "confidence": confidence}
# Six made-up judged notes, shaped like Unit 1's output. Replace or extend them with your own.
SAMPLES = [
{"id": "t1", "order": order("9000001", "CUST-A", "18250.00", "01", "", "B"),
"draft": draft("Both a delivery block and a credit check status are set, so credit is worth "
"checking first.", "Check the customer's credit exposure and open items.", "medium"),
"verdict": "right", "reason": "Uses only the fields given and points to the right next check."},
{"id": "t2", "order": order("9000002", "CUST-B", "940.00", "", "02", ""),
"draft": draft("A billing block is set; the data doesn't say why.",
"Ask billing which block code 02 stands for in this system.", "low"),
"verdict": "right", "reason": "Admits what the data can't tell and asks for the code meaning."},
{"id": "t3", "order": order("9000004", "CUST-D", "72400.00", "01", "", "B"),
"draft": draft("Delivery block 01 means the customer is over the credit limit.",
"Release the order in credit management.", "high"),
"verdict": "wrong", "reason": "Guesses what code 01 means and tells the analyst to release, not to check."},
{"id": "t4", "order": order("9000005", "CUST-E", "3150.00", "03", "02", ""),
"draft": draft("A delivery block is set.", "Check the delivery block with the sales team.", "medium"),
"verdict": "partly", "reason": "Misses the billing block on the same order."},
{"id": "t5", "order": order("9000006", "CUST-A", "12900.00", "01", "", "B"),
"draft": draft("A delivery block with a non-blank credit status; possibly a credit hold.",
"Check credit exposure; CUST-A also has order 9000001 blocked.", "medium"),
"verdict": "partly", "reason": "Good check, but it cites another order it was never given."},
{"id": "t6", "order": order("9000007", "CUST-F", "560.00", "", "", ""),
"draft": draft("No block fields are set, so the data doesn't show why this order is held.",
"Confirm with the analyst why it was flagged.", "low"),
"verdict": "right", "reason": "Says the data is insufficient instead of inventing a cause."},
]
def validate(path: Path) -> list:
"""Read the golden set and return its records, or stop with the line number of the first problem."""
if not path.exists():
sys.exit(f"No golden set yet at {path}. Run: python unit08/golden_set.py --sample")
records, seen = [], set()
for number, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
if not line.strip():
continue # blank lines are allowed
try:
record = json.loads(line)
except json.JSONDecodeError as error:
sys.exit(f"Line {number} is not valid JSON ({error.msg}). Each line must be one complete object.")
missing = [field for field in REQUIRED if field not in record]
if missing:
sys.exit(f"Line {number} is missing: {', '.join(missing)}")
if record["verdict"] not in VERDICTS:
sys.exit(f"Line {number}: verdict must be one of {', '.join(VERDICTS)}, not {record['verdict']!r}")
if record["id"] in seen:
sys.exit(f"Line {number}: the id {record['id']!r} is used twice")
seen.add(record["id"])
records.append(record)
if not records:
sys.exit("The golden set is empty. Run with --sample, or add your own judged notes.")
return records
def main() -> None:
parser = argparse.ArgumentParser(description="Create or check the Unit 8 golden set of judged triage notes.")
parser.add_argument("--sample", action="store_true", help="write six made-up judged notes")
parser.add_argument("--force", action="store_true", help="with --sample, overwrite an existing file")
args = parser.parse_args()
if args.sample:
if GOLDEN.exists() and not args.force:
sys.exit(f"{GOLDEN.name} already exists. Add --force to overwrite it with the samples.")
GOLDEN.parent.mkdir(parents=True, exist_ok=True)
GOLDEN.write_text("".join(json.dumps(r) + "\n" for r in SAMPLES), encoding="utf-8")
print(f"Wrote {len(SAMPLES)} made-up judged notes to unit08/data/{GOLDEN.name}")
records = validate(GOLDEN)
counts = Counter(r["verdict"] for r in records)
print(f"\nGolden set: {len(records)} judged notes")
for verdict in VERDICTS:
print(f" {verdict:7} {counts[verdict]}")
share = counts["right"] / len(records)
print(f"\nShare of drafts you judged right: {share:.0%}")
print("That share is your first evaluation number: the quality of the Unit 1 drafts, judged by a person.")
if __name__ == "__main__":
main()
Write the samples and check them:
python unit08/golden_set.py --sample
What success looks like (from our test):
Wrote 6 made-up judged notes to unit08/data/triage_golden.jsonl
Golden set: 6 judged notes
right 3
partly 2
wrong 1
Share of drafts you judged right: 50%
That share is your first evaluation number: the quality of the Unit 1 drafts, judged by a person.
Run it once more without --sample. It only checks the file. If you edit the file and break a line, it names the line, for example Line 7: verdict must be one of right, partly, wrong, not 'Right'.
Look at the six samples. Note t3: the draft guesses what code 01 means, which Unit 1 told the model not to do. Note t5: the draft mentions an order it was never given. Both kinds of error come back in the hallucination topic.
What each part of the script does:
Part
What it does
GOLDEN
Where the file lives: unit08/data/triage_golden.jsonl
order, draft
Build records with the same fields as Unit 1's triage.py
SAMPLES
Six made-up judged notes: three right, two partly, one wrong
validate
Reads each line, checks the fields and verdict, and stops with the line number of the first problem
--sample, --force
Write the samples; refuse to overwrite your file unless you add --force
main
Prints counts by verdict and the share judged right
#Step 4: Score a search result by hand and with ranx
The script takes three questions about the Unit 7 help notes (n1 to n8), the notes that are relevant to each, and what a search returned. It computes three metrics in plain Python, then asks ranx.
In unit08, create retrieval_metrics_hello.py, paste the code below and save.
"""Unit 8 smoke test: score a tiny search result with three retrieval metrics, by hand and with ranx.
Three questions were asked of the Unit 7 help notes (IDs n1 to n8). For each, we know which notes
are relevant (the "qrels") and which notes the search returned, best first (the "run").
The script computes Recall@3, MRR and nDCG@3 in plain Python, then asks the ranx library for the
same numbers. If the two columns agree, the library is installed and you read it correctly.
How to run (from your course folder, with .venv turned on):
python unit08/retrieval_metrics_hello.py # by hand and with ranx
python unit08/retrieval_metrics_hello.py --no-ranx # by hand only (no library needed)
"""
import argparse
import math
import sys
import warnings
K = 3
# Which notes are relevant for each question. 2 = answers it fully, 1 = helps a bit.
QRELS = {
"q1 why can't this order be delivered": {"n1": 2, "n3": 1},
"q2 invoice quantity higher than goods receipt": {"n4": 2, "n6": 1},
"q3 planned order arrives too late": {"n7": 2},
}
# What a search returned for each question, best first, with its similarity score.
RUN = {
"q1 why can't this order be delivered": {"n3": 0.81, "n2": 0.77, "n1": 0.74, "n6": 0.40},
"q2 invoice quantity higher than goods receipt": {"n4": 0.88, "n5": 0.71, "n6": 0.69, "n1": 0.30},
"q3 planned order arrives too late": {"n8": 0.79, "n2": 0.55, "n5": 0.41, "n7": 0.39},
}
def ranked(results: dict) -> list:
"""Note IDs sorted by score, best first."""
return sorted(results, key=results.get, reverse=True)
def recall_at_k(relevant: dict, results: dict, k: int) -> float:
"""Share of the relevant notes that appear in the top k."""
top = ranked(results)[:k]
return sum(1 for note in relevant if note in top) / len(relevant)
def reciprocal_rank(relevant: dict, results: dict) -> float:
"""1 divided by the position of the first relevant note (0 if none is returned)."""
for position, note in enumerate(ranked(results), 1):
if note in relevant:
return 1 / position
return 0.0
def ndcg_at_k(relevant: dict, results: dict, k: int) -> float:
"""Graded gain, discounted by position, divided by the best possible ordering."""
gains = [relevant.get(note, 0) for note in ranked(results)[:k]]
dcg = sum(g / math.log2(i + 2) for i, g in enumerate(gains))
ideal = sorted(relevant.values(), reverse=True)[:k]
idcg = sum(g / math.log2(i + 2) for i, g in enumerate(ideal))
return dcg / idcg if idcg else 0.0
def by_hand() -> dict:
n = len(QRELS)
return {
f"recall@{K}": sum(recall_at_k(QRELS[q], RUN[q], K) for q in QRELS) / n,
"mrr": sum(reciprocal_rank(QRELS[q], RUN[q]) for q in QRELS) / n,
f"ndcg@{K}": sum(ndcg_at_k(QRELS[q], RUN[q], K) for q in QRELS) / n,
}
def with_ranx() -> dict:
try:
from ranx import Qrels, Run, evaluate
except ImportError:
sys.exit("ranx is not installed. Run: pip install -r requirements.txt (or add --no-ranx)")
warnings.filterwarnings("ignore", message="unsafe cast") # a harmless notice from ranx's compiler
print("Asking ranx (the first call compiles some code, so it can take several seconds)...")
return evaluate(Qrels(QRELS), Run(RUN), [f"recall@{K}", "mrr", f"ndcg@{K}"])
def main() -> None:
parser = argparse.ArgumentParser(description="Compute retrieval metrics by hand and with ranx.")
parser.add_argument("--no-ranx", action="store_true", help="skip the ranx library")
args = parser.parse_args()
print("Per question (by hand):")
for q in QRELS:
print(f" {q[:2]} top {K}: {', '.join(ranked(RUN[q])[:K]):12} "
f"recall {recall_at_k(QRELS[q], RUN[q], K):.2f} RR {reciprocal_rank(QRELS[q], RUN[q]):.2f} "
f"nDCG {ndcg_at_k(QRELS[q], RUN[q], K):.2f}")
mine = by_hand()
theirs = {} if args.no_ranx else with_ranx()
print(f"\n{'metric':10} {'by hand':>8} {'ranx':>8}")
for metric, value in mine.items():
other = f"{theirs[metric]:8.3f}" if metric in theirs else " -"
print(f"{metric:10} {value:8.3f} {other}")
if theirs:
same = all(abs(mine[m] - theirs[m]) < 1e-6 for m in mine)
print("\nThe two columns agree." if same else "\nThe columns differ: check the data you changed.")
if __name__ == "__main__":
main()
Run it:
python unit08/retrieval_metrics_hello.py
What success looks like (from our test; the first run took about 40 seconds, later runs about 5):
Per question (by hand):
q1 top 3: n3, n2, n1 recall 1.00 RR 1.00 nDCG 0.76
q2 top 3: n4, n5, n6 recall 1.00 RR 1.00 nDCG 0.95
q3 top 3: n8, n2, n5 recall 0.00 RR 0.25 nDCG 0.00
Asking ranx (the first call compiles some code, so it can take several seconds)...
metric by hand ranx
recall@3 0.667 0.667
mrr 0.750 0.750
ndcg@3 0.570 0.570
The two columns agree.
Read question q3: the one relevant note, n7, came fourth, outside the top 3. Recall@3 is 0, but MRR still gives a little credit (1/4). That is why teams report more than one metric.
If ranx won't install or run, use --no-ranx. The by-hand column is enough to follow the Evaluating retrieval topic.
What each part of the script does:
Part
What it does
QRELS
The relevant notes per question; 2 = answers it, 1 = helps
RUN
What the search returned, with scores; higher is better
recall_at_k, reciprocal_rank, ndcg_at_k
The three metrics in plain Python
with_ranx
Builds Qrels and Run from the same dictionaries and calls evaluate
warnings.filterwarnings(...)
Hides a harmless notice that ranx's compiler prints
The script turns each golden-set line into an Inspect sample. A judge reads the order and the draft and replies with a verdict. Inspect compares it with yours. Without --model, the judge is a mock that always says "partly" and makes no calls.
In unit08, create inspect_hello.py, paste the code below and save.
"""Unit 8 smoke test: run a first evaluation with Inspect, using your golden set.
The task asks a model to act as a judge: for each judged triage note in
unit08/data/triage_golden.jsonl, it reads the order and the draft note and replies right, partly
or wrong. Inspect compares the model's verdict with yours and reports the share that agree.
How to run (from your course folder, with .venv turned on):
python unit08/inspect_hello.py # no account: a mock judge
python unit08/inspect_hello.py --model anthropic/claude-opus-5-5 # a real model (small charge)
Then look at the results in your browser:
inspect view --log-dir unit08/logs
"""
import argparse
import json
import os
import sys
from pathlib import Path
from dotenv import load_dotenv
HERE = Path(__file__).resolve().parent
GOLDEN = HERE / "data" / "triage_golden.jsonl"
LOGS = HERE / "logs"
JUDGE = (
"You review draft notes that an assistant wrote for credit and order-management analysts. "
"A good note uses only the order fields given, does not guess what block or status codes mean, "
"says when the data is insufficient, and suggests a sensible next check. "
"Judge the draft as right, partly (useful but incomplete or with a small error) or wrong. "
"Think briefly, then end with one line exactly like: VERDICT: right"
)
def load_samples() -> list:
"""Turn each golden-set line into an Inspect Sample: input = order + draft, target = your verdict."""
from inspect_ai.dataset import Sample
if not GOLDEN.exists():
sys.exit("No golden set yet. Run: python unit08/golden_set.py --sample")
samples = []
for line in GOLDEN.read_text(encoding="utf-8").splitlines():
if line.strip():
record = json.loads(line)
text = (f"Order fields:\n{json.dumps(record['order'], indent=2)}\n\n"
f"Draft note:\n{json.dumps(record['draft'], indent=2)}")
samples.append(Sample(id=record["id"], input=text, target=record["verdict"]))
return samples
def lazy_judge():
"""A stand-in model for the no-account path: it answers 'partly' every time, without reading."""
from inspect_ai.model import ModelOutput, ModelUsage, get_model
def always_partly(messages, tools, tool_choice, config):
output = ModelOutput.from_content(model="mockllm", content="I did not read it.\nVERDICT: partly")
output.usage = ModelUsage(input_tokens=0, output_tokens=0, total_tokens=0) # nothing to count offline
return output
return get_model("mockllm/model", custom_outputs=always_partly)
def main() -> None:
parser = argparse.ArgumentParser(description="Run a first Inspect evaluation over the golden set.")
parser.add_argument("--model", help="for example anthropic/claude-opus-5-5; leave out for the mock judge")
args = parser.parse_args()
load_dotenv() # ANTHROPIC_API_KEY comes from .env, as in Unit 1
if args.model and args.model.startswith("anthropic/") and not os.getenv("ANTHROPIC_API_KEY"):
sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or leave out --model.")
try:
from inspect_ai import Task, eval
from inspect_ai.scorer import pattern
from inspect_ai.solver import generate, system_message
except ImportError:
sys.exit("inspect-ai is not installed. Run: pip install -r requirements.txt")
task = Task(
dataset=load_samples(),
solver=[system_message(JUDGE), generate()],
scorer=pattern(r"VERDICT:\s*(right|partly|wrong)"), # pull the verdict out of the reply
name="triage_judge",
)
model = args.model or lazy_judge()
print(f"Judging {len(task.dataset)} notes with {args.model or 'a mock judge that always says partly'}...")
logs = eval(task, model=model, log_dir=str(LOGS), display="plain")
log = logs[0]
if log.status != "success":
sys.exit(f"The evaluation did not finish ({log.status}). Read the error above, or run: inspect view --log-dir unit08/logs")
accuracy = log.results.scores[0].metrics["accuracy"].value
print(f"\nAgreement with your verdicts: {accuracy:.0%} of {len(task.dataset)} notes")
print("Log saved in unit08/logs. Browse it with: inspect view --log-dir unit08/logs")
if __name__ == "__main__":
main()
Run it with the mock judge:
python unit08/inspect_hello.py
What success looks like (trimmed; from our test):
Judging 6 notes with a mock judge that always says partly...
...
triage_judge (6 samples): mockllm/model
...
pattern
accuracy 0.333
stderr 0.211
Log:
unit08/logs/2026-10-04T18-24-25-00-00_triage-judge_....eval
Agreement with your verdicts: 33% of 6 notes
Log saved in unit08/logs. Browse it with: inspect view --log-dir unit08/logs
The mock agrees on the two notes you judged "partly", so 2 of 6. Any real judge must beat 33% on this set. stderr is Inspect's estimate of how much that number could move with different cases; with six cases it is large.
Look at the log in your browser:
inspect view --log-dir unit08/logs
The terminal shows Running on http://127.0.0.1:7575. Open that address. Click the run, then a sample, to see the input, the reply and the score. Press Ctrl+C in the terminal to stop the viewer.
The logs are rebuilt by every run and will hold real prompts later, so keep them out of Git. Open .gitignore, add this line at the end and save:
unit08/logs/
Optional, small charge: with ANTHROPIC_API_KEY in .env from Unit 1, run a real judge:
This sends six short requests. You can use any model your key supports; write it after anthropic/. Note the agreement number. The LLM evaluation fundamentals topic is about making it trustworthy.
What each part of the script does:
Part
What it does
JUDGE
The judge's instructions, which mirror the rules Unit 1 gave the triage model
load_samples
One Sample per golden-set line; input = order + draft, target = your verdict
lazy_judge
Inspect's mockllm/model with a fixed reply, VERDICT: partly, and zero token usage
In the course folder (not in unit08), create check_unit08.py, paste the code below and save. It uses only built-in Python, like the earlier checks.
"""Check that your computer is ready for Unit 8 (evaluation).
Run it from your course folder: python check_unit08.py
It uses built-in Python only. It looks for the Unit 8 libraries and files, checks that your golden
set is well formed, and reads the names (not the values) of the keys in .env. It changes nothing.
"""
import importlib.metadata
import importlib.util
import json
import os
import sys
problems = 0
SAMPLE_IDS = {"t1", "t2", "t3", "t4", "t5", "t6"}
def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
"""Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
global problems
if ok:
print(f" OK {label}")
elif optional:
print(f" LATER {label} -> {fix}")
else:
problems += 1
print(f" MISSING {label} -> {fix}")
def library(module: str, package: str) -> str:
"""Return the installed version of a library, or '' if it isn't installed."""
if importlib.util.find_spec(module) is None:
return ""
try:
return importlib.metadata.version(package)
except importlib.metadata.PackageNotFoundError:
return "installed"
def env_names(path: str = ".env") -> set:
"""Names of the settings in .env that have a value (the values are never printed)."""
names = set()
if os.path.exists(path):
with open(path, encoding="utf-8") as handle:
for line in handle:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
name, value = line.split("=", 1)
if value.strip().strip('"').strip("'"):
names.add(name.strip())
return names
def golden_set(path: str):
"""Return (number of records, number of your own records, error text)."""
if not os.path.exists(path):
return 0, 0, "not found"
total = own = 0
with open(path, encoding="utf-8") as handle:
for number, line in enumerate(handle, 1):
if not line.strip():
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
return total, own, f"line {number} is not valid JSON"
if record.get("verdict") not in ("right", "partly", "wrong"):
return total, own, f"line {number} has no valid verdict"
total += 1
own += record.get("id") not in SAMPLE_IDS
return total, own, ""
print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
"the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")
print("\n2. Python libraries")
for module, package, fix in [
("inspect_ai", "inspect-ai", "pip install -r requirements.txt (Step 2)"),
("ranx", "ranx", "pip install -r requirements.txt (Step 2)"),
("sklearn", "scikit-learn", "added in Set up for Unit 2; pip install -r requirements.txt"),
("pytest", "pytest", "added in Testing AI applications (Unit 6); add pytest to requirements.txt"),
("dotenv", "python-dotenv", "added in Set up your computer; pip install -r requirements.txt"),
("anthropic", "anthropic", "added in Set up your computer; pip install -r requirements.txt"),
]:
version = library(module, package)
report(bool(version), f"{package} {version}".strip(), fix)
print("\n3. Course folder")
for name, step in [("golden_set.py", "3"), ("retrieval_metrics_hello.py", "4"), ("inspect_hello.py", "5")]:
path = os.path.join("unit08", name)
report(os.path.exists(path), path, f"create it (Step {step})")
total, own, error = golden_set(os.path.join("unit08", "data", "triage_golden.jsonl"))
report(total > 0 and not error, f"unit08/data/triage_golden.jsonl ({total} judged notes)",
error or "run python unit08/golden_set.py --sample (Step 3)")
report(own >= 5, f"your own judged notes in the golden set ({own})",
"add at least 5 of your own (Exercise)", optional=True)
report(os.path.isdir(os.path.join("unit08", "logs")), "unit08/logs (your first Inspect run)",
"run python unit08/inspect_hello.py (Step 5)", optional=True)
ignored = os.path.exists(".gitignore") and "unit08/logs" in open(".gitignore", encoding="utf-8").read()
report(ignored, ".gitignore leaves out unit08/logs", "add the line unit08/logs/ (Step 5)", optional=True)
print("\n4. A model for judging (needed from LLM evaluation fundamentals)")
names = env_names()
report("ANTHROPIC_API_KEY" in names, "ANTHROPIC_API_KEY in .env",
"add it as in Set up your computer; until then use the mock judge", optional=True)
print()
if problems:
print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
sys.exit(1)
print("All set. Your computer is ready for Unit 8.")
Run it:
python check_unit08.py
What success looks like (from our test, before adding your own notes or key; your versions will differ):
1. Python
OK Python 3.13.16
OK virtual environment is active
2. Python libraries
OK inspect-ai 0.3.276
OK ranx 0.3.21
OK scikit-learn 1.9.1
OK pytest 9.1.1
OK python-dotenv 1.2.4
OK anthropic 1.11.0
3. Course folder
OK unit08/golden_set.py
OK unit08/retrieval_metrics_hello.py
OK unit08/inspect_hello.py
OK unit08/data/triage_golden.jsonl (6 judged notes)
LATER your own judged notes in the golden set (0) -> add at least 5 of your own (Exercise)
OK unit08/logs (your first Inspect run)
OK .gitignore leaves out unit08/logs
4. A model for judging (needed from LLM evaluation fundamentals)
LATER ANTHROPIC_API_KEY in .env -> add it as in Set up your computer; until then use the mock judge
All set. Your computer is ready for Unit 8.
LATER lines don't stop you. The Exercise turns the first one into OK.
This section is short on purpose. Later Unit 8 topics show SAP's path in more detail.
Evaluations in the generative AI hub. As of SAP's SAP AI Core guide of 4 September 2026, Evaluations benchmarks models and prompts as orchestration configurations. It offers system-defined metrics, naming ROUGE, BLEU, COMET and tool-calling metrics, and custom LLM-as-a-judge metrics with rating criteria. It needs the extended plan.
The SDK you already have. The SAP Cloud SDK for AI (sap-ai-sdk-gen) from Unit 5 contains a gen_ai_hub.evaluations package. SAP's reference shows helpers that register datasets in SAP AI Core and read CSV, JSON and JSONL files from an object store. That is why the golden set is JSONL: the same file can move to SAP's service later.
What a trial may lack. SAP's path needs an object store connected to SAP AI Core. Treat it as a sketch until you have that access.
Need
Use
Why
Learn the habits on made-up data today
Inspect and ranx on your laptop
No account, logs you can read, free
Compare prompts and models for an SAP AI Core project
Evaluations in the generative AI hub
Runs the same orchestration configurations you deploy
Data in logs. An evaluation log holds every input and output. With real data, store logs where the data is allowed to live, with the same access rules.
Approved test data. Golden sets built from real orders copy business data out of SAP. Get approval, remove what the test doesn't need, and record where each case came from.
Judge cost. A model judge doubles the model calls of a run. Keep the golden set focused, and use rule-based scorers where a rule is enough.
Versioning. Keep the golden set in Git. A changed test set makes old and new numbers incomparable, so change it on purpose and note why.
Credentials. Keys stay in .env locally and in a secret store in the cloud, as in Unit 1.
Replace the made-up cases with your own. The result is the golden set the rest of Unit 8 uses.
Find the five triage notes you judged in Unit 1's exercise. If you didn't keep them, run Unit 1's triage.py --llm --sample again and judge the drafts now.
Open unit08/data/triage_golden.jsonl in VS Code.
For each note, add one line at the end, in the same shape as the samples. Use new IDs such as u1, u2. Copy the order fields and the draft from Unit 1's output, set verdict to right, partly or wrong, and write a one-sentence reason.
Save, then check the file:
python unit08/golden_set.py
Fix any line it reports.
Rerun the mock judge and note the new floor:
python unit08/inspect_hello.py
Run python check_unit08.py, then commit: git add unit08/data/triage_golden.jsonl and git commit -m "Unit 8: add my judged triage notes".
Done whencheck_unit08.py shows OK for "your own judged notes in the golden set" with at least 5, and golden_set.py prints counts that include your notes.
Pick one answer for each question. The explanation appears after you choose.
1In the Inspect task, what is the target of each sample?
Answer: C. Each sample's input is the order plus the draft, and its target is your verdict. The scorer compares the judge's verdict with that target, so accuracy here means agreement with you.
2Why does the mock judge that always says "partly" score 33% on the sample set?
Answer: B. The mock never reads the input. It matches only the cases whose target is "partly", two of six. That constant answer is the floor a real judge must beat.
3For question q3, the only relevant note came fourth. What do Recall@3 and MRR show?
Answer: D. Recall@3 only looks at the top 3, so it finds nothing. MRR uses the position of the first relevant result anywhere in the list, which is 4.
4ranx and the by-hand column print different numbers after you edit QRELS. What is the most likely cause?
Answer: A. Both columns use the same dictionaries and the same metric definitions, so they agreed in the test. A difference after an edit points to the data you changed, as the script's message says.
5Your colleague wants to commit unit08/logs so the team can see results. What do you suggest?
Answer: C. Logs save every input and reply, which may be sensitive once real data is used, and each run creates new ones. Commit the golden set and the code; share results through a report.
6inspect_hello.py fails with an error naming tiktoken and a blocked download. What is happening?
Answer: B. Inspect can count tokens with a tokenizer file it downloads once. A blocked network stops that; the script's mock judge sets usage itself to avoid it, and IT can allow the address.
7Why is the golden set stored as JSONL rather than in a spreadsheet?
Answer: D. Each line is one complete case, so additions and changes show clearly in Git. SAP's evaluations helpers also read JSONL, so the same file can move to SAP's service later.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Inspect (UK AI Security Institute documentation)— developed by the UK AI Security Institute and Meridian Labs; pip install inspect-ai; a task combines dataset, solver and scorer; inspect eval with --model; inspect view; built-in support for over 20 providers including Anthropic
Tutorial (Inspect documentation)— Task with Sample(input, target), system_message and generate solvers, match scorer; json_dataset and csv_dataset; logs saved to ./logs by default; inspect view opens a browser viewer
Scorers (Inspect documentation)— built-in scorers include includes, match, pattern, answer, exact, f1, choice, model_graded_qa and model_graded_fact; accuracy and stderr metrics
inspect-ai (PyPI)— version 0.3.276 of 2 October 2026; MIT licence; Python 3.10 or newer; author UK AI Security Institute
ranx (GitHub)— Qrels and Run built from dictionaries; evaluate(qrels, run, metrics); MIT licence; uses Numba for speed
Metrics (ranx documentation)— metric names such as hit_rate, precision, recall, mrr, map, ndcg, with @k cut-offs like recall@5 and ndcg@10
SAP AI Core service guide (PDF, 4 September 2026)— Evaluations benchmarks models and prompts as orchestration configurations; system-defined metrics such as ROUGE, BLEU, COMET and tool-calling metrics; custom LLM-as-a-judge metrics with rating criteria; generative AI hub only in the extended plan