Orchestrate

Set up for Unit 8: evaluation tools

Add Inspect and ranx to your course folder, turn your judged Unit 1 triage notes into a golden set, and run a first evaluation with no new accounts.

Updated Oct 4, 2026Foundational 7 minDeep 35 min
Foundational layer · 7 min read

The 60-second version

Unit 8 is about evaluation: proving with numbers whether an AI system gives the right answers, and noticing when a change makes it worse. Without it, every demo is an opinion.

This setup adds two free, open-source tools and one file:

  • Inspect, an evaluation framework from the UK AI Security Institute. It runs a list of test cases through a model, scores each answer and keeps a log you can browse.
  • ranx, a small library that scores search results with the standard measures of retrieval quality.
  • A golden set: a short file of test cases with known right answers. You build it from the triage notes you judged at the end of What is an SAP FDE? in Unit 1.

No new accounts. Setup takes 30 to 50 minutes and costs nothing. A real model call later in the unit is a small per-request charge on the model key you already have.

Why it matters to the business

Leaders hear two kinds of claims about AI: "it works great" and "it hallucinates". Neither helps a decision. Evaluation turns both into a number on a fixed set of cases, such as "on 50 past blocked orders, the draft note was right 41 times". That number can be compared before and after a change, and between vendors.

Take the running example. In Unit 1 an AI drafted notes for blocked sales orders, and you judged five of them as right, partly right or wrong. Those judgments are the start of an evaluation set. Each time someone changes the prompt, the model or the data the AI reads, the team reruns the same cases. If the share of right notes drops, the change doesn't ship.

Three points for a leader:

  • Evaluation starts with people, not tools. Someone who knows the process must judge real cases first. The tools only repeat that judgment at scale.
  • The first number is often unflattering. That is useful. It is the baseline that later work must beat.
  • Open-source tools cost nothing to try. The cost is in the expert time to build and maintain good test cases.

How SAP does it

As of September 2026, SAP's service guide for SAP AI Core describes Evaluations in the generative AI hub:

  • It benchmarks models and prompts by running them as orchestration configurations, the same setup you met in Unit 5.
  • You can choose system-defined metrics. SAP names ROUGE, BLEU and COMET (scores that compare text with a reference answer) and metrics for tool calling.
  • You can define custom LLM-as-a-judge metrics, where a model grades answers against rating criteria you write.
  • The generative AI hub is available only in the extended service plan of SAP AI Core.

The SAP Cloud SDK for AI that you installed in Set up for Unit 5 already contains an evaluations module. SAP's reference shows it registering test data as datasets in SAP AI Core and reading files from an object store. That needs storage and access that a personal trial may not have, so this course teaches the ideas with open tools first. Later Unit 8 topics show the SAP path as a sketch.

What this unit adds

Item What it is Cost Used in
Inspect (inspect-ai) Open-source framework that runs test cases through a model, scores them and logs everything Free LLM evaluation fundamentals; Building an evaluation harness
ranx Library that scores search results (recall, MRR, nDCG) Free Evaluating retrieval
Golden set (triage_golden.jsonl) Your judged triage notes, as test cases Your time Every Unit 8 topic
Model key from Unit 1 Lets a model act as a judge Small per-request charge LLM evaluation fundamentals; Measuring hallucinations
scikit-learn and pytest (from Units 2 and 6) Metrics and an automatic test runner Free Reused across the unit

Time and money

  • Time: 30 to 50 minutes. Installing Inspect takes a few minutes because it brings many helper libraries. ranx is slow the first time it runs, then fast.
  • Money: nothing for the setup. The mock judge in the setup makes no model calls.
  • Later in the unit: running six cases through a real model is a small per-request charge. A golden set of 50 cases, rerun after each change, adds up; check the price of the model you pick.

Questions to ask IT

  • May learners install inspect-ai and ranx from the Python package index?
  • Does the network allow a one-time download from openaipublic.blob.core.windows.net? Inspect may fetch a small tokenizer file from there to count tokens.
  • May evaluation logs, which contain every prompt and answer, be stored on laptops? Where should they live for a real project?
  • Which past cases may be used as test data, and who approves that? For the setup, the answer is: made-up data and your own judged notes only.
  • Does the company have SAP AI Core with the extended plan, and an object store connected to it, for evaluations on SAP's side?

Common misconceptions

  • "Evaluation needs a big labelled dataset before we can start." Five judged cases are enough to start. The set grows as the team finds new failure cases.
  • "A model can grade itself, so we don't need experts." A model judge only helps once its verdicts agree with an expert's on known cases. Measuring that agreement is the first job.
  • "An evaluation tool is a quality guarantee." The tool counts. The test cases decide what it counts. Weak cases give a confident but meaningless number.
  • "We need SAP's evaluation service to start." SAP's service is useful in production on SAP AI Core. The habits (golden set, metrics, reruns) start on a laptop.

Key terms

  • Evaluation: measuring how often an AI system gives acceptable answers on a fixed set of cases.
  • Golden set: test cases with known right answers, agreed by people who know the process.
  • Baseline: the first measured number, which later changes must beat.
  • Judge (LLM-as-a-judge): a model that grades another model's answer against criteria.
  • Agreement: how often the judge's verdict matches the expert's.
  • Retrieval metrics: scores for search results, such as recall (did we find the right items?) and MRR (how high was the first right item?).
  • Inspect: an open-source evaluation framework from the UK AI Security Institute.
  • ranx: an open-source library for retrieval metrics.
  • Evaluation log: the saved record of every case, answer and score in a run.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does the Unit 8 setup turn your judged Unit 1 triage notes into?

    Answer: B. The judged notes become test cases that every later change is scored against. Nothing is retrained; the cases measure the system.
  2. 2The first evaluation shows only half of the draft notes are right. What is the best reading?

    Answer: C. An unflattering first number is still the baseline. Its value is that every later change can be compared with it on the same cases.
  3. 3A vendor says their model can grade its own answers, so no expert time is needed. What do you ask?

    Answer: B. A model judge is only useful once its verdicts match an expert's on cases with known answers. Agreement is the number that earns trust.
  4. 4Which SAP offering covers evaluating models and prompts, as of September 2026?

    Answer: C. SAP's service guide describes Evaluations in the generative AI hub, with system-defined and custom LLM-as-a-judge metrics. The generative AI hub is only in the extended service plan of SAP AI Core.
  5. 5Why does the course start with open-source tools instead of SAP's evaluation service?

    Answer: B. SAP's evaluations module works with datasets in SAP AI Core and files in an object store. The habits are the same, so the course teaches them on a laptop first and shows SAP's path later.
  6. 6Which question to IT matters for evaluation logs specifically?

    Answer: D. An evaluation log saves every input and output of a run. With real data, that is sensitive, so where logs live is a data protection decision.
Deep layer · 35 min read

Mental model: a fixed exam, rerun after every change

Evaluation is an exam with fixed questions. The golden set holds the questions and the right answers. A solver (your prompt plus a model) writes the answers. A scorer marks them. You rerun the same exam after every change and compare the marks.

flowchart LR
  G[(Golden set<br/>cases + right answers)] --> S[Solver<br/>prompt + model]
  S --> A[Answers]
  A --> SC[Scorer<br/>rule or judge model]
  G --> SC
  SC --> M[Metrics<br/>accuracy, recall, MRR]
  M --> L[(Log<br/>every case, every score)]

This setup builds each box once, small. Inspect provides the solver, scorer and log. ranx provides retrieval metrics. You provide the golden set.

How it works

Inspect: dataset, solver, scorer

Inspect is an evaluation framework developed by the UK AI Security Institute and Meridian Labs, published on PyPI under the MIT licence. As of 2 October 2026 the current version is 0.3.276, for Python 3.10 or newer. Its documentation builds every evaluation from three parts:

Part What it is In this setup
Dataset A list of Sample(input=..., target=...) One sample per judged triage note; the target is your verdict
Solver Steps that produce an answer, such as system_message(...) then generate() The judge instructions, then one model call
Scorer A function that marks the answer pattern(...), which pulls VERDICT: right out of the reply and compares it with the target

Inspect's scorer list also includes match, includes, exact, f1, choice and model-graded scorers such as model_graded_qa. The tutorial says logs go to ./logs by default and that inspect view opens a viewer in your browser. Models are named by provider, such as anthropic/<model>, which reads ANTHROPIC_API_KEY.

For the no-account path, the setup uses Inspect's mock model, mockllm/model, with a fixed reply. It behaves like a lazy judge that says "partly" to everything. That is useful, not just a stand-in: a constant answer is the floor any real judge must beat.

ranx: scoring search results

Retrieval is scored differently from answers. You need two things per question: which documents are relevant (qrels, short for relevance judgments) and what the search returned, in order (run). ranx builds both from Python dictionaries and computes metrics by name, such as recall@3, mrr and ndcg@3. It is MIT-licensed and uses Numba, a compiler for numeric Python, so its first call is slow while it compiles and later calls are fast.

Metric Plain meaning
Recall@k Share of the relevant documents found in the top k
MRR Average of 1 / position of the first relevant result
nDCG@k Rewards relevant results near the top, and very relevant ones more

The Evaluating retrieval topic explains them in depth. Here you only prove ranx gives the same numbers as a few lines of plain Python.

The golden set file

The golden set is a JSONL file: one JSON object per line, easy to add to and easy to compare in Git. Each line holds one judged note:

{"id": "t2", "order": {"SalesOrder": "9000002", "HeaderBillingBlockReason": "02"}, "draft": {"likely_cause": "...", "next_check": "...", "confidence": "low"}, "verdict": "right", "reason": "Admits what the data can't tell."}

The order fields and draft keys match what triage.py printed in Unit 1. verdict is your judgment: right, partly or wrong. reason is one sentence on why. That sentence matters: it is what turns a verdict into a rule a judge model can follow later.

Build it yourself: install the Unit 8 tools and run a first evaluation

You will install Inspect and ranx, create a golden set, score a small search result two ways, run an evaluation with a mock judge, and run a check. Every step works without an account.

Before you start: complete Set up your computer for this course and Set up for Unit 2. They install Python, VS Code and Git, and create your orchestrate-course folder with its .venv, .env and .gitignore. Testing AI applications in Unit 6 added pytest. This walkthrough doesn't repeat those steps.

flowchart LR
  S1[Steps 1-2<br/>install] --> S3[Step 3<br/>golden set]
  S3 --> S4[Step 4<br/>retrieval metrics]
  S4 --> S5[Step 5<br/>first Inspect run]
  S5 --> S6[Step 6<br/>check_unit08.py]

What you need

  • Your course folder from earlier units.
  • About 30 to 50 minutes.
  • No new accounts. Your Anthropic key from Unit 1 is optional, for one real-model run in Step 5.
  • Cost: free. The optional real-model run is a small per-request charge.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate

Run every command in this topic from the course folder.

Step 2: Add Inspect and ranx

  1. Open requirements.txt and add these two lines at the end, then save:

    inspect-ai
    ranx

    If pytest isn't in the file yet, add it too.

  2. Install (the same on every system):

    pip install -r requirements.txt

    Inspect brings many helper libraries, so this can take a few minutes.

  3. Check both:

    pip show inspect-ai ranx

What success looks like (trimmed; from our test, your versions may be newer):

Name: inspect_ai
Version: 0.3.276
...
Name: ranx
Version: 0.3.21

Inspect also installs a command, inspect. Check it:

inspect --version

You should see a version number such as 0.3.276.

Step 3: Create your golden set

The script writes six made-up judged notes, shaped like Unit 1's output, and checks the file. You will add your own notes in the Exercise.

  1. Make the Unit 8 folder:

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit08
    • macOS / Linux:

      mkdir -p unit08
  2. In VS Code, right-click unit08, choose New File, name it golden_set.py, paste the code below and save.

"""Unit 8: turn judged triage notes into a golden set, and check that the file is well formed.

A golden set is a fixed list of cases with a known right answer. Here each case is one blocked
sales order, the draft note an AI wrote for it in Unit 1, and your verdict on that note.
The file is unit08/data/triage_golden.jsonl: one JSON object per line.

How to run (from your course folder, with .venv turned on):
    python unit08/golden_set.py --sample     # write six made-up judged notes to start from
    python unit08/golden_set.py              # check the file and print a summary
    python unit08/golden_set.py --sample --force   # overwrite the file with the samples again
"""
import argparse
import json
import sys
from collections import Counter
from pathlib import Path

GOLDEN = Path(__file__).resolve().parent / "data" / "triage_golden.jsonl"
VERDICTS = ("right", "partly", "wrong")
REQUIRED = ("id", "order", "draft", "verdict", "reason")


def order(number, customer, amount, delivery_block, billing_block, credit_status):
    """Build the header fields Unit 1's triage.py read. The codes are placeholders, not SAP meanings."""
    return {"SalesOrder": number, "SoldToParty": customer, "TotalNetAmount": amount,
            "TransactionCurrency": "USD", "TotalCreditCheckStatus": credit_status,
            "DeliveryBlockReason": delivery_block, "HeaderBillingBlockReason": billing_block}


def draft(cause, check, confidence):
    return {"likely_cause": cause, "next_check": check, "confidence": confidence}


# Six made-up judged notes, shaped like Unit 1's output. Replace or extend them with your own.
SAMPLES = [
    {"id": "t1", "order": order("9000001", "CUST-A", "18250.00", "01", "", "B"),
     "draft": draft("Both a delivery block and a credit check status are set, so credit is worth "
                    "checking first.", "Check the customer's credit exposure and open items.", "medium"),
     "verdict": "right", "reason": "Uses only the fields given and points to the right next check."},
    {"id": "t2", "order": order("9000002", "CUST-B", "940.00", "", "02", ""),
     "draft": draft("A billing block is set; the data doesn't say why.",
                    "Ask billing which block code 02 stands for in this system.", "low"),
     "verdict": "right", "reason": "Admits what the data can't tell and asks for the code meaning."},
    {"id": "t3", "order": order("9000004", "CUST-D", "72400.00", "01", "", "B"),
     "draft": draft("Delivery block 01 means the customer is over the credit limit.",
                    "Release the order in credit management.", "high"),
     "verdict": "wrong", "reason": "Guesses what code 01 means and tells the analyst to release, not to check."},
    {"id": "t4", "order": order("9000005", "CUST-E", "3150.00", "03", "02", ""),
     "draft": draft("A delivery block is set.", "Check the delivery block with the sales team.", "medium"),
     "verdict": "partly", "reason": "Misses the billing block on the same order."},
    {"id": "t5", "order": order("9000006", "CUST-A", "12900.00", "01", "", "B"),
     "draft": draft("A delivery block with a non-blank credit status; possibly a credit hold.",
                    "Check credit exposure; CUST-A also has order 9000001 blocked.", "medium"),
     "verdict": "partly", "reason": "Good check, but it cites another order it was never given."},
    {"id": "t6", "order": order("9000007", "CUST-F", "560.00", "", "", ""),
     "draft": draft("No block fields are set, so the data doesn't show why this order is held.",
                    "Confirm with the analyst why it was flagged.", "low"),
     "verdict": "right", "reason": "Says the data is insufficient instead of inventing a cause."},
]


def validate(path: Path) -> list:
    """Read the golden set and return its records, or stop with the line number of the first problem."""
    if not path.exists():
        sys.exit(f"No golden set yet at {path}. Run: python unit08/golden_set.py --sample")
    records, seen = [], set()
    for number, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
        if not line.strip():
            continue  # blank lines are allowed
        try:
            record = json.loads(line)
        except json.JSONDecodeError as error:
            sys.exit(f"Line {number} is not valid JSON ({error.msg}). Each line must be one complete object.")
        missing = [field for field in REQUIRED if field not in record]
        if missing:
            sys.exit(f"Line {number} is missing: {', '.join(missing)}")
        if record["verdict"] not in VERDICTS:
            sys.exit(f"Line {number}: verdict must be one of {', '.join(VERDICTS)}, not {record['verdict']!r}")
        if record["id"] in seen:
            sys.exit(f"Line {number}: the id {record['id']!r} is used twice")
        seen.add(record["id"])
        records.append(record)
    if not records:
        sys.exit("The golden set is empty. Run with --sample, or add your own judged notes.")
    return records


def main() -> None:
    parser = argparse.ArgumentParser(description="Create or check the Unit 8 golden set of judged triage notes.")
    parser.add_argument("--sample", action="store_true", help="write six made-up judged notes")
    parser.add_argument("--force", action="store_true", help="with --sample, overwrite an existing file")
    args = parser.parse_args()

    if args.sample:
        if GOLDEN.exists() and not args.force:
            sys.exit(f"{GOLDEN.name} already exists. Add --force to overwrite it with the samples.")
        GOLDEN.parent.mkdir(parents=True, exist_ok=True)
        GOLDEN.write_text("".join(json.dumps(r) + "\n" for r in SAMPLES), encoding="utf-8")
        print(f"Wrote {len(SAMPLES)} made-up judged notes to unit08/data/{GOLDEN.name}")

    records = validate(GOLDEN)
    counts = Counter(r["verdict"] for r in records)
    print(f"\nGolden set: {len(records)} judged notes")
    for verdict in VERDICTS:
        print(f"  {verdict:7} {counts[verdict]}")
    share = counts["right"] / len(records)
    print(f"\nShare of drafts you judged right: {share:.0%}")
    print("That share is your first evaluation number: the quality of the Unit 1 drafts, judged by a person.")


if __name__ == "__main__":
    main()
  1. Write the samples and check them:

    python unit08/golden_set.py --sample

What success looks like (from our test):

Wrote 6 made-up judged notes to unit08/data/triage_golden.jsonl

Golden set: 6 judged notes
  right   3
  partly  2
  wrong   1

Share of drafts you judged right: 50%
That share is your first evaluation number: the quality of the Unit 1 drafts, judged by a person.
  1. Run it once more without --sample. It only checks the file. If you edit the file and break a line, it names the line, for example Line 7: verdict must be one of right, partly, wrong, not 'Right'.

Look at the six samples. Note t3: the draft guesses what code 01 means, which Unit 1 told the model not to do. Note t5: the draft mentions an order it was never given. Both kinds of error come back in the hallucination topic.

What each part of the script does:

Part What it does
GOLDEN Where the file lives: unit08/data/triage_golden.jsonl
order, draft Build records with the same fields as Unit 1's triage.py
SAMPLES Six made-up judged notes: three right, two partly, one wrong
validate Reads each line, checks the fields and verdict, and stops with the line number of the first problem
--sample, --force Write the samples; refuse to overwrite your file unless you add --force
main Prints counts by verdict and the share judged right

Step 4: Score a search result by hand and with ranx

The script takes three questions about the Unit 7 help notes (n1 to n8), the notes that are relevant to each, and what a search returned. It computes three metrics in plain Python, then asks ranx.

  1. In unit08, create retrieval_metrics_hello.py, paste the code below and save.
"""Unit 8 smoke test: score a tiny search result with three retrieval metrics, by hand and with ranx.

Three questions were asked of the Unit 7 help notes (IDs n1 to n8). For each, we know which notes
are relevant (the "qrels") and which notes the search returned, best first (the "run").
The script computes Recall@3, MRR and nDCG@3 in plain Python, then asks the ranx library for the
same numbers. If the two columns agree, the library is installed and you read it correctly.

How to run (from your course folder, with .venv turned on):
    python unit08/retrieval_metrics_hello.py              # by hand and with ranx
    python unit08/retrieval_metrics_hello.py --no-ranx    # by hand only (no library needed)
"""
import argparse
import math
import sys
import warnings

K = 3

# Which notes are relevant for each question. 2 = answers it fully, 1 = helps a bit.
QRELS = {
    "q1 why can't this order be delivered": {"n1": 2, "n3": 1},
    "q2 invoice quantity higher than goods receipt": {"n4": 2, "n6": 1},
    "q3 planned order arrives too late": {"n7": 2},
}

# What a search returned for each question, best first, with its similarity score.
RUN = {
    "q1 why can't this order be delivered": {"n3": 0.81, "n2": 0.77, "n1": 0.74, "n6": 0.40},
    "q2 invoice quantity higher than goods receipt": {"n4": 0.88, "n5": 0.71, "n6": 0.69, "n1": 0.30},
    "q3 planned order arrives too late": {"n8": 0.79, "n2": 0.55, "n5": 0.41, "n7": 0.39},
}


def ranked(results: dict) -> list:
    """Note IDs sorted by score, best first."""
    return sorted(results, key=results.get, reverse=True)


def recall_at_k(relevant: dict, results: dict, k: int) -> float:
    """Share of the relevant notes that appear in the top k."""
    top = ranked(results)[:k]
    return sum(1 for note in relevant if note in top) / len(relevant)


def reciprocal_rank(relevant: dict, results: dict) -> float:
    """1 divided by the position of the first relevant note (0 if none is returned)."""
    for position, note in enumerate(ranked(results), 1):
        if note in relevant:
            return 1 / position
    return 0.0


def ndcg_at_k(relevant: dict, results: dict, k: int) -> float:
    """Graded gain, discounted by position, divided by the best possible ordering."""
    gains = [relevant.get(note, 0) for note in ranked(results)[:k]]
    dcg = sum(g / math.log2(i + 2) for i, g in enumerate(gains))
    ideal = sorted(relevant.values(), reverse=True)[:k]
    idcg = sum(g / math.log2(i + 2) for i, g in enumerate(ideal))
    return dcg / idcg if idcg else 0.0


def by_hand() -> dict:
    n = len(QRELS)
    return {
        f"recall@{K}": sum(recall_at_k(QRELS[q], RUN[q], K) for q in QRELS) / n,
        "mrr": sum(reciprocal_rank(QRELS[q], RUN[q]) for q in QRELS) / n,
        f"ndcg@{K}": sum(ndcg_at_k(QRELS[q], RUN[q], K) for q in QRELS) / n,
    }


def with_ranx() -> dict:
    try:
        from ranx import Qrels, Run, evaluate
    except ImportError:
        sys.exit("ranx is not installed. Run: pip install -r requirements.txt  (or add --no-ranx)")
    warnings.filterwarnings("ignore", message="unsafe cast")  # a harmless notice from ranx's compiler
    print("Asking ranx (the first call compiles some code, so it can take several seconds)...")
    return evaluate(Qrels(QRELS), Run(RUN), [f"recall@{K}", "mrr", f"ndcg@{K}"])


def main() -> None:
    parser = argparse.ArgumentParser(description="Compute retrieval metrics by hand and with ranx.")
    parser.add_argument("--no-ranx", action="store_true", help="skip the ranx library")
    args = parser.parse_args()

    print("Per question (by hand):")
    for q in QRELS:
        print(f"  {q[:2]}  top {K}: {', '.join(ranked(RUN[q])[:K]):12}  "
              f"recall {recall_at_k(QRELS[q], RUN[q], K):.2f}  RR {reciprocal_rank(QRELS[q], RUN[q]):.2f}  "
              f"nDCG {ndcg_at_k(QRELS[q], RUN[q], K):.2f}")

    mine = by_hand()
    theirs = {} if args.no_ranx else with_ranx()
    print(f"\n{'metric':10} {'by hand':>8} {'ranx':>8}")
    for metric, value in mine.items():
        other = f"{theirs[metric]:8.3f}" if metric in theirs else "       -"
        print(f"{metric:10} {value:8.3f} {other}")
    if theirs:
        same = all(abs(mine[m] - theirs[m]) < 1e-6 for m in mine)
        print("\nThe two columns agree." if same else "\nThe columns differ: check the data you changed.")


if __name__ == "__main__":
    main()
  1. Run it:

    python unit08/retrieval_metrics_hello.py

What success looks like (from our test; the first run took about 40 seconds, later runs about 5):

Per question (by hand):
  q1  top 3: n3, n2, n1    recall 1.00  RR 1.00  nDCG 0.76
  q2  top 3: n4, n5, n6    recall 1.00  RR 1.00  nDCG 0.95
  q3  top 3: n8, n2, n5    recall 0.00  RR 0.25  nDCG 0.00
Asking ranx (the first call compiles some code, so it can take several seconds)...

metric      by hand     ranx
recall@3      0.667    0.667
mrr           0.750    0.750
ndcg@3        0.570    0.570

The two columns agree.

Read question q3: the one relevant note, n7, came fourth, outside the top 3. Recall@3 is 0, but MRR still gives a little credit (1/4). That is why teams report more than one metric.

  1. If ranx won't install or run, use --no-ranx. The by-hand column is enough to follow the Evaluating retrieval topic.

What each part of the script does:

Part What it does
QRELS The relevant notes per question; 2 = answers it, 1 = helps
RUN What the search returned, with scores; higher is better
recall_at_k, reciprocal_rank, ndcg_at_k The three metrics in plain Python
with_ranx Builds Qrels and Run from the same dictionaries and calls evaluate
warnings.filterwarnings(...) Hides a harmless notice that ranx's compiler prints

Step 5: Run a first evaluation with Inspect

The script turns each golden-set line into an Inspect sample. A judge reads the order and the draft and replies with a verdict. Inspect compares it with yours. Without --model, the judge is a mock that always says "partly" and makes no calls.

  1. In unit08, create inspect_hello.py, paste the code below and save.
"""Unit 8 smoke test: run a first evaluation with Inspect, using your golden set.

The task asks a model to act as a judge: for each judged triage note in
unit08/data/triage_golden.jsonl, it reads the order and the draft note and replies right, partly
or wrong. Inspect compares the model's verdict with yours and reports the share that agree.

How to run (from your course folder, with .venv turned on):
    python unit08/inspect_hello.py                                    # no account: a mock judge
    python unit08/inspect_hello.py --model anthropic/claude-opus-5-5  # a real model (small charge)
Then look at the results in your browser:
    inspect view --log-dir unit08/logs
"""
import argparse
import json
import os
import sys
from pathlib import Path

from dotenv import load_dotenv

HERE = Path(__file__).resolve().parent
GOLDEN = HERE / "data" / "triage_golden.jsonl"
LOGS = HERE / "logs"

JUDGE = (
    "You review draft notes that an assistant wrote for credit and order-management analysts. "
    "A good note uses only the order fields given, does not guess what block or status codes mean, "
    "says when the data is insufficient, and suggests a sensible next check. "
    "Judge the draft as right, partly (useful but incomplete or with a small error) or wrong. "
    "Think briefly, then end with one line exactly like: VERDICT: right"
)


def load_samples() -> list:
    """Turn each golden-set line into an Inspect Sample: input = order + draft, target = your verdict."""
    from inspect_ai.dataset import Sample

    if not GOLDEN.exists():
        sys.exit("No golden set yet. Run: python unit08/golden_set.py --sample")
    samples = []
    for line in GOLDEN.read_text(encoding="utf-8").splitlines():
        if line.strip():
            record = json.loads(line)
            text = (f"Order fields:\n{json.dumps(record['order'], indent=2)}\n\n"
                    f"Draft note:\n{json.dumps(record['draft'], indent=2)}")
            samples.append(Sample(id=record["id"], input=text, target=record["verdict"]))
    return samples


def lazy_judge():
    """A stand-in model for the no-account path: it answers 'partly' every time, without reading."""
    from inspect_ai.model import ModelOutput, ModelUsage, get_model

    def always_partly(messages, tools, tool_choice, config):
        output = ModelOutput.from_content(model="mockllm", content="I did not read it.\nVERDICT: partly")
        output.usage = ModelUsage(input_tokens=0, output_tokens=0, total_tokens=0)  # nothing to count offline
        return output

    return get_model("mockllm/model", custom_outputs=always_partly)


def main() -> None:
    parser = argparse.ArgumentParser(description="Run a first Inspect evaluation over the golden set.")
    parser.add_argument("--model", help="for example anthropic/claude-opus-5-5; leave out for the mock judge")
    args = parser.parse_args()

    load_dotenv()  # ANTHROPIC_API_KEY comes from .env, as in Unit 1
    if args.model and args.model.startswith("anthropic/") and not os.getenv("ANTHROPIC_API_KEY"):
        sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or leave out --model.")
    try:
        from inspect_ai import Task, eval
        from inspect_ai.scorer import pattern
        from inspect_ai.solver import generate, system_message
    except ImportError:
        sys.exit("inspect-ai is not installed. Run: pip install -r requirements.txt")

    task = Task(
        dataset=load_samples(),
        solver=[system_message(JUDGE), generate()],
        scorer=pattern(r"VERDICT:\s*(right|partly|wrong)"),  # pull the verdict out of the reply
        name="triage_judge",
    )
    model = args.model or lazy_judge()
    print(f"Judging {len(task.dataset)} notes with {args.model or 'a mock judge that always says partly'}...")
    logs = eval(task, model=model, log_dir=str(LOGS), display="plain")

    log = logs[0]
    if log.status != "success":
        sys.exit(f"The evaluation did not finish ({log.status}). Read the error above, or run: inspect view --log-dir unit08/logs")
    accuracy = log.results.scores[0].metrics["accuracy"].value
    print(f"\nAgreement with your verdicts: {accuracy:.0%} of {len(task.dataset)} notes")
    print("Log saved in unit08/logs. Browse it with: inspect view --log-dir unit08/logs")


if __name__ == "__main__":
    main()
  1. Run it with the mock judge:

    python unit08/inspect_hello.py

What success looks like (trimmed; from our test):

Judging 6 notes with a mock judge that always says partly...
...
triage_judge (6 samples): mockllm/model
...
pattern
accuracy  0.333
stderr    0.211
Log:
unit08/logs/2026-10-04T18-24-25-00-00_triage-judge_....eval

Agreement with your verdicts: 33% of 6 notes
Log saved in unit08/logs. Browse it with: inspect view --log-dir unit08/logs

The mock agrees on the two notes you judged "partly", so 2 of 6. Any real judge must beat 33% on this set. stderr is Inspect's estimate of how much that number could move with different cases; with six cases it is large.

  1. Look at the log in your browser:

    inspect view --log-dir unit08/logs

    The terminal shows Running on http://127.0.0.1:7575. Open that address. Click the run, then a sample, to see the input, the reply and the score. Press Ctrl+C in the terminal to stop the viewer.

  2. The logs are rebuilt by every run and will hold real prompts later, so keep them out of Git. Open .gitignore, add this line at the end and save:

    unit08/logs/
  3. Optional, small charge: with ANTHROPIC_API_KEY in .env from Unit 1, run a real judge:

    python unit08/inspect_hello.py --model anthropic/claude-opus-5-5

    This sends six short requests. You can use any model your key supports; write it after anthropic/. Note the agreement number. The LLM evaluation fundamentals topic is about making it trustworthy.

What each part of the script does:

Part What it does
JUDGE The judge's instructions, which mirror the rules Unit 1 gave the triage model
load_samples One Sample per golden-set line; input = order + draft, target = your verdict
lazy_judge Inspect's mockllm/model with a fixed reply, VERDICT: partly, and zero token usage
Task(...) Dataset + solver (system_message, generate) + scorer
pattern(r"VERDICT:...") Pulls the verdict out of the reply and compares it with the target
eval(..., log_dir=...) Runs the task and saves the log in unit08/logs
log.results.scores[0].metrics["accuracy"] Reads the agreement number back from the log

Step 6: Run the Unit 8 check

  1. In the course folder (not in unit08), create check_unit08.py, paste the code below and save. It uses only built-in Python, like the earlier checks.
"""Check that your computer is ready for Unit 8 (evaluation).

Run it from your course folder:  python check_unit08.py
It uses built-in Python only. It looks for the Unit 8 libraries and files, checks that your golden
set is well formed, and reads the names (not the values) of the keys in .env. It changes nothing.
"""
import importlib.metadata
import importlib.util
import json
import os
import sys

problems = 0
SAMPLE_IDS = {"t1", "t2", "t3", "t4", "t5", "t6"}


def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
    """Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
    global problems
    if ok:
        print(f"  OK       {label}")
    elif optional:
        print(f"  LATER    {label}  ->  {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label}  ->  {fix}")


def library(module: str, package: str) -> str:
    """Return the installed version of a library, or '' if it isn't installed."""
    if importlib.util.find_spec(module) is None:
        return ""
    try:
        return importlib.metadata.version(package)
    except importlib.metadata.PackageNotFoundError:
        return "installed"


def env_names(path: str = ".env") -> set:
    """Names of the settings in .env that have a value (the values are never printed)."""
    names = set()
    if os.path.exists(path):
        with open(path, encoding="utf-8") as handle:
            for line in handle:
                line = line.strip()
                if line and not line.startswith("#") and "=" in line:
                    name, value = line.split("=", 1)
                    if value.strip().strip('"').strip("'"):
                        names.add(name.strip())
    return names


def golden_set(path: str):
    """Return (number of records, number of your own records, error text)."""
    if not os.path.exists(path):
        return 0, 0, "not found"
    total = own = 0
    with open(path, encoding="utf-8") as handle:
        for number, line in enumerate(handle, 1):
            if not line.strip():
                continue
            try:
                record = json.loads(line)
            except json.JSONDecodeError:
                return total, own, f"line {number} is not valid JSON"
            if record.get("verdict") not in ("right", "partly", "wrong"):
                return total, own, f"line {number} has no valid verdict"
            total += 1
            own += record.get("id") not in SAMPLE_IDS
    return total, own, ""


print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
       "the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")

print("\n2. Python libraries")
for module, package, fix in [
    ("inspect_ai", "inspect-ai", "pip install -r requirements.txt (Step 2)"),
    ("ranx", "ranx", "pip install -r requirements.txt (Step 2)"),
    ("sklearn", "scikit-learn", "added in Set up for Unit 2; pip install -r requirements.txt"),
    ("pytest", "pytest", "added in Testing AI applications (Unit 6); add pytest to requirements.txt"),
    ("dotenv", "python-dotenv", "added in Set up your computer; pip install -r requirements.txt"),
    ("anthropic", "anthropic", "added in Set up your computer; pip install -r requirements.txt"),
]:
    version = library(module, package)
    report(bool(version), f"{package} {version}".strip(), fix)

print("\n3. Course folder")
for name, step in [("golden_set.py", "3"), ("retrieval_metrics_hello.py", "4"), ("inspect_hello.py", "5")]:
    path = os.path.join("unit08", name)
    report(os.path.exists(path), path, f"create it (Step {step})")
total, own, error = golden_set(os.path.join("unit08", "data", "triage_golden.jsonl"))
report(total > 0 and not error, f"unit08/data/triage_golden.jsonl ({total} judged notes)",
       error or "run python unit08/golden_set.py --sample (Step 3)")
report(own >= 5, f"your own judged notes in the golden set ({own})",
       "add at least 5 of your own (Exercise)", optional=True)
report(os.path.isdir(os.path.join("unit08", "logs")), "unit08/logs (your first Inspect run)",
       "run python unit08/inspect_hello.py (Step 5)", optional=True)
ignored = os.path.exists(".gitignore") and "unit08/logs" in open(".gitignore", encoding="utf-8").read()
report(ignored, ".gitignore leaves out unit08/logs", "add the line unit08/logs/ (Step 5)", optional=True)

print("\n4. A model for judging (needed from LLM evaluation fundamentals)")
names = env_names()
report("ANTHROPIC_API_KEY" in names, "ANTHROPIC_API_KEY in .env",
       "add it as in Set up your computer; until then use the mock judge", optional=True)

print()
if problems:
    print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
    sys.exit(1)
print("All set. Your computer is ready for Unit 8.")
  1. Run it:

    python check_unit08.py

What success looks like (from our test, before adding your own notes or key; your versions will differ):

1. Python
  OK       Python 3.13.16
  OK       virtual environment is active

2. Python libraries
  OK       inspect-ai 0.3.276
  OK       ranx 0.3.21
  OK       scikit-learn 1.9.1
  OK       pytest 9.1.1
  OK       python-dotenv 1.2.4
  OK       anthropic 1.11.0

3. Course folder
  OK       unit08/golden_set.py
  OK       unit08/retrieval_metrics_hello.py
  OK       unit08/inspect_hello.py
  OK       unit08/data/triage_golden.jsonl (6 judged notes)
  LATER    your own judged notes in the golden set (0)  ->  add at least 5 of your own (Exercise)
  OK       unit08/logs (your first Inspect run)
  OK       .gitignore leaves out unit08/logs

4. A model for judging (needed from LLM evaluation fundamentals)
  LATER    ANTHROPIC_API_KEY in .env  ->  add it as in Set up your computer; until then use the mock judge

All set. Your computer is ready for Unit 8.

LATER lines don't stop you. The Exercise turns the first one into OK.

Step 7: Save your work in Git

  1. Check what Git sees:

    git status

    You should see requirements.txt, .gitignore, check_unit08.py and unit08/. You must not see .env or unit08/logs.

  2. Save. The golden set is saved: it is your test data and should be versioned.

    git add requirements.txt .gitignore check_unit08.py unit08
    git commit -m "Set up Unit 8: Inspect, ranx and a golden set"

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
inspect-ai is not installed or ModuleNotFoundError: No module named 'ranx' The library isn't in the Python you are using Check for (.venv) in the prompt, then pip install -r requirements.txt
inspect is not recognized as a command The virtual environment isn't active, so its commands aren't found Turn on .venv (Step 1) and try again
pip stops while building numba or llvmlite No ready-made package for your Python version or chip Check python --version; use 3.11 to 3.14. Meanwhile run Step 4 with --no-ranx
ranx takes 30 seconds or more It compiles code on first use Wait; later runs are fast
No golden set yet Step 3 wasn't run, or not from the course folder Run python unit08/golden_set.py --sample from the course folder
Line N is not valid JSON A line you edited lost a quote, comma or brace Fix that line; each line must be one complete {...} object
An error naming openaipublic.blob.core.windows.net or tiktoken Inspect tried to download a tokenizer file and the network blocked it Try another network, or ask IT to allow that address. The mock judge in this script avoids the download
ANTHROPIC_API_KEY is not set in .env The key from Unit 1 is missing Add it as in Set up your computer, or leave out --model
Could not resolve authentication method or HTTP 401 The key is missing or wrong Copy the key again into .env and save
ConnectionError, proxy or timeout errors with --model Your network blocks the model API Try a home network, or ask IT to allow the provider's API address
inspect view says the port is in use A viewer is already running Close the other terminal, or add --port 7576
Windows: Activate.ps1 cannot be loaded PowerShell blocks scripts Run Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser, answer Y, and try again

Where this shows up in SAP

This section is short on purpose. Later Unit 8 topics show SAP's path in more detail.

  • Evaluations in the generative AI hub. As of SAP's SAP AI Core guide of 4 September 2026, Evaluations benchmarks models and prompts as orchestration configurations. It offers system-defined metrics, naming ROUGE, BLEU, COMET and tool-calling metrics, and custom LLM-as-a-judge metrics with rating criteria. It needs the extended plan.
  • The SDK you already have. The SAP Cloud SDK for AI (sap-ai-sdk-gen) from Unit 5 contains a gen_ai_hub.evaluations package. SAP's reference shows helpers that register datasets in SAP AI Core and read CSV, JSON and JSONL files from an object store. That is why the golden set is JSONL: the same file can move to SAP's service later.
  • What a trial may lack. SAP's path needs an object store connected to SAP AI Core. Treat it as a sketch until you have that access.
Need Use Why
Learn the habits on made-up data today Inspect and ranx on your laptop No account, logs you can read, free
Compare prompts and models for an SAP AI Core project Evaluations in the generative AI hub Runs the same orchestration configurations you deploy
Score search quality ranx, or plain Python Standard metrics, no service needed
Gate every change in CI pytest calling your evaluation Unit 6's habits, applied to quality

Production concerns

  • Data in logs. An evaluation log holds every input and output. With real data, store logs where the data is allowed to live, with the same access rules.
  • Approved test data. Golden sets built from real orders copy business data out of SAP. Get approval, remove what the test doesn't need, and record where each case came from.
  • Judge cost. A model judge doubles the model calls of a run. Keep the golden set focused, and use rule-based scorers where a rule is enough.
  • Versioning. Keep the golden set in Git. A changed test set makes old and new numbers incomparable, so change it on purpose and note why.
  • Credentials. Keys stay in .env locally and in a secret store in the cloud, as in Unit 1.

Pitfalls

  • Judging only the easy cases. A golden set of clean examples gives a flattering number. Add the cases that went wrong.
  • Changing the test set and the system at once. You can't tell which change moved the number.
  • Trusting a judge you haven't checked. Measure agreement with your own verdicts first, as Step 5 does.
  • Reading one number. As q3 showed, recall and MRR can tell different stories about the same search.
  • Committing logs. They are rebuilt by every run and may hold sensitive text. Keep unit08/logs/ in .gitignore.

Exercise

Replace the made-up cases with your own. The result is the golden set the rest of Unit 8 uses.

  1. Find the five triage notes you judged in Unit 1's exercise. If you didn't keep them, run Unit 1's triage.py --llm --sample again and judge the drafts now.

  2. Open unit08/data/triage_golden.jsonl in VS Code.

  3. For each note, add one line at the end, in the same shape as the samples. Use new IDs such as u1, u2. Copy the order fields and the draft from Unit 1's output, set verdict to right, partly or wrong, and write a one-sentence reason.

  4. Save, then check the file:

    python unit08/golden_set.py

    Fix any line it reports.

  5. Rerun the mock judge and note the new floor:

    python unit08/inspect_hello.py
  6. Run python check_unit08.py, then commit: git add unit08/data/triage_golden.jsonl and git commit -m "Unit 8: add my judged triage notes".

Done when check_unit08.py shows OK for "your own judged notes in the golden set" with at least 5, and golden_set.py prints counts that include your notes.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the Inspect task, what is the target of each sample?

    Answer: C. Each sample's input is the order plus the draft, and its target is your verdict. The scorer compares the judge's verdict with that target, so accuracy here means agreement with you.
  2. 2Why does the mock judge that always says "partly" score 33% on the sample set?

    Answer: B. The mock never reads the input. It matches only the cases whose target is "partly", two of six. That constant answer is the floor a real judge must beat.
  3. 3For question q3, the only relevant note came fourth. What do Recall@3 and MRR show?

    Answer: D. Recall@3 only looks at the top 3, so it finds nothing. MRR uses the position of the first relevant result anywhere in the list, which is 4.
  4. 4ranx and the by-hand column print different numbers after you edit QRELS. What is the most likely cause?

    Answer: A. Both columns use the same dictionaries and the same metric definitions, so they agreed in the test. A difference after an edit points to the data you changed, as the script's message says.
  5. 5Your colleague wants to commit unit08/logs so the team can see results. What do you suggest?

    Answer: C. Logs save every input and reply, which may be sensitive once real data is used, and each run creates new ones. Commit the golden set and the code; share results through a report.
  6. 6inspect_hello.py fails with an error naming tiktoken and a blocked download. What is happening?

    Answer: B. Inspect can count tokens with a tokenizer file it downloads once. A blocked network stops that; the script's mock judge sets usage itself to avoid it, and IT can allow the address.
  7. 7Why is the golden set stored as JSONL rather than in a spreadsheet?

    Answer: D. Each line is one complete case, so additions and changes show clearly in Git. SAP's evaluations helpers also read JSONL, so the same file can move to SAP's service later.

Sources

  • Inspect (UK AI Security Institute documentation) — developed by the UK AI Security Institute and Meridian Labs; pip install inspect-ai; a task combines dataset, solver and scorer; inspect eval with --model; inspect view; built-in support for over 20 providers including Anthropic
  • Tutorial (Inspect documentation) — Task with Sample(input, target), system_message and generate solvers, match scorer; json_dataset and csv_dataset; logs saved to ./logs by default; inspect view opens a browser viewer
  • Scorers (Inspect documentation) — built-in scorers include includes, match, pattern, answer, exact, f1, choice, model_graded_qa and model_graded_fact; accuracy and stderr metrics
  • Model providers (Inspect documentation) — Anthropic models are named anthropic/<model>; needs ANTHROPIC_API_KEY
  • inspect-ai (PyPI) — version 0.3.276 of 2 October 2026; MIT licence; Python 3.10 or newer; author UK AI Security Institute
  • ranx (GitHub) — Qrels and Run built from dictionaries; evaluate(qrels, run, metrics); MIT licence; uses Numba for speed
  • Metrics (ranx documentation) — metric names such as hit_rate, precision, recall, mrr, map, ndcg, with @k cut-offs like recall@5 and ndcg@10
  • SAP AI Core service guide (PDF, 4 September 2026) — Evaluations benchmarks models and prompts as orchestration configurations; system-defined metrics such as ROUGE, BLEU, COMET and tool-calling metrics; custom LLM-as-a-judge metrics with rating criteria; generative AI hub only in the extended plan
  • gen_ai_hub.evaluations.helpers package (SAP Cloud SDK for AI reference) — helpers for evaluation configurations, dataset artifacts registered in SAP AI Core, and reading and writing CSV, JSON, JSONL files in an S3 object store

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in