Orchestrate

LLM evaluation fundamentals

Why judging AI answers is hard, and how golden sets, human review, LLM judges, pairwise comparison and regression tests make quality measurable.

Updated Oct 4, 2026Foundational 9 minDeep 38 min
Foundational layer · 9 min read

The 60-second version

A normal program gives the same answer every time, and a test can check it. A language model writes free text. Two answers can use different words and both be right. The same question can get different answers on different days. So "is it right?" needs a method, not a glance.

That method has four parts:

  • A golden set: real cases with answers that experts agreed are right.
  • Human review: an expert reads answers and grades them. It is the reference everything else is checked against.
  • An AI judge: a second model grades answers against written rules, so you can grade hundreds of cases cheaply. It is only trusted after it agrees with the experts.
  • Regression tests: the same cases, rerun after every change to the prompt or model, so nothing that worked quietly breaks.

Evaluation happens offline, before release, on the golden set. It continues online, after release, on real use.

Why it matters to the business

Without evaluation, every decision about an AI feature rests on a demo. Demos show the best case. Evaluation shows the rate of good answers on cases you chose in advance.

Take the running example from Unit 1: an AI drafts a note for each blocked sales order, and a credit analyst reads it. The team wants a shorter note. They change the prompt. On the five cases someone tries, the new notes look better. They ship.

A week later an analyst finds a note that says a block code "means a credit hold" and suggests releasing the order. The old prompt never did that. Nobody had rerun the old cases. That is a regression: a change that fixed some cases and broke one that worked.

Evaluation would have caught it in minutes. It pays off in three places:

  • Choosing. Comparing two prompts, two models or two vendors on the same cases, instead of on two demos.
  • Changing safely. Every prompt or model change is rerun against the golden set before release. Model providers retire versions, so this happens more often than teams expect.
  • Proving value. A business case needs a number: "on 50 past blocked orders, 44 notes were judged right". Unit 8 ends with linking such numbers to business metrics.

The cost is mostly expert time. Someone who knows the process must judge real cases at the start, and spot-check the AI judge later.

How SAP does it

As of SAP's SAP AI Core service guide of 4 September 2026, the generative AI hub includes Evaluations:

  • It benchmarks prompt templates and models as orchestration configurations, the same setup you met in Unit 5. You compare configurations to find the best combination for a use case.
  • You can pick system-defined metrics. SAP names ROUGE, BLEU and COMET, which compare text with a reference answer, and metrics for tool calling.
  • You can define custom LLM-as-a-judge metrics with your own rating criteria.
  • Prompt optimization evaluates prompt variants against a metric you choose and saves the improved prompt. SAP states that only judge metrics with numerical or Boolean outputs can drive it.
  • The generative AI hub is part of the extended service plan of SAP AI Core.

An SAP product expert's community post groups evaluation by output type. Structured output, such as a JSON field or a classification, is checked by exact match and validators. Free text is checked by reference metrics or an LLM judge. Agents are checked by comparing the tools they called with the expected calls.

The ideas in this topic apply whichever tool runs them. The course teaches them with open tools first, as set up in Set up for Unit 8.

Choosing a grading method

Method Good for Watch out for Rough cost
Rules in code (exact match, a schema check, "no release advice") Structured output, hard rules, safety lines Misses anything the rules don't name Near zero
Reference metrics (ROUGE, BLEU) Text that should be close to a known answer Two correct answers with different words score low Near zero
Human expert review The reference grade; new or risky cases Slow, costly, experts disagree with each other too Expert hours
AI judge, one answer at a time Grading many free-text answers against a rubric Biases (below); must be checked against experts Model calls per case
AI judge, two answers side by side Choosing between two prompts or models Prefers whichever answer it sees first Two model calls per pair

A sensible mix: rules for what can be ruled, an AI judge for the rest, and a person who reviews a sample of the judge's work every release.

What can go wrong with an AI judge

Research on AI judges, such as Zheng and colleagues' 2023 study, found strong judges agreed with human preferences more than 80% of the time. That was about as often as humans agreed with each other. The same work, and a paper by Wang and colleagues, documented biases a leader should know by name:

  • Position bias. Shown two answers, a judge may favour the first or the second regardless of content. Wang's team made a weaker model "beat" a stronger one on 66 of 80 questions, just by changing the order.
  • Verbosity bias. A judge may prefer the longer answer, even when the extra length adds nothing.
  • Self-enhancement bias. A judge may favour answers written by its own model family. Zheng's team saw signs of this but called the evidence limited.
  • Weak on what it can't do. A judge grading a calculation can be misled by a wrong answer. Giving it a reference answer helped in the study.

The fixes are routine: ask twice with the order swapped and call it a tie if the pick changes, write a precise rubric, give a reference answer, and measure agreement with your experts.

Questions to ask

  • Which cases are in the golden set, who judged them, and do they include past failures, not only clean examples?
  • How often does the AI judge agree with your experts, and how was that measured?
  • When two prompts or models were compared, was each pair judged in both orders?
  • Is the judge from the same model family as the system it grades?
  • What runs automatically before each prompt or model change, and what blocks a release?
  • How do we learn about bad answers after go-live, and who reviews them?
  • Which of SAP's evaluation metrics, if any, fit our use case, and do we have the extended plan?

Common misconceptions

  • "The model is good, so the feature is good." Public benchmarks measure general skill. Your golden set measures your process, your data and your prompt.
  • "An AI judge removes the need for experts." The judge copies the experts' standard. Without their verdicts, you can't tell a good judge from a confident one.
  • "Agreement of 80% means the judge is good." It depends on the cases. If most answers are right, a judge that always says "right" scores high too. Unit 8 shows a measure that corrects for this.
  • "The new prompt wins more comparisons, so ship it." It can still break cases the old one handled. Wins and regressions are separate questions.
  • "Evaluation is a project phase." It runs before every change and continues after go-live.

Key terms

  • Golden set: cases with answers experts agreed are right; the fixed exam.
  • Human review: an expert grading answers; the reference for every other method.
  • LLM-as-a-judge: a model that grades answers against a written rubric.
  • Rubric: the written criteria a grader applies.
  • Pairwise comparison: a judge picks the better of two answers to the same case.
  • Position bias: a judge's preference for an answer because of where it appears.
  • Offline evaluation: testing on a fixed set before release.
  • Online evaluation: measuring quality on real use after release, through ratings, reviews and monitoring.
  • Regression: a case that worked before a change and fails after it.
  • Regression test: rerunning known cases after each change to catch regressions.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why can't you test an AI-written note the way you test a normal program?

    Answer: B. A normal program gives one fixed output that a test can compare. A model writes free text that can be right in many wordings, and it can answer differently on different runs, so you need a method such as a golden set and a grader.
  2. 2What is the role of human expert review once an AI judge is in place?

    Answer: D. The judge copies the experts' standard. Experts grade a golden set first and then spot-check samples, so you can measure whether the judge still agrees with them.
  3. 3A team changed the triage prompt to get shorter notes. The new notes look better on five cases. What should happen before release?

    Answer: A. A change can improve some cases and break others. Rerunning all known cases, a regression test, is what would have caught the note that guessed a code's meaning and suggested a release.
  4. 4Your partner compared two prompts with an AI judge, showing each pair once. What is the main risk?

    Answer: C. Position bias is documented: Wang's team made a weaker model beat a stronger one on 66 of 80 questions just by changing the order. Judging each pair in both orders, and calling it a tie if the pick changes, guards against it.
  5. 5A vendor says its AI judge agrees with your experts 80% of the time. What is the best follow-up question?

    Answer: B. If most answers in the set are right, a judge that always says "right" also scores high. Agreement means little until you know the cases and compare it with such a lazy judge.
  6. 6Which SAP offering covers benchmarking prompts and models with custom judge metrics, as of September 2026?

    Answer: C. SAP's service guide describes Evaluations in the generative AI hub. It compares prompt templates and models as orchestration configurations, with system-defined metrics and custom LLM-as-a-judge metrics, in the extended plan.
  7. 7What does online evaluation add that offline evaluation can't?

    Answer: D. Offline evaluation runs on a fixed set before release. Online evaluation, through ratings, reviews and monitoring, shows real behaviour, and its bad cases become new golden-set entries.
Deep layer · 38 min read

Mental model: grade the grader, then gate the change

Evaluation has two loops. The inner loop asks "is this answer right?". The outer loop asks "is the grader right?". Most teams build the first and skip the second.

flowchart LR
  E[Expert verdicts<br/>golden set] --> A{Judge agrees<br/>with experts?}
  J[Judge<br/>rules or model] --> A
  A -- yes --> P[Compare versions<br/>pairwise, both orders]
  A -- no --> R[Fix rubric<br/>or judge]
  R --> J
  P --> G{Any case<br/>got worse?}
  G -- no --> S[Ship, then watch<br/>online signals]
  G -- yes --> X[Block release]
  S --> N[New failures]
  N --> E

Read it left to right. Experts judge a golden set. A judge is trusted only once it agrees with them. The trusted judge then compares versions and gates releases. Failures found after release go back into the golden set.

This topic builds each box once, small, on the golden set from Set up for Unit 8.

How it works

Why LLM output is hard to evaluate

Four properties make the usual "expected equals actual" test fail:

Property Example in the triage notes Consequence
Many right answers "Check credit exposure" and "Review the customer's open items and limit" are both fine Exact match is useless for free text
Partly right A note covers the delivery block and misses the billing block You need grades, not pass or fail
Non-deterministic The same order gets two different notes on two runs One run is a sample, not a fact
Fails silently A fluent note invents what code 01 means The failure looks like success

Anthropic's evaluation guidance suggests turning vague goals into specific, measurable criteria. "Good notes" is not testable. "Never states what a block code means" and "covers every block that is set" are.

Some outputs avoid the problem. If the model returns a JSON field, a classification or a tool call, you can check it with code. SAP's Felix Bartler makes the same split: exact match and validators for structured output, a judge or reference metrics for free text, and comparing call traces for agents. Push as much of your output into checkable structure as the task allows.

The golden set and human review

The golden set is the fixed exam. Each case holds an input, the output under test or a reference answer, and an expert's verdict with a reason. In this course the cases are blocked sales orders, Unit 1's draft notes, and your verdict: right, partly or wrong.

Three rules make it useful:

  • Seed it with failures. Anthropic's engineering team suggests starting with 20 to 50 tasks drawn from real failures. Clean examples flatter the system.
  • Write the reason. "Guesses what code 01 means" becomes a rubric line. A bare verdict teaches nothing.
  • Version it. Keep it in Git. A changed exam makes old and new scores incomparable, so change it on purpose.

Human review sets the standard but is slow. The evaluation guides agree: use it to build the golden set, to check the automated graders, and to read samples of real output. Don't use it to grade every run.

LLM-as-a-judge

A judge is a model given a rubric, the input and the answer, and asked for a grade. Zheng and colleagues describe three forms:

Form Judge sees Good for Weakness they report
Single-answer grading One answer, a scale or labels Scoring many answers, tracking a number over time Scores drift more when the judge model changes
Pairwise comparison Two answers to the same input Choosing between two versions Pairs grow quickly with many versions; position bias
Reference-guided grading One answer plus a reference answer Tasks with a checkable right answer Needs a reference for each case

The guidance from Anthropic's documentation on writing judge prompts is practical:

  • A detailed rubric with concrete conditions, not "is this good?".
  • A constrained output, such as one of three labels or a 1 to 5 score, so code can read it.
  • Reasoning before the grade. Letting the judge think first improves grades on complex judgments. Keep the reasoning in the log, and parse only the final grade.

Judge biases and their fixes

Bias What it looks like Fix
Position In Zheng's study, GPT-4 gave the same pick after swapping the order in 65% of pairs with the default prompt Judge each pair in both orders; a win counts only if both agree, otherwise a tie
Verbosity A padded answer with repeated list items beats the plain one Rubric says length is not a merit; test with padded answers
Self-enhancement The judge rates its own model family higher Use a judge from a different family, or a panel
Limited reasoning The judge is talked into accepting a wrong calculation Give a reference answer (reference-guided grading)

Wang and colleagues proposed three calibrations: ask for evidence before the score, average over both positions, and send cases where the judge is inconsistent to a human. That last one is a useful pattern in itself: disagreement between the two orderings is a signal that a person should look.

A panel of judges is another fix. In Inspect, model_graded_qa accepts a list of grader models. Its documentation says a grade wins only when more than half the panel returns it.

Measuring the judge: agreement and Cohen's kappa

Run the judge on the golden set and compare its verdicts with the experts'. Two numbers matter:

  • Agreement: the share of cases where judge and expert gave the same verdict.
  • Cohen's kappa: agreement corrected for chance. scikit-learn defines it as (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected if both picked labels at random with their own label frequencies. 1 is complete agreement; 0 is no better than chance.

Why both? A judge that always says "partly" agrees with you on every case you judged "partly". On a set where most cases are "partly", it looks good. Its kappa is 0. You will see exactly this in Step 2.

Also read the confusion matrix: rows are the expert's verdicts, columns the judge's. Which mistakes the judge makes matters more than how many. A judge that calls a wrong note right is worse than one that calls a right note partly.

Comparing versions: pairwise, both orders

To choose between prompt v1 and v2, show the judge both drafts for the same order and ask which is better. Do it twice, with the order swapped. Zheng's conservative rule: a version wins only if it is picked in both orders; otherwise the pair is a tie. Report the position consistency too: the share of pairs where both orders agreed. Low consistency means the judge can't tell the versions apart on those cases.

Regression testing

A regression is a case that passed before a change and fails after it. Anthropic's engineers separate capability evals, hard tasks where scores start low, from regression evals, which should pass nearly always. A regression suite answers one question: did anything that worked break?

The key point is that "v2 scores higher overall" and "v2 breaks nothing" are different tests. A prompt can raise the average and still break the case that matters most. So the gate compares case by case, not only averages. It returns a non-zero exit code on failure, so pytest or a CI pipeline can block the release, as in Testing AI applications.

Because models are non-deterministic, one run of a case is a sample. Anthropic's engineers describe pass@k (at least one of k tries succeeds) and pass^k (all k tries succeed). For a note an analyst relies on every time, consistency matters, so pass^k is the stricter and more honest measure.

Offline and online evaluation

flowchart LR
  subgraph Offline[Offline: before release]
    GS[(Golden set)] --> RUN[Run v2] --> GATE{Regression gate}
  end
  subgraph Online[Online: after release]
    USE[Real use] --> SIG[Ratings, overrides,<br/>sampled reviews]
  end
  GATE -- pass --> USE
  SIG -- bad cases --> GS

Offline evaluation runs on the golden set, before release, at no risk to users. Online evaluation measures real use: clerks' "helpful or not" ratings, how often an analyst overrides the AI's suggestion, a weekly expert review of a random sample. Anthropic's engineers recommend combining automated evals for fast iteration, production monitoring for ground truth, and periodic human review for calibration.

You already built one online signal. In CAP and side-by-side extensions, clerks could rate an explanation. Every "not helpful" rating is a candidate golden-set case.

Build it yourself: check a judge, compare two prompts, block a regression

Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder with its .venv, install pytest and the Unit 8 tools, and create unit08/data/triage_golden.jsonl with golden_set.py. This walkthrough doesn't repeat those steps.

You will write one script, judge_lab.py, with three commands. agree checks whether a judge agrees with your verdicts. pairwise compares drafts from two prompt versions, judged in both orders. regress finds cases the new version broke and fails if there are any. A small pytest file turns that into an automatic gate.

flowchart LR
  G[(triage_golden.jsonl<br/>orders, v1 drafts,<br/>your verdicts)] --> A[agree<br/>judge vs you]
  G --> P[pairwise<br/>v1 vs v2]
  V[(drafts_v2.jsonl<br/>new prompt)] --> P
  G --> R[regress<br/>case by case]
  V --> R
  R --> T[pytest gate<br/>pass or fail]

The script has three judges. rules is a few written checks: free, offline and repeatable. lazy always answers "partly", the floor any judge must beat. llm asks a real model, as Unit 1's triage.py did. Every step works without an account. Only Step 6 calls a model.

What you need

  • Your course folder from the setup topics, with unit08/data/triage_golden.jsonl.
  • No new accounts and no new libraries. The optional Step 6 uses the Anthropic key in .env from Unit 1: a small per-request charge (12 short requests for the six sample cases).
  • About 60 to 75 minutes.

Step 1: Open your course folder and check the golden set

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal with Terminal > New Terminal.

  3. Turn on the virtual environment:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate

    The prompt now starts with (.venv).

  4. Check the golden set (the same command on every system):

    python unit08/golden_set.py

What success looks like (with the six samples; your counts differ if you added your own notes):

Golden set: 6 judged notes
  right   3
  partly  2
  wrong   1

Share of drafts you judged right: 50%

If it says No golden set yet, run python unit08/golden_set.py --sample first.

Step 2: Check a judge against your verdicts

  1. In VS Code, right-click unit08, choose New File, name it judge_lab.py, paste the code below and save.
"""Unit 8: LLM evaluation fundamentals. Check a judge, compare two prompt versions, catch regressions.

Three commands, run in this order:
    agree     Does the judge agree with your verdicts in the golden set? (agreement and Cohen's kappa)
    pairwise  Which prompt version wrote better drafts? Each pair is judged twice, in both orders.
    regress   Did the new prompt version break a case the old one got right? Exits with code 1 if so.

Judges:
    --judge rules   a few written rules, no account, no cost (the default)
    --judge lazy    always says "partly": the floor any judge must beat
    --judge llm     a real model through Anthropic's API (small per-request charge; key in .env)

How to run (from your course folder, with .venv turned on):
    python unit08/judge_lab.py agree
    python unit08/judge_lab.py agree --judge lazy
    python unit08/judge_lab.py pairwise --sample      # also writes made-up "version 2" drafts
    python unit08/judge_lab.py regress
    python unit08/judge_lab.py agree --judge llm      # needs ANTHROPIC_API_KEY in .env
"""
import argparse
import json
import os
import re
import sys
from collections import Counter
from pathlib import Path

HERE = Path(__file__).resolve().parent
GOLDEN = HERE / "data" / "triage_golden.jsonl"   # from Set up for Unit 8: orders, v1 drafts, your verdicts
DRAFTS_V2 = HERE / "data" / "drafts_v2.jsonl"     # drafts from a changed prompt, for the same orders
VERDICTS = ("right", "partly", "wrong")

RUBRIC = (
    "You review draft notes that an assistant wrote for credit and order-management analysts. "
    "A right note uses only the order fields given, does not guess what block or status codes mean, "
    "covers every block that is set, says when the data is insufficient, and suggests a check, "
    "not an action such as releasing the order. "
    "partly = useful but incomplete or with a small error. wrong = misleading or unsafe."
)

# Made-up drafts from "prompt version 2", which asked for shorter notes. Same orders as the golden set.
SAMPLE_V2 = {
    "t1": {"likely_cause": "Delivery block 01 means a credit hold.",
           "next_check": "Release the order once credit confirms.", "confidence": "high"},
    "t2": {"likely_cause": "Billing block 02 is set; the data doesn't say why.",
           "next_check": "Ask billing what 02 stands for here.", "confidence": "low"},
    "t3": {"likely_cause": "A delivery block and a credit status are set; credit is worth checking first.",
           "next_check": "Check the customer's credit exposure.", "confidence": "medium"},
    "t4": {"likely_cause": "A delivery block and a billing block are both set.",
           "next_check": "Check both blocks with sales and billing.", "confidence": "medium"},
    "t5": {"likely_cause": "Delivery block with a non-blank credit status.",
           "next_check": "Check credit exposure; CUST-A also has order 9000001 blocked.", "confidence": "medium"},
    "t6": {"likely_cause": "No block fields are set, so the data doesn't show why it is held.",
           "next_check": "Ask the analyst why it was flagged.", "confidence": "low"},
}


# ---------- data ----------

def read_jsonl(path: Path, hint: str) -> list:
    if not path.exists():
        sys.exit(f"Missing {path.relative_to(HERE.parent)}. {hint}")
    return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]


def golden() -> list:
    return read_jsonl(GOLDEN, "Run: python unit08/golden_set.py --sample")


def drafts_v2(write_sample: bool) -> dict:
    if write_sample:
        DRAFTS_V2.write_text("".join(json.dumps({"id": k, "draft": v}) + "\n" for k, v in SAMPLE_V2.items()),
                             encoding="utf-8")
        print(f"Wrote {len(SAMPLE_V2)} made-up version 2 drafts to unit08/data/{DRAFTS_V2.name}\n")
    rows = read_jsonl(DRAFTS_V2, "Add --sample to write made-up version 2 drafts.")
    return {row["id"]: row["draft"] for row in rows}


# ---------- judges: each one takes (order, draft) and returns right, partly or wrong ----------

GUESSES_MEANING = re.compile(r"\b(block|code|status)\s+[0-9A-Z]{1,3}\s+(means|indicates)\b", re.I)
SUGGESTS_ACTION = re.compile(r"\brelease\b", re.I)


def rules_judge(order: dict, draft: dict) -> str:
    """Three written rules. Cheap and repeatable, but blind to anything the rules don't name."""
    text = f"{draft.get('likely_cause', '')} {draft.get('next_check', '')}"
    if GUESSES_MEANING.search(text) or SUGGESTS_ACTION.search(text):
        return "wrong"
    both_set = order.get("DeliveryBlockReason") and order.get("HeaderBillingBlockReason")
    if both_set and not ("delivery" in text.lower() and "billing" in text.lower()):
        return "partly"  # two blocks are set but the note covers only one
    return "right"


def lazy_judge(order: dict, draft: dict) -> str:
    return "partly"


def call_llm(system: str, user: str) -> str:
    """One request to a model through Anthropic's SDK, as in Unit 1. Replace only this to change provider."""
    import anthropic

    client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
    message = client.messages.create(
        model=os.environ.get("LLM_MODEL", "claude-opus-5-5"), max_tokens=400,
        system=system, messages=[{"role": "user", "content": user}])
    return "".join(block.text for block in message.content if block.type == "text")


def llm_judge(order: dict, draft: dict) -> str:
    reply = call_llm(RUBRIC + " Think briefly, then end with one line exactly like: VERDICT: right",
                     f"Order fields:\n{json.dumps(order, indent=2)}\n\nDraft note:\n{json.dumps(draft, indent=2)}")
    found = re.findall(r"VERDICT:\s*(right|partly|wrong)", reply, re.I)
    return found[-1].lower() if found else "unparsed"  # the last verdict line counts


JUDGES = {"rules": rules_judge, "lazy": lazy_judge, "llm": llm_judge}
SCORE = {"right": 2, "partly": 1, "wrong": 0, "unparsed": 0}


def rules_pairwise(order: dict, first: dict, second: dict) -> str:
    """Prefer the draft the rules score higher. On a tie it picks whichever it saw first: a position bias."""
    a, b = SCORE[rules_judge(order, first)], SCORE[rules_judge(order, second)]
    return "first" if a >= b else "second"


def llm_pairwise(order: dict, first: dict, second: dict) -> str:
    reply = call_llm(RUBRIC + " You see two drafts for the same order. Decide which is better. "
                     "Ignore length and order of presentation. Think briefly, then end with one line "
                     "exactly like: BETTER: 1  (or BETTER: 2, or BETTER: tie)",
                     f"Order fields:\n{json.dumps(order, indent=2)}\n\nDraft 1:\n{json.dumps(first, indent=2)}"
                     f"\n\nDraft 2:\n{json.dumps(second, indent=2)}")
    found = re.findall(r"BETTER:\s*(1|2|tie)", reply, re.I)
    return {"1": "first", "2": "second"}.get(found[-1].lower(), "tie") if found else "tie"


# ---------- metrics ----------

def cohen_kappa(human: list, judge: list) -> float:
    """Agreement beyond chance: (observed - expected) / (1 - expected). 1 = perfect, 0 = no better than chance."""
    n = len(human)
    observed = sum(h == j for h, j in zip(human, judge)) / n
    h_counts, j_counts = Counter(human), Counter(judge)
    expected = sum(h_counts[label] * j_counts[label] for label in set(human) | set(judge)) / (n * n)
    return 1.0 if expected == 1 else (observed - expected) / (1 - expected)


# ---------- commands ----------

def cmd_agree(judge_name: str) -> None:
    judge, cases = JUDGES[judge_name], golden()
    human = [c["verdict"] for c in cases]
    verdicts = [judge(c["order"], c["draft"]) for c in cases]
    print(f"Judge: {judge_name}   cases: {len(cases)}\n")
    print(f"{'id':5}{'you':9}{'judge':9}")
    for case, said in zip(cases, verdicts):
        flag = "" if said == case["verdict"] else "   <- disagree: " + case["reason"]
        print(f"{case['id']:5}{case['verdict']:9}{said:9}{flag}")
    labels = list(VERDICTS) + (["unparsed"] if "unparsed" in verdicts else [])
    print("\nConfusion matrix (rows = you, columns = judge)")
    print(" " * 9 + "".join(f"{label:>9}" for label in labels))
    for row in VERDICTS:
        counts = [sum(h == row and j == col for h, j in zip(human, verdicts)) for col in labels]
        print(f"{row:9}" + "".join(f"{count:>9}" for count in counts))
    agreement = sum(h == j for h, j in zip(human, verdicts)) / len(cases)
    print(f"\nAgreement: {agreement:.0%}   Cohen's kappa: {cohen_kappa(human, verdicts):.2f}")
    print("Kappa near 0 means no better than chance, however high the agreement looks.")


def cmd_pairwise(judge_name: str, write_sample: bool) -> None:
    if judge_name == "lazy":
        sys.exit("The lazy judge can't compare. Use --judge rules or --judge llm.")
    compare = llm_pairwise if judge_name == "llm" else rules_pairwise
    cases, v2 = golden(), drafts_v2(write_sample)
    tally, consistent = Counter(), 0
    print(f"Judge: {judge_name}. Each pair is shown twice: v1 first, then v2 first.\n")
    print(f"{'id':5}{'v1 first':11}{'v2 first':11}result")
    for case in cases:
        if case["id"] not in v2:
            continue
        one = compare(case["order"], case["draft"], v2[case["id"]])     # v1 shown first
        two = compare(case["order"], v2[case["id"]], case["draft"])     # v2 shown first
        pick_one = {"first": "v1", "second": "v2"}.get(one, "tie")
        pick_two = {"first": "v2", "second": "v1"}.get(two, "tie")
        result = pick_one if pick_one == pick_two else "tie (order changed the pick)"
        consistent += pick_one == pick_two
        tally[result if result in ("v1", "v2") else "tie"] += 1
        print(f"{case['id']:5}{pick_one:11}{pick_two:11}{result}")
    total = sum(tally.values())
    print(f"\nv2 wins {tally['v2']}, v1 wins {tally['v1']}, ties {tally['tie']} of {total}")
    print(f"Position consistency: {consistent} of {total} pairs gave the same pick in both orders")


def cmd_regress(judge_name: str, write_sample: bool) -> int:
    judge, cases, v2 = JUDGES[judge_name], golden(), drafts_v2(write_sample)
    before, after, broken = {}, {}, []
    print(f"Judge: {judge_name}\n\n{'id':5}{'v1':9}{'v2':9}")
    for case in cases:
        if case["id"] not in v2:
            continue
        before[case["id"]] = judge(case["order"], case["draft"])
        after[case["id"]] = judge(case["order"], v2[case["id"]])
        if before[case["id"]] == "right" and after[case["id"]] != "right":
            broken.append(case["id"])
        mark = "   <- REGRESSION" if case["id"] in broken else ""
        print(f"{case['id']:5}{before[case['id']]:9}{after[case['id']]:9}{mark}")
    rate = lambda results: sum(v == "right" for v in results.values()) / len(results)
    print(f"\nJudged right: v1 {rate(before):.0%}  ->  v2 {rate(after):.0%}")
    if broken or rate(after) < rate(before):
        print(f"FAIL: v2 breaks {len(broken)} case(s) v1 got right: {', '.join(broken) or '-'}. Don't ship v2.")
        return 1
    print("PASS: no case got worse. v2 may ship (after a person reads a sample of its drafts).")
    return 0


def main() -> None:
    parser = argparse.ArgumentParser(description="Check a judge, compare prompt versions, catch regressions.")
    parser.add_argument("command", choices=["agree", "pairwise", "regress"])
    parser.add_argument("--judge", choices=sorted(JUDGES), default="rules", help="default: rules")
    parser.add_argument("--sample", action="store_true", help="write made-up version 2 drafts first")
    args = parser.parse_args()

    if args.judge == "llm":
        try:
            from dotenv import load_dotenv
            import anthropic  # noqa: F401  (only checks that it is installed)
        except ImportError:
            sys.exit("Missing library. Run: pip install -r requirements.txt")
        load_dotenv()  # ANTHROPIC_API_KEY comes from .env, as in Unit 1
        if not os.getenv("ANTHROPIC_API_KEY"):
            sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or use --judge rules.")

    if args.command == "agree":
        cmd_agree(args.judge)
    elif args.command == "pairwise":
        cmd_pairwise(args.judge, args.sample)
    else:
        sys.exit(cmd_regress(args.judge, args.sample))


if __name__ == "__main__":
    main()
  1. Run the lazy judge first. It is the floor:

    python unit08/judge_lab.py agree --judge lazy

What success looks like (trimmed; from our test):

Confusion matrix (rows = you, columns = judge)
             right   partly    wrong
right            0        3        0
partly           0        2        0
wrong            0        1        0

Agreement: 33%   Cohen's kappa: 0.00
Kappa near 0 means no better than chance, however high the agreement looks.

The lazy judge agrees on the two notes you judged "partly". Its kappa is 0: it knows nothing. If your own golden set had mostly "partly" verdicts, its agreement would look high. Its kappa would still be 0.

  1. Now run the rules judge:

    python unit08/judge_lab.py agree

What success looks like (from our test):

Judge: rules   cases: 6

id   you      judge    
t1   right    right    
t2   right    right    
t3   wrong    wrong    
t4   partly   partly   
t5   partly   right       <- disagree: Good check, but it cites another order it was never given.
t6   right    right    

Confusion matrix (rows = you, columns = judge)
             right   partly    wrong
right            3        0        0
partly           1        1        0
wrong            0        0        1

Agreement: 83%   Cohen's kappa: 0.71
Kappa near 0 means no better than chance, however high the agreement looks.

Read the disagreement. In t5 the draft mentions order 9000001, which it was never given. None of the three rules looks for invented facts, so the rules judge calls it right. That is the typical weakness of code-based graders: fast and repeatable, blind to what they don't name. You can check the kappa yourself with scikit-learn's cohen_kappa_score; it gives the same 0.71.

Step 3: Compare two prompt versions, in both orders

Imagine the team changed the triage prompt to ask for shorter notes. --sample writes made-up drafts from that "version 2" prompt for the same six orders, to unit08/data/drafts_v2.jsonl.

  1. Run:

    python unit08/judge_lab.py pairwise --sample

What success looks like (from our test):

Wrote 6 made-up version 2 drafts to unit08/data/drafts_v2.jsonl

Judge: rules. Each pair is shown twice: v1 first, then v2 first.

id   v1 first   v2 first   result
t1   v1         v1         v1
t2   v1         v2         tie (order changed the pick)
t3   v2         v2         v2
t4   v2         v2         v2
t5   v1         v2         tie (order changed the pick)
t6   v1         v2         tie (order changed the pick)

v2 wins 2, v1 wins 1, ties 3 of 6
Position consistency: 3 of 6 pairs gave the same pick in both orders

The rules-based comparer has a built-in position bias: when it can't separate two drafts, it picks the first one it saw. That is how real judges fail too. Look at t2, t5 and t6. Shown once, each would count as a win for whichever version came first. Shown twice, the pick flips, so they are ties. Without the swap, the result would depend on the order you happened to use.

  1. Note what the summary hides. v2 wins more pairs, but t1 went to v1. A win count can't tell you whether that loss matters. Step 4 can.

Step 4: Find regressions, case by case

  1. Run:

    python unit08/judge_lab.py regress

What success looks like (from our test):

Judge: rules

id   v1       v2       
t1   right    wrong       <- REGRESSION
t2   right    right    
t3   wrong    right    
t4   partly   right    
t5   right    right    
t6   right    right    

Judged right: v1 67%  ->  v2 83%
FAIL: v2 breaks 1 case(s) v1 got right: t1. Don't ship v2.

The average went up from 67% to 83%. The gate still fails, because the v2 draft for t1 says "Delivery block 01 means a credit hold" and suggests releasing the order. That is the exact failure from the foundational layer's story.

  1. See the exit code. A program returns 0 on success and another number on failure, and test tools read it:

    • Windows (PowerShell):

      python unit08/judge_lab.py regress; $LASTEXITCODE
    • macOS / Linux:

      python unit08/judge_lab.py regress; echo $?

    The last line is 1.

Note that the v1 column is the judge's view, not yours. In Step 2 you saw the rules judge calls t5 right, while you said partly. That is why you check the judge before you let it gate a release.

Step 5: Make it an automatic gate with pytest

  1. In unit08, create test_regression.py, paste the code below and save.
"""Unit 8: a regression gate for prompt changes. Run from your course folder:  pytest unit08 -q"""
import subprocess
import sys
from pathlib import Path

LAB = Path(__file__).resolve().parent / "judge_lab.py"


def test_new_prompt_breaks_no_case_the_old_one_got_right():
    run = subprocess.run([sys.executable, str(LAB), "regress"], capture_output=True, text=True)
    assert run.returncode == 0, run.stdout + run.stderr
  1. Run it from the course folder:

    pytest unit08 -q

What success looks like (trimmed; from our test):

F                                                                        [100%]
...
E         t1   right    wrong       <- REGRESSION
...
E         FAIL: v2 breaks 1 case(s) v1 got right: t1. Don't ship v2.
...
FAILED unit08/test_regression.py::test_new_prompt_breaks_no_case_the_old_one_got_right
1 failed in 0.10s

A failing test is the right result here: the gate did its job. In the Exercise you fix the v2 draft and watch it pass. A CI pipeline, as in Testing AI applications, would run the same command on every prompt change.

Step 6 (optional, small charge): Try a real model as judge

  1. Check that .env in your course folder has your key from Unit 1:

    ANTHROPIC_API_KEY=sk-ant-...
  2. Run the agreement check with a model judge:

    python unit08/judge_lab.py agree --judge llm

    This sends six requests. The model name comes from LLM_MODEL in .env, as in Unit 1, and defaults to claude-opus-5-5.

  3. Compare its agreement and kappa with the rules judge. Look at t5: a model reading the order can notice the invented order number. Run it twice. If a verdict changes between runs, you have seen non-determinism in your own judge.

  4. Optionally run the comparison with the model, which sends 12 requests:

    python unit08/judge_lab.py pairwise --judge llm

    Its prompt tells the judge to ignore length and order. Position consistency tells you whether it managed.

We could not run Step 6 in our test environment without a key, so we have no sample output for it. Your numbers will differ from run to run.

Step 7: Save your work in Git

git add unit08/judge_lab.py unit08/test_regression.py unit08/data/drafts_v2.jsonl
git commit -m "Unit 8: judge agreement, pairwise comparison and a regression gate"

What each part of the script does

Part What it does
RUBRIC The judge's criteria, written from the reasons in your golden set
SAMPLE_V2 Made-up drafts from a "version 2" prompt, for the same six orders
rules_judge Three checks: guesses a code's meaning or says "release" (wrong); two blocks set but one missed (partly); otherwise right
lazy_judge Always "partly"; the floor any judge must beat
call_llm, llm_judge One Anthropic API call, as in Unit 1; reads the last VERDICT: line, or unparsed if there is none
rules_pairwise Compares two drafts by their rules grade; on a tie it picks the first one, a deliberate position bias
llm_pairwise Asks the model which draft is better; reads the last BETTER: line
cohen_kappa Agreement corrected for chance, in plain Python
cmd_agree Per-case table, confusion matrix, agreement and kappa
cmd_pairwise Judges each pair in both orders; counts a win only if both agree
cmd_regress Grades v1 and v2 case by case; returns exit code 1 if any right case got worse or the rate fell
test_regression.py Runs regress and fails the test if the exit code isn't 0

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
Missing unit08/data/triage_golden.jsonl The setup step wasn't run, or you're not in the course folder Run python unit08/golden_set.py --sample from orchestrate-course
Missing unit08/data/drafts_v2.jsonl No version 2 drafts yet Add --sample once: python unit08/judge_lab.py pairwise --sample
json.decoder.JSONDecodeError A line you edited in a .jsonl file is broken Each line must be one complete {...} object; python unit08/golden_set.py names the bad line in the golden set
pytest is not recognized .venv isn't active, or pytest isn't installed Turn on .venv; run pip install -r requirements.txt
Missing library. Run: pip install -r requirements.txt anthropic or python-dotenv isn't in this Python Check for (.venv), then pip install -r requirements.txt
ANTHROPIC_API_KEY is not set in .env The key from Unit 1 is missing Add it as in Set up your computer, or use the rules judge
AuthenticationError: Error code: 401 The key is wrong or was revoked Copy the key again from the Anthropic console into .env and save
NotFoundError naming the model Your account can't use that model name Set LLM_MODEL= in .env to a model your account lists
APIConnectionError, proxy or timeout errors Your network blocks the model API Try another network, or ask IT to allow the provider's API address; Steps 1 to 5 work offline
Many unparsed verdicts with --judge llm The model didn't end with a VERDICT: line Run again; if it persists, make the last sentence of the instructions stricter

The SAP way

As of SAP's SAP AI Core service guide of 4 September 2026, the same ideas map onto the generative AI hub:

Idea in this topic SAP AI Core, as documented Notes
Compare prompt v1 and v2, or two models Evaluations benchmarks prompt templates and models as orchestration configurations The configurations you deploy are the ones you test
Rules and reference metrics System-defined metrics; SAP names ROUGE, BLEU, COMET and metrics for tool calling Reference metrics need a reference answer per case
LLM-as-a-judge with a rubric Custom LLM-as-a-judge metrics with rating criteria Write criteria the way you wrote RUBRIC
Improve a prompt against the golden set Prompt optimization evaluates prompts against a chosen metric and saves the improved prompt SAP says only judge metrics with numerical or Boolean outputs work there
Golden set file Test data in an object store connected to SAP AI Core Set up for Unit 8 explains why the course keeps it as JSONL

What we could not confirm from SAP's documentation in this run: whether Evaluations offers a built-in pairwise mode with order swapping, and how its judge metrics handle position bias. If you need pairwise comparison on SAP AI Core, run each configuration with a pointwise judge metric and compare per case, as regress does. Check the current SAP guide before you design around a feature.

Build vs. SAP

Situation Use Why
Learning, prototyping, made-up data Plain Python or Inspect on a laptop Free, transparent, every grade readable
Production on SAP AI Core orchestration Evaluations in the generative AI hub Tests the exact configurations you deploy, close to the data
Structured output (JSON, a classification) Code checks, anywhere Exact and free; no judge needed
Many free-text answers to grade A judge: custom metric on SAP, or model_graded_qa in Inspect Scales; calibrate against experts first
Choosing between two versions Pairwise in both orders, or per-case pointwise grades Avoids position bias
Release gating in CI pytest calling your evaluation, or an SAP evaluation run triggered from the pipeline Fails the build on any regression

Production concerns

  • Data in prompts and logs. A judge sees the order data and the answer. Send it only where that data may go, and store evaluation logs with the same access rules as the source data. Golden sets built from real SAP orders need approval and should hold only the fields the test needs.
  • Authorizations. A golden set copied out of SAP no longer carries SAP's authorization checks. Restrict who can read it, and record where each case came from.
  • Judge drift. If the judge model is updated or retired, its grades change. Pin the judge model version, and rerun the agreement check whenever it changes.
  • Separate judge and system. Use a different model family for the judge where you can, or a panel, to reduce self-enhancement bias.
  • Cost. A judge adds at least one model call per case, and pairwise doubles it. Use code checks where a rule is enough, and keep the regression suite focused.
  • Non-determinism. Run important cases several times and gate on the share that pass every time, not on one lucky run.
  • Online signals. Log ratings and overrides with the prompt and model version, so a drop can be traced to a change. Feed bad cases back into the golden set.
  • Clean core. Evaluation runs outside S/4HANA, on copies or on API reads. It never needs custom code in the ERP.

Pitfalls

  • Grading the judge on its own training cases. Rules or prompts tuned on six cases score well on those six. Test on cases the judge's author hasn't seen.
  • Reporting agreement without kappa. On an unbalanced set, a lazy judge looks good.
  • Pairwise in one order only. The result can reflect position, not quality.
  • Gating on the average. An average can rise while a critical case breaks. Compare case by case.
  • Changing the golden set and the system in one step. You can't tell which change moved the number.
  • Not reading the transcripts. Anthropic's engineers stress reading the judge's reasoning. A judge can be right for the wrong reason.
  • Treating an LLM judge's verdict as ground truth. It is an estimate. Disagreements between orderings or between judges should go to a person.

Exercise

Fix the regression, then test the judge on cases it has never seen. The report you write feeds the evaluation harness later in Unit 8.

  1. Open unit08/data/drafts_v2.jsonl in VS Code. The first line is t1.

  2. Replace the t1 draft with one that follows the rubric. Keep it one line, for example:

    {"id": "t1", "draft": {"likely_cause": "A delivery block and a credit status are set; credit is worth checking first.", "next_check": "Check the customer's credit exposure and open items.", "confidence": "medium"}}
  3. Save, then run the gate:

    python unit08/judge_lab.py regress
    pytest unit08 -q

    regress now ends with PASS: no case got worse. and pytest prints 1 passed.

  4. Open unit08/data/triage_golden.jsonl. Add two new judged notes at the end, in the same shape, with IDs h1 and h2. Make them cases the rules can't see, for example a draft that invents an amount or a customer name, or one that says "probably a credit hold" for an order with no blocks set. Judge them yourself and write a reason.

  5. Check the file with python unit08/golden_set.py, then run python unit08/judge_lab.py agree. Note the new agreement and kappa.

  6. If you have a key, run python unit08/judge_lab.py agree --judge llm twice and note both results.

  7. Create unit08/judge_report.md with four lines: the rules judge's agreement and kappa before and after your new cases, the model judge's numbers (or "not run"), and one sentence on which judge you would trust to gate a release, and why.

  8. Commit: git add unit08 and git commit -m "Unit 8: fix regression, test judges on new cases".

Done when pytest unit08 -q prints 1 passed, your golden set has at least two new cases that you judged, and unit08/judge_report.md records the agreement and kappa for each judge you ran.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1The lazy judge agrees with you on 33% of the six cases. Why is its Cohen's kappa 0?

    Answer: C. Kappa is (p_o - p_e) / (1 - p_e). A judge that always says "partly" agrees only on the cases you judged "partly", which is exactly the agreement expected by chance, so the numerator is 0.
  2. 2In Step 2 the rules judge calls t5 right, while you judged it partly. What does that show?

    Answer: A. None of the three rules checks for facts the draft was never given, so the invented order 9000001 passes. Code checks are fast and repeatable but blind to what they don't encode, which is where a model judge or a person adds value.
  3. 3Why does cmd_pairwise judge each pair twice, with the order swapped?

    Answer: D. Judges can favour whichever answer they see first. Following Zheng and colleagues' conservative rule, a version wins only if it is picked in both orders; otherwise the pair is a tie.
  4. 4regress shows v2 judged right on 83% of cases versus 67% for v1, but returns FAIL. Why is that the correct behaviour?

    Answer: B. A regression is a case that worked before and fails after. The average rose, but the v2 note for t1 guesses a code's meaning and suggests a release, so the gate blocks it.
  5. 5You switch the judge from one model version to a newer one. What should you do before trusting its grades in the gate?

    Answer: C. Judge grades drift when the judge model changes. Rerunning agree against the experts' verdicts shows whether the new judge still matches the standard; rewriting the golden set with its verdicts would remove the standard.
  6. 6A teammate wants the same model family to write the notes and judge them. Which risk should you raise?

    Answer: D. Zheng and colleagues saw signs that judges favour answers from their own model family, though they called the evidence limited. Using a different family or a panel reduces the risk.
  7. 7Your team runs prompt changes on SAP AI Core. Which SAP feature lets you score orchestration configurations with your own rubric, as of September 2026?

    Answer: B. SAP's service guide says Evaluations benchmarks prompt templates and models as orchestration configurations and supports custom LLM-as-a-judge metrics with rating criteria. It needs the extended plan.
  8. 8After go-live, clerks mark 15 notes as "not helpful" in a week. What should happen to those cases?

    Answer: C. Online signals find failures the offline set missed. A person confirms them, and they become regression cases so the same failure is caught before the next release.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in