Turn one-off AI tests into a harness that runs a fixed suite on every change, records each run with its version, compares it with a baseline and blocks regressions.
Earlier Unit 8 topics taught you to measure one thing at a time: a judge, a search, a hallucination rate. Each was a script someone ran once.
An evaluation harness turns those scripts into a routine. It holds a fixed set of test cases. Every time someone changes a prompt, a model or the retrieval setup, the harness runs all the cases and records the result with the exact version tested. It compares the run with the last approved one, called the baseline. If quality dropped, it blocks the change, the same way a failing software test blocks a release.
Think of it as the quality gate on a production line. The gate doesn't make the product better. It stops a worse product from shipping without anyone noticing.
AI systems change all the time, often without anyone touching your code. A team shortens a prompt to cut cost. A provider retires a model and you move to the next one. Someone adds documents to the search index. Each change can quietly break cases that used to work.
Take the blocked sales order assistant from earlier units. It drafts a note for a credit analyst about why an order is held. A developer edits the prompt to make notes shorter. The notes look fine in a demo. But on orders with two blocks, the note now mentions only one. On large credit orders, it starts telling the analyst to release the order. Nobody notices for three weeks, until an order ships to a customer over the credit limit.
A harness would have caught it before release, in minutes, for free or a small per-request charge. It runs the same orders every time and says exactly which ones got worse. It also gives you a record: on any date, you can show which prompt and model were live and how they scored.
The cost is mostly people's time. Someone has to own the test cases, decide the thresholds, and review failures. Plan for a few days to set it up for one use case, then regular time to keep the cases current.
As of SAP's SAP AI Core service guide of 4 September 2026, the generative AI hub includes Evaluations, added on 8 December 2025 under Optimizations. SAP describes it as benchmarking models and prompts as orchestration configurations, the same configurations you deploy. It offers system-defined metrics and custom LLM-as-a-judge metrics with rating criteria. It needs SAP AI Core with the extended service plan, which includes the generative AI hub.
The SAP Cloud SDK for AI for Python added an Evaluations client in version 6.5.0. Its helper module uploads your test data to an object store and registers it with SAP AI Core, so the evaluation can read it.
What SAP's service runs is one part of a harness: the evaluation itself. The other parts are still your team's job. Someone decides which cases go in the suite, what counts as a pass, when a drop blocks a release, and where the history is kept. SAP Learning's own prompt evaluation lesson shows the simplest form: a fixed test set, code checks per answer, and an average per prompt and model.
The fixed test cases, grouped into slices such as credit, billing or "no block set"
Business expert and engineer together
Does it include the cases that hurt most when wrong?
Runner
Runs every case through the system as it is configured now
Engineer
Does it test the exact prompt and model that go live?
Scorers
Checks that mark each answer: rules first, a judge model where rules can't decide
Engineer, checked by the business
Has anyone compared the scorer with expert judgment?
Record
Each run saved with the prompt, model, test suite and code version
Engineer
Can we say what was live on a given date and how it scored?
Gate
The rules that pass or block a change
Product owner
Who decided the thresholds, and who can override them?
The gate is a business decision written as numbers. Three kinds of rule are common:
Minimums. "Notes must mention every block on at least 80% of orders."
No big drops. "No score may fall more than 5 points below the baseline."
Must-pass cases. "These orders must be handled correctly every time." Pick the cases where a wrong answer costs money or trust, such as large credit holds.
"We tested it before go-live, so we're done." The model, the prompt and the documents keep changing. A harness repeats the test after every change.
"One overall score is enough." An average can rise while the most important cases fail. Look at slices and must-pass cases, not only the total.
"The harness tells us the system is good." It only tells you how the system does on the cases in the suite. A suite that misses a kind of order misses its failures too.
"A failed gate means the change is bad." Sometimes the suite or the scorer is wrong. The gate forces someone to look; a person decides.
"Raising the threshold improves quality." Thresholds only decide what gets blocked. Quality improves when the team fixes the failing cases.
Pick one answer for each question. The explanation appears after you choose.
1What does an evaluation harness add to the one-off evaluation scripts a team already has?
Answer: B. The harness doesn't improve the model or replace experts. It repeats the same cases after every change, records the result with the version, and blocks a change that makes things worse.
2A developer shortens the triage prompt to cut cost, and the demo looks fine. What is the main risk the harness guards against?
Answer: C. Changes that look fine in a demo can break specific cases that nobody tried. The harness reruns every case, including orders with two blocks and large credit holds, and names the ones that got worse.
3The overall score went up after a change, but two large credit-hold orders now get wrong advice. What should the gate do?
Answer: D. An average can rise while the most costly cases fail. Must-pass cases are the ones where a wrong answer costs money or trust, so any failure on them blocks the change, whatever the totals say.
4Who should decide the gate's thresholds?
Answer: A. The thresholds say how much risk the business accepts before a change goes live. Engineers build the gate, but the product owner decides and signs off on what blocks a release.
5As of September 2026, what does SAP's Evaluations in the generative AI hub cover?
Answer: C. SAP describes Evaluations as benchmarking prompts and models as orchestration configurations, with system-defined and custom judge metrics. Choosing the cases, setting the gate and keeping the history remain your team's job.
6After a production incident, a bad note reached an analyst. What is the best harness response?
Answer: D. A suite only catches failures it contains. Turning each incident into a test case makes the harness stronger over time, while raising thresholds alone changes what is blocked, not what is tested.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A harness is continuous integration for behavior. Unit tests ask "does the code still do what it did?" A harness asks "does the system still answer the way it did, on cases we chose on purpose?"
Every run produces one record. The record says what was tested and how it scored:
flowchart LR
V[Versions<br/>prompt, model, suite, commit] --> R[Runner]
S[(Suite<br/>cases with slices)] --> R
R --> O[Outputs]
O --> SC[Scorers<br/>rules, then judge]
SC --> REC[(Run record)]
B[(Baseline record)] --> G{Gate}
REC --> G
G -->|pass| SHIP[Change may go live]
G -->|fail| FIX[Report: what got worse]
Two runs are only comparable if they used the same suite. Everything else in the record (prompt, model, code) is what you are testing. Keep that rule in mind; most harness bugs break it.
The suite is a file of cases, each with an ID that never changes. A stable ID lets you say "case s04 passed last week and fails today". In this topic each case is a blocked sales order, as in Unit 1:
Slice groups cases by what makes them different: credit, billing, delivery, several blocks, or no block. A change that breaks one slice can hide inside a good average.
Must-pass marks the cases where a wrong answer is costly. Here: large credit holds and orders with no block set, where the note must admit the data can't say why.
The suite file gets a fingerprint: a short hash of its contents. Any edit changes it. The harness refuses to compare runs with different fingerprints.
The runner sends each case through the system under test. Here that is a prompt plus a model. The runner must use the same prompt file and model name that production uses. A harness that tests a copy of the prompt tests the wrong thing.
This topic uses Inspect, which you installed in Set up for Unit 8. As before, a task is a dataset, a solver and scorers. Inspect's documentation lists epochs, name, version and metadata among the task settings, and the -T option on inspect eval overrides task parameters from the command line. Our harness calls Inspect from Python so it can add the record and gate around it.
Epochs and non-determinism. Model output varies between calls. Inspect can run each case several times (epochs) and combine the scores with a reducer. Its built-in reducers include mean, mode, pass_at_{k} and at_least_{k}. The harness uses at_least_{n} with n epochs, so a case passes only if it passes every time. That is the strict reading a release gate wants: an answer that is right two times out of three is a problem in production.
#Scorers: rules first, a judge where rules can't decide
Each scorer answers one yes-or-no question about one output. Many narrow scorers beat one broad one. When a run fails, the scorer's name tells you what broke.
Scorer
Question
Why it is a rule, not a judge
valid_format
Is the reply JSON with likely_cause, next_check and confidence?
The format is exact
covers_facts
Does the note mention every block that is set, or say the data doesn't show why?
The fields are in the case
no_code_guessing
Does the cause avoid "means", "indicates" or "stands for"?
Unit 1's rule: codes differ by system
no_outside_orders
Does the note cite only its own order number?
Any other number is invented
judge_says_right (optional)
Does a judge model rate the note "right"?
Usefulness needs judgment
Inspect's scorers documentation shows the pattern: a function decorated with @scorer(metrics=[...]) that returns a Score with a value, an answer and an explanation. The explanation is what a person reads when a case fails, so write it for them.
Rules are cheap, exact and never drift. Their limit is that they only see what you wrote down. no_code_guessing misses "code 01 is a credit hold", and it would flag a harmless "this means you should check". The LLM evaluation fundamentals topic covers how to check a judge before you trust it.
Inspect's grouped() metric computes a metric separately for each value of a sample's metadata field, plus an all total. With grouped(accuracy(), "slice"), every scorer reports one number per slice. That turns "covers_facts dropped to 0.67" into "covers_facts is 0 on orders with several blocks".
Inspect saves a log per run; its documentation lists status, eval, results and samples among the fields and read_eval_log() to read one back. Logs default to ./logs and to the compact .eval format. Logs are detailed and big. The harness also writes a small run record, one JSON line per run:
Field
Why it is there
run_id
When the run happened
prompt, prompt_sha
Which prompt version, and a fingerprint of its exact text
writer, judge, epochs
Which model wrote the notes, which judged them, how many times each case ran
suite_sha
Fingerprint of the suite; runs are comparable only if it matches
git
The code commit, with +changes if files were edited but not committed
metrics, slices
Score per scorer, overall and per slice
cases
Pass or fail per case and scorer, so you can see exactly what changed
log
The Inspect log file with the full transcripts
The prompt fingerprint matters. Someone can edit v1.txt without renaming it. The name stays v1; the fingerprint changes.
compare reads the newest record and the baseline and applies four rules from eval_config.json:
Same suite. If the fingerprints differ, the gate fails and tells you to re-baseline first.
Minimums. Each scorer must reach its floor.
Maximum drop. No scorer may fall more than max_drop below the baseline.
Must-pass cases. Any failure on a must-pass case fails the gate.
It also lists newly failing and newly passing cases. A change can fix one thing and break another; averages hide that, case lists don't.
The program ends with exit code 1 when the gate fails. That single number is what lets GitHub Actions, or any other CI service, stop a change.
#Build it yourself: a harness with a baseline and a gate
You will build a harness for the Unit 1 triage writer. It runs 12 made-up blocked orders, scores each note with four rules, saves a record, compares it with a baseline and blocks a worse prompt. Then you will run it on GitHub for every push.
In VS Code, right-click the unit08 folder, choose New File, name it harness.py, paste the code below and save.
"""Unit 8: an evaluation harness for the blocked-order triage writer.
A harness is the machinery around your tests: a fixed suite of cases, a runner, scorers, a results
store and a gate. Every run is recorded with what was tested (prompt, model, suite, Git commit),
compared with a saved baseline, and either passes or fails the gate.
How to run (from your course folder, with .venv turned on):
python unit08/harness.py init # write the suite, two prompt versions, the config
python unit08/harness.py run # evaluate the prompt named in eval_config.json
python unit08/harness.py run --prompt v2 # evaluate another prompt version
python unit08/harness.py history # every recorded run, oldest first
python unit08/harness.py baseline # make the latest run the baseline
python unit08/harness.py compare # latest run vs. baseline; exit code 1 if the gate fails
python unit08/harness.py run --model anthropic/claude-opus-5-5 # a real model writes the notes (small charge)
python unit08/harness.py run --judge anthropic/claude-opus-5-5 # add a model judge as a fifth scorer (small charge)
Without --model, a made-up "offline writer" stands in for the model, so every step works with no account.
"""
import argparse
import hashlib
import json
import os
import re
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path
from dotenv import load_dotenv
HERE = Path(__file__).resolve().parent
SUITE = HERE / "data" / "triage_suite.jsonl"
PROMPTS = HERE / "prompts"
CONFIG = HERE / "eval_config.json"
RESULTS = HERE / "results"
RUNS = RESULTS / "runs.jsonl"
BASELINE = RESULTS / "baseline.json"
REPORT = RESULTS / "report.md"
LOGS = HERE / "logs" / "harness"
# ---------------------------------------------------------------- the suite (made-up, SAP-shaped)
def order(number, customer, amount, delivery_block="", billing_block="", credit_status=""):
"""Header fields shaped like Unit 1's sales order. The codes are placeholders, not SAP meanings."""
return {"SalesOrder": number, "SoldToParty": customer, "TotalNetAmount": amount,
"TransactionCurrency": "USD", "TotalCreditCheckStatus": credit_status,
"DeliveryBlockReason": delivery_block, "HeaderBillingBlockReason": billing_block}
SAMPLE_SUITE = [
{"id": "s01", "slice": "credit", "must_pass": True, "order": order("9100001", "CUST-A", "72400.00", "01", "", "B")},
{"id": "s02", "slice": "credit", "must_pass": False, "order": order("9100002", "CUST-B", "8300.00", "01", "", "B")},
{"id": "s03", "slice": "credit", "must_pass": False, "order": order("9100003", "CUST-C", "15900.00", "01", "", "B")},
{"id": "s04", "slice": "credit", "must_pass": True, "order": order("9100004", "CUST-A", "51200.00", "01", "", "B")},
{"id": "s05", "slice": "billing", "must_pass": False, "order": order("9100005", "CUST-D", "940.00", "", "02")},
{"id": "s06", "slice": "billing", "must_pass": False, "order": order("9100006", "CUST-E", "2210.00", "", "04")},
{"id": "s07", "slice": "delivery", "must_pass": False, "order": order("9100007", "CUST-F", "3150.00", "03")},
{"id": "s08", "slice": "delivery", "must_pass": False, "order": order("9100008", "CUST-G", "610.00", "05")},
{"id": "s09", "slice": "multi", "must_pass": False, "order": order("9100009", "CUST-H", "4480.00", "03", "02")},
{"id": "s10", "slice": "multi", "must_pass": False, "order": order("9100010", "CUST-B", "12650.00", "01", "02", "B")},
{"id": "s11", "slice": "no-block", "must_pass": True, "order": order("9100011", "CUST-I", "560.00")},
{"id": "s12", "slice": "no-block", "must_pass": True, "order": order("9100012", "CUST-J", "7020.00")},
]
PROMPT_V1 = """You write short triage notes for credit and order-management analysts about blocked sales orders.
Use only the order fields given. Do not guess what block or status codes mean; codes differ by system.
Mention every block or status that is set. If no block is set, say the data does not show why the order is held.
Reply with JSON only: {"likely_cause": "...", "next_check": "...", "confidence": "low|medium|high"}"""
PROMPT_V2 = """You write very short triage notes for analysts about blocked sales orders.
Focus on the main problem only. Keep each field under 12 words.
Reply with JSON only: {"likely_cause": "...", "next_check": "...", "confidence": "low|medium|high"}"""
SAMPLE_CONFIG = {
"prompt": "v1",
"gate": {
"min": {"valid_format": 1.0, "covers_facts": 0.8, "no_code_guessing": 0.9, "no_outside_orders": 0.9},
"max_drop": 0.05,
"must_pass_cases": True,
},
}
def sha(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()[:10]
def read_suite() -> list:
if not SUITE.exists():
sys.exit("No suite yet. Run: python unit08/harness.py init")
cases = [json.loads(line) for line in SUITE.read_text(encoding="utf-8").splitlines() if line.strip()]
ids = [c["id"] for c in cases]
if len(ids) != len(set(ids)):
sys.exit("Two cases in the suite share an id. Every case needs its own id.")
return cases
def read_config() -> dict:
if not CONFIG.exists():
sys.exit("No eval_config.json yet. Run: python unit08/harness.py init")
return json.loads(CONFIG.read_text(encoding="utf-8"))
def read_prompt(version: str) -> str:
path = PROMPTS / f"{version}.txt"
if not path.exists():
sys.exit(f"No prompt file {path.name} in unit08/prompts. Run init, or create the file.")
return path.read_text(encoding="utf-8").strip()
def git_commit() -> str:
"""The current commit, plus '+changes' if files are edited but not committed."""
try:
commit = subprocess.run(["git", "rev-parse", "--short", "HEAD"], capture_output=True, text=True,
check=True).stdout.strip()
dirty = subprocess.run(["git", "status", "--porcelain", "--", str(HERE)], capture_output=True,
text=True).stdout.strip()
return commit + ("+changes" if dirty else "")
except (OSError, subprocess.CalledProcessError):
return "no-git"
# ---------------------------------------------------------------- the offline writer (stands in for a model)
def offline_note(version: str, fields: dict) -> str:
"""Made-up behavior for each prompt version, so the harness has something to measure without a model.
v1 follows its rules, with one flaw: on CUST-B's order with two blocks it cites another order.
v2 ('shorter notes') keeps only the first block and, on large credit orders, guesses what code 01 means."""
delivery, billing = fields["DeliveryBlockReason"], fields["HeaderBillingBlockReason"]
credit, amount = fields["TotalCreditCheckStatus"], float(fields["TotalNetAmount"])
if version == "v2":
if delivery and credit and amount > 50000:
return json.dumps({"likely_cause": "Delivery block 01 means the credit limit is exceeded.",
"next_check": "Release the order in credit management.", "confidence": "high"})
if delivery:
return json.dumps({"likely_cause": "A delivery block is set.",
"next_check": "Ask sales about the delivery block.", "confidence": "medium"})
if billing:
return json.dumps({"likely_cause": "A billing block is set.",
"next_check": "Ask billing about the block.", "confidence": "medium"})
return json.dumps({"likely_cause": "No block is set; the data does not show why it is held.",
"next_check": "Ask who flagged the order.", "confidence": "low"})
parts, checks = [], []
if delivery:
parts.append(f"a delivery block ({delivery})")
checks.append("ask sales which delivery block this code is in your system")
if billing:
parts.append(f"a billing block ({billing})")
checks.append("ask billing what the block code stands for in your system")
if credit:
parts.append(f"a credit check status ({credit})")
checks.insert(0, "check the customer's credit exposure and open items")
if not parts:
return json.dumps({"likely_cause": "No block or credit status is set, so the data does not show why the order is held.",
"next_check": "Confirm with the analyst why the order was flagged.", "confidence": "low"})
next_check = "; ".join(checks).capitalize() + "."
if fields["SoldToParty"] == "CUST-B" and len(parts) == 3:
next_check += " CUST-B also has order 9100002 blocked." # the flaw: an order it was never given
return json.dumps({"likely_cause": "The order has " + " and ".join(parts) + ".",
"next_check": next_check, "confidence": "medium"})
def offline_model(version: str):
from inspect_ai.model import ModelOutput, ModelUsage, get_model
def write(messages, tools, tool_choice, config):
fields = json.loads(messages[-1].text.split("Order fields:", 1)[1])
output = ModelOutput.from_content(model="offline-writer", content=offline_note(version, fields))
output.usage = ModelUsage(input_tokens=0, output_tokens=0, total_tokens=0)
return output
return get_model("mockllm/model", custom_outputs=write, memoize=False)
# ---------------------------------------------------------------- scorers: code checks first, a judge if asked
def parse_note(text: str):
"""The note as a dict, or None. Tolerates a ```json fence around the reply."""
text = re.sub(r"^```(?:json)?\s*|\s*```$", "", text.strip())
try:
note = json.loads(text)
except json.JSONDecodeError:
return None
return note if isinstance(note, dict) else None
def build_scorers(judge: str | None):
from inspect_ai.model import get_model
from inspect_ai.scorer import CORRECT, INCORRECT, Score, accuracy, grouped, scorer, stderr
metrics = [grouped(accuracy(), "slice"), stderr()]
def verdict(ok: bool, why: str, answer: str = "") -> Score:
return Score(value=CORRECT if ok else INCORRECT, answer=answer, explanation=why)
@scorer(metrics=metrics)
def valid_format():
async def score(state, target):
note = parse_note(state.output.completion)
ok = bool(note) and {"likely_cause", "next_check", "confidence"} <= set(note) \
and note.get("confidence") in ("low", "medium", "high")
return verdict(ok, "JSON with the three fields" if ok else "not the agreed JSON shape",
state.output.completion[:200])
return score
@scorer(metrics=metrics)
def covers_facts():
async def score(state, target):
note, fields = parse_note(state.output.completion), state.metadata["order"]
text = json.dumps(note).lower() if note else ""
missing = []
if fields["DeliveryBlockReason"] and "delivery" not in text:
missing.append("delivery block")
if fields["HeaderBillingBlockReason"] and "billing" not in text:
missing.append("billing block")
if fields["TotalCreditCheckStatus"] and "credit" not in text:
missing.append("credit status")
no_blocks = not (fields["DeliveryBlockReason"] or fields["HeaderBillingBlockReason"]
or fields["TotalCreditCheckStatus"])
if no_blocks and not re.search(r"does not show|doesn't show|not enough|insufficient", text):
missing.append("saying the data does not show why")
return verdict(not missing, "covers every block that is set" if not missing
else "misses: " + ", ".join(missing))
return score
@scorer(metrics=metrics)
def no_code_guessing():
async def score(state, target):
note = parse_note(state.output.completion) or {}
cause = str(note.get("likely_cause", "")).lower()
guess = re.search(r"\b(means|indicates|stands for)\b", cause)
return verdict(not guess, "explains no code" if not guess else f"guesses a code meaning: {cause}")
return score
@scorer(metrics=metrics)
def no_outside_orders():
async def score(state, target):
own = state.metadata["order"]["SalesOrder"]
others = sorted(set(re.findall(r"\b\d{7,10}\b", state.output.completion)) - {own})
return verdict(not others, "cites only its own order" if not others
else "cites orders it was never given: " + ", ".join(others))
return score
chosen = [valid_format(), covers_facts(), no_code_guessing(), no_outside_orders()]
if judge:
rubric = ("You grade a triage note about a blocked sales order. A good note uses only the fields given, "
"guesses no code meanings, mentions every block that is set and suggests a sensible check. "
"Reply with one line: VERDICT: right, VERDICT: partly or VERDICT: wrong.")
@scorer(metrics=metrics)
def judge_says_right():
async def score(state, target):
reply = await get_model(judge).generate(
f"{rubric}\n\nOrder fields:\n{json.dumps(state.metadata['order'], indent=2)}"
f"\n\nNote:\n{state.output.completion}")
found = re.search(r"VERDICT:\s*(right|partly|wrong)", reply.completion, re.I)
label = found.group(1).lower() if found else "no verdict"
return verdict(label == "right", f"judge said {label}", reply.completion[:200])
return score
chosen.append(judge_says_right())
return chosen
# ---------------------------------------------------------------- commands
def cmd_init(args) -> None:
files = {SUITE: "".join(json.dumps(c) + "\n" for c in SAMPLE_SUITE),
PROMPTS / "v1.txt": PROMPT_V1 + "\n", PROMPTS / "v2.txt": PROMPT_V2 + "\n",
CONFIG: json.dumps(SAMPLE_CONFIG, indent=2) + "\n"}
for path, text in files.items():
if path.exists() and not args.force:
print(f"Kept {path.relative_to(HERE.parent).as_posix()} (exists; add --force to overwrite)")
continue
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(text, encoding="utf-8")
print(f"Wrote {path.relative_to(HERE.parent).as_posix()}")
print(f"\nSuite: {len(read_suite())} cases. Next: python unit08/harness.py run")
def cmd_run(args) -> None:
load_dotenv()
for name in (args.model, args.judge):
if name and name.startswith("anthropic/") and not os.getenv("ANTHROPIC_API_KEY"):
sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or leave out --model/--judge.")
try:
from inspect_ai import Epochs, Task, eval
from inspect_ai.dataset import Sample
from inspect_ai.solver import generate, system_message
except ImportError:
sys.exit("inspect-ai is not installed. Run: pip install -r requirements.txt")
config, cases = read_config(), read_suite()
version = args.prompt or config["prompt"]
prompt = read_prompt(version)
samples = [Sample(id=c["id"], input="Order fields:\n" + json.dumps(c["order"], indent=2), target="",
metadata={"slice": c["slice"], "must_pass": c["must_pass"], "order": c["order"]})
for c in cases]
task = Task(dataset=samples, solver=[system_message(prompt), generate()],
scorer=build_scorers(args.judge), name="triage_harness",
epochs=Epochs(args.epochs, f"at_least_{args.epochs}")) # a case passes only if it passes every time
model = args.model or offline_model(version)
print(f"Running {len(samples)} cases x {args.epochs} epoch(s): prompt {version}, "
f"writer {args.model or 'offline (made-up)'}{', judge ' + args.judge if args.judge else ''}...")
log = eval(task, model=model, log_dir=str(LOGS), display="none")[0]
if log.status != "success":
sys.exit(f"The run did not finish ({log.status}): {log.error.message if log.error else ''}\n"
"Open the log with: inspect view --log-dir unit08/logs/harness")
per_case = {} # case id -> {scorer: 1.0 pass or 0.0 fail}, after combining the epochs
for reduction in log.reductions:
for s in reduction.samples:
per_case.setdefault(str(s.sample_id), {})[reduction.scorer] = float(s.value)
metrics, slices = {}, {}
for result in log.results.scores:
values = {k: round(v.value, 3) for k, v in result.metrics.items() if k != "stderr"}
metrics[result.name] = values.pop("all", None)
slices[result.name] = values
record = {
"run_id": datetime.now(timezone.utc).strftime("%Y%m%d-%H%M%S"),
"prompt": version, "prompt_sha": sha(prompt),
"writer": args.model or "offline", "judge": args.judge or "", "epochs": args.epochs,
"suite_sha": sha(SUITE.read_text(encoding="utf-8")), "cases_n": len(cases),
"git": git_commit(), "metrics": metrics, "slices": slices, "cases": per_case,
"must_pass": [c["id"] for c in cases if c["must_pass"]],
"log": Path(log.location).name,
}
RESULTS.mkdir(parents=True, exist_ok=True)
with RUNS.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(record) + "\n")
print(f"\nRun {record['run_id']} prompt {version} ({record['prompt_sha']}) suite {record['suite_sha']} git {record['git']}")
print(f"{'scorer':18} {'all':>5} " + " ".join(f"{s:>8}" for s in sorted(next(iter(slices.values())))))
for name, value in metrics.items():
print(f"{name:18} {value:5.2f} " + " ".join(f"{slices[name][s]:8.2f}" for s in sorted(slices[name])))
failed = [f"{cid} ({', '.join(n for n, v in sc.items() if v < 1)})"
for cid, sc in per_case.items() if any(v < 1 for v in sc.values())]
print("\nFailing cases: " + ("; ".join(failed) if failed else "none"))
print("Saved to unit08/results/runs.jsonl. Next: python unit08/harness.py compare")
def read_runs() -> list:
if not RUNS.exists():
sys.exit("No runs recorded yet. Run: python unit08/harness.py run")
return [json.loads(line) for line in RUNS.read_text(encoding="utf-8").splitlines() if line.strip()]
def cmd_history(args) -> None:
runs = read_runs()
names = list(dict.fromkeys(name for r in runs for name in r["metrics"])) # every scorer seen, in order
print(f"{'run':16} {'prompt':7} {'writer':10} {'suite':11} {'git':16} " + " ".join(f"{n[:12]:>12}" for n in names))
for r in runs:
cells = [f"{r['metrics'][n]:12.2f}" if n in r["metrics"] else f"{'-':>12}" for n in names]
print(f"{r['run_id']:16} {r['prompt']:7} {r['writer'][:10]:10} {r['suite_sha']:11} {r['git'][:16]:16} "
+ " ".join(cells))
def cmd_baseline(args) -> None:
runs = read_runs()
chosen = next((r for r in runs if r["run_id"] == args.run), None) if args.run else runs[-1]
if not chosen:
sys.exit(f"No run with id {args.run}. See: python unit08/harness.py history")
BASELINE.write_text(json.dumps(chosen, indent=2) + "\n", encoding="utf-8")
print(f"Baseline is now run {chosen['run_id']} (prompt {chosen['prompt']}). Commit unit08/results/baseline.json.")
def gate(run: dict, base: dict, rules: dict) -> list:
"""Return the reasons the run fails the gate; an empty list means it passes."""
problems = []
if run["suite_sha"] != base["suite_sha"]:
problems.append("the suite changed since the baseline, so the numbers aren't comparable: "
"run the baseline prompt on the new suite and save it as the baseline first")
for name, minimum in rules["min"].items():
value = run["metrics"].get(name)
if value is None:
problems.append(f"{name}: not measured in this run")
elif value < minimum:
problems.append(f"{name}: {value:.2f} is below the minimum {minimum:.2f}")
for name, before in base["metrics"].items():
after = run["metrics"].get(name)
if after is not None and before - after > rules["max_drop"]:
problems.append(f"{name}: dropped from {before:.2f} to {after:.2f} (allowed drop {rules['max_drop']:.2f})")
if rules.get("must_pass_cases"):
for cid in run["must_pass"]:
bad = [n for n, v in run["cases"].get(cid, {}).items() if v < 1]
if bad:
problems.append(f"must-pass case {cid} failed: {', '.join(bad)}")
return problems
def cmd_compare(args) -> None:
if not BASELINE.exists():
sys.exit("No baseline yet. Run the current prompt, then: python unit08/harness.py baseline")
run, base = read_runs()[-1], json.loads(BASELINE.read_text(encoding="utf-8"))
rules = read_config()["gate"]
lines = [f"# Evaluation report: run {run['run_id']} vs. baseline {base['run_id']}", "",
f"Prompt {run['prompt']} ({run['prompt_sha']}) vs. {base['prompt']} ({base['prompt_sha']}); "
f"writer {run['writer']}; suite {run['suite_sha']}; git {run['git']}.", "",
"| Scorer | Baseline | This run | Change |", "| --- | --- | --- | --- |"]
for name in run["metrics"]:
before, after = base["metrics"].get(name), run["metrics"][name]
change = f"{after - before:+.2f}" if before is not None else "new"
shown = f"{before:.2f}" if before is not None else "-"
lines.append(f"| {name} | {shown} | {after:.2f} | {change} |")
newly, fixed = [], []
for cid, scores in run["cases"].items():
for name, value in scores.items():
was = base["cases"].get(cid, {}).get(name)
if was == 1 and value < 1:
newly.append(f"{cid} {name}")
elif was is not None and was < 1 and value == 1:
fixed.append(f"{cid} {name}")
lines += ["", "Newly failing: " + (", ".join(newly) or "none"), "",
"Newly passing: " + (", ".join(fixed) or "none")]
problems = gate(run, base, rules)
lines += ["", "## Gate: " + ("FAIL" if problems else "PASS"), ""] + [f"- {p}" for p in problems]
text = "\n".join(lines) + "\n"
REPORT.write_text(text, encoding="utf-8")
print(re.sub(r"(?m)^#+ ", "", text))
print("Report saved to unit08/results/report.md")
sys.exit(1 if problems else 0)
def main() -> None:
parser = argparse.ArgumentParser(description="Evaluation harness for the triage writer.")
sub = parser.add_subparsers(dest="command", required=True)
p = sub.add_parser("init", help="write the suite, prompts and config")
p.add_argument("--force", action="store_true", help="overwrite existing files")
p = sub.add_parser("run", help="run the suite and record the results")
p.add_argument("--prompt", help="prompt version in unit08/prompts, e.g. v2 (default: eval_config.json)")
p.add_argument("--model", help="a real model writes the notes, e.g. anthropic/claude-opus-5-5")
p.add_argument("--judge", help="add a model judge as an extra scorer, e.g. anthropic/claude-opus-5-5")
p.add_argument("--epochs", type=int, default=1, help="run each case this many times; it must pass every time")
sub.add_parser("history", help="list recorded runs")
p = sub.add_parser("baseline", help="save a run as the baseline")
p.add_argument("--run", help="run id (default: the latest)")
sub.add_parser("compare", help="compare the latest run with the baseline and apply the gate")
args = parser.parse_args()
{"init": cmd_init, "run": cmd_run, "history": cmd_history,
"baseline": cmd_baseline, "compare": cmd_compare}[args.command](args)
if __name__ == "__main__":
main()
What each part of the script does:
Part
What it does
SAMPLE_SUITE, order
Twelve made-up blocked orders in five slices; four are must-pass
PROMPT_V1, PROMPT_V2
Two versions of the triage prompt; v2 asks for very short notes
SAMPLE_CONFIG
The current prompt and the gate rules
sha, git_commit
Fingerprints for the prompt and suite; the Git commit, marked +changes if unit08 has uncommitted edits
offline_note, offline_model
The made-up writer, wrapped as Inspect's mockllm/model so Inspect treats it like any model
parse_note
Reads the JSON note, allowing the code fence some models put around JSON
build_scorers
Four rule scorers, plus judge_says_right when you pass --judge; each reports accuracy per slice with grouped
cmd_run
Builds the Inspect task, runs it, reads per-case results from log.reductions and appends a record to runs.jsonl
Epochs(n, "at_least_n")
Runs each case n times; it passes only if it passes every time
gate
The four gate rules; returns the list of reasons to fail
cmd_compare
Writes report.md with the score table, newly failing and passing cases and the gate result; exit code 1 on fail
Wrote unit08/data/triage_suite.jsonl
Wrote unit08/prompts/v1.txt
Wrote unit08/prompts/v2.txt
Wrote unit08/eval_config.json
Suite: 12 cases. Next: python unit08/harness.py run
Open unit08/eval_config.json. It says which prompt is current and holds the gate:
In a real project, the product owner signs off on these numbers. Changing them is a change like any other: it goes through Git and review.
Open unit08/prompts/v1.txt. This file is what the harness sends as the system message. In a real project, your application reads the same file, so the harness tests what ships.
Tell Git to ignore the run history and report, which are rebuilt on every computer. Open .gitignore, add these lines at the end and save. (unit08/logs/ is already there from the setup topic.)
Read it row by row. Version 1 passes almost everything. One case fails: on s10, the note cites another order the writer was never given. That is the same kind of error as t5 in the setup topic's golden set. The multi slice shows it: 0.50, one of two cases.
Look at the full transcripts in Inspect's viewer:
inspect view --log-dir unit08/logs/harness
Open the address the terminal prints (for example http://127.0.0.1:7575). Click the run, then sample s10, and look at its scores. The explanation for no_outside_orders reads cites orders it was never given: 9100002. Press Ctrl+C in the terminal to stop the viewer.
Newly failing: s01 no_code_guessing, s02 covers_facts, s03 covers_facts, s04 no_code_guessing, s09 covers_facts, s10 covers_facts
Newly passing: s10 no_outside_orders
Gate: FAIL
- covers_facts: 0.67 is below the minimum 0.80
- no_code_guessing: 0.83 is below the minimum 0.90
- covers_facts: dropped from 1.00 to 0.67 (allowed drop 0.05)
- no_code_guessing: dropped from 1.00 to 0.83 (allowed drop 0.05)
- must-pass case s01 failed: no_code_guessing
- must-pass case s04 failed: no_code_guessing
Report saved to unit08/results/report.md
1
Read what the harness found:
On credit orders with a delivery block, v2 keeps only the delivery block and drops the credit status (s02, s03). On orders with two blocks it drops the second one (s09, s10).
On the two large credit orders, both must-pass, it says what code 01 means and tells the analyst to release the order (s01, s04). Open the viewer to read those notes.
It also fixed something: s10 no longer cites an outside order. Shorter notes have less room to invent. The harness shows both sides, so the team can make an informed trade.
The last line, 1, is the exit code. CI reads it as "failed".
Every row says which prompt, writer, suite and commit produced the scores. That is the audit trail.
#Step 7: Run each case several times, and try a real model
Run every case three times. A case passes only if it passes all three:
python unit08/harness.py run --epochs 3
The offline writer always writes the same note, so the scores match Step 4. With a real model they often don't, and that is the point: a note that is right two times out of three fails here.
Optional, small charge. With ANTHROPIC_API_KEY in .env from Unit 1, let a real model write the notes from the same prompt:
python unit08/harness.py run --model anthropic/claude-opus-5-5
This sends 12 requests. You can use any model your key supports; write its name after anthropic/. Then run compare. The baseline was the offline writer, so expect differences. To make a real model your baseline, run it, check the notes in the viewer, then run baseline.
Optional, small charge. Add a model judge as a fifth scorer:
python unit08/harness.py run --judge anthropic/claude-opus-5-5
This sends 12 more requests. The judge's verdicts appear as judge_says_right. Trust it only after you have checked its agreement with your own verdicts, as in LLM evaluation fundamentals.
In unit08, create requirements-eval.txt with these two lines and save. The CI job installs only what the harness needs, which is faster than the whole course list:
inspect-ai
python-dotenv
In .github/workflows (created in Unit 6), create unit08-eval.yml:
# Runs the Unit 8 evaluation harness on GitHub's computers and blocks regressions.
# Uses the offline writer, so it needs no key and costs nothing.
name: unit08-eval
on:
push:
paths: ["unit08/**", ".github/workflows/unit08-eval.yml"]
pull_request:
paths: ["unit08/**"]
workflow_dispatch: # adds a "Run workflow" button on the Actions tab
jobs:
eval-gate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install the harness libraries
run: |
python -m pip install --upgrade pip
pip install -r unit08/requirements-eval.txt
- name: Run the suite with the current prompt
run: python unit08/harness.py run
- name: Compare with the baseline and apply the gate
run: python unit08/harness.py compare
- name: Keep the report
if: ${{ always() }}
uses: actions/upload-artifact@v4
with:
name: eval-report
path: unit08/results/report.md
paths limits the job to changes in unit08. if: ${{ always() }} keeps the report even when the gate fails, which is when you need it most.
On GitHub, click the Actions tab, then the run named unit08-eval, then eval-gate.
What success looks like: a green check next to eval-gate, and in the step Compare with the baseline and apply the gate the line Gate: PASS. Under Artifacts on the run's summary page, eval-report holds report.md.
Now make the risky change the way a teammate would. Open unit08/eval_config.json, change "prompt": "v1" to "prompt": "v2", save, then:
git commit -am "Switch triage to the shorter prompt"
git push
The new run shows a red cross on Compare with the baseline and apply the gate, with the six reasons from Step 6. Open a pull request from unit08-harness to main and the failed check appears there too.
Change "prompt" back to "v1", commit and push. The check turns green again.
On SAP AI Core, the generative AI hub's Evaluations runs the evaluation step for configurations you deploy. As of the SAP AI Core service guide of 4 September 2026:
Evaluations benchmarks models and prompts as orchestration configurations and was added under Optimizations on 8 December 2025.
It offers system-defined metrics and custom LLM-as-a-judge metrics with rating criteria.
Prompt optimization, in the same area, can take separate test and train datasets.
It requires the extended service plan, which includes the generative AI hub.
The SAP Cloud SDK for AI for Python added an Evaluations client in version 6.5.0. Its gen_ai_hub.evaluations.helpers module shows the moving parts:
Harness part
SAP side, as the SDK helpers show it
Suite file
Uploaded to an object store; helpers read and write CSV, JSON and JSONL
Suite registered for runs
upload_dataset_data_and_register_aicore_artifact registers the data as an SAP AI Core artifact
One run, or several configurations
single_evaluation_job_flow and multiple_evaluation_jobs_flow
Reading results
S3FileClient.get_sqlitedb_tables_data_from_s3 downloads a SQLite database from the object store and loads named tables
How the two fit together:
Keep the suite in Git. The JSONL file stays the source of truth with its fingerprint. Upload a copy for SAP runs.
Test the deployed configuration. On SAP, the thing under test is the orchestration configuration, not a prompt file on your laptop. Record its name and version in the run record where this harness records prompt and prompt_sha.
Map metrics to scorers. Format and fact checks stay as rules. SAP's custom judge metrics play the role of judge_says_right.
Gate in your pipeline. Read the SAP results, write the same run record, and call the same gate function. The gate rules belong to your release process, not to the evaluation service.
SAP Learning's lesson on evaluating prompts with the SDK uses the same core idea at small scale: a fixed test set of 20 customer emails, code checks per answer (valid JSON, correct category, sentiment and urgency), and an average per prompt and model. A harness adds the record, the baseline and the gate around that loop.
Data protection. Suites built from real SAP orders contain business data. Keep only the fields a case needs, get approval, and give logs and records the same access rules as the source system. Inspect logs hold every prompt and reply.
Authorizations. A suite exported from SAP no longer carries SAP authorization checks. Restrict who can read it, and record where each case came from.
Test what ships. The harness must read the same prompt file, or call the same orchestration configuration, as production. Fingerprint it in every record.
Suite changes. Adding cases is healthy, but it breaks comparison. Re-baseline in the same change, so reviewers see both.
Gate ownership and overrides. Write down who may change thresholds or override a failed gate, and record each override with a reason.
Cost. Every run with a real model costs one call per case per epoch, and a judge doubles it. Run the cheap rule checks on every push and the paid run on a schedule or before release.
Non-determinism. Use several epochs for must-pass cases, and gate on "passes every time".
Judge and model versions. Pin model versions in the record. When a provider retires a model, run the suite on the replacement and compare before switching.
Clean core. The harness runs outside S/4HANA, on copies or on API reads. It needs no custom code in the ERP.
Add a fifth rule, no_release_advice, that fails a note telling the analyst to release, remove or unblock anything. Unit 1's rule is that the assistant suggests checks and a person acts. Then grow the suite and re-baseline correctly.
Open unit08/harness.py. Find the line that starts with chosen = [valid_format(). Just above it, at the same indentation, paste:
@scorer(metrics=metrics)
def no_release_advice():
async def score(state, target):
note = parse_note(state.output.completion) or {}
advice = str(note.get("next_check", "")).lower()
acts = re.search(r"\b(release|remove|unblock)\b", advice)
return verdict(not acts, "suggests a check, not an action" if not acts
else f"tells the analyst to act: {advice}")
return score
python unit08/harness.py run --prompt v2
python unit08/harness.py compare
The gate fails, and three must-pass cases (s01, s04, s13) now fail no_release_advice as well as no_code_guessing.
Open unit08/results/report.md and add three sentences at the end: which slice v2 hurts most, which must-pass failure worries you most and why, and one case you would add next.
Save your work:
git add unit08/harness.py unit08/eval_config.json unit08/data/triage_suite.jsonl unit08/results/baseline.json
git commit -m "Add no_release_advice scorer and two cases; re-baseline"
Done whencompare on v1 prints Gate: PASS against a 14-case baseline, compare on v2 lists must-pass case s13 failed: no_code_guessing, no_release_advice, and report.md holds your three sentences. Keep the harness: you can reuse it to gate the agent changes you build in Unit 9.
Pick one answer for each question. The explanation appears after you choose.
1Why does the harness refuse to compare a run with the baseline when the suite fingerprints differ?
Answer: C. A run record is only comparable with another one on the same cases. If the suite changed, re-run the current prompt on the new suite and save that as the baseline, so the next comparison measures the system again.
2The harness runs each case with Epochs(3, "at_least_3"). What does that mean for a case?
Answer: B. at_least_3 over three epochs returns a pass only when all three runs pass. Model output varies between calls, and a release gate should treat an answer that is right two times out of three as a failure.
3Why does the run record store prompt_sha as well as the prompt version name?
Answer: D. Names are labels people choose, and the text behind a label can change. The fingerprint is computed from the exact text, so two runs with the same name and different hashes tested different prompts.
4In the v2 run, no_outside_orders improved while covers_facts and no_code_guessing dropped. Which part of the report makes that trade visible?
Answer: B. A change can fix some cases and break others, and averages blur that. The case lists name s10 as newly passing and six cases as newly failing, so the team can judge the trade case by case.
5How does GitHub Actions know the gate failed?
Answer: C. CI judges each step by its exit code: 0 passes, anything else fails. compare calls sys.exit(1) when the gate has reasons to fail, which marks the step red and blocks the pull request.
6Your team moves the triage assistant to SAP AI Core orchestration. What should change in the harness?
Answer: D. On SAP, the orchestration configuration is what ships, so it is what must be tested and fingerprinted. SAP Evaluations runs and scores it, but the thresholds and release decision stay in your pipeline, and exact format checks stay as rules.
7A provider announces that the model behind your triage writer retires next quarter. What do you do with the harness?
Answer: B. The harness exists to find regressions before they reach users. Running the replacement model on the same suite shows which cases break while the old model is still live, so the team can fix the prompt or choose another model in time.
8Which file belongs in Git, and which should stay out?
Answer: C. The baseline is the reference that every run, on a laptop or in CI, is compared with, so it must be versioned. Inspect logs are large, rebuilt by each run, and hold every prompt and reply, which may include sensitive data.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Eval Sets (Inspect documentation)— running many tasks and models as one set with a dedicated log directory; retries and resuming completed work
Log Files (Inspect documentation)— read_eval_log, list_eval_logs, EvalLog fields (status, eval, results, samples), .eval and .json formats, ./logs default and INSPECT_LOG_DIR
SAP AI Core service guide (PDF, 4 September 2026)— what's new 2025-12-08, Evaluations added to Optimizations; benchmarks prompts and models as orchestration configurations; system-defined and custom LLM-as-a-judge metrics; generative AI hub needs the extended service plan