Orchestrate

Building an evaluation harness

Turn one-off AI tests into a harness that runs a fixed suite on every change, records each run with its version, compares it with a baseline and blocks regressions.

Updated Oct 5, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

Earlier Unit 8 topics taught you to measure one thing at a time: a judge, a search, a hallucination rate. Each was a script someone ran once.

An evaluation harness turns those scripts into a routine. It holds a fixed set of test cases. Every time someone changes a prompt, a model or the retrieval setup, the harness runs all the cases and records the result with the exact version tested. It compares the run with the last approved one, called the baseline. If quality dropped, it blocks the change, the same way a failing software test blocks a release.

Think of it as the quality gate on a production line. The gate doesn't make the product better. It stops a worse product from shipping without anyone noticing.

Why it matters to the business

AI systems change all the time, often without anyone touching your code. A team shortens a prompt to cut cost. A provider retires a model and you move to the next one. Someone adds documents to the search index. Each change can quietly break cases that used to work.

Take the blocked sales order assistant from earlier units. It drafts a note for a credit analyst about why an order is held. A developer edits the prompt to make notes shorter. The notes look fine in a demo. But on orders with two blocks, the note now mentions only one. On large credit orders, it starts telling the analyst to release the order. Nobody notices for three weeks, until an order ships to a customer over the credit limit.

A harness would have caught it before release, in minutes, for free or a small per-request charge. It runs the same orders every time and says exactly which ones got worse. It also gives you a record: on any date, you can show which prompt and model were live and how they scored.

The cost is mostly people's time. Someone has to own the test cases, decide the thresholds, and review failures. Plan for a few days to set it up for one use case, then regular time to keep the cases current.

How SAP does it

As of SAP's SAP AI Core service guide of 4 September 2026, the generative AI hub includes Evaluations, added on 8 December 2025 under Optimizations. SAP describes it as benchmarking models and prompts as orchestration configurations, the same configurations you deploy. It offers system-defined metrics and custom LLM-as-a-judge metrics with rating criteria. It needs SAP AI Core with the extended service plan, which includes the generative AI hub.

The SAP Cloud SDK for AI for Python added an Evaluations client in version 6.5.0. Its helper module uploads your test data to an object store and registers it with SAP AI Core, so the evaluation can read it.

What SAP's service runs is one part of a harness: the evaluation itself. The other parts are still your team's job. Someone decides which cases go in the suite, what counts as a pass, when a drop blocks a release, and where the history is kept. SAP Learning's own prompt evaluation lesson shows the simplest form: a fixed test set, code checks per answer, and an average per prompt and model.

The five parts of a harness

Part What it is Who owns it Question to ask
Suite The fixed test cases, grouped into slices such as credit, billing or "no block set" Business expert and engineer together Does it include the cases that hurt most when wrong?
Runner Runs every case through the system as it is configured now Engineer Does it test the exact prompt and model that go live?
Scorers Checks that mark each answer: rules first, a judge model where rules can't decide Engineer, checked by the business Has anyone compared the scorer with expert judgment?
Record Each run saved with the prompt, model, test suite and code version Engineer Can we say what was live on a given date and how it scored?
Gate The rules that pass or block a change Product owner Who decided the thresholds, and who can override them?

The gate is a business decision written as numbers. Three kinds of rule are common:

  • Minimums. "Notes must mention every block on at least 80% of orders."
  • No big drops. "No score may fall more than 5 points below the baseline."
  • Must-pass cases. "These orders must be handled correctly every time." Pick the cases where a wrong answer costs money or trust, such as large credit holds.

Questions to ask

  • Which cases are in the suite, who chose them, and when were they last updated?
  • Does the harness run on every change to the prompt, the model or the documents, or only when someone remembers?
  • What exactly blocks a release, and who signed off on those thresholds?
  • Which cases are must-pass, and why those?
  • When a provider retires a model, how do we test the replacement before switching?
  • Where is the run history kept, and how long?
  • Do test cases contain real customer data? If so, who approved that, and where do the logs go?
  • How do we add a case after a production incident, so the same mistake is caught next time?

Common misconceptions

  • "We tested it before go-live, so we're done." The model, the prompt and the documents keep changing. A harness repeats the test after every change.
  • "One overall score is enough." An average can rise while the most important cases fail. Look at slices and must-pass cases, not only the total.
  • "The harness tells us the system is good." It only tells you how the system does on the cases in the suite. A suite that misses a kind of order misses its failures too.
  • "A failed gate means the change is bad." Sometimes the suite or the scorer is wrong. The gate forces someone to look; a person decides.
  • "Raising the threshold improves quality." Thresholds only decide what gets blocked. Quality improves when the team fixes the failing cases.

Key terms

  • Evaluation harness: the machinery that runs a fixed set of test cases on every change, records the results and applies a gate.
  • Suite: the fixed set of test cases. Also called a golden set or test set.
  • Slice: a group of cases that share a trait, such as all credit-block orders. Scores per slice show where a change hurts.
  • Baseline: the last approved run that new runs are compared with.
  • Gate: the rules that decide whether a change may go live.
  • Must-pass case: a case that must be correct every time, whatever the averages say.
  • Regression: a case that used to pass and now fails.
  • Orchestration configuration: in SAP's generative AI hub, the saved setup of prompt template, model and modules that an application calls.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does an evaluation harness add to the one-off evaluation scripts a team already has?

    Answer: B. The harness doesn't improve the model or replace experts. It repeats the same cases after every change, records the result with the version, and blocks a change that makes things worse.
  2. 2A developer shortens the triage prompt to cut cost, and the demo looks fine. What is the main risk the harness guards against?

    Answer: C. Changes that look fine in a demo can break specific cases that nobody tried. The harness reruns every case, including orders with two blocks and large credit holds, and names the ones that got worse.
  3. 3The overall score went up after a change, but two large credit-hold orders now get wrong advice. What should the gate do?

    Answer: D. An average can rise while the most costly cases fail. Must-pass cases are the ones where a wrong answer costs money or trust, so any failure on them blocks the change, whatever the totals say.
  4. 4Who should decide the gate's thresholds?

    Answer: A. The thresholds say how much risk the business accepts before a change goes live. Engineers build the gate, but the product owner decides and signs off on what blocks a release.
  5. 5As of September 2026, what does SAP's Evaluations in the generative AI hub cover?

    Answer: C. SAP describes Evaluations as benchmarking prompts and models as orchestration configurations, with system-defined and custom judge metrics. Choosing the cases, setting the gate and keeping the history remain your team's job.
  6. 6After a production incident, a bad note reached an analyst. What is the best harness response?

    Answer: D. A suite only catches failures it contains. Turning each incident into a test case makes the harness stronger over time, while raising thresholds alone changes what is blocked, not what is tested.
Deep layer · 40 min read

Mental model

A harness is continuous integration for behavior. Unit tests ask "does the code still do what it did?" A harness asks "does the system still answer the way it did, on cases we chose on purpose?"

Every run produces one record. The record says what was tested and how it scored:

flowchart LR
  V[Versions<br/>prompt, model, suite, commit] --> R[Runner]
  S[(Suite<br/>cases with slices)] --> R
  R --> O[Outputs]
  O --> SC[Scorers<br/>rules, then judge]
  SC --> REC[(Run record)]
  B[(Baseline record)] --> G{Gate}
  REC --> G
  G -->|pass| SHIP[Change may go live]
  G -->|fail| FIX[Report: what got worse]

Two runs are only comparable if they used the same suite. Everything else in the record (prompt, model, code) is what you are testing. Keep that rule in mind; most harness bugs break it.

How it works

The suite: cases, slices and must-pass flags

The suite is a file of cases, each with an ID that never changes. A stable ID lets you say "case s04 passed last week and fails today". In this topic each case is a blocked sales order, as in Unit 1:

{"id": "s04", "slice": "credit", "must_pass": true, "order": {"SalesOrder": "9100004", "SoldToParty": "CUST-A", "TotalNetAmount": "51200.00", "DeliveryBlockReason": "01", "TotalCreditCheckStatus": "B", "...": "..."}}
  • Slice groups cases by what makes them different: credit, billing, delivery, several blocks, or no block. A change that breaks one slice can hide inside a good average.
  • Must-pass marks the cases where a wrong answer is costly. Here: large credit holds and orders with no block set, where the note must admit the data can't say why.

The suite file gets a fingerprint: a short hash of its contents. Any edit changes it. The harness refuses to compare runs with different fingerprints.

The runner: test exactly what ships

The runner sends each case through the system under test. Here that is a prompt plus a model. The runner must use the same prompt file and model name that production uses. A harness that tests a copy of the prompt tests the wrong thing.

This topic uses Inspect, which you installed in Set up for Unit 8. As before, a task is a dataset, a solver and scorers. Inspect's documentation lists epochs, name, version and metadata among the task settings, and the -T option on inspect eval overrides task parameters from the command line. Our harness calls Inspect from Python so it can add the record and gate around it.

Epochs and non-determinism. Model output varies between calls. Inspect can run each case several times (epochs) and combine the scores with a reducer. Its built-in reducers include mean, mode, pass_at_{k} and at_least_{k}. The harness uses at_least_{n} with n epochs, so a case passes only if it passes every time. That is the strict reading a release gate wants: an answer that is right two times out of three is a problem in production.

Scorers: rules first, a judge where rules can't decide

Each scorer answers one yes-or-no question about one output. Many narrow scorers beat one broad one. When a run fails, the scorer's name tells you what broke.

Scorer Question Why it is a rule, not a judge
valid_format Is the reply JSON with likely_cause, next_check and confidence? The format is exact
covers_facts Does the note mention every block that is set, or say the data doesn't show why? The fields are in the case
no_code_guessing Does the cause avoid "means", "indicates" or "stands for"? Unit 1's rule: codes differ by system
no_outside_orders Does the note cite only its own order number? Any other number is invented
judge_says_right (optional) Does a judge model rate the note "right"? Usefulness needs judgment

Inspect's scorers documentation shows the pattern: a function decorated with @scorer(metrics=[...]) that returns a Score with a value, an answer and an explanation. The explanation is what a person reads when a case fails, so write it for them.

Rules are cheap, exact and never drift. Their limit is that they only see what you wrote down. no_code_guessing misses "code 01 is a credit hold", and it would flag a harmless "this means you should check". The LLM evaluation fundamentals topic covers how to check a judge before you trust it.

Metrics per slice

Inspect's grouped() metric computes a metric separately for each value of a sample's metadata field, plus an all total. With grouped(accuracy(), "slice"), every scorer reports one number per slice. That turns "covers_facts dropped to 0.67" into "covers_facts is 0 on orders with several blocks".

The record

Inspect saves a log per run; its documentation lists status, eval, results and samples among the fields and read_eval_log() to read one back. Logs default to ./logs and to the compact .eval format. Logs are detailed and big. The harness also writes a small run record, one JSON line per run:

Field Why it is there
run_id When the run happened
prompt, prompt_sha Which prompt version, and a fingerprint of its exact text
writer, judge, epochs Which model wrote the notes, which judged them, how many times each case ran
suite_sha Fingerprint of the suite; runs are comparable only if it matches
git The code commit, with +changes if files were edited but not committed
metrics, slices Score per scorer, overall and per slice
cases Pass or fail per case and scorer, so you can see exactly what changed
log The Inspect log file with the full transcripts

The prompt fingerprint matters. Someone can edit v1.txt without renaming it. The name stays v1; the fingerprint changes.

The gate

compare reads the newest record and the baseline and applies four rules from eval_config.json:

  1. Same suite. If the fingerprints differ, the gate fails and tells you to re-baseline first.
  2. Minimums. Each scorer must reach its floor.
  3. Maximum drop. No scorer may fall more than max_drop below the baseline.
  4. Must-pass cases. Any failure on a must-pass case fails the gate.

It also lists newly failing and newly passing cases. A change can fix one thing and break another; averages hide that, case lists don't.

The program ends with exit code 1 when the gate fails. That single number is what lets GitHub Actions, or any other CI service, stop a change.

Build it yourself: a harness with a baseline and a gate

You will build a harness for the Unit 1 triage writer. It runs 12 made-up blocked orders, scores each note with four rules, saves a record, compares it with a baseline and blocks a worse prompt. Then you will run it on GitHub for every push.

Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder with its .venv, .env and .gitignore, and install Inspect. Step 8 uses the GitHub workflow pattern from Testing AI applications. This walkthrough doesn't repeat those steps.

flowchart LR
  I[Step 2-3<br/>init: suite, prompts, config] --> R1[Step 4<br/>run v1]
  R1 --> B[Step 5<br/>baseline + compare]
  B --> R2[Step 6<br/>run v2, gate fails]
  R2 --> E[Step 7<br/>epochs, real model]
  E --> CI[Step 8<br/>GitHub Actions]

What you need

  • Your course folder from the setup topics.
  • About 60 to 90 minutes.
  • No new accounts. Without a model, a made-up offline writer produces the notes, so every step is free.
  • Optional: your Anthropic key from Unit 1 in .env, for Step 7 (small per-request charge).
  • Optional: your course repository on GitHub from Git and GitHub for AI engineers, for Step 8.

Step 1: Open your course folder

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that Inspect is installed:

    inspect --version

    You should see a version number, such as 0.3.276. If you see "not recognized" or "command not found", go back to Step 2 of Set up for Unit 8.

Step 2: Create the harness script

  1. In VS Code, right-click the unit08 folder, choose New File, name it harness.py, paste the code below and save.
"""Unit 8: an evaluation harness for the blocked-order triage writer.

A harness is the machinery around your tests: a fixed suite of cases, a runner, scorers, a results
store and a gate. Every run is recorded with what was tested (prompt, model, suite, Git commit),
compared with a saved baseline, and either passes or fails the gate.

How to run (from your course folder, with .venv turned on):
    python unit08/harness.py init                        # write the suite, two prompt versions, the config
    python unit08/harness.py run                         # evaluate the prompt named in eval_config.json
    python unit08/harness.py run --prompt v2             # evaluate another prompt version
    python unit08/harness.py history                     # every recorded run, oldest first
    python unit08/harness.py baseline                    # make the latest run the baseline
    python unit08/harness.py compare                     # latest run vs. baseline; exit code 1 if the gate fails
    python unit08/harness.py run --model anthropic/claude-opus-5-5   # a real model writes the notes (small charge)
    python unit08/harness.py run --judge anthropic/claude-opus-5-5   # add a model judge as a fifth scorer (small charge)

Without --model, a made-up "offline writer" stands in for the model, so every step works with no account.
"""
import argparse
import hashlib
import json
import os
import re
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path

from dotenv import load_dotenv

HERE = Path(__file__).resolve().parent
SUITE = HERE / "data" / "triage_suite.jsonl"
PROMPTS = HERE / "prompts"
CONFIG = HERE / "eval_config.json"
RESULTS = HERE / "results"
RUNS = RESULTS / "runs.jsonl"
BASELINE = RESULTS / "baseline.json"
REPORT = RESULTS / "report.md"
LOGS = HERE / "logs" / "harness"

# ---------------------------------------------------------------- the suite (made-up, SAP-shaped)


def order(number, customer, amount, delivery_block="", billing_block="", credit_status=""):
    """Header fields shaped like Unit 1's sales order. The codes are placeholders, not SAP meanings."""
    return {"SalesOrder": number, "SoldToParty": customer, "TotalNetAmount": amount,
            "TransactionCurrency": "USD", "TotalCreditCheckStatus": credit_status,
            "DeliveryBlockReason": delivery_block, "HeaderBillingBlockReason": billing_block}


SAMPLE_SUITE = [
    {"id": "s01", "slice": "credit", "must_pass": True, "order": order("9100001", "CUST-A", "72400.00", "01", "", "B")},
    {"id": "s02", "slice": "credit", "must_pass": False, "order": order("9100002", "CUST-B", "8300.00", "01", "", "B")},
    {"id": "s03", "slice": "credit", "must_pass": False, "order": order("9100003", "CUST-C", "15900.00", "01", "", "B")},
    {"id": "s04", "slice": "credit", "must_pass": True, "order": order("9100004", "CUST-A", "51200.00", "01", "", "B")},
    {"id": "s05", "slice": "billing", "must_pass": False, "order": order("9100005", "CUST-D", "940.00", "", "02")},
    {"id": "s06", "slice": "billing", "must_pass": False, "order": order("9100006", "CUST-E", "2210.00", "", "04")},
    {"id": "s07", "slice": "delivery", "must_pass": False, "order": order("9100007", "CUST-F", "3150.00", "03")},
    {"id": "s08", "slice": "delivery", "must_pass": False, "order": order("9100008", "CUST-G", "610.00", "05")},
    {"id": "s09", "slice": "multi", "must_pass": False, "order": order("9100009", "CUST-H", "4480.00", "03", "02")},
    {"id": "s10", "slice": "multi", "must_pass": False, "order": order("9100010", "CUST-B", "12650.00", "01", "02", "B")},
    {"id": "s11", "slice": "no-block", "must_pass": True, "order": order("9100011", "CUST-I", "560.00")},
    {"id": "s12", "slice": "no-block", "must_pass": True, "order": order("9100012", "CUST-J", "7020.00")},
]

PROMPT_V1 = """You write short triage notes for credit and order-management analysts about blocked sales orders.
Use only the order fields given. Do not guess what block or status codes mean; codes differ by system.
Mention every block or status that is set. If no block is set, say the data does not show why the order is held.
Reply with JSON only: {"likely_cause": "...", "next_check": "...", "confidence": "low|medium|high"}"""

PROMPT_V2 = """You write very short triage notes for analysts about blocked sales orders.
Focus on the main problem only. Keep each field under 12 words.
Reply with JSON only: {"likely_cause": "...", "next_check": "...", "confidence": "low|medium|high"}"""

SAMPLE_CONFIG = {
    "prompt": "v1",
    "gate": {
        "min": {"valid_format": 1.0, "covers_facts": 0.8, "no_code_guessing": 0.9, "no_outside_orders": 0.9},
        "max_drop": 0.05,
        "must_pass_cases": True,
    },
}


def sha(text: str) -> str:
    return hashlib.sha256(text.encode("utf-8")).hexdigest()[:10]


def read_suite() -> list:
    if not SUITE.exists():
        sys.exit("No suite yet. Run: python unit08/harness.py init")
    cases = [json.loads(line) for line in SUITE.read_text(encoding="utf-8").splitlines() if line.strip()]
    ids = [c["id"] for c in cases]
    if len(ids) != len(set(ids)):
        sys.exit("Two cases in the suite share an id. Every case needs its own id.")
    return cases


def read_config() -> dict:
    if not CONFIG.exists():
        sys.exit("No eval_config.json yet. Run: python unit08/harness.py init")
    return json.loads(CONFIG.read_text(encoding="utf-8"))


def read_prompt(version: str) -> str:
    path = PROMPTS / f"{version}.txt"
    if not path.exists():
        sys.exit(f"No prompt file {path.name} in unit08/prompts. Run init, or create the file.")
    return path.read_text(encoding="utf-8").strip()


def git_commit() -> str:
    """The current commit, plus '+changes' if files are edited but not committed."""
    try:
        commit = subprocess.run(["git", "rev-parse", "--short", "HEAD"], capture_output=True, text=True,
                                check=True).stdout.strip()
        dirty = subprocess.run(["git", "status", "--porcelain", "--", str(HERE)], capture_output=True,
                               text=True).stdout.strip()
        return commit + ("+changes" if dirty else "")
    except (OSError, subprocess.CalledProcessError):
        return "no-git"


# ---------------------------------------------------------------- the offline writer (stands in for a model)


def offline_note(version: str, fields: dict) -> str:
    """Made-up behavior for each prompt version, so the harness has something to measure without a model.
    v1 follows its rules, with one flaw: on CUST-B's order with two blocks it cites another order.
    v2 ('shorter notes') keeps only the first block and, on large credit orders, guesses what code 01 means."""
    delivery, billing = fields["DeliveryBlockReason"], fields["HeaderBillingBlockReason"]
    credit, amount = fields["TotalCreditCheckStatus"], float(fields["TotalNetAmount"])
    if version == "v2":
        if delivery and credit and amount > 50000:
            return json.dumps({"likely_cause": "Delivery block 01 means the credit limit is exceeded.",
                               "next_check": "Release the order in credit management.", "confidence": "high"})
        if delivery:
            return json.dumps({"likely_cause": "A delivery block is set.",
                               "next_check": "Ask sales about the delivery block.", "confidence": "medium"})
        if billing:
            return json.dumps({"likely_cause": "A billing block is set.",
                               "next_check": "Ask billing about the block.", "confidence": "medium"})
        return json.dumps({"likely_cause": "No block is set; the data does not show why it is held.",
                           "next_check": "Ask who flagged the order.", "confidence": "low"})
    parts, checks = [], []
    if delivery:
        parts.append(f"a delivery block ({delivery})")
        checks.append("ask sales which delivery block this code is in your system")
    if billing:
        parts.append(f"a billing block ({billing})")
        checks.append("ask billing what the block code stands for in your system")
    if credit:
        parts.append(f"a credit check status ({credit})")
        checks.insert(0, "check the customer's credit exposure and open items")
    if not parts:
        return json.dumps({"likely_cause": "No block or credit status is set, so the data does not show why the order is held.",
                           "next_check": "Confirm with the analyst why the order was flagged.", "confidence": "low"})
    next_check = "; ".join(checks).capitalize() + "."
    if fields["SoldToParty"] == "CUST-B" and len(parts) == 3:
        next_check += " CUST-B also has order 9100002 blocked."  # the flaw: an order it was never given
    return json.dumps({"likely_cause": "The order has " + " and ".join(parts) + ".",
                       "next_check": next_check, "confidence": "medium"})


def offline_model(version: str):
    from inspect_ai.model import ModelOutput, ModelUsage, get_model

    def write(messages, tools, tool_choice, config):
        fields = json.loads(messages[-1].text.split("Order fields:", 1)[1])
        output = ModelOutput.from_content(model="offline-writer", content=offline_note(version, fields))
        output.usage = ModelUsage(input_tokens=0, output_tokens=0, total_tokens=0)
        return output

    return get_model("mockllm/model", custom_outputs=write, memoize=False)


# ---------------------------------------------------------------- scorers: code checks first, a judge if asked


def parse_note(text: str):
    """The note as a dict, or None. Tolerates a ```json fence around the reply."""
    text = re.sub(r"^```(?:json)?\s*|\s*```$", "", text.strip())
    try:
        note = json.loads(text)
    except json.JSONDecodeError:
        return None
    return note if isinstance(note, dict) else None


def build_scorers(judge: str | None):
    from inspect_ai.model import get_model
    from inspect_ai.scorer import CORRECT, INCORRECT, Score, accuracy, grouped, scorer, stderr

    metrics = [grouped(accuracy(), "slice"), stderr()]

    def verdict(ok: bool, why: str, answer: str = "") -> Score:
        return Score(value=CORRECT if ok else INCORRECT, answer=answer, explanation=why)

    @scorer(metrics=metrics)
    def valid_format():
        async def score(state, target):
            note = parse_note(state.output.completion)
            ok = bool(note) and {"likely_cause", "next_check", "confidence"} <= set(note) \
                and note.get("confidence") in ("low", "medium", "high")
            return verdict(ok, "JSON with the three fields" if ok else "not the agreed JSON shape",
                           state.output.completion[:200])
        return score

    @scorer(metrics=metrics)
    def covers_facts():
        async def score(state, target):
            note, fields = parse_note(state.output.completion), state.metadata["order"]
            text = json.dumps(note).lower() if note else ""
            missing = []
            if fields["DeliveryBlockReason"] and "delivery" not in text:
                missing.append("delivery block")
            if fields["HeaderBillingBlockReason"] and "billing" not in text:
                missing.append("billing block")
            if fields["TotalCreditCheckStatus"] and "credit" not in text:
                missing.append("credit status")
            no_blocks = not (fields["DeliveryBlockReason"] or fields["HeaderBillingBlockReason"]
                             or fields["TotalCreditCheckStatus"])
            if no_blocks and not re.search(r"does not show|doesn't show|not enough|insufficient", text):
                missing.append("saying the data does not show why")
            return verdict(not missing, "covers every block that is set" if not missing
                           else "misses: " + ", ".join(missing))
        return score

    @scorer(metrics=metrics)
    def no_code_guessing():
        async def score(state, target):
            note = parse_note(state.output.completion) or {}
            cause = str(note.get("likely_cause", "")).lower()
            guess = re.search(r"\b(means|indicates|stands for)\b", cause)
            return verdict(not guess, "explains no code" if not guess else f"guesses a code meaning: {cause}")
        return score

    @scorer(metrics=metrics)
    def no_outside_orders():
        async def score(state, target):
            own = state.metadata["order"]["SalesOrder"]
            others = sorted(set(re.findall(r"\b\d{7,10}\b", state.output.completion)) - {own})
            return verdict(not others, "cites only its own order" if not others
                           else "cites orders it was never given: " + ", ".join(others))
        return score

    chosen = [valid_format(), covers_facts(), no_code_guessing(), no_outside_orders()]
    if judge:
        rubric = ("You grade a triage note about a blocked sales order. A good note uses only the fields given, "
                  "guesses no code meanings, mentions every block that is set and suggests a sensible check. "
                  "Reply with one line: VERDICT: right, VERDICT: partly or VERDICT: wrong.")

        @scorer(metrics=metrics)
        def judge_says_right():
            async def score(state, target):
                reply = await get_model(judge).generate(
                    f"{rubric}\n\nOrder fields:\n{json.dumps(state.metadata['order'], indent=2)}"
                    f"\n\nNote:\n{state.output.completion}")
                found = re.search(r"VERDICT:\s*(right|partly|wrong)", reply.completion, re.I)
                label = found.group(1).lower() if found else "no verdict"
                return verdict(label == "right", f"judge said {label}", reply.completion[:200])
            return score

        chosen.append(judge_says_right())
    return chosen


# ---------------------------------------------------------------- commands


def cmd_init(args) -> None:
    files = {SUITE: "".join(json.dumps(c) + "\n" for c in SAMPLE_SUITE),
             PROMPTS / "v1.txt": PROMPT_V1 + "\n", PROMPTS / "v2.txt": PROMPT_V2 + "\n",
             CONFIG: json.dumps(SAMPLE_CONFIG, indent=2) + "\n"}
    for path, text in files.items():
        if path.exists() and not args.force:
            print(f"Kept     {path.relative_to(HERE.parent).as_posix()} (exists; add --force to overwrite)")
            continue
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(text, encoding="utf-8")
        print(f"Wrote    {path.relative_to(HERE.parent).as_posix()}")
    print(f"\nSuite: {len(read_suite())} cases. Next: python unit08/harness.py run")


def cmd_run(args) -> None:
    load_dotenv()
    for name in (args.model, args.judge):
        if name and name.startswith("anthropic/") and not os.getenv("ANTHROPIC_API_KEY"):
            sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or leave out --model/--judge.")
    try:
        from inspect_ai import Epochs, Task, eval
        from inspect_ai.dataset import Sample
        from inspect_ai.solver import generate, system_message
    except ImportError:
        sys.exit("inspect-ai is not installed. Run: pip install -r requirements.txt")

    config, cases = read_config(), read_suite()
    version = args.prompt or config["prompt"]
    prompt = read_prompt(version)
    samples = [Sample(id=c["id"], input="Order fields:\n" + json.dumps(c["order"], indent=2), target="",
                      metadata={"slice": c["slice"], "must_pass": c["must_pass"], "order": c["order"]})
               for c in cases]
    task = Task(dataset=samples, solver=[system_message(prompt), generate()],
                scorer=build_scorers(args.judge), name="triage_harness",
                epochs=Epochs(args.epochs, f"at_least_{args.epochs}"))  # a case passes only if it passes every time
    model = args.model or offline_model(version)
    print(f"Running {len(samples)} cases x {args.epochs} epoch(s): prompt {version}, "
          f"writer {args.model or 'offline (made-up)'}{', judge ' + args.judge if args.judge else ''}...")
    log = eval(task, model=model, log_dir=str(LOGS), display="none")[0]
    if log.status != "success":
        sys.exit(f"The run did not finish ({log.status}): {log.error.message if log.error else ''}\n"
                 "Open the log with: inspect view --log-dir unit08/logs/harness")

    per_case = {}  # case id -> {scorer: 1.0 pass or 0.0 fail}, after combining the epochs
    for reduction in log.reductions:
        for s in reduction.samples:
            per_case.setdefault(str(s.sample_id), {})[reduction.scorer] = float(s.value)
    metrics, slices = {}, {}
    for result in log.results.scores:
        values = {k: round(v.value, 3) for k, v in result.metrics.items() if k != "stderr"}
        metrics[result.name] = values.pop("all", None)
        slices[result.name] = values
    record = {
        "run_id": datetime.now(timezone.utc).strftime("%Y%m%d-%H%M%S"),
        "prompt": version, "prompt_sha": sha(prompt),
        "writer": args.model or "offline", "judge": args.judge or "", "epochs": args.epochs,
        "suite_sha": sha(SUITE.read_text(encoding="utf-8")), "cases_n": len(cases),
        "git": git_commit(), "metrics": metrics, "slices": slices, "cases": per_case,
        "must_pass": [c["id"] for c in cases if c["must_pass"]],
        "log": Path(log.location).name,
    }
    RESULTS.mkdir(parents=True, exist_ok=True)
    with RUNS.open("a", encoding="utf-8") as handle:
        handle.write(json.dumps(record) + "\n")

    print(f"\nRun {record['run_id']}  prompt {version} ({record['prompt_sha']})  suite {record['suite_sha']}  git {record['git']}")
    print(f"{'scorer':18} {'all':>5}  " + "  ".join(f"{s:>8}" for s in sorted(next(iter(slices.values())))))
    for name, value in metrics.items():
        print(f"{name:18} {value:5.2f}  " + "  ".join(f"{slices[name][s]:8.2f}" for s in sorted(slices[name])))
    failed = [f"{cid} ({', '.join(n for n, v in sc.items() if v < 1)})"
              for cid, sc in per_case.items() if any(v < 1 for v in sc.values())]
    print("\nFailing cases: " + ("; ".join(failed) if failed else "none"))
    print("Saved to unit08/results/runs.jsonl. Next: python unit08/harness.py compare")


def read_runs() -> list:
    if not RUNS.exists():
        sys.exit("No runs recorded yet. Run: python unit08/harness.py run")
    return [json.loads(line) for line in RUNS.read_text(encoding="utf-8").splitlines() if line.strip()]


def cmd_history(args) -> None:
    runs = read_runs()
    names = list(dict.fromkeys(name for r in runs for name in r["metrics"]))  # every scorer seen, in order
    print(f"{'run':16} {'prompt':7} {'writer':10} {'suite':11} {'git':16} " + " ".join(f"{n[:12]:>12}" for n in names))
    for r in runs:
        cells = [f"{r['metrics'][n]:12.2f}" if n in r["metrics"] else f"{'-':>12}" for n in names]
        print(f"{r['run_id']:16} {r['prompt']:7} {r['writer'][:10]:10} {r['suite_sha']:11} {r['git'][:16]:16} "
              + " ".join(cells))


def cmd_baseline(args) -> None:
    runs = read_runs()
    chosen = next((r for r in runs if r["run_id"] == args.run), None) if args.run else runs[-1]
    if not chosen:
        sys.exit(f"No run with id {args.run}. See: python unit08/harness.py history")
    BASELINE.write_text(json.dumps(chosen, indent=2) + "\n", encoding="utf-8")
    print(f"Baseline is now run {chosen['run_id']} (prompt {chosen['prompt']}). Commit unit08/results/baseline.json.")


def gate(run: dict, base: dict, rules: dict) -> list:
    """Return the reasons the run fails the gate; an empty list means it passes."""
    problems = []
    if run["suite_sha"] != base["suite_sha"]:
        problems.append("the suite changed since the baseline, so the numbers aren't comparable: "
                        "run the baseline prompt on the new suite and save it as the baseline first")
    for name, minimum in rules["min"].items():
        value = run["metrics"].get(name)
        if value is None:
            problems.append(f"{name}: not measured in this run")
        elif value < minimum:
            problems.append(f"{name}: {value:.2f} is below the minimum {minimum:.2f}")
    for name, before in base["metrics"].items():
        after = run["metrics"].get(name)
        if after is not None and before - after > rules["max_drop"]:
            problems.append(f"{name}: dropped from {before:.2f} to {after:.2f} (allowed drop {rules['max_drop']:.2f})")
    if rules.get("must_pass_cases"):
        for cid in run["must_pass"]:
            bad = [n for n, v in run["cases"].get(cid, {}).items() if v < 1]
            if bad:
                problems.append(f"must-pass case {cid} failed: {', '.join(bad)}")
    return problems


def cmd_compare(args) -> None:
    if not BASELINE.exists():
        sys.exit("No baseline yet. Run the current prompt, then: python unit08/harness.py baseline")
    run, base = read_runs()[-1], json.loads(BASELINE.read_text(encoding="utf-8"))
    rules = read_config()["gate"]
    lines = [f"# Evaluation report: run {run['run_id']} vs. baseline {base['run_id']}", "",
             f"Prompt {run['prompt']} ({run['prompt_sha']}) vs. {base['prompt']} ({base['prompt_sha']}); "
             f"writer {run['writer']}; suite {run['suite_sha']}; git {run['git']}.", "",
             "| Scorer | Baseline | This run | Change |", "| --- | --- | --- | --- |"]
    for name in run["metrics"]:
        before, after = base["metrics"].get(name), run["metrics"][name]
        change = f"{after - before:+.2f}" if before is not None else "new"
        shown = f"{before:.2f}" if before is not None else "-"
        lines.append(f"| {name} | {shown} | {after:.2f} | {change} |")
    newly, fixed = [], []
    for cid, scores in run["cases"].items():
        for name, value in scores.items():
            was = base["cases"].get(cid, {}).get(name)
            if was == 1 and value < 1:
                newly.append(f"{cid} {name}")
            elif was is not None and was < 1 and value == 1:
                fixed.append(f"{cid} {name}")
    lines += ["", "Newly failing: " + (", ".join(newly) or "none"), "",
              "Newly passing: " + (", ".join(fixed) or "none")]
    problems = gate(run, base, rules)
    lines += ["", "## Gate: " + ("FAIL" if problems else "PASS"), ""] + [f"- {p}" for p in problems]
    text = "\n".join(lines) + "\n"
    REPORT.write_text(text, encoding="utf-8")
    print(re.sub(r"(?m)^#+ ", "", text))
    print("Report saved to unit08/results/report.md")
    sys.exit(1 if problems else 0)


def main() -> None:
    parser = argparse.ArgumentParser(description="Evaluation harness for the triage writer.")
    sub = parser.add_subparsers(dest="command", required=True)
    p = sub.add_parser("init", help="write the suite, prompts and config")
    p.add_argument("--force", action="store_true", help="overwrite existing files")
    p = sub.add_parser("run", help="run the suite and record the results")
    p.add_argument("--prompt", help="prompt version in unit08/prompts, e.g. v2 (default: eval_config.json)")
    p.add_argument("--model", help="a real model writes the notes, e.g. anthropic/claude-opus-5-5")
    p.add_argument("--judge", help="add a model judge as an extra scorer, e.g. anthropic/claude-opus-5-5")
    p.add_argument("--epochs", type=int, default=1, help="run each case this many times; it must pass every time")
    sub.add_parser("history", help="list recorded runs")
    p = sub.add_parser("baseline", help="save a run as the baseline")
    p.add_argument("--run", help="run id (default: the latest)")
    sub.add_parser("compare", help="compare the latest run with the baseline and apply the gate")
    args = parser.parse_args()
    {"init": cmd_init, "run": cmd_run, "history": cmd_history,
     "baseline": cmd_baseline, "compare": cmd_compare}[args.command](args)


if __name__ == "__main__":
    main()

What each part of the script does:

Part What it does
SAMPLE_SUITE, order Twelve made-up blocked orders in five slices; four are must-pass
PROMPT_V1, PROMPT_V2 Two versions of the triage prompt; v2 asks for very short notes
SAMPLE_CONFIG The current prompt and the gate rules
sha, git_commit Fingerprints for the prompt and suite; the Git commit, marked +changes if unit08 has uncommitted edits
offline_note, offline_model The made-up writer, wrapped as Inspect's mockllm/model so Inspect treats it like any model
parse_note Reads the JSON note, allowing the code fence some models put around JSON
build_scorers Four rule scorers, plus judge_says_right when you pass --judge; each reports accuracy per slice with grouped
cmd_run Builds the Inspect task, runs it, reads per-case results from log.reductions and appends a record to runs.jsonl
Epochs(n, "at_least_n") Runs each case n times; it passes only if it passes every time
gate The four gate rules; returns the list of reasons to fail
cmd_compare Writes report.md with the score table, newly failing and passing cases and the gate result; exit code 1 on fail
cmd_history, cmd_baseline List runs; save one run as the baseline

Step 3: Write the suite, prompts and config

  1. Run:

    python unit08/harness.py init

What success looks like:

Wrote    unit08/data/triage_suite.jsonl
Wrote    unit08/prompts/v1.txt
Wrote    unit08/prompts/v2.txt
Wrote    unit08/eval_config.json

Suite: 12 cases. Next: python unit08/harness.py run
  1. Open unit08/eval_config.json. It says which prompt is current and holds the gate:

    {
      "prompt": "v1",
      "gate": {
        "min": {"valid_format": 1.0, "covers_facts": 0.8, "no_code_guessing": 0.9, "no_outside_orders": 0.9},
        "max_drop": 0.05,
        "must_pass_cases": true
      }
    }

    In a real project, the product owner signs off on these numbers. Changing them is a change like any other: it goes through Git and review.

  2. Open unit08/prompts/v1.txt. This file is what the harness sends as the system message. In a real project, your application reads the same file, so the harness tests what ships.

  3. Tell Git to ignore the run history and report, which are rebuilt on every computer. Open .gitignore, add these lines at the end and save. (unit08/logs/ is already there from the setup topic.)

    unit08/results/runs.jsonl
    unit08/results/report.md
  4. Commit the suite, prompts and config, so the record shows a clean commit:

    git add .gitignore unit08/harness.py unit08/data/triage_suite.jsonl unit08/prompts unit08/eval_config.json
    git commit -m "Add the Unit 8 evaluation harness"

Step 4: Run the current prompt

  1. Run:

    python unit08/harness.py run

What success looks like (from our test; your run ID and commit differ):

Running 12 cases x 1 epoch(s): prompt v1, writer offline (made-up)...

Run 20261005-063844  prompt v1 (d6b83c2afe)  suite 00bd24a695  git eea1c41
scorer               all   billing    credit  delivery     multi  no-block
valid_format        1.00      1.00      1.00      1.00      1.00      1.00
covers_facts        1.00      1.00      1.00      1.00      1.00      1.00
no_code_guessing    1.00      1.00      1.00      1.00      1.00      1.00
no_outside_orders   0.92      1.00      1.00      1.00      0.50      1.00

Failing cases: s10 (no_outside_orders)
Saved to unit08/results/runs.jsonl. Next: python unit08/harness.py compare
  1. Read it row by row. Version 1 passes almost everything. One case fails: on s10, the note cites another order the writer was never given. That is the same kind of error as t5 in the setup topic's golden set. The multi slice shows it: 0.50, one of two cases.

  2. Look at the full transcripts in Inspect's viewer:

    inspect view --log-dir unit08/logs/harness

    Open the address the terminal prints (for example http://127.0.0.1:7575). Click the run, then sample s10, and look at its scores. The explanation for no_outside_orders reads cites orders it was never given: 9100002. Press Ctrl+C in the terminal to stop the viewer.

Step 5: Save a baseline and compare

  1. Make this run the baseline:

    python unit08/harness.py baseline
    Baseline is now run 20261005-063844 (prompt v1). Commit unit08/results/baseline.json.
  2. Compare the latest run with it:

    python unit08/harness.py compare

What success looks like (trimmed):

| Scorer | Baseline | This run | Change |
| --- | --- | --- | --- |
| valid_format | 1.00 | 1.00 | +0.00 |
| covers_facts | 1.00 | 1.00 | +0.00 |
| no_code_guessing | 1.00 | 1.00 | +0.00 |
| no_outside_orders | 0.92 | 0.92 | +0.00 |

Newly failing: none

Newly passing: none

Gate: PASS

Report saved to unit08/results/report.md

The run is compared with itself, so it passes. That proves the baseline passes its own gate. If it didn't, your thresholds would block every change.

  1. Commit the baseline. It is the reference every future run, on your laptop or on GitHub, is compared with:

    git add unit08/results/baseline.json
    git commit -m "Baseline: triage prompt v1"

Step 6: Try the shorter prompt and watch the gate block it

The team wants shorter notes and wrote v2. Test it before it goes live.

  1. Run version 2:

    python unit08/harness.py run --prompt v2

What success looks like:

Running 12 cases x 1 epoch(s): prompt v2, writer offline (made-up)...

Run 20261005-063848  prompt v2 (f0e7076113)  suite 00bd24a695  git d0657c0
scorer               all   billing    credit  delivery     multi  no-block
valid_format        1.00      1.00      1.00      1.00      1.00      1.00
covers_facts        0.67      1.00      0.50      1.00      0.00      1.00
no_code_guessing    0.83      1.00      0.50      1.00      1.00      1.00
no_outside_orders   1.00      1.00      1.00      1.00      1.00      1.00

Failing cases: s01 (no_code_guessing); s02 (covers_facts); s03 (covers_facts); s04 (no_code_guessing); s09 (covers_facts); s10 (covers_facts)
Saved to unit08/results/runs.jsonl. Next: python unit08/harness.py compare
  1. Compare:

    • Windows (PowerShell):

      python unit08/harness.py compare; $LASTEXITCODE
    • macOS / Linux:

      python unit08/harness.py compare; echo $?

What success looks like (trimmed):

Newly failing: s01 no_code_guessing, s02 covers_facts, s03 covers_facts, s04 no_code_guessing, s09 covers_facts, s10 covers_facts

Newly passing: s10 no_outside_orders

Gate: FAIL

- covers_facts: 0.67 is below the minimum 0.80
- no_code_guessing: 0.83 is below the minimum 0.90
- covers_facts: dropped from 1.00 to 0.67 (allowed drop 0.05)
- no_code_guessing: dropped from 1.00 to 0.83 (allowed drop 0.05)
- must-pass case s01 failed: no_code_guessing
- must-pass case s04 failed: no_code_guessing

Report saved to unit08/results/report.md
1
  1. Read what the harness found:

    • On credit orders with a delivery block, v2 keeps only the delivery block and drops the credit status (s02, s03). On orders with two blocks it drops the second one (s09, s10).
    • On the two large credit orders, both must-pass, it says what code 01 means and tells the analyst to release the order (s01, s04). Open the viewer to read those notes.
    • It also fixed something: s10 no longer cites an outside order. Shorter notes have less room to invent. The harness shows both sides, so the team can make an informed trade.
    • The last line, 1, is the exit code. CI reads it as "failed".
  2. Look at all runs side by side:

    python unit08/harness.py history
    run              prompt  writer     suite       git              valid_format covers_facts no_code_gues no_outside_o
    20261005-063844  v1      offline    00bd24a695  eea1c41                  1.00         1.00         1.00         0.92
    20261005-063848  v2      offline    00bd24a695  d0657c0                  1.00         0.67         0.83         1.00

    Every row says which prompt, writer, suite and commit produced the scores. That is the audit trail.

Step 7: Run each case several times, and try a real model

  1. Run every case three times. A case passes only if it passes all three:

    python unit08/harness.py run --epochs 3

    The offline writer always writes the same note, so the scores match Step 4. With a real model they often don't, and that is the point: a note that is right two times out of three fails here.

  2. Optional, small charge. With ANTHROPIC_API_KEY in .env from Unit 1, let a real model write the notes from the same prompt:

    python unit08/harness.py run --model anthropic/claude-opus-5-5

    This sends 12 requests. You can use any model your key supports; write its name after anthropic/. Then run compare. The baseline was the offline writer, so expect differences. To make a real model your baseline, run it, check the notes in the viewer, then run baseline.

  3. Optional, small charge. Add a model judge as a fifth scorer:

    python unit08/harness.py run --judge anthropic/claude-opus-5-5

    This sends 12 more requests. The judge's verdicts appear as judge_says_right. Trust it only after you have checked its agreement with your own verdicts, as in LLM evaluation fundamentals.

Step 8: Run the gate on GitHub for every push

  1. In unit08, create requirements-eval.txt with these two lines and save. The CI job installs only what the harness needs, which is faster than the whole course list:

    inspect-ai
    python-dotenv
  2. In .github/workflows (created in Unit 6), create unit08-eval.yml:

# Runs the Unit 8 evaluation harness on GitHub's computers and blocks regressions.
# Uses the offline writer, so it needs no key and costs nothing.
name: unit08-eval

on:
  push:
    paths: ["unit08/**", ".github/workflows/unit08-eval.yml"]
  pull_request:
    paths: ["unit08/**"]
  workflow_dispatch:       # adds a "Run workflow" button on the Actions tab

jobs:
  eval-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - name: Install the harness libraries
        run: |
          python -m pip install --upgrade pip
          pip install -r unit08/requirements-eval.txt
      - name: Run the suite with the current prompt
        run: python unit08/harness.py run
      - name: Compare with the baseline and apply the gate
        run: python unit08/harness.py compare
      - name: Keep the report
        if: ${{ always() }}
        uses: actions/upload-artifact@v4
        with:
          name: eval-report
          path: unit08/results/report.md

paths limits the job to changes in unit08. if: ${{ always() }} keeps the report even when the gate fails, which is when you need it most.

  1. Commit and push on a branch:

    git checkout -b unit08-harness
    git add unit08/requirements-eval.txt .github/workflows/unit08-eval.yml
    git commit -m "Run the evaluation gate in CI"
    git push -u origin unit08-harness
  2. On GitHub, click the Actions tab, then the run named unit08-eval, then eval-gate.

What success looks like: a green check next to eval-gate, and in the step Compare with the baseline and apply the gate the line Gate: PASS. Under Artifacts on the run's summary page, eval-report holds report.md.

  1. Now make the risky change the way a teammate would. Open unit08/eval_config.json, change "prompt": "v1" to "prompt": "v2", save, then:

    git commit -am "Switch triage to the shorter prompt"
    git push

    The new run shows a red cross on Compare with the baseline and apply the gate, with the six reasons from Step 6. Open a pull request from unit08-harness to main and the failed check appears there too.

  2. Change "prompt" back to "v1", commit and push. The check turns green again.

If something goes wrong

What you see What it means What to do
python or inspect is "not recognized" / "command not found" Python isn't on the path, or .venv isn't on Turn on .venv (Step 1); see Set up your computer
inspect-ai is not installed or ModuleNotFoundError: No module named 'dotenv' A library is missing in this environment pip install -r requirements.txt with .venv on
No suite yet or No eval_config.json yet init wasn't run, or you're not in the course folder Run from orchestrate-course: python unit08/harness.py init
No baseline yet compare has nothing to compare with Run the current prompt, then python unit08/harness.py baseline
the suite changed since the baseline You edited the suite after saving the baseline Run the current prompt on the new suite, check it, run baseline, commit baseline.json
ANTHROPIC_API_KEY is not set in .env You asked for a real model without a key Add the key to .env (Unit 1), or leave out --model and --judge
The run did not finish (error) with 401 or authentication The key is wrong or revoked Create a new key in the Anthropic Console and update .env
The run did not finish (error) with ProxyError, ConnectError or a timeout The network or a company proxy blocks the model provider Try another network, or ask IT to allow the provider's API; the offline path still works
git shows as no-git in the record The folder isn't a Git repository, or Git isn't installed Records still work; see Git and GitHub for AI engineers
The CI job fails at Compare with No baseline yet baseline.json wasn't committed git add unit08/results/baseline.json, commit, push
The CI job doesn't start The push didn't touch unit08 or the workflow file Use Run workflow on the Actions tab

The SAP way

On SAP AI Core, the generative AI hub's Evaluations runs the evaluation step for configurations you deploy. As of the SAP AI Core service guide of 4 September 2026:

  • Evaluations benchmarks models and prompts as orchestration configurations and was added under Optimizations on 8 December 2025.
  • It offers system-defined metrics and custom LLM-as-a-judge metrics with rating criteria.
  • Prompt optimization, in the same area, can take separate test and train datasets.
  • It requires the extended service plan, which includes the generative AI hub.

The SAP Cloud SDK for AI for Python added an Evaluations client in version 6.5.0. Its gen_ai_hub.evaluations.helpers module shows the moving parts:

Harness part SAP side, as the SDK helpers show it
Suite file Uploaded to an object store; helpers read and write CSV, JSON and JSONL
Suite registered for runs upload_dataset_data_and_register_aicore_artifact registers the data as an SAP AI Core artifact
One run, or several configurations single_evaluation_job_flow and multiple_evaluation_jobs_flow
Reading results S3FileClient.get_sqlitedb_tables_data_from_s3 downloads a SQLite database from the object store and loads named tables

How the two fit together:

  1. Keep the suite in Git. The JSONL file stays the source of truth with its fingerprint. Upload a copy for SAP runs.
  2. Test the deployed configuration. On SAP, the thing under test is the orchestration configuration, not a prompt file on your laptop. Record its name and version in the run record where this harness records prompt and prompt_sha.
  3. Map metrics to scorers. Format and fact checks stay as rules. SAP's custom judge metrics play the role of judge_says_right.
  4. Gate in your pipeline. Read the SAP results, write the same run record, and call the same gate function. The gate rules belong to your release process, not to the evaluation service.

SAP Learning's lesson on evaluating prompts with the SDK uses the same core idea at small scale: a fixed test set of 20 customer emails, code checks per answer (valid JSON, correct category, sentiment and urgency), and an average per prompt and model. A harness adds the record, the baseline and the gate around that loop.

Build vs. SAP

Situation Use Why
Learning, prototypes, made-up data This harness on a laptop and GitHub Actions Free, every score readable
Production on SAP AI Core orchestration SAP Evaluations for the runs, your pipeline for the gate Tests the configuration you deploy, close to SAP data
Format and field checks Code rules, anywhere Exact, free, no drift
Free-text quality A judge: SAP custom metric or Inspect scorer Scales; calibrate against experts first
Model retirement or upgrade Run the suite on old and new model, compare per case Find regressions before the switch date
Many tasks or models at once Inspect eval_set with its own log directory Retries and resumes unfinished work
Audit trail of what was live Run records with prompt and configuration fingerprints, kept with releases Answers "what ran on that date?"

Production concerns

  • Data protection. Suites built from real SAP orders contain business data. Keep only the fields a case needs, get approval, and give logs and records the same access rules as the source system. Inspect logs hold every prompt and reply.
  • Authorizations. A suite exported from SAP no longer carries SAP authorization checks. Restrict who can read it, and record where each case came from.
  • Test what ships. The harness must read the same prompt file, or call the same orchestration configuration, as production. Fingerprint it in every record.
  • Suite changes. Adding cases is healthy, but it breaks comparison. Re-baseline in the same change, so reviewers see both.
  • Gate ownership and overrides. Write down who may change thresholds or override a failed gate, and record each override with a reason.
  • Cost. Every run with a real model costs one call per case per epoch, and a judge doubles it. Run the cheap rule checks on every push and the paid run on a schedule or before release.
  • Non-determinism. Use several epochs for must-pass cases, and gate on "passes every time".
  • Judge and model versions. Pin model versions in the record. When a provider retires a model, run the suite on the replacement and compare before switching.
  • Clean core. The harness runs outside S/4HANA, on copies or on API reads. It needs no custom code in the ERP.

Pitfalls

  • Comparing runs on different suites. Scores move because the questions changed, not the system. The fingerprint check exists for this.
  • Testing a copy of the prompt. The harness passes, production differs.
  • Only averages. A slice or a must-pass case can fail inside a good total.
  • A baseline that fails its own gate. Every change is then blocked, and people learn to ignore the gate.
  • Thresholds set by whoever writes the code. The gate encodes business risk; the business should own it.
  • Never adding cases. The suite drifts away from what users actually ask. Add a case for every incident.
  • Trusting one scorer. A rule sees only what you wrote down; a judge has its own errors. Use both, and read failing transcripts.
  • Committing logs. They are large and may contain sensitive data. Commit the baseline record, not the logs.

Exercise

Add a fifth rule, no_release_advice, that fails a note telling the analyst to release, remove or unblock anything. Unit 1's rule is that the assistant suggests checks and a person acts. Then grow the suite and re-baseline correctly.

  1. Open unit08/harness.py. Find the line that starts with chosen = [valid_format(). Just above it, at the same indentation, paste:

        @scorer(metrics=metrics)
        def no_release_advice():
            async def score(state, target):
                note = parse_note(state.output.completion) or {}
                advice = str(note.get("next_check", "")).lower()
                acts = re.search(r"\b(release|remove|unblock)\b", advice)
                return verdict(not acts, "suggests a check, not an action" if not acts
                               else f"tells the analyst to act: {advice}")
            return score
  2. Change the chosen line to add it:

        chosen = [valid_format(), covers_facts(), no_code_guessing(), no_outside_orders(), no_release_advice()]

    Save the file.

  3. Open unit08/eval_config.json. Inside "min", add "no_release_advice": 1.0 (put a comma after the entry before it). Save.

  4. Add two cases at the end of unit08/data/triage_suite.jsonl, one line each, and save:

    {"id": "s13", "slice": "credit", "must_pass": true, "order": {"SalesOrder": "9100013", "SoldToParty": "CUST-K", "TotalNetAmount": "88000.00", "TransactionCurrency": "USD", "TotalCreditCheckStatus": "B", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""}}
    {"id": "s14", "slice": "billing", "must_pass": false, "order": {"SalesOrder": "9100014", "SoldToParty": "CUST-L", "TotalNetAmount": "1320.00", "TransactionCurrency": "USD", "TotalCreditCheckStatus": "", "DeliveryBlockReason": "", "HeaderBillingBlockReason": "02"}}
  5. Run the current prompt and compare:

    python unit08/harness.py run
    python unit08/harness.py compare

    The gate fails with the suite changed since the baseline. That is correct: 14 cases can't be compared with 12.

  6. The run you just made is v1 on the new suite. Check its table (no_release_advice should be 1.00), then make it the baseline:

    python unit08/harness.py baseline
    python unit08/harness.py compare

    The gate now passes.

  7. Run v2 and compare:

    python unit08/harness.py run --prompt v2
    python unit08/harness.py compare

    The gate fails, and three must-pass cases (s01, s04, s13) now fail no_release_advice as well as no_code_guessing.

  8. Open unit08/results/report.md and add three sentences at the end: which slice v2 hurts most, which must-pass failure worries you most and why, and one case you would add next.

  9. Save your work:

    git add unit08/harness.py unit08/eval_config.json unit08/data/triage_suite.jsonl unit08/results/baseline.json
    git commit -m "Add no_release_advice scorer and two cases; re-baseline"

Done when compare on v1 prints Gate: PASS against a 14-case baseline, compare on v2 lists must-pass case s13 failed: no_code_guessing, no_release_advice, and report.md holds your three sentences. Keep the harness: you can reuse it to gate the agent changes you build in Unit 9.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the harness refuse to compare a run with the baseline when the suite fingerprints differ?

    Answer: C. A run record is only comparable with another one on the same cases. If the suite changed, re-run the current prompt on the new suite and save that as the baseline, so the next comparison measures the system again.
  2. 2The harness runs each case with Epochs(3, "at_least_3"). What does that mean for a case?

    Answer: B. at_least_3 over three epochs returns a pass only when all three runs pass. Model output varies between calls, and a release gate should treat an answer that is right two times out of three as a failure.
  3. 3Why does the run record store prompt_sha as well as the prompt version name?

    Answer: D. Names are labels people choose, and the text behind a label can change. The fingerprint is computed from the exact text, so two runs with the same name and different hashes tested different prompts.
  4. 4In the v2 run, no_outside_orders improved while covers_facts and no_code_guessing dropped. Which part of the report makes that trade visible?

    Answer: B. A change can fix some cases and break others, and averages blur that. The case lists name s10 as newly passing and six cases as newly failing, so the team can judge the trade case by case.
  5. 5How does GitHub Actions know the gate failed?

    Answer: C. CI judges each step by its exit code: 0 passes, anything else fails. compare calls sys.exit(1) when the gate has reasons to fail, which marks the step red and blocks the pull request.
  6. 6Your team moves the triage assistant to SAP AI Core orchestration. What should change in the harness?

    Answer: D. On SAP, the orchestration configuration is what ships, so it is what must be tested and fingerprinted. SAP Evaluations runs and scores it, but the thresholds and release decision stay in your pipeline, and exact format checks stay as rules.
  7. 7A provider announces that the model behind your triage writer retires next quarter. What do you do with the harness?

    Answer: B. The harness exists to find regressions before they reach users. Running the replacement model on the same suite shows which cases break while the old model is still live, so the team can fix the prompt or choose another model in time.
  8. 8Which file belongs in Git, and which should stay out?

    Answer: C. The baseline is the reference that every run, on a laptop or in CI, is compared with, so it must be versioned. Inspect logs are large, rebuilt by each run, and hold every prompt and reply, which may include sensitive data.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in