Orchestrate

Defining success metrics with the business

Agree one business metric with a baseline and target, guardrails that must not get worse, and a fair pilot comparison, so an AI result can be proven.

Updated Oct 5, 2026Foundational 9 minDeep 35 min
Foundational layer · 9 min read

The 60-second version

Every AI project promises something: faster decisions, fewer errors, lower cost. A success metric turns that promise into a number both sides agree on before the work starts.

A good agreement has four parts. One primary metric the business cares about, such as hours to decide a blocked order. A baseline: today's value, measured from real data. A target: the value you promise, by a date. And guardrails: things that must not get worse on the way, such as bad credit decisions.

Unit 8 so far measured whether the AI's answers are right. This topic measures whether the business is better off. Both matter. A model can score well and still change nothing, if nobody uses it or the bottleneck is elsewhere.

Why it matters to the business

Most AI pilots end with a demo, a happy sponsor and no proof. Six months later, when budget is renewed, nobody can say what changed. The project gets cut, or worse, it keeps running at a cost nobody can justify.

Take the blocked sales order assistant used throughout this course. Credit analysts get a drafted note for each order on credit hold. The pitch is "analysts decide faster". Without an agreed metric, three problems appear:

  • No baseline. Nobody measured how long decisions took before. Any number after go-live has nothing to compare with.
  • The wrong comparison. Decisions got faster in the pilot month. But order volume also dropped that month, so everyone was faster, with or without the assistant.
  • A hidden cost. Analysts decide faster because they release more orders without checking. Overdue receivables rise a quarter later.

Agreeing the metrics first fixes all three. It costs a workshop and a few days of data work. It also forces the most useful question of the whole project: what decision will we make based on this number?

How SAP does it

SAP offers several places where process numbers already live. Use them for baselines before you build anything. As of October 2026:

  • SAP S/4HANA, SAP Smart Business. SAP Learning describes the Sales Order Fulfillment Issues app for analyzing and processing blocked sales orders. It covers issues such as billing blocks, delivery blocks and incomplete data. Smart Business shows KPIs with colors based on defined targets and thresholds. The Manage KPIs and Reports app, available since S/4HANA 1909, lets an analytics specialist configure new KPIs.
  • SAP Signavio Process Insights. SAP Learning lists more than 200 process performance indicators, with comparison inside your company and against industry peers. Standard indicator types include throughput, backlog, exceptions and automation rate.
  • SAP Signavio Process Intelligence. It holds a library of metrics, predefined calculations that are reused to compute KPIs. Every process gets an average cycle time metric by default.
  • SAP Business AI Catalog in SAP Discovery Center. SAP News reported almost 400 features and agents in June 2026, plus cost and ROI estimators. In August 2025 SAP said its value estimates rest on SAP's own expertise, benchmarking, third-party research and early customer feedback.

Treat SAP's value figures as a reason to look, not as your baseline. At Sapphire 2026, SAP quoted a customer cutting wholesale order processing from days to minutes. That is their process and their starting point. Your number has to come from your own system.

The metric agreement, on one page

Write this down and get the sponsor to sign it before the pilot starts.

Part What it says Blocked-order example
Decision What the numbers will decide, and who decides Roll out to all credit analysts, or stop. Head of Credit Management decides.
Primary metric One business number, defined exactly Median hours from credit block to release or rejection
Baseline Today's value, from system data, with the period 19.7 hours over the last 8 weeks
Target The promised change and the date 25% faster in a 4-week pilot
Guardrails What must not get worse, with a limit Released orders overdue after 30 days: no more than 1 point higher
Supporting metrics Numbers that explain the result Notes opened, notes rated useful, cost per order
Comparison What the pilot is compared with Analysts without the assistant in the same weeks
Owner and source Who owns each number, where it comes from Credit team for decisions, finance for overdue items, IT for cost

Three layers of numbers sit behind this. Model quality (is the note right?) is what earlier Unit 8 topics measured. Adoption (do analysts open the note?) tells you whether people use it. Business outcome (do decisions get faster, safely?) is what the sponsor pays for. If the outcome doesn't move, the lower layers tell you why.

Questions to ask

  • What decision will this number drive, and who makes it?
  • What is the baseline, from which system, over which weeks? Who checked it?
  • How much does the number move week to week without any change? Is our target bigger than that?
  • What are we comparing the pilot with? A group without the AI in the same weeks, or only "before"?
  • What could get worse if the main number improves? Who owns each guardrail?
  • How many cases does the pilot need before the result means anything?
  • Which numbers come from SAP, and which from the AI system's own logs?
  • What happens if the result is real but smaller than the target?
  • Is SAP's published value for a similar feature measured on processes like ours?

Common misconceptions

  • "Model accuracy is the success metric." Accuracy says the answers are right. It doesn't say anyone decides faster or better.
  • "We'll measure the value after go-live." Without a baseline from before, there is nothing to compare with. The "before" data must be captured first.
  • "Before and after is good enough." Seasons, volume and staffing change too. A group without the AI in the same weeks removes most of that noise.
  • "More metrics give a fuller picture." Ten equal metrics let everyone pick their favorite. Keep one primary, a few guardrails, and the rest as supporting.
  • "If the target is missed, the project failed." A real but smaller gain may still pay off. The agreement says in advance who decides in that case.
  • "SAP's published figures will apply to us." They describe other customers' processes and starting points. Use them to choose where to look.

Key terms

  • Primary metric: the one business number the project is judged on. Also called the north star or overall evaluation criterion.
  • Baseline: the value of a metric before the change, measured over a stated period.
  • Target: the promised value or change, with a date.
  • Guardrail metric: a number that must not get worse, with an agreed limit.
  • Supporting metric: a number that helps explain the result but doesn't decide it.
  • Control group: cases handled without the AI in the same period, used for a fair comparison.
  • Leading and lagging metrics: leading metrics move quickly (notes opened); lagging ones move later (overdue receivables).
  • KPI: key performance indicator, a metric the business already tracks and reports.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1An AI note-drafting tool scores 92% on the evaluation set. Why is that not enough to call the project a success?

    Answer: B. Evaluation shows the answers are right. The sponsor pays for a business outcome, such as faster decisions, and that depends on adoption and on the process too. You need both layers of numbers.
  2. 2What should a metric agreement settle first, before any number is chosen?

    Answer: C. The decision shapes everything else. If the numbers will decide rollout or stop, the primary metric, target and guardrails all follow from that. Without a decision, metrics become reporting nobody acts on.
  3. 3Decisions on blocked orders got 34% faster in the pilot month. Order volume also dropped that month. What is the fairest way to judge the assistant?

    Answer: D. A control group in the same weeks sees the same season and volume. The difference between the two groups is what the assistant caused. Before-and-after numbers mix the assistant with everything else that changed.
  4. 4Analysts decide faster with the assistant, but overdue receivables on released orders start rising. Which part of the agreement should catch this?

    Answer: A. Guardrails name what must not get worse while the main number improves. Here, faster releases could mean fewer checks. A guardrail on overdue items, with an owner and a limit, stops the rollout if it is breached.
  5. 5Your vendor quotes another SAP customer's 70% time saving for a similar feature. How should you use that figure?

    Answer: C. Published figures describe other processes and starting points. They help you choose where to look. Your baseline and target must come from your own system data.
  6. 6The pilot shows a real 20% improvement, but the target was 25%. What should happen?

    Answer: D. A smaller real gain may still be worth it. The agreement names who decides, so the outcome is a business decision, not an argument about the target after the fact.
Deep layer · 35 min read

Mental model

A success metric is a test you write before the code. It has an input (the cases in a period), an expected result (the target), and failure conditions (the guardrails). Like a test, it is only useful if it is written down before anyone sees the answer.

The numbers form a chain. Each layer explains the one above it:

flowchart BT
  Q[Model quality<br/>note right, grounded] --> A[Adoption<br/>note opened, rated useful]
  A --> P[Process<br/>decided within 24 h]
  P --> B[Business outcome<br/>median hours to decision]
  G[Guardrails<br/>overdue after release, cost] -.must not get worse.-> B

Earlier Unit 8 topics built the bottom layer: LLM evaluation, groundedness and the evaluation harness. This topic adds the top: the number the sponsor pays for, measured fairly. In the first topic of the course, an FDE's job was one sentence: move metric M in process P from X to Y by date D. This is how you make that sentence measurable.

How it works

Step one: write the metric as a definition, not a phrase

"Faster decisions" is a phrase. A definition says exactly what is counted:

Field Example
Name Median hours from block to decision
Formula median(decided_at minus blocked_at), credit-blocked orders decided in the period
Unit and direction hours; lower is better
Population credit blocks only; one sales organization; orders decided in the period
Source order block and release events from SAP, agreed with the SAP team
Owner Head of Credit Management

The population line prevents most arguments later. Does a rejected order count? An order blocked twice? An order still blocked at the end of the period? Decide once, in writing.

Why the median? Decision times are skewed: most orders take hours, a few take weeks. The mean jumps when one stuck order is resolved. The median, the middle value, doesn't.

Step two: pick the layers

Microsoft's experimentation team recommends a core metric set with user satisfaction metrics, guardrail metrics, feature and engagement metrics, and data quality metrics. It describes guardrails as aspects you don't want to degrade but won't necessarily improve. Mapped to an SAP AI assistant:

Role Purpose Blocked-order example Typical source
Primary The one number the decision rests on Median hours to decision SAP order events
Business guardrail Stops a "win" that harms the business Released orders overdue after 30 days Receivables aging
Quality guardrail Stops a rollout of a tool people distrust Notes rated useful at least 75% Assistant feedback log
Cost guardrail Keeps unit cost within the business case Model cost per order at most $0.05 Provider usage export
Supporting Explains movement Notes opened; share decided within 24 hours Assistant log; SAP events

The Goals-Signals-Metrics process from Google's HEART work gives the order to fill this table in: state the goal, name the signal that would show it, then pick the metric that measures the signal. Start from the goal, not from the data you happen to have.

One primary metric. With several equal metrics, some will improve by chance, and people will quote those. One primary metric plus explicit guardrails makes the decision rule clear.

Leading and lagging. Notes opened moves on day one. Overdue receivables take 30 days or more to show. A pilot shorter than the slowest guardrail can't check it, so plan the pilot length around the guardrail, not only the primary metric.

Step three: measure the baseline and its noise

Take the last 8 to 12 weeks of data. Compute the metric for the whole period and for each week. The spread between weeks is the noise: how much the number moves with no change at all.

If your target is smaller than the weekly spread, a before-and-after chart can't show it. You will need a control group and enough cases.

Some metrics have no baseline. "Notes opened" doesn't exist before the assistant. That's fine: those are supporting metrics measured only in the pilot.

Step four: compare fairly

Before-and-after comparisons mix your change with everything else: season, volume, staffing, a new credit policy. The fix is a control group: analysts or orders without the assistant, in the same weeks.

How you split matters:

  • By analyst is usual for an assistant. Each analyst always or never has it. That avoids people copying habits between orders.
  • By order is fairer statistically, but one analyst then sees both versions and can learn from the notes.
  • By team or site is easiest to run but leaves fewer groups, so differences between teams can masquerade as an effect.

Microsoft's guidance also warns about sample ratio mismatch: if you planned a 50/50 split and get 60/40, something in the assignment is broken, and the result can't be trusted. Check group sizes before reading any result.

Step five: size the pilot

A pilot that is too small can't distinguish a real effect from noise. Microsoft's guidance says a power calculation sets the lower bound on how many units the test needs. For a median of skewed times, there is no simple formula, so simulate it: draw fake pilots from your baseline data, apply the target change to one group, and count how often the difference stands out from the noise.

Step six: decide with a pre-agreed rule

The scorecard reports the pilot-versus-control change with a 95% interval, and each guardrail against its limit. The verdict follows a rule written before the pilot:

flowchart TD
  S[Pilot results] --> G{Any guardrail<br/>breached?}
  G -->|yes| STOP[Stop and investigate]
  G -->|no| R{Interval shows<br/>a real improvement?}
  R -->|no| N[No clear effect:<br/>extend or rethink]
  R -->|yes| T{Point estimate<br/>meets target?}
  T -->|yes| GO[Target met:<br/>recommend rollout]
  T -->|no| D[Improved, below target:<br/>sponsor decides]

This is one simple rule, suitable for a pilot. Statisticians may prefer stricter versions. What matters is that it is agreed in advance.

Build it yourself: a metric spec, a baseline and a pilot scorecard

You will write a metric spec for the blocked-order assistant, measure its baseline from order events, estimate how many orders a pilot needs, and score a pilot against a control group. The output is a one-page scorecard a sponsor can read.

Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder, its .venv and the unit08 folder. This script uses only Python's built-in modules, so there is nothing new to install.

flowchart LR
  E[(Order events CSV)] --> B[baseline]
  SP[Metric spec JSON] --> C[check-spec]
  SP --> B
  B --> Z[size]
  E --> P[pilot]
  SP --> P
  P --> SC[scorecard.md]

What you need

  • Your course folder from the setup topics.
  • About 45 to 60 minutes.
  • No accounts and no cost. The script makes up its own order events.
  • Optional: an export of real block and release times from your SAP team, for the --csv option.

Step 1: Open your course folder

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate

Step 2: Create the script

  1. In VS Code, right-click the unit08 folder, choose New File, name it success_metrics.py, paste the code below and save.
"""Unit 8: define success metrics with the business, measure a baseline, size a pilot and score it.

The use case is the blocked sales order assistant from earlier units. A metric spec (JSON) says what
"success" means: one primary business metric with a target, guardrails that must not get worse, and
supporting metrics. This script checks the spec, measures the baseline from order events, estimates
how many orders a pilot needs, and writes a scorecard comparing pilot orders with control orders.

How to run (from your course folder, with .venv turned on):
    python unit08/success_metrics.py sample          # write made-up events and a starter metric spec
    python unit08/success_metrics.py check-spec      # is every metric fully defined?
    python unit08/success_metrics.py baseline        # measure the "before" weeks
    python unit08/success_metrics.py size            # how many orders per group does the pilot need?
    python unit08/success_metrics.py pilot           # pilot vs. control, guardrails, verdict, scorecard.md
    python unit08/success_metrics.py baseline --csv path/to/your_export.csv   # use your own events

Standard library only. All data from "sample" is made up; codes and numbers are not SAP meanings.
"""
import argparse
import csv
import json
import random
import statistics
import sys
from datetime import datetime, timedelta
from pathlib import Path

HERE = Path(__file__).resolve().parent
EVENTS = HERE / "data" / "blocked_orders_events.csv"
SPEC = HERE / "metrics" / "blocked_orders_metrics.json"
RESULTS = HERE / "results"
BASELINE_OUT = RESULTS / "baseline_metrics.json"
SCORECARD = RESULTS / "scorecard.md"

COLUMNS = ["SalesOrder", "period", "group", "week", "blocked_at", "decided_at", "decision",
           "TotalNetAmount", "note_opened", "note_rating", "overdue_30d", "model_cost_usd"]
REQUIRED = ["id", "name", "layer", "role", "direction", "unit", "formula", "source", "owner"]
LAYERS = {"business", "process", "adoption", "quality", "cost"}
ROLES = {"primary", "guardrail", "supporting"}

STARTER_SPEC = {
    "use_case": "Blocked sales order assistant: draft notes that help credit analysts decide faster",
    "decision": "Roll out to all credit analysts in the sales organization, or stop",
    "sponsor": "Head of Credit Management",
    "metrics": [
        {"id": "hours_to_decision", "name": "Median hours from block to decision", "layer": "business",
         "role": "primary", "direction": "down", "unit": "hours", "target_change_pct": -25,
         "formula": "median(decided_at - blocked_at) over credit-blocked orders decided in the period",
         "source": "order block and release events exported from SAP (agree the source with the SAP team)",
         "owner": "Head of Credit Management"},
        {"id": "decided_within_24h", "name": "Share decided within 24 hours", "layer": "process",
         "role": "supporting", "direction": "up", "unit": "share",
         "formula": "count(hours <= 24) / count(decided orders)",
         "source": "same events as hours_to_decision", "owner": "Credit team lead"},
        {"id": "overdue_after_release", "name": "Released orders overdue more than 30 days",
         "layer": "business", "role": "guardrail", "direction": "down", "unit": "share",
         "max_worsening": 0.01,
         "formula": "count(released and overdue_30d) / count(released)",
         "source": "receivables aging joined to released orders", "owner": "Head of Credit Management"},
        {"id": "note_opened", "name": "Notes opened by the analyst", "layer": "adoption",
         "role": "supporting", "direction": "up", "unit": "share", "target": 0.7,
         "formula": "count(note_opened) / count(pilot orders)",
         "source": "assistant usage log", "owner": "Product owner"},
        {"id": "note_useful", "name": "Notes rated useful", "layer": "quality",
         "role": "guardrail", "direction": "up", "unit": "share", "min_value": 0.75,
         "formula": "count(rating = useful) / count(rated notes)",
         "source": "thumbs up/down in the assistant", "owner": "Product owner"},
        {"id": "cost_per_order", "name": "Model cost per order", "layer": "cost",
         "role": "guardrail", "direction": "down", "unit": "usd", "max_value": 0.05,
         "formula": "sum(model_cost_usd) / count(pilot orders)",
         "source": "model provider usage export", "owner": "IT finance"},
    ],
}

# ---------------------------------------------------------------- made-up events


def make_sample(seed=8):
    """Eight 'before' weeks, then four pilot weeks split into control and pilot analysts."""
    rng = random.Random(seed)
    start = datetime(2026, 6, 1, 8, 0)
    rows, number = [], 9200000

    def add(period, group, week, median_hours, overdue_rate, assisted):
        nonlocal number
        number += 1
        blocked = start + timedelta(days=7 * week + rng.uniform(0, 5), hours=rng.uniform(0, 9))
        hours = rng.lognormvariate(0, 0.75) * median_hours
        released = rng.random() < 0.88
        opened = assisted and rng.random() < 0.78
        rating = ""
        if opened and rng.random() < 0.6:
            rating = "useful" if rng.random() < 0.83 else "not useful"
        rows.append({
            "SalesOrder": str(number), "period": period, "group": group, "week": week + 1,
            "blocked_at": blocked.isoformat(timespec="minutes"),
            "decided_at": (blocked + timedelta(hours=hours)).isoformat(timespec="minutes"),
            "decision": "released" if released else "rejected",
            "TotalNetAmount": f"{rng.uniform(2000, 90000):.2f}",
            "note_opened": "yes" if opened else ("no" if assisted else ""),
            "note_rating": rating,
            "overdue_30d": "yes" if released and rng.random() < overdue_rate else "no",
            "model_cost_usd": f"{rng.uniform(0.01, 0.03):.4f}" if assisted else "0",
        })

    for week in range(8):                              # before: one team, no assistant
        for _ in range(rng.randint(52, 64)):
            add("before", "all", week, 20.0, 0.035, False)
    for week in range(8, 12):                          # pilot weeks: quieter season for everyone
        for _ in range(rng.randint(26, 32)):
            add("pilot", "control", week, 17.0, 0.035, False)
        for _ in range(rng.randint(26, 32)):
            add("pilot", "pilot", week, 12.0, 0.035, True)
    return rows


# ---------------------------------------------------------------- loading


def load_events(path):
    if not path.exists():
        sys.exit(f"No events file at {path}. Run: python unit08/success_metrics.py sample")
    with open(path, newline="", encoding="utf-8") as f:
        rows = list(csv.DictReader(f))
    missing = [c for c in COLUMNS if rows and c not in rows[0]]
    if not rows or missing:
        sys.exit(f"{path} has no rows or lacks columns: {', '.join(missing) or 'all'}")
    for r in rows:
        r["hours"] = (datetime.fromisoformat(r["decided_at"])
                      - datetime.fromisoformat(r["blocked_at"])).total_seconds() / 3600
    return rows


def load_spec():
    if not SPEC.exists():
        sys.exit(f"No metric spec at {SPEC}. Run: python unit08/success_metrics.py sample")
    return json.loads(SPEC.read_text(encoding="utf-8"))


# ---------------------------------------------------------------- metric formulas


def share(rows, test, base=lambda r: True):
    pool = [r for r in rows if base(r)]
    return sum(1 for r in pool if test(r)) / len(pool) if pool else None


FORMULAS = {
    "hours_to_decision": lambda rows: statistics.median(r["hours"] for r in rows) if rows else None,
    "decided_within_24h": lambda rows: share(rows, lambda r: r["hours"] <= 24),
    "overdue_after_release": lambda rows: share(rows, lambda r: r["overdue_30d"] == "yes",
                                                lambda r: r["decision"] == "released"),
    "note_opened": lambda rows: share(rows, lambda r: r["note_opened"] == "yes",
                                      lambda r: r["note_opened"] != ""),
    "note_useful": lambda rows: share(rows, lambda r: r["note_rating"] == "useful",
                                      lambda r: r["note_rating"] != ""),
    "cost_per_order": lambda rows: (sum(float(r["model_cost_usd"]) for r in rows) / len(rows)) if rows else None,
}


def fmt(value, unit):
    if value is None:
        return "n/a"
    if unit == "share":
        return f"{value:.1%}"
    if unit == "usd":
        return f"${value:.3f}"
    return f"{value:.1f} h" if unit == "hours" else f"{value:.2f}"


# ---------------------------------------------------------------- commands


def cmd_sample(args):
    EVENTS.parent.mkdir(parents=True, exist_ok=True)
    SPEC.parent.mkdir(parents=True, exist_ok=True)
    rows = make_sample()
    with open(EVENTS, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=COLUMNS)
        writer.writeheader()
        writer.writerows(rows)
    print(f"Wrote {len(rows)} made-up order events to unit08/data/{EVENTS.name}")
    if SPEC.exists() and not args.force:
        print(f"Kept your existing unit08/metrics/{SPEC.name} (add --force to overwrite it)")
    else:
        SPEC.write_text(json.dumps(STARTER_SPEC, indent=2) + "\n", encoding="utf-8")
        print(f"Wrote the starter metric spec to unit08/metrics/{SPEC.name}")


def spec_problems(spec):
    problems = []
    for key in ("use_case", "decision", "sponsor"):
        if not str(spec.get(key, "")).strip():
            problems.append(f"spec: '{key}' is empty; say who decides what, based on these numbers")
    metrics = spec.get("metrics", [])
    for m in metrics:
        label = m.get("id", "?")
        for field in REQUIRED:
            if not str(m.get(field, "")).strip():
                problems.append(f"{label}: '{field}' is missing")
        if m.get("layer") and m["layer"] not in LAYERS:
            problems.append(f"{label}: layer must be one of {sorted(LAYERS)}")
        if m.get("role") and m["role"] not in ROLES:
            problems.append(f"{label}: role must be one of {sorted(ROLES)}")
        if m.get("direction") not in ("up", "down"):
            problems.append(f"{label}: direction must be 'up' or 'down'")
        if m.get("id") and m["id"] not in FORMULAS:
            problems.append(f"{label}: no formula in the script for this id (add one to FORMULAS)")
        if m.get("role") == "primary" and "target_change_pct" not in m:
            problems.append(f"{label}: the primary metric needs a target_change_pct")
        if m.get("role") == "guardrail" and not any(k in m for k in ("max_worsening", "max_value", "min_value")):
            problems.append(f"{label}: a guardrail needs max_worsening, max_value or min_value")
    primaries = [m for m in metrics if m.get("role") == "primary"]
    if len(primaries) != 1:
        problems.append(f"spec: needs exactly one primary metric, found {len(primaries)}")
    if not any(m.get("role") == "guardrail" and m.get("layer") == "business" for m in metrics):
        problems.append("spec: add at least one business guardrail (what must not get worse)")
    return problems


def cmd_check_spec(args):
    spec = load_spec()
    problems = spec_problems(spec)
    for m in spec.get("metrics", []):
        print(f"  {m.get('role', '?'):<10} {m.get('layer', '?'):<9} {m.get('id', '?')}")
    if problems:
        print(f"\n{len(problems)} problem(s):")
        for p in problems:
            print(f"  - {p}")
        sys.exit(1)
    print("\nSpec OK: one primary metric with a target, guardrails with limits, every metric owned.")


def cmd_baseline(args):
    spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
    before = [r for r in rows if r["period"] == "before"]
    if not before:
        sys.exit("No rows with period=before. The baseline needs the weeks before any change.")
    weeks = sorted({int(r["week"]) for r in before})
    print(f"Baseline: {len(before)} orders over {len(weeks)} weeks (period=before)\n")
    print(f"{'metric':<24}{'baseline':>10}{'lowest wk':>11}{'highest wk':>12}")
    out = {"orders": len(before), "weeks": len(weeks), "metrics": {}}
    for m in spec["metrics"]:
        value = FORMULAS[m["id"]](before)
        weekly = [FORMULAS[m["id"]]([r for r in before if int(r["week"]) == w]) for w in weeks]
        weekly = [v for v in weekly if v is not None]
        out["metrics"][m["id"]] = value
        if value is None:
            print(f"{m['id']:<24}{'n/a':>10}   (not measurable before the assistant exists)")
            continue
        print(f"{m['id']:<24}{fmt(value, m['unit']):>10}{fmt(min(weekly), m['unit']):>11}"
              f"{fmt(max(weekly), m['unit']):>12}")
    RESULTS.mkdir(parents=True, exist_ok=True)
    BASELINE_OUT.write_text(json.dumps(out, indent=2) + "\n", encoding="utf-8")
    print(f"\nSaved to unit08/results/{BASELINE_OUT.name}. Week-to-week spread is normal noise:"
          "\na change smaller than that spread can't be seen in a before/after chart.")


def simulated_power(hours, change_pct, n, rng, sims=1000):
    """Share of simulated pilots of n orders per group whose median gap beats the noise level."""
    factor = 1 + change_pct / 100

    def gap(a, b):
        return statistics.median(b) - statistics.median(a)

    null = sorted(gap(rng.choices(hours, k=n), rng.choices(hours, k=n)) for _ in range(sims))
    cutoff = null[int(0.025 * sims)]              # 2.5% of no-effect pilots look this good by chance
    hits = sum(gap(rng.choices(hours, k=n), [h * factor for h in rng.choices(hours, k=n)]) < cutoff
               for _ in range(sims))
    return hits / sims


def cmd_size(args):
    spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
    primary = next(m for m in spec["metrics"] if m["role"] == "primary")
    hours = [r["hours"] for r in rows if r["period"] == "before"]
    change = args.change if args.change is not None else primary["target_change_pct"]
    rng = random.Random(42)
    print(f"Orders per group needed to see a {change:+.0f}% change in {primary['id']}"
          f" (from {len(hours)} baseline orders)\n")
    print(f"{'orders per group':>17}{'chance to detect':>18}")
    enough = None
    for n in (30, 60, 100, 150, 200, 300, 400):
        power = simulated_power(hours, change, n, rng)
        print(f"{n:>17}{power:>18.0%}")
        if enough is None and power >= 0.8:
            enough = n
    if enough:
        print(f"\nAbout {enough} orders per group give an 80% chance to detect the change."
              "\nFewer, and a real improvement can easily look like 'no clear effect'.")
    else:
        print("\nEven 400 orders per group give less than 80%. Plan a longer pilot or a bigger change.")


def bootstrap_relative(control, pilot, func, rng, reps=1000):
    """95% interval for (pilot - control) / control, resampling each group with replacement."""
    values = []
    for _ in range(reps):
        c = func(rng.choices(control, k=len(control)))
        p = func(rng.choices(pilot, k=len(pilot)))
        if c:
            values.append((p - c) / c)
    values.sort()
    return values[int(0.025 * len(values))], values[int(0.975 * len(values)) - 1]


def cmd_pilot(args):
    spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
    before = [r for r in rows if r["period"] == "before"]
    control = [r for r in rows if r["period"] == "pilot" and r["group"] == "control"]
    pilot = [r for r in rows if r["period"] == "pilot" and r["group"] == "pilot"]
    if not control or not pilot:
        sys.exit("Need rows with period=pilot and group=control and group=pilot.")
    rng = random.Random(7)
    primary = next(m for m in spec["metrics"] if m["role"] == "primary")
    f = FORMULAS[primary["id"]]
    b, c, p = f(before), f(control), f(pilot)
    naive, controlled = (p - b) / b, (p - c) / c
    low, high = bootstrap_relative(control, pilot, f, rng)
    target = primary["target_change_pct"] / 100
    better = (high < 0) if primary["direction"] == "down" else (low > 0)
    meets = (controlled <= target) if primary["direction"] == "down" else (controlled >= target)

    lines = [f"# Pilot scorecard: {spec['use_case']}", "",
             f"Decision this informs: {spec['decision']} (sponsor: {spec['sponsor']})", "",
             f"Orders: {len(before)} before, {len(control)} control, {len(pilot)} pilot", "",
             "## Primary metric", "",
             f"{primary['name']}: before {fmt(b, primary['unit'])}, control {fmt(c, primary['unit'])},"
             f" pilot {fmt(p, primary['unit'])}", "",
             f"- Before/after change (misleading, includes the season): {naive:+.0%}",
             f"- Pilot vs. control, same weeks: {controlled:+.0%}"
             f" (95% interval {low:+.0%} to {high:+.0%}); target {target:+.0%}", "",
             "## Guardrails and supporting metrics", "",
             "| Metric | Role | Control | Pilot | Limit | Status |", "| --- | --- | --- | --- | --- | --- |"]
    breached = []
    for m in spec["metrics"]:
        if m is primary:
            continue
        fm = FORMULAS[m["id"]]
        cv, pv = fm(control), fm(pilot)
        status, limit = "info", ""
        if "max_worsening" in m and cv is not None and pv is not None:
            worse = (pv - cv) if m["direction"] == "down" else (cv - pv)
            limit = (f"no worse than +{m['max_worsening'] * 100:.1f} points" if m["unit"] == "share"
                     else f"no worse than +{fmt(m['max_worsening'], m['unit'])}")
            status = "BREACHED" if worse > m["max_worsening"] else "ok"
        elif "max_value" in m and pv is not None:
            limit, status = f"at most {fmt(m['max_value'], m['unit'])}", "BREACHED" if pv > m["max_value"] else "ok"
        elif "min_value" in m and pv is not None:
            limit, status = f"at least {fmt(m['min_value'], m['unit'])}", "BREACHED" if pv < m["min_value"] else "ok"
        elif "target" in m and pv is not None:
            limit, status = f"target {fmt(m['target'], m['unit'])}", "ok" if pv >= m["target"] else "below target"
        if status == "BREACHED" and m["role"] == "guardrail":
            breached.append(m["id"])
        lines.append(f"| {m['name']} | {m['role']} | {fmt(cv, m['unit'])} | {fmt(pv, m['unit'])} | {limit} | {status} |")

    if breached:
        verdict = f"STOP: guardrail breached ({', '.join(breached)}). Investigate before any rollout."
    elif better and meets:
        verdict = "TARGET MET: the improvement is real and reaches the target. Recommend rollout."
    elif better:
        verdict = "IMPROVED, BELOW TARGET: real improvement, smaller than promised. Decide with the sponsor."
    else:
        verdict = "NO CLEAR EFFECT: the interval includes no change. Extend the pilot or rethink."
    lines += ["", "## Verdict", "", verdict, ""]
    RESULTS.mkdir(parents=True, exist_ok=True)
    SCORECARD.write_text("\n".join(lines), encoding="utf-8")
    print("\n".join(lines[4:]))
    print(f"Saved to unit08/results/{SCORECARD.name}")


def main():
    parser = argparse.ArgumentParser(description="Success metrics for the blocked-order assistant")
    parser.add_argument("command", choices=["sample", "check-spec", "baseline", "size", "pilot"])
    parser.add_argument("--csv", help="your own events file with the same columns")
    parser.add_argument("--change", type=float, help="size: change to detect in percent, e.g. -15")
    parser.add_argument("--force", action="store_true", help="sample: overwrite the metric spec")
    args = parser.parse_args()
    {"sample": cmd_sample, "check-spec": cmd_check_spec, "baseline": cmd_baseline,
     "size": cmd_size, "pilot": cmd_pilot}[args.command](args)


if __name__ == "__main__":
    main()

Step 3: Write the sample events and the metric spec

  1. In the terminal, run:

    python unit08/success_metrics.py sample
  2. You should see:

    Wrote 688 made-up order events to unit08/data/blocked_orders_events.csv
    Wrote the starter metric spec to unit08/metrics/blocked_orders_metrics.json
  3. Open unit08/metrics/blocked_orders_metrics.json in VS Code. This is the metric agreement from the foundational layer, in a form a program can check. Each metric has a role (primary, guardrail or supporting), a layer, a direction, a formula in words, a source and an owner.

  4. Open unit08/data/blocked_orders_events.csv. Each row is one credit-blocked order: when it was blocked, when it was decided, whether it was released, and for pilot orders whether the analyst opened and rated the note. The period column says before or pilot; the group column says all, control or pilot.

If you run sample again, it rewrites the events but keeps your edited spec. Add --force to reset the spec too.

Step 4: Check the spec

  1. Run:

    python unit08/success_metrics.py check-spec
  2. You should see the six metrics and:

    Spec OK: one primary metric with a target, guardrails with limits, every metric owned.
  3. Now break it on purpose. In the JSON file, change the first metric's "owner" to "" and save. Run check-spec again. You should see:

    1 problem(s):
      - hours_to_decision: 'owner' is missing

    The script ends with exit code 1, so you could run it in CI like the evaluation harness. Put the owner back and save.

The checks encode the agreement rules: exactly one primary metric with a target, at least one business guardrail, a limit on every guardrail, and an owner on every metric.

Step 5: Measure the baseline

  1. Run:

    python unit08/success_metrics.py baseline
  2. You should see:

    Baseline: 459 orders over 8 weeks (period=before)
    
    metric                    baseline  lowest wk  highest wk
    hours_to_decision           19.7 h     18.2 h      22.8 h
    decided_within_24h           59.9%      52.7%       63.9%
    overdue_after_release         2.3%       0.0%        6.2%
    note_opened                    n/a   (not measurable before the assistant exists)
    note_useful                    n/a   (not measurable before the assistant exists)
    cost_per_order              $0.000     $0.000      $0.000
  3. Read the spread. The median decision time moved between 18.2 and 22.8 hours from week to week, with no change at all. The overdue rate swung between 0% and 6.2%, because only a few released orders go overdue each week. A one-week before-and-after comparison of either number would mean nothing.

n/a is a valid result here: adoption and rating metrics don't exist before the assistant. The baseline is saved in unit08/results/baseline_metrics.json.

Step 6: Size the pilot

  1. Run:

    python unit08/success_metrics.py size
  2. You should see:

    Orders per group needed to see a -25% change in hours_to_decision (from 459 baseline orders)
    
     orders per group  chance to detect
                   30               18%
                   60               30%
                  100               38%
                  150               59%
                  200               83%
                  300               94%
                  400               99%
    
    About 200 orders per group give an 80% chance to detect the change.
  3. Try a bigger effect: python unit08/success_metrics.py size --change -40. The script says about 100 orders per group are enough. Bigger effects need smaller pilots.

This is the conversation to have with the sponsor before the pilot. With about 30 credit blocks per group per week, 200 orders per group means about 7 weeks, not 4.

Step 7: Score the pilot

  1. Run:

    python unit08/success_metrics.py pilot
  2. You should see (shortened):

    Orders: 459 before, 112 control, 117 pilot
    
    ## Primary metric
    
    Median hours from block to decision: before 19.7 h, control 17.0 h, pilot 13.0 h
    
    - Before/after change (misleading, includes the season): -34%
    - Pilot vs. control, same weeks: -24% (95% interval -46% to -5%); target -25%
    
    | Metric | Role | Control | Pilot | Limit | Status |
    | --- | --- | --- | --- | --- | --- |
    | Released orders overdue more than 30 days | guardrail | 4.0% | 4.7% | no worse than +1.0 points | ok |
    | Notes rated useful | guardrail | n/a | 80.7% | at least 75.0% | ok |
    | Model cost per order | guardrail | $0.000 | $0.021 | at most $0.050 | ok |
    
    ## Verdict
    
    IMPROVED, BELOW TARGET: real improvement, smaller than promised. Decide with the sponsor.
  3. Open unit08/results/scorecard.md. VS Code shows a preview with the Open Preview to the Side button at the top right.

Read the result the way a sponsor would:

  • Before and after says 34%. But the control group also got faster, from 19.7 to 17.0 hours, because the made-up pilot weeks were a quieter season. The fair number is 24%.
  • The interval is wide, from 46% to 5% faster. The improvement is real (the whole range is faster), but the pilot had about 115 orders per group, while Step 6 said about 200 were needed. The point estimate just misses the 25% target.
  • No guardrail was breached. The overdue rate is 0.7 points higher in the pilot, inside the 1-point limit, but a 4-week pilot is short for a 30-day lagging metric.

The honest recommendation is to extend the pilot to the planned size before deciding, or to let the sponsor accept a real 24% gain. Either way, the decision is the sponsor's, and the scorecard says so.

How the code works

Part of the script What it does
STARTER_SPEC The metric agreement as data: use case, decision, sponsor and six metrics with role, layer, limits and owner
make_sample() Makes up 8 "before" weeks and 4 pilot weeks with a control and a pilot group; the seed makes it repeatable
load_events() Reads the CSV, checks the columns and computes hours from block to decision
FORMULAS One small function per metric id; this is the code version of each formula line in the spec
spec_problems() The agreement rules: one primary with a target, a business guardrail, limits on guardrails, owners
cmd_baseline() Computes each metric for the before period and per week, and saves the baseline
simulated_power() Draws fake pilots from baseline times, applies the target change to one group, and counts how often the gap beats chance
bootstrap_relative() Resamples control and pilot orders 1,000 times to get a 95% interval for the relative change
cmd_pilot() Compares pilot with control, checks each guardrail, applies the verdict rule and writes scorecard.md

Use your own data

If your SAP team can export real block and decision times, save them as a CSV with the same column names and run:

python unit08/success_metrics.py baseline --csv path/to/your_export.csv

Columns the baseline doesn't use, such as note_opened, can be empty. Keep the export to the fields the metrics need, and treat it as business data: don't commit it to a public repository.

If something goes wrong

What you see What it means What to do
python is "not recognized" or "command not found" Python isn't installed or isn't on the path Turn on .venv (Step 1); see Set up your computer
can't open file ... success_metrics.py You aren't in the course folder, or the file has another name Run from orchestrate-course; check the file is unit08/success_metrics.py
No events file or No metric spec sample hasn't run yet python unit08/success_metrics.py sample
lacks columns: ... Your own CSV uses other column names Rename the columns to match the sample file's header
json.decoder.JSONDecodeError The spec has a typo, often a missing comma or quote Undo your last edit, or run sample --force to reset it
no formula in the script for this id You added a metric to the spec but not to FORMULAS Add a matching function, as in the Exercise
ModuleNotFoundError You copied code that imports a library this script doesn't use Re-paste the script; it needs only built-in modules
A network or proxy error Not expected: this script makes no network calls Check you're running this script, not another unit's

The SAP way

You rarely need to invent the business metric. In SAP landscapes, process numbers often already exist. As of October 2026, these are the places to look before building your own baseline.

KPIs already in SAP S/4HANA

SAP Learning's lesson on SAP Smart Business for sales order fulfillment describes:

  • The Sales Order Fulfillment Issues app, which lists blocked sales orders for analysis and processing. Issue types include incomplete data, unconfirmed quantities, billing blocks and delivery blocks.
  • KPI visualizations with semantic colors based on defined targets and thresholds.
  • The Manage KPIs and Reports app, available since SAP S/4HANA 1909, for configuring new KPIs. The lesson names the business role SAP_BR_ANALYTICS_SPECIALIST for it.

For the blocked-order assistant, that means the people who own blocked orders may already watch a KPI on them. Use their definition as your primary metric if it fits. A metric the business already reports is easier to trust than one you invent.

Process mining for the baseline

SAP Signavio Process Intelligence computes metrics on process event data. SAP Learning describes a metric library of predefined calculations written as SiGNAL expressions, which can be reused to calculate KPIs. Every process gets an average cycle time metric by default. One practical warning from the same lesson: process views control data access, and a metric can become invalid if the view hides data it needs. Check that your baseline isn't silently filtered.

SAP Signavio Process Insights comes with prebuilt content. SAP Learning lists more than 200 process performance indicators, with benchmarking inside the company and against industry peers. Standard indicator types include throughput, backlog, exceptions and automation rate. Backlog and automation rate are natural guardrail or supporting metrics for an AI assistant.

SAP's own value estimates

SAP Discovery Center hosts the SAP Business AI Catalog. SAP News reported almost 400 features and agents in it in June 2026, with cost and ROI estimator tools. In August 2025, SAP said its value estimates combine efficiency and effectiveness, and rest on SAP's industry expertise, benchmarking data, third-party research and early customer feedback.

Use these the way you used SAP's feature tables in Where AI creates value in S/4HANA: to choose candidates and frame the target conversation. Don't paste them into your metric agreement as a baseline or a promise.

Measuring an SAP feature you switch on

The same method works for SAP's own AI features, not only custom builds. If you switch on a standard feature for one team first, that team is the pilot group and the others are the control. The agreement, the baseline and the guardrails are the same; only the system under test changes.

Build vs. SAP

Situation Lean towards Why
The business already tracks a KPI for the process in S/4HANA Use that KPI as the primary metric Trusted definition, owner already exists
Your company runs SAP Signavio Baseline from Process Intelligence or Process Insights Event data and cycle times already computed
No process analytics available An export of the events plus this script Cheap, transparent, enough for a pilot
Adoption, ratings and cost Your AI system's own logs SAP doesn't see what happens inside your assistant
Estimating value before a pilot SAP Business AI Catalog estimates, then your own baseline Estimates point the way; your data decides
Business case for a budget The pilot scorecard, carried into Unit 10 Measured effect, interval and unit cost, not projections

Production concerns

  • Authorizations and data protection. Order events carry customer names and amounts. Export only the fields the metrics need, under the same access rules as the source, and keep exports out of public repositories. SAP Signavio's process views control which data a metric can see; respect them rather than working around them.
  • Metric definitions are versioned. Keep the spec in Git with the code. If the definition changes mid-pilot, results before and after the change can't be compared.
  • Data quality first. Check group sizes against the planned split, missing timestamps, and orders decided before they were blocked. Microsoft's guidance treats a sample ratio mismatch as a sign the result can't be trusted.
  • Lagging guardrails. Keep measuring guardrails such as overdue receivables after the pilot ends, until their lag has passed.
  • Cost per unit. Report model cost per order next to the business gain. Unit 10 turns that into a business case.
  • Ownership after go-live. The metric needs an owner after the project team leaves. Put it on the process owner's existing KPI report if possible.
  • Clean core. Measurement reads data through exports, released APIs or process mining. It needs no modification to SAP standard.

Pitfalls

  • Choosing the metric after seeing the data. Some number always improved. Write the spec first, and commit it.
  • Several primary metrics. Everyone quotes their favorite. One primary, explicit guardrails.
  • Before and after only. Seasons and volume change too. Use a control group in the same weeks.
  • A pilot too small to answer. Size it first; otherwise "no clear effect" is the likely result even when the AI works.
  • A guardrail nobody owns. It won't be measured. Every guardrail gets an owner and a source.
  • Ignoring adoption. If analysts don't open the notes, the outcome can't move. Low adoption is a different problem from a bad model.
  • Gaming the metric. Once people are judged on a number, they find ways to move it. Guardrails such as release rate catch the obvious shortcuts.
  • Mean of skewed times. A few stuck orders dominate the mean. Use the median or a share such as "decided within 24 hours".

Exercise

Add a guardrail against rubber-stamping: analysts might decide faster simply by releasing more orders. Then tighten it to see a guardrail stop the rollout.

  1. Open unit08/success_metrics.py. Find the line that starts with "note_opened": lambda rows: inside FORMULAS. Just above it, at the same indentation, add:

        "released_share": lambda rows: share(rows, lambda r: r["decision"] == "released"),

    Save the file.

  2. Open unit08/metrics/blocked_orders_metrics.json. After the last metric's closing }, add a comma and this new metric, then save:

    {"id": "released_share", "name": "Share of blocked orders released", "layer": "business",
     "role": "guardrail", "direction": "down", "unit": "share", "max_worsening": 0.03,
     "formula": "count(decision = released) / count(decided orders)",
     "source": "same events as hours_to_decision", "owner": "Head of Credit Management"}
  3. Run python unit08/success_metrics.py check-spec. You should see seven metrics and Spec OK.

  4. Run python unit08/success_metrics.py pilot. A new row shows the share released: 89.3% for control, 91.5% for pilot, status ok.

  5. Change "max_worsening": 0.03 to 0.02 and save. Run pilot again. The row now says BREACHED and the verdict says STOP: guardrail breached (released_share).

  6. Put the limit back to 0.03. In unit08/results/, create metric_agreement.md with five lines: the decision, the primary metric with baseline and target, your guardrails with limits, the comparison group, and the pilot size from Step 6.

  7. Save your work:

    git add unit08/success_metrics.py unit08/metrics/blocked_orders_metrics.json unit08/results/metric_agreement.md
    git commit -m "Add success metrics, baseline and pilot scorecard"

Done when check-spec passes with seven metrics, the 0.02 limit produced a STOP verdict, and metric_agreement.md holds your five lines. Keep it: Unit 10's business case starts from this agreement and the scorecard.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the script use the median of decision hours rather than the mean?

    Answer: B. Decision times are skewed: most take hours, a few take weeks. The mean jumps when one stuck order closes; the median, the middle value, stays stable, so changes reflect typical orders.
  2. 2The baseline shows weekly median decision times between 18.2 and 22.8 hours with no change. What does that spread tell you?

    Answer: C. The spread is noise, how much the number moves on its own. An effect smaller than that range disappears in a simple before-and-after view, which is why the pilot needs a control group and enough orders.
  3. 3In the sample pilot, before-and-after shows -34% but pilot vs. control shows -24%. What explains the gap?

    Answer: A. Both groups worked in the same quieter weeks, and the control group went from 19.7 to 17.0 hours without the assistant. Comparing pilot with control in the same weeks removes that shared change; before-and-after counts it as the assistant's effect.
  4. 4The size command says about 200 orders per group give an 80% chance to detect a 25% change. The pilot had about 115 per group. What is the right reading?

    Answer: D. Power says how likely a pilot of a given size is to detect a real effect. With fewer orders than planned, the interval widens, here from 46% to 5% faster, so the point estimate can easily miss the target by chance.
  5. 5You plan a 4-week pilot. One guardrail is "released orders overdue after 30 days". What should you do?

    Answer: C. Overdue receivables are a lagging metric; harm from careless releases shows up weeks later. Dropping or swapping it removes the protection it gives, so keep tracking it past the pilot before final rollout.
  6. 6A colleague adds a new metric id to the JSON spec, and check-spec fails with "no formula in the script for this id". Why does the script insist on this?

    Answer: B. The spec's formula line is words; FORMULAS is the code that computes it. A metric with no code would appear in the agreement but never in the scorecard, so the check catches the gap before the pilot.
  7. 7Your company runs SAP Signavio Process Intelligence, and its average cycle time for credit blocks looks far lower than your export. What should you check first?

    Answer: D. SAP Learning warns that process views control data access, and a metric can become invalid if the view restricts data it needs. A filtered baseline can look better than reality, so confirm the population before you trust either number.
  8. 8Your assistant cuts decision time by 30% in a well-sized pilot, but the share of orders released rises 4 points against a 3-point limit. What should the scorecard verdict be?

    Answer: C. The decision rule checks guardrails first. A breached guardrail stops the rollout whatever the primary metric says, because faster decisions may come from releasing orders without proper checks.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in