Every AI project promises something: faster decisions, fewer errors, lower cost. A success metric turns that promise into a number both sides agree on before the work starts.
A good agreement has four parts. One primary metric the business cares about, such as hours to decide a blocked order. A baseline: today's value, measured from real data. A target: the value you promise, by a date. And guardrails: things that must not get worse on the way, such as bad credit decisions.
Unit 8 so far measured whether the AI's answers are right. This topic measures whether the business is better off. Both matter. A model can score well and still change nothing, if nobody uses it or the bottleneck is elsewhere.
Most AI pilots end with a demo, a happy sponsor and no proof. Six months later, when budget is renewed, nobody can say what changed. The project gets cut, or worse, it keeps running at a cost nobody can justify.
Take the blocked sales order assistant used throughout this course. Credit analysts get a drafted note for each order on credit hold. The pitch is "analysts decide faster". Without an agreed metric, three problems appear:
No baseline. Nobody measured how long decisions took before. Any number after go-live has nothing to compare with.
The wrong comparison. Decisions got faster in the pilot month. But order volume also dropped that month, so everyone was faster, with or without the assistant.
A hidden cost. Analysts decide faster because they release more orders without checking. Overdue receivables rise a quarter later.
Agreeing the metrics first fixes all three. It costs a workshop and a few days of data work. It also forces the most useful question of the whole project: what decision will we make based on this number?
SAP offers several places where process numbers already live. Use them for baselines before you build anything. As of October 2026:
SAP S/4HANA, SAP Smart Business. SAP Learning describes the Sales Order Fulfillment Issues app for analyzing and processing blocked sales orders. It covers issues such as billing blocks, delivery blocks and incomplete data. Smart Business shows KPIs with colors based on defined targets and thresholds. The Manage KPIs and Reports app, available since S/4HANA 1909, lets an analytics specialist configure new KPIs.
SAP Signavio Process Insights. SAP Learning lists more than 200 process performance indicators, with comparison inside your company and against industry peers. Standard indicator types include throughput, backlog, exceptions and automation rate.
SAP Signavio Process Intelligence. It holds a library of metrics, predefined calculations that are reused to compute KPIs. Every process gets an average cycle time metric by default.
SAP Business AI Catalog in SAP Discovery Center. SAP News reported almost 400 features and agents in June 2026, plus cost and ROI estimators. In August 2025 SAP said its value estimates rest on SAP's own expertise, benchmarking, third-party research and early customer feedback.
Treat SAP's value figures as a reason to look, not as your baseline. At Sapphire 2026, SAP quoted a customer cutting wholesale order processing from days to minutes. That is their process and their starting point. Your number has to come from your own system.
Write this down and get the sponsor to sign it before the pilot starts.
Part
What it says
Blocked-order example
Decision
What the numbers will decide, and who decides
Roll out to all credit analysts, or stop. Head of Credit Management decides.
Primary metric
One business number, defined exactly
Median hours from credit block to release or rejection
Baseline
Today's value, from system data, with the period
19.7 hours over the last 8 weeks
Target
The promised change and the date
25% faster in a 4-week pilot
Guardrails
What must not get worse, with a limit
Released orders overdue after 30 days: no more than 1 point higher
Supporting metrics
Numbers that explain the result
Notes opened, notes rated useful, cost per order
Comparison
What the pilot is compared with
Analysts without the assistant in the same weeks
Owner and source
Who owns each number, where it comes from
Credit team for decisions, finance for overdue items, IT for cost
Three layers of numbers sit behind this. Model quality (is the note right?) is what earlier Unit 8 topics measured. Adoption (do analysts open the note?) tells you whether people use it. Business outcome (do decisions get faster, safely?) is what the sponsor pays for. If the outcome doesn't move, the lower layers tell you why.
"Model accuracy is the success metric." Accuracy says the answers are right. It doesn't say anyone decides faster or better.
"We'll measure the value after go-live." Without a baseline from before, there is nothing to compare with. The "before" data must be captured first.
"Before and after is good enough." Seasons, volume and staffing change too. A group without the AI in the same weeks removes most of that noise.
"More metrics give a fuller picture." Ten equal metrics let everyone pick their favorite. Keep one primary, a few guardrails, and the rest as supporting.
"If the target is missed, the project failed." A real but smaller gain may still pay off. The agreement says in advance who decides in that case.
"SAP's published figures will apply to us." They describe other customers' processes and starting points. Use them to choose where to look.
Pick one answer for each question. The explanation appears after you choose.
1An AI note-drafting tool scores 92% on the evaluation set. Why is that not enough to call the project a success?
Answer: B. Evaluation shows the answers are right. The sponsor pays for a business outcome, such as faster decisions, and that depends on adoption and on the process too. You need both layers of numbers.
2What should a metric agreement settle first, before any number is chosen?
Answer: C. The decision shapes everything else. If the numbers will decide rollout or stop, the primary metric, target and guardrails all follow from that. Without a decision, metrics become reporting nobody acts on.
3Decisions on blocked orders got 34% faster in the pilot month. Order volume also dropped that month. What is the fairest way to judge the assistant?
Answer: D. A control group in the same weeks sees the same season and volume. The difference between the two groups is what the assistant caused. Before-and-after numbers mix the assistant with everything else that changed.
4Analysts decide faster with the assistant, but overdue receivables on released orders start rising. Which part of the agreement should catch this?
Answer: A. Guardrails name what must not get worse while the main number improves. Here, faster releases could mean fewer checks. A guardrail on overdue items, with an owner and a limit, stops the rollout if it is breached.
5Your vendor quotes another SAP customer's 70% time saving for a similar feature. How should you use that figure?
Answer: C. Published figures describe other processes and starting points. They help you choose where to look. Your baseline and target must come from your own system data.
6The pilot shows a real 20% improvement, but the target was 25%. What should happen?
Answer: D. A smaller real gain may still be worth it. The agreement names who decides, so the outcome is a business decision, not an argument about the target after the fact.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A success metric is a test you write before the code. It has an input (the cases in a period), an expected result (the target), and failure conditions (the guardrails). Like a test, it is only useful if it is written down before anyone sees the answer.
The numbers form a chain. Each layer explains the one above it:
flowchart BT
Q[Model quality<br/>note right, grounded] --> A[Adoption<br/>note opened, rated useful]
A --> P[Process<br/>decided within 24 h]
P --> B[Business outcome<br/>median hours to decision]
G[Guardrails<br/>overdue after release, cost] -.must not get worse.-> B
Earlier Unit 8 topics built the bottom layer: LLM evaluation, groundedness and the evaluation harness. This topic adds the top: the number the sponsor pays for, measured fairly. In the first topic of the course, an FDE's job was one sentence: move metric M in process P from X to Y by date D. This is how you make that sentence measurable.
#Step one: write the metric as a definition, not a phrase
"Faster decisions" is a phrase. A definition says exactly what is counted:
Field
Example
Name
Median hours from block to decision
Formula
median(decided_at minus blocked_at), credit-blocked orders decided in the period
Unit and direction
hours; lower is better
Population
credit blocks only; one sales organization; orders decided in the period
Source
order block and release events from SAP, agreed with the SAP team
Owner
Head of Credit Management
The population line prevents most arguments later. Does a rejected order count? An order blocked twice? An order still blocked at the end of the period? Decide once, in writing.
Why the median? Decision times are skewed: most orders take hours, a few take weeks. The mean jumps when one stuck order is resolved. The median, the middle value, doesn't.
Microsoft's experimentation team recommends a core metric set with user satisfaction metrics, guardrail metrics, feature and engagement metrics, and data quality metrics. It describes guardrails as aspects you don't want to degrade but won't necessarily improve. Mapped to an SAP AI assistant:
Role
Purpose
Blocked-order example
Typical source
Primary
The one number the decision rests on
Median hours to decision
SAP order events
Business guardrail
Stops a "win" that harms the business
Released orders overdue after 30 days
Receivables aging
Quality guardrail
Stops a rollout of a tool people distrust
Notes rated useful at least 75%
Assistant feedback log
Cost guardrail
Keeps unit cost within the business case
Model cost per order at most $0.05
Provider usage export
Supporting
Explains movement
Notes opened; share decided within 24 hours
Assistant log; SAP events
The Goals-Signals-Metrics process from Google's HEART work gives the order to fill this table in: state the goal, name the signal that would show it, then pick the metric that measures the signal. Start from the goal, not from the data you happen to have.
One primary metric. With several equal metrics, some will improve by chance, and people will quote those. One primary metric plus explicit guardrails makes the decision rule clear.
Leading and lagging. Notes opened moves on day one. Overdue receivables take 30 days or more to show. A pilot shorter than the slowest guardrail can't check it, so plan the pilot length around the guardrail, not only the primary metric.
Take the last 8 to 12 weeks of data. Compute the metric for the whole period and for each week. The spread between weeks is the noise: how much the number moves with no change at all.
If your target is smaller than the weekly spread, a before-and-after chart can't show it. You will need a control group and enough cases.
Some metrics have no baseline. "Notes opened" doesn't exist before the assistant. That's fine: those are supporting metrics measured only in the pilot.
Before-and-after comparisons mix your change with everything else: season, volume, staffing, a new credit policy. The fix is a control group: analysts or orders without the assistant, in the same weeks.
How you split matters:
By analyst is usual for an assistant. Each analyst always or never has it. That avoids people copying habits between orders.
By order is fairer statistically, but one analyst then sees both versions and can learn from the notes.
By team or site is easiest to run but leaves fewer groups, so differences between teams can masquerade as an effect.
Microsoft's guidance also warns about sample ratio mismatch: if you planned a 50/50 split and get 60/40, something in the assignment is broken, and the result can't be trusted. Check group sizes before reading any result.
A pilot that is too small can't distinguish a real effect from noise. Microsoft's guidance says a power calculation sets the lower bound on how many units the test needs. For a median of skewed times, there is no simple formula, so simulate it: draw fake pilots from your baseline data, apply the target change to one group, and count how often the difference stands out from the noise.
The scorecard reports the pilot-versus-control change with a 95% interval, and each guardrail against its limit. The verdict follows a rule written before the pilot:
flowchart TD
S[Pilot results] --> G{Any guardrail<br/>breached?}
G -->|yes| STOP[Stop and investigate]
G -->|no| R{Interval shows<br/>a real improvement?}
R -->|no| N[No clear effect:<br/>extend or rethink]
R -->|yes| T{Point estimate<br/>meets target?}
T -->|yes| GO[Target met:<br/>recommend rollout]
T -->|no| D[Improved, below target:<br/>sponsor decides]
This is one simple rule, suitable for a pilot. Statisticians may prefer stricter versions. What matters is that it is agreed in advance.
#Build it yourself: a metric spec, a baseline and a pilot scorecard
You will write a metric spec for the blocked-order assistant, measure its baseline from order events, estimate how many orders a pilot needs, and score a pilot against a control group. The output is a one-page scorecard a sponsor can read.
Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder, its .venv and the unit08 folder. This script uses only Python's built-in modules, so there is nothing new to install.
flowchart LR
E[(Order events CSV)] --> B[baseline]
SP[Metric spec JSON] --> C[check-spec]
SP --> B
B --> Z[size]
E --> P[pilot]
SP --> P
P --> SC[scorecard.md]
In VS Code, right-click the unit08 folder, choose New File, name it success_metrics.py, paste the code below and save.
"""Unit 8: define success metrics with the business, measure a baseline, size a pilot and score it.
The use case is the blocked sales order assistant from earlier units. A metric spec (JSON) says what
"success" means: one primary business metric with a target, guardrails that must not get worse, and
supporting metrics. This script checks the spec, measures the baseline from order events, estimates
how many orders a pilot needs, and writes a scorecard comparing pilot orders with control orders.
How to run (from your course folder, with .venv turned on):
python unit08/success_metrics.py sample # write made-up events and a starter metric spec
python unit08/success_metrics.py check-spec # is every metric fully defined?
python unit08/success_metrics.py baseline # measure the "before" weeks
python unit08/success_metrics.py size # how many orders per group does the pilot need?
python unit08/success_metrics.py pilot # pilot vs. control, guardrails, verdict, scorecard.md
python unit08/success_metrics.py baseline --csv path/to/your_export.csv # use your own events
Standard library only. All data from "sample" is made up; codes and numbers are not SAP meanings.
"""
import argparse
import csv
import json
import random
import statistics
import sys
from datetime import datetime, timedelta
from pathlib import Path
HERE = Path(__file__).resolve().parent
EVENTS = HERE / "data" / "blocked_orders_events.csv"
SPEC = HERE / "metrics" / "blocked_orders_metrics.json"
RESULTS = HERE / "results"
BASELINE_OUT = RESULTS / "baseline_metrics.json"
SCORECARD = RESULTS / "scorecard.md"
COLUMNS = ["SalesOrder", "period", "group", "week", "blocked_at", "decided_at", "decision",
"TotalNetAmount", "note_opened", "note_rating", "overdue_30d", "model_cost_usd"]
REQUIRED = ["id", "name", "layer", "role", "direction", "unit", "formula", "source", "owner"]
LAYERS = {"business", "process", "adoption", "quality", "cost"}
ROLES = {"primary", "guardrail", "supporting"}
STARTER_SPEC = {
"use_case": "Blocked sales order assistant: draft notes that help credit analysts decide faster",
"decision": "Roll out to all credit analysts in the sales organization, or stop",
"sponsor": "Head of Credit Management",
"metrics": [
{"id": "hours_to_decision", "name": "Median hours from block to decision", "layer": "business",
"role": "primary", "direction": "down", "unit": "hours", "target_change_pct": -25,
"formula": "median(decided_at - blocked_at) over credit-blocked orders decided in the period",
"source": "order block and release events exported from SAP (agree the source with the SAP team)",
"owner": "Head of Credit Management"},
{"id": "decided_within_24h", "name": "Share decided within 24 hours", "layer": "process",
"role": "supporting", "direction": "up", "unit": "share",
"formula": "count(hours <= 24) / count(decided orders)",
"source": "same events as hours_to_decision", "owner": "Credit team lead"},
{"id": "overdue_after_release", "name": "Released orders overdue more than 30 days",
"layer": "business", "role": "guardrail", "direction": "down", "unit": "share",
"max_worsening": 0.01,
"formula": "count(released and overdue_30d) / count(released)",
"source": "receivables aging joined to released orders", "owner": "Head of Credit Management"},
{"id": "note_opened", "name": "Notes opened by the analyst", "layer": "adoption",
"role": "supporting", "direction": "up", "unit": "share", "target": 0.7,
"formula": "count(note_opened) / count(pilot orders)",
"source": "assistant usage log", "owner": "Product owner"},
{"id": "note_useful", "name": "Notes rated useful", "layer": "quality",
"role": "guardrail", "direction": "up", "unit": "share", "min_value": 0.75,
"formula": "count(rating = useful) / count(rated notes)",
"source": "thumbs up/down in the assistant", "owner": "Product owner"},
{"id": "cost_per_order", "name": "Model cost per order", "layer": "cost",
"role": "guardrail", "direction": "down", "unit": "usd", "max_value": 0.05,
"formula": "sum(model_cost_usd) / count(pilot orders)",
"source": "model provider usage export", "owner": "IT finance"},
],
}
# ---------------------------------------------------------------- made-up events
def make_sample(seed=8):
"""Eight 'before' weeks, then four pilot weeks split into control and pilot analysts."""
rng = random.Random(seed)
start = datetime(2026, 6, 1, 8, 0)
rows, number = [], 9200000
def add(period, group, week, median_hours, overdue_rate, assisted):
nonlocal number
number += 1
blocked = start + timedelta(days=7 * week + rng.uniform(0, 5), hours=rng.uniform(0, 9))
hours = rng.lognormvariate(0, 0.75) * median_hours
released = rng.random() < 0.88
opened = assisted and rng.random() < 0.78
rating = ""
if opened and rng.random() < 0.6:
rating = "useful" if rng.random() < 0.83 else "not useful"
rows.append({
"SalesOrder": str(number), "period": period, "group": group, "week": week + 1,
"blocked_at": blocked.isoformat(timespec="minutes"),
"decided_at": (blocked + timedelta(hours=hours)).isoformat(timespec="minutes"),
"decision": "released" if released else "rejected",
"TotalNetAmount": f"{rng.uniform(2000, 90000):.2f}",
"note_opened": "yes" if opened else ("no" if assisted else ""),
"note_rating": rating,
"overdue_30d": "yes" if released and rng.random() < overdue_rate else "no",
"model_cost_usd": f"{rng.uniform(0.01, 0.03):.4f}" if assisted else "0",
})
for week in range(8): # before: one team, no assistant
for _ in range(rng.randint(52, 64)):
add("before", "all", week, 20.0, 0.035, False)
for week in range(8, 12): # pilot weeks: quieter season for everyone
for _ in range(rng.randint(26, 32)):
add("pilot", "control", week, 17.0, 0.035, False)
for _ in range(rng.randint(26, 32)):
add("pilot", "pilot", week, 12.0, 0.035, True)
return rows
# ---------------------------------------------------------------- loading
def load_events(path):
if not path.exists():
sys.exit(f"No events file at {path}. Run: python unit08/success_metrics.py sample")
with open(path, newline="", encoding="utf-8") as f:
rows = list(csv.DictReader(f))
missing = [c for c in COLUMNS if rows and c not in rows[0]]
if not rows or missing:
sys.exit(f"{path} has no rows or lacks columns: {', '.join(missing) or 'all'}")
for r in rows:
r["hours"] = (datetime.fromisoformat(r["decided_at"])
- datetime.fromisoformat(r["blocked_at"])).total_seconds() / 3600
return rows
def load_spec():
if not SPEC.exists():
sys.exit(f"No metric spec at {SPEC}. Run: python unit08/success_metrics.py sample")
return json.loads(SPEC.read_text(encoding="utf-8"))
# ---------------------------------------------------------------- metric formulas
def share(rows, test, base=lambda r: True):
pool = [r for r in rows if base(r)]
return sum(1 for r in pool if test(r)) / len(pool) if pool else None
FORMULAS = {
"hours_to_decision": lambda rows: statistics.median(r["hours"] for r in rows) if rows else None,
"decided_within_24h": lambda rows: share(rows, lambda r: r["hours"] <= 24),
"overdue_after_release": lambda rows: share(rows, lambda r: r["overdue_30d"] == "yes",
lambda r: r["decision"] == "released"),
"note_opened": lambda rows: share(rows, lambda r: r["note_opened"] == "yes",
lambda r: r["note_opened"] != ""),
"note_useful": lambda rows: share(rows, lambda r: r["note_rating"] == "useful",
lambda r: r["note_rating"] != ""),
"cost_per_order": lambda rows: (sum(float(r["model_cost_usd"]) for r in rows) / len(rows)) if rows else None,
}
def fmt(value, unit):
if value is None:
return "n/a"
if unit == "share":
return f"{value:.1%}"
if unit == "usd":
return f"${value:.3f}"
return f"{value:.1f} h" if unit == "hours" else f"{value:.2f}"
# ---------------------------------------------------------------- commands
def cmd_sample(args):
EVENTS.parent.mkdir(parents=True, exist_ok=True)
SPEC.parent.mkdir(parents=True, exist_ok=True)
rows = make_sample()
with open(EVENTS, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=COLUMNS)
writer.writeheader()
writer.writerows(rows)
print(f"Wrote {len(rows)} made-up order events to unit08/data/{EVENTS.name}")
if SPEC.exists() and not args.force:
print(f"Kept your existing unit08/metrics/{SPEC.name} (add --force to overwrite it)")
else:
SPEC.write_text(json.dumps(STARTER_SPEC, indent=2) + "\n", encoding="utf-8")
print(f"Wrote the starter metric spec to unit08/metrics/{SPEC.name}")
def spec_problems(spec):
problems = []
for key in ("use_case", "decision", "sponsor"):
if not str(spec.get(key, "")).strip():
problems.append(f"spec: '{key}' is empty; say who decides what, based on these numbers")
metrics = spec.get("metrics", [])
for m in metrics:
label = m.get("id", "?")
for field in REQUIRED:
if not str(m.get(field, "")).strip():
problems.append(f"{label}: '{field}' is missing")
if m.get("layer") and m["layer"] not in LAYERS:
problems.append(f"{label}: layer must be one of {sorted(LAYERS)}")
if m.get("role") and m["role"] not in ROLES:
problems.append(f"{label}: role must be one of {sorted(ROLES)}")
if m.get("direction") not in ("up", "down"):
problems.append(f"{label}: direction must be 'up' or 'down'")
if m.get("id") and m["id"] not in FORMULAS:
problems.append(f"{label}: no formula in the script for this id (add one to FORMULAS)")
if m.get("role") == "primary" and "target_change_pct" not in m:
problems.append(f"{label}: the primary metric needs a target_change_pct")
if m.get("role") == "guardrail" and not any(k in m for k in ("max_worsening", "max_value", "min_value")):
problems.append(f"{label}: a guardrail needs max_worsening, max_value or min_value")
primaries = [m for m in metrics if m.get("role") == "primary"]
if len(primaries) != 1:
problems.append(f"spec: needs exactly one primary metric, found {len(primaries)}")
if not any(m.get("role") == "guardrail" and m.get("layer") == "business" for m in metrics):
problems.append("spec: add at least one business guardrail (what must not get worse)")
return problems
def cmd_check_spec(args):
spec = load_spec()
problems = spec_problems(spec)
for m in spec.get("metrics", []):
print(f" {m.get('role', '?'):<10} {m.get('layer', '?'):<9} {m.get('id', '?')}")
if problems:
print(f"\n{len(problems)} problem(s):")
for p in problems:
print(f" - {p}")
sys.exit(1)
print("\nSpec OK: one primary metric with a target, guardrails with limits, every metric owned.")
def cmd_baseline(args):
spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
before = [r for r in rows if r["period"] == "before"]
if not before:
sys.exit("No rows with period=before. The baseline needs the weeks before any change.")
weeks = sorted({int(r["week"]) for r in before})
print(f"Baseline: {len(before)} orders over {len(weeks)} weeks (period=before)\n")
print(f"{'metric':<24}{'baseline':>10}{'lowest wk':>11}{'highest wk':>12}")
out = {"orders": len(before), "weeks": len(weeks), "metrics": {}}
for m in spec["metrics"]:
value = FORMULAS[m["id"]](before)
weekly = [FORMULAS[m["id"]]([r for r in before if int(r["week"]) == w]) for w in weeks]
weekly = [v for v in weekly if v is not None]
out["metrics"][m["id"]] = value
if value is None:
print(f"{m['id']:<24}{'n/a':>10} (not measurable before the assistant exists)")
continue
print(f"{m['id']:<24}{fmt(value, m['unit']):>10}{fmt(min(weekly), m['unit']):>11}"
f"{fmt(max(weekly), m['unit']):>12}")
RESULTS.mkdir(parents=True, exist_ok=True)
BASELINE_OUT.write_text(json.dumps(out, indent=2) + "\n", encoding="utf-8")
print(f"\nSaved to unit08/results/{BASELINE_OUT.name}. Week-to-week spread is normal noise:"
"\na change smaller than that spread can't be seen in a before/after chart.")
def simulated_power(hours, change_pct, n, rng, sims=1000):
"""Share of simulated pilots of n orders per group whose median gap beats the noise level."""
factor = 1 + change_pct / 100
def gap(a, b):
return statistics.median(b) - statistics.median(a)
null = sorted(gap(rng.choices(hours, k=n), rng.choices(hours, k=n)) for _ in range(sims))
cutoff = null[int(0.025 * sims)] # 2.5% of no-effect pilots look this good by chance
hits = sum(gap(rng.choices(hours, k=n), [h * factor for h in rng.choices(hours, k=n)]) < cutoff
for _ in range(sims))
return hits / sims
def cmd_size(args):
spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
primary = next(m for m in spec["metrics"] if m["role"] == "primary")
hours = [r["hours"] for r in rows if r["period"] == "before"]
change = args.change if args.change is not None else primary["target_change_pct"]
rng = random.Random(42)
print(f"Orders per group needed to see a {change:+.0f}% change in {primary['id']}"
f" (from {len(hours)} baseline orders)\n")
print(f"{'orders per group':>17}{'chance to detect':>18}")
enough = None
for n in (30, 60, 100, 150, 200, 300, 400):
power = simulated_power(hours, change, n, rng)
print(f"{n:>17}{power:>18.0%}")
if enough is None and power >= 0.8:
enough = n
if enough:
print(f"\nAbout {enough} orders per group give an 80% chance to detect the change."
"\nFewer, and a real improvement can easily look like 'no clear effect'.")
else:
print("\nEven 400 orders per group give less than 80%. Plan a longer pilot or a bigger change.")
def bootstrap_relative(control, pilot, func, rng, reps=1000):
"""95% interval for (pilot - control) / control, resampling each group with replacement."""
values = []
for _ in range(reps):
c = func(rng.choices(control, k=len(control)))
p = func(rng.choices(pilot, k=len(pilot)))
if c:
values.append((p - c) / c)
values.sort()
return values[int(0.025 * len(values))], values[int(0.975 * len(values)) - 1]
def cmd_pilot(args):
spec, rows = load_spec(), load_events(Path(args.csv) if args.csv else EVENTS)
before = [r for r in rows if r["period"] == "before"]
control = [r for r in rows if r["period"] == "pilot" and r["group"] == "control"]
pilot = [r for r in rows if r["period"] == "pilot" and r["group"] == "pilot"]
if not control or not pilot:
sys.exit("Need rows with period=pilot and group=control and group=pilot.")
rng = random.Random(7)
primary = next(m for m in spec["metrics"] if m["role"] == "primary")
f = FORMULAS[primary["id"]]
b, c, p = f(before), f(control), f(pilot)
naive, controlled = (p - b) / b, (p - c) / c
low, high = bootstrap_relative(control, pilot, f, rng)
target = primary["target_change_pct"] / 100
better = (high < 0) if primary["direction"] == "down" else (low > 0)
meets = (controlled <= target) if primary["direction"] == "down" else (controlled >= target)
lines = [f"# Pilot scorecard: {spec['use_case']}", "",
f"Decision this informs: {spec['decision']} (sponsor: {spec['sponsor']})", "",
f"Orders: {len(before)} before, {len(control)} control, {len(pilot)} pilot", "",
"## Primary metric", "",
f"{primary['name']}: before {fmt(b, primary['unit'])}, control {fmt(c, primary['unit'])},"
f" pilot {fmt(p, primary['unit'])}", "",
f"- Before/after change (misleading, includes the season): {naive:+.0%}",
f"- Pilot vs. control, same weeks: {controlled:+.0%}"
f" (95% interval {low:+.0%} to {high:+.0%}); target {target:+.0%}", "",
"## Guardrails and supporting metrics", "",
"| Metric | Role | Control | Pilot | Limit | Status |", "| --- | --- | --- | --- | --- | --- |"]
breached = []
for m in spec["metrics"]:
if m is primary:
continue
fm = FORMULAS[m["id"]]
cv, pv = fm(control), fm(pilot)
status, limit = "info", ""
if "max_worsening" in m and cv is not None and pv is not None:
worse = (pv - cv) if m["direction"] == "down" else (cv - pv)
limit = (f"no worse than +{m['max_worsening'] * 100:.1f} points" if m["unit"] == "share"
else f"no worse than +{fmt(m['max_worsening'], m['unit'])}")
status = "BREACHED" if worse > m["max_worsening"] else "ok"
elif "max_value" in m and pv is not None:
limit, status = f"at most {fmt(m['max_value'], m['unit'])}", "BREACHED" if pv > m["max_value"] else "ok"
elif "min_value" in m and pv is not None:
limit, status = f"at least {fmt(m['min_value'], m['unit'])}", "BREACHED" if pv < m["min_value"] else "ok"
elif "target" in m and pv is not None:
limit, status = f"target {fmt(m['target'], m['unit'])}", "ok" if pv >= m["target"] else "below target"
if status == "BREACHED" and m["role"] == "guardrail":
breached.append(m["id"])
lines.append(f"| {m['name']} | {m['role']} | {fmt(cv, m['unit'])} | {fmt(pv, m['unit'])} | {limit} | {status} |")
if breached:
verdict = f"STOP: guardrail breached ({', '.join(breached)}). Investigate before any rollout."
elif better and meets:
verdict = "TARGET MET: the improvement is real and reaches the target. Recommend rollout."
elif better:
verdict = "IMPROVED, BELOW TARGET: real improvement, smaller than promised. Decide with the sponsor."
else:
verdict = "NO CLEAR EFFECT: the interval includes no change. Extend the pilot or rethink."
lines += ["", "## Verdict", "", verdict, ""]
RESULTS.mkdir(parents=True, exist_ok=True)
SCORECARD.write_text("\n".join(lines), encoding="utf-8")
print("\n".join(lines[4:]))
print(f"Saved to unit08/results/{SCORECARD.name}")
def main():
parser = argparse.ArgumentParser(description="Success metrics for the blocked-order assistant")
parser.add_argument("command", choices=["sample", "check-spec", "baseline", "size", "pilot"])
parser.add_argument("--csv", help="your own events file with the same columns")
parser.add_argument("--change", type=float, help="size: change to detect in percent, e.g. -15")
parser.add_argument("--force", action="store_true", help="sample: overwrite the metric spec")
args = parser.parse_args()
{"sample": cmd_sample, "check-spec": cmd_check_spec, "baseline": cmd_baseline,
"size": cmd_size, "pilot": cmd_pilot}[args.command](args)
if __name__ == "__main__":
main()
#Step 3: Write the sample events and the metric spec
In the terminal, run:
python unit08/success_metrics.py sample
You should see:
Wrote 688 made-up order events to unit08/data/blocked_orders_events.csv
Wrote the starter metric spec to unit08/metrics/blocked_orders_metrics.json
Open unit08/metrics/blocked_orders_metrics.json in VS Code. This is the metric agreement from the foundational layer, in a form a program can check. Each metric has a role (primary, guardrail or supporting), a layer, a direction, a formula in words, a source and an owner.
Open unit08/data/blocked_orders_events.csv. Each row is one credit-blocked order: when it was blocked, when it was decided, whether it was released, and for pilot orders whether the analyst opened and rated the note. The period column says before or pilot; the group column says all, control or pilot.
If you run sample again, it rewrites the events but keeps your edited spec. Add --force to reset the spec too.
Spec OK: one primary metric with a target, guardrails with limits, every metric owned.
Now break it on purpose. In the JSON file, change the first metric's "owner" to "" and save. Run check-spec again. You should see:
1 problem(s):
- hours_to_decision: 'owner' is missing
The script ends with exit code 1, so you could run it in CI like the evaluation harness. Put the owner back and save.
The checks encode the agreement rules: exactly one primary metric with a target, at least one business guardrail, a limit on every guardrail, and an owner on every metric.
Baseline: 459 orders over 8 weeks (period=before)
metric baseline lowest wk highest wk
hours_to_decision 19.7 h 18.2 h 22.8 h
decided_within_24h 59.9% 52.7% 63.9%
overdue_after_release 2.3% 0.0% 6.2%
note_opened n/a (not measurable before the assistant exists)
note_useful n/a (not measurable before the assistant exists)
cost_per_order $0.000 $0.000 $0.000
Read the spread. The median decision time moved between 18.2 and 22.8 hours from week to week, with no change at all. The overdue rate swung between 0% and 6.2%, because only a few released orders go overdue each week. A one-week before-and-after comparison of either number would mean nothing.
n/a is a valid result here: adoption and rating metrics don't exist before the assistant. The baseline is saved in unit08/results/baseline_metrics.json.
Orders per group needed to see a -25% change in hours_to_decision (from 459 baseline orders)
orders per group chance to detect
30 18%
60 30%
100 38%
150 59%
200 83%
300 94%
400 99%
About 200 orders per group give an 80% chance to detect the change.
Try a bigger effect: python unit08/success_metrics.py size --change -40. The script says about 100 orders per group are enough. Bigger effects need smaller pilots.
This is the conversation to have with the sponsor before the pilot. With about 30 credit blocks per group per week, 200 orders per group means about 7 weeks, not 4.
Orders: 459 before, 112 control, 117 pilot
## Primary metric
Median hours from block to decision: before 19.7 h, control 17.0 h, pilot 13.0 h
- Before/after change (misleading, includes the season): -34%
- Pilot vs. control, same weeks: -24% (95% interval -46% to -5%); target -25%
| Metric | Role | Control | Pilot | Limit | Status |
| --- | --- | --- | --- | --- | --- |
| Released orders overdue more than 30 days | guardrail | 4.0% | 4.7% | no worse than +1.0 points | ok |
| Notes rated useful | guardrail | n/a | 80.7% | at least 75.0% | ok |
| Model cost per order | guardrail | $0.000 | $0.021 | at most $0.050 | ok |
## Verdict
IMPROVED, BELOW TARGET: real improvement, smaller than promised. Decide with the sponsor.
Open unit08/results/scorecard.md. VS Code shows a preview with the Open Preview to the Side button at the top right.
Read the result the way a sponsor would:
Before and after says 34%. But the control group also got faster, from 19.7 to 17.0 hours, because the made-up pilot weeks were a quieter season. The fair number is 24%.
The interval is wide, from 46% to 5% faster. The improvement is real (the whole range is faster), but the pilot had about 115 orders per group, while Step 6 said about 200 were needed. The point estimate just misses the 25% target.
No guardrail was breached. The overdue rate is 0.7 points higher in the pilot, inside the 1-point limit, but a 4-week pilot is short for a 30-day lagging metric.
The honest recommendation is to extend the pilot to the planned size before deciding, or to let the sponsor accept a real 24% gain. Either way, the decision is the sponsor's, and the scorecard says so.
Columns the baseline doesn't use, such as note_opened, can be empty. Keep the export to the fields the metrics need, and treat it as business data: don't commit it to a public repository.
You rarely need to invent the business metric. In SAP landscapes, process numbers often already exist. As of October 2026, these are the places to look before building your own baseline.
SAP Learning's lesson on SAP Smart Business for sales order fulfillment describes:
The Sales Order Fulfillment Issues app, which lists blocked sales orders for analysis and processing. Issue types include incomplete data, unconfirmed quantities, billing blocks and delivery blocks.
KPI visualizations with semantic colors based on defined targets and thresholds.
The Manage KPIs and Reports app, available since SAP S/4HANA 1909, for configuring new KPIs. The lesson names the business role SAP_BR_ANALYTICS_SPECIALIST for it.
For the blocked-order assistant, that means the people who own blocked orders may already watch a KPI on them. Use their definition as your primary metric if it fits. A metric the business already reports is easier to trust than one you invent.
SAP Signavio Process Intelligence computes metrics on process event data. SAP Learning describes a metric library of predefined calculations written as SiGNAL expressions, which can be reused to calculate KPIs. Every process gets an average cycle time metric by default. One practical warning from the same lesson: process views control data access, and a metric can become invalid if the view hides data it needs. Check that your baseline isn't silently filtered.
SAP Signavio Process Insights comes with prebuilt content. SAP Learning lists more than 200 process performance indicators, with benchmarking inside the company and against industry peers. Standard indicator types include throughput, backlog, exceptions and automation rate. Backlog and automation rate are natural guardrail or supporting metrics for an AI assistant.
SAP Discovery Center hosts the SAP Business AI Catalog. SAP News reported almost 400 features and agents in it in June 2026, with cost and ROI estimator tools. In August 2025, SAP said its value estimates combine efficiency and effectiveness, and rest on SAP's industry expertise, benchmarking data, third-party research and early customer feedback.
Use these the way you used SAP's feature tables in Where AI creates value in S/4HANA: to choose candidates and frame the target conversation. Don't paste them into your metric agreement as a baseline or a promise.
The same method works for SAP's own AI features, not only custom builds. If you switch on a standard feature for one team first, that team is the pilot group and the others are the control. The agreement, the baseline and the guardrails are the same; only the system under test changes.
Authorizations and data protection. Order events carry customer names and amounts. Export only the fields the metrics need, under the same access rules as the source, and keep exports out of public repositories. SAP Signavio's process views control which data a metric can see; respect them rather than working around them.
Metric definitions are versioned. Keep the spec in Git with the code. If the definition changes mid-pilot, results before and after the change can't be compared.
Data quality first. Check group sizes against the planned split, missing timestamps, and orders decided before they were blocked. Microsoft's guidance treats a sample ratio mismatch as a sign the result can't be trusted.
Lagging guardrails. Keep measuring guardrails such as overdue receivables after the pilot ends, until their lag has passed.
Cost per unit. Report model cost per order next to the business gain. Unit 10 turns that into a business case.
Ownership after go-live. The metric needs an owner after the project team leaves. Put it on the process owner's existing KPI report if possible.
Clean core. Measurement reads data through exports, released APIs or process mining. It needs no modification to SAP standard.
Add a guardrail against rubber-stamping: analysts might decide faster simply by releasing more orders. Then tighten it to see a guardrail stop the rollout.
Open unit08/success_metrics.py. Find the line that starts with "note_opened": lambda rows: inside FORMULAS. Just above it, at the same indentation, add:
Run python unit08/success_metrics.py check-spec. You should see seven metrics and Spec OK.
Run python unit08/success_metrics.py pilot. A new row shows the share released: 89.3% for control, 91.5% for pilot, status ok.
Change "max_worsening": 0.03 to 0.02 and save. Run pilot again. The row now says BREACHED and the verdict says STOP: guardrail breached (released_share).
Put the limit back to 0.03. In unit08/results/, create metric_agreement.md with five lines: the decision, the primary metric with baseline and target, your guardrails with limits, the comparison group, and the pilot size from Step 6.
Save your work:
git add unit08/success_metrics.py unit08/metrics/blocked_orders_metrics.json unit08/results/metric_agreement.md
git commit -m "Add success metrics, baseline and pilot scorecard"
Done whencheck-spec passes with seven metrics, the 0.02 limit produced a STOP verdict, and metric_agreement.md holds your five lines. Keep it: Unit 10's business case starts from this agreement and the scorecard.
Pick one answer for each question. The explanation appears after you choose.
1Why does the script use the median of decision hours rather than the mean?
Answer: B. Decision times are skewed: most take hours, a few take weeks. The mean jumps when one stuck order closes; the median, the middle value, stays stable, so changes reflect typical orders.
2The baseline shows weekly median decision times between 18.2 and 22.8 hours with no change. What does that spread tell you?
Answer: C. The spread is noise, how much the number moves on its own. An effect smaller than that range disappears in a simple before-and-after view, which is why the pilot needs a control group and enough orders.
3In the sample pilot, before-and-after shows -34% but pilot vs. control shows -24%. What explains the gap?
Answer: A. Both groups worked in the same quieter weeks, and the control group went from 19.7 to 17.0 hours without the assistant. Comparing pilot with control in the same weeks removes that shared change; before-and-after counts it as the assistant's effect.
4The size command says about 200 orders per group give an 80% chance to detect a 25% change. The pilot had about 115 per group. What is the right reading?
Answer: D. Power says how likely a pilot of a given size is to detect a real effect. With fewer orders than planned, the interval widens, here from 46% to 5% faster, so the point estimate can easily miss the target by chance.
5You plan a 4-week pilot. One guardrail is "released orders overdue after 30 days". What should you do?
Answer: C. Overdue receivables are a lagging metric; harm from careless releases shows up weeks later. Dropping or swapping it removes the protection it gives, so keep tracking it past the pilot before final rollout.
6A colleague adds a new metric id to the JSON spec, and check-spec fails with "no formula in the script for this id". Why does the script insist on this?
Answer: B. The spec's formula line is words; FORMULAS is the code that computes it. A metric with no code would appear in the agreement but never in the scorecard, so the check catches the gap before the pilot.
7Your company runs SAP Signavio Process Intelligence, and its average cycle time for credit blocks looks far lower than your export. What should you check first?
Answer: D. SAP Learning warns that process views control data access, and a metric can become invalid if the view restricts data it needs. A filtered baseline can look better than reality, so confirm the population before you trust either number.
8Your assistant cuts decision time by 30% in a well-sized pilot, but the share of orders released rises 4 points against a 3-point limit. What should the scorecard verdict be?
Answer: C. The decision rule checks guardrails first. A breached guardrail stops the rollout whatever the primary metric says, because faster decisions may come from releasing orders without proper checks.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Using SAP Smart Business for Sales Order Fulfillment (SAP Learning)— Sales Order Fulfillment Issues app for blocked orders (billing blocks, delivery blocks, incomplete data); Smart Business KPIs with targets and thresholds; Manage KPIs and Reports app since S/4HANA 1909
The Business Value of SAP AI Use Cases (SAP News, August 2025)— catalog of AI use cases with business benefits and estimated value, based on SAP expertise, benchmarking, third-party research and early customer feedback; efficiency and effectiveness view