Orchestrate

Classification and the metrics that matter

Why accuracy misleads, how precision, recall and the decision threshold trade off, and how to pick a metric and threshold from business costs.

Updated Sep 30, 2026Foundational 9 minDeep 38 min
Foundational layer · 9 min read

The 60-second version

Most business predictions are classification: pick one outcome from a short list. Will this delivery be late? Is this invoice a duplicate? Which cost center does this line belong to?

A classifier doesn't really answer "yes" or "no". It gives each case a score, say 0.72 for "late". A person or a system then applies a threshold: flag everything scoring 0.5 or more, for example. Changing that threshold changes which cases get flagged, without retraining anything.

Once the flags are out, there are only four things that can happen to each case:

Really late Really on time
Flagged Caught it (true positive) False alarm (false positive)
Not flagged Missed it (false negative) Correctly left alone (true negative)

That table is the confusion matrix. Every metric in this topic is a way of summarizing it. The one idea to keep: the right metric is the one that matches what a false alarm and a miss cost your business. Accuracy, the metric most people quote, is usually the wrong one.

Why it matters to the business

Take order-to-cash. A planner can check only so many deliveries a day. You want a model that flags deliveries at risk of being late so the planner can act: switch carrier, split the shipment, warn the customer.

Suppose 15 in every 100 deliveries end up late. A "model" that never flags anything is right 85% of the time. That is 85% accuracy, and it catches zero late deliveries. In the deep layer of this topic, an actual trained tree scores 86% accuracy on made-up data like this, barely above doing nothing. Google's Machine Learning Crash Course warns about exactly this: accuracy can mislead when one outcome is much rarer than the other.

So you need metrics that talk about the rare outcome directly:

  • Recall: of all the deliveries that really were late, what share did we flag? A low recall means late deliveries slip through.
  • Precision: of the deliveries we flagged, what share really were late? A low precision means planners waste time on false alarms.

The two pull against each other. Lower the threshold and you catch more late deliveries (recall up), but you also flag more on-time ones (precision down). Google's course describes this as an inverse relationship driven by the threshold.

The good news is that the business can settle the trade-off with money. In the deep layer's example, a missed late delivery costs 250 EUR and a planner check costs 50 EUR. At the default threshold of 0.5 the model's flags cost 25,600 EUR on 1,000 test deliveries. At a threshold chosen from those two costs, about 0.2, the cost drops to about 24,000 EUR, even though accuracy falls from 0.89 to 0.81. Lower accuracy, lower cost. That is the whole lesson in one line.

How SAP does it

As of September 2026, the same metrics appear wherever SAP runs classification:

  • Embedded in S/4HANA, through ISLM. SAP Learning describes Intelligent Scenario Lifecycle Management (ISLM) training and activating customer-specific models. In S/4HANA Utilities, for example, a model computes a release confidence value for implausible meter readings and outsorted billing documents, and clerks see it as a column. That value is a score. Somebody still decides at what value a document is released without a human look: a threshold.
  • In SAP HANA, through PAL and the hana-ml client. An SAP blog on the UnifiedClassification class of the hana-ml Python client shows it reporting AUC, recall, precision, F1 score, accuracy and two more agreement measures (kappa and MCC), plus a confusion matrix and an HTML model report. The vocabulary in this topic is the vocabulary of those reports.
  • Pretrained for tables: SAP-RPT-1. SAP Learning lists classification tasks such as cost centers, payment terms and sales groups. It returns a confidence score between 0.0 and 1.0 for each predicted value. Whether 0.8 really means "right about 80% of the time" on your data is something you check on your own held-back rows, as described below.

Whichever route you take, SAP gives you scores and standard metrics. Choosing the metric and the threshold is still your team's job, because only the business knows what a miss costs.

Choosing the metric: a decision guide

Your situation Watch this Why
Misses are expensive (late delivery, fraud, a duplicate payment going out) Recall You want to catch as many real cases as possible
False alarms are expensive (a blocked order, an angry customer, a clerk's hour) Precision Every flag should be worth acting on
Both matter and you need one number F1 score Balances precision and recall
You want to compare models before anyone picks a threshold ROC AUC or average precision Measures how well the scores rank real cases above the rest
The score itself drives a decision, such as a priority queue or a price Calibration A score of 0.3 should mean about 30% of such cases turn out positive
You know the cost of each kind of error Total cost Converts the confusion matrix straight into euros
The outcomes are roughly balanced and errors cost about the same Accuracy is acceptable The one case where it doesn't mislead

A practical rule: agree the cost of a miss and the cost of a false alarm with the process owner before the project starts. Everything else follows from those two numbers.

Questions to ask

  • What share of cases is positive (late, duplicate, fraudulent)? What accuracy does "always say no" get?
  • Which is worse for this process: a miss or a false alarm? Roughly what does each cost?
  • What threshold are you using, and how was it chosen? Was it chosen without looking at the final test set?
  • How many cases per day will be flagged at that threshold? Can the team handle that volume?
  • What recall and precision do you get at that threshold, on data the model never saw?
  • If we act on the score directly, is it calibrated? When it says 0.8, are about 80% of those cases really positive?
  • How will we notice when the metrics drift after go-live?

Common misconceptions

  • "95% accuracy means the model works." If 95% of cases are negative, a model that always says "no" gets 95%. Ask for recall and precision on the rare outcome.
  • "The threshold is 0.5, because that's the default." Libraries use 0.5 by default. The right threshold depends on your costs and is often far from 0.5.
  • "A better model always has higher accuracy." In the deep layer's example, the cheapest set of flags has lower accuracy than the default one.
  • "A confidence score is a probability." Some models' scores are well calibrated, many are not. Check before using a score as a percentage.
  • "One number is enough." Precision without recall hides misses; recall without precision hides noise. Report both, plus the flag volume.
  • "Once chosen, the threshold is fixed." When costs, volumes or data change, the best threshold moves too.

Key terms

  • Classification: predicting one category from a fixed set, such as late or on time.
  • Score: the model's number for "how likely is the positive outcome". Often between 0 and 1.
  • Threshold: the cut-off score at which a case is flagged.
  • Confusion matrix: the table of caught, false alarms, missed and correctly left alone.
  • Accuracy: share of all cases the model got right.
  • Recall: share of real positive cases that were flagged.
  • Precision: share of flagged cases that were really positive.
  • F1 score: one number that balances precision and recall.
  • ROC AUC: how well the scores rank positives above negatives; 0.5 is random, 1.0 perfect.
  • Calibration: whether a score of 0.7 really means about 70% of such cases are positive.
  • Class imbalance: one outcome is much rarer than the other.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 115% of deliveries are late. A vendor reports 86% accuracy for their late-delivery model. What should you ask first?

    Answer: B. With 15% late, "always on time" is right 85% of the time and catches nothing. Accuracy hides that. Recall and precision describe the rare outcome you actually care about.
  2. 2A model gives each delivery a score. What does lowering the flagging threshold from 0.5 to 0.2 usually do?

    Answer: C. A lower threshold flags more deliveries. More real late ones get caught, so recall rises, but more on-time ones get flagged too, so precision falls. The model itself is unchanged.
  3. 3In accounts payable, a duplicate payment that goes out is costly, and a clerk's check is cheap. Which metric should lead?

    Answer: D. When misses cost far more than false alarms, you want to catch as many real cases as possible. That is recall. Precision would lead if false alarms were the expensive error.
  4. 4In the example, a threshold chosen from business costs lowered accuracy from 0.89 to 0.81. Why can that be the better choice?

    Answer: B. Accuracy treats a miss and a false alarm as equally bad. Here a miss costs 250 EUR and a check 50 EUR, so flagging more deliveries lowered the total cost even though more flags were wrong.
  5. 5SAP-RPT-1 returns a confidence score of 0.8 for a predicted payment term. What should your team do before treating it as "80% sure"?

    Answer: D. A score between 0 and 1 is not automatically calibrated. Comparing scores with how often predictions were actually right, on your own data, tells you whether you can read it as a percentage.
  6. 6Which two numbers should you agree with the process owner before choosing a threshold?

    Answer: A. The threshold is a business decision about which error to accept more of. Once the two error costs are known, the threshold can be chosen to minimize the total, and the rest of the metrics follow.
Deep layer · 38 min read

Mental model: the model ranks, the business decides

A classifier does two separate jobs, and most confusion about metrics comes from mixing them up. scikit-learn's own guide makes this split explicit: first a statistical problem, predicting a probability or score; then a decision problem, turning that score into an action.

flowchart LR
  X[Delivery features] --> M[Model]
  M --> S[Score 0 to 1]
  S --> T{Score >= threshold?}
  T -->|yes| F[Flag for planner]
  T -->|no| N[Leave alone]

That gives three families of metrics, each answering a different question:

Question Metrics Depends on the threshold?
Are the scores in the right order? ROC AUC, average precision No
Are the scores honest numbers? Calibration curve, Brier score, log loss No
Are the decisions good? Confusion matrix, precision, recall, F1, total cost Yes

Pick a model on the first two. Pick a threshold on the third, using costs. Never judge the first two by the third.

How it works

Logistic regression: the workhorse classifier

This topic uses logistic regression, the classifier promised at the end of Gradient descent, by hand. It computes the same weighted sum as the freight-cost line, w1*x1 + w2*x2 + ... + b, then squeezes it through the sigmoid function 1 / (1 + e^-z) so the result lands between 0 and 1. Training uses a gradient-based method to adjust the weights. The loss is log loss, which punishes confident wrong scores heavily.

In scikit-learn, LogisticRegression uses the lbfgs solver by default, with max_iter defaulting to 100. Its predict_proba returns one column per class; column 1 holds the score for "late". The script also offers SGDClassifier(loss="log_loss"), which is the plain stochastic gradient descent you wrote yourself, so you can reuse the learning rate from your notes.

The confusion matrix and its metrics

With the counts TP (caught), FP (false alarm), FN (missed) and TN (correctly left alone), the metrics are simple ratios. The definitions below follow Google's crash course and scikit-learn's guide.

Metric Formula Plain question
Accuracy (TP + TN) / all How often was the flag right?
Recall (true positive rate) TP / (TP + FN) Of the late deliveries, how many did we flag?
Precision TP / (TP + FP) Of the flags, how many were late?
False positive rate FP / (FP + TN) Of the on-time deliveries, how many did we flag?
F1 2 × precision × recall / (precision + recall) One number balancing the two
Balanced accuracy (recall + TN / (TN + FP)) / 2 Accuracy that weighs both classes equally

In code, sklearn.metrics.confusion_matrix(y_true, y_pred).ravel() returns the four counts in the order tn, fp, fn, tp. Getting that order wrong is a classic bug; the script names them explicitly.

Why accuracy fails on imbalanced classes

Accuracy counts every row equally. When 85% of rows are "on time", the on-time rows dominate the number. A model can ignore the late class entirely and still look good. Google's crash course recommends F1 over accuracy for imbalanced data, and scikit-learn's guide points to balanced accuracy for the same reason.

The fix is not a cleverer average. It is to look at the rare class directly (recall, precision) and to always print the baseline next to the model, as in Machine learning in plain terms.

Ranking metrics: ROC AUC and average precision

Before any threshold is chosen, you can ask how well the scores rank late deliveries above on-time ones.

  • The ROC curve plots recall (true positive rate) against the false positive rate at every possible threshold. Google's crash course gives the key reading of the area under it: AUC is the probability that the model scores a randomly chosen positive higher than a randomly chosen negative. 0.5 is a coin flip, 1.0 is perfect ranking.
  • The precision-recall curve plots precision against recall at every threshold. Its summary, average precision, starts at the share of positives for random scores (0.15 in our data), not at 0.5. Google's course notes that precision-recall curves can give a better comparison when classes are imbalanced, because they ignore the large pile of easy true negatives.

Use these to compare models. Don't report them to a process owner as "the model is 88% good". AUC is a ranking probability, not a hit rate.

Choosing the threshold from costs

scikit-learn's predict uses a threshold of 0.5 on predict_proba by default. Its user guide states plainly that this default is often not what the task needs.

Suppose each flagged delivery costs a planner check (c_check), and a check always prevents the loss. Each missed late delivery costs c_miss. For one delivery with a well-calibrated score p:

  • Flag it: cost c_check, whatever the outcome.
  • Don't flag it: expected cost p × c_miss.

Flag when p × c_miss ≥ c_check, which means a threshold of c_check / c_miss. With 50 EUR and 250 EUR, that is 0.2. This break-even rule only holds when the scores are calibrated and the cost model is this simple. In practice you let cross-validation search for the threshold with the real cost function, and use the formula as a sanity check.

scikit-learn does that search with TunedThresholdClassifierCV. By default it maximizes balanced accuracy using 5-fold stratified cross-validation; you pass your own score with make_scorer. Its guide warns against tuning the threshold on the same data used to train the model. FixedThresholdClassifier wraps a model with a threshold you set by hand.

flowchart LR
  TR[Training rows] --> CV[5 folds: train on 4, score thresholds on 1]
  CV --> BT[Threshold with lowest average cost]
  BT --> RF[Refit model on all training rows]
  RF --> TE[Report once on test rows]

Calibration

A score is well calibrated when, among cases scored around 0.8, about 80% are really positive. That is scikit-learn's definition. Ranking can be excellent while calibration is poor: multiply every score by 0.5 and the order, and so the AUC, doesn't change.

You check it with a calibration curve (reliability diagram): sort test rows by score into bins, then compare each bin's average score with its actual share of positives. On a perfect model the points lie on the diagonal. scikit-learn's guide notes that LogisticRegression tends to be well calibrated out of the box, while models such as random forests and naive Bayes are not. When calibration matters, CalibratedClassifierCV fits a correction: method="sigmoid" for small data, or method="isotonic", which the guide warns can overfit below about 1,000 rows.

Two summary numbers measure scores directly: the Brier score (mean squared difference between score and outcome) and log loss. Lower is better for both. scikit-learn calls them strictly proper scoring rules. They mix calibration with ranking quality, so read them with the calibration curve, not instead of it.

Class weights and resampling

LogisticRegression(class_weight="balanced") weights each class by n_samples / (n_classes × count of that class), so rare late rows count more in the loss. That usually raises recall at the default threshold. It also pushes every score upward, so the scores are no longer calibrated. The script lets you see this. Resampling tricks, such as creating synthetic minority rows, have the same side effect.

Often the simpler route gives the same result: train without weights, keep calibrated scores, and move the threshold. Either way, re-tune the threshold after changing the weights.

Build it yourself: judge a late-delivery classifier by cost

You will train a logistic regression and a small tree to flag late deliveries on 4,000 made-up SAP-shaped deliveries, of which about 15% are late. You will print the confusion matrix and every metric above, see accuracy mislead, pick a threshold from business costs with cross-validation, and check calibration. One script, about 40 minutes.

Before you start: complete Set up your computer for this course and Set up for Unit 2: data science tools. They give you the orchestrate-course folder, its .venv, the unit02 folder, and numpy, pandas, scikit-learn and matplotlib. This walkthrough doesn't repeat those steps.

flowchart LR
  G[4,000 made-up deliveries] --> S[Split 75/25, stratified]
  S --> M[Train: baseline, tree, logistic regression]
  M --> C[Confusion matrix and metrics at one threshold]
  M --> R[ROC AUC, average precision]
  M --> T[Threshold chosen by cost, with cross-validation]
  M --> K[Calibration check]

What you need

  • The course folder, .venv and unit02 folder from the two setup topics.
  • About 40 minutes. No accounts, no API keys, no cost.
  • scikit-learn 1.5 or newer, for TunedThresholdClassifierCV. The Unit 2 setup installs a current version.
  • All data is made up. Nothing is sent anywhere, and the script needs no internet.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 2 folder:

    cd unit02
  5. Check your scikit-learn version:

    python -c "import sklearn; print(sklearn.__version__)"

    Anything from 1.5 upward is fine. If it's older, run pip install --upgrade scikit-learn.

Step 2: Save the script

  1. In VS Code's file list, right-click unit02, choose New File and name it classify_late.py.
  2. Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""Classification and the metrics that matter: flag late deliveries, then judge the flags honestly.

How to run (from the unit02 folder, with the course .venv turned on):
    python classify_late.py                          # 4,000 made-up deliveries, about 1 in 7 late
    python classify_late.py --threshold 0.3          # flag more deliveries: recall up, precision down
    python classify_late.py --cost-miss 800 --cost-check 25    # change the business costs
    python classify_late.py --balanced               # weight the rare late class more in training
    python classify_late.py --sgd --lr 0.01          # train with SGD and your learning rate from the last topic
    python classify_late.py --plot                   # also save classification_metrics.png

All data is made up. Nothing is sent anywhere.
"""
import argparse
from pathlib import Path

import numpy as np
import pandas as pd
from sklearn.calibration import calibration_curve
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression, SGDClassifier
from sklearn.metrics import (average_precision_score, brier_score_loss, confusion_matrix,
                             make_scorer, roc_auc_score)
from sklearn.model_selection import TunedThresholdClassifierCV, train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier

HERE = Path(__file__).parent
PLANTS = ["1010", "1020", "1710"]           # made-up plant codes
CARRIERS = ["ROAD", "RAIL", "EXPRESS"]      # made-up shipping types


def make_deliveries(rows: int, seed: int = 42) -> pd.DataFrame:
    """Same columns as rules_vs_learning.py, but late deliveries are rarer: about 1 in 7."""
    rng = np.random.default_rng(seed)
    df = pd.DataFrame({
        "Plant": rng.choice(PLANTS, size=rows),
        "ShippingType": rng.choice(CARRIERS, size=rows, p=[0.6, 0.3, 0.1]),
        "Items": rng.integers(1, 12, size=rows),
        "WeightKg": rng.gamma(2.0, 150.0, size=rows).round(1),
        "LeadTimeDays": rng.integers(1, 15, size=rows),
    })
    risk = (
        0.25 * df["Items"]
        + 0.004 * df["WeightKg"]
        - 0.35 * df["LeadTimeDays"]
        + np.where(df["Plant"] == "1710", 1.0, 0.0)
        + np.where(df["ShippingType"] == "EXPRESS", -1.5, 0.0)
        - 3.0                                   # shifts the mix so fewer deliveries are late
        + rng.logistic(0, 1.0, size=rows)
    )
    df["Late"] = (risk > 0).astype(int)
    return df


def counts(y_true, flagged):
    """The four cells of the confusion matrix: tn, fp, fn, tp."""
    tn, fp, fn, tp = confusion_matrix(y_true, flagged, labels=[0, 1]).ravel()
    return int(tn), int(fp), int(fn), int(tp)


def safe_div(a, b):
    return a / b if b else 0.0


def cost_of(y_true, flagged, cost_miss, cost_check):
    """Business cost: every flagged delivery gets a planner check; every missed late one costs cost_miss."""
    _, fp, fn, tp = counts(y_true, flagged)
    return (tp + fp) * cost_check + fn * cost_miss


def negative_cost_per_delivery(y_true, flagged, cost_miss, cost_check):
    """Scikit-learn maximizes scores, so return minus the average cost."""
    return -cost_of(y_true, flagged, cost_miss, cost_check) / len(y_true)


def report(name, y_true, flagged, cost_miss, cost_check):
    """One line per model: the confusion matrix cells plus the metrics built from them."""
    tn, fp, fn, tp = counts(y_true, flagged)
    accuracy = (tp + tn) / len(y_true)
    precision = safe_div(tp, tp + fp)
    recall = safe_div(tp, tp + fn)
    f1 = safe_div(2 * precision * recall, precision + recall)
    cost = cost_of(y_true, flagged, cost_miss, cost_check)
    print(f"  {name:<38}{tp:>5}{fp:>5}{fn:>5}{tn:>6}   {accuracy:.2f}   {precision:.2f}   "
          f"{recall:.2f}  {f1:.2f}  {cost:>8,.0f}")


def main() -> None:
    parser = argparse.ArgumentParser(description="Train a late-delivery classifier and judge it with the right metrics.")
    parser.add_argument("--rows", type=int, default=4000, help="how many made-up deliveries")
    parser.add_argument("--threshold", type=float, default=0.5, help="flag a delivery when its score is at least this")
    parser.add_argument("--cost-miss", type=float, default=250.0, help="EUR lost when a late delivery is not flagged")
    parser.add_argument("--cost-check", type=float, default=50.0, help="EUR a planner spends checking one flagged delivery")
    parser.add_argument("--balanced", action="store_true", help="use class_weight='balanced' in training")
    parser.add_argument("--sgd", action="store_true", help="train logistic regression with SGDClassifier instead of lbfgs")
    parser.add_argument("--lr", type=float, default=0.01, help="learning rate for --sgd")
    parser.add_argument("--plot", action="store_true", help="save classification_metrics.png")
    args = parser.parse_args()
    if not 0.0 < args.threshold < 1.0:
        parser.error("--threshold must be between 0 and 1, for example 0.3")

    df = make_deliveries(args.rows)
    label = df["Late"]
    features = pd.get_dummies(df.drop(columns=["Late"]), dtype=int)  # text columns become 0/1 columns

    # Hold back a quarter as the test set. stratify keeps the same share of late rows on both sides.
    x_train, x_test, y_train, y_test = train_test_split(
        features, label, test_size=0.25, random_state=0, stratify=label)
    print(f"{len(df)} deliveries, {label.mean():.0%} late. "
          f"{len(x_train)} to learn from, {len(x_test)} held back for testing.")
    print(f"Costs: a missed late delivery {args.cost_miss:,.0f} EUR, a planner check {args.cost_check:,.0f} EUR\n")

    # The models. Logistic regression outputs a score between 0 and 1 for "late".
    weight = "balanced" if args.balanced else None
    if args.sgd:
        learner = SGDClassifier(loss="log_loss", learning_rate="constant", eta0=args.lr,
                                class_weight=weight, max_iter=1000, random_state=0)
        logreg_name = f"Logistic regression, SGD lr {args.lr}"
    else:
        learner = LogisticRegression(max_iter=1000, class_weight=weight)
        logreg_name = "Logistic regression"
    logreg = make_pipeline(StandardScaler(), learner).fit(x_train, y_train)
    tree = DecisionTreeClassifier(max_depth=5, class_weight=weight, random_state=0).fit(x_train, y_train)
    baseline = DummyClassifier(strategy="most_frequent").fit(x_train, y_train)

    p_logreg = logreg.predict_proba(x_test)[:, 1]   # column 1 = chance of class 1, "late"
    p_tree = tree.predict_proba(x_test)[:, 1]

    # 1. The confusion matrix and the metrics built from it, at one threshold.
    t = args.threshold
    print(f"At threshold {t}: a delivery is flagged when its score is at least {t}")
    print(f"  {'model':<38}{'TP':>5}{'FP':>5}{'FN':>5}{'TN':>6}   acc    prec   rec   F1   cost EUR")
    report("Baseline (always 'on time')", y_test, baseline.predict(x_test), args.cost_miss, args.cost_check)
    report("Decision tree (depth 5)", y_test, (p_tree >= t).astype(int), args.cost_miss, args.cost_check)
    report(logreg_name, y_test, (p_logreg >= t).astype(int), args.cost_miss, args.cost_check)

    # 2. Threshold-free ranking metrics: how well do the scores sort late above on time?
    print("\nRanking quality (no threshold needed; higher is better)")
    print(f"  {'model':<38}ROC AUC   avg precision")
    print(f"  {'Random scores would get':<38}  0.50       {y_test.mean():.2f}")
    for name, p in [("Decision tree (depth 5)", p_tree), (logreg_name, p_logreg)]:
        print(f"  {name:<38}  {roc_auc_score(y_test, p):.2f}       {average_precision_score(y_test, p):.2f}")

    # 3. The threshold trade-off, shown on the test set so you can see it. Don't pick from this table.
    print(f"\nThreshold trade-off for {logreg_name.lower()} (test set, for reading only)")
    print("  threshold  flagged  caught  precision  recall   cost EUR")
    for th in [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8]:
        flagged = (p_logreg >= th).astype(int)
        _, fp, fn, tp = counts(y_test, flagged)
        print(f"  {th:>9.1f}  {tp + fp:>7}  {tp:>3}/{tp + fn:<3} {safe_div(tp, tp + fp):>8.2f}  "
              f"{safe_div(tp, tp + fn):>6.2f}  {cost_of(y_test, flagged, args.cost_miss, args.cost_check):>9,.0f}")

    # 4. Choose the threshold properly: cross-validation on the training rows only, scored by business cost.
    scorer = make_scorer(negative_cost_per_delivery, cost_miss=args.cost_miss, cost_check=args.cost_check)
    tuned = TunedThresholdClassifierCV(make_pipeline(StandardScaler(), learner), scoring=scorer, cv=5)
    tuned.fit(x_train, y_train)
    best = float(tuned.best_threshold_)
    flagged = tuned.predict(x_test)
    print(f"\nThreshold chosen by 5-fold cross-validation on training rows: {best:.2f}")
    print(f"  {'model':<38}{'TP':>5}{'FP':>5}{'FN':>5}{'TN':>6}   acc    prec   rec   F1   cost EUR")
    report(f"Same model, threshold {best:.2f}", y_test, flagged, args.cost_miss, args.cost_check)
    worst = cost_of(y_test, np.zeros(len(y_test), dtype=int), args.cost_miss, args.cost_check)
    print(f"  Cost if nothing is flagged: {worst:,.0f} EUR. If everything is flagged: "
          f"{cost_of(y_test, np.ones(len(y_test), dtype=int), args.cost_miss, args.cost_check):,.0f} EUR.")

    # 5. Calibration: when the model says 0.7, are about 70% of those deliveries late?
    print("\nCalibration on the test set (mean score in a bin vs share actually late)")
    print(f"  {'model':<38}Brier score (lower is better)")
    for name, p in [("Decision tree (depth 5)", p_tree), (logreg_name, p_logreg)]:
        print(f"  {name:<38}{brier_score_loss(y_test, p):.3f}")
    share, mean_score = calibration_curve(y_test, p_logreg, n_bins=5, strategy="quantile")
    print(f"  {logreg_name.lower()}, test rows in 5 equal-sized groups by score:")
    print("    mean score   share late")
    for m, s in zip(mean_score, share):
        print(f"       {m:.2f}         {s:.2f}")

    if args.plot:
        import matplotlib
        matplotlib.use("Agg")  # draw into a file, no window needed
        import matplotlib.pyplot as plt
        from sklearn.metrics import precision_recall_curve

        fig, (left, right) = plt.subplots(1, 2, figsize=(10, 4))
        for name, p in [("tree", p_tree), ("logistic", p_logreg)]:
            prec, rec, _ = precision_recall_curve(y_test, p)
            left.plot(rec, prec, label=name)
            s, m = calibration_curve(y_test, p, n_bins=5, strategy="quantile")
            right.plot(m, s, marker="o", label=name)
        left.axhline(y_test.mean(), linestyle=":", color="grey", label="random")
        left.set(xlabel="recall", ylabel="precision", title="Precision-recall curve")
        right.plot([0, 1], [0, 1], linestyle=":", color="grey", label="perfect")
        right.set(xlabel="mean score", ylabel="share actually late", title="Calibration")
        left.legend()
        right.legend()
        fig.tight_layout()
        out = HERE / "classification_metrics.png"
        fig.savefig(out, dpi=120)
        print(f"\nSaved {out.name}")


if __name__ == "__main__":
    main()

Step 3: Run it

In the terminal, still inside unit02, run:

python classify_late.py

It takes a few seconds. What success looks like:

4000 deliveries, 15% late. 3000 to learn from, 1000 held back for testing.
Costs: a missed late delivery 250 EUR, a planner check 50 EUR

At threshold 0.5: a delivery is flagged when its score is at least 0.5
  model                                    TP   FP   FN    TN   acc    prec   rec   F1   cost EUR
  Baseline (always 'on time')               0    0  151   849   0.85   0.00   0.00  0.00    37,750
  Decision tree (depth 5)                  54   43   97   806   0.86   0.56   0.36  0.44    29,100
  Logistic regression                      67   25   84   824   0.89   0.73   0.44  0.55    25,600

Ranking quality (no threshold needed; higher is better)
  model                                 ROC AUC   avg precision
  Random scores would get                 0.50       0.15
  Decision tree (depth 5)                 0.82       0.43
  Logistic regression                     0.88       0.63

Threshold trade-off for logistic regression (test set, for reading only)
  threshold  flagged  caught  precision  recall   cost EUR
        0.1      401  131/151     0.33    0.87     25,050
        0.2      258  110/151     0.43    0.73     23,150
        0.3      173   90/151     0.52    0.60     23,900
        0.4      122   75/151     0.61    0.50     25,100
        0.5       92   67/151     0.73    0.44     25,600
        0.6       63   49/151     0.78    0.32     28,650
        0.7       43   37/151     0.86    0.25     30,650
        0.8       28   24/151     0.86    0.16     33,150

Threshold chosen by 5-fold cross-validation on training rows: 0.21
  model                                    TP   FP   FN    TN   acc    prec   rec   F1   cost EUR
  Same model, threshold 0.21              103  138   48   711   0.81   0.43   0.68  0.53    24,050
  Cost if nothing is flagged: 37,750 EUR. If everything is flagged: 50,000 EUR.

Calibration on the test set (mean score in a bin vs share actually late)
  model                                 Brier score (lower is better)
  Decision tree (depth 5)               0.109
  Logistic regression                   0.086
  logistic regression, test rows in 5 equal-sized groups by score:
    mean score   share late
       0.00         0.01
       0.02         0.01
       0.06         0.08
       0.17         0.17
       0.52         0.49

Your numbers should match, because the data comes from a fixed random seed. A different last decimal is fine; library versions can cause that.

Step 4: Read the results

Take the output one block at a time.

  1. The first table. TP, FP, FN and TN are the four cells of the confusion matrix, on the 1,000 held-back deliveries. The baseline never flags anything: 0.85 accuracy, zero recall, 37,750 EUR of missed late deliveries. The tree reaches only 0.86 accuracy. Looking at accuracy alone, you would call the tree useless. Its cost column tells a different story: it saves 8,650 EUR against the baseline.
  2. Logistic regression at 0.5. Precision 0.73: when it flags, it's right about three times in four. Recall 0.44: it misses more than half of the late deliveries. That is what the default threshold does on imbalanced data.
  3. Ranking quality. Logistic regression ranks better than the tree on both measures: 0.88 against 0.82 AUC, and 0.63 against 0.43 average precision. Note the "random" line: average precision starts at 0.15, the share of late deliveries, not at 0.5.
  4. The trade-off table. Read down the columns. As the threshold rises, fewer deliveries are flagged, precision climbs and recall falls. The cost is lowest around 0.2. This table uses the test set so you can see the shape; choosing a threshold from it would leak test information into your decision.
  5. The chosen threshold, 0.21. Cross-validation on the training rows alone found 0.21, next to the break-even value of 50 / 250 = 0.2 from "How it works". On the test set it catches 103 of 151 late deliveries, and the cost falls from 25,600 to 24,050 EUR. Accuracy fell to 0.81. For this process, that is the better model.
  6. Why 0.21 isn't the very cheapest row. The 0.2 row of the trade-off table costs 23,150 EUR. With 151 late deliveries in the test set, a few rows either way move the cost by hundreds of euros. That is noise, and chasing it is exactly the test-set tuning you are avoiding.
  7. Calibration. Logistic regression's Brier score is lower than the tree's. In the five score groups, the mean score and the share actually late are close: 0.17 against 0.17, 0.52 against 0.49. These scores can be read as percentages, which is why the break-even formula worked.

Step 5: Move the threshold by hand

python classify_late.py --threshold 0.3

The first table now uses 0.3. The logistic regression row becomes:

  Logistic regression                      90   83   61   766   0.86   0.52   0.60  0.56    23,900

More deliveries caught (90 instead of 67), more false alarms (83 instead of 25), lower cost. The model didn't change; only the decision did.

Step 6: Change the costs

python classify_late.py --cost-miss 800 --cost-check 25

Now a miss is 32 times as expensive as a check. The break-even threshold is 25 / 800, about 0.03, and cross-validation picks 0.05:

Threshold chosen by 5-fold cross-validation on training rows: 0.05
  model                                    TP   FP   FN    TN   acc    prec   rec   F1   cost EUR
  Same model, threshold 0.05              145  393    6   456   0.60   0.27   0.96  0.42    18,250

It flags over half the deliveries and misses only 6 late ones. Accuracy is 0.60, and it is the cheapest option by far. When costs change, the threshold must change with them.

Step 7: Try class weights

python classify_late.py --balanced

With class_weight="balanced", the logistic regression at 0.5 now has recall 0.82 instead of 0.44. Look at the calibration block at the end:

    mean score   share late
       0.02         0.01
       0.10         0.01
       0.26         0.07
       0.52         0.17
       0.83         0.49

Scores of about 0.5 now mean 17% late. The weights moved every score upward, so the break-even formula no longer applies directly. Cross-validation compensates by choosing a threshold of 0.56. Weighting and threshold tuning are two ways to reach a similar decision; only the unweighted model keeps honest scores.

Step 8: Use your learning rate from the last topic

python classify_late.py --sgd --lr 0.01
python classify_late.py --sgd --lr 1.0

This trains the same logistic model with SGDClassifier, plain stochastic gradient descent with a constant step. At 0.01, the result is close to the default solver: AUC 0.88, and a tuned threshold of 0.18. At 1.0, the steps are too big, recall at 0.5 drops to 0.26, the Brier score rises from 0.087 to 0.125, and the cost goes up. Compare this with the working range you wrote down in your gradient descent notes.

Step 9: Draw the curves

python classify_late.py --plot

Open classification_metrics.png from the VS Code file list. On the left, the precision-recall curves: logistic regression's line sits above the tree's almost everywhere, and both sit well above the dotted "random" line at 0.15. On the right, the calibration curves: logistic regression follows the dotted diagonal closely; the tree's top point, at a mean score of about 0.66, is late only about 46% of the time.

Step 10: Save your work in Git

From the course folder:

cd ..
git add unit02/classify_late.py
git commit -m "Late-delivery classifier with cost-based threshold"

classification_metrics.png is an output you can recreate at any time, so you don't need to commit it.

What each part of the script does

Part What it does
make_deliveries Makes deliveries with the same columns as rules_vs_learning.py, shifted so about 15% are late
pd.get_dummies Turns Plant and ShippingType into 0/1 columns
train_test_split(..., stratify=label) Holds back 25% as a test set with the same share of late rows
counts Returns the confusion matrix cells as tn, fp, fn, tp
cost_of Business cost: every flag costs a check, every miss costs --cost-miss
report Prints the four cells, accuracy, precision, recall, F1 and cost for one set of flags
make_pipeline(StandardScaler(), learner) Scales features, then fits logistic regression (lbfgs, or SGDClassifier with --sgd)
predict_proba(...)[:, 1] The score for "late"
Ranking block roc_auc_score and average_precision_score, which need scores, not flags
Trade-off loop Precision, recall and cost at thresholds 0.1 to 0.8, for reading only
TunedThresholdClassifierCV Searches the threshold with 5-fold cross-validation on training rows, scored by make_scorer(negative_cost_per_delivery)
Calibration block brier_score_loss, and calibration_curve with 5 equal-sized groups
--plot block Saves precision-recall and calibration curves to classification_metrics.png

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of the Unit 1 setup, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'sklearn' (or pandas, matplotlib) The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
ImportError: cannot import name 'TunedThresholdClassifierCV' Your scikit-learn is older than 1.5 Run pip install --upgrade scikit-learn, then run the script again
pip shows ProxyError, SSLError or Could not fetch URL Your network or company proxy blocks PyPI Try another network, or ask IT for access to pypi.org and files.pythonhosted.org or an internal mirror. The script itself needs no network
can't open file ... classify_late.py The terminal is not in unit02, or the file has another name Run cd unit02 from the course folder, and check the file name
error: --threshold must be between 0 and 1 You passed a value such as 30 or 1.5 Use a decimal: --threshold 0.3
ConvergenceWarning with --sgd The SGD steps didn't settle within the iteration limit Expected with a large --lr. Lower it, for example --lr 0.01
No API key is asked for Correct: nothing in this topic uses a key or account Nothing to do

Where this shows up in SAP

Unit 2 is a maths-and-mechanics unit, so this section is short. The point is that SAP's tools speak the same language, and the decisions you just practised still fall to your team.

Embedded: scores inside S/4HANA

SAP Learning describes ISLM being used to train and activate customer-specific predictive models. In S/4HANA Utilities, those models produce a release confidence value for implausible meter readings and outsorted billing documents. That is a classifier score shown to a clerk. The questions from this topic apply unchanged: at what value can a document be released without review, what does a wrong release cost, and how many documents land in the review queue at that value? Using ISLM needs an S/4HANA system with the scenario set up, which this course doesn't assume.

SAP HANA: PAL through hana-ml

An SAP blog on UnifiedClassification in the hana-ml Python client shows training with a stratified partition and a report of AUC, RECALL, PRECISION, F1_SCORE, ACCURACY, KAPPA and MCC, a confusion matrix with actual and predicted class counts, ROC and gains data, and an HTML model report. The blog was written for an older hana-ml release, so check the current reference before relying on exact output names. You now know how to read every one of those numbers, and why ACCURACY alone can't decide.

Pretrained: SAP-RPT-1

SAP Learning describes SAP-RPT-1 classification tasks such as cost centers, payment terms and sales groups, reached through a REST endpoint in SAP AI Core, with a confidence score between 0.0 and 1.0 for each predicted cell. Two things follow from this topic:

  • Treat the confidence as a score until you've checked it. Hold back labelled rows, send the rest as context, and draw the same calibration table the script prints.
  • You still choose the threshold. For example: auto-fill the cost center when confidence is high enough, otherwise route to a person. Pick that cut-off from the cost of a wrong auto-fill against the cost of a manual check.

SAP's open-source research variant, sap-rpt-1-oss on SAP-samples, exposes a scikit-learn style SAP_RPT_OSS_Classifier with fit, predict and predict_proba. So the metric functions in the script work on it unchanged. Its README recommends a large GPU and limits the checkpoints to research use, so it isn't a laptop exercise.

Build, library or service

Situation Use Why
Learning the metrics, auditing a vendor's numbers Your own script Every count is visible
A classifier on a table you already have scikit-learn: LogisticRegression, TunedThresholdClassifierCV Tested metrics, cost-based threshold search
Data already in SAP HANA, analysts who want reports PAL through hana-ml Training and standard metric reports next to the data
A standard S/4HANA prediction scenario Embedded ML with ISLM Runs inside the standard lifecycle; you still own the threshold
Table predictions without a training pipeline SAP-RPT-1 Pretrained; you check calibration and set the cut-off

Production concerns

  • Write down the metric contract before training. The positive class, the costs of a miss and a false alarm, the flag volume the team can handle, and the baseline. In the FDE's terms from What an SAP forward deployed engineer does, this is how "move metric M from X to Y" becomes concrete.
  • Choose thresholds on training or validation data only. Report the test set once. If you keep adjusting after seeing test results, you need a fresh test set.
  • Split by time for real data. Test on the newest deliveries, as the previous topics explained. Class balance also shifts over time, and so does the best threshold.
  • Version the threshold with the model. The threshold is a model setting. Store it with the model, log it with each prediction, and change it through the same review as the model.
  • Watch flag volume. A threshold that is cheap on paper can flood a planner queue. Put the expected flags per day in front of the process owner.
  • Monitor recall, precision and calibration after go-live. Labels arrive late (you only know a delivery was late after it arrives), so plan when each week's metrics can be computed.
  • Authorizations. Scores about customers, suppliers or employees are business data. Show them only to users allowed to see the underlying documents, and log who acted on a flag. Unit 7 covers grounding without breaking authorizations.
  • Explain flags. A planner will ask why a delivery was flagged. Logistic regression weights and small trees can answer; keep that in mind before choosing a more complex model.

Pitfalls

  • Quoting accuracy on imbalanced data. Always print the "always no" baseline next to it.
  • Leaving the threshold at 0.5. It is a library default, not a business decision.
  • Tuning the threshold on the test set. The trade-off table is for reading. The choice comes from cross-validation on training rows.
  • Reading AUC as a hit rate. AUC is a ranking probability. A model with 0.88 AUC can still miss most late deliveries at a bad threshold.
  • Mixing up the confusion matrix order. scikit-learn's ravel() gives tn, fp, fn, tp. Name the variables.
  • Using class weights and then reading scores as probabilities. Weights distort calibration, as Step 7 showed.
  • Calibrating with isotonic regression on small data. scikit-learn warns it can overfit below about 1,000 rows. Use the sigmoid method or more data.
  • Forgetting the positive class. Precision and recall of "on time" are different numbers from those of "late". State which class is positive.

Exercise: write the metric contract for late deliveries

You will turn this topic into a one-page metric contract for the next topic in Unit 2, Predicting late deliveries with SAP data. The contract fixes the metric, the costs and the threshold rule before any real data arrives.

  1. In the terminal, inside unit02 with (.venv) showing, run the script with three cost settings and note the chosen threshold, the recall, the precision, the cost and the number of flagged deliveries (TP + FP) for each:

    python classify_late.py --cost-miss 250 --cost-check 50
    python classify_late.py --cost-miss 500 --cost-check 50
    python classify_late.py --cost-miss 150 --cost-check 50
  2. For each run, compute the break-even threshold cost_check / cost_miss by hand and compare it with the threshold cross-validation chose.

  3. Run python classify_late.py --balanced and note the chosen threshold and the calibration table. Write one sentence on why its threshold is higher.

  4. In VS Code, create unit02/metric_contract_late_deliveries.md with these sections:

    • Positive class: which outcome is "positive".
    • Costs: your assumed cost of a miss and of a planner check, and who in the business should confirm them.
    • Primary metric: total cost, plus which of recall or precision you will report next to it, and why.
    • Threshold rule: how the threshold will be chosen (cross-validation on training rows, scored by cost) and the break-even check.
    • Guard rails: the baseline to beat, and the maximum flags per 1,000 deliveries the planners can handle.
    • Results table: the three runs from step 1.
  5. Save your work:

    git add unit02/metric_contract_late_deliveries.md
    git commit -m "Metric contract for late-delivery prediction"

Done when: the contract names the positive class, both costs, a primary metric and a threshold rule; the results table has all three runs with the chosen and break-even thresholds side by side; the --balanced sentence explains the higher threshold; and git log shows the commit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Which metrics can you compute before choosing any threshold?

    Answer: B. ROC AUC and average precision measure how well the scores rank late above on-time deliveries, and the Brier score compares scores with outcomes directly. Precision, recall, F1, accuracy and cost all need flags, which need a threshold.
  2. 2In the script, confusion_matrix(y_true, flagged).ravel() returns four numbers. In what order?

    Answer: D. scikit-learn orders rows and columns by label, 0 then 1, so the flattened matrix is true negatives, false positives, false negatives, true positives. The script names each variable to avoid mixing them up.
  3. 3A miss costs 250 EUR and a planner check costs 50 EUR. With calibrated scores, what is the break-even threshold, and why?

    Answer: A. Flag when the expected cost of not flagging, p × 250, is at least the 50 EUR check. That gives p ≥ 50 / 250 = 0.2. Cross-validation on the training rows found 0.21, which agrees.
  4. 4A colleague picks the threshold by choosing the cheapest row in the test-set trade-off table. What is wrong with that?

    Answer: C. Anything chosen by looking at the test set is tuned to it. scikit-learn's guide warns against tuning the threshold on the data used to train, and the same logic applies to test data. Use cross-validation on training rows, then report the test set once.
  5. 5After switching on class_weight="balanced", recall at 0.5 jumped from 0.44 to 0.82, but deliveries scored about 0.5 were late only 17% of the time. What happened?

    Answer: B. Balanced weights make rare late rows count more in training, which raises every late score. Ranking can stay good, but scores stop meaning percentages. Cross-validation compensated by choosing 0.56.
  6. 6SAP-RPT-1 returns a confidence of 0.0 to 1.0 per predicted cost center. You want to auto-fill high-confidence predictions. What do you do first?

    Answer: D. A confidence score is only a percentage if the calibration table says so on your data. The cut-off is then a cost decision, exactly like the late-delivery threshold. AUC is a ranking probability, not a hit rate.
  7. 7Your model beats the baseline on cost, but at the chosen threshold it flags 400 deliveries a day, and planners can check 150. What should you do?

    Answer: D. The cost model assumed every flag gets checked. If only 150 can be, the unchecked flags behave like misses. Flag volume is a guard rail in the metric contract, so the threshold or the work queue must respect it.
  8. 8A hana-ml report for a late-delivery model shows ACCURACY 0.92 and RECALL 0.34 for the late class. How do you read it?

    Answer: C. On imbalanced data, accuracy is dominated by the common class. A recall of 0.34 means most real positives go unflagged. Look at recall and precision for the positive class, and at the baseline.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in