Most business predictions are classification: pick one outcome from a short list. Will this delivery be late? Is this invoice a duplicate? Which cost center does this line belong to?
A classifier doesn't really answer "yes" or "no". It gives each case a score, say 0.72 for "late". A person or a system then applies a threshold: flag everything scoring 0.5 or more, for example. Changing that threshold changes which cases get flagged, without retraining anything.
Once the flags are out, there are only four things that can happen to each case:
Really late
Really on time
Flagged
Caught it (true positive)
False alarm (false positive)
Not flagged
Missed it (false negative)
Correctly left alone (true negative)
That table is the confusion matrix. Every metric in this topic is a way of summarizing it. The one idea to keep: the right metric is the one that matches what a false alarm and a miss cost your business. Accuracy, the metric most people quote, is usually the wrong one.
Take order-to-cash. A planner can check only so many deliveries a day. You want a model that flags deliveries at risk of being late so the planner can act: switch carrier, split the shipment, warn the customer.
Suppose 15 in every 100 deliveries end up late. A "model" that never flags anything is right 85% of the time. That is 85% accuracy, and it catches zero late deliveries. In the deep layer of this topic, an actual trained tree scores 86% accuracy on made-up data like this, barely above doing nothing. Google's Machine Learning Crash Course warns about exactly this: accuracy can mislead when one outcome is much rarer than the other.
So you need metrics that talk about the rare outcome directly:
Recall: of all the deliveries that really were late, what share did we flag? A low recall means late deliveries slip through.
Precision: of the deliveries we flagged, what share really were late? A low precision means planners waste time on false alarms.
The two pull against each other. Lower the threshold and you catch more late deliveries (recall up), but you also flag more on-time ones (precision down). Google's course describes this as an inverse relationship driven by the threshold.
The good news is that the business can settle the trade-off with money. In the deep layer's example, a missed late delivery costs 250 EUR and a planner check costs 50 EUR. At the default threshold of 0.5 the model's flags cost 25,600 EUR on 1,000 test deliveries. At a threshold chosen from those two costs, about 0.2, the cost drops to about 24,000 EUR, even though accuracy falls from 0.89 to 0.81. Lower accuracy, lower cost. That is the whole lesson in one line.
As of September 2026, the same metrics appear wherever SAP runs classification:
Embedded in S/4HANA, through ISLM. SAP Learning describes Intelligent Scenario Lifecycle Management (ISLM) training and activating customer-specific models. In S/4HANA Utilities, for example, a model computes a release confidence value for implausible meter readings and outsorted billing documents, and clerks see it as a column. That value is a score. Somebody still decides at what value a document is released without a human look: a threshold.
In SAP HANA, through PAL and the hana-ml client. An SAP blog on the UnifiedClassification class of the hana-ml Python client shows it reporting AUC, recall, precision, F1 score, accuracy and two more agreement measures (kappa and MCC), plus a confusion matrix and an HTML model report. The vocabulary in this topic is the vocabulary of those reports.
Pretrained for tables: SAP-RPT-1. SAP Learning lists classification tasks such as cost centers, payment terms and sales groups. It returns a confidence score between 0.0 and 1.0 for each predicted value. Whether 0.8 really means "right about 80% of the time" on your data is something you check on your own held-back rows, as described below.
Whichever route you take, SAP gives you scores and standard metrics. Choosing the metric and the threshold is still your team's job, because only the business knows what a miss costs.
Misses are expensive (late delivery, fraud, a duplicate payment going out)
Recall
You want to catch as many real cases as possible
False alarms are expensive (a blocked order, an angry customer, a clerk's hour)
Precision
Every flag should be worth acting on
Both matter and you need one number
F1 score
Balances precision and recall
You want to compare models before anyone picks a threshold
ROC AUC or average precision
Measures how well the scores rank real cases above the rest
The score itself drives a decision, such as a priority queue or a price
Calibration
A score of 0.3 should mean about 30% of such cases turn out positive
You know the cost of each kind of error
Total cost
Converts the confusion matrix straight into euros
The outcomes are roughly balanced and errors cost about the same
Accuracy is acceptable
The one case where it doesn't mislead
A practical rule: agree the cost of a miss and the cost of a false alarm with the process owner before the project starts. Everything else follows from those two numbers.
"95% accuracy means the model works." If 95% of cases are negative, a model that always says "no" gets 95%. Ask for recall and precision on the rare outcome.
"The threshold is 0.5, because that's the default." Libraries use 0.5 by default. The right threshold depends on your costs and is often far from 0.5.
"A better model always has higher accuracy." In the deep layer's example, the cheapest set of flags has lower accuracy than the default one.
"A confidence score is a probability." Some models' scores are well calibrated, many are not. Check before using a score as a percentage.
"One number is enough." Precision without recall hides misses; recall without precision hides noise. Report both, plus the flag volume.
"Once chosen, the threshold is fixed." When costs, volumes or data change, the best threshold moves too.
Pick one answer for each question. The explanation appears after you choose.
115% of deliveries are late. A vendor reports 86% accuracy for their late-delivery model. What should you ask first?
Answer: B. With 15% late, "always on time" is right 85% of the time and catches nothing. Accuracy hides that. Recall and precision describe the rare outcome you actually care about.
2A model gives each delivery a score. What does lowering the flagging threshold from 0.5 to 0.2 usually do?
Answer: C. A lower threshold flags more deliveries. More real late ones get caught, so recall rises, but more on-time ones get flagged too, so precision falls. The model itself is unchanged.
3In accounts payable, a duplicate payment that goes out is costly, and a clerk's check is cheap. Which metric should lead?
Answer: D. When misses cost far more than false alarms, you want to catch as many real cases as possible. That is recall. Precision would lead if false alarms were the expensive error.
4In the example, a threshold chosen from business costs lowered accuracy from 0.89 to 0.81. Why can that be the better choice?
Answer: B. Accuracy treats a miss and a false alarm as equally bad. Here a miss costs 250 EUR and a check 50 EUR, so flagging more deliveries lowered the total cost even though more flags were wrong.
5SAP-RPT-1 returns a confidence score of 0.8 for a predicted payment term. What should your team do before treating it as "80% sure"?
Answer: D. A score between 0 and 1 is not automatically calibrated. Comparing scores with how often predictions were actually right, on your own data, tells you whether you can read it as a percentage.
6Which two numbers should you agree with the process owner before choosing a threshold?
Answer: A. The threshold is a business decision about which error to accept more of. Once the two error costs are known, the threshold can be chosen to minimize the total, and the rest of the metrics follow.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 38 min read
#Mental model: the model ranks, the business decides
A classifier does two separate jobs, and most confusion about metrics comes from mixing them up. scikit-learn's own guide makes this split explicit: first a statistical problem, predicting a probability or score; then a decision problem, turning that score into an action.
flowchart LR
X[Delivery features] --> M[Model]
M --> S[Score 0 to 1]
S --> T{Score >= threshold?}
T -->|yes| F[Flag for planner]
T -->|no| N[Leave alone]
That gives three families of metrics, each answering a different question:
Question
Metrics
Depends on the threshold?
Are the scores in the right order?
ROC AUC, average precision
No
Are the scores honest numbers?
Calibration curve, Brier score, log loss
No
Are the decisions good?
Confusion matrix, precision, recall, F1, total cost
Yes
Pick a model on the first two. Pick a threshold on the third, using costs. Never judge the first two by the third.
This topic uses logistic regression, the classifier promised at the end of Gradient descent, by hand. It computes the same weighted sum as the freight-cost line, w1*x1 + w2*x2 + ... + b, then squeezes it through the sigmoid function 1 / (1 + e^-z) so the result lands between 0 and 1. Training uses a gradient-based method to adjust the weights. The loss is log loss, which punishes confident wrong scores heavily.
In scikit-learn, LogisticRegression uses the lbfgs solver by default, with max_iter defaulting to 100. Its predict_proba returns one column per class; column 1 holds the score for "late". The script also offers SGDClassifier(loss="log_loss"), which is the plain stochastic gradient descent you wrote yourself, so you can reuse the learning rate from your notes.
With the counts TP (caught), FP (false alarm), FN (missed) and TN (correctly left alone), the metrics are simple ratios. The definitions below follow Google's crash course and scikit-learn's guide.
Metric
Formula
Plain question
Accuracy
(TP + TN) / all
How often was the flag right?
Recall (true positive rate)
TP / (TP + FN)
Of the late deliveries, how many did we flag?
Precision
TP / (TP + FP)
Of the flags, how many were late?
False positive rate
FP / (FP + TN)
Of the on-time deliveries, how many did we flag?
F1
2 × precision × recall / (precision + recall)
One number balancing the two
Balanced accuracy
(recall + TN / (TN + FP)) / 2
Accuracy that weighs both classes equally
In code, sklearn.metrics.confusion_matrix(y_true, y_pred).ravel() returns the four counts in the order tn, fp, fn, tp. Getting that order wrong is a classic bug; the script names them explicitly.
Accuracy counts every row equally. When 85% of rows are "on time", the on-time rows dominate the number. A model can ignore the late class entirely and still look good. Google's crash course recommends F1 over accuracy for imbalanced data, and scikit-learn's guide points to balanced accuracy for the same reason.
The fix is not a cleverer average. It is to look at the rare class directly (recall, precision) and to always print the baseline next to the model, as in Machine learning in plain terms.
Before any threshold is chosen, you can ask how well the scores rank late deliveries above on-time ones.
The ROC curve plots recall (true positive rate) against the false positive rate at every possible threshold. Google's crash course gives the key reading of the area under it: AUC is the probability that the model scores a randomly chosen positive higher than a randomly chosen negative. 0.5 is a coin flip, 1.0 is perfect ranking.
The precision-recall curve plots precision against recall at every threshold. Its summary, average precision, starts at the share of positives for random scores (0.15 in our data), not at 0.5. Google's course notes that precision-recall curves can give a better comparison when classes are imbalanced, because they ignore the large pile of easy true negatives.
Use these to compare models. Don't report them to a process owner as "the model is 88% good". AUC is a ranking probability, not a hit rate.
scikit-learn's predict uses a threshold of 0.5 on predict_proba by default. Its user guide states plainly that this default is often not what the task needs.
Suppose each flagged delivery costs a planner check (c_check), and a check always prevents the loss. Each missed late delivery costs c_miss. For one delivery with a well-calibrated score p:
Flag it: cost c_check, whatever the outcome.
Don't flag it: expected cost p × c_miss.
Flag when p × c_miss ≥ c_check, which means a threshold of c_check / c_miss. With 50 EUR and 250 EUR, that is 0.2. This break-even rule only holds when the scores are calibrated and the cost model is this simple. In practice you let cross-validation search for the threshold with the real cost function, and use the formula as a sanity check.
scikit-learn does that search with TunedThresholdClassifierCV. By default it maximizes balanced accuracy using 5-fold stratified cross-validation; you pass your own score with make_scorer. Its guide warns against tuning the threshold on the same data used to train the model. FixedThresholdClassifier wraps a model with a threshold you set by hand.
flowchart LR
TR[Training rows] --> CV[5 folds: train on 4, score thresholds on 1]
CV --> BT[Threshold with lowest average cost]
BT --> RF[Refit model on all training rows]
RF --> TE[Report once on test rows]
A score is well calibrated when, among cases scored around 0.8, about 80% are really positive. That is scikit-learn's definition. Ranking can be excellent while calibration is poor: multiply every score by 0.5 and the order, and so the AUC, doesn't change.
You check it with a calibration curve (reliability diagram): sort test rows by score into bins, then compare each bin's average score with its actual share of positives. On a perfect model the points lie on the diagonal. scikit-learn's guide notes that LogisticRegression tends to be well calibrated out of the box, while models such as random forests and naive Bayes are not. When calibration matters, CalibratedClassifierCV fits a correction: method="sigmoid" for small data, or method="isotonic", which the guide warns can overfit below about 1,000 rows.
Two summary numbers measure scores directly: the Brier score (mean squared difference between score and outcome) and log loss. Lower is better for both. scikit-learn calls them strictly proper scoring rules. They mix calibration with ranking quality, so read them with the calibration curve, not instead of it.
LogisticRegression(class_weight="balanced") weights each class by n_samples / (n_classes × count of that class), so rare late rows count more in the loss. That usually raises recall at the default threshold. It also pushes every score upward, so the scores are no longer calibrated. The script lets you see this. Resampling tricks, such as creating synthetic minority rows, have the same side effect.
Often the simpler route gives the same result: train without weights, keep calibrated scores, and move the threshold. Either way, re-tune the threshold after changing the weights.
#Build it yourself: judge a late-delivery classifier by cost
You will train a logistic regression and a small tree to flag late deliveries on 4,000 made-up SAP-shaped deliveries, of which about 15% are late. You will print the confusion matrix and every metric above, see accuracy mislead, pick a threshold from business costs with cross-validation, and check calibration. One script, about 40 minutes.
flowchart LR
G[4,000 made-up deliveries] --> S[Split 75/25, stratified]
S --> M[Train: baseline, tree, logistic regression]
M --> C[Confusion matrix and metrics at one threshold]
M --> R[ROC AUC, average precision]
M --> T[Threshold chosen by cost, with cross-validation]
M --> K[Calibration check]
In VS Code's file list, right-click unit02, choose New File and name it classify_late.py.
Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""Classification and the metrics that matter: flag late deliveries, then judge the flags honestly.
How to run (from the unit02 folder, with the course .venv turned on):
python classify_late.py # 4,000 made-up deliveries, about 1 in 7 late
python classify_late.py --threshold 0.3 # flag more deliveries: recall up, precision down
python classify_late.py --cost-miss 800 --cost-check 25 # change the business costs
python classify_late.py --balanced # weight the rare late class more in training
python classify_late.py --sgd --lr 0.01 # train with SGD and your learning rate from the last topic
python classify_late.py --plot # also save classification_metrics.png
All data is made up. Nothing is sent anywhere.
"""
import argparse
from pathlib import Path
import numpy as np
import pandas as pd
from sklearn.calibration import calibration_curve
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression, SGDClassifier
from sklearn.metrics import (average_precision_score, brier_score_loss, confusion_matrix,
make_scorer, roc_auc_score)
from sklearn.model_selection import TunedThresholdClassifierCV, train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
HERE = Path(__file__).parent
PLANTS = ["1010", "1020", "1710"] # made-up plant codes
CARRIERS = ["ROAD", "RAIL", "EXPRESS"] # made-up shipping types
def make_deliveries(rows: int, seed: int = 42) -> pd.DataFrame:
"""Same columns as rules_vs_learning.py, but late deliveries are rarer: about 1 in 7."""
rng = np.random.default_rng(seed)
df = pd.DataFrame({
"Plant": rng.choice(PLANTS, size=rows),
"ShippingType": rng.choice(CARRIERS, size=rows, p=[0.6, 0.3, 0.1]),
"Items": rng.integers(1, 12, size=rows),
"WeightKg": rng.gamma(2.0, 150.0, size=rows).round(1),
"LeadTimeDays": rng.integers(1, 15, size=rows),
})
risk = (
0.25 * df["Items"]
+ 0.004 * df["WeightKg"]
- 0.35 * df["LeadTimeDays"]
+ np.where(df["Plant"] == "1710", 1.0, 0.0)
+ np.where(df["ShippingType"] == "EXPRESS", -1.5, 0.0)
- 3.0 # shifts the mix so fewer deliveries are late
+ rng.logistic(0, 1.0, size=rows)
)
df["Late"] = (risk > 0).astype(int)
return df
def counts(y_true, flagged):
"""The four cells of the confusion matrix: tn, fp, fn, tp."""
tn, fp, fn, tp = confusion_matrix(y_true, flagged, labels=[0, 1]).ravel()
return int(tn), int(fp), int(fn), int(tp)
def safe_div(a, b):
return a / b if b else 0.0
def cost_of(y_true, flagged, cost_miss, cost_check):
"""Business cost: every flagged delivery gets a planner check; every missed late one costs cost_miss."""
_, fp, fn, tp = counts(y_true, flagged)
return (tp + fp) * cost_check + fn * cost_miss
def negative_cost_per_delivery(y_true, flagged, cost_miss, cost_check):
"""Scikit-learn maximizes scores, so return minus the average cost."""
return -cost_of(y_true, flagged, cost_miss, cost_check) / len(y_true)
def report(name, y_true, flagged, cost_miss, cost_check):
"""One line per model: the confusion matrix cells plus the metrics built from them."""
tn, fp, fn, tp = counts(y_true, flagged)
accuracy = (tp + tn) / len(y_true)
precision = safe_div(tp, tp + fp)
recall = safe_div(tp, tp + fn)
f1 = safe_div(2 * precision * recall, precision + recall)
cost = cost_of(y_true, flagged, cost_miss, cost_check)
print(f" {name:<38}{tp:>5}{fp:>5}{fn:>5}{tn:>6} {accuracy:.2f} {precision:.2f} "
f"{recall:.2f} {f1:.2f} {cost:>8,.0f}")
def main() -> None:
parser = argparse.ArgumentParser(description="Train a late-delivery classifier and judge it with the right metrics.")
parser.add_argument("--rows", type=int, default=4000, help="how many made-up deliveries")
parser.add_argument("--threshold", type=float, default=0.5, help="flag a delivery when its score is at least this")
parser.add_argument("--cost-miss", type=float, default=250.0, help="EUR lost when a late delivery is not flagged")
parser.add_argument("--cost-check", type=float, default=50.0, help="EUR a planner spends checking one flagged delivery")
parser.add_argument("--balanced", action="store_true", help="use class_weight='balanced' in training")
parser.add_argument("--sgd", action="store_true", help="train logistic regression with SGDClassifier instead of lbfgs")
parser.add_argument("--lr", type=float, default=0.01, help="learning rate for --sgd")
parser.add_argument("--plot", action="store_true", help="save classification_metrics.png")
args = parser.parse_args()
if not 0.0 < args.threshold < 1.0:
parser.error("--threshold must be between 0 and 1, for example 0.3")
df = make_deliveries(args.rows)
label = df["Late"]
features = pd.get_dummies(df.drop(columns=["Late"]), dtype=int) # text columns become 0/1 columns
# Hold back a quarter as the test set. stratify keeps the same share of late rows on both sides.
x_train, x_test, y_train, y_test = train_test_split(
features, label, test_size=0.25, random_state=0, stratify=label)
print(f"{len(df)} deliveries, {label.mean():.0%} late. "
f"{len(x_train)} to learn from, {len(x_test)} held back for testing.")
print(f"Costs: a missed late delivery {args.cost_miss:,.0f} EUR, a planner check {args.cost_check:,.0f} EUR\n")
# The models. Logistic regression outputs a score between 0 and 1 for "late".
weight = "balanced" if args.balanced else None
if args.sgd:
learner = SGDClassifier(loss="log_loss", learning_rate="constant", eta0=args.lr,
class_weight=weight, max_iter=1000, random_state=0)
logreg_name = f"Logistic regression, SGD lr {args.lr}"
else:
learner = LogisticRegression(max_iter=1000, class_weight=weight)
logreg_name = "Logistic regression"
logreg = make_pipeline(StandardScaler(), learner).fit(x_train, y_train)
tree = DecisionTreeClassifier(max_depth=5, class_weight=weight, random_state=0).fit(x_train, y_train)
baseline = DummyClassifier(strategy="most_frequent").fit(x_train, y_train)
p_logreg = logreg.predict_proba(x_test)[:, 1] # column 1 = chance of class 1, "late"
p_tree = tree.predict_proba(x_test)[:, 1]
# 1. The confusion matrix and the metrics built from it, at one threshold.
t = args.threshold
print(f"At threshold {t}: a delivery is flagged when its score is at least {t}")
print(f" {'model':<38}{'TP':>5}{'FP':>5}{'FN':>5}{'TN':>6} acc prec rec F1 cost EUR")
report("Baseline (always 'on time')", y_test, baseline.predict(x_test), args.cost_miss, args.cost_check)
report("Decision tree (depth 5)", y_test, (p_tree >= t).astype(int), args.cost_miss, args.cost_check)
report(logreg_name, y_test, (p_logreg >= t).astype(int), args.cost_miss, args.cost_check)
# 2. Threshold-free ranking metrics: how well do the scores sort late above on time?
print("\nRanking quality (no threshold needed; higher is better)")
print(f" {'model':<38}ROC AUC avg precision")
print(f" {'Random scores would get':<38} 0.50 {y_test.mean():.2f}")
for name, p in [("Decision tree (depth 5)", p_tree), (logreg_name, p_logreg)]:
print(f" {name:<38} {roc_auc_score(y_test, p):.2f} {average_precision_score(y_test, p):.2f}")
# 3. The threshold trade-off, shown on the test set so you can see it. Don't pick from this table.
print(f"\nThreshold trade-off for {logreg_name.lower()} (test set, for reading only)")
print(" threshold flagged caught precision recall cost EUR")
for th in [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8]:
flagged = (p_logreg >= th).astype(int)
_, fp, fn, tp = counts(y_test, flagged)
print(f" {th:>9.1f} {tp + fp:>7} {tp:>3}/{tp + fn:<3} {safe_div(tp, tp + fp):>8.2f} "
f"{safe_div(tp, tp + fn):>6.2f} {cost_of(y_test, flagged, args.cost_miss, args.cost_check):>9,.0f}")
# 4. Choose the threshold properly: cross-validation on the training rows only, scored by business cost.
scorer = make_scorer(negative_cost_per_delivery, cost_miss=args.cost_miss, cost_check=args.cost_check)
tuned = TunedThresholdClassifierCV(make_pipeline(StandardScaler(), learner), scoring=scorer, cv=5)
tuned.fit(x_train, y_train)
best = float(tuned.best_threshold_)
flagged = tuned.predict(x_test)
print(f"\nThreshold chosen by 5-fold cross-validation on training rows: {best:.2f}")
print(f" {'model':<38}{'TP':>5}{'FP':>5}{'FN':>5}{'TN':>6} acc prec rec F1 cost EUR")
report(f"Same model, threshold {best:.2f}", y_test, flagged, args.cost_miss, args.cost_check)
worst = cost_of(y_test, np.zeros(len(y_test), dtype=int), args.cost_miss, args.cost_check)
print(f" Cost if nothing is flagged: {worst:,.0f} EUR. If everything is flagged: "
f"{cost_of(y_test, np.ones(len(y_test), dtype=int), args.cost_miss, args.cost_check):,.0f} EUR.")
# 5. Calibration: when the model says 0.7, are about 70% of those deliveries late?
print("\nCalibration on the test set (mean score in a bin vs share actually late)")
print(f" {'model':<38}Brier score (lower is better)")
for name, p in [("Decision tree (depth 5)", p_tree), (logreg_name, p_logreg)]:
print(f" {name:<38}{brier_score_loss(y_test, p):.3f}")
share, mean_score = calibration_curve(y_test, p_logreg, n_bins=5, strategy="quantile")
print(f" {logreg_name.lower()}, test rows in 5 equal-sized groups by score:")
print(" mean score share late")
for m, s in zip(mean_score, share):
print(f" {m:.2f} {s:.2f}")
if args.plot:
import matplotlib
matplotlib.use("Agg") # draw into a file, no window needed
import matplotlib.pyplot as plt
from sklearn.metrics import precision_recall_curve
fig, (left, right) = plt.subplots(1, 2, figsize=(10, 4))
for name, p in [("tree", p_tree), ("logistic", p_logreg)]:
prec, rec, _ = precision_recall_curve(y_test, p)
left.plot(rec, prec, label=name)
s, m = calibration_curve(y_test, p, n_bins=5, strategy="quantile")
right.plot(m, s, marker="o", label=name)
left.axhline(y_test.mean(), linestyle=":", color="grey", label="random")
left.set(xlabel="recall", ylabel="precision", title="Precision-recall curve")
right.plot([0, 1], [0, 1], linestyle=":", color="grey", label="perfect")
right.set(xlabel="mean score", ylabel="share actually late", title="Calibration")
left.legend()
right.legend()
fig.tight_layout()
out = HERE / "classification_metrics.png"
fig.savefig(out, dpi=120)
print(f"\nSaved {out.name}")
if __name__ == "__main__":
main()
The first table. TP, FP, FN and TN are the four cells of the confusion matrix, on the 1,000 held-back deliveries. The baseline never flags anything: 0.85 accuracy, zero recall, 37,750 EUR of missed late deliveries. The tree reaches only 0.86 accuracy. Looking at accuracy alone, you would call the tree useless. Its cost column tells a different story: it saves 8,650 EUR against the baseline.
Logistic regression at 0.5. Precision 0.73: when it flags, it's right about three times in four. Recall 0.44: it misses more than half of the late deliveries. That is what the default threshold does on imbalanced data.
Ranking quality. Logistic regression ranks better than the tree on both measures: 0.88 against 0.82 AUC, and 0.63 against 0.43 average precision. Note the "random" line: average precision starts at 0.15, the share of late deliveries, not at 0.5.
The trade-off table. Read down the columns. As the threshold rises, fewer deliveries are flagged, precision climbs and recall falls. The cost is lowest around 0.2. This table uses the test set so you can see the shape; choosing a threshold from it would leak test information into your decision.
The chosen threshold, 0.21. Cross-validation on the training rows alone found 0.21, next to the break-even value of 50 / 250 = 0.2 from "How it works". On the test set it catches 103 of 151 late deliveries, and the cost falls from 25,600 to 24,050 EUR. Accuracy fell to 0.81. For this process, that is the better model.
Why 0.21 isn't the very cheapest row. The 0.2 row of the trade-off table costs 23,150 EUR. With 151 late deliveries in the test set, a few rows either way move the cost by hundreds of euros. That is noise, and chasing it is exactly the test-set tuning you are avoiding.
Calibration. Logistic regression's Brier score is lower than the tree's. In the five score groups, the mean score and the share actually late are close: 0.17 against 0.17, 0.52 against 0.49. These scores can be read as percentages, which is why the break-even formula worked.
Now a miss is 32 times as expensive as a check. The break-even threshold is 25 / 800, about 0.03, and cross-validation picks 0.05:
Threshold chosen by 5-fold cross-validation on training rows: 0.05
model TP FP FN TN acc prec rec F1 cost EUR
Same model, threshold 0.05 145 393 6 456 0.60 0.27 0.96 0.42 18,250
It flags over half the deliveries and misses only 6 late ones. Accuracy is 0.60, and it is the cheapest option by far. When costs change, the threshold must change with them.
With class_weight="balanced", the logistic regression at 0.5 now has recall 0.82 instead of 0.44. Look at the calibration block at the end:
mean score share late
0.02 0.01
0.10 0.01
0.26 0.07
0.52 0.17
0.83 0.49
Scores of about 0.5 now mean 17% late. The weights moved every score upward, so the break-even formula no longer applies directly. Cross-validation compensates by choosing a threshold of 0.56. Weighting and threshold tuning are two ways to reach a similar decision; only the unweighted model keeps honest scores.
#Step 8: Use your learning rate from the last topic
This trains the same logistic model with SGDClassifier, plain stochastic gradient descent with a constant step. At 0.01, the result is close to the default solver: AUC 0.88, and a tuned threshold of 0.18. At 1.0, the steps are too big, recall at 0.5 drops to 0.26, the Brier score rises from 0.087 to 0.125, and the cost goes up. Compare this with the working range you wrote down in your gradient descent notes.
Open classification_metrics.png from the VS Code file list. On the left, the precision-recall curves: logistic regression's line sits above the tree's almost everywhere, and both sit well above the dotted "random" line at 0.15. On the right, the calibration curves: logistic regression follows the dotted diagonal closely; the tree's top point, at a mean score of about 0.66, is late only about 46% of the time.
Unit 2 is a maths-and-mechanics unit, so this section is short. The point is that SAP's tools speak the same language, and the decisions you just practised still fall to your team.
SAP Learning describes ISLM being used to train and activate customer-specific predictive models. In S/4HANA Utilities, those models produce a release confidence value for implausible meter readings and outsorted billing documents. That is a classifier score shown to a clerk. The questions from this topic apply unchanged: at what value can a document be released without review, what does a wrong release cost, and how many documents land in the review queue at that value? Using ISLM needs an S/4HANA system with the scenario set up, which this course doesn't assume.
An SAP blog on UnifiedClassification in the hana-ml Python client shows training with a stratified partition and a report of AUC, RECALL, PRECISION, F1_SCORE, ACCURACY, KAPPA and MCC, a confusion matrix with actual and predicted class counts, ROC and gains data, and an HTML model report. The blog was written for an older hana-ml release, so check the current reference before relying on exact output names. You now know how to read every one of those numbers, and why ACCURACY alone can't decide.
SAP Learning describes SAP-RPT-1 classification tasks such as cost centers, payment terms and sales groups, reached through a REST endpoint in SAP AI Core, with a confidence score between 0.0 and 1.0 for each predicted cell. Two things follow from this topic:
Treat the confidence as a score until you've checked it. Hold back labelled rows, send the rest as context, and draw the same calibration table the script prints.
You still choose the threshold. For example: auto-fill the cost center when confidence is high enough, otherwise route to a person. Pick that cut-off from the cost of a wrong auto-fill against the cost of a manual check.
SAP's open-source research variant, sap-rpt-1-oss on SAP-samples, exposes a scikit-learn style SAP_RPT_OSS_Classifier with fit, predict and predict_proba. So the metric functions in the script work on it unchanged. Its README recommends a large GPU and limits the checkpoints to research use, so it isn't a laptop exercise.
Write down the metric contract before training. The positive class, the costs of a miss and a false alarm, the flag volume the team can handle, and the baseline. In the FDE's terms from What an SAP forward deployed engineer does, this is how "move metric M from X to Y" becomes concrete.
Choose thresholds on training or validation data only. Report the test set once. If you keep adjusting after seeing test results, you need a fresh test set.
Split by time for real data. Test on the newest deliveries, as the previous topics explained. Class balance also shifts over time, and so does the best threshold.
Version the threshold with the model. The threshold is a model setting. Store it with the model, log it with each prediction, and change it through the same review as the model.
Watch flag volume. A threshold that is cheap on paper can flood a planner queue. Put the expected flags per day in front of the process owner.
Monitor recall, precision and calibration after go-live. Labels arrive late (you only know a delivery was late after it arrives), so plan when each week's metrics can be computed.
Authorizations. Scores about customers, suppliers or employees are business data. Show them only to users allowed to see the underlying documents, and log who acted on a flag. Unit 7 covers grounding without breaking authorizations.
Explain flags. A planner will ask why a delivery was flagged. Logistic regression weights and small trees can answer; keep that in mind before choosing a more complex model.
Quoting accuracy on imbalanced data. Always print the "always no" baseline next to it.
Leaving the threshold at 0.5. It is a library default, not a business decision.
Tuning the threshold on the test set. The trade-off table is for reading. The choice comes from cross-validation on training rows.
Reading AUC as a hit rate. AUC is a ranking probability. A model with 0.88 AUC can still miss most late deliveries at a bad threshold.
Mixing up the confusion matrix order. scikit-learn's ravel() gives tn, fp, fn, tp. Name the variables.
Using class weights and then reading scores as probabilities. Weights distort calibration, as Step 7 showed.
Calibrating with isotonic regression on small data. scikit-learn warns it can overfit below about 1,000 rows. Use the sigmoid method or more data.
Forgetting the positive class. Precision and recall of "on time" are different numbers from those of "late". State which class is positive.
#Exercise: write the metric contract for late deliveries
You will turn this topic into a one-page metric contract for the next topic in Unit 2, Predicting late deliveries with SAP data. The contract fixes the metric, the costs and the threshold rule before any real data arrives.
In the terminal, inside unit02 with (.venv) showing, run the script with three cost settings and note the chosen threshold, the recall, the precision, the cost and the number of flagged deliveries (TP + FP) for each:
Done when: the contract names the positive class, both costs, a primary metric and a threshold rule; the results table has all three runs with the chosen and break-even thresholds side by side; the --balanced sentence explains the higher threshold; and git log shows the commit.
Pick one answer for each question. The explanation appears after you choose.
1Which metrics can you compute before choosing any threshold?
Answer: B. ROC AUC and average precision measure how well the scores rank late above on-time deliveries, and the Brier score compares scores with outcomes directly. Precision, recall, F1, accuracy and cost all need flags, which need a threshold.
2In the script, confusion_matrix(y_true, flagged).ravel() returns four numbers. In what order?
Answer: D. scikit-learn orders rows and columns by label, 0 then 1, so the flattened matrix is true negatives, false positives, false negatives, true positives. The script names each variable to avoid mixing them up.
3A miss costs 250 EUR and a planner check costs 50 EUR. With calibrated scores, what is the break-even threshold, and why?
Answer: A. Flag when the expected cost of not flagging, p × 250, is at least the 50 EUR check. That gives p ≥ 50 / 250 = 0.2. Cross-validation on the training rows found 0.21, which agrees.
4A colleague picks the threshold by choosing the cheapest row in the test-set trade-off table. What is wrong with that?
Answer: C. Anything chosen by looking at the test set is tuned to it. scikit-learn's guide warns against tuning the threshold on the data used to train, and the same logic applies to test data. Use cross-validation on training rows, then report the test set once.
5After switching on class_weight="balanced", recall at 0.5 jumped from 0.44 to 0.82, but deliveries scored about 0.5 were late only 17% of the time. What happened?
Answer: B. Balanced weights make rare late rows count more in training, which raises every late score. Ranking can stay good, but scores stop meaning percentages. Cross-validation compensated by choosing 0.56.
6SAP-RPT-1 returns a confidence of 0.0 to 1.0 per predicted cost center. You want to auto-fill high-confidence predictions. What do you do first?
Answer: D. A confidence score is only a percentage if the calibration table says so on your data. The cut-off is then a cost decision, exactly like the late-delivery threshold. AUC is a ranking probability, not a hit rate.
7Your model beats the baseline on cost, but at the chosen threshold it flags 400 deliveries a day, and planners can check 150. What should you do?
Answer: D. The cost model assumed every flag gets checked. If only 150 can be, the unchecked flags behave like misses. Flag volume is a guard rail in the metric contract, so the threshold or the work queue must respect it.
8A hana-ml report for a late-delivery model shows ACCURACY 0.92 and RECALL 0.34 for the late class. How do you read it?
Answer: C. On imbalanced data, accuracy is dominated by the common class. A recall of 0.34 means most real positives go unflagged. Look at recall and precision for the positive class, and at the baseline.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Classification: ROC and AUC (Google Machine Learning Crash Course)— ROC plots true positive rate against false positive rate across thresholds; AUC as the chance a random positive is ranked above a random negative; 0.5 is random; precision-recall curves for imbalanced data; pick the threshold by error costs
Tuning the decision threshold for class prediction (scikit-learn 1.9 user guide)— predict uses a probability threshold of 0.5 by default; TunedThresholdClassifierCV (default scoring balanced accuracy, 5-fold stratified cross-validation) and FixedThresholdClassifier; never tune the threshold on the training data; business scores with make_scorer
Probability calibration (scikit-learn 1.9 user guide)— meaning of a well-calibrated score; calibration curves; LogisticRegression well calibrated by default, random forests and naive Bayes not; CalibratedClassifierCV sigmoid and isotonic, isotonic can overfit below about 1,000 rows; Brier score and log loss
Introduction to SAP-RPT-1 (SAP Learning)— classification examples such as cost centers, payment terms and sales groups; confidence scores between 0.0 and 1.0 for each predicted cell; REST endpoint through SAP AI Core
sap-rpt-1-oss (SAP-samples on GitHub)— SAP_RPT_OSS_Classifier with scikit-learn style fit, predict and predict_proba; checkpoints limited to research use; large GPU recommended