Traditional software follows rules that a person writes: "if lead time is under five days, flag the delivery as at risk". Machine learning turns this around. You show the computer many past examples where you already know the answer, and it works out the rules itself.
The result is called a model. You give it a new case, and it returns a prediction, often with a score for how likely that prediction is.
Three ideas carry most of the weight:
It learns from examples with known answers. Past deliveries marked "late" or "on time" teach it what late looks like. No examples with answers, no learning.
It is judged on cases it hasn't seen. A model that only does well on the examples it learned from has memorized, not learned. Teams always hold some examples back to test it.
It must beat a simple starting point. If 90% of deliveries are on time, a "model" that always says "on time" is 90% right and useless. That simple starting point is the baseline.
Many SAP processes are full of rules that people wrote years ago: credit limits, tolerance checks, planning parameters. Rules are clear and auditable, but they miss patterns nobody wrote down. A planner may know that heavy orders from one plant slip more often. No rule in the system says so.
Take the running example of this unit: late outbound deliveries in order-to-cash. Every late delivery can mean a penalty, an expedited freight bill or an unhappy customer. If a model can flag the risky 10% of deliveries a few days early, planners can act on those and ignore the rest.
The value depends on three questions a leader can ask without any maths:
Question
Why it matters
Do we have enough history, with the right answer recorded?
Learning needs examples. If "late" was never recorded consistently, there is nothing to learn from
What would we do with the prediction?
A flag nobody acts on is worth nothing. Name the action and the person first
How much better than today is it?
Compare with the current rule or the "always on time" guess, on data the model has never seen
The main risks are also simple. A model trained on old patterns can be wrong when the business changes: a new carrier, a new plant, a strike. And a model that looked great in a demo may have been tested on the same data it learned from, which proves nothing.
SAP offers machine learning at three levels. As of September 2026:
Built into S/4HANA ("embedded"). Some predictions run inside S/4HANA itself, using machine learning libraries in the SAP HANA database. SAP describes a predictive delivery delays capability for S/4HANA that does exactly the kind of job this unit teaches. Customers manage these scenarios, and create their own, with a framework called Intelligent Scenario Lifecycle Management (ISLM).
Beside S/4HANA ("side-by-side"). Heavier jobs, such as reading images or text, run on SAP Business Technology Platform (BTP), in services such as SAP AI Core, and send results back.
A pre-trained model for business tables. SAP offers SAP-RPT-1, a model built for predictions on table-shaped business data. SAP says it predicts from example rows you send with the request, without a separate training step. SAP lists it as generally available through SAP AI Core.
The SAP Business AI landscape topic shows where these pieces sit. This topic is about the idea underneath all three: learning from examples.
"Machine learning means the system keeps learning on its own." Most business models learn once, from a fixed set of examples, and stay the same until someone retrains them.
"More complex models are always better." A very complex model can memorize its examples and do worse on new ones. This is called overfitting.
"95% accuracy is great." Not if 95% of cases have the same answer. Always compare with the baseline.
"ML replaces the rules we have." Usually it sits beside them. Rules handle what is known; the model handles what is fuzzy.
"Generative AI made classic ML obsolete." Predicting a yes/no or a number from table data is still classic ML's home ground. SAP-RPT-1 is itself a model aimed at this job.
Pick one answer for each question. The explanation appears after you choose.
1What is the core difference between traditional software and machine learning?
Answer: B. Traditional software follows rules a person writes. Machine learning learns the rules from many past examples where the answer is already known. That is why recorded history, with the right answer, is the first thing to check.
2A vendor says their late-delivery model is 92% accurate. What should you ask first?
Answer: D. If 92% of deliveries are on time, always guessing "on time" also scores 92%. And a score measured on the training examples proves nothing. Ask for the baseline and for results on held-back data.
3What is overfitting, in business terms?
Answer: C. An overfitted model looks excellent on the examples it learned from and disappoints on new ones. It is the most common reason a strong demo fails in real use.
4Your team has no consistent record of which past deliveries were late. What is the sensible first step?
Answer: A. Learning needs past examples with the right answer. Without them there is nothing to learn from. A written rule covers the gap while the history builds up.
5When are hand-written rules the better choice?
Answer: B. Rules are clear, testable and easy to change on purpose. That suits known, stable logic that must be explained. Fuzzy, many-factor patterns are where learning helps.
6Which SAP option runs predictions inside S/4HANA, using libraries in the SAP HANA database?
Answer: C. Embedded scenarios run in S/4HANA itself, with machine learning libraries in SAP HANA, and are managed through ISLM. Side-by-side services on BTP handle heavier jobs such as images and text.
7What makes a late-delivery prediction worth money?
Answer: D. A flag nobody acts on is worth nothing. Name the action and the person first, then measure whether late deliveries actually fall compared with the baseline.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 35 min read
#Mental model: a function you fit, not a program you write
In ordinary programming you write a function by hand: inputs go in, your logic runs, an answer comes out. In machine learning you choose a family of functions (for example, "trees of yes/no questions, at most three deep") and let an algorithm pick the member of that family that best matches past examples.
flowchart LR
subgraph Rules
A[Inputs] --> B[Logic you wrote] --> C[Answer]
end
subgraph Learning
D[Past inputs + known answers] --> E[Training algorithm] --> F[Model]
G[New inputs] --> F --> H[Prediction]
end
Everything else in this unit hangs off that picture:
The features are the inputs. The label is the known answer. Together, many rows of them are the training data.
Training searches the family for the function that fits the training data best. The next topic in this unit, on gradient descent, shows how that search works for one family.
The fitted function is only useful if it also works on rows it never saw. That property is called generalization, and it is the whole game.
This course spends most of Unit 2 on supervised classification, because most business predictions are "which of these outcomes will happen?". SAP uses the same split: SAP Learning's material on SAP-RPT-1 describes classification examples such as cost centers or payment terms, and regression examples such as prices and quantities.
scikit-learn's own documentation calls training and testing on the same data "a methodological mistake". A model that simply repeated the labels it had seen would score perfectly and predict nothing. That is overfitting. The fix is to split the data before training:
flowchart LR
D[All labelled rows] --> S{Split}
S -->|75%| T[Training set]
S -->|25%| X[Test set]
T --> M[Train model]
M --> E[Score on test set]
X --> E
The test set plays the role of the future. You never let the model learn from it. If you use it to pick between many models, it slowly stops being "unseen", so larger projects add a third validation set or use k-fold cross-validation, where the training data is split into k parts and each part takes a turn as the check. The classification topic later in this unit uses this.
Before trusting any score, compute what a trivial predictor gets. scikit-learn provides dummy estimators for this purpose. DummyClassifier(strategy="most_frequent") always predicts the most common label. If your model can't beat it clearly, the model has learned nothing useful, whatever its accuracy says.
A decision tree is a good first family because you can read it. It asks a series of yes/no questions ("is lead time at most 8.5 days?") and ends in an answer. scikit-learn describes trees as a white-box model: you can explain any prediction by following the questions.
Trees also show overfitting clearly. Each extra level of questions lets the tree carve the training data more finely. With no limit, it can put every training row in its own box and score 100% on training data. The scikit-learn documentation warns that over-complex trees "do not generalize the data well" and recommends limiting size with max_depth, starting at 3.
Model setting
Training score
Test score
What is happening
Too simple (depth 1)
Low
Low
Underfitting: can't capture the pattern
About right
Good
Good, close to training
Learns the pattern, not the noise
Too complex (no limit)
Near 100%
Lower than it could be
Overfitting: memorized noise
You will see all three rows appear in the script below.
#Build it yourself: rules versus learning on late deliveries
You will compare three ways to flag late deliveries on made-up SAP-shaped data: always guessing, a rule a planner might write, and a decision tree that learns from examples. Then you will make the tree bigger and watch it overfit. It takes one script and about 30 minutes.
flowchart LR
G[Make 2,000 made-up deliveries] --> S[Split: learn 75% / test 25%]
S --> B[Baseline: always 'late']
S --> R[Rule: lead time up to 5 days]
S --> T[Learned tree, depth 3]
B & R & T --> C[Compare accuracy on test rows]
T --> O[Bigger trees: watch overfitting]
T --> P[Score two new deliveries]
In VS Code's file list, right-click unit02, choose New File and name it rules_vs_learning.py.
Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""Rules versus learning: three ways to flag late deliveries, on made-up SAP-shaped data.
How to run (from the unit02 folder, with the course .venv turned on):
python rules_vs_learning.py # 2,000 made-up deliveries
python rules_vs_learning.py --rows 300 # fewer rows: watch overfitting get worse
python rules_vs_learning.py --csv my_deliveries.csv # use your own sample file
All data is made up. Nothing is sent anywhere.
"""
import argparse
from pathlib import Path
import numpy as np
import pandas as pd
from sklearn.dummy import DummyClassifier
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier, export_text
HERE = Path(__file__).parent
PLANTS = ["1010", "1020", "1710"] # made-up plant codes
CARRIERS = ["ROAD", "RAIL", "EXPRESS"] # made-up shipping types
def make_deliveries(rows: int, seed: int = 42) -> pd.DataFrame:
"""Same generator as hello_data.py: made-up deliveries with a 'Late' flag."""
rng = np.random.default_rng(seed)
df = pd.DataFrame({
"Delivery": [str(80000000 + i) for i in range(rows)],
"Plant": rng.choice(PLANTS, size=rows),
"ShippingType": rng.choice(CARRIERS, size=rows, p=[0.6, 0.3, 0.1]),
"Items": rng.integers(1, 12, size=rows),
"WeightKg": rng.gamma(2.0, 150.0, size=rows).round(1),
"LeadTimeDays": rng.integers(1, 15, size=rows),
})
risk = (
0.25 * df["Items"]
+ 0.004 * df["WeightKg"]
- 0.35 * df["LeadTimeDays"]
+ np.where(df["Plant"] == "1710", 1.0, 0.0)
+ np.where(df["ShippingType"] == "EXPRESS", -1.5, 0.0)
+ rng.normal(0, 1.0, size=rows)
)
df["Late"] = (risk > 0).astype(int)
return df
def hand_written_rule(df: pd.DataFrame) -> pd.Series:
"""What a planner might write from experience: short lead time means late."""
return (df["LeadTimeDays"] <= 5).astype(int)
def main() -> None:
parser = argparse.ArgumentParser(description="Compare a rule, a baseline and learned models.")
parser.add_argument("--rows", type=int, default=2000, help="how many made-up deliveries")
parser.add_argument("--csv", help="read deliveries from this CSV instead of making them")
args = parser.parse_args()
if args.csv:
if not (HERE / args.csv).exists():
parser.error(f"can't find {args.csv} next to this script; run the script that makes it first")
df = pd.read_csv(HERE / args.csv, dtype={"Delivery": str, "Plant": str, "Customer": str})
else:
df = make_deliveries(args.rows)
# Features (the inputs) and the label (the answer we want to predict).
label = df["Late"]
inputs = df.drop(columns=["Delivery", "Late"])
features = pd.get_dummies(inputs, dtype=int) # text columns become 0/1 columns
# Hold back a quarter of the rows as a test set the model never sees in training.
x_train, x_test, y_train, y_test, _, raw_test = train_test_split(
features, label, inputs, test_size=0.25, random_state=0, stratify=label
)
print(f"{len(df)} deliveries: {len(x_train)} to learn from, {len(x_test)} held back for testing")
print(f"Share late overall: {label.mean():.2f}\n")
# 1. Baseline: always predict the most common answer.
baseline = DummyClassifier(strategy="most_frequent").fit(x_train, y_train)
base_acc = baseline.score(x_test, y_test)
# 2. Rules: a person writes the logic.
rule_acc = (hand_written_rule(raw_test) == y_test).mean()
# 3. Learning: the computer finds the logic from labelled examples.
tree = DecisionTreeClassifier(max_depth=3, random_state=0).fit(x_train, y_train)
tree_acc = tree.score(x_test, y_test)
always = "late" if baseline.predict(x_test[:1])[0] == 1 else "on time"
print("Accuracy on the held-back deliveries")
print(f" {'Baseline (always ' + repr(always) + ')':<42}{base_acc:.2f}")
print(f" {'Hand-written rule (lead time <= 5 days)':<42}{rule_acc:.2f}")
print(f" {'Learned decision tree (depth 3)':<42}{tree_acc:.2f}")
print("\nThe rules the tree learned (read top to bottom):")
print(export_text(tree, feature_names=list(features.columns), decimals=1))
# 4. Overfitting: a bigger tree memorises the training rows.
print("Tree depth vs accuracy (train = rows it learned from, test = held back)")
print(" depth train test")
for depth in [1, 2, 3, 5, 8, 12, None]:
model = DecisionTreeClassifier(max_depth=depth, random_state=0).fit(x_train, y_train)
name = "no limit" if depth is None else str(depth)
print(f" {name:<10}{model.score(x_train, y_train):.2f} {model.score(x_test, y_test):.2f}")
# 5. Inference: use the learned model on deliveries it has never seen.
new = pd.DataFrame({
"Plant": ["1710", "1010"], "ShippingType": ["ROAD", "EXPRESS"],
"Items": [9, 2], "WeightKg": [620.0, 80.0], "LeadTimeDays": [3, 12],
})
new_features = pd.get_dummies(new, dtype=int).reindex(columns=features.columns, fill_value=0)
chance = tree.predict_proba(new_features)[:, 1]
print("\nTwo new deliveries, scored by the depth-3 tree")
for (_, row), p in zip(new.iterrows(), chance):
print(f" Plant {row.Plant}, {row.ShippingType + ',':<8} {row.Items:>2} items, "
f"{row.WeightKg:>5.0f} kg, {row.LeadTimeDays:>2} days lead time -> chance late {p:.2f}")
if __name__ == "__main__":
main()
What success looks like (the tree printout is shortened here; yours shows all of it):
2000 deliveries: 1500 to learn from, 500 held back for testing
Share late overall: 0.53
Accuracy on the held-back deliveries
Baseline (always 'late') 0.53
Hand-written rule (lead time <= 5 days) 0.68
Learned decision tree (depth 3) 0.76
The rules the tree learned (read top to bottom):
|--- LeadTimeDays <= 8.5
| |--- Items <= 3.5
| | |--- WeightKg <= 423.6
| | | |--- class: 0
| | |--- WeightKg > 423.6
| | | |--- class: 1
...
Tree depth vs accuracy (train = rows it learned from, test = held back)
depth train test
1 0.76 0.74
2 0.77 0.76
3 0.81 0.76
5 0.84 0.80
8 0.92 0.82
12 0.98 0.79
no limit 1.00 0.80
Two new deliveries, scored by the depth-3 tree
Plant 1710, ROAD, 9 items, 620 kg, 3 days lead time -> chance late 0.95
Plant 1010, EXPRESS, 2 items, 80 kg, 12 days lead time -> chance late 0.05
Your numbers should match, because the data comes from a fixed random seed. A different last decimal is fine; library versions can cause that.
Baseline, 0.53. Just over half the made-up deliveries are late, so "always late" is right 53% of the time. Anything worth using must beat this clearly.
Rule, 0.68. The planner's rule is much better than guessing. Rules are not stupid; they capture real knowledge.
Tree, 0.76. The learned tree beats the rule. Read its printout: the first question it learned is "lead time at most 8.5 days?". The planner's instinct about lead time was right, but the data puts the line at a different place, and the tree adds weight and item count.
Some branches end in the same class on both sides. That is normal. The tree splits where it makes the groups purer, even if both groups still lean the same way.
Depth table. Training accuracy climbs steadily to 1.00. Test accuracy rises, peaks around depth 8, then slips. The gap between the two columns is the overfitting you read about. The "no limit" tree is perfect on data it has seen and worse than depth 8 on data it hasn't.
Two new deliveries. This is inference. A heavy, many-item delivery from plant 1710 with three days' lead time gets 0.95. A small express delivery with twelve days gets 0.05. The number is the share of late deliveries in the training rows that ended in the same leaf of the tree, so treat it as a rough score, not a calibrated probability. The classification topic later in this unit deals with that.
#Step 5: Watch overfitting get worse with less data
Run it again with far fewer deliveries:
python rules_vs_learning.py --rows 300
Look at the depth table near the end:
Tree depth vs accuracy (train = rows it learned from, test = held back)
depth train test
1 0.76 0.75
2 0.78 0.76
3 0.82 0.75
5 0.89 0.67
8 0.98 0.71
12 1.00 0.71
no limit 1.00 0.71
With only 225 rows to learn from, the bigger trees memorize almost at once, and their test scores fall well below the small trees. This is why small datasets call for simple models, and why "how many examples do we have?" is the first question in the foundational layer.
If the customer you made "more often late" matters, it appears in the tree printout as a column such as Customer_17100003. You didn't tell the model about that customer; it found it in the examples.
SAP Learning describes embedded machine learning as running in the same stack as S/4HANA, using two SAP HANA libraries: the Automated Predictive Library (APL), aimed at business analysts, and the Predictive Analysis Library (PAL), with algorithms for data scientists. The model trains and predicts inside the database, next to the data.
These scenarios are run through Intelligent Scenario Lifecycle Management (ISLM). SAP's ISLM topic page names two Fiori apps: Intelligent Scenarios, to see and create custom scenarios, and Intelligent Scenario Management, to manage both custom and SAP-shipped ones. That is where the same lifecycle you practised (train, check, use for prediction) is run in a live system. SAP also describes a predictive delivery delays capability for S/4HANA, which is this topic's running example in product form.
For work that needs heavier models, such as images, sentiment or deep learning, SAP Learning describes a side-by-side setup: the model runs on SAP BTP, for example in SAP AI Core, and S/4HANA calls it. ISLM can also manage side-by-side scenarios. Unit 3 sets up BTP; AI Core access comes in Unit 5.
As of September 2026, SAP offers SAP-RPT-1, which it describes as a relational pretrained model for predictions on structured business data. According to SAP's product page and SAP Learning:
It handles classification and regression, the same two supervised tasks in the table above.
It uses in-context learning: you send example rows with known answers along with the rows to predict, and it returns predictions without a separate training step.
It is offered in small and large versions through SAP AI Core, with a free playground for trying it, and an open-source variant for non-commercial research.
SAP's Architecture Center lists it as generally available, deployed through SAP AI Core; commercial use needs AI Core's extended plan.
That changes who trains the model, not the rules of the game. You still need labelled examples, a held-back test set and a baseline to know whether its predictions help. The classification and late-delivery topics later in this unit give you the tools to judge any of these options on equal terms.
Time-based testing. Test on the latest period, not random rows. Deliveries from next month are what the model will face.
Leakage. Never use a column that is only known after the fact. A goods-issue date that is later than planned already tells you the delivery was late. A model with such a feature looks brilliant in testing and is useless in real use.
Drift. Carriers, plants and customers change. Track the model's accuracy on new data every month and compare with the baseline. Plan who retrains it.
Authorizations. Training data pulled from S/4HANA must respect the same access rules as the source. Extracting delivery data for a notebook is still extracting customer data.
Explainability. A planner will ask "why is this one flagged?". Simple models such as small trees can answer directly; complex ones need extra tooling to explain their answers.
Clean core. Custom models should live beside S/4HANA or in the standard ISLM framework, not in modified standard code.
Scoring on training data. The "no limit" tree scored 1.00 there. That number means nothing.
Skipping the baseline. Accuracy without a baseline hides useless models, especially when one answer dominates.
Tuning on the test set. If you try twenty depths and keep the best test score, the test set is no longer unseen. Use a validation set or cross-validation for choices.
Treating codes as numbers. Plant 1710 is not "bigger" than plant 1010. Turn codes into one-hot columns, as the script does.
Trusting the score as a probability. A tree's leaf share is a rough score. Calibration comes later in this unit.
Believing the tree's rule is the truth. It is the pattern in these examples. With other examples, the split points move.
You will add a validation check to choose the tree depth without touching the test set, then write down what you found. The output feeds the late-delivery prediction topic at the end of this unit.
Copy rules_vs_learning.py to choose_depth.py in unit02.
At the top, add this import below the other sklearn imports:
from sklearn.model_selection import cross_val_score
In the depth loop, after model = ..., add a line that computes the average 5-fold cross-validation accuracy on the training rows only:
Change the print line in the loop so it also shows cv, for example f" {name:<10}{model.score(x_train, y_train):.2f} {cv:.2f} {model.score(x_test, y_test):.2f}", and add cv to the header line.
Run python choose_depth.py and python choose_depth.py --rows 300.
For each run, pick the depth with the best cv column; if two tie, pick the smaller tree. Then look at its test score.
Create unit02/notes_ml_plain_terms.md and write four lines: the baseline accuracy, the rule accuracy, the depth you chose and why, and that depth's test accuracy, for each run.
Save your work:
git add unit02/choose_depth.py unit02/notes_ml_plain_terms.md
git commit -m "Choose tree depth with cross-validation"
Done when:choose_depth.py prints a table with train, cv and test columns for both runs; your notes name a depth chosen from the cv column, not the test column; and git log shows the commit.
Pick one answer for each question. The explanation appears after you choose.
1In supervised learning, what are features and what is the label?
Answer: B. Features are the inputs, such as lead time and weight. The label is the known answer in past rows, such as Late. Training fits a function from features to label.
2Why does the script split the data before training anything?
Answer: D. Testing on the training data is a methodological mistake: a model that repeats seen labels scores perfectly and predicts nothing. The held-back 25% plays the role of new deliveries.
3The depth table shows training accuracy 1.00 and test accuracy 0.80 for the "no limit" tree. What does the gap tell you?
Answer: A. A perfect training score with a lower test score is overfitting. Depth 8 did better on unseen rows, so the extra complexity was learning noise, not pattern.
4What does DummyClassifier(strategy="most_frequent") give you?
Answer: C. It always predicts the majority label. Here that is "late", scoring 0.53. A model that can't clearly beat it has learned nothing useful.
5With --rows 300, the bigger trees score worse on test data than with 2,000 rows. Why?
Answer: B. Fewer rows give a flexible model less pattern and relatively more noise to fit. Small datasets call for simpler models, which is why "how many examples?" comes first.
6You try 20 tree depths and keep the one with the best test score. What is wrong with that?
Answer: D. Choosing by test score leaks the test set into the model design. Use a validation set or k-fold cross-validation on the training rows to choose, and keep the test set for the final check.
7A colleague adds "actual goods issue date minus planned date" as a feature and test accuracy jumps to 0.99. What should you do?
Answer: C. A late goods issue already tells you the delivery is late, and it isn't known at prediction time. This is leakage: brilliant in testing, useless in real use.
8How do embedded and side-by-side machine learning differ in SAP?
Answer: B. SAP Learning describes embedded ML as running in the same stack as S/4HANA with APL and PAL, and side-by-side as running on BTP services such as SAP AI Core for heavier cases like images and deep learning. ISLM manages both.
9SAP-RPT-1 predicts from example rows without a training step. Does that remove the need for a test set and baseline?
Answer: D. In-context learning changes who fits the model, not how you judge it. You still compare its predictions on unseen rows with a baseline, exactly as with the tree.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Decision Trees (scikit-learn 1.9 docs)— trees are simple to interpret (white box); over-complex trees overfit; use max_depth to control size, start at 3
SAP-RPT-1 (sap.com product page)— relational pretrained model for classification and regression on business tables; in-context learning with no model training; small and large versions; playground; open-source variant