Orchestrate

Predicting late deliveries with SAP data

Turn S/4HANA outbound delivery records into a late-delivery model you can trust, with an honest label, no leakage, a time-based test and a planner worklist.

Updated Sep 30, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

This topic joins the rest of Unit 2 together. You have seen what learning from examples means, how a model is trained and how to judge its flags by cost. Now the examples come from the shape of real SAP records: outbound deliveries in S/4HANA.

Three decisions make or break a late-delivery model, and none of them is the algorithm:

  1. What counts as "late". S/4HANA stores several dates on a delivery. You have to pick one comparison and write it down. This topic uses "goods left the warehouse after the planned goods issue date", because both dates sit on the delivery itself.
  2. What the model may know. A model that predicts lateness at order time may only use facts known at order time. Some fields on a delivery are only filled after it ships. If they sneak into training, the model looks perfect and is useless. This is called leakage.
  3. How you test it. The test must imitate real use: train on the past, test on the future. A shuffled test mixes next month's patterns into training and flatters the model.

Get those three right and a plain model is useful. Get one wrong and no model will save you.

Why it matters to the business

Order-to-cash depends on goods leaving on time. A late goods issue can mean a missed truck, a missed delivery window, a penalty, or revenue slipping into the next period.

Planners already know some deliveries are risky. What they lack is a ranked list each morning: which of today's 300 open deliveries deserve a look? That is the product of this topic, a worklist of open deliveries with a score and a "check" flag.

In the deep layer, on 3,000 made-up deliveries shaped like SAP data, such a worklist cuts the cost of late deliveries from 32,000 EUR to 22,700 EUR over three months of test data. It does that with about 11 flags a week, a load one planner can carry. The same run also shows the traps. A model given one field that is only updated at goods issue scored a perfect 1.00 in testing. On the open deliveries, the ones planners care about, it flagged nothing at all.

That gap between a perfect test score and a useless tool is the main risk for a leader to understand. It is also why these projects should be judged on a time-based test and on the live worklist, not on a slide with one accuracy number.

How SAP does it

As of September 2026, there are three routes, and your team may use more than one:

  • Embedded in S/4HANA. SAP publishes an asset called SAP S/4HANA Predictive Delivery Delays, describing machine learning that anticipates delays in order deliveries. A partner blog from NTT DATA describes the matching sales scenario and a Predicted Delivery Delay app, trained inside S/4HANA through Intelligent Scenario Lifecycle Management (ISLM). For purchasing, SAP's product page for Supplier Delivery Date Prediction describes predicting delivery dates for purchase order items from historical data, also through ISLM.
  • Pretrained for tables: SAP-RPT-1. SAP Learning describes a model that predicts a column of a business table from example rows, without a training project. You send labelled deliveries as context and mark the rows to predict.
  • Your own model. Extract deliveries through the Outbound Delivery API, train a model on SAP BTP or elsewhere, and write a worklist back. That is what the deep layer builds, on made-up data.

Whichever route you choose, the three decisions above stay with your team. SAP's scenario defines its own target and inputs. Check that its definition of "late" matches the one your process owner means.

Which route fits: a decision guide

Your situation Start with Why
S/4HANA, standard delivery process, you want it inside the planner's apps The embedded scenario, through ISLM Runs where the data lives; training and activation are part of the standard lifecycle
Late supplier deliveries in procure-to-pay Supplier Delivery Date Prediction Same idea, pointed at purchase order items
You want a quick answer on a table, without a training project SAP-RPT-1 Learns from example rows at request time
Your own definition of late, extra data (carrier events, weather), or a non-standard process Your own model on extracted data Full control of label, features and threshold
Any of the above A time-based test and a cost-based threshold The only honest way to compare routes

Questions to ask

  • Which dates define "late" here, and who in the business agreed to that definition?
  • Is the model predicting at order time, at delivery creation, or the day before goods issue? Which fields exist at that moment?
  • Were any fields used that change when the delivery ships, such as statuses or "last changed" dates?
  • Was the test set the most recent period, with training strictly before it?
  • What are the recall, precision and cost on that test period, against "flag nothing"?
  • How many flags per week does the chosen threshold produce, and can the planners handle them?
  • How often will the model be retrained, and who notices when a route or carrier suddenly changes?
  • Who may see the scores? Delivery data carries customer names, addresses and volumes.

Common misconceptions

  • "We have years of deliveries, so the data is ready." The records exist; the label doesn't. Someone has to define late and confirm that the dates are filled consistently.
  • "More fields make a better model." Fields filled after goods issue make a worse one. In the deep layer, one such field produced a perfect test score and an empty worklist.
  • "A shuffled 75/25 split is standard, so it's fine." For data that changes over time, a shuffle lets the model peek at the future. Test on the latest period.
  • "The model will spot new problems by itself." It only knows patterns from its training period. When a route got worse after the cut-off, the time-split model didn't rank that route near the top. Retraining and monitoring catch this; the model doesn't.
  • "A late-delivery model needs a complex algorithm." In the deep layer, plain logistic regression beat gradient boosting in time-ordered cross-validation. Definitions and data matter more.

Key terms

  • Outbound delivery: the S/4HANA document for shipping goods to a customer, created from a sales order.
  • Planned goods issue date: the date goods are scheduled to leave the warehouse.
  • Actual goods movement date: the date goods issue was actually posted.
  • Label: the answer the model learns to predict, here "late" or "on time".
  • Feature: a fact the model uses, such as shipping point or number of items.
  • Leakage: training with information that would not exist when the prediction is made.
  • Time-based split: train on older records, test on newer ones.
  • Drift: the patterns in the data change after training, for example a route gets slower.
  • Worklist: the ranked list of open deliveries a planner checks.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A team proposes a late-delivery model. Which decision should be settled before anyone picks an algorithm?

    Answer: B. The label is the definition the whole model learns. S/4HANA stores several dates on a delivery, so "late" must be chosen and agreed. Algorithm choice mattered far less in this topic's example.
  2. 2A vendor's model scores a perfect 1.00 on its test data. What is the most likely explanation to check first?

    Answer: C. Perfect scores on business data are a warning sign. In the deep layer, one field that changes at goods issue gave a perfect test score, and the model then flagged no open deliveries at all.
  3. 3Why should the test set be the most recent deliveries rather than a random quarter?

    Answer: D. In production the model always predicts deliveries it has never seen, from a later period. A shuffled split mixes future patterns into training and makes the model look better than it will be.
  4. 4A route got slower after the training cut-off. What does that mean for a model that is not retrained?

    Answer: A. A model only knows the patterns in its training period. Changes after the cut-off are drift, and they are caught by monitoring and retraining, not by the model itself.
  5. 5Your S/4HANA system offers an embedded delivery-delay scenario. What should you still check before relying on it?

    Answer: C. SAP's scenario defines its own target and inputs. The business definition of late, the flag volume and the cost trade-off are still your team's decisions, whichever route produces the score.
  6. 6Which report best shows whether a late-delivery model is worth deploying?

    Answer: B. The latest period imitates real use, and cost and flag volume connect the model to the planners' work. Accuracy on a shuffled sample hides both the rare late class and any drift.
Deep layer · 40 min read

Mental model: every row is a snapshot at prediction time

A training row is not "everything the system knows about this delivery". It is what the system knew at the moment you would have made the prediction, plus the answer that came later.

flowchart LR
  C[Delivery created] --> P[Prediction moment]
  P --> G[Goods issue posted]
  G --> A[Answer known: late or on time]
  P -.-> F[Features: only fields filled by now]
  A -.-> L[Label]

Three rules follow, and the rest of this topic is about applying them to SAP data:

Rule What it prevents In the script
Define the label from dates, in writing A model of the wrong question Late = goods issue after planned goods issue date
Use only fields that exist at the prediction moment Leakage --leak shows what one wrong field does
Test on the future, train on the past Flattering scores Time split, and TimeSeriesSplit inside cross-validation

How it works

The data: outbound deliveries in S/4HANA

S/4HANA exposes outbound deliveries through the OData V2 service API_OUTBOUND_DELIVERY_SRV, at the service path /sap/opu/odata/sap/API_OUTBOUND_DELIVERY_SRV. The field names below come from SAP Cloud SDK's generated package for this service (version 2.1.0), which is built from the service's metadata.

  • A_OutbDeliveryHeader: one record per delivery. Key DeliveryDocument.
  • A_OutbDeliveryItem: the lines, reached through the navigation property to_DeliveryDocumentItem.

The fields this topic uses:

Field Entity Used as
CreationDate Header Orders rows in time; start of the planned lead time
PlannedGoodsIssueDate Header Label (planned side); lead time; weekday
ActualGoodsMovementDate Header Label (actual side). Never a feature
ShippingPoint, ShippingType, ShippingCondition, DeliveryPriority, ProposedDeliveryRoute Header Features
HeaderGrossWeight, HeaderWeightUnit Header Feature, with a unit check
ActualDeliveryQuantity Item Feature, summed per delivery; item count
LastChangeDate Header Leakage demo only

Dates arrive in the OData V2 JSON form /Date(milliseconds since 1 January 1970)/, as you saw in Calling your first SAP API. The script converts them. Codes such as shipping types and routes are configuration: their values and meanings differ between systems, so the model treats them as labels, never as numbers.

Where the dates come from

The planned dates are not guesses. SAP Learning's lesson on delivery and transportation scheduling explains that S/4HANA schedules backwards from the customer's requested delivery date, through the goods issue date, to the material availability and transportation planning dates. The gaps are the pick/pack time, loading time, transit time and transportation lead time. Loading and pick/pack times come from the shipping point; transit time and transportation lead time come from the route.

That is why shipping point and route are strong features: they carry the planned durations. It also means a late goods issue often says the durations in configuration no longer match reality, which is a finding worth taking to the process owner in itself.

Defining "late"

A delivery has several "late" candidates. Choose one per model and write it into the metric contract from Classification and the metrics that matter.

Definition Compares Available from Catch
Late goods issue (this topic) ActualGoodsMovementDate vs PlannedGoodsIssueDate Delivery header Says the warehouse was late, not that the customer was
Late arrival Proof of delivery or carrier arrival vs DeliveryDate ProofOfDeliveryDate exists on the header, but is only useful if your process records proof of delivery; otherwise carrier data Often empty or outside S/4HANA
Late against the customer's request Delivery vs the requested date on the sales order Sales order schedule lines Needs a second API and a join

Three details matter in the code:

  • Only rows with an actual goods issue date get a label. The rest are open deliveries. They are not "on time"; they are unknown, and they become the worklist.
  • Days, not hours. Both fields are dates. The header also has ActualGoodsMovementTime; this topic ignores time of day.
  • Don't decode statuses by habit. OverallGoodsMovementStatus exists, but status values can vary by system and release. The script uses "has an actual goods movement date" instead.

Features known at prediction time

The prediction moment here is delivery creation. So the features are what a delivery shows when it is created: where it ships from, how, on which route, how heavy, how many items, how many days until planned goods issue, and the weekday of that date.

The dangerous fields are the ones that change as the delivery moves through picking, goods issue and billing:

Field Safe at creation? Why
ShippingPoint, ProposedDeliveryRoute, ShippingCondition Yes Set when the delivery is created
PlannedGoodsIssueDate Yes Scheduled at creation (though it can be rescheduled; see Pitfalls)
ActualGoodsMovementDate No It is the answer
OverallGoodsMovementStatus, OverallPickingStatus No Change as work is done
LastChangeDate No Moves when goods issue is posted
ProofOfDeliveryDate, BillingDocumentDate No Belong to steps after shipping
ActualDeliveryRoute Check May be set later than the proposed route

scikit-learn's guide defines leakage in exactly these terms: using information that would not be available at prediction time, which makes development scores optimistic and real performance worse. LastChangeDate is the sneaky kind. It sounds like housekeeping, but in the sample data it moves on the day goods issue is posted, so it holds the answer. Step 6 shows the effect.

Splitting by time, and deliveries in flight

A time split picks a cut-off date. Test rows are deliveries created on or after it. Training rows are deliveries whose outcome was known before it: their goods issue was posted before the cut-off.

Some deliveries were created before the cut-off but shipped after it. On the cut-off date, their answer did not exist yet. The script leaves them out of both sets and prints how many ("in flight"). That mirrors a real retraining run, which can only learn from outcomes that have happened.

Inside the training set, model choice and threshold tuning use cross-validation too. A normal k-fold shuffles time again, so the script uses TimeSeriesSplit. Its documentation describes it: the rows must be in time order, each training fold is a superset of the one before, and the test fold always comes later. The default is 5 splits.

flowchart LR
  subgraph TR["Training rows, oldest to newest"]
    F1[Fold 1: train, test] --> F2[Fold 2: train more, test later] --> F5[Fold 5]
  end
  F5 --> M[Pick model and threshold]
  M --> T[Test once on rows after the cut-off]
  T --> W[Score open deliveries: worklist]

Two models and a threshold

The script compares two classifiers, both wrapped in a pipeline so that encoding and scaling are learned from training rows only. That is the practice scikit-learn's pitfalls guide recommends: never call fit on test data, and let a pipeline apply each step to the right subset.

  • Logistic regression, from the last topic. Its weights are readable, which planners appreciate.
  • HistGradientBoostingClassifier, a tree ensemble that captures interactions such as "heavy and on route R17102". Its reference says it is much faster than GradientBoostingClassifier from about 10,000 rows and handles missing values natively. Defaults include learning_rate=0.1 and max_iter=100; the script limits trees to max_depth=3.

Text columns become 0/1 columns with OneHotEncoder(handle_unknown="ignore"). The ignore matters: a route that first appears after the cut-off must not crash the scoring.

The model with the higher cross-validated average precision wins. Then TunedThresholdClassifierCV searches the threshold that minimizes the business cost, with the same time-ordered folds. Its cv parameter accepts a splitter object, and by default it refits on the whole training set once the threshold is found.

Drift

The sample data has a built-in surprise: from mid-July 2026, route R17102 gets much slower. The cut-off falls in late June. So the model trains on a world where that route was ordinary, and is tested on one where it isn't. That is drift, and it is normal in logistics: a carrier changes, a warehouse is rebuilt, a customer's volumes jump. A time split shows its cost. A shuffled split hides it, as Step 5 shows.

Build it yourself: from delivery records to a planner worklist

You will run one script that reads outbound deliveries in the exact JSON shape of API_OUTBOUND_DELIVERY_SRV, defines the label, builds features known at creation, splits by time, picks a model and a cost-based threshold, and writes a ranked worklist of open deliveries. Then you will break it on purpose twice, to see why the rules matter.

Before you start: complete Set up your computer for this course and Set up for Unit 2: data science tools. They give you the orchestrate-course folder, its .venv, the unit02 folder, and numpy, pandas, scikit-learn, requests and python-dotenv. This walkthrough doesn't repeat those steps. Doing Classification and the metrics that matter first helps: this script reuses its costs and threshold search.

flowchart LR
  S[OData JSON: sample, file or sandbox] --> T[One row per delivery]
  T --> L[Label from two dates]
  T --> F[Features known at creation]
  L --> X[Time split]
  F --> X
  X --> M[Model and threshold, time-ordered folds]
  M --> R[Test report]
  M --> W[at_risk_deliveries.csv]

What you need

  • The course folder, .venv and unit02 folder from the two setup topics.
  • About 40 minutes. The main path needs no account, no key and no internet: --sample builds 3,000 made-up deliveries on your computer.
  • Optional: your free SAP Business Accelerator Hub key in .env as SAP_API_KEY, from the Unit 1 setup, to read the sandbox's deliveries. Free.
  • scikit-learn 1.5 or newer, for TunedThresholdClassifierCV. The Unit 2 setup installs a current version.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 2 folder:

    cd unit02

Step 2: Save the script

  1. In VS Code's file list, right-click unit02, choose New File and name it predict_late.py.
  2. Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""Predicting late deliveries with SAP data: from outbound delivery records to a planner worklist.

How to run (from the unit02 folder, with the course .venv turned on):
    python predict_late.py --sample                   # 3,000 made-up deliveries shaped like API_OUTBOUND_DELIVERY_SRV
    python predict_late.py --sample --random-split    # shuffle instead of splitting by time, and compare
    python predict_late.py --sample --leak            # add a column that secretly contains the answer
    python predict_late.py --sample --cost-miss 500   # change a cost from your metric contract
    python predict_late.py                            # read deliveries from SAP's sandbox (needs SAP_API_KEY in .env)
    python predict_late.py --file my_extract.json     # read a saved OData V2 JSON file

Made-up data never leaves your computer. Without --sample the script only reads (GET) from the sandbox.
"""
import argparse
import json
import os
import re
import sys
from pathlib import Path

import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyClassifier
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import average_precision_score, confusion_matrix, make_scorer, roc_auc_score
from sklearn.model_selection import (TimeSeriesSplit, TunedThresholdClassifierCV, cross_val_score,
                                     train_test_split)
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

HERE = Path(__file__).parent
SERVICE = "/sap/opu/odata/sap/API_OUTBOUND_DELIVERY_SRV"
SANDBOX = "https://sandbox.api.sap.com/s4hanacloud" + SERVICE
TODAY = pd.Timestamp("2026-09-18")   # the made-up data's "today": deliveries due after this are still open

CATEGORICAL = ["ShippingPoint", "ShippingType", "ShippingCondition", "DeliveryPriority",
               "ProposedDeliveryRoute", "PlannedGIWeekday"]
NUMERIC = ["Items", "TotalQuantity", "GrossWeight", "PlannedLeadDays"]
MIN_LABELLED = 200                   # below this, training and testing honestly isn't possible


# ---------------------------------------------------------------- 1. Get the data

def odata_date(ts: pd.Timestamp) -> str:
    """Write a date the way OData V2 JSON does: /Date(milliseconds since 1 Jan 1970)/."""
    return f"/Date({int(ts.timestamp() * 1000)})/"


def make_sample_payload(rows: int, seed: int = 7) -> dict:
    """Made-up outbound deliveries in the same JSON shape the OData V2 API returns.

    Field names follow A_OutbDeliveryHeader and A_OutbDeliveryItem. All codes and values are made up.
    """
    rng = np.random.default_rng(seed)
    start = pd.Timestamp("2025-10-01")
    created = np.sort(start + pd.to_timedelta(rng.integers(0, 360, size=rows), unit="D"))
    points = rng.choice(["1010", "1020", "1710"], size=rows, p=[0.45, 0.35, 0.20])
    routes = {"1010": ["R10101", "R10102"], "1020": ["R10201", "R10202"], "1710": ["R17101", "R17102"]}
    results = []
    for i in range(rows):
        point = points[i]
        route = rng.choice(routes[point])
        shipping_type = rng.choice(["01", "02", "03"], p=[0.6, 0.3, 0.1])
        condition = rng.choice(["01", "02"], p=[0.8, 0.2])
        priority = rng.choice(["02", "01"], p=[0.85, 0.15])
        lead = int(rng.integers(1, 11))
        creation = pd.Timestamp(created[i])
        planned_gi = creation + pd.Timedelta(days=lead)
        n_items = int(rng.poisson(2.5)) + 1
        items, weight = [], 0.0
        for j in range(n_items):
            qty = int(rng.integers(1, 40))
            item_weight = round(qty * float(rng.gamma(2.0, 4.0)), 3)
            weight += item_weight
            items.append({
                "DeliveryDocument": str(80000000 + i),
                "DeliveryDocumentItem": f"{(j + 1) * 10:06d}",
                "Material": f"TG{int(rng.integers(10, 30))}",
                "Plant": point,
                "ActualDeliveryQuantity": f"{qty:.3f}",
                "ItemGrossWeight": f"{item_weight:.3f}",
                "ItemWeightUnit": "KG",
            })
        # A made-up rule plus noise decides how late goods issue is. From mid-July 2026, route R17102 gets much worse.
        risk = (-2.4 + 0.22 * n_items + 0.0012 * weight - 0.20 * lead
                + (0.9 if point == "1710" else 0.0)
                + (-0.8 if condition == "02" else 0.0)
                + (-0.5 if priority == "01" else 0.0)
                + (0.6 if planned_gi.dayofweek == 4 else 0.0)            # goods issue planned on a Friday
                + (2.5 if route == "R17102" and creation >= pd.Timestamp("2026-07-15") else 0.0)
                + rng.logistic(0, 0.8))
        delay = int(rng.geometric(0.45)) if risk > 0 else -int(rng.integers(0, 2))
        actual_gi = planned_gi + pd.Timedelta(days=delay)
        is_open = planned_gi > TODAY
        last_change = creation + pd.Timedelta(days=int(rng.integers(0, 2))) if is_open else actual_gi
        results.append({
            "DeliveryDocument": str(80000000 + i),
            "SalesOrganization": point,
            "ShippingPoint": point,
            "ShippingType": shipping_type,
            "ShippingCondition": condition,
            "DeliveryPriority": priority,
            "ProposedDeliveryRoute": route,
            "ShipToParty": f"10{int(rng.integers(100, 160)):05d}",
            "CreationDate": odata_date(creation),
            "PlannedGoodsIssueDate": odata_date(planned_gi),
            "DeliveryDate": odata_date(planned_gi + pd.Timedelta(days=int(rng.integers(1, 4)))),
            "ActualGoodsMovementDate": None if is_open else odata_date(actual_gi),
            "LastChangeDate": odata_date(last_change),
            "HeaderGrossWeight": f"{weight:.3f}",
            "HeaderWeightUnit": "KG",
            "to_DeliveryDocumentItem": {"results": items},
        })
    return {"d": {"results": results}}


def fetch_sandbox(key: str, top: int, max_pages: int) -> dict:
    """Read outbound deliveries with their items from the SAP Business Accelerator Hub sandbox (GET only)."""
    try:
        import requests  # imported here so --sample works even without the library
    except ModuleNotFoundError:
        sys.exit("The requests library is missing. Run: pip install -r requirements.txt (or use --sample).")
    session = requests.Session()
    session.headers.update({"APIKey": key, "Accept": "application/json"})
    url = f"{SANDBOX}/A_OutbDeliveryHeader"
    params = {"$top": top, "$expand": "to_DeliveryDocumentItem", "$format": "json"}
    results = []
    for _ in range(max_pages):
        try:
            response = session.get(url, params=params, timeout=30)
        except requests.exceptions.RequestException as err:
            sys.exit(f"Could not reach sandbox.api.sap.com ({type(err).__name__}). "
                     "Check your network or proxy, or run with --sample.")
        if response.status_code != 200:
            hint = {401: "the key is missing or wrong: check SAP_API_KEY in .env",
                    403: "the key is not allowed to call this API",
                    404: "the service path is wrong or the API is not on the sandbox"}
            sys.exit(f"Stopped: HTTP {response.status_code}, {hint.get(response.status_code, 'see the message above')}.")
        body = response.json()["d"]
        results.extend(body.get("results", []))
        if "__next" not in body:
            break
        url, params = body["__next"], None   # the next link already carries the query
    return {"d": {"results": results}}


# ---------------------------------------------------------------- 2. Turn records into a table

def parse_date(value):
    """'/Date(1759190400000)/' -> Timestamp. Empty values become NaT (not a time)."""
    if not value:
        return pd.NaT
    match = re.search(r"/Date\((-?\d+)", str(value))
    return pd.to_datetime(int(match.group(1)), unit="ms") if match else pd.NaT


def flatten(payload: dict) -> pd.DataFrame:
    """One row per delivery: header fields plus a few item totals."""
    rows = []
    for header in payload["d"]["results"]:
        items = (header.get("to_DeliveryDocumentItem") or {}).get("results", [])
        rows.append({
            "DeliveryDocument": header.get("DeliveryDocument"),
            "ShipToParty": header.get("ShipToParty"),
            "ShippingPoint": header.get("ShippingPoint") or "(blank)",
            "ShippingType": header.get("ShippingType") or "(blank)",
            "ShippingCondition": header.get("ShippingCondition") or "(blank)",
            "DeliveryPriority": header.get("DeliveryPriority") or "(blank)",
            "ProposedDeliveryRoute": header.get("ProposedDeliveryRoute") or "(blank)",
            "CreationDate": parse_date(header.get("CreationDate")),
            "PlannedGoodsIssueDate": parse_date(header.get("PlannedGoodsIssueDate")),
            "ActualGoodsMovementDate": parse_date(header.get("ActualGoodsMovementDate")),
            "LastChangeDate": parse_date(header.get("LastChangeDate")),
            "GrossWeight": float(header.get("HeaderGrossWeight") or 0),
            "WeightUnit": header.get("HeaderWeightUnit") or "",
            "Items": len(items),
            "TotalQuantity": sum(float(item.get("ActualDeliveryQuantity") or 0) for item in items),
        })
    return pd.DataFrame(rows)


def add_label_and_features(df: pd.DataFrame, leak: bool) -> pd.DataFrame:
    """Label = goods issue posted after the planned date. Features = what is known when the delivery is created."""
    df = df.dropna(subset=["CreationDate", "PlannedGoodsIssueDate"]).copy()
    df["GIDelayDays"] = (df["ActualGoodsMovementDate"] - df["PlannedGoodsIssueDate"]).dt.days
    df["Late"] = np.where(df["GIDelayDays"].isna(), np.nan, (df["GIDelayDays"] > 0).astype(float))
    df["PlannedLeadDays"] = (df["PlannedGoodsIssueDate"] - df["CreationDate"]).dt.days
    df["PlannedGIWeekday"] = df["PlannedGoodsIssueDate"].dt.day_name().str[:3]
    if leak:
        # LastChangeDate moves when goods issue is posted, so this column quietly contains the answer.
        df["DaysPlannedGIToLastChange"] = (df["LastChangeDate"] - df["PlannedGoodsIssueDate"]).dt.days
    return df.sort_values("CreationDate").reset_index(drop=True)


# ---------------------------------------------------------------- 3. Split, train, judge

def split_by_time(labelled: pd.DataFrame, share_test: float = 0.25):
    """Train on what was known before a cut-off date; test on deliveries created after it."""
    cutoff = labelled["CreationDate"].quantile(1 - share_test).normalize()
    train = labelled[labelled["ActualGoodsMovementDate"] < cutoff]      # outcome known before the cut-off
    test = labelled[labelled["CreationDate"] >= cutoff]
    in_flight = len(labelled) - len(train) - len(test)
    return train, test, cutoff, in_flight


def preprocessor(numeric):
    return ColumnTransformer([
        ("categories", OneHotEncoder(handle_unknown="ignore", sparse_output=False), CATEGORICAL),
        ("numbers", StandardScaler(), numeric),
    ])


def counts(y_true, flagged):
    tn, fp, fn, tp = confusion_matrix(y_true, flagged, labels=[0, 1]).ravel()
    return int(tn), int(fp), int(fn), int(tp)


def cost_of(y_true, flagged, cost_miss, cost_check):
    _, fp, fn, tp = counts(y_true, flagged)
    return (tp + fp) * cost_check + fn * cost_miss


def negative_cost_per_delivery(y_true, flagged, cost_miss, cost_check):
    return -cost_of(y_true, flagged, cost_miss, cost_check) / len(y_true)


def main() -> None:
    parser = argparse.ArgumentParser(description="Predict late goods issue for outbound deliveries.")
    parser.add_argument("--sample", action="store_true", help="use made-up deliveries; no key or internet")
    parser.add_argument("--file", help="read a saved OData V2 JSON file instead of calling the sandbox")
    parser.add_argument("--rows", type=int, default=3000, help="how many made-up deliveries with --sample")
    parser.add_argument("--top", type=int, default=200, help="deliveries per sandbox page")
    parser.add_argument("--max-pages", type=int, default=5, help="safety limit on sandbox pages")
    parser.add_argument("--cost-miss", type=float, default=250.0, help="EUR lost when a late delivery is not flagged")
    parser.add_argument("--cost-check", type=float, default=50.0, help="EUR for a planner to check one flag")
    parser.add_argument("--random-split", action="store_true", help="shuffle rows instead of splitting by time")
    parser.add_argument("--leak", action="store_true", help="add a column that contains the answer (a lesson)")
    args = parser.parse_args()

    # 1. Get the data: made-up, a saved file, or the sandbox.
    if args.sample:
        payload = make_sample_payload(args.rows)
        saved = HERE / "sample_outbound_deliveries.json"
        source = "made-up sample"
    elif args.file:
        payload = json.loads(Path(args.file).read_text(encoding="utf-8"))
        saved, source = None, args.file
    else:
        try:
            from dotenv import load_dotenv
            load_dotenv()
        except ModuleNotFoundError:
            pass  # without python-dotenv, the key can still come from the environment
        key = os.environ.get("SAP_API_KEY")
        if not key:
            sys.exit("SAP_API_KEY is not set. Add it to .env, or run with --sample.")
        payload = fetch_sandbox(key, args.top, args.max_pages)
        saved = HERE / "sandbox_outbound_deliveries.json"
        source = "SAP Business Accelerator Hub sandbox"
    if saved:
        saved.write_text(json.dumps(payload, indent=1), encoding="utf-8")

    # 2. Turn the records into one row per delivery, with a label and features.
    table = flatten(payload)
    if table.empty:
        print(f"Source: {source}. The answer held no deliveries. That is a valid, empty result: "
              "the call worked but found nothing to learn from.")
        return
    df = add_label_and_features(table, args.leak)
    labelled = df[df["Late"].notna()].copy()
    labelled["Late"] = labelled["Late"].astype(int)
    open_deliveries = df[df["Late"].isna()]
    print(f"Source: {source}. {len(df)} deliveries read"
          + (f", raw JSON saved as {saved.name}" if saved else "") + ".")
    print(f"  With a goods issue date (labelled): {len(labelled)}. Still open: {len(open_deliveries)}.")
    if len(labelled):
        print(f"  {labelled['Late'].mean():.0%} of labelled deliveries had goods issue after the planned date.")
    if df["WeightUnit"].nunique() > 1:
        print(f"  Warning: mixed weight units {sorted(df['WeightUnit'].unique())}; convert before trusting GrossWeight.")
    if len(labelled) < MIN_LABELLED or labelled["Late"].nunique() < 2:
        print(f"\nOnly {len(labelled)} labelled deliveries. That is too few to train and test honestly "
              f"(this script wants {MIN_LABELLED}).\nThe extraction worked; to see the full flow, "
              "run with --sample, or use a larger extract with --file.")
        return

    numeric = NUMERIC + (["DaysPlannedGIToLastChange"] if args.leak else [])
    features = CATEGORICAL + numeric

    # 3. Split. By time unless you ask for the shuffle.
    if args.random_split:
        train, test = train_test_split(labelled, test_size=0.25, random_state=0, stratify=labelled["Late"])
        train = train.sort_values("CreationDate")
        print(f"\nRandom split: {len(train)} to learn from, {len(test)} to test, mixed across all dates.")
    else:
        train, test, cutoff, in_flight = split_by_time(labelled)
        print(f"\nTime split at {cutoff.date()}: learn from {len(train)} deliveries whose goods issue was "
              f"posted before it,\n  test on {len(test)} created on or after it. "
              f"{in_flight} were still in flight at the cut-off and are left out.")
    x_train, y_train = train[features], train["Late"]
    x_test, y_test = test[features], test["Late"]
    print(f"  Late share: {y_train.mean():.0%} in training, {y_test.mean():.0%} in testing.")

    # 4. Compare models by cross-validation on the training rows only, folds in time order.
    candidates = {
        "Logistic regression": make_pipeline(preprocessor(numeric), LogisticRegression(max_iter=1000)),
        "Gradient boosting": make_pipeline(preprocessor(numeric),
                                           HistGradientBoostingClassifier(max_depth=3, random_state=0)),
    }
    folds = TimeSeriesSplit(n_splits=5)
    print("\nModel choice: average precision, 5 time-ordered folds on training rows (higher is better)")
    cv_scores = {}
    for name, model in candidates.items():
        cv_scores[name] = cross_val_score(model, x_train, y_train, cv=folds, scoring="average_precision").mean()
        print(f"  {name:<22}{cv_scores[name]:.2f}")
    best_name = max(cv_scores, key=cv_scores.get)
    print(f"  Chosen: {best_name}")

    # 5. Choose the threshold from the metric contract's costs, again on training rows only.
    scorer = make_scorer(negative_cost_per_delivery, cost_miss=args.cost_miss, cost_check=args.cost_check)
    tuned = TunedThresholdClassifierCV(candidates[best_name], scoring=scorer, cv=folds).fit(x_train, y_train)
    print(f"  Threshold from costs (miss {args.cost_miss:,.0f} EUR, check {args.cost_check:,.0f} EUR): "
          f"{tuned.best_threshold_:.2f}  (break-even {args.cost_check / args.cost_miss:.2f})")

    # 6. Report once on the test rows.
    scores = tuned.predict_proba(x_test)[:, 1]
    flagged = tuned.predict(x_test)
    tn, fp, fn, tp = counts(y_test, flagged)
    baseline = DummyClassifier(strategy="most_frequent").fit(x_train, y_train)
    print(f"\nTest result ({len(test)} deliveries, {int(y_test.sum())} late)")
    print(f"  ROC AUC {roc_auc_score(y_test, scores):.2f}   average precision {average_precision_score(y_test, scores):.2f}"
          f"   (random scores: 0.50 and {y_test.mean():.2f})")
    print(f"  Flagged {tp + fp}: caught {tp} of {tp + fn} late, {fp} false alarms, "
          f"recall {tp / max(tp + fn, 1):.2f}, precision {tp / max(tp + fp, 1):.2f}")
    if not args.random_split:  # a shuffled test set spreads over the whole year, so a weekly rate means little
        weeks = max(1.0, (test["CreationDate"].max() - test["CreationDate"].min()).days / 7)
        print(f"  About {(tp + fp) / weeks:.0f} flags per week for the planners")
    print(f"  Cost {cost_of(y_test, flagged, args.cost_miss, args.cost_check):,.0f} EUR, against "
          f"{cost_of(y_test, baseline.predict(x_test), args.cost_miss, args.cost_check):,.0f} EUR if nothing is flagged")

    # 7. What drives the score? Logistic regression weights are readable; refit one on the training rows.
    readable = make_pipeline(preprocessor(numeric), LogisticRegression(max_iter=1000)).fit(x_train, y_train)
    names = readable[0].get_feature_names_out()
    weights = pd.Series(readable[-1].coef_[0], index=[n.split("__", 1)[1] for n in names]).sort_values()
    print("\nStrongest signals (logistic regression weights; + pushes towards late)")
    for name, weight in pd.concat([weights.tail(4)[::-1], weights.head(2)]).items():
        print(f"  {name:<32}{weight:+.2f}")

    # 8. Score the open deliveries and write the planner worklist.
    if len(open_deliveries):
        worklist = open_deliveries[["DeliveryDocument", "ShippingPoint", "ProposedDeliveryRoute",
                                    "PlannedGoodsIssueDate"]].copy()
        worklist["PlannedGoodsIssueDate"] = worklist["PlannedGoodsIssueDate"].dt.date
        worklist["Score"] = tuned.predict_proba(open_deliveries[features])[:, 1].round(2)
        worklist["Flag"] = np.where(worklist["Score"] >= tuned.best_threshold_, "CHECK", "")
        worklist = worklist.sort_values("Score", ascending=False)
        worklist.to_csv(HERE / "at_risk_deliveries.csv", index=False)
        print(f"\nOpen deliveries scored: {len(worklist)}, flagged for a check: {(worklist['Flag'] == 'CHECK').sum()}."
              " Top five:")
        print(worklist.head(5).to_string(index=False))
        print("Wrote at_risk_deliveries.csv")


if __name__ == "__main__":
    main()

Step 3: Run it on sample data

In the terminal, still inside unit02, run:

python predict_late.py --sample

It takes a few seconds. What success looks like:

Source: made-up sample. 3000 deliveries read, raw JSON saved as sample_outbound_deliveries.json.
  With a goods issue date (labelled): 2892. Still open: 108.
  14% of labelled deliveries had goods issue after the planned date.

Time split at 2026-06-22: learn from 2119 deliveries whose goods issue was posted before it,
  test on 727 created on or after it. 46 were still in flight at the cut-off and are left out.
  Late share: 13% in training, 18% in testing.

Model choice: average precision, 5 time-ordered folds on training rows (higher is better)
  Logistic regression   0.42
  Gradient boosting     0.37
  Chosen: Logistic regression
  Threshold from costs (miss 250 EUR, check 50 EUR): 0.20  (break-even 0.20)

Test result (727 deliveries, 128 late)
  ROC AUC 0.78   average precision 0.47   (random scores: 0.50 and 0.18)
  Flagged 134: caught 64 of 128 late, 70 false alarms, recall 0.50, precision 0.48
  About 11 flags per week for the planners
  Cost 22,700 EUR, against 32,000 EUR if nothing is flagged

Strongest signals (logistic regression weights; + pushes towards late)
  PlannedGIWeekday_Fri            +0.83
  ShippingPoint_1710              +0.49
  GrossWeight                     +0.49
  ShippingCondition_01            +0.43
  PlannedLeadDays                 -0.70
  PlannedGIWeekday_Sat            -0.44

Open deliveries scored: 108, flagged for a check: 15. Top five:
DeliveryDocument ShippingPoint ProposedDeliveryRoute PlannedGoodsIssueDate  Score  Flag
        80002918          1710                R17102            2026-09-20   0.88 CHECK
        80002970          1010                R10102            2026-09-26   0.50 CHECK
        80002930          1710                R17101            2026-09-21   0.48 CHECK
        80002935          1010                R10102            2026-09-21   0.45 CHECK
        80002972          1710                R17102            2026-09-23   0.44 CHECK
Wrote at_risk_deliveries.csv

Your numbers should match, because the data comes from a fixed random seed. A different last decimal is fine; library versions can cause that. Two files appear in unit02: sample_outbound_deliveries.json, the raw records, and at_risk_deliveries.csv, the worklist.

Step 4: Read the results

  1. The data block. 2,892 deliveries have an actual goods issue date and get a label; 108 are still open and are not treated as "on time". 14% of labelled deliveries left after the planned date. Open sample_outbound_deliveries.json in VS Code and find PlannedGoodsIssueDate and ActualGoodsMovementDate in the first record: the label is just those two dates compared.
  2. The split. The cut-off is 22 June 2026. 2,119 deliveries shipped before it and train the model; 727 created after it test it; 46 were in flight and are left out. Note the late share rises from 13% to 18%: the test period is harder. That is the drift you read about.
  3. Model choice. Logistic regression beat gradient boosting on time-ordered cross-validation, 0.42 against 0.37 average precision. More flexible isn't always better, especially on a few thousand rows.
  4. The threshold. Cross-validation chose 0.20, the same as the break-even 50 / 250 from the last topic.
  5. The test result. ROC AUC 0.78, average precision 0.47 against 0.18 for random scores. At the chosen threshold it catches 64 of 128 late deliveries with 70 false alarms, about 11 flags a week. Cost falls from 32,000 EUR (flag nothing) to 22,700 EUR.
  6. Strongest signals. A planned goods issue on a Friday, shipping point 1710, weight and the standard shipping condition push towards late; a longer planned lead time pushes away. With one-hot columns, each weight compares a value with the others, and small weights such as the Saturday one can be noise. Note what is missing: route R17102, the route that got worse after the cut-off, isn't among the top signals. The model can't know what happened after its training period.
  7. The worklist. 108 open deliveries scored, 15 flagged CHECK. The top one, on route R17102, scores 0.88. Open at_risk_deliveries.csv from the VS Code file list: that ranked table is what a planner would see.

Step 5: Compare with a shuffled split

python predict_late.py --sample --random-split

The test and signals blocks now read:

Test result (723 deliveries, 101 late)
  ROC AUC 0.84   average precision 0.49   (random scores: 0.50 and 0.14)
  Flagged 133: caught 60 of 101 late, 73 false alarms, recall 0.59, precision 0.45
  Cost 16,900 EUR, against 25,250 EUR if nothing is flagged

Strongest signals (logistic regression weights; + pushes towards late)
  ProposedDeliveryRoute_R17102    +0.72
  PlannedGIWeekday_Fri            +0.60
  ShippingPoint_1710              +0.56
  Items                           +0.44
  PlannedLeadDays                 -0.68
  PlannedGIWeekday_Tue            -0.48

AUC 0.84 instead of 0.78, recall 0.59 instead of 0.50, and route R17102 jumps to the strongest signal. The shuffle put deliveries from August and September into training, so the model "knew" about the slow route before the test. On real data, you would report the shuffled numbers, deploy, and then watch the model underperform. The time split's lower numbers are the honest ones. There is no flags-per-week line here, because a shuffled test set spreads over the whole year.

Step 6: Leak the answer on purpose

python predict_late.py --sample --leak

This adds one feature: the days from planned goods issue to LastChangeDate. The test block:

Test result (727 deliveries, 128 late)
  ROC AUC 1.00   average precision 1.00   (random scores: 0.50 and 0.18)
  Flagged 128: caught 128 of 128 late, 0 false alarms, recall 1.00, precision 1.00
  About 10 flags per week for the planners
  Cost 6,400 EUR, against 32,000 EUR if nothing is flagged
...
Open deliveries scored: 108, flagged for a check: 0. Top five:

A perfect score: every late delivery caught, no false alarms. Now look at the worklist: not a single open delivery is flagged. For an open delivery, LastChangeDate hasn't moved to the goods issue day yet, because goods issue hasn't happened. The model learned to read the answer from a field that only contains it after the fact. When you see a near-perfect score on business data, look for a field like this first.

Step 7: Change the costs

python predict_late.py --sample --cost-miss 500

With a miss twice as expensive, the break-even is 0.10 and cross-validation picks 0.13. The test block shows 219 flags, 81 of 128 late deliveries caught, and about 18 flags a week. Before accepting that, ask the planners whether 18 checks a week fit their day. The metric contract from the last topic has a guard rail for exactly this.

Step 8 (optional): Read the sandbox's deliveries

This step checks that the same code works on records from SAP's sandbox. The sandbox holds demo data, so expect far fewer deliveries than a company system.

  1. Check that .env in your course folder has a line SAP_API_KEY="...". If not, follow the key steps in the setup topic.

  2. Optional: look at the API first. Go to https://api.sap.com, search for API_OUTBOUND_DELIVERY_SRV, and open the API with that technical name. Its API Reference lists A_OutbDeliveryHeader and to_DeliveryDocumentItem.

  3. Run, without --sample:

    python predict_late.py
  4. What success looks like. The first lines say Source: SAP Business Accelerator Hub sandbox, how many deliveries were read, and how many have a goods issue date. The raw answer is saved as sandbox_outbound_deliveries.json. With fewer than 200 labelled deliveries, the script stops with a message like this one (the count is an example):

    Only 29 labelled deliveries. That is too few to train and test honestly (this script wants 200).
    The extraction worked. To see the full flow, run python predict_late.py --sample, or use a larger extract with --file.

    Your count will differ. That message is a success: the call, the paging and the flattening worked. A message that the answer held no deliveries is also a valid, empty result.

Step 9: Save your work in Git

From the course folder:

cd ..
git add unit02/predict_late.py
git commit -m "Late-delivery model on outbound delivery records"

The JSON and CSV files are outputs you can recreate, so you don't need to commit them.

What each part of the script does

Part What it does
make_sample_payload Builds made-up deliveries in the OData V2 JSON shape, with items under to_DeliveryDocumentItem, dates as /Date(...)/ and a slow route from mid-July
fetch_sandbox Reads A_OutbDeliveryHeader with $expand=to_DeliveryDocumentItem, sends the key as the APIKey header, follows __next links up to --max-pages
parse_date Turns /Date(ms)/ into a pandas date; empty becomes "not a time"
flatten One row per delivery: header fields plus item count and total quantity
add_label_and_features Late from the two dates (open deliveries stay unlabelled), lead days, weekday; --leak adds the leaky column
split_by_time Cut-off at the 75% point of creation dates; drops deliveries in flight
preprocessor One-hot encodes codes (ignoring unseen values) and scales numbers, inside a pipeline
cross_val_score(..., cv=TimeSeriesSplit(5)) Compares the two models on time-ordered folds of training rows
TunedThresholdClassifierCV Picks the threshold that minimizes cost, on the same folds
Test block ROC AUC, average precision, the confusion counts, flags per week and cost, once
Signals block Logistic regression weights, largest first
Worklist block Scores open deliveries and writes at_risk_deliveries.csv, highest score first

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of the Unit 1 setup, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'sklearn' (or pandas) The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
ImportError: cannot import name 'TunedThresholdClassifierCV' Your scikit-learn is older than 1.5 Run pip install --upgrade scikit-learn, then run the script again
The requests library is missing Only matters for the sandbox Run pip install -r requirements.txt, or use --sample
SAP_API_KEY is not set The script can't find your key Check .env is in the course folder with SAP_API_KEY="..." on its own line, or run with --sample
Stopped: HTTP 401 The key is missing, wrong or incomplete Copy it again with Show API Key on the hub and update .env
Could not reach sandbox.api.sap.com Your network or a company proxy blocks the site Try another network, or ask IT to allow sandbox.api.sap.com. --sample still works
pip shows ProxyError, SSLError or Could not fetch URL Your network blocks PyPI Ask IT for access to pypi.org and files.pythonhosted.org or an internal mirror
Only N labelled deliveries Too few rows with a goods issue date to train Expected on the sandbox. Use --sample, or a larger extract with --file
Warning: mixed weight units Deliveries use different weight units Convert weights to one unit before relying on GrossWeight

The SAP way

Getting the data out

In a company system, the same GET calls go to your S/4HANA host with a communication user or certificate, set up through communication arrangements, as described in Calling your first SAP API. Three practical points for training data:

  • Page through history. A year of deliveries is many pages. Follow __next links, set a safety limit, and store the raw JSON with the extraction date, so a training run can be repeated.
  • Filter on the server. Ask for a date range with $filter on CreationDate, rather than downloading everything. Check the exact filter syntax your system accepts on the API reference.
  • Keep it read-only. The training extract needs GET only. Ask for a communication user limited to reading deliveries.

Embedded: delivery delay scenarios in S/4HANA

SAP's asset SAP S/4HANA Predictive Delivery Delays describes predictive analytics and machine learning that anticipate potential delays in order deliveries. NTT DATA, an SAP partner, describes the matching sales scenario: predicted delays for outbound shipments, shown in a Predicted Delivery Delay app, trained through ISLM, with a catalogue of about 130 standard input fields that can be narrowed during training. That description comes from a partner, not SAP documentation, so confirm names and scope in your release's documentation before you plan around it.

For procure-to-pay, SAP's product page for Supplier Delivery Date Prediction describes predicting delivery dates for purchase order items from historical data, through Intelligent Scenario Lifecycle Management. As of September 2026, the page says AI units are not currently required for it, subject to change.

What this topic gives you for either scenario: the questions to ask of it. What is its target, exactly? Which inputs does it use, and are any filled after the prediction moment? How was it tested? Which threshold turns its prediction into a planner's action?

Pretrained: SAP-RPT-1

SAP Learning describes SAP-RPT-1 as tabular in-context learning: you send a table with labelled context rows and mark the cells to predict with [PREDICT]. It supports classification, returns a confidence between 0.0 and 1.0 per prediction, and runs through a REST endpoint in SAP AI Core. The recommended context is 500 to 2,000 rows for the small model and 4,000 to 8,000 for the large one.

For late deliveries, that means: labelled deliveries from before the cut-off as context, open deliveries with Late set to [PREDICT]. Every rule of this topic still applies. Context rows must not carry leaky columns, the test must use later deliveries, and the confidence needs the calibration check from the last topic before you set a threshold. SAP AI Core access is set up in Unit 5.

Build vs. SAP

Situation Use Why
Learning, or auditing someone else's model This script Every step is visible and repeatable
Standard S/4HANA delivery process, planners live in S/4HANA apps The embedded scenario through ISLM Data stays in the system; lifecycle is standard
Supplier-side delays in purchasing Supplier Delivery Date Prediction Built for purchase order items
A quick first model on an extract, no training pipeline SAP-RPT-1 Learns from context rows at request time
Your own label, carrier or external data, custom worklist Own model on BTP or elsewhere, fed by the API Full control of every decision

Production concerns

  • Agree the label and the prediction moment in writing. Put them in the metric contract next to the costs. A model predicting at order time and one predicting the day before goods issue are different models with different features.
  • Audit every feature for leakage. For each field, ask: when is it filled, and can it change afterwards? Keep the answer in a feature list reviewed with a functional consultant.
  • Snapshot features at prediction time. Today's record of an old delivery shows its final state, not its state at creation. Planned dates can be rescheduled. For production, store the features you scored with, and train on those snapshots.
  • Retrain on a schedule, test by time. Retrain monthly or when monitoring says so, with the newest period as the test. Compare against the previous model before switching.
  • Monitor recall and flag volume. Labels arrive after goods issue, so this week's recall is known next week. Watch the late share by shipping point and route; a jump is drift.
  • Authorizations. Deliveries carry customers, addresses and volumes. Show scores only to users allowed to see the deliveries, and use a read-only technical user for extraction.
  • Clean core. Read through released APIs such as API_OUTBOUND_DELIVERY_SRV, run the model side by side, and write results back through released interfaces or show them in a separate app. Don't modify standard delivery tables.
  • Explain flags. A planner will ask why a delivery was flagged. Keep a readable model, or at least report the top signals per delivery.

Pitfalls

  • Treating open deliveries as on time. No goods issue date means unknown, not "not late". Counting them as on time lowers the late share and teaches the model the wrong thing.
  • Using the final state of a record. Rescheduled planned dates, changed routes and updated weights overwrite what was known at creation.
  • Innocent-looking dates. LastChangeDate leaked the answer in Step 6. Any "changed on", status or follow-on document date deserves suspicion.
  • Shuffled cross-validation inside a time split. A time split is undone if the threshold search shuffles again. Use TimeSeriesSplit.
  • Forgetting deliveries in flight. Training on outcomes that happened after the cut-off leaks the future too.
  • Mixed units. Gross weight in KG and LB in one column is noise. Check HeaderWeightUnit first.
  • Reading codes as numbers. Shipping point 1710 is not bigger than 1010. Encode codes as categories.
  • Chasing the model instead of the process. If one route is always late, the fix may be its transit time in configuration, not a better model.

Exercise: add a feature and write the results note

You will test one new feature honestly and record the results in a note that Unit 6 uses when you wrap this model in an API.

  1. Make a copy of the script to experiment on. In the terminal, inside unit02 with (.venv) showing:

    • Windows (PowerShell):

      Copy-Item predict_late.py predict_late_v2.py
    • macOS / Linux:

      cp predict_late.py predict_late_v2.py
  2. Open predict_late_v2.py. Add the planned transit time as a feature, in three edits:

    • In flatten, below the line "PlannedGoodsIssueDate": parse_date(header.get("PlannedGoodsIssueDate")), add:

      "DeliveryDate": parse_date(header.get("DeliveryDate")),
    • In add_label_and_features, below the line that computes df["PlannedLeadDays"], add:

      df["PlannedTransitDays"] = (df["DeliveryDate"] - df["PlannedGoodsIssueDate"]).dt.days
    • Near the top, change the NUMERIC line to:

      NUMERIC = ["Items", "TotalQuantity", "GrossWeight", "PlannedLeadDays", "PlannedTransitDays"]
  3. Before running, write down in one sentence whether DeliveryDate is known at delivery creation, and why that makes it a safe or unsafe feature.

  4. Run both versions and note the test AUC, recall, precision, flags per week and cost of each:

    python predict_late.py --sample
    python predict_late_v2.py --sample
  5. In VS Code, create unit02/late_delivery_results.md with these sections:

    • Label: the definition of late, in one sentence with the two field names.
    • Prediction moment and features: when the model predicts, and the feature list.
    • Excluded fields: at least three fields you would never use as features, and why.
    • Results: a table with the time split, the shuffled split (Step 5), the leak run (Step 6) and your v2 run.
    • Decision: keep or drop PlannedTransitDays, based on the time-split cost, not the shuffled one.
  6. Save your work:

    git add unit02/predict_late_v2.py unit02/late_delivery_results.md
    git commit -m "Late-delivery results note and transit-time feature test"

Done when: predict_late_v2.py runs with the new feature; the note states the label, the prediction moment, three excluded fields and a results table with all four runs; the keep-or-drop decision cites the time-split numbers; and git log shows the commit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the script, which deliveries get a Late label?

    Answer: B. The label compares the actual goods movement date with the planned goods issue date, so it needs both. Open deliveries have no answer yet; they are scored in the worklist instead of being counted as on time.
  2. 2With --leak, the test score was perfect but no open delivery was flagged. Why?

    Answer: C. The leaky column encodes the answer after the fact. For completed deliveries it gives the outcome away; for open ones the event hasn't happened, so the model sees nothing and scores them low.
  3. 3Why does the time split leave out deliveries created before the cut-off but shipped after it?

    Answer: D. A retraining run on the cut-off date can only learn from outcomes that have already happened. Keeping in-flight deliveries would add answers from after the cut-off to the training data.
  4. 4The shuffled split scored AUC 0.84 and the time split 0.78. Which should you report to the process owner, and why?

    Answer: A. In production the model predicts later deliveries. The shuffle put August and September into training, so the model knew about the slow route before the test. The time split's lower number is the honest one.
  5. 5Why does the script use TimeSeriesSplit inside cross_val_score and TunedThresholdClassifierCV?

    Answer: C. The rows are sorted by creation date, and TimeSeriesSplit always tests each fold on later rows than it trained on. Shuffled folds would reintroduce the look-ahead that the time split removed.
  6. 6Why does the preprocessing use OneHotEncoder(handle_unknown="ignore")?

    Answer: B. Configuration grows over time. A route that didn't exist in training becomes all zeros in its one-hot columns instead of raising an error. Codes are categories, never numbers to compare.
  7. 7A consultant suggests adding OverallPickingStatus as a feature because it improves test AUC. What do you do?

    Answer: D. A model that predicts at delivery creation may only use fields known then. A status that changes during picking describes work done after the prediction moment, which is leakage, however much it lifts the test score.
  8. 8After go-live, late deliveries on one route double, but the model rarely flags that route. What is the most likely cause, and the fix?

    Answer: B. The model scores by patterns from its training period, as the time split showed with route R17102. Monitoring late share by route catches the change, and retraining on the newest outcomes teaches the model the new pattern.
  9. 9You want SAP-RPT-1 to score open deliveries. How do you prepare the request?

    Answer: C. SAP-RPT-1 learns from labelled context rows at request time, and [PREDICT] marks the cells to fill. The rules on leakage and time still apply, so context rows must not carry fields filled after goods issue.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in