Orchestrate

Gradient descent, by hand

How a model actually learns, one small downhill step at a time, and what the learning rate, batch size and loss curve tell you about a training run.

Updated Sep 30, 2026Foundational 8 minDeep 35 min
Foundational layer · 8 min read

The 60-second version

In Machine learning in plain terms you saw that a model learns its rules from past examples. This topic is about how that learning happens for most modern models, from simple price estimates to large language models.

The method is called gradient descent. Picture standing on a hillside in thick fog, trying to reach the valley floor. You can't see the valley. You can only feel which way the ground slopes under your feet. So you take a small step downhill, feel again, and take another step. Repeat until the ground is flat.

In a model:

  • The height is the error. A number called the loss says how wrong the model's predictions are on past examples. High loss means very wrong.
  • Your position is the model's settings. A model is a set of numbers, often called weights. Changing them changes the predictions.
  • Feeling the slope is the gradient. Maths tells the computer which way to nudge each number so the error drops fastest.
  • The step size is the learning rate. Too small and training takes forever. Too big and you leap past the valley and end up higher than you started.

Training a model is just this loop, run thousands or millions of times.

Why it matters to the business

You will almost never write gradient descent yourself. But it sits under most of the AI you will buy or build, and it explains things that otherwise look mysterious in project status meetings.

Take a simple job in order-to-cash: estimating freight cost for a shipment from its weight, so a sales rep can quote a delivered price. A model starts with a bad guess, such as "every shipment costs zero". Gradient descent then adjusts two numbers, a fixed charge and a price per kilogram, step by step, until the estimates match past freight invoices as closely as possible. The deep layer of this topic does exactly that on made-up data.

Knowing the mechanism helps you read four things correctly:

What you hear What it means Why you care
"Training takes six hours on GPUs" The loop runs many steps; each step costs compute Training cost scales with data size and number of steps. Ask how often it must be repeated
"The training run diverged" The steps were too big and the error grew instead of shrinking A configuration problem, usually fixable. Not a sign the use case is impossible
"Loss has flattened out" Further steps no longer reduce the error Training is done. More time won't help; better data or a better model might
"We need to clean and scale the data first" Columns on very different scales, such as kilograms next to a yes/no flag, make the steps unstable Data preparation is real work, not a delay tactic

The main risk for leaders is mistaking "training finished" for "model is good". Gradient descent reduces error on the examples it was given. It says nothing about new cases, which is why the held-back test set and the baseline from the previous topic still decide whether a model is worth using.

How SAP does it

As of September 2026, SAP hides this loop in three different places, depending on the offering:

  • Inside S/4HANA and SAP HANA ("embedded"). SAP Learning describes embedded machine learning that runs in the same stack as S/4HANA, using the SAP HANA Automated Predictive Library (APL) and Predictive Analysis Library (PAL). Training runs in the database, and Intelligent Scenario Lifecycle Management (ISLM) is the self-service tool for training and managing these scenarios.
  • On SAP BTP ("side-by-side"). In SAP AI Core, your team supplies the training code in a container. SAP Learning describes a configuration as "a set of parameters which can be changed for every run". That is where settings such as the learning rate live. AI Core stores the trained model and offers APIs to record metrics, such as the loss, so runs can be compared.
  • Already trained by SAP. SAP-RPT-1 is pretrained for business tables. SAP says it removes the need for "costly and time-consuming model training": you send example rows with your request instead. The large language models in SAP's generative AI hub are also pretrained by their providers. The gradient descent happened before you arrived.

For a leader, the practical question is which of these three your use case needs, because it decides who owns training, tuning and retraining.

Reading a training run: a guide for non-specialists

When a team shows you a loss curve (error on the vertical axis, training steps on the horizontal), you can read it without any maths:

Shape of the curve What is happening Reasonable next question
Drops steeply, then flattens Healthy training that has converged How does it score on held-back data, against the baseline?
Drops very slowly, still falling at the end Steps too small, or training stopped early Would more steps or a larger learning rate help? What does that cost?
Jumps up and down, or climbs Steps too big: the run is unstable Was the learning rate lowered, or the data scaled?
Flat from the start at a high level Nothing is being learned Is the data connected correctly? Is the target column right?
Wobbly but trending down Normal for training on small random batches of rows Is the trend clear over the whole run?

Questions to ask

  • Who trains this model: SAP, the model provider, a partner, or our own team?
  • How long does one training run take, what does it cost, and how often must it be repeated?
  • Can we see the loss curve of the final run? Did it clearly flatten?
  • Which settings were tuned, such as the learning rate and number of steps, and how were they chosen?
  • Were those choices made without looking at the final test set?
  • Is the result reproducible if someone reruns the training next month?
  • If we use a pretrained model such as SAP-RPT-1 or an LLM, which part, if any, do we still train ourselves?

Common misconceptions

  • "Training is a one-shot calculation." For most modern models it is a long series of small corrections. That is why it takes time and compute.
  • "More training always means a better model." Once the loss flattens, extra steps add cost, not quality. On new data, too much fitting can even hurt, as the overfitting example in the previous topic showed.
  • "A failed training run means the idea doesn't work." A run that diverges usually has a step size or data scaling problem, which is fixable.
  • "Low training loss proves the model is good." It proves the model fits the examples it saw. Only held-back data and a baseline prove it is useful.
  • "Pretrained models need no data from us." SAP-RPT-1 skips training, but it still needs good example rows from your data in each request.

Key terms

  • Loss: a single number measuring how wrong the model is on the training examples. Lower is better.
  • Weights (parameters): the numbers inside a model that training adjusts.
  • Gradient: the direction and steepness of the loss for each weight; it says which way to nudge.
  • Learning rate: how big each step is.
  • Step (iteration): one update of the weights.
  • Epoch: one full pass over all training examples.
  • Batch: the rows used to compute one step. All rows, one row, or a small group (mini-batch).
  • Convergence: the point where more steps no longer reduce the loss.
  • Divergence: the loss grows instead of shrinking, usually because steps are too big.
  • Feature scaling: putting input columns on similar scales so steps behave well.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the "downhill in fog" picture of gradient descent, what does the height of the ground stand for?

    Answer: B. Gradient descent walks downhill on the loss. Each step changes the weights so the error on training examples drops. When the ground is flat, the loss has stopped falling.
  2. 2A partner reports that last night's training run "diverged". What is the most likely explanation?

    Answer: C. Divergence means the loss grew. The usual causes are a learning rate that is too large or input columns on very different scales. Both are configuration problems that a team can fix, not proof the idea fails.
  3. 3A loss curve dropped steeply and has been flat for the last half of the run. The team asks for budget to train three times longer. What should you ask?

    Answer: D. A flat curve means training has converged. Extra steps add cost, not quality. The useful question is whether the model beats the baseline on data it never saw.
  4. 4Which SAP option removes the training step for predictions on business tables?

    Answer: A. SAP describes SAP-RPT-1 as pretrained for structured data, so you send example rows at request time instead of training. AI Core and PAL both run training, either your code in containers or algorithms in the database.
  5. 5Your team trains its own model in SAP AI Core. Where do settings like the learning rate belong?

    Answer: B. SAP Learning describes an AI Core configuration as a set of parameters that can change for every run. Keeping the learning rate there, and recording the loss as a metric, lets the team compare runs.
  6. 6Why is "training finished with a low loss" not enough to approve a model?

    Answer: C. Gradient descent only reduces error on the examples it was given. Whether the model is useful is decided on held-back data and against a baseline, exactly as in the previous topic.
Deep layer · 35 min read

Mental model: the loss is a landscape, the gradient points uphill

A model is a function with adjustable numbers, its parameters. For every setting of those numbers you can compute one loss: how wrong the model is on the training data. Picture every possible setting laid out on a map, with the loss as the height. Training means finding the lowest point on that map.

You can't see the whole map. For a model with two parameters you could draw it; for an LLM with billions you can't even store it. What you can compute cheaply is the gradient: at the point where you stand, which direction goes uphill fastest, and how steeply. Gradient descent steps the opposite way:

new parameter = old parameter - learning_rate * gradient

That one line is the heart of training for linear models, logistic regression, neural networks and transformers. The rest of this topic is about what goes into it, and what goes wrong.

How it works

The model and the loss

This topic uses the smallest useful model: a straight line that estimates freight cost from shipment weight.

predicted_freight = w * weight + b

w (the slope, or weight) is the price per kilogram. b (the intercept, or bias) is the fixed charge. Training must find good values for both.

The loss is mean squared error (MSE): for every shipment, take the prediction minus the actual cost, square it, and average. Squaring makes every error positive and punishes big misses more than small ones. Google's Machine Learning Crash Course uses the same loss for its gradient descent walkthrough.

The gradient, derived once

Calculus gives the slope of the MSE with respect to each parameter. For n rows, with error = prediction - actual:

gradient for w = 2 * average(error * weight)
gradient for b = 2 * average(error)

Read them in plain words. If predictions are too high on average, error is positive, the gradient for b is positive, and the update lowers b. If the errors are largest on heavy shipments, the gradient for w is large, and w moves the most. You derive these formulas once, by hand or with a library, and the computer applies them millions of times.

One step, by hand

Take three tiny shipments. To keep the arithmetic easy, weight is in hundreds of kilograms and freight in tens of euros. The true pattern happens to be freight = 2 * weight + 1.

Weight (100 kg) Freight (10 EUR) Prediction at w=0, b=0 Error
1 3 0 -3
2 5 0 -5
3 7 0 -7
  1. Loss. Average of 9, 25 and 49 is 27.67.
  2. Gradient for b. 2 × average(-3, -5, -7) = 2 × -5 = -10.
  3. Gradient for w. 2 × average(-3×1, -5×2, -7×3) = 2 × average(-3, -10, -21) = -22.67.
  4. Step with learning rate 0.1: w = 0 - 0.1 × -22.67 = 2.267, b = 0 - 0.1 × -10 = 1.0.
  5. New loss. Predictions are now 3.27, 5.53 and 7.80. Errors are 0.27, 0.53 and 0.80. Loss is 0.33.

One step cut the loss from 27.67 to 0.33. The next step moves w back to 2.02 and b to 0.89, and the loss drops to 0.005. The script in this topic prints the same kind of table for 200 shipments.

The loop

flowchart LR
  S[Start: w=0, b=0] --> P[Predict every row]
  P --> L[Compute loss]
  L --> G[Compute gradients]
  G --> U[Step: w, b minus lr x gradient]
  U --> C{Loss still falling?}
  C -->|yes| P
  C -->|no| D[Stop: converged]

Google's crash course describes the same loop: compute the loss, find the direction that reduces it, move a small amount that way, and stop at convergence, when more steps don't reduce the loss.

The learning rate

The learning rate scales every step. Google's crash course summarizes the trade-off: too small and the model "can take a long time to converge"; too large and it "bounces around" and never converges.

For this particular model there is a neat way to see why. When the weight column is scaled to average 0 and spread 1 (next section), each step shrinks the distance to the best w and the best b by the same factor, 1 - 2 × learning_rate. So:

Learning rate Factor per step What you will see
0.01 0.98 Crawls: after 60 steps, still far from the answer
0.1 0.8 Smooth, fast convergence
0.5 0 Lands on the answer in one step
0.9 -0.8 Overshoots back and forth, but converges
1.0 -1 Bounces between two points forever
1.1 -1.2 Every step overshoots further: divergence

This neat formula only holds for a one-feature line with a scaled feature. For real models nobody knows the best learning rate in advance. Teams try a few values, watch the loss curve and keep the best. That is why libraries expose it as a setting: eta0 in scikit-learn's SGDRegressor, lr in PyTorch's torch.optim.SGD.

Feature scaling

The gradient for w is multiplied by the feature values. With raw kilograms (values from 20 to about 1,000), the gradient for w comes out hundreds of times larger than the gradient for b. Any learning rate small enough to keep w stable barely moves b. You will see this in the script: without scaling, a learning rate of 0.1 explodes in one step, and a tiny one "converges" to the wrong answer because the fixed charge never gets a chance to move.

The fix is standardization: subtract the column's average and divide by its standard deviation, so every feature has average 0 and spread 1. scikit-learn's documentation states plainly that SGD is sensitive to feature scaling and recommends StandardScaler in a pipeline, fitted on the training data only. Its LogisticRegression reference adds that the sag and saga solvers only converge fast on features of about the same scale.

Batch, mini-batch and stochastic

The gradient formula averages over rows. Which rows?

Variant Rows per step Behaviour Where it is used
Full batch All of them Smooth, exact gradient; each step is expensive on big data Small tables, teaching
Stochastic (SGD) One Very cheap steps, very noisy scikit-learn's SGDRegressor and SGDClassifier
Mini-batch A small group, such as 32 A compromise: cheap and fairly stable Almost all neural network and LLM training

Google's crash course notes that small mini-batches behave like SGD and large ones like full batch. One epoch is one full pass over the training data. With mini-batches, one epoch contains many steps.

Convex and not convex

For a straight line with MSE, the loss landscape is a single bowl. Google's crash course calls it convex: there is one lowest point and gradient descent will find it. You can check the answer with an exact formula, which the script does with numpy.polyfit.

Neural networks, covered in Unit 3, have no such guarantee. Their landscape has many dips, and where you end up depends on the starting point, the learning rate and the batches. The update rule is the same one you are about to run.

Checking your gradient

A wrong gradient formula is a silent bug: training still runs, just badly. The classic test is a numerical gradient check. Nudge a parameter up and down by a tiny amount h, measure how much the loss changes, and divide by 2h. If that brute-force slope matches your formula, the formula is right. The script does this once at the start.

Build it yourself: fit freight cost one step at a time

You will write the gradient descent loop yourself in about 60 lines of Python, fit a freight-cost line to 200 made-up shipments, and then break it on purpose: steps too small, too big, unscaled data and noisy batches. At the end you compare your answer with the exact formula and with scikit-learn.

Before you start: complete Set up your computer for this course and Set up for Unit 2: data science tools. They give you the orchestrate-course folder, its .venv, the unit02 folder, and numpy, scikit-learn and matplotlib. This walkthrough doesn't repeat those steps.

flowchart LR
  G[200 made-up shipments] --> S[Scale weight]
  S --> L[Loop: predict, loss, gradient, step]
  L --> T[Table of w, b, loss]
  L --> E[Compare: exact answer and scikit-learn]
  L --> P[Optional: loss_curve.png]

What you need

  • The course folder, .venv and unit02 folder from the two setup topics.
  • About 40 minutes. No accounts, no API keys, no cost.
  • All data is made up. Nothing is sent anywhere, and the script needs no internet.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 2 folder:

    cd unit02

Step 2: Save the script

  1. In VS Code's file list, right-click unit02, choose New File and name it gradient_descent.py.
  2. Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""Gradient descent by hand: fit freight cost from shipment weight, one small step at a time.

How to run (from the unit02 folder, with the course .venv turned on):
    python gradient_descent.py                   # learning rate 0.1, 60 steps
    python gradient_descent.py --lr 0.01         # too small: slow
    python gradient_descent.py --lr 1.1          # too big: blows up
    python gradient_descent.py --batch 16        # mini-batch: 16 rows per step
    python gradient_descent.py --no-scale        # skip feature scaling and see what breaks
    python gradient_descent.py --plot            # also save loss_curve.png
    python gradient_descent.py --sklearn         # compare with scikit-learn's SGDRegressor

All data is made up. Nothing is sent anywhere.
"""
import argparse
from pathlib import Path

import numpy as np

HERE = Path(__file__).parent


def make_shipments(rows: int, seed: int = 7):
    """Made-up shipments: freight cost in EUR grows with weight, plus noise."""
    rng = np.random.default_rng(seed)
    weight_kg = rng.gamma(2.0, 150.0, size=rows).round(1)
    freight_eur = (35 + 0.11 * weight_kg + rng.normal(0, 12, size=rows)).round(2)
    return weight_kg, freight_eur


def predict(x, w, b):
    """The model: a straight line. w is the slope (weight), b is the intercept (bias)."""
    return w * x + b


def mse(x, y, w, b):
    """Mean squared error: the average of (prediction - actual) squared."""
    return float(np.mean((predict(x, w, b) - y) ** 2))


def gradients(x, y, w, b):
    """Slope of the loss with respect to w and b (calculus done once, by hand)."""
    error = predict(x, w, b) - y
    grad_w = 2 * np.mean(error * x)
    grad_b = 2 * np.mean(error)
    return float(grad_w), float(grad_b)


def check_gradients(x, y, w, b, h=1e-5):
    """Compare the formula with a brute-force estimate: nudge each value a tiny bit."""
    num_w = (mse(x, y, w + h, b) - mse(x, y, w - h, b)) / (2 * h)
    num_b = (mse(x, y, w, b + h) - mse(x, y, w, b - h)) / (2 * h)
    return num_w, num_b


def main() -> None:
    parser = argparse.ArgumentParser(description="Fit a line with gradient descent, step by step.")
    parser.add_argument("--lr", type=float, default=0.1, help="learning rate: how big each step is")
    parser.add_argument("--steps", type=int, default=60, help="how many passes over the data")
    parser.add_argument("--batch", type=int, default=0, help="rows per update; 0 means all rows")
    parser.add_argument("--rows", type=int, default=200, help="how many made-up shipments")
    parser.add_argument("--no-scale", action="store_true", help="use raw kilograms instead of scaled values")
    parser.add_argument("--plot", action="store_true", help="save the loss curve as loss_curve.png")
    parser.add_argument("--sklearn", action="store_true", help="also fit scikit-learn's SGDRegressor for comparison")
    args = parser.parse_args()

    weight_kg, freight = make_shipments(args.rows)
    print(f"{args.rows} made-up shipments. Weight {weight_kg.min():.0f} to {weight_kg.max():.0f} kg, "
          f"freight {freight.min():.0f} to {freight.max():.0f} EUR\n")

    # Feature scaling: turn kilograms into "how many standard deviations from the average".
    mean_kg, std_kg = weight_kg.mean(), weight_kg.std()
    x = weight_kg if args.no_scale else (weight_kg - mean_kg) / std_kg
    y = freight

    # Baseline: always predict the average freight. Any model must beat this loss.
    print(f"Baseline loss (always predict the average, {y.mean():.2f} EUR): {mse(x, y, 0.0, y.mean()):,.1f}")

    # Start from a guess of zero and check the gradient formula once.
    w, b = 0.0, 0.0
    gw, gb = gradients(x, y, w, b)
    nw, nb = check_gradients(x, y, w, b)
    print(f"Gradient check at w=0, b=0: formula ({gw:.3f}, {gb:.3f})  brute force ({nw:.3f}, {nb:.3f})\n")

    rng = np.random.default_rng(0)
    batch = args.batch if 0 < args.batch < len(x) else len(x)
    mode = "full batch" if batch == len(x) else f"mini-batch of {batch}"
    print(f"Learning rate {args.lr}, {args.steps} steps, {mode}, "
          f"{'raw kilograms' if args.no_scale else 'scaled weight'}")
    print("  step        w          b        loss")

    history = []
    report_at = {0, 1, 2, 3, 5, 10, 20, 40, args.steps}
    for step in range(args.steps + 1):
        loss = mse(x, y, w, b)
        history.append(loss)
        if step in report_at:
            print(f"  {step:>4} {w:>10.3f} {b:>10.3f} {loss:>11,.1f}")
        if not np.isfinite(loss) or loss > 1e12:
            print(f"  Loss exploded at step {step}. The steps are too big for this data: "
                  "lower --lr, or scale the feature.")
            break
        if step == args.steps:
            break
        # One pass over the data ("epoch"): shuffle, then one update per batch.
        order = rng.permutation(len(x)) if batch < len(x) else np.arange(len(x))
        for start in range(0, len(x), batch):
            rows = order[start:start + batch]
            gw, gb = gradients(x[rows], y[rows], w, b)
            w -= args.lr * gw          # step downhill for the slope
            b -= args.lr * gb          # step downhill for the intercept

    if np.isfinite(history[-1]) and history[-1] < 1e12:
        if history[-1] > history[0]:
            print("\n  Warning: the loss ended higher than it started. This run diverged,"
                  " so the rule below is meaningless.")
        # Translate back into business units.
        per_kg = w if args.no_scale else w / std_kg
        base = round(b if args.no_scale else b - w * mean_kg / std_kg, 6) + 0.0  # avoids printing -0.00
        exact_slope, exact_base = np.polyfit(weight_kg, freight, 1)   # closed-form answer, for comparison
        print("\nLearned rule:   freight = {:.2f} EUR + {:.4f} EUR per kg".format(base, per_kg))
        print("Exact answer:   freight = {:.2f} EUR + {:.4f} EUR per kg  (least squares, numpy.polyfit)".format(
            exact_base, exact_slope))
        print(f"A 500 kg shipment: predicted {base + per_kg * 500:.2f} EUR")

    if args.sklearn:
        # The same idea, done by a library: scale the feature, then stochastic gradient descent.
        from sklearn.linear_model import SGDRegressor
        from sklearn.pipeline import make_pipeline
        from sklearn.preprocessing import StandardScaler

        model = make_pipeline(StandardScaler(), SGDRegressor(random_state=0))
        model.fit(weight_kg.reshape(-1, 1), freight)
        at_0, at_500 = model.predict(np.array([[0.0], [500.0]]))
        print(f"scikit-learn:   freight = {at_0:.2f} EUR + {(at_500 - at_0) / 500:.4f} EUR per kg  (SGDRegressor)")

    if args.plot:
        import matplotlib
        matplotlib.use("Agg")  # draw into a file, no window needed
        import matplotlib.pyplot as plt

        finite = [v for v in history if np.isfinite(v)]
        plt.figure(figsize=(6, 3.5))
        plt.plot(range(len(finite)), finite, marker="o", markersize=3)
        plt.yscale("log")
        plt.xlabel("step")
        plt.ylabel("loss (mean squared error, log scale)")
        plt.title(f"Loss curve, learning rate {args.lr}")
        plt.tight_layout()
        out = HERE / "loss_curve.png"
        plt.savefig(out, dpi=120)
        print(f"\nSaved {out.name}")


if __name__ == "__main__":
    main()

Step 3: Run it with the default settings

In the terminal, still inside unit02, run:

python gradient_descent.py

What success looks like:

200 made-up shipments. Weight 20 to 992 kg, freight 17 to 135 EUR

Baseline loss (always predict the average, 61.85 EUR): 506.0
Gradient check at w=0, b=0: formula (-39.264, -123.702)  brute force (-39.264, -123.702)

Learning rate 0.1, 60 steps, full batch, scaled weight
  step        w          b        loss
     0      0.000      0.000     4,331.6
     1      3.926     12.370     2,815.6
     2      7.068     22.266     1,845.4
     3      9.580     30.183     1,224.5
     5     13.199     41.584       572.7
    10     17.524     55.210       169.1
    20     19.406     61.138       121.2
    40     19.630     61.843       120.6
    60     19.632     61.851       120.6

Learned rule:   freight = 34.08 EUR + 0.1043 EUR per kg
Exact answer:   freight = 34.08 EUR + 0.1043 EUR per kg  (least squares, numpy.polyfit)
A 500 kg shipment: predicted 86.22 EUR

Your numbers should match, because the data comes from a fixed random seed. A different last decimal is fine.

Step 4: Read the results

  1. Baseline, 506.0. Always guessing the average freight gives this loss. The trained line ends at 120.6, about four times lower, so weight really does explain freight.
  2. Gradient check. The formula and the brute-force estimate agree to three decimals. Your gradient maths is right.
  3. The table. The loss falls steeply at first, then more slowly, and stops changing between step 40 and step 60. That is convergence.
  4. w and b are in scaled units. w = 19.632 means "euros per standard deviation of weight". The script converts it back to euros per kilogram for the "Learned rule" line.
  5. Learned rule versus exact answer. They match. For a straight line there is an exact formula, and gradient descent found the same bottom of the bowl. The made-up data was generated with 35 EUR plus 0.11 EUR per kg and random noise, so 34.08 and 0.1043 are close to the truth without being equal. That gap is the noise, not a bug.

Step 5: Break the learning rate on purpose

Run three more times:

python gradient_descent.py --lr 0.01
python gradient_descent.py --lr 1.0
python gradient_descent.py --lr 1.1

What you will see (the lines that matter):

Learning rate 0.01 ...
    60     13.791     43.447       493.4
Learned rule:   freight = 23.94 EUR + 0.0732 EUR per kg

Learning rate 1.0 ...
     1     39.264    123.702     4,331.6
     2      0.000      0.000     4,331.6

Learning rate 1.1 ...
    20   -733.016  -2309.375 6,189,315.9
  Loss exploded at step 53. The steps are too big for this data: lower --lr, or scale the feature.
  • At 0.01 the loss is still falling after 60 steps. The rule it learned is wrong only because training stopped too soon. Try --lr 0.01 --steps 400 to see it get there.
  • At 1.0 the parameters jump between two points and the loss never moves. This matches the factor of -1 in the learning-rate table above.
  • At 1.1 each step overshoots further than the last, and the loss explodes. That is divergence. With a learning rate just above 1, such as 1.05, the loss grows more slowly and the script warns that the run diverged instead of stopping.

Step 6: Skip the scaling

python gradient_descent.py --no-scale
python gradient_descent.py --no-scale --lr 0.000001 --steps 200

What you will see:

Gradient check at w=0, b=0: formula (-40334.854, -123.702)  brute force (-40334.854, -123.702)
...
     1   4033.485     12.370 1,730,332,118,033.7
  Loss exploded at step 1. ...

Learning rate 1e-06, 200 steps, full batch, raw kilograms
...
   200      0.190      0.005       507.6
Learned rule:   freight = 0.01 EUR + 0.1896 EUR per kg

Look at the gradient check line: the gradient for w is about 300 times the gradient for b, because it is multiplied by raw kilograms. At 0.1 the first step explodes. At a tiny learning rate, w settles but b barely moves from zero. The loss stalls at 507.6, worse than the baseline. The run looks finished, but the answer is wrong: no fixed charge, and a price per kilogram almost twice the real one. Scaling isn't cosmetic.

Step 7: Try mini-batches

python gradient_descent.py --batch 16
python gradient_descent.py --batch 1

With --batch 16, each pass over the data makes about 13 small updates instead of one. The loss is close to its minimum within two passes, then wobbles slightly around 121. With --batch 1 (pure stochastic gradient descent) and the same learning rate, it wobbles much more: the loss at step 60 is 281.8 and the learned rule is off. Noisy steps need a smaller learning rate. This is why libraries offer learning-rate schedules that shrink the step over time; scikit-learn's SGD estimators have a learning_rate option for exactly that.

Step 8: Compare with scikit-learn and draw the curve

python gradient_descent.py --sklearn --plot

The last lines are:

scikit-learn:   freight = 34.10 EUR + 0.1043 EUR per kg  (SGDRegressor)

Saved loss_curve.png

scikit-learn's SGDRegressor, with its own scaling and step schedule, lands on nearly the same line. Open loss_curve.png from the VS Code file list. It shows the steep drop, then the flat part: the "healthy" row of the loss-curve guide in the foundational layer.

Step 9: Save your work in Git

From the course folder:

cd ..
git add unit02/gradient_descent.py
git commit -m "Fit freight cost with gradient descent by hand"

loss_curve.png is an output you can recreate at any time, so you don't need to commit it.

What each part of the script does

Part What it does
make_shipments Makes 200 shipments with a fixed seed: freight is 35 EUR plus 0.11 EUR per kg plus random noise
predict The model: w * x + b
mse The loss: average squared error
gradients The two formulas from "The gradient, derived once"
check_gradients Nudges w and b by a tiny h to confirm the formulas numerically
Scaling lines Standardize weight to average 0, spread 1, unless --no-scale is given
Baseline line Loss of always predicting the average freight
Main loop For each pass: shuffle if using mini-batches, then compute gradients and step w and b downhill
Explosion check Stops cleanly when the loss becomes huge, and warns if a run ends with a higher loss than it started
"Learned rule" block Converts w and b back to euros and kilograms, and compares with numpy.polyfit
--sklearn block Fits StandardScaler plus SGDRegressor in a pipeline, as scikit-learn's guide recommends
--plot block Saves the loss per step as loss_curve.png, on a log scale so late small changes are visible

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of the Unit 1 setup, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'numpy' (or sklearn, matplotlib) The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
pip shows ProxyError, SSLError or Could not fetch URL while installing Your network or company proxy blocks PyPI Try another network, or ask IT for access to pypi.org and files.pythonhosted.org or an internal mirror. The script itself needs no network
can't open file ... gradient_descent.py The terminal is not in unit02, or the file has another name Run cd unit02 from the course folder, and check the file name in VS Code
RuntimeWarning: overflow A very large learning rate or unscaled data blew up the numbers before the check caught it Expected in the "break it" steps. Lower --lr or drop --no-scale
error: argument --lr: invalid float value A typo in the option, such as --lr 0,1 Use a dot for decimals: --lr 0.1
Numbers differ a lot from the page You changed the generator or passed different options Run exactly python gradient_descent.py with the script as published
No API key is asked for Correct: nothing in this topic uses a key or account Nothing to do

Where this shows up in SAP

Unit 2 is a maths-and-mechanics unit, so this section is short. The loop you wrote runs, hidden, in each of SAP's three routes to machine learning.

Embedded: training inside SAP HANA

SAP Learning describes embedded scenarios that run in the same stack as S/4HANA, using SAP HANA's APL and PAL, managed with ISLM. You don't write the loop; you pick an algorithm and its settings, and training runs next to the data. The ideas from this topic still decide whether a run went well: did it converge, was the input scaled, and does it beat a baseline on held-back rows?

Side-by-side: your training code in SAP AI Core

In SAP AI Core, SAP Learning describes training as follows. An executable declares the training inputs and outputs, the containers to run and the resources needed. A configuration binds an executable to a dataset and holds parameters "which can be changed for every run". Training runs in Docker containers on Argo workflows, with resource plans that differ in CPU, GPU and memory. The trained model is stored and registered for deployment, and AI Core provides APIs to register metrics.

Mapped onto this topic:

This topic SAP AI Core equivalent
--lr, --steps, --batch options Parameters in a configuration, changed per run
gradient_descent.py Your training code, packaged in a container image
The printed loss table Metrics registered through AI Core's metrics APIs
CPU on your laptop A resource plan, with GPUs for large models

Pretrained: SAP-RPT-1 and LLMs

As of September 2026, SAP describes SAP-RPT-1 as pretrained for structured data, able to handle classification and regression from example rows sent at request time, and offered in two commercial versions in the generative AI hub plus an open-source release. SAP Learning says it needs no additional training or fine-tuning. The large language models in the generative AI hub are also trained by their providers. In both cases gradient descent has already run, at a scale no project would repeat, and your work moves to choosing good examples and prompts.

Outside SAP: the same loop in libraries

Library What you write What it does for you
scikit-learn SGDRegressor, SGDClassifier fit(X, y) with eta0 and learning_rate options Stochastic gradient descent with step schedules and stopping rules
scikit-learn LogisticRegression fit(X, y) Uses the lbfgs solver by default, a smarter gradient-based method; max_iter defaults to 100
PyTorch torch.optim.SGD zero_grad(), loss.backward(), step() backward() computes the gradients automatically; step() applies the update with lr and optional momentum

PyTorch's three calls are exactly your loop: clear the old gradients, compute new ones, take the step. Unit 3 uses them to train a neural network.

Build, library or service

Situation Use Why
Learning how training works, debugging a strange run Your own loop Every number is visible
A linear or logistic model on a table you already have scikit-learn Tested solvers, scaling pipelines, sensible defaults
Neural networks, custom architectures PyTorch, run locally or in SAP AI Core Automatic gradients and GPU support
A standard S/4HANA prediction scenario Embedded ML with ISLM Training runs next to the data, inside the standard framework
Table predictions without a training pipeline SAP-RPT-1 Pretrained; you supply example rows per request

Production concerns

  • Record every run. Log the learning rate, batch size, number of steps, data version and final loss. Without them, "the model got worse" can't be investigated. In AI Core this means configuration parameters plus registered metrics.
  • Fit scaling on training data only. Compute the average and spread on the training rows, then apply the same numbers to test rows and to live requests. scikit-learn's guide shows exactly this. Scaling with statistics from the test set leaks information into training.
  • Save the scaler with the model. Live predictions must use the same average and spread as training. A model served without its scaler returns confident nonsense.
  • Watch for divergence automatically. A training job should stop and alert when the loss becomes huge or not a number, as the script does, rather than store a broken model.
  • Stopping rules. Stop when the loss stops improving, or when a separate validation set stops improving (early stopping). scikit-learn's SGD estimators offer both a tolerance and an early_stopping option.
  • Cost. Training cost is roughly steps times the cost of one step. Batch size, data size and hardware are the levers. Ask whether the model must be retrained weekly or yearly before choosing hardware.
  • Authorizations and data handling. Training data pulled from S/4HANA is still business data. The access rules from Calling your first SAP API and the previous topic still apply, wherever the loop runs.
  • Reproducibility. Random shuffling and random starting points change results slightly from run to run. Fix seeds where you can, as the script does, so a rerun can be compared.

Pitfalls

  • Unscaled features. The most common reason a simple model trains badly. Step 6 showed a run that looks finished and is simply wrong.
  • Declaring victory on training loss. Low training loss only says the model fits what it saw. Judge it on held-back data against a baseline.
  • Tuning the learning rate on the test set. Choosing the learning rate is a model choice, like choosing tree depth. Use a validation set or cross-validation.
  • Stopping too early. A curve still falling at the last step means training was cut short, not that the model is as good as it gets.
  • Hand-written gradients without a check. A sign error still trains, just badly. Always run a numerical gradient check once.
  • Copying a learning rate between projects. A value that works on one dataset, scale or batch size can diverge on another.

Exercise: find the working range of the learning rate

You will measure how the learning rate and batch size change training, then write down a rule of thumb. The notes feed the classification topic next in this unit, where you will train a logistic regression model and need a starting learning rate.

  1. In the terminal, inside unit02 with (.venv) showing, run the script with each learning rate below and write down the loss at step 10 and whether the run converged, crawled, bounced or exploded:

    python gradient_descent.py --lr 0.03
    python gradient_descent.py --lr 0.1
    python gradient_descent.py --lr 0.3
    python gradient_descent.py --lr 0.5
    python gradient_descent.py --lr 0.9
    python gradient_descent.py --lr 1.05
  2. Find the smallest number of steps at which the default run (learning rate 0.1) reaches a loss of 121.0 or less. Try --steps 15, --steps 20 and --steps 25 and read the last row of each table, then narrow it down one step at a time (for example --steps 21).

  3. Run python gradient_descent.py --batch 1 --lr 0.01 and python gradient_descent.py --batch 1 --lr 0.1. Note which one ends closer to the exact answer.

  4. Run python gradient_descent.py --lr 0.3 --plot and look at loss_curve.png. Then run python gradient_descent.py --lr 0.9 --plot and compare the shape.

  5. In VS Code, create unit02/notes_gradient_descent.md with three short sections:

    • Learning rate: a table of the six runs from step 1, and the range you would call "works well" for this scaled data.
    • Batch size: one sentence on what happened with batch size 1 at the two learning rates.
    • Rule of thumb: one sentence you will use next time you train a model, in your own words.
  6. Save your work:

    git add unit02/notes_gradient_descent.md
    git commit -m "Notes: learning rate and batch size experiments"

Done when: your notes list all six learning rates with the step-10 loss and a label for each; they state the step count from step 2; the batch-size sentence says which learning rate worked better with batch size 1; and git log shows the commit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does one gradient descent update do to a parameter such as w?

    Answer: C. The update is w = w - learning_rate * gradient. The gradient points uphill on the loss, so subtracting it moves downhill, and the learning rate sets how far.
  2. 2In the three-shipment example, the gradient for b at w=0, b=0 is -10. What does the update do, and why?

    Answer: A. The gradient for b is 2 times the average error, and every prediction was below the actual cost. Subtracting 0.1 times -10 adds 1.0 to b, raising all predictions.
  3. 3With a scaled feature, learning rate 1.0 made the parameters jump between two points forever. What explains it?

    Answer: D. For this model with a standardized feature, each step multiplies the distance to the best values by 1 - 2 × learning rate. At 1.0 that factor is -1: the step overshoots by exactly the distance it should have covered.
  4. 4Without scaling, a tiny learning rate gave a final loss of 507.6, worse than the baseline of 506.0. Why?

    Answer: B. The gradient for w is multiplied by raw kilograms. Any step small enough to keep w stable is far too small for b, so the fixed charge stayed near zero. Standardizing the feature puts both gradients on a similar scale.
  5. 5Why does the script compare the gradient formula with a brute-force estimate before training?

    Answer: C. Nudging a parameter by a tiny h and measuring the loss change gives the true slope. If it matches the formula, the maths is right. A sign or factor error would otherwise go unnoticed.
  6. 6Batch size 1 at learning rate 0.1 ended far from the exact answer, while the full batch converged. What would you try first?

    Answer: D. One-row steps are noisy, so a step size that suits the full batch keeps bouncing. Smaller or shrinking steps, or bigger batches that average out the noise, settle it down.
  7. 7Your team trains a model in SAP AI Core and wants to compare runs with different learning rates. What should the setup include?

    Answer: C. SAP Learning describes AI Core configurations as parameters that can change for every run, and AI Core provides APIs to register metrics. That makes runs comparable without rebuilding the image.
  8. 8A model is deployed, but live predictions are wildly off even though test results were good. Training used standardized features. What do you check first?

    Answer: A. A model trained on standardized inputs expects standardized inputs. If the service sends raw kilograms, or recomputes scaling from live data, predictions are nonsense. Save the scaler with the model and apply it at prediction time.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in