Orchestrate

Neural networks from scratch

What a neural network is, why hidden layers let it learn patterns a straight line can't, and how to build, train and check one yourself on SAP-style invoice data.

Updated Sep 30, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

A neural network is a prediction model built from many small, simple calculators called units (or neurons). Each unit does one thing: it multiplies its inputs by some numbers, adds them up, and passes on the result, but only after a small twist.

Units are arranged in layers. The first layer reads the inputs, such as the price difference on a supplier invoice. The last layer gives the answer, such as "likely to be blocked". The layers in between are hidden layers. They build their own intermediate signals, such as "price is well above the order" or "quantity is above what we received".

The twist in each unit is what matters. Without it, stacking layers gives you nothing more than a single straight-line rule. With it, the network can learn bends, corners and combinations that a straight line can't draw.

Training works exactly like the gradient descent you did by hand: measure the error, work out which way to nudge every number, take a small step, repeat. The only new piece is a bookkeeping method called backpropagation, which works out those nudges for every layer at once.

Large language models, image readers and the embedding models in the next topic are all neural networks. They are much bigger, but the building block is the one in this topic.

Why it matters to the business

Most of the AI your company will buy runs on neural networks. Knowing what they are good and bad at helps you judge proposals.

Take the procure-to-pay running example: supplier invoices blocked in the three-way match. In SAP S/4HANA, invoice verification compares the invoice with the purchase order and the goods receipt. SAP Learning explains that when a variance exceeds the upper tolerance limit, the invoice is blocked for payment. Price variances and quantity variances each have their own tolerance keys, such as PP and DQ.

That gives a pattern with a corner in it: an invoice is blocked if the price is too high or the quantity is too high. A single straight-line rule can't draw that corner well. In the deep layer, a straight-line model catches only 62% of the blocked invoices in made-up test data. A small neural network with 8 hidden units catches 90%.

That example teaches a second lesson, and it is the more important one for a leader. You would never train a network to rediscover your own tolerance rules. They are configured in the system; read the configuration. The example uses a known rule only so you can check what the network learned. Neural networks earn their place when nobody can write the rule down: text, images, scanned documents, and patterns across many columns.

Where neural networks help Where they usually don't
Reading text, such as notes, emails and item descriptions Rules that are already configured, such as tolerance limits
Scanned documents and images Small tables where a simpler model scores just as well
Patterns across many columns that interact Decisions that must be explained line by line to an auditor
Tasks where a pretrained model already exists Use cases with no labelled history and no pretrained model

How SAP does it

As of September 2026, SAP offers neural networks in three places. SAP's Architecture Center page on classic machine learning (last updated April 2026) lays out when to use each:

  • Pretrained, for business tables: SAP-RPT-1. SAP describes it as a relational pretrained transformer, a kind of neural network. You send example rows with each request instead of training. SAP's guidance is to start here for classification and regression on tables.
  • Inside SAP HANA: the Predictive Analysis Library (PAL). The hana-ml Python client for SAP HANA includes multi-layer perceptron classes, the textbook neural network this topic builds. Training runs in the database, next to the data.
  • Your own network: SAP AI Core. SAP's guidance is to choose AI Core when deep learning or large-scale neural networks are needed, or custom models in frameworks such as PyTorch. Your team supplies the training code, and AI Core runs it in containers, with GPU options.

The language models in SAP's generative AI hub are neural networks too, trained by their providers. Unit 4 opens them up.

For a leader, the practical question is the same as in Unit 2: who trains the network, and who keeps it working?

Neural network or not? A decision guide

Your situation Sensible first choice Why
The rule is known and configured in SAP No model; use or fix the configuration Cheaper, exact and auditable
A table of business data, a yes/no or number to predict A simple model, or SAP-RPT-1 Often as good as a network, and faster to deliver
Text, documents or images A pretrained neural network Nobody can write these rules by hand
A pattern with interactions that simple models miss A small network, compared against a simple model Prove the gain on held-back data before paying for it
Very large data and a custom architecture Your own network in SAP AI Core Full control, full responsibility

Questions to ask

  • What does a simple model score on the same held-back data? How much does the network add?
  • Is there already a pretrained model, from SAP or elsewhere, for this task?
  • Is the rule we are "learning" already written down in configuration?
  • Who retrains the network when the data changes, and how often?
  • How will we explain a single prediction to a user or an auditor?
  • What does training and running it cost, and does it need GPUs?
  • If two training runs give different results, which one ships, and why?

Common misconceptions

  • "A neural network works like a brain." It borrows the name. Each unit is a weighted sum and a simple function, nothing more.
  • "More layers always means better." In the deep layer, 2 hidden units do as well as 32 on held-back data. Bigger networks cost more and can memorize noise.
  • "Neural networks don't need clean data." They are sensitive to how inputs are scaled and to label errors, like any model.
  • "A network will find our business rules for us." If the rule is configured, read it. Learning it back from data is slower and less exact.
  • "Training is repeatable." Networks start from random numbers. Two runs can end in different places; the deep layer shows one that gets stuck.

Key terms

  • Neural network: a model made of layers of simple units, trained by gradient descent.
  • Unit (neuron): multiplies inputs by weights, adds a bias, then applies an activation function.
  • Hidden layer: a layer between the inputs and the output that builds its own intermediate signals.
  • Activation function: the "twist" in each unit that lets a network draw bends and corners. ReLU is the most common.
  • Backpropagation: the method that works out how to nudge every weight in every layer.
  • Weights and biases (parameters): the numbers training adjusts.
  • Multi-layer perceptron (MLP): the classic network of fully connected layers, the kind built in this topic.
  • Deep learning: neural networks with many layers.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What makes a neural network able to learn patterns a straight-line rule can't?

    Answer: B. Without an activation function, stacked layers collapse into one straight-line rule. The small twist in each unit is what lets hidden layers combine into shapes such as the corner in the invoice-blocking pattern.
  2. 2How is a neural network trained?

    Answer: C. Training is the same loop as in Unit 2: measure the error, find the direction that reduces it, take a small step. Backpropagation is the bookkeeping that finds that direction for every layer at once.
  3. 3A team proposes a neural network to predict which supplier invoices S/4HANA will block for price or quantity variances. What is your first question?

    Answer: D. SAP blocks an invoice for payment when a variance exceeds the configured upper tolerance limit. A known, configured rule should be read, not learned back from data. Neural networks earn their place where nobody can write the rule down.
  4. 4Your use case is a yes/no prediction on a table of business data. Per SAP's own guidance, where do you start?

    Answer: A. SAP's Architecture Center advises starting with RPT-1 for classification and regression on tables, because it needs no training. AI Core is for cases that really need deep learning or custom frameworks.
  5. 5When is SAP AI Core the right home for a neural network?

    Answer: C. SAP's guidance is to choose AI Core for deep learning, large-scale neural networks, or custom models in frameworks such as PyTorch. The team then owns the training code and its upkeep.
  6. 6A vendor says their network with 32 hidden units is "clearly better" than one with 2. What do you ask for?

    Answer: D. A bigger network can fit its training data better and still do no better, or worse, on new cases. In the deep layer, 2 hidden units match 32 on held-back data. Only held-back results and a simple baseline settle it.
  7. 7Two training runs of the same network give quite different results. What does that tell you?

    Answer: B. Training starts from random weights, and a network's error landscape has more than one valley. Teams should fix seeds, record runs, and choose on held-back data. The deep layer shows a run that gets stuck.
Deep layer · 40 min read

Mental model: logistic regression, stacked, with a bend in between

You already know the output end of a neural network. In Classification and the metrics that matter a model computed a weighted sum and squeezed it through a sigmoid into a probability. That is one unit.

A neural network puts a layer of such units in front of it. Each hidden unit computes its own weighted sum and passes it through an activation function, usually ReLU: keep the number if it is positive, otherwise output zero. The output unit then combines the hidden units' signals.

hidden_j = ReLU(w_j1 * price_var + w_j2 * qty_var + b_j)      for each hidden unit j
P(blocked) = sigmoid(v_1 * hidden_1 + ... + v_H * hidden_H + c)

The bend in ReLU is the whole trick. Google's crash course puts it plainly: linear operations performed on linear operations are still linear. Take ReLU out and any stack of layers collapses into one weighted sum. Keep it, and each hidden unit contributes one "hinge" that the output can combine with others into corners and curves.

Training is unchanged from Gradient descent, by hand: weight = weight - learning_rate * gradient. Backpropagation is how you get the gradient for weights that sit behind other layers.

How it works

The data: invoices blocked in the three-way match

SAP Learning describes how S/4HANA invoice verification handles variances. Tolerance limits are set in Customizing, per tolerance key: for example PP for price variances and DQ for quantity variances. If a variance exceeds the upper limit, the invoice is blocked for payment. If it falls below the lower limit, the system only issues a message. The block applies to the whole invoice, even if only one item varies.

This topic uses made-up invoices with one item each and two inputs:

Input Meaning Range in the data
price_var % the invoice price is above (+) or below (-) the purchase order price about -15 to +15
qty_var % the invoiced quantity is above (+) or below (-) the goods receipt quantity about -15 to +15

The label is "blocked": price_var > 5 or qty_var > 3, with 2% of labels flipped at random to stand in for manual blocks and data errors. The 5% and 3% limits are made up; real limits depend on your configuration. Negative variances never block, which matches the "below the lower limit, only a message" behaviour.

Why a straight line fails here

Draw the invoices on a chart with price variance across and quantity variance up. The blocked region is an L shape: everything right of the price limit plus everything above the quantity limit. Logistic regression draws one straight line across that chart. Whatever angle it picks, it cuts through the L and gets one of the two arms wrong. In this topic's script it catches only 62% of blocked invoices.

Two ReLU units are enough to fix it. Here is a network designed by hand, to show the idea:

h1 = ReLU(price_var - 5)        zero until price is 5% over, then grows
h2 = ReLU(qty_var - 3)          zero until quantity is 3% over, then grows
P(blocked) = sigmoid(-2 + 4*h1 + 4*h2)
Invoice h1 h2 Weighted sum P(blocked)
price +2%, qty +1% 0 0 -2 0.12
price +7%, qty 0% 2 0 6 1.00
price 0%, qty +5% 0 2 6 1.00
price -8%, qty 0% 0 0 -2 0.12

Each hidden unit has learned one "arm" of the L. The output adds them. Training finds weights like these on its own; you will see it do so with 2 hidden units.

The forward pass

Running inputs through the layers is called the forward pass. For a batch of rows it is two matrix multiplications:

flowchart LR
  X[Inputs<br/>price_var, qty_var] --> H[Hidden layer<br/>8 units: weighted sum, ReLU]
  H --> O[Output unit<br/>weighted sum, sigmoid]
  O --> P[P blocked]
  P --> L[Loss<br/>cross-entropy]

Counting parameters: each of the 8 hidden units has 2 weights and 1 bias (24), and the output has 8 weights and 1 bias (9). That is 33 numbers to learn. An LLM has billions, arranged differently, but each one is learned the same way.

Activation functions

Google's crash course lists three common activation functions:

Function Formula Output range Where you meet it
Sigmoid 1 / (1 + e^-x) 0 to 1 The output unit, for a probability
Tanh tanh(x) -1 to 1 Older networks; a common option in libraries
ReLU max(0, x) 0 and up Hidden layers in most modern networks

Google notes that ReLU is easier to compute and less likely to suffer from vanishing gradients (below) than sigmoid and tanh.

The loss: cross-entropy

For a yes/no answer, the loss is binary cross-entropy: for each row, -log(p) if the invoice was blocked and -log(1 - p) if it wasn't, averaged. A confident wrong answer costs a lot; a confident right answer costs almost nothing. A model that says 0.5 for everything scores 0.693, a number you will see in the script.

Backpropagation: the chain rule, done backwards

The output unit's weights are easy: their gradient is the same one you would compute for logistic regression. The hidden weights are harder, because they affect the loss only through the output. Backpropagation handles this by starting at the loss and passing blame backwards, layer by layer:

  1. At the output: blame = (p - y) / n. With sigmoid and cross-entropy together, the slope at the output's weighted sum simplifies to exactly this.
  2. Output weights: gradient = hidden signals times blame.
  3. Pass blame back: each hidden unit gets its share, blame * its outgoing weight.
  4. Through ReLU: a hidden unit that output zero passes no blame back (ReLU is flat there).
  5. Hidden weights: gradient = inputs times the hidden unit's blame.
sequenceDiagram
  participant I as Inputs
  participant H as Hidden layer
  participant O as Output
  participant L as Loss
  I->>H: forward: weighted sums, ReLU
  H->>O: forward: weighted sum, sigmoid
  O->>L: compare with the label
  L-->>O: blame = p - y
  O-->>H: blame x output weights, zero where ReLU was off
  H-->>I: gradients for the hidden weights

Google's crash course calls backpropagation the most common training algorithm for neural networks, and the thing that makes gradient descent feasible for many layers. PyTorch automates it. Its autograd records each operation in a graph as the forward pass runs, then applies the chain rule when you call loss.backward(). The script does it by hand first, then shows PyTorch getting the same numbers.

Three ways training goes wrong

Google's crash course names three failure cases. You will trigger two of them on purpose.

Problem What happens Usual fix
Vanishing gradients In deep networks, blame shrinks as it passes back, so early layers barely learn ReLU instead of sigmoid or tanh in hidden layers
Exploding gradients Blame grows as it passes back, and training never settles Lower learning rate, batch normalization
Dead ReLU units A unit's weighted sum stays below zero, it always outputs 0, and no blame reaches it again Lower learning rate, or a ReLU variant such as LeakyReLU

A fourth problem is specific to how you start. If every weight starts at the same value, every hidden unit computes the same thing and gets the same update, forever. With all zeros and ReLU it is worse: no blame flows back at all. That is why networks start from small random numbers.

Not one valley, but many

For logistic regression the loss is a single bowl. For a neural network it isn't. scikit-learn's user guide lists this as the first drawback of its multi-layer perceptron: the loss is non-convex, has more than one local minimum, and different random starts can give different results. The same guide warns that the models are sensitive to feature scaling and need tuning of hidden units, layers and iterations. The script lets you see all three.

Build it yourself: a neural network for blocked invoices

You will write a neural network in about 150 lines of numpy, with backpropagation by hand, and train it to predict which made-up invoices get blocked. You will compare it with a straight-line model, break it in three ways, then train the same network with PyTorch and confirm both give identical numbers.

Before you start: complete Set up your computer for this course and Set up for Unit 3. They give you the orchestrate-course folder with its .venv, numpy and matplotlib from Unit 2, PyTorch, and the unit03 folder. This walkthrough doesn't repeat those steps.

flowchart LR
  D[2,000 made-up invoices] --> S[Scale with training rows]
  S --> N[numpy network<br/>forward, backprop, step]
  N --> R[Accuracy, recall, precision]
  N --> T[Optional: same network in PyTorch]
  N --> F[Optional: invoice_net.json and decision_map.png]

What you need

  • The course folder, .venv and unit03 folder from the setup topics.
  • About 45 minutes. No accounts, no API keys, no cost.
  • All data is made up. Nothing is sent anywhere, and the script needs no internet.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 3 folder:

    cd unit03

Step 2: Save the script

  1. In VS Code's file list, right-click unit03, choose New File and name it neural_net.py.
  2. Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""A neural network from scratch: learn which supplier invoices get blocked for payment.

How to run (from the unit03 folder, with the course .venv turned on):
    python neural_net.py                   # 8 hidden units, trained with numpy only
    python neural_net.py --hidden 0        # no hidden layer: plain logistic regression
    python neural_net.py --hidden 2        # a very small network
    python neural_net.py --zeros           # start every weight at zero and see what happens
    python neural_net.py --lr 20           # steps far too big
    python neural_net.py --torch           # train the same network with PyTorch and compare
    python neural_net.py --plot            # also save decision_map.png
    python neural_net.py --save            # also save the trained network to invoice_net.json

All data is made up. Nothing is sent anywhere.
"""
import argparse
import json
from pathlib import Path

import numpy as np

HERE = Path(__file__).parent
PRICE_LIMIT = 5.0   # made-up upper tolerance: invoice price more than 5% above the PO price
QTY_LIMIT = 3.0     # made-up upper tolerance: invoiced quantity more than 3% above goods receipt


def make_invoices(rows: int, seed: int = 7):
    """Made-up invoice items. Two inputs, one yes/no answer: was the invoice blocked for payment?"""
    rng = np.random.default_rng(seed)
    price_var = rng.normal(0.0, 4.0, rows).clip(-15, 15)   # % invoice price above (+) or below (-) PO price
    qty_var = rng.normal(0.0, 3.0, rows).clip(-15, 15)     # % invoiced quantity above (+) or below (-) goods receipt
    blocked = (price_var > PRICE_LIMIT) | (qty_var > QTY_LIMIT)
    flip = rng.random(rows) < 0.02                           # 2% noise: manual blocks, data errors
    blocked = blocked ^ flip
    x = np.column_stack([price_var, qty_var]).round(2)
    return x, blocked.astype(float)


def init_params(n_in: int, hidden: int, seed: int, zeros: bool):
    """Starting weights. Small random numbers, unless --zeros is given."""
    rng = np.random.default_rng(seed)
    sizes = [n_in, hidden, 1] if hidden > 0 else [n_in, 1]
    params = []
    for fan_in, fan_out in zip(sizes[:-1], sizes[1:]):
        if zeros:
            w = np.zeros((fan_in, fan_out))
        else:
            w = rng.normal(0.0, np.sqrt(2.0 / fan_in), (fan_in, fan_out))  # "He" scaling, common with ReLU
        params.append([w, np.zeros(fan_out)])
    return params


def sigmoid(z):
    return 1.0 / (1.0 + np.exp(-np.clip(z, -500, 500)))


def forward(params, x):
    """Run inputs through the layers. Returns the output probability and what each layer computed."""
    cache = []
    a = x
    for i, (w, b) in enumerate(params):
        z = a @ w + b                     # weighted sum plus bias
        cache.append((a, z))
        last = i == len(params) - 1
        a = sigmoid(z) if last else np.maximum(0.0, z)  # ReLU in hidden layers, sigmoid at the end
    return a[:, 0], cache


def bce(p, y):
    """Binary cross-entropy: the usual loss for yes/no answers. Lower is better."""
    p = np.clip(p, 1e-12, 1 - 1e-12)
    return float(-np.mean(y * np.log(p) + (1 - y) * np.log(1 - p)))


def backward(params, cache, p, y):
    """Backpropagation: the chain rule, applied from the output layer back to the first layer."""
    grads = [None] * len(params)
    dz = ((p - y) / len(y))[:, None]      # slope of the loss at the output's weighted sum
    for i in reversed(range(len(params))):
        a_in, _ = cache[i]
        w, _ = params[i]
        grads[i] = [a_in.T @ dz, dz.sum(axis=0)]   # slopes for this layer's weights and biases
        if i > 0:
            _, z_prev = cache[i - 1]
            dz = (dz @ w.T) * (z_prev > 0)        # pass the blame back through ReLU
    return grads


def gradient_check(params, x, y, h=1e-5):
    """Compare backprop with a brute-force slope for a few weights. Returns the largest gap."""
    p, cache = forward(params, x)
    grads = backward(params, cache, p, y)
    worst = 0.0
    for layer, (w, _) in enumerate(params):
        for idx in [(0, 0), (w.shape[0] - 1, w.shape[1] - 1)]:
            old = w[idx]
            w[idx] = old + h
            up = bce(forward(params, x)[0], y)
            w[idx] = old - h
            down = bce(forward(params, x)[0], y)
            w[idx] = old
            numeric = (up - down) / (2 * h)
            worst = max(worst, abs(numeric - grads[layer][0][idx]))
    return worst


def scores(p, y, threshold=0.5):
    flagged = p >= threshold
    truth = y == 1
    tp, fp, fn = int((flagged & truth).sum()), int((flagged & ~truth).sum()), int((~flagged & truth).sum())
    acc = float((flagged == truth).mean())
    recall = tp / (tp + fn) if tp + fn else 0.0
    precision = tp / (tp + fp) if tp + fp else 0.0
    return acc, recall, precision


def train_numpy(params, x, y, lr, epochs):
    history = []
    for epoch in range(epochs + 1):
        p, cache = forward(params, x)
        loss = bce(p, y)
        history.append(loss)
        if epoch == epochs:
            break
        grads = backward(params, cache, p, y)
        for (w, b), (gw, gb) in zip(params, grads):
            w -= lr * gw                  # the gradient descent step from Unit 2
            b -= lr * gb
    return history


def train_torch(params, x, y, lr, epochs):
    """The same network, same starting weights, same steps, but PyTorch works out the gradients."""
    try:
        import torch
    except ImportError:
        raise SystemExit("\nPyTorch is not installed. See Set up for Unit 3, Steps 2 and 3.")

    layers = []
    for i, (w, b) in enumerate(params):
        linear = torch.nn.Linear(w.shape[0], w.shape[1]).double()
        with torch.no_grad():
            linear.weight.copy_(torch.from_numpy(w.T.copy()))
            linear.bias.copy_(torch.from_numpy(b.copy()))
        layers.append(linear)
        if i < len(params) - 1:
            layers.append(torch.nn.ReLU())
    model = torch.nn.Sequential(*layers)
    optimizer = torch.optim.SGD(model.parameters(), lr=lr)
    loss_fn = torch.nn.BCEWithLogitsLoss()   # sigmoid and cross-entropy in one, numerically stable
    xt, yt = torch.from_numpy(x), torch.from_numpy(y)
    history = []
    for epoch in range(epochs + 1):
        optimizer.zero_grad()
        loss = loss_fn(model(xt)[:, 0], yt)
        history.append(loss.item())
        if epoch == epochs:
            break
        loss.backward()
        optimizer.step()

    def predict(new_x):
        with torch.no_grad():             # predicting only: no gradients needed
            return torch.sigmoid(model(torch.from_numpy(new_x))[:, 0]).numpy()

    return history, predict, model


def main() -> None:
    parser = argparse.ArgumentParser(description="Train a small neural network on made-up invoice data.")
    parser.add_argument("--hidden", type=int, default=8, help="hidden units; 0 means no hidden layer")
    parser.add_argument("--lr", type=float, default=0.5, help="learning rate")
    parser.add_argument("--epochs", type=int, default=2000, help="training steps over all rows")
    parser.add_argument("--rows", type=int, default=2000, help="how many made-up invoice items")
    parser.add_argument("--seed", type=int, default=1, help="seed for the starting weights")
    parser.add_argument("--zeros", action="store_true", help="start every weight at zero")
    parser.add_argument("--torch", action="store_true", help="also train the same network with PyTorch")
    parser.add_argument("--plot", action="store_true", help="save decision_map.png")
    parser.add_argument("--save", action="store_true", help="save the trained network to invoice_net.json")
    args = parser.parse_args()

    x_raw, y = make_invoices(args.rows)
    cut = int(len(y) * 0.75)
    x_train_raw, x_test_raw, y_train, y_test = x_raw[:cut], x_raw[cut:], y[:cut], y[cut:]
    print(f"{args.rows} made-up invoice items, {y.mean():.0%} blocked. "
          f"Training on {cut}, testing on {len(y) - cut}.")

    # Scale with the training rows only, then apply the same numbers everywhere.
    mean, std = x_train_raw.mean(axis=0), x_train_raw.std(axis=0)
    x_train, x_test = (x_train_raw - mean) / std, (x_test_raw - mean) / std

    base_acc = float((y_test == 0).mean())
    print(f"Baseline (never block): test accuracy {base_acc:.1%}, catches 0% of blocked invoices\n")

    params = init_params(2, args.hidden, args.seed, args.zeros)
    n_weights = sum(w.size + b.size for w, b in params)
    shape = f"2 inputs -> {args.hidden} hidden (ReLU) -> 1 output" if args.hidden else "2 inputs -> 1 output"
    print(f"Network: {shape}, {n_weights} weights and biases, "
          f"{'all starting at zero' if args.zeros else 'random start'}")
    print(f"Gradient check: largest gap between backprop and brute force = "
          f"{gradient_check(params, x_train, y_train):.1e}\n")

    start = [[w.copy(), b.copy()] for w, b in params]
    history = train_numpy(params, x_train, y_train, args.lr, args.epochs)
    print(f"Learning rate {args.lr}, {args.epochs} epochs, full batch")
    print("  epoch    loss")
    for epoch in sorted({0, 1, 10, 100, 500, 1000, args.epochs} & set(range(len(history)))):
        print(f"  {epoch:>5}  {history[epoch]:.4f}")
    if not np.isfinite(history[-1]) or history[-1] > history[0]:
        print("  Warning: the loss ended higher than it started. Lower --lr.")

    p_train, _ = forward(params, x_train)
    p_test, _ = forward(params, x_test)
    tr, te = scores(p_train, y_train), scores(p_test, y_test)
    print(f"\nTrain accuracy {tr[0]:.1%}   Test accuracy {te[0]:.1%}")
    print(f"Test: catches {te[1]:.0%} of blocked invoices; {te[2]:.0%} of its flags are right")

    if args.hidden and not args.zeros:
        first = params[0][0]
        dead = int((np.maximum(0, x_train @ first + params[0][1]).max(axis=0) == 0).sum())
        print(f"Hidden units that never switch on: {dead} of {args.hidden}")
    if args.zeros and args.hidden:
        spread = float(np.ptp(params[0][0]))
        print(f"Spread between hidden-unit weights after training: {spread:.4f} "
              "(0 means every unit is identical, or none learned)")

    examples = np.array([[2.0, 1.0], [7.0, 0.0], [0.0, 5.0], [-8.0, 0.0], [4.5, 2.5]])
    p_ex, _ = forward(params, (examples - mean) / std)
    print("\nWhat the network says about new invoice items:")
    print("  price var  qty var  P(blocked)")
    for (pv, qv), prob in zip(examples, p_ex):
        print(f"  {pv:+8.1f}%  {qv:+6.1f}%  {prob:9.2f}")

    if args.torch:
        t_hist, t_predict, _ = train_torch(start, x_train, y_train, args.lr, args.epochs)
        t_acc = scores(t_predict(x_test), y_test)[0]
        print(f"\nPyTorch, same start and steps: final loss {t_hist[-1]:.6f} "
              f"(numpy: {history[-1]:.6f}), test accuracy {t_acc:.1%}")

    if args.save:
        out = HERE / "invoice_net.json"
        out.write_text(json.dumps({
            "inputs": ["price_variance_pct", "quantity_variance_pct"],
            "scaler": {"mean": mean.tolist(), "std": std.tolist()},
            "layers": [{"weights": w.tolist(), "biases": b.tolist()} for w, b in params],
            "hidden_activation": "relu", "output_activation": "sigmoid",
            "test_accuracy": round(te[0], 4),
        }, indent=2))
        print(f"\nSaved {out.name} (weights plus the scaler: both are needed to predict)")

    if args.plot:
        import matplotlib
        matplotlib.use("Agg")  # draw into a file, no window needed
        import matplotlib.pyplot as plt

        gx, gy = np.meshgrid(np.linspace(-12, 12, 200), np.linspace(-10, 10, 200))
        grid = np.column_stack([gx.ravel(), gy.ravel()])
        prob, _ = forward(params, (grid - mean) / std)
        plt.figure(figsize=(6, 4.5))
        plt.contourf(gx, gy, prob.reshape(gx.shape), levels=20, cmap="RdYlGn_r", alpha=0.7)
        plt.colorbar(label="P(blocked)")
        plt.scatter(x_test_raw[:, 0], x_test_raw[:, 1], c=y_test, cmap="coolwarm", s=8, edgecolors="none")
        plt.axvline(PRICE_LIMIT, color="black", lw=0.8, ls="--")
        plt.axhline(QTY_LIMIT, color="black", lw=0.8, ls="--")
        plt.xlabel("price variance, % above PO price")
        plt.ylabel("quantity variance, % above goods receipt")
        plt.title(f"{args.hidden} hidden units: test accuracy {te[0]:.1%}")
        plt.tight_layout()
        plt.savefig(HERE / "decision_map.png", dpi=120)
        print("\nSaved decision_map.png")


if __name__ == "__main__":
    main()

Step 3: Run it with the default settings

In the terminal, still inside unit03, run:

python neural_net.py

It takes about a second. What success looks like:

2000 made-up invoice items, 26% blocked. Training on 1500, testing on 500.
Baseline (never block): test accuracy 74.0%, catches 0% of blocked invoices

Network: 2 inputs -> 8 hidden (ReLU) -> 1 output, 33 weights and biases, random start
Gradient check: largest gap between backprop and brute force = 6.6e-12

Learning rate 0.5, 2000 epochs, full batch
  epoch    loss
      0  0.7561
      1  0.6742
     10  0.4853
    100  0.2203
    500  0.1747
   1000  0.1679
   2000  0.1654

Train accuracy 96.6%   Test accuracy 96.6%
Test: catches 90% of blocked invoices; 97% of its flags are right
Hidden units that never switch on: 0 of 8

What the network says about new invoice items:
  price var  qty var  P(blocked)
      +2.0%    +1.0%       0.03
      +7.0%    +0.0%       0.99
      +0.0%    +5.0%       1.00
      -8.0%    +0.0%       0.03
      +4.5%    +2.5%       0.31

Your numbers should match, because the data and starting weights come from fixed seeds. The gradient check gap may show a slightly different tiny number; anything below about 1e-6 is fine.

Step 4: Read the results

  1. Baseline, 74.0%. Never blocking anything is right 74% of the time, because 74% of test invoices weren't blocked. That is why accuracy alone misleads, as the classification topic showed. Recall ("catches") and precision ("flags are right") tell the real story.
  2. Gradient check. Backpropagation and the brute-force slope agree to about 12 decimal places. Your chain-rule code is right.
  3. The loss. It starts near 0.69, the "don't know" value, and falls to 0.17. Most of the progress happens in the first 100 epochs.
  4. Train versus test. Both are 96.6%. The network isn't memorizing; it learned something that carries over to new invoices. With 2% of labels flipped at random, about 98% is the best any model could do here.
  5. New invoices. Price +7% and quantity +5% are clearly blocked; price -8% is not, because under-billing only triggers a message. Price +4.5% with quantity +2.5% is near both limits, and the network is unsure (0.31). That is honest: it is close to the corner.

Step 5: Compare with a straight line

python neural_net.py --hidden 0

With no hidden layer, the network is exactly logistic regression: 3 parameters. What you will see (the lines that matter):

Network: 2 inputs -> 1 output, 3 weights and biases, random start
...
Train accuracy 85.9%   Test accuracy 85.8%
Test: catches 62% of blocked invoices; 79% of its flags are right
...
      +7.0%    +0.0%       0.58
      +4.5%    +2.5%       0.75

The straight line misses 38% of blocked invoices. It thinks price +7% alone is a coin flip, and it is more worried about the in-tolerance invoice at +4.5% and +2.5% than about the one clearly over the price limit. A line can't draw the L.

Step 6: Try a tiny network, and a bad start

python neural_net.py --hidden 2
python neural_net.py --hidden 2 --seed 2

What you will see:

Network: 2 inputs -> 2 hidden (ReLU) -> 1 output, 9 weights and biases, random start
...
Train accuracy 96.7%   Test accuracy 96.6%

Network: 2 inputs -> 2 hidden (ReLU) -> 1 output, 9 weights and biases, random start
...
Train accuracy 86.7%   Test accuracy 85.6%
Test: catches 61% of blocked invoices; 79% of its flags are right

Two hidden units are enough, just like the hand-designed network in "Why a straight line fails here". But with --seed 2, the same network starts from different random weights and gets stuck at straight-line quality. The loss stops falling at 0.3334 and never recovers. This is the non-convex loss from scikit-learn's warning: more than one valley, and where you land depends on where you start. With 8 units there are more ways to find a good valley, which is one reason networks are usually built larger than the minimum.

Step 7: Break the training on purpose

python neural_net.py --zeros
python neural_net.py --lr 20

What you will see:

Network: 2 inputs -> 8 hidden (ReLU) -> 1 output, 33 weights and biases, all starting at zero
...
      0  0.6931
   2000  0.5660
Train accuracy 74.7%   Test accuracy 74.0%
Test: catches 0% of blocked invoices; 0% of its flags are right
Spread between hidden-unit weights after training: 0.0000 (0 means every unit is identical, or none learned)
...
Learning rate 20.0, 2000 epochs, full batch
...
      1  1.9933
   2000  1.7034
  Warning: the loss ended higher than it started. Lower --lr.
Train accuracy 74.7%   Test accuracy 74.0%
Hidden units that never switch on: 6 of 8
  • All zeros. Every hidden unit outputs 0, and ReLU passes no blame back through a zero. The hidden weights never move. Only the output bias learns, so the network says 0.25 for every invoice: the share of blocked invoices. It is no better than the baseline.
  • Learning rate 20. The loss jumps up on the first step and never settles. Worse, 6 of the 8 hidden units have been pushed so far negative that they output 0 for every training invoice. These are the dead ReLU units from Google's list, and they can't come back, because no blame reaches them.

Try --lr 0.05 as well: it trains, but slower, and ends at a loss of 0.1931 instead of 0.1654 after the same 2,000 epochs.

Step 8: Train the same network with PyTorch

python neural_net.py --torch

The last line is:

PyTorch, same start and steps: final loss 0.165382 (numpy: 0.165382), test accuracy 96.6%

The script copies your starting weights into a PyTorch model and runs the same 2,000 steps. PyTorch's autograd works out the gradients instead of your backward function, and the final loss matches to six decimal places. Your backpropagation is exactly what the library does. From here on, the course lets PyTorch do it.

Note two PyTorch details in train_torch. The loss is BCEWithLogitsLoss, which takes the output before the sigmoid; PyTorch's documentation says combining the two is more numerically stable. And optimizer.zero_grad() runs every step, because the autograd tutorial explains that PyTorch adds new gradients to old ones unless you clear them.

Step 9: Draw the decision map and save the network

python neural_net.py --plot --save

The last lines are:

Saved invoice_net.json (weights plus the scaler: both are needed to predict)

Saved decision_map.png

Open decision_map.png from the VS Code file list. Green means "not blocked" and red means "blocked"; the dashed lines are the made-up limits. You should see an L-shaped red region hugging both lines, with rounded edges near the corner where the network is unsure. Now run python neural_net.py --hidden 0 --plot and open the picture again: a single straight edge, cutting through the L.

invoice_net.json holds the 33 learned numbers plus the scaling averages and spreads. The exercise uses it.

Step 10: Save your work in Git

From the course folder:

cd ..
git add unit03/neural_net.py
git commit -m "Build a neural network from scratch for blocked invoices"

decision_map.png and invoice_net.json are outputs you can recreate at any time, so you don't need to commit them.

What each part of the script does

Part What it does
make_invoices Makes 2,000 invoices with fixed seeds: blocked if price is over 5% or quantity over 3%, with 2% of labels flipped
init_params Creates the weight matrices: small random numbers scaled for ReLU, or zeros with --zeros
forward Weighted sums through each layer, ReLU in the hidden layer, sigmoid at the output
bce Binary cross-entropy loss
backward Backpropagation: blame at the output, passed back through the weights and ReLU
gradient_check Nudges a few weights by a tiny h and compares the slope with backward
scores Accuracy, recall ("catches") and precision ("flags are right") at a 0.5 threshold
train_numpy Full-batch gradient descent: forward, loss, backward, step
train_torch The same network in PyTorch, starting from the same weights, with autograd and SGD
Scaling lines Standardize both inputs with averages and spreads from the training rows only
Dead-unit line Counts hidden units whose output is 0 for every training invoice
--save block Writes weights, biases and the scaler to invoice_net.json
--plot block Colours a grid of invoices by predicted probability and saves decision_map.png

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of the Unit 1 setup, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'numpy' (or matplotlib) The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
PyTorch is not installed with --torch PyTorch isn't in this .venv Follow Steps 2 and 3 of Set up for Unit 3. Everything except --torch works without it
pip shows ProxyError, SSLError or Could not fetch URL while installing Your network or company proxy blocks the package sites Try another network, or ask IT for access or an internal mirror. The script itself needs no network
can't open file ... neural_net.py The terminal isn't in unit03, or the file has another name Run cd unit03 from the course folder, and check the file name
RuntimeWarning: overflow A very large learning rate pushed the numbers too far Expected with --lr 20. Use a smaller --lr
Numbers differ a lot from the page You changed the script or passed other options Run exactly python neural_net.py with the script as published. Tiny differences in the last decimal are normal across library versions
Asked for an API key Nothing in this topic uses a key or account Check you are running neural_net.py, not another script

Where this shows up in SAP

Unit 3 is a mechanics unit, so this section is short. It maps what you built onto SAP's three routes, as described by SAP's Architecture Center page on classic machine learning (last updated April 23, 2026).

Inside SAP HANA: PAL's multi-layer perceptron

The hana-ml Python client for SAP HANA (release 2.30, September 2026) has a neural_network module under hana_ml.algorithms.pal with MLPClassifier, MLPRegressor, MLPMultiTaskClassifier and MLPMultiTaskRegressor. Training runs in the database. The parameters map directly onto this topic:

This topic MLPClassifier parameter in hana-ml 2.30
--hidden 8 hidden_layer_size, a tuple of sizes, one per hidden layer
ReLU in the hidden layer activation; options include 'relu', 'tanh' and several sigmoid variants
Sigmoid at the output output_activation
--lr, full batch learning_rate, momentum, and training_style ('batch' or 'stochastic')
Scaling with training averages normalization, with options 'no', 'z-transform' and 'scalar'
Random or zero start weight_init, with options such as 'all-zeros', 'normal' and 'uniform'
--epochs max_iter

The client's own example sets weight_init='normal' and normalization='z-transform'. After this topic you know why both matter. The multi-task classes add options such as batch normalization, a batch size and several optimizers.

SAP's Architecture Center recommends PAL and APL mainly for time series, anomaly detection, clustering and other in-database machine learning. For classification and regression on tables, it recommends starting with SAP-RPT-1.

Pretrained: SAP-RPT-1

As of September 2026, SAP describes SAP-RPT-1 as a relational pretrained transformer for structured business data, offered as a small and a large version in the generative AI hub, plus an open-source release. You send example rows with the request instead of training. The Architecture Center adds that it can be called from SAP HANA Cloud through a SQL stored procedure. Transformers are the neural network architecture of Unit 4.

Your own network: SAP AI Core

The Architecture Center's guidance is to choose SAP AI Core when deep learning or large-scale neural networks are required, or custom models in TensorFlow, PyTorch or similar frameworks. SAP Learning describes how training runs there: your code in a container, orchestrated by Argo workflows; configurations holding parameters that can change for every run; resource plans with different CPU, GPU and memory; and APIs to register metrics.

This topic SAP AI Core equivalent
neural_net.py with --torch Your PyTorch training code, packaged in a container image
--hidden, --lr, --epochs, --seed Parameters in a configuration, changed per run
The printed loss and test scores Metrics registered through AI Core's metrics APIs
invoice_net.json The trained model, stored for serving
Your laptop's CPU A resource plan, with GPUs for larger networks

Build, library or service

Situation Use Why
Learning how networks train, debugging a strange run Your own numpy network Every number is visible
A custom network or a new architecture PyTorch, locally or in SAP AI Core Autograd, GPU support, the whole ecosystem
A small network on data already in SAP HANA PAL MLPClassifier through hana-ml Trains next to the data, no data movement
Classification or regression on business tables SAP-RPT-1 first, per SAP's guidance Pretrained, no training pipeline
A rule that is already configured, such as tolerance limits No model Read the configuration

Production concerns

  • Beat a simple model first. Always report a straight-line model on the same held-back data. If the network's gain is small, the simpler model is cheaper to run and easier to explain.
  • Save the scaler with the weights. invoice_net.json stores both. A network served without its scaling numbers returns confident nonsense, because it expects inputs on the training scale.
  • Record every run. Seed, hidden size, learning rate, epochs, data version, loss curve, test scores. Step 6 showed two runs of the same network ending 11 points apart. In AI Core, that means configuration parameters plus registered metrics.
  • Choose on held-back data, not the test set. Picking the hidden size or the seed by test score leaks the test into training. Use a validation split or cross-validation, as scikit-learn's guide suggests for tuning.
  • Watch for dead units and blow-ups. Stop a training job and alert when the loss grows or becomes not-a-number. Counting units that never switch on is a cheap health check.
  • Explainability. A network's 33 weights don't read like a rule. For decisions such as payment blocks, auditors need a reason per invoice. Plan how you will explain predictions before you build, or keep the decision in configured rules and use the model only to prioritize work.
  • Authorizations and data. Real invoice data carries supplier names, prices and bank details. The access rules from Calling your first SAP API apply wherever training runs; Unit 11 goes deeper.
  • Clean core. Train and serve side by side (AI Core) or in the database (PAL). Don't change standard invoice verification logic to host a model.

Pitfalls

  • Training a model to learn a configured rule. It is slower, less exact and harder to audit than reading the configuration.
  • Trusting one run. A bad starting point can leave a network stuck at straight-line quality. Try several seeds and record them.
  • Starting every weight at the same value. Identical units learn identical things; with zeros and ReLU, nothing learns at all.
  • A learning rate copied from another project. Too large kills ReLU units for good. Watch the first epochs of the loss curve.
  • Forgetting to scale, or scaling with the wrong numbers. scikit-learn's guide calls networks sensitive to feature scaling. Fit the scaler on training rows and reuse it everywhere.
  • Reading accuracy alone. With 74% of invoices not blocked, "74% accurate" is what doing nothing scores.
  • Forgetting zero_grad() in PyTorch. Gradients add up across steps, and training quietly goes wrong.

Exercise: load your network and score invoices without training

You will use the saved network to score new invoices in a separate script, as a service would, and record which settings train well. The prediction script is the seed of the AI API you build in Unit 6.

  1. In the terminal, inside unit03 with (.venv) showing, run each command below and write down the test accuracy and the "catches" percentage:

    python neural_net.py --hidden 0
    python neural_net.py --hidden 2
    python neural_net.py --hidden 8
    python neural_net.py --hidden 32
    python neural_net.py --hidden 2 --seed 2
    python neural_net.py --hidden 2 --seed 3
  2. Run python neural_net.py --save to write invoice_net.json with the default 8-unit network.

  3. In VS Code, create unit03/predict_invoice.py, paste the code below and save:

    """Score one invoice item with the network saved by neural_net.py --save.
    
    How to run (from the unit03 folder):
        python predict_invoice.py --price 7 --qty 0
    
    Uses only numpy and the saved file. No training happens here.
    """
    import argparse
    import json
    from pathlib import Path
    
    import numpy as np
    
    HERE = Path(__file__).parent
    
    
    def main() -> None:
        parser = argparse.ArgumentParser(description="Score one invoice item with invoice_net.json.")
        parser.add_argument("--price", type=float, required=True, help="% invoice price above the PO price")
        parser.add_argument("--qty", type=float, required=True, help="% invoiced quantity above goods receipt")
        args = parser.parse_args()
    
        path = HERE / "invoice_net.json"
        if not path.exists():
            raise SystemExit("invoice_net.json not found. Run: python neural_net.py --save")
        net = json.loads(path.read_text())
    
        x = np.array([args.price, args.qty])
        a = (x - np.array(net["scaler"]["mean"])) / np.array(net["scaler"]["std"])   # same scaling as training
        layers = net["layers"]
        for i, layer in enumerate(layers):
            z = a @ np.array(layer["weights"]) + np.array(layer["biases"])            # weighted sum plus bias
            a = 1 / (1 + np.exp(-z)) if i == len(layers) - 1 else np.maximum(0.0, z)  # sigmoid last, ReLU before
        print(f"Price {args.price:+.1f}%, quantity {args.qty:+.1f}%: P(blocked) = {a[0]:.2f}")
    
    
    if __name__ == "__main__":
        main()
  4. Run it for two invoices from Step 3's table:

    python predict_invoice.py --price 7 --qty 0
    python predict_invoice.py --price 4.5 --qty 2.5

    You should see P(blocked) = 0.99 and P(blocked) = 0.31, the same as in neural_net.py's output. If they differ, check that you ran Step 2 of this exercise with no other options.

  5. Score three invoices of your own choosing: one clearly fine, one clearly over a limit, and one right at a corner, such as --price 5 --qty 3.

  6. Create unit03/notes_neural_networks.md with three short sections:

    • Hidden units: a table of the six runs from step 1 with test accuracy and catches, and one sentence on the smallest network you would trust.
    • Seeds: one sentence on what --seed 2 and --seed 3 showed, and what you would do about it in a real project.
    • Scoring: your three invoices from step 5 with their probabilities, and one sentence on why predict_invoice.py needs the scaler numbers.
  7. Save your work:

    git add unit03/predict_invoice.py unit03/notes_neural_networks.md
    git commit -m "Score invoices with a saved neural network, plus notes"

Done when: predict_invoice.py prints 0.99 and 0.31 for the two invoices in step 4; your notes list all six runs with test accuracy and catches; the seed and scoring sections are filled in; and git log shows the commit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does a stack of layers without activation functions learn nothing a single layer couldn't?

    Answer: B. Linear operations on linear operations stay linear, as Google's crash course puts it. The ReLU bend in each hidden unit is what lets the network combine hinges into shapes such as the L of the blocked-invoice region.
  2. 2In backward, what does the line dz = (dz @ w.T) * (z_prev > 0) do?

    Answer: D. Blame from the output reaches each hidden unit in proportion to its outgoing weight. ReLU is flat below zero, so units that output 0 get no blame. The weight update itself happens later, in train_numpy.
  3. 3The straight-line model (--hidden 0) catches 62% of blocked invoices; 2 hidden units catch 90%. Why?

    Answer: A. Blocked means price over its limit or quantity over its limit. A single line can't draw that corner. Each hidden unit acts as a hinge at one limit, and the output adds them, as in the hand-designed network.
  4. 4You train with --zeros and every invoice gets P(blocked) = 0.25. What happened?

    Answer: C. With all weights at zero, every hidden unit's output is 0 and ReLU passes no gradient through it. The hidden weights never change. The output bias alone learns the share of blocked invoices, 0.25, so the network matches the do-nothing baseline.
  5. 5After a run with --lr 20, 6 of 8 hidden units never switch on. What are they, and what fixes it?

    Answer: D. Huge steps pushed their weighted sums below zero for every invoice, so they always output 0 and get no gradient. Google's crash course names lowering the learning rate, or a ReLU variant such as LeakyReLU, as fixes.
  6. 6The same 2-unit network scores 96.6% with --seed 1 and 85.6% with --seed 2. What would you do in a real project?

    Answer: B. A network's loss has more than one valley, and the start decides where it lands. Recording runs and choosing on validation data keeps the test set honest. Picking by test score leaks the test into model selection.
  7. 7Why does train_torch use BCEWithLogitsLoss on the raw output instead of applying a sigmoid first?

    Answer: C. PyTorch's documentation says combining the sigmoid and binary cross-entropy in one class is more numerically stable than applying them separately. The prediction function still applies a sigmoid, inside torch.no_grad().
  8. 8Your team wants a neural network on invoice data already in SAP HANA, trained next to the data. Which SAP option fits?

    Answer: D. hana-ml's neural_network module includes MLPClassifier, which runs PAL's multi-layer perceptron in the database. SAP-RPT-1 is pretrained and needs no training, and AI Core is for custom deep learning code.
  9. 9A model served in production returns near-certain "blocked" for most invoices, even small variances, though test results were good. What do you check first?

    Answer: A. The network was trained on standardized inputs. If the service sends raw percentages, weighted sums land far from the training range and outputs saturate. That is why invoice_net.json stores the scaler with the weights.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in