What a neural network is, why hidden layers let it learn patterns a straight line can't, and how to build, train and check one yourself on SAP-style invoice data.
A neural network is a prediction model built from many small, simple calculators called units (or neurons). Each unit does one thing: it multiplies its inputs by some numbers, adds them up, and passes on the result, but only after a small twist.
Units are arranged in layers. The first layer reads the inputs, such as the price difference on a supplier invoice. The last layer gives the answer, such as "likely to be blocked". The layers in between are hidden layers. They build their own intermediate signals, such as "price is well above the order" or "quantity is above what we received".
The twist in each unit is what matters. Without it, stacking layers gives you nothing more than a single straight-line rule. With it, the network can learn bends, corners and combinations that a straight line can't draw.
Training works exactly like the gradient descent you did by hand: measure the error, work out which way to nudge every number, take a small step, repeat. The only new piece is a bookkeeping method called backpropagation, which works out those nudges for every layer at once.
Large language models, image readers and the embedding models in the next topic are all neural networks. They are much bigger, but the building block is the one in this topic.
Most of the AI your company will buy runs on neural networks. Knowing what they are good and bad at helps you judge proposals.
Take the procure-to-pay running example: supplier invoices blocked in the three-way match. In SAP S/4HANA, invoice verification compares the invoice with the purchase order and the goods receipt. SAP Learning explains that when a variance exceeds the upper tolerance limit, the invoice is blocked for payment. Price variances and quantity variances each have their own tolerance keys, such as PP and DQ.
That gives a pattern with a corner in it: an invoice is blocked if the price is too high or the quantity is too high. A single straight-line rule can't draw that corner well. In the deep layer, a straight-line model catches only 62% of the blocked invoices in made-up test data. A small neural network with 8 hidden units catches 90%.
That example teaches a second lesson, and it is the more important one for a leader. You would never train a network to rediscover your own tolerance rules. They are configured in the system; read the configuration. The example uses a known rule only so you can check what the network learned. Neural networks earn their place when nobody can write the rule down: text, images, scanned documents, and patterns across many columns.
Where neural networks help
Where they usually don't
Reading text, such as notes, emails and item descriptions
Rules that are already configured, such as tolerance limits
Scanned documents and images
Small tables where a simpler model scores just as well
Patterns across many columns that interact
Decisions that must be explained line by line to an auditor
Tasks where a pretrained model already exists
Use cases with no labelled history and no pretrained model
As of September 2026, SAP offers neural networks in three places. SAP's Architecture Center page on classic machine learning (last updated April 2026) lays out when to use each:
Pretrained, for business tables: SAP-RPT-1. SAP describes it as a relational pretrained transformer, a kind of neural network. You send example rows with each request instead of training. SAP's guidance is to start here for classification and regression on tables.
Inside SAP HANA: the Predictive Analysis Library (PAL). The hana-ml Python client for SAP HANA includes multi-layer perceptron classes, the textbook neural network this topic builds. Training runs in the database, next to the data.
Your own network: SAP AI Core. SAP's guidance is to choose AI Core when deep learning or large-scale neural networks are needed, or custom models in frameworks such as PyTorch. Your team supplies the training code, and AI Core runs it in containers, with GPU options.
The language models in SAP's generative AI hub are neural networks too, trained by their providers. Unit 4 opens them up.
For a leader, the practical question is the same as in Unit 2: who trains the network, and who keeps it working?
"A neural network works like a brain." It borrows the name. Each unit is a weighted sum and a simple function, nothing more.
"More layers always means better." In the deep layer, 2 hidden units do as well as 32 on held-back data. Bigger networks cost more and can memorize noise.
"Neural networks don't need clean data." They are sensitive to how inputs are scaled and to label errors, like any model.
"A network will find our business rules for us." If the rule is configured, read it. Learning it back from data is slower and less exact.
"Training is repeatable." Networks start from random numbers. Two runs can end in different places; the deep layer shows one that gets stuck.
Pick one answer for each question. The explanation appears after you choose.
1What makes a neural network able to learn patterns a straight-line rule can't?
Answer: B. Without an activation function, stacked layers collapse into one straight-line rule. The small twist in each unit is what lets hidden layers combine into shapes such as the corner in the invoice-blocking pattern.
2How is a neural network trained?
Answer: C. Training is the same loop as in Unit 2: measure the error, find the direction that reduces it, take a small step. Backpropagation is the bookkeeping that finds that direction for every layer at once.
3A team proposes a neural network to predict which supplier invoices S/4HANA will block for price or quantity variances. What is your first question?
Answer: D. SAP blocks an invoice for payment when a variance exceeds the configured upper tolerance limit. A known, configured rule should be read, not learned back from data. Neural networks earn their place where nobody can write the rule down.
4Your use case is a yes/no prediction on a table of business data. Per SAP's own guidance, where do you start?
Answer: A. SAP's Architecture Center advises starting with RPT-1 for classification and regression on tables, because it needs no training. AI Core is for cases that really need deep learning or custom frameworks.
5When is SAP AI Core the right home for a neural network?
Answer: C. SAP's guidance is to choose AI Core for deep learning, large-scale neural networks, or custom models in frameworks such as PyTorch. The team then owns the training code and its upkeep.
6A vendor says their network with 32 hidden units is "clearly better" than one with 2. What do you ask for?
Answer: D. A bigger network can fit its training data better and still do no better, or worse, on new cases. In the deep layer, 2 hidden units match 32 on held-back data. Only held-back results and a simple baseline settle it.
7Two training runs of the same network give quite different results. What does that tell you?
Answer: B. Training starts from random weights, and a network's error landscape has more than one valley. Teams should fix seeds, record runs, and choose on held-back data. The deep layer shows a run that gets stuck.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: logistic regression, stacked, with a bend in between
You already know the output end of a neural network. In Classification and the metrics that matter a model computed a weighted sum and squeezed it through a sigmoid into a probability. That is one unit.
A neural network puts a layer of such units in front of it. Each hidden unit computes its own weighted sum and passes it through an activation function, usually ReLU: keep the number if it is positive, otherwise output zero. The output unit then combines the hidden units' signals.
hidden_j = ReLU(w_j1 * price_var + w_j2 * qty_var + b_j) for each hidden unit j
P(blocked) = sigmoid(v_1 * hidden_1 + ... + v_H * hidden_H + c)
The bend in ReLU is the whole trick. Google's crash course puts it plainly: linear operations performed on linear operations are still linear. Take ReLU out and any stack of layers collapses into one weighted sum. Keep it, and each hidden unit contributes one "hinge" that the output can combine with others into corners and curves.
Training is unchanged from Gradient descent, by hand: weight = weight - learning_rate * gradient. Backpropagation is how you get the gradient for weights that sit behind other layers.
#The data: invoices blocked in the three-way match
SAP Learning describes how S/4HANA invoice verification handles variances. Tolerance limits are set in Customizing, per tolerance key: for example PP for price variances and DQ for quantity variances. If a variance exceeds the upper limit, the invoice is blocked for payment. If it falls below the lower limit, the system only issues a message. The block applies to the whole invoice, even if only one item varies.
This topic uses made-up invoices with one item each and two inputs:
Input
Meaning
Range in the data
price_var
% the invoice price is above (+) or below (-) the purchase order price
about -15 to +15
qty_var
% the invoiced quantity is above (+) or below (-) the goods receipt quantity
about -15 to +15
The label is "blocked": price_var > 5orqty_var > 3, with 2% of labels flipped at random to stand in for manual blocks and data errors. The 5% and 3% limits are made up; real limits depend on your configuration. Negative variances never block, which matches the "below the lower limit, only a message" behaviour.
Draw the invoices on a chart with price variance across and quantity variance up. The blocked region is an L shape: everything right of the price limit plus everything above the quantity limit. Logistic regression draws one straight line across that chart. Whatever angle it picks, it cuts through the L and gets one of the two arms wrong. In this topic's script it catches only 62% of blocked invoices.
Two ReLU units are enough to fix it. Here is a network designed by hand, to show the idea:
h1 = ReLU(price_var - 5) zero until price is 5% over, then grows
h2 = ReLU(qty_var - 3) zero until quantity is 3% over, then grows
P(blocked) = sigmoid(-2 + 4*h1 + 4*h2)
Invoice
h1
h2
Weighted sum
P(blocked)
price +2%, qty +1%
0
0
-2
0.12
price +7%, qty 0%
2
0
6
1.00
price 0%, qty +5%
0
2
6
1.00
price -8%, qty 0%
0
0
-2
0.12
Each hidden unit has learned one "arm" of the L. The output adds them. Training finds weights like these on its own; you will see it do so with 2 hidden units.
Running inputs through the layers is called the forward pass. For a batch of rows it is two matrix multiplications:
flowchart LR
X[Inputs<br/>price_var, qty_var] --> H[Hidden layer<br/>8 units: weighted sum, ReLU]
H --> O[Output unit<br/>weighted sum, sigmoid]
O --> P[P blocked]
P --> L[Loss<br/>cross-entropy]
Counting parameters: each of the 8 hidden units has 2 weights and 1 bias (24), and the output has 8 weights and 1 bias (9). That is 33 numbers to learn. An LLM has billions, arranged differently, but each one is learned the same way.
For a yes/no answer, the loss is binary cross-entropy: for each row, -log(p) if the invoice was blocked and -log(1 - p) if it wasn't, averaged. A confident wrong answer costs a lot; a confident right answer costs almost nothing. A model that says 0.5 for everything scores 0.693, a number you will see in the script.
The output unit's weights are easy: their gradient is the same one you would compute for logistic regression. The hidden weights are harder, because they affect the loss only through the output. Backpropagation handles this by starting at the loss and passing blame backwards, layer by layer:
At the output: blame = (p - y) / n. With sigmoid and cross-entropy together, the slope at the output's weighted sum simplifies to exactly this.
Output weights: gradient = hidden signals times blame.
Pass blame back: each hidden unit gets its share, blame * its outgoing weight.
Through ReLU: a hidden unit that output zero passes no blame back (ReLU is flat there).
Hidden weights: gradient = inputs times the hidden unit's blame.
sequenceDiagram
participant I as Inputs
participant H as Hidden layer
participant O as Output
participant L as Loss
I->>H: forward: weighted sums, ReLU
H->>O: forward: weighted sum, sigmoid
O->>L: compare with the label
L-->>O: blame = p - y
O-->>H: blame x output weights, zero where ReLU was off
H-->>I: gradients for the hidden weights
Google's crash course calls backpropagation the most common training algorithm for neural networks, and the thing that makes gradient descent feasible for many layers. PyTorch automates it. Its autograd records each operation in a graph as the forward pass runs, then applies the chain rule when you call loss.backward(). The script does it by hand first, then shows PyTorch getting the same numbers.
Google's crash course names three failure cases. You will trigger two of them on purpose.
Problem
What happens
Usual fix
Vanishing gradients
In deep networks, blame shrinks as it passes back, so early layers barely learn
ReLU instead of sigmoid or tanh in hidden layers
Exploding gradients
Blame grows as it passes back, and training never settles
Lower learning rate, batch normalization
Dead ReLU units
A unit's weighted sum stays below zero, it always outputs 0, and no blame reaches it again
Lower learning rate, or a ReLU variant such as LeakyReLU
A fourth problem is specific to how you start. If every weight starts at the same value, every hidden unit computes the same thing and gets the same update, forever. With all zeros and ReLU it is worse: no blame flows back at all. That is why networks start from small random numbers.
For logistic regression the loss is a single bowl. For a neural network it isn't. scikit-learn's user guide lists this as the first drawback of its multi-layer perceptron: the loss is non-convex, has more than one local minimum, and different random starts can give different results. The same guide warns that the models are sensitive to feature scaling and need tuning of hidden units, layers and iterations. The script lets you see all three.
#Build it yourself: a neural network for blocked invoices
You will write a neural network in about 150 lines of numpy, with backpropagation by hand, and train it to predict which made-up invoices get blocked. You will compare it with a straight-line model, break it in three ways, then train the same network with PyTorch and confirm both give identical numbers.
Before you start: complete Set up your computer for this course and Set up for Unit 3. They give you the orchestrate-course folder with its .venv, numpy and matplotlib from Unit 2, PyTorch, and the unit03 folder. This walkthrough doesn't repeat those steps.
flowchart LR
D[2,000 made-up invoices] --> S[Scale with training rows]
S --> N[numpy network<br/>forward, backprop, step]
N --> R[Accuracy, recall, precision]
N --> T[Optional: same network in PyTorch]
N --> F[Optional: invoice_net.json and decision_map.png]
In VS Code's file list, right-click unit03, choose New File and name it neural_net.py.
Paste the code below and save (Ctrl+S, or Cmd+S on Mac).
"""A neural network from scratch: learn which supplier invoices get blocked for payment.
How to run (from the unit03 folder, with the course .venv turned on):
python neural_net.py # 8 hidden units, trained with numpy only
python neural_net.py --hidden 0 # no hidden layer: plain logistic regression
python neural_net.py --hidden 2 # a very small network
python neural_net.py --zeros # start every weight at zero and see what happens
python neural_net.py --lr 20 # steps far too big
python neural_net.py --torch # train the same network with PyTorch and compare
python neural_net.py --plot # also save decision_map.png
python neural_net.py --save # also save the trained network to invoice_net.json
All data is made up. Nothing is sent anywhere.
"""
import argparse
import json
from pathlib import Path
import numpy as np
HERE = Path(__file__).parent
PRICE_LIMIT = 5.0 # made-up upper tolerance: invoice price more than 5% above the PO price
QTY_LIMIT = 3.0 # made-up upper tolerance: invoiced quantity more than 3% above goods receipt
def make_invoices(rows: int, seed: int = 7):
"""Made-up invoice items. Two inputs, one yes/no answer: was the invoice blocked for payment?"""
rng = np.random.default_rng(seed)
price_var = rng.normal(0.0, 4.0, rows).clip(-15, 15) # % invoice price above (+) or below (-) PO price
qty_var = rng.normal(0.0, 3.0, rows).clip(-15, 15) # % invoiced quantity above (+) or below (-) goods receipt
blocked = (price_var > PRICE_LIMIT) | (qty_var > QTY_LIMIT)
flip = rng.random(rows) < 0.02 # 2% noise: manual blocks, data errors
blocked = blocked ^ flip
x = np.column_stack([price_var, qty_var]).round(2)
return x, blocked.astype(float)
def init_params(n_in: int, hidden: int, seed: int, zeros: bool):
"""Starting weights. Small random numbers, unless --zeros is given."""
rng = np.random.default_rng(seed)
sizes = [n_in, hidden, 1] if hidden > 0 else [n_in, 1]
params = []
for fan_in, fan_out in zip(sizes[:-1], sizes[1:]):
if zeros:
w = np.zeros((fan_in, fan_out))
else:
w = rng.normal(0.0, np.sqrt(2.0 / fan_in), (fan_in, fan_out)) # "He" scaling, common with ReLU
params.append([w, np.zeros(fan_out)])
return params
def sigmoid(z):
return 1.0 / (1.0 + np.exp(-np.clip(z, -500, 500)))
def forward(params, x):
"""Run inputs through the layers. Returns the output probability and what each layer computed."""
cache = []
a = x
for i, (w, b) in enumerate(params):
z = a @ w + b # weighted sum plus bias
cache.append((a, z))
last = i == len(params) - 1
a = sigmoid(z) if last else np.maximum(0.0, z) # ReLU in hidden layers, sigmoid at the end
return a[:, 0], cache
def bce(p, y):
"""Binary cross-entropy: the usual loss for yes/no answers. Lower is better."""
p = np.clip(p, 1e-12, 1 - 1e-12)
return float(-np.mean(y * np.log(p) + (1 - y) * np.log(1 - p)))
def backward(params, cache, p, y):
"""Backpropagation: the chain rule, applied from the output layer back to the first layer."""
grads = [None] * len(params)
dz = ((p - y) / len(y))[:, None] # slope of the loss at the output's weighted sum
for i in reversed(range(len(params))):
a_in, _ = cache[i]
w, _ = params[i]
grads[i] = [a_in.T @ dz, dz.sum(axis=0)] # slopes for this layer's weights and biases
if i > 0:
_, z_prev = cache[i - 1]
dz = (dz @ w.T) * (z_prev > 0) # pass the blame back through ReLU
return grads
def gradient_check(params, x, y, h=1e-5):
"""Compare backprop with a brute-force slope for a few weights. Returns the largest gap."""
p, cache = forward(params, x)
grads = backward(params, cache, p, y)
worst = 0.0
for layer, (w, _) in enumerate(params):
for idx in [(0, 0), (w.shape[0] - 1, w.shape[1] - 1)]:
old = w[idx]
w[idx] = old + h
up = bce(forward(params, x)[0], y)
w[idx] = old - h
down = bce(forward(params, x)[0], y)
w[idx] = old
numeric = (up - down) / (2 * h)
worst = max(worst, abs(numeric - grads[layer][0][idx]))
return worst
def scores(p, y, threshold=0.5):
flagged = p >= threshold
truth = y == 1
tp, fp, fn = int((flagged & truth).sum()), int((flagged & ~truth).sum()), int((~flagged & truth).sum())
acc = float((flagged == truth).mean())
recall = tp / (tp + fn) if tp + fn else 0.0
precision = tp / (tp + fp) if tp + fp else 0.0
return acc, recall, precision
def train_numpy(params, x, y, lr, epochs):
history = []
for epoch in range(epochs + 1):
p, cache = forward(params, x)
loss = bce(p, y)
history.append(loss)
if epoch == epochs:
break
grads = backward(params, cache, p, y)
for (w, b), (gw, gb) in zip(params, grads):
w -= lr * gw # the gradient descent step from Unit 2
b -= lr * gb
return history
def train_torch(params, x, y, lr, epochs):
"""The same network, same starting weights, same steps, but PyTorch works out the gradients."""
try:
import torch
except ImportError:
raise SystemExit("\nPyTorch is not installed. See Set up for Unit 3, Steps 2 and 3.")
layers = []
for i, (w, b) in enumerate(params):
linear = torch.nn.Linear(w.shape[0], w.shape[1]).double()
with torch.no_grad():
linear.weight.copy_(torch.from_numpy(w.T.copy()))
linear.bias.copy_(torch.from_numpy(b.copy()))
layers.append(linear)
if i < len(params) - 1:
layers.append(torch.nn.ReLU())
model = torch.nn.Sequential(*layers)
optimizer = torch.optim.SGD(model.parameters(), lr=lr)
loss_fn = torch.nn.BCEWithLogitsLoss() # sigmoid and cross-entropy in one, numerically stable
xt, yt = torch.from_numpy(x), torch.from_numpy(y)
history = []
for epoch in range(epochs + 1):
optimizer.zero_grad()
loss = loss_fn(model(xt)[:, 0], yt)
history.append(loss.item())
if epoch == epochs:
break
loss.backward()
optimizer.step()
def predict(new_x):
with torch.no_grad(): # predicting only: no gradients needed
return torch.sigmoid(model(torch.from_numpy(new_x))[:, 0]).numpy()
return history, predict, model
def main() -> None:
parser = argparse.ArgumentParser(description="Train a small neural network on made-up invoice data.")
parser.add_argument("--hidden", type=int, default=8, help="hidden units; 0 means no hidden layer")
parser.add_argument("--lr", type=float, default=0.5, help="learning rate")
parser.add_argument("--epochs", type=int, default=2000, help="training steps over all rows")
parser.add_argument("--rows", type=int, default=2000, help="how many made-up invoice items")
parser.add_argument("--seed", type=int, default=1, help="seed for the starting weights")
parser.add_argument("--zeros", action="store_true", help="start every weight at zero")
parser.add_argument("--torch", action="store_true", help="also train the same network with PyTorch")
parser.add_argument("--plot", action="store_true", help="save decision_map.png")
parser.add_argument("--save", action="store_true", help="save the trained network to invoice_net.json")
args = parser.parse_args()
x_raw, y = make_invoices(args.rows)
cut = int(len(y) * 0.75)
x_train_raw, x_test_raw, y_train, y_test = x_raw[:cut], x_raw[cut:], y[:cut], y[cut:]
print(f"{args.rows} made-up invoice items, {y.mean():.0%} blocked. "
f"Training on {cut}, testing on {len(y) - cut}.")
# Scale with the training rows only, then apply the same numbers everywhere.
mean, std = x_train_raw.mean(axis=0), x_train_raw.std(axis=0)
x_train, x_test = (x_train_raw - mean) / std, (x_test_raw - mean) / std
base_acc = float((y_test == 0).mean())
print(f"Baseline (never block): test accuracy {base_acc:.1%}, catches 0% of blocked invoices\n")
params = init_params(2, args.hidden, args.seed, args.zeros)
n_weights = sum(w.size + b.size for w, b in params)
shape = f"2 inputs -> {args.hidden} hidden (ReLU) -> 1 output" if args.hidden else "2 inputs -> 1 output"
print(f"Network: {shape}, {n_weights} weights and biases, "
f"{'all starting at zero' if args.zeros else 'random start'}")
print(f"Gradient check: largest gap between backprop and brute force = "
f"{gradient_check(params, x_train, y_train):.1e}\n")
start = [[w.copy(), b.copy()] for w, b in params]
history = train_numpy(params, x_train, y_train, args.lr, args.epochs)
print(f"Learning rate {args.lr}, {args.epochs} epochs, full batch")
print(" epoch loss")
for epoch in sorted({0, 1, 10, 100, 500, 1000, args.epochs} & set(range(len(history)))):
print(f" {epoch:>5} {history[epoch]:.4f}")
if not np.isfinite(history[-1]) or history[-1] > history[0]:
print(" Warning: the loss ended higher than it started. Lower --lr.")
p_train, _ = forward(params, x_train)
p_test, _ = forward(params, x_test)
tr, te = scores(p_train, y_train), scores(p_test, y_test)
print(f"\nTrain accuracy {tr[0]:.1%} Test accuracy {te[0]:.1%}")
print(f"Test: catches {te[1]:.0%} of blocked invoices; {te[2]:.0%} of its flags are right")
if args.hidden and not args.zeros:
first = params[0][0]
dead = int((np.maximum(0, x_train @ first + params[0][1]).max(axis=0) == 0).sum())
print(f"Hidden units that never switch on: {dead} of {args.hidden}")
if args.zeros and args.hidden:
spread = float(np.ptp(params[0][0]))
print(f"Spread between hidden-unit weights after training: {spread:.4f} "
"(0 means every unit is identical, or none learned)")
examples = np.array([[2.0, 1.0], [7.0, 0.0], [0.0, 5.0], [-8.0, 0.0], [4.5, 2.5]])
p_ex, _ = forward(params, (examples - mean) / std)
print("\nWhat the network says about new invoice items:")
print(" price var qty var P(blocked)")
for (pv, qv), prob in zip(examples, p_ex):
print(f" {pv:+8.1f}% {qv:+6.1f}% {prob:9.2f}")
if args.torch:
t_hist, t_predict, _ = train_torch(start, x_train, y_train, args.lr, args.epochs)
t_acc = scores(t_predict(x_test), y_test)[0]
print(f"\nPyTorch, same start and steps: final loss {t_hist[-1]:.6f} "
f"(numpy: {history[-1]:.6f}), test accuracy {t_acc:.1%}")
if args.save:
out = HERE / "invoice_net.json"
out.write_text(json.dumps({
"inputs": ["price_variance_pct", "quantity_variance_pct"],
"scaler": {"mean": mean.tolist(), "std": std.tolist()},
"layers": [{"weights": w.tolist(), "biases": b.tolist()} for w, b in params],
"hidden_activation": "relu", "output_activation": "sigmoid",
"test_accuracy": round(te[0], 4),
}, indent=2))
print(f"\nSaved {out.name} (weights plus the scaler: both are needed to predict)")
if args.plot:
import matplotlib
matplotlib.use("Agg") # draw into a file, no window needed
import matplotlib.pyplot as plt
gx, gy = np.meshgrid(np.linspace(-12, 12, 200), np.linspace(-10, 10, 200))
grid = np.column_stack([gx.ravel(), gy.ravel()])
prob, _ = forward(params, (grid - mean) / std)
plt.figure(figsize=(6, 4.5))
plt.contourf(gx, gy, prob.reshape(gx.shape), levels=20, cmap="RdYlGn_r", alpha=0.7)
plt.colorbar(label="P(blocked)")
plt.scatter(x_test_raw[:, 0], x_test_raw[:, 1], c=y_test, cmap="coolwarm", s=8, edgecolors="none")
plt.axvline(PRICE_LIMIT, color="black", lw=0.8, ls="--")
plt.axhline(QTY_LIMIT, color="black", lw=0.8, ls="--")
plt.xlabel("price variance, % above PO price")
plt.ylabel("quantity variance, % above goods receipt")
plt.title(f"{args.hidden} hidden units: test accuracy {te[0]:.1%}")
plt.tight_layout()
plt.savefig(HERE / "decision_map.png", dpi=120)
print("\nSaved decision_map.png")
if __name__ == "__main__":
main()
2000 made-up invoice items, 26% blocked. Training on 1500, testing on 500.
Baseline (never block): test accuracy 74.0%, catches 0% of blocked invoices
Network: 2 inputs -> 8 hidden (ReLU) -> 1 output, 33 weights and biases, random start
Gradient check: largest gap between backprop and brute force = 6.6e-12
Learning rate 0.5, 2000 epochs, full batch
epoch loss
0 0.7561
1 0.6742
10 0.4853
100 0.2203
500 0.1747
1000 0.1679
2000 0.1654
Train accuracy 96.6% Test accuracy 96.6%
Test: catches 90% of blocked invoices; 97% of its flags are right
Hidden units that never switch on: 0 of 8
What the network says about new invoice items:
price var qty var P(blocked)
+2.0% +1.0% 0.03
+7.0% +0.0% 0.99
+0.0% +5.0% 1.00
-8.0% +0.0% 0.03
+4.5% +2.5% 0.31
Your numbers should match, because the data and starting weights come from fixed seeds. The gradient check gap may show a slightly different tiny number; anything below about 1e-6 is fine.
Baseline, 74.0%. Never blocking anything is right 74% of the time, because 74% of test invoices weren't blocked. That is why accuracy alone misleads, as the classification topic showed. Recall ("catches") and precision ("flags are right") tell the real story.
Gradient check. Backpropagation and the brute-force slope agree to about 12 decimal places. Your chain-rule code is right.
The loss. It starts near 0.69, the "don't know" value, and falls to 0.17. Most of the progress happens in the first 100 epochs.
Train versus test. Both are 96.6%. The network isn't memorizing; it learned something that carries over to new invoices. With 2% of labels flipped at random, about 98% is the best any model could do here.
New invoices. Price +7% and quantity +5% are clearly blocked; price -8% is not, because under-billing only triggers a message. Price +4.5% with quantity +2.5% is near both limits, and the network is unsure (0.31). That is honest: it is close to the corner.
With no hidden layer, the network is exactly logistic regression: 3 parameters. What you will see (the lines that matter):
Network: 2 inputs -> 1 output, 3 weights and biases, random start
...
Train accuracy 85.9% Test accuracy 85.8%
Test: catches 62% of blocked invoices; 79% of its flags are right
...
+7.0% +0.0% 0.58
+4.5% +2.5% 0.75
The straight line misses 38% of blocked invoices. It thinks price +7% alone is a coin flip, and it is more worried about the in-tolerance invoice at +4.5% and +2.5% than about the one clearly over the price limit. A line can't draw the L.
Network: 2 inputs -> 2 hidden (ReLU) -> 1 output, 9 weights and biases, random start
...
Train accuracy 96.7% Test accuracy 96.6%
Network: 2 inputs -> 2 hidden (ReLU) -> 1 output, 9 weights and biases, random start
...
Train accuracy 86.7% Test accuracy 85.6%
Test: catches 61% of blocked invoices; 79% of its flags are right
Two hidden units are enough, just like the hand-designed network in "Why a straight line fails here". But with --seed 2, the same network starts from different random weights and gets stuck at straight-line quality. The loss stops falling at 0.3334 and never recovers. This is the non-convex loss from scikit-learn's warning: more than one valley, and where you land depends on where you start. With 8 units there are more ways to find a good valley, which is one reason networks are usually built larger than the minimum.
Network: 2 inputs -> 8 hidden (ReLU) -> 1 output, 33 weights and biases, all starting at zero
...
0 0.6931
2000 0.5660
Train accuracy 74.7% Test accuracy 74.0%
Test: catches 0% of blocked invoices; 0% of its flags are right
Spread between hidden-unit weights after training: 0.0000 (0 means every unit is identical, or none learned)
...
Learning rate 20.0, 2000 epochs, full batch
...
1 1.9933
2000 1.7034
Warning: the loss ended higher than it started. Lower --lr.
Train accuracy 74.7% Test accuracy 74.0%
Hidden units that never switch on: 6 of 8
All zeros. Every hidden unit outputs 0, and ReLU passes no blame back through a zero. The hidden weights never move. Only the output bias learns, so the network says 0.25 for every invoice: the share of blocked invoices. It is no better than the baseline.
Learning rate 20. The loss jumps up on the first step and never settles. Worse, 6 of the 8 hidden units have been pushed so far negative that they output 0 for every training invoice. These are the dead ReLU units from Google's list, and they can't come back, because no blame reaches them.
Try --lr 0.05 as well: it trains, but slower, and ends at a loss of 0.1931 instead of 0.1654 after the same 2,000 epochs.
PyTorch, same start and steps: final loss 0.165382 (numpy: 0.165382), test accuracy 96.6%
The script copies your starting weights into a PyTorch model and runs the same 2,000 steps. PyTorch's autograd works out the gradients instead of your backward function, and the final loss matches to six decimal places. Your backpropagation is exactly what the library does. From here on, the course lets PyTorch do it.
Note two PyTorch details in train_torch. The loss is BCEWithLogitsLoss, which takes the output before the sigmoid; PyTorch's documentation says combining the two is more numerically stable. And optimizer.zero_grad() runs every step, because the autograd tutorial explains that PyTorch adds new gradients to old ones unless you clear them.
#Step 9: Draw the decision map and save the network
python neural_net.py --plot --save
The last lines are:
Saved invoice_net.json (weights plus the scaler: both are needed to predict)
Saved decision_map.png
Open decision_map.png from the VS Code file list. Green means "not blocked" and red means "blocked"; the dashed lines are the made-up limits. You should see an L-shaped red region hugging both lines, with rounded edges near the corner where the network is unsure. Now run python neural_net.py --hidden 0 --plot and open the picture again: a single straight edge, cutting through the L.
invoice_net.json holds the 33 learned numbers plus the scaling averages and spreads. The exercise uses it.
Unit 3 is a mechanics unit, so this section is short. It maps what you built onto SAP's three routes, as described by SAP's Architecture Center page on classic machine learning (last updated April 23, 2026).
The hana-ml Python client for SAP HANA (release 2.30, September 2026) has a neural_network module under hana_ml.algorithms.pal with MLPClassifier, MLPRegressor, MLPMultiTaskClassifier and MLPMultiTaskRegressor. Training runs in the database. The parameters map directly onto this topic:
This topic
MLPClassifier parameter in hana-ml 2.30
--hidden 8
hidden_layer_size, a tuple of sizes, one per hidden layer
ReLU in the hidden layer
activation; options include 'relu', 'tanh' and several sigmoid variants
Sigmoid at the output
output_activation
--lr, full batch
learning_rate, momentum, and training_style ('batch' or 'stochastic')
Scaling with training averages
normalization, with options 'no', 'z-transform' and 'scalar'
Random or zero start
weight_init, with options such as 'all-zeros', 'normal' and 'uniform'
--epochs
max_iter
The client's own example sets weight_init='normal' and normalization='z-transform'. After this topic you know why both matter. The multi-task classes add options such as batch normalization, a batch size and several optimizers.
SAP's Architecture Center recommends PAL and APL mainly for time series, anomaly detection, clustering and other in-database machine learning. For classification and regression on tables, it recommends starting with SAP-RPT-1.
As of September 2026, SAP describes SAP-RPT-1 as a relational pretrained transformer for structured business data, offered as a small and a large version in the generative AI hub, plus an open-source release. You send example rows with the request instead of training. The Architecture Center adds that it can be called from SAP HANA Cloud through a SQL stored procedure. Transformers are the neural network architecture of Unit 4.
The Architecture Center's guidance is to choose SAP AI Core when deep learning or large-scale neural networks are required, or custom models in TensorFlow, PyTorch or similar frameworks. SAP Learning describes how training runs there: your code in a container, orchestrated by Argo workflows; configurations holding parameters that can change for every run; resource plans with different CPU, GPU and memory; and APIs to register metrics.
This topic
SAP AI Core equivalent
neural_net.py with --torch
Your PyTorch training code, packaged in a container image
Beat a simple model first. Always report a straight-line model on the same held-back data. If the network's gain is small, the simpler model is cheaper to run and easier to explain.
Save the scaler with the weights.invoice_net.json stores both. A network served without its scaling numbers returns confident nonsense, because it expects inputs on the training scale.
Record every run. Seed, hidden size, learning rate, epochs, data version, loss curve, test scores. Step 6 showed two runs of the same network ending 11 points apart. In AI Core, that means configuration parameters plus registered metrics.
Choose on held-back data, not the test set. Picking the hidden size or the seed by test score leaks the test into training. Use a validation split or cross-validation, as scikit-learn's guide suggests for tuning.
Watch for dead units and blow-ups. Stop a training job and alert when the loss grows or becomes not-a-number. Counting units that never switch on is a cheap health check.
Explainability. A network's 33 weights don't read like a rule. For decisions such as payment blocks, auditors need a reason per invoice. Plan how you will explain predictions before you build, or keep the decision in configured rules and use the model only to prioritize work.
Authorizations and data. Real invoice data carries supplier names, prices and bank details. The access rules from Calling your first SAP API apply wherever training runs; Unit 11 goes deeper.
Clean core. Train and serve side by side (AI Core) or in the database (PAL). Don't change standard invoice verification logic to host a model.
Training a model to learn a configured rule. It is slower, less exact and harder to audit than reading the configuration.
Trusting one run. A bad starting point can leave a network stuck at straight-line quality. Try several seeds and record them.
Starting every weight at the same value. Identical units learn identical things; with zeros and ReLU, nothing learns at all.
A learning rate copied from another project. Too large kills ReLU units for good. Watch the first epochs of the loss curve.
Forgetting to scale, or scaling with the wrong numbers. scikit-learn's guide calls networks sensitive to feature scaling. Fit the scaler on training rows and reuse it everywhere.
Reading accuracy alone. With 74% of invoices not blocked, "74% accurate" is what doing nothing scores.
Forgetting zero_grad() in PyTorch. Gradients add up across steps, and training quietly goes wrong.
#Exercise: load your network and score invoices without training
You will use the saved network to score new invoices in a separate script, as a service would, and record which settings train well. The prediction script is the seed of the AI API you build in Unit 6.
In the terminal, inside unit03 with (.venv) showing, run each command below and write down the test accuracy and the "catches" percentage:
Run python neural_net.py --save to write invoice_net.json with the default 8-unit network.
In VS Code, create unit03/predict_invoice.py, paste the code below and save:
"""Score one invoice item with the network saved by neural_net.py --save.
How to run (from the unit03 folder):
python predict_invoice.py --price 7 --qty 0
Uses only numpy and the saved file. No training happens here.
"""
import argparse
import json
from pathlib import Path
import numpy as np
HERE = Path(__file__).parent
def main() -> None:
parser = argparse.ArgumentParser(description="Score one invoice item with invoice_net.json.")
parser.add_argument("--price", type=float, required=True, help="% invoice price above the PO price")
parser.add_argument("--qty", type=float, required=True, help="% invoiced quantity above goods receipt")
args = parser.parse_args()
path = HERE / "invoice_net.json"
if not path.exists():
raise SystemExit("invoice_net.json not found. Run: python neural_net.py --save")
net = json.loads(path.read_text())
x = np.array([args.price, args.qty])
a = (x - np.array(net["scaler"]["mean"])) / np.array(net["scaler"]["std"]) # same scaling as training
layers = net["layers"]
for i, layer in enumerate(layers):
z = a @ np.array(layer["weights"]) + np.array(layer["biases"]) # weighted sum plus bias
a = 1 / (1 + np.exp(-z)) if i == len(layers) - 1 else np.maximum(0.0, z) # sigmoid last, ReLU before
print(f"Price {args.price:+.1f}%, quantity {args.qty:+.1f}%: P(blocked) = {a[0]:.2f}")
if __name__ == "__main__":
main()
You should see P(blocked) = 0.99 and P(blocked) = 0.31, the same as in neural_net.py's output. If they differ, check that you ran Step 2 of this exercise with no other options.
Score three invoices of your own choosing: one clearly fine, one clearly over a limit, and one right at a corner, such as --price 5 --qty 3.
Create unit03/notes_neural_networks.md with three short sections:
Hidden units: a table of the six runs from step 1 with test accuracy and catches, and one sentence on the smallest network you would trust.
Seeds: one sentence on what --seed 2 and --seed 3 showed, and what you would do about it in a real project.
Scoring: your three invoices from step 5 with their probabilities, and one sentence on why predict_invoice.py needs the scaler numbers.
Save your work:
git add unit03/predict_invoice.py unit03/notes_neural_networks.md
git commit -m "Score invoices with a saved neural network, plus notes"
Done when:predict_invoice.py prints 0.99 and 0.31 for the two invoices in step 4; your notes list all six runs with test accuracy and catches; the seed and scoring sections are filled in; and git log shows the commit.
Pick one answer for each question. The explanation appears after you choose.
1Why does a stack of layers without activation functions learn nothing a single layer couldn't?
Answer: B. Linear operations on linear operations stay linear, as Google's crash course puts it. The ReLU bend in each hidden unit is what lets the network combine hinges into shapes such as the L of the blocked-invoice region.
2In backward, what does the line dz = (dz @ w.T) * (z_prev > 0) do?
Answer: D. Blame from the output reaches each hidden unit in proportion to its outgoing weight. ReLU is flat below zero, so units that output 0 get no blame. The weight update itself happens later, in train_numpy.
3The straight-line model (--hidden 0) catches 62% of blocked invoices; 2 hidden units catch 90%. Why?
Answer: A. Blocked means price over its limit or quantity over its limit. A single line can't draw that corner. Each hidden unit acts as a hinge at one limit, and the output adds them, as in the hand-designed network.
4You train with --zeros and every invoice gets P(blocked) = 0.25. What happened?
Answer: C. With all weights at zero, every hidden unit's output is 0 and ReLU passes no gradient through it. The hidden weights never change. The output bias alone learns the share of blocked invoices, 0.25, so the network matches the do-nothing baseline.
5After a run with --lr 20, 6 of 8 hidden units never switch on. What are they, and what fixes it?
Answer: D. Huge steps pushed their weighted sums below zero for every invoice, so they always output 0 and get no gradient. Google's crash course names lowering the learning rate, or a ReLU variant such as LeakyReLU, as fixes.
6The same 2-unit network scores 96.6% with --seed 1 and 85.6% with --seed 2. What would you do in a real project?
Answer: B. A network's loss has more than one valley, and the start decides where it lands. Recording runs and choosing on validation data keeps the test set honest. Picking by test score leaks the test into model selection.
7Why does train_torch use BCEWithLogitsLoss on the raw output instead of applying a sigmoid first?
Answer: C. PyTorch's documentation says combining the sigmoid and binary cross-entropy in one class is more numerically stable than applying them separately. The prediction function still applies a sigmoid, inside torch.no_grad().
8Your team wants a neural network on invoice data already in SAP HANA, trained next to the data. Which SAP option fits?
Answer: D. hana-ml's neural_network module includes MLPClassifier, which runs PAL's multi-layer perceptron in the database. SAP-RPT-1 is pretrained and needs no training, and AI Core is for custom deep learning code.
9A model served in production returns near-certain "blocked" for most invoices, even small variances, though test results were good. What do you check first?
Answer: A. The network was trained on standardized inputs. If the service sends raw percentages, weighted sums land far from the training range and outputs saturate. That is why invoice_net.json stores the scaler with the weights.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Neural network models (supervised) (scikit-learn 1.9 user guide)— multi-layer perceptron; non-convex loss with more than one local minimum, so different random starts give different results; needs tuning of hidden units, layers and iterations; sensitive to feature scaling, fit the scaler on training data only
BCEWithLogitsLoss (PyTorch 2.13 documentation)— combines a sigmoid and binary cross-entropy in one class; more numerically stable than the two separately; pos_weight for class imbalance
Classic ML Scenarios (SAP Architecture Center, last updated April 23, 2026)— three routes; start with RPT-1 for classification and regression; HANA PAL/APL for time series, anomaly detection, clustering and in-database ML; AI Core when deep learning or large-scale neural networks, or TensorFlow and PyTorch models, are needed
Entering Invoices with Variances (SAP Learning, Invoice Verification in SAP S/4HANA)— quantity, price, order price quantity and date variances; tolerance limits set in Customizing; above the upper limit the invoice is blocked for payment, below the lower limit only a message; tolerance keys such as PP and DQ; the whole invoice is blocked even if one item varies