Train a small GPT on your laptop to write SAP-style process notes, and see first-hand how language models learn, invent details and memorize their training data.
A GPT is a transformer that has learned one skill: guess the next piece of text. In this topic, learners build a tiny one and train it on their own laptop.
The tiny GPT reads made-up SAP process notes, such as "Sales order 14212 for customer C-1032 blocked: credit limit exceeded". It sees one character at a time. Its only task is to guess the next character. At first it guesses at random and writes gibberish. After a few minutes of practice, it writes new notes that look real.
That is how large language models (LLMs) are made, too. The design is the same as the transformer from the previous topic. The task is the same: guess the next token. What differs is scale. The tiny GPT has under a million learned numbers and reads about 320,000 characters. Large models have billions of numbers and read trillions of tokens.
Three lessons come out of the tiny GPT that matter to every leader:
It learns patterns, not facts. It writes notes that look right, with numbers that are often wrong.
It can memorize. Trained too long on too little data, it repeats training notes word for word, customer numbers included.
Training from scratch is the wrong default. Even this toy takes minutes. Real models take data centres.
Fluent is not the same as correct. In our test, the trained model wrote an invoice note with a goods receipt of 1,460 pieces against an invoice of 230. Every note it saw had a receipt slightly below the invoice. The sentence was fluent; the number was invented. Large models do the same thing with more polish. This is why SAP processes such as three-way match need checks against system data, not trust in generated text.
Models can leak what they learned. Trained on only 100 notes, our tiny GPT copied 42% of its outputs straight from the training data, customer numbers and amounts included. Researchers showed in 2021 that GPT-2 could be made to repeat names, phone numbers and email addresses from its training data. Some appeared in just one document. If a team proposes training or fine-tuning a model on customer data, ask what stops it from repeating that data to the wrong person.
Scale drives cost. A 2022 study trained over 400 models and concluded that training data should grow with model size. Their 70-billion-parameter model, Chinchilla, read 1.4 trillion tokens. Very few companies should pay for that. Most SAP projects call a model someone else trained, or use one SAP pretrained.
A concrete order-to-cash example. A team wants drafts of customer messages about blocked sales orders. Option one: train a model on five years of order notes. Option two: call a hosted LLM through SAP's generative AI hub, and give it the facts of each order at request time. The tiny GPT shows why option two is the usual starting point. Training needs data, hardware and time. It also bakes customer details into the model, where no authorization check can reach them.
SAP doesn't ask customers to pretrain language models. As of October 2026, SAP offers three paths:
Pretrained models in the generative AI hub. The generative AI hub, part of SAP AI Core, gives access to LLMs from several providers. SAP's AI Core guide says it is available only in the extended service plan. Unit 5 sets up access.
SAP's own pretrained models. SAP's release highlights for Q4 2025 describe SAP-ABAP-1, trained on more than 250 million lines of ABAP code, 30 million lines of CDS code and technical documentation. They also describe SAP-RPT-1, a pretrained model for business tables. Both are offered through the generative AI hub, and both can be tried in its trial. SAP did the expensive training; customers call the result.
Your own training jobs in SAP AI Core. SAP AI Core runs training pipelines as batch jobs. You package your code, choose a resource plan with the CPUs, GPUs and memory you need, and SAP AI Core stores the trained model. This suits small custom models on your own data, such as a table model or a classifier, more than pretraining an LLM.
The pattern is clear. Pretraining a useful LLM is a research-lab activity. A project team's job is to choose a model, give it the right context, and check its output.
Are we training a model, fine-tuning one, or calling a pretrained one? Why that choice, and what would the simpler option cost?
If we train or fine-tune on our data, which records go in? Could the model repeat customer, supplier or employee data to a user who isn't allowed to see it?
How do we check generated numbers, such as quantities, amounts and order numbers, against the SAP system before anyone acts on them?
Which SAP service plan does the proposal need: standard for training jobs in SAP AI Core, extended for the generative AI hub?
How was the model tested on data it didn't train on, and what was the result?
Who owns a model we train ourselves: retraining, monitoring and deletion when data must be removed?
"The model looks up answers." A GPT stores no records to look up. It learned patterns in its weights and writes text that fits them, which is why it invents plausible numbers.
"A model that writes fluent notes understands the process." Our tiny GPT writes perfect-looking three-way match notes and breaks the rule behind them.
"Training data stays private inside the model." A model can repeat training text word for word, especially with little data or long training.
"Training our own LLM gives us control at a fair price." Pretraining needs vast data and compute. Most teams get better value from a pretrained model plus their own data at request time.
"Lower training loss always means a better model." In the exercise, training loss keeps falling while the model gets worse on new text and starts copying.
Pick one answer for each question. The explanation appears after you choose.
1What is the one task a GPT is trained to do?
Answer: B. A GPT is trained only to predict the next token. Writing whole notes, answers or emails comes from repeating that one guess many times. It has no database to look things up in.
2The tiny GPT wrote an invoice note with a goods receipt of 1,460 pieces against an invoice of 230. What does that show?
Answer: D. Every training note had a receipt slightly below the invoice, yet the model wrote a fluent note that breaks that pattern. Language models learn what text looks like, not the facts behind it. Generated numbers need checks against system data.
3A team plans to fine-tune a model on five years of customer order notes. What risk should you raise first?
Answer: A. Models can memorize and repeat training text. The tiny GPT copied 42% of its outputs from 100 training notes, and researchers extracted names and phone numbers from GPT-2. No authorization check works inside a model's weights.
4Why is pretraining your own LLM the wrong default for most SAP projects?
Answer: C. The Chinchilla study trained a 70-billion-parameter model on 1.4 trillion tokens. Few companies should pay for that. The generative AI hub and SAP's own pretrained models offer the result of that work as a service.
5Which path does SAP offer for running your own small training job?
Answer: D. SAP AI Core runs pipelines as batch jobs, for example to train models, with resource plans that differ in CPUs, GPUs and memory. The generative AI hub, in the extended plan, is for calling pretrained models.
6In the exercise, training loss keeps falling while loss on new text rises. What is happening?
Answer: B. Falling training loss with rising loss on unseen text is the sign of overfitting. In the exercise, the same model also started copying training notes word for word. Judge a model on data it didn't train on.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A GPT does exactly one thing: given some text, it scores every possible next token. Everything else is a loop around that one guess.
Training is the loop of guessing, measuring how wrong the guess was, and nudging the weights. You did this in Gradient descent by hand. Here the model and the data are bigger, but the loop is the same.
Writing is the loop of guessing, picking one token, adding it to the text, and guessing again.
The trick that makes training efficient is the causal mask from the transformer. Because each position sees only earlier tokens, one window of 64 characters gives 64 guesses at once. Each position predicts the character after it. One forward pass, 64 lessons.
The tiny GPT uses the simplest tokenizer there is: one token per character. It collects every different character in the training file, sorts them and numbers them. Our made-up SAP notes use 57 different characters, so the vocabulary has 57 entries. Tiny Shakespeare, a classic practice text, uses 65.
Real LLMs split text into word pieces instead. GPT-2 had a vocabulary of 50,257 pieces, as you saw in the previous topic's parameter count. Word pieces make text shorter in tokens, so the same context window holds more meaning. Characters keep this topic simple: no extra library, and every output is readable.
#Inputs and targets: the same text, shifted by one
Training takes random windows of the text. The target for each window is the same text, moved one character to the right.
flowchart LR
T["Text: Order 4711 on hold"] --> X["Input: Order 4711 on hol"]
T --> Y["Target: rder 4711 on hold"]
X --> M[Tiny GPT]
M --> P[A guess at every position]
P --> L[Loss: guesses vs targets]
Y --> L
At the position of O, the target is r. At Order 47, the target is 1. With the causal mask, no position can peek at its own target. So one window teaches as many lessons as it has characters.
The loss is cross-entropy: for each position, minus the log of the probability the model gave the correct next character. You met its yes/no version in Neural networks from scratch. Here it covers 57 choices instead of two.
That gives a useful check before training. An untrained model spreads its probability about evenly, about 1/57 for each character. Its loss should then be close to ln(57), which is 4.04. The script prints this number, and the first loss it measures is 4.09. If your first loss is far above ln(vocabulary size), something is wrong with the setup, before any training happens.
The loss also has a floor. Our notes contain random order numbers and amounts. No model can predict a random digit, so the loss can never reach zero. On the made-up notes, it settles near 0.31. On Shakespeare it stays far higher, because real language is less predictable than templates.
Take a batch of 32 random windows of 64 characters from the training part of the text.
Run the model to get logits of shape (32, 64, 57): a score for every vocabulary character, at every position, in every window.
Compute the cross-entropy loss against the shifted targets.
Run backward() to get the gradients, and clip them to a maximum overall size of 1.0, a common guard against a single wild step.
Let the AdamW optimizer update the weights. AdamW is a refined form of gradient descent that adapts the step size for each weight. It is the usual choice for transformers.
Every few hundred steps, the script measures the loss on 20 batches from the training part and 20 from a check part the model never trains on. That second number is the honest one.
The script keeps the first 90% of the text for training and the last 10% for checking, the same split nanoGPT uses. Two curves come out:
What you see
What it means
Both losses fall together
The model is learning patterns that carry over to new text
Both flatten at the same level
The model has learned what it can at this size; more steps add little
Training loss falls, check loss rises
Overfitting: the model fits its training text and gets worse on new text
With 3,000 notes, both losses flatten together near 0.31. With 100 notes, training loss keeps falling to 0.13 while check loss climbs above 0.6. The exercise shows what the model does in that state: it copies.
Once trained, the model writes by repeating its one guess:
flowchart LR
P[Prompt] --> C[Keep the last<br/>64 characters]
C --> M[Tiny GPT]
M --> S[Scores for the<br/>next character]
S --> D[Draw one character<br/>by probability]
D --> A[Append to the text]
A -->|repeat| C
A -->|line ends| E[Done]
Two details matter:
The context limit. The model has 64 position slots. Before each guess, the script keeps only the last 64 characters. nanoGPT's generate does the same: if the text grows too long, it is cropped to the context size. Anything earlier is gone from the model's view. Large models have the same limit, just bigger.
Drawing, not picking the top. The script draws the next character at random, weighted by the probabilities. That is why two runs give different notes. The --temperature option sharpens or flattens the probabilities before the draw. The next topic, How LLMs generate text, explores these choices in depth.
The trained weights go into a checkpoint file with torch.save. The script saves the weights (the state dict), the character list and the model size, as PyTorch's serialization notes recommend, rather than the whole Python object.
Loading uses torch.load(..., weights_only=True). PyTorch's documentation explains that this setting, the default since version 2.6, restricts loading to what a state dict needs. A normal Python "pickle" file can run code when it is loaded. So a model file from an unknown source is a security risk unless it is loaded this way. Treat model files like any other software you install.
The small size is close to nanoGPT's suggested CPU settings, which its README says reach a validation loss of about 1.88 on Tiny Shakespeare in about 3 minutes. The medium size has the shape of nanoGPT's "baby GPT", which the README reports reaches 1.4697 in about 3 minutes on one A100 GPU. The README now points readers to a newer project, nanochat, for anything beyond learning.
#Build it yourself: train a tiny GPT on SAP-style notes
You will generate a file of made-up SAP process notes, train the transformer from the previous topic on it, and make it write new notes. Then you will check whether the new notes are original or copied. Everything runs on your laptop with made-up data.
Before you start: complete Set up your computer for this course and Set up for Unit 4. You also need transformer_block.py from The transformer architecture, saved in your unit04 folder: this topic imports its TinyTransformer class. Your timing_notes.md from the setup topic tells you whether the small size is quick enough on your computer.
flowchart LR
G[make-data] --> N[sap_notes.txt<br/>3,000 made-up notes]
N --> T[train] --> K[tiny_gpt.pt]
K --> W[generate] --> O[New notes]
K --> C[copies] --> R[Copied or original?]
In VS Code's file list, right-click unit04, choose New File and name it tiny_gpt.py.
Paste the code below and save with File > Save. The next topic loads the model this script trains, so keep the name.
"""Unit 4: build a tiny GPT. Train the transformer from transformer_block.py to write text.
The model reads characters and learns one thing: predict the next character.
Trained on made-up SAP-style process notes, it starts writing notes of its own.
How to run (from the unit04 folder, with the course .venv turned on):
python tiny_gpt.py make-data # write sap_notes.txt: 3,000 made-up SAP process notes
python tiny_gpt.py train # train the small model on sap_notes.txt, save tiny_gpt.pt
python tiny_gpt.py generate # write 5 new notes with the trained model
python tiny_gpt.py generate --prompt "Invoice" # start the notes with your own text
python tiny_gpt.py copies # how many generated notes copy a training note word for word
Optional:
python tiny_gpt.py train --steps 300 # a quick run to see it work (about a minute)
python tiny_gpt.py train --size medium # a bigger model; use a GPU notebook for this one
python tiny_gpt.py get-shakespeare # download a classic practice text (needs internet)
python tiny_gpt.py train --data shakespeare.txt --model shakespeare.pt --prompt "ROMEO:"
python tiny_gpt.py generate --model shakespeare.pt --prompt "ROMEO:" --no-stop --chars 300
"""
import argparse
import math
import random
import time
import urllib.request
from pathlib import Path
import torch
import torch.nn.functional as F
from transformer_block import TinyTransformer
SIZES = { # the same two sizes as bench_tiny_gpt.py in the Unit 4 setup topic
"small": dict(n_layer=4, n_embd=128, n_head=4, context=64, batch=32),
"medium": dict(n_layer=6, n_embd=384, n_head=6, context=256, batch=64),
}
SHAKESPEARE_URL = "https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt"
# ---------------------------------------------------------------- 1. data: made-up SAP-style notes
def make_note(r: random.Random) -> str:
"""One made-up process note. Numbers, names and rules are invented for practice."""
customer, vendor = f"C-{r.randint(1000, 1999)}", f"V-{r.randint(2000, 2999)}"
material, plant = f"M-{r.randint(4000, 4999)}", r.choice(["1010", "1020", "1710", "2010"])
amount = f"{r.randint(1, 99)},{r.randint(0, 999):03d}.{r.choice(['00', '50'])} {r.choice(['EUR', 'USD'])}"
qty_inv = r.randint(5, 50) * 10
qty_gr = qty_inv - r.randint(1, 4) * 10
kind = r.choice(["o2c", "p2p", "mrp"])
if kind == "o2c": # order-to-cash: blocked sales orders
so = f"{r.randint(10000, 19999)}"
reason = r.choice(["credit limit exceeded by " + amount, "overdue items on account",
"missing export license", "price below minimum"])
action = r.choice(["Credit team to review before release.", "Sales rep to call the customer.",
"Release after payment arrives.", "Check with the credit manager today."])
start = r.choice([f"Sales order {so} for customer {customer} blocked: ",
f"Order {so} ({customer}) on hold, ",
f"Delivery block on sales order {so}, customer {customer}: "])
return start + reason + ". " + action
if kind == "p2p": # procure-to-pay: three-way match exceptions
inv, po = f"{r.randint(510000, 519999)}", f"45{r.randint(10000, 99999)}"
issue = r.choice([f"invoice quantity {qty_inv} PC, goods receipt {qty_gr} PC",
f"price differs from the purchase order by {amount}",
"no goods receipt posted yet"])
action = r.choice([f"Buyer to confirm with supplier {vendor}.", "AP clerk to park the invoice.",
"Warehouse to check the receipt.", f"Ask {vendor} for a credit memo."])
start = r.choice([f"Invoice {inv} for PO {po} held: ", f"Three-way match failed on PO {po}, ",
f"Payment block on invoice {inv} (PO {po}): "])
return start + issue + ". " + action
order = f"{r.randint(1000000, 1999999)}" # plan-to-produce: MRP exceptions
issue = r.choice(["start date in the past", "stock below safety stock",
"reschedule in", "reschedule out", "opening date in the past"])
action = r.choice([f"Planner to reschedule order {order}.", "Expedite with the supplier.",
f"Check capacity at plant {plant}.", "Convert the planned order today."])
start = r.choice([f"MRP exception for material {material} at plant {plant}: ",
f"Material {material}, plant {plant}, exception: "])
return start + issue + ". " + action
def make_data(path: Path, notes: int, seed: int) -> None:
r = random.Random(seed)
lines = [make_note(r) for _ in range(notes)]
path.write_text("\n".join(lines) + "\n", encoding="utf-8")
print(f"Wrote {notes:,} made-up notes ({path.stat().st_size:,} characters) to {path}")
print("First three:")
for line in lines[:3]:
print(" " + line)
def get_shakespeare(path: Path) -> None:
print(f"Downloading {SHAKESPEARE_URL}")
with urllib.request.urlopen(SHAKESPEARE_URL, timeout=30) as response:
path.write_bytes(response.read())
print(f"Saved {path.stat().st_size:,} characters to {path}")
# ---------------------------------------------------------------- 2. tokenizer: one token per character
class CharTokenizer:
def __init__(self, chars: list):
self.chars = chars
self.stoi = {c: i for i, c in enumerate(chars)}
def encode(self, text: str) -> list:
return [self.stoi[c] for c in text]
def decode(self, ids: list) -> str:
return "".join(self.chars[i] for i in ids)
# ---------------------------------------------------------------- 3. device and batches
def pick_device(choice: str) -> torch.device:
if choice != "auto":
return torch.device(choice)
if torch.cuda.is_available():
return torch.device("cuda")
if torch.backends.mps.is_available():
return torch.device("mps")
return torch.device("cpu")
def get_batch(data: torch.Tensor, context: int, batch: int, device: torch.device):
"""Random windows of text. The target is the same window shifted one character to the right."""
starts = torch.randint(0, len(data) - context - 1, (batch,))
x = torch.stack([data[s:s + context] for s in starts])
y = torch.stack([data[s + 1:s + context + 1] for s in starts])
return x.to(device), y.to(device)
def loss_on(model, data, cfg, device, batches: int = 20) -> float:
model.eval()
total = 0.0
with torch.no_grad():
for _ in range(batches):
x, y = get_batch(data, cfg["context"], cfg["batch"], device)
logits = model(x)
total += F.cross_entropy(logits.reshape(-1, logits.shape[-1]), y.reshape(-1)).item()
model.train()
return total / batches
# ---------------------------------------------------------------- 4. generation: one character at a time
@torch.no_grad()
def generate(model, tok: CharTokenizer, prompt: str, max_chars: int, context: int,
temperature: float, device: torch.device, stop_at_newline: bool = True) -> str:
model.eval()
idx = torch.tensor([tok.encode(prompt)], device=device)
for _ in range(max_chars):
window = idx[:, -context:] # the model only has `context` position slots
logits = model(window)[:, -1, :] / temperature # scores for the next character only
probs = F.softmax(logits, dim=-1)
nxt = torch.multinomial(probs, num_samples=1) # draw one character at random, by probability
idx = torch.cat([idx, nxt], dim=1)
if stop_at_newline and tok.chars[nxt.item()] == "\n":
break # one note per line: stop at the line end
return tok.decode(idx[0].tolist())
# ---------------------------------------------------------------- 5. training
def indent(text: str) -> str:
return "\n".join(" | " + line for line in text.strip("\n").splitlines())
def train(args) -> None:
path = Path(args.data)
if not path.exists():
raise SystemExit(f"{path} not found. Run: python tiny_gpt.py make-data")
text = path.read_text(encoding="utf-8")
tok = CharTokenizer(sorted(set(text)))
data = torch.tensor(tok.encode(text), dtype=torch.long)
split = int(0.9 * len(data)) # first 90% to learn from, last 10% to check on
train_data, val_data = data[:split], data[split:]
cfg = SIZES[args.size]
device = pick_device(args.device)
torch.manual_seed(args.seed)
model = TinyTransformer(len(tok.chars), cfg["context"], cfg["n_embd"], cfg["n_head"],
cfg["n_layer"]).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=args.lr)
params = sum(p.numel() for p in model.parameters())
print(f"Data: {path} ({len(text):,} characters, {len(tok.chars)} different ones)")
print(f"Model '{args.size}': {params / 1e6:.2f} million parameters, {cfg['n_layer']} blocks, "
f"context {cfg['context']} characters, on {device.type}")
print(f"Loss if guessing evenly: ln({len(tok.chars)}) = {math.log(len(tok.chars)):.2f}")
before = generate(model, tok, args.prompt, 80, cfg["context"], 1.0, device, stop_at_newline=False)
print("Before training, it writes:\n" + indent(before))
every = max(1, args.steps // 5)
start = time.time()
for step in range(args.steps + 1):
if step % every == 0 or step == args.steps:
tr, va = loss_on(model, train_data, cfg, device), loss_on(model, val_data, cfg, device)
print(f" step {step:>5} train loss {tr:.2f} check loss {va:.2f} [{time.time() - start:.0f} s]")
if step == args.steps:
break
x, y = get_batch(train_data, cfg["context"], cfg["batch"], device)
logits = model(x) # (batch, context, vocabulary)
loss = F.cross_entropy(logits.reshape(-1, logits.shape[-1]), y.reshape(-1))
opt.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step()
after = generate(model, tok, args.prompt, 240, cfg["context"], 1.0, device, stop_at_newline=False)
print("After training, it writes:\n" + indent(after))
torch.save({"model": model.state_dict(), "chars": tok.chars, "size": args.size,
"data": str(path)}, args.model)
print(f"Saved the trained model to {args.model}")
def load(model_path: str, device: torch.device):
if not Path(model_path).exists():
raise SystemExit(f"{model_path} not found. Run: python tiny_gpt.py train")
ckpt = torch.load(model_path, map_location=device, weights_only=True) # numbers and text only, no code
cfg = SIZES[ckpt["size"]]
tok = CharTokenizer(ckpt["chars"])
model = TinyTransformer(len(tok.chars), cfg["context"], cfg["n_embd"], cfg["n_head"],
cfg["n_layer"]).to(device)
model.load_state_dict(ckpt["model"])
return model, tok, cfg, ckpt
def generate_cmd(args) -> None:
device = pick_device(args.device)
model, tok, cfg, _ = load(args.model, device)
unknown = sorted(set(args.prompt) - set(tok.chars))
if unknown:
raise SystemExit(f"The prompt uses characters the model never saw: {unknown}")
torch.manual_seed(args.seed)
for _ in range(args.count):
print(generate(model, tok, args.prompt, args.chars, cfg["context"], args.temperature, device,
stop_at_newline=not args.no_stop).strip())
def copies_cmd(args) -> None:
device = pick_device(args.device)
model, tok, cfg, ckpt = load(args.model, device)
text = Path(ckpt["data"]).read_text(encoding="utf-8")
split = int(0.9 * len(text))
seen = set(text[:split].splitlines()) # notes the model trained on
torch.manual_seed(args.seed)
copies = []
for _ in range(args.count):
note = generate(model, tok, "\n", args.chars, cfg["context"], 1.0, device).strip()
if note in seen:
copies.append(note)
print(f"Generated {args.count} notes; {len(copies)} are word-for-word copies of a training note "
f"({len(copies) / args.count:.0%}).")
print(f"The training part of {ckpt['data']} holds {len(seen):,} different notes.")
for note in sorted(set(copies))[:2]:
print(f" copied: {note}")
def main() -> None:
p = argparse.ArgumentParser(description="Build, train and sample a tiny character-level GPT.")
sub = p.add_subparsers(dest="command", required=True)
d = sub.add_parser("make-data", help="write made-up SAP-style notes to a text file")
d.add_argument("--notes", type=int, default=3000)
d.add_argument("--out", default="sap_notes.txt")
d.add_argument("--seed", type=int, default=0)
s = sub.add_parser("get-shakespeare", help="download the Tiny Shakespeare practice text")
s.add_argument("--out", default="shakespeare.txt")
t = sub.add_parser("train", help="train the model and save it")
t.add_argument("--data", default="sap_notes.txt")
t.add_argument("--size", choices=sorted(SIZES), default="small")
t.add_argument("--steps", type=int, default=2000)
t.add_argument("--lr", type=float, default=1e-3, help="learning rate")
t.add_argument("--prompt", default="Sales order", help="text to start the before/after samples")
t.add_argument("--model", default="tiny_gpt.pt", help="file to save the trained model in")
t.add_argument("--device", choices=["auto", "cpu", "cuda", "mps"], default="auto")
t.add_argument("--seed", type=int, default=0)
g = sub.add_parser("generate", help="write new text with a trained model")
g.add_argument("--prompt", default="\n", help="text to start with (default: a new line)")
g.add_argument("--count", type=int, default=5)
g.add_argument("--chars", type=int, default=200, help="maximum characters per note")
g.add_argument("--temperature", type=float, default=1.0)
g.add_argument("--no-stop", action="store_true", help="keep writing past the end of a line")
c = sub.add_parser("copies", help="count generated notes that copy a training note exactly")
c.add_argument("--count", type=int, default=200)
c.add_argument("--chars", type=int, default=200)
for sp in (g, c):
sp.add_argument("--model", default="tiny_gpt.pt")
sp.add_argument("--device", choices=["auto", "cpu", "cuda", "mps"], default="auto")
sp.add_argument("--seed", type=int, default=0)
args = p.parse_args()
if args.command == "make-data":
make_data(Path(args.out), args.notes, args.seed)
elif args.command == "get-shakespeare":
get_shakespeare(Path(args.out))
elif args.command == "train":
train(args)
elif args.command == "generate":
generate_cmd(args)
else:
copies_cmd(args)
if __name__ == "__main__":
main()
Wrote 3,000 made-up notes (319,555 characters) to sap_notes.txt
First three:
Three-way match failed on PO 4538631, no goods receipt posted yet. AP clerk to park the invoice.
Material M-4097, plant 1710, exception: reschedule out. Planner to reschedule order 1346236.
Order 16534 (C-1444) on hold, credit limit exceeded by 71,488.50 USD. Check with the credit manager today.
The notes mix three running examples of this course: blocked sales orders (order-to-cash), three-way match exceptions (procure-to-pay) and MRP exceptions (plan-to-produce). Customer numbers, amounts and the rules are invented. One rule is hidden in the data on purpose: in every quantity note, the goods receipt is 10 to 40 pieces below the invoice quantity. Watch whether the model learns it.
Open sap_notes.txt in VS Code to look at it. It is plain text, one note per line.
What success looks like (times and sample notes differ between computers):
Data: sap_notes.txt (319,555 characters, 57 different ones)
Model 'small': 0.81 million parameters, 4 blocks, context 64 characters, on cpu
Loss if guessing evenly: ln(57) = 4.04
Before training, it writes:
| Sales order2D vlD.0aywBTC99qhunB:Catol
| h
| Pq)fhnTck37cEfuq0pVPl9Br-qE-esvy)B6cSs)WkO00C1C64
step 0 train loss 4.09 check loss 4.09 [1 s]
step 400 train loss 0.39 check loss 0.39 [41 s]
step 800 train loss 0.32 check loss 0.33 [81 s]
step 1200 train loss 0.31 check loss 0.32 [121 s]
step 1600 train loss 0.30 check loss 0.32 [161 s]
step 2000 train loss 0.31 check loss 0.31 [203 s]
After training, it writes:
| Sales order 16161, customer C-1398: overdue items on account. Sales rep to call the customer.
| Delivery block on sales order 13710, customer C-1978: overdue items on account. Release after payment arrives.
| MRP exception for material M-4048 at plant 171
Saved the trained model to tiny_gpt.pt
How to read it:
The first loss, 4.09, is close to ln(57) = 4.04. The untrained model guesses evenly, as expected. Its "before" text is random characters after the prompt.
Both losses drop fast, then flatten near 0.31. The train and check numbers stay together, so the model learns patterns that carry over to new notes. The floor comes from the random digits no model can predict.
The "after" text reads like real notes. Look closely at the first one: "Sales order 16161, customer C-1398:" isn't one of the templates in make_note. The model mixed two openings it had seen into a new one.
The last note stops mid-word because the sample is capped at 240 characters.
Material M-4805, plant 2010, exception: stock below safety stock. Expedite with the supplier.
Payment block on invoice 513371 (PO 4503891): price differs from the purchase order by 21,117.00 EUR. AP clerk to park the invoice.
Material M-4169, plant 1010, exception: reschedule out. Expedite with the supplier.
Material M-4769, plant 2010, exception: stock below safety stock.
MRP exception for material M-4112 at plant 1020: opening date in the past. Convert the planned order today.
Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 1460 PC. AP clerk to park the invoice.
Invoice 511680 for PO 4518851 held: invoice quantity 210 PC, goods receipt 110 PC.
Invoice 517902 for PO 4536921 held: invoice quantity 200 PC, goods receipt 110 PC. Buyer to confirm with supplier V-2707.
Check them against the hidden rule from Step 3. In the training data, the goods receipt is always 10 to 40 pieces below the invoice. All three notes break it: 1,460 against 230, and 110 against 210 and 200. The sentences are perfect; the numbers are invented. The model learned what a three-way match note looks like, not the rule behind it. That is a hallucination, in miniature.
If the prompt uses a character the model never saw, such as é, the script stops with a clear message. A character-level model can only write the 57 characters it knows.
Generated 200 notes; 0 are word-for-word copies of a training note (0%).
The training part of sap_notes.txt holds 2,699 different notes.
The script writes 200 notes and looks each one up in the training part of the file. With 3,000 notes, a model this small learns the patterns but copies none of them. Remember this number; the exercise changes it.
In our run, the check loss reached 1.72 after 2,000 steps, in about 3.5 minutes. The model writes play-shaped text with line breaks, capitals and many invented words:
ROMEO:
Why moves had selfess Hate exal's,
Nake to the Moccius your namerer? hear this being suigness
Bustints, your wake a whats oard; and not my father?
Compare 1.72 with the 0.31 on the SAP notes. Real language is far less predictable than templates. A model needs much more size and data to write it well, which is the scale lesson of this unit in one number.
Use this only if your timing_notes.md says the medium size is too slow on your laptop. On our two-core test machine, 2,000 medium steps would take hours.
Open Google Colab and switch on a GPU as in Set up for Unit 4, Step 6, items 1 to 3.
In the first code cell, type %%writefile transformer_block.py on the first line, paste the whole of your transformer_block.py below it, and run the cell.
Click + Code, type %%writefile tiny_gpt.py on the first line, paste the whole of tiny_gpt.py below it, and run the cell.
What success looks like: the second line of the training output ends with on cuda, and the losses fall as in Step 4. Colab deletes its machine when you leave. To keep the model, download tiny_gpt_medium.pt from the Files pane on the left before you close the tab.
cd ..
git add unit04/tiny_gpt.py
git commit -m "Train a tiny GPT on made-up SAP notes"
Don't commit tiny_gpt.pt or the text files: they are rebuilt by the script in minutes, and model files don't belong in a code history. If you use a .gitignore file, add the line *.pt to it.
The generative AI hub gives access to pretrained LLMs from several providers. SAP's AI Core guide says it is available only in the extended service plan. SAP's own models follow the same pattern. SAP's Q4 2025 release highlights describe SAP-ABAP-1, trained on more than 250 million lines of ABAP code, 30 million lines of CDS code and technical documentation, and available on the generative AI hub. They describe SAP-RPT-1 as pretrained, to spare customers costly training. SAP doesn't publish the training recipes behind these models, so don't assume they match this topic's code. The idea is the same: an expensive training run once, then many cheap calls. Check the generative AI hub's current model list before planning around a specific model.
For your own small models, SAP AI Core runs training as a pipeline. SAP Learning describes the flow: put the training data in an object store, create a configuration with the parameters for one run, start an execution, and collect the trained model as an artifact that SAP AI Core registers. A resource plan picks the hardware. The AI Core guide of September 2026 lists plans from Starter (one CPU core) up to Train-L (an "advanced GPU"). The guide also says SAP AI Core integrates with Argo Workflow, an open-source workflow engine.
Your code goes into a Docker image, and an Argo WorkflowTemplate tells SAP AI Core how to run it. This is how tiny_gpt.py could look as a training pipeline:
# Sketch: tiny_gpt.py as an SAP AI Core training pipeline (not tested)
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata:
name: tiny-gpt-train
annotations:
scenarios.ai.sap.com/name: "Tiny GPT (course)"
scenarios.ai.sap.com/description: "Character-level GPT on made-up SAP notes"
executables.ai.sap.com/name: "tiny-gpt-train"
executables.ai.sap.com/description: "Trains the Unit 4 tiny GPT"
labels:
scenarios.ai.sap.com/id: "tiny-gpt"
ai.sap.com/version: "1.0"
spec:
imagePullSecrets:
- name: <your-registry-secret>
entrypoint: train
templates:
- name: train
container:
image: docker.io/<YOUR_DOCKER_USERNAME>/tiny-gpt:01
command: ["/bin/sh", "-c"]
args:
- "python tiny_gpt.py make-data && python tiny_gpt.py train"
A real pipeline would also declare the checkpoint as an output artifact so it lands in the object store, and choose a resource plan. SAP documents both; Unit 10 walks through them.
Memorization is a data protection issue. Carlini and colleagues extracted names, phone numbers and email addresses from GPT-2, some from a single document, and found larger models more vulnerable. The exercise reproduces this in miniature. Keep personal and confidential data out of training sets unless you have a tested plan for it. Unit 11 covers data security and PII.
Authorizations don't reach inside weights. SAP authorizations decide who may see a sales order. Once order details are in a model's weights, any user of the model might get them back. Give data to a model at request time, filtered by the user's authorizations, instead. Unit 7 covers this.
Generated numbers need verification. The tiny GPT broke the quantity rule while writing fluent notes. Check every generated amount, quantity and document number against the system of record before anyone acts on it.
Judge on held-out data. Training loss says little. Always keep a check set the model never trains on, and evaluate on it. Unit 8 builds an evaluation harness.
Model files are code-adjacent. Load checkpoints with weights_only=True, as PyTorch recommends, and only from sources you trust.
Training cost scales fast. The Chinchilla study says data should grow with model size. Budget hardware, time and data together, and compare with the cost of calling a pretrained model.
Clean core still applies. Training jobs run beside S/4HANA, in SAP AI Core or elsewhere, never inside it. Read SAP data through released APIs, as in Calling your first SAP API.
The tiny GPT copied nothing when it trained on 3,000 notes. Now give it only 100 notes, train it too long, and watch it start copying. Then find the fix. Your notes from this exercise feed Unit 11, where data security and PII come back.
flowchart LR
F[100 notes] --> L[Train 2,000 steps]
F --> S[Train 400 steps]
L --> C1[copies]
S --> C2[copies]
C1 --> N[memorization_notes.md]
C2 --> N
Before you start: finish Build it yourself above. You need tiny_gpt.py and transformer_block.py in unit04, and the terminal in unit04 with .venv on. Each training run takes 1 to 4 minutes on a laptop CPU.
What success looks like (your numbers may differ a little):
step 0 train loss 4.09 check loss 4.09 [1 s]
step 400 train loss 0.34 check loss 0.43 [43 s]
step 800 train loss 0.19 check loss 0.52 [83 s]
step 1200 train loss 0.14 check loss 0.55 [123 s]
step 1600 train loss 0.13 check loss 0.63 [163 s]
step 2000 train loss 0.13 check loss 0.62 [204 s]
Training loss falls below the 0.31 floor of Step 4, while check loss climbs. That is overfitting: the model has learned the random digits of these exact 90 training notes.
Count the copies:
python tiny_gpt.py copies --model few.pt
What success looks like:
Generated 200 notes; 85 are word-for-word copies of a training note (42%).
The training part of few_notes.txt holds 91 different notes.
copied: Delivery block on sales order 10693, customer C-1281: credit limit exceeded by 33,136.50 EUR. Sales rep to call the customer.
copied: Delivery block on sales order 11727, customer C-1369: price below minimum. Check with the credit manager today.
Open few_notes.txt and search for 10693 with Edit > Find. The note appears once. The model repeats a made-up customer's number and exact open amount, from a single training line.
Now stop training early. Train on the same 100 notes for 400 steps only:
In our run, the check loss ended at 0.44 and the model copied 0 of 200 notes. Stopping when the check loss stops improving is called early stopping. It doesn't make memorization impossible, but it removes the worst of it here.
In VS Code, create unit04/memorization_notes.md with this table, filled in with your numbers:
# Memorization in my tiny GPT
| Notes | Steps | Final train loss | Final check loss | Copies out of 200 |
| --- | --- | --- | --- | --- |
| 3,000 | 2,000 | | | |
| 100 | 2,000 | | | |
| 100 | 400 | | | |
Below the table, write three sentences: why the 100-note model copied, what early stopping changed, and what this means for training a model on real customer data.
Save your work:
cd ..
git add unit04/memorization_notes.md
git commit -m "Measure memorization in a tiny GPT"
If the 100-note model copies far fewer notes on your computer, run step 2 again with --seed 1 and note the seed in your file. Keep tiny_gpt.pt from Step 4 of Build it yourself: the next topic, How LLMs generate text, uses it to explore temperature and other ways of choosing the next token.
Done when:tiny_gpt.py generate writes readable notes from your tiny_gpt.pt, memorization_notes.md holds the three-row table with your numbers and your three sentences, and both tiny_gpt.py and memorization_notes.md are committed in Git.
Pick one answer for each question. The explanation appears after you choose.
1Why can one 64-character window give the model 64 lessons in a single forward pass?
Answer: C. The targets are the inputs shifted by one, and the causal mask stops each position from seeing later characters. So every position makes a fair guess at its own next character, and the loss averages all of them at once.
2Your untrained model with a 57-character vocabulary reports a first loss of 9.5. What should you conclude?
Answer: B. An untrained model spreads its probability about evenly, so its loss should be close to ln(vocabulary size). A first loss far above that means the model starts out confidently wrong, which points to a bug in the setup, not a training problem.
3On the made-up notes, both losses flatten near 0.31 and never reach zero. Why?
Answer: D. Order numbers, customer numbers and amounts are random digits, so some guesses must stay uncertain. That sets a floor. A training loss well below it, as in the exercise, means the model has memorized specific numbers.
4What does generate do before each guess, and why?
Answer: A. The tiny GPT learned position embeddings for 64 slots, so a longer input can't be processed. nanoGPT crops the same way. Anything earlier than the window is invisible to the model, the same limit large models have at a larger size.
5The model writes an invoice note with a goods receipt of 1,460 against an invoice of 230. What explains it best?
Answer: C. Every training note had a receipt slightly below the invoice, and this note breaks that rule while looking perfect. The model predicts plausible characters, not facts. That is why generated numbers must be checked against the system of record.
6Why does the script load checkpoints with weights_only=True?
Answer: D. PyTorch's documentation explains that weights_only=True, the default since 2.6, restricts the unpickler to what state dicts need. A normal pickle file can run code when loaded, so an untrusted model file is a security risk otherwise.
7A team trained a model on 500 real service notes for 20,000 steps. Training loss is tiny and check loss has risen for hours. What do you do?
Answer: B. Falling training loss with rising check loss is overfitting. In the exercise, the same pattern came with 42% word-for-word copies, while stopping early cut copies to zero. With real notes, those copies could contain customer data.
8Your company wants a model that drafts replies about blocked sales orders. Which approach fits SAP's offering best?
Answer: D. Pretraining needs vast data and compute, and puts customer data into the weights where authorizations can't reach. The generative AI hub, in the extended plan, gives access to pretrained models, and facts filtered by the user's authorizations can be given with each request.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
nanoGPT README (karpathy/nanoGPT on GitHub)— character-level Shakespeare quick start; baby GPT with context 256, 384 channels, 6 layers, 6 heads reaches best validation loss 1.4697 in about 3 minutes on one A100; CPU settings with 4 layers, width 128, context 64 reach about 1.88 in about 3 minutes; nanoGPT points to its newer cousin nanochat
Serialization semantics (PyTorch main documentation)— since 2.6, torch.load uses weights_only=True by default; it restricts unpickling to what state dicts need; saving the state dict rather than the whole module is recommended
SAP AI Core (SAP Help Portal PDF, version of 2026-09-04)— executes pipelines as batch jobs, for example to train models; integrates with Argo Workflow and KServe; free, standard and extended plans; generative AI hub only in the extended plan; resource plans from Starter (1 CPU core) to Train-L (advanced GPU)
Build a House Price Predictor with SAP AI Core (SAP Developers tutorial)— Argo WorkflowTemplate with scenarios.ai.sap.com and executables.ai.sap.com annotations, scenarios.ai.sap.com/id and ai.sap.com/version labels; Docker image built and pushed to a registry; standard plan for narrow AI, extended plan for the generative AI hub
SAP Business AI: Release Highlights Q4 2025 (SAP News, 14 January 2026)— SAP-ABAP-1 trained on more than 250 million lines of ABAP code, 30 million lines of CDS code and technical documentation, available on the generative AI hub; SAP-RPT-1 comes pretrained in small and large versions; both can be tried in the generative AI hub trial