How LLMs generate text, and what that means for enterprise use
See how a language model picks each next token, what temperature and top-p change, and why that shapes cost, speed, consistency and checks in SAP processes.
A large language model (LLM) writes one small piece of text at a time. That piece is a token: a word, part of a word or a punctuation mark. For each next token, the model gives every token it knows a probability. Then a separate rule, the decoding setting, picks one. The chosen token is added to the text, and the model goes again.
That simple loop explains most of what enterprises see when they use LLMs:
The same question can get different answers. Most settings draw the next token at random, weighted by probability. Even the "always pick the top token" setting isn't fully repeatable on large hosted models.
Settings change the style, not the knowledge. A setting called temperature makes the output safer or more varied. No setting makes the model know something it doesn't.
Cost and wait time grow with tokens. Every token in the question and every token in the answer is processed and billed. Answers arrive one token at a time.
In the previous topic, learners trained a tiny GPT on made-up SAP notes. In this one they turn its decoding knobs and measure what changes.
Consistency is a design choice, not a model feature. Picture an AI assistant that suggests why a sales order is blocked. One clerk asks and gets "credit limit exceeded". A colleague asks the same question a minute later and gets "overdue items". Both answers came from the same model. Auditors and users will ask why. A team can lower the randomness, but it can't remove it. Researchers at Thinking Machines Lab sent one prompt 1,000 times to a large open model at temperature 0, the "least random" setting. They got 80 different answers. The cause was the serving system, not the model. So processes that need the same answer every time need more than a setting: fixed answer formats, checks against SAP data, and a log of what was asked and answered.
Lower randomness doesn't mean more correct. In this topic's lab, the tiny GPT broke a simple business rule at every temperature it was tested at. Every training note had a goods receipt slightly below the invoice quantity. The model wrote notes that broke that rule whether it was set to cautious or creative. Turning randomness down made notes more uniform, not more true. In a three-way match process, the check against purchase order and goods receipt data must come from the system, not from the model.
Tokens are the unit of cost and speed. SAP's AI Core guide says use of LLMs in the generative AI hub is metered in input tokens (the prompt) and output tokens (the answer), converted at rates that vary by model. A prompt stuffed with ten pages of order history costs more on every call. A model that writes a long answer takes longer, because each token is generated after the one before. Ask for short, structured answers where a process needs only a decision.
A length limit can cut answers off. Every call sets a maximum number of output tokens. If the answer is longer, it stops mid-sentence, or mid-record. A program that expected a complete result then breaks or, worse, carries on with half of one.
SAP's main route to LLMs is the generative AI hub in SAP AI Core. As of October 2026, it works like this for decoding:
You choose the model and its settings per request. SAP Learning's course on the generative AI hub shows a model set up with a maximum output length and a temperature of 0.2. It notes that "lower temperature means more deterministic outputs". SAP's Python SDK reference shows the same idea in its newer orchestration API.
The orchestration service wraps the call. Templates, filters and output formats sit around the model call. Unit 5 covers it in depth.
Usage is metered in tokens. SAP's AI Core guide describes input and output tokens converted into capacity units, with conversion rates per model. Output tokens cost more than input tokens in the guide's example.
Streaming is available for selected models. The orchestration service can send the answer piece by piece as it is generated, so users see text sooner.
SAP doesn't change how the underlying models generate text. The loop in this topic is the same whether the model is called directly or through SAP. What SAP adds is a governed path: one place for model access, settings, filters and metering.
Use this as a starting point, then test with your own data. Not every model accepts every setting. Anthropic's documentation, for example, says its Claude models from version 4.7 accept only the default temperature.
Task
Example in SAP
Randomness
Why
Pick from fixed options
Classify a block reason, choose a three-way match exception category
As low as the model allows
You want the most likely option every time; check it against SAP data anyway
Extract fields
Pull PO number, quantity and amount from a supplier email
As low as the model allows, plus a fixed output format
Variety adds nothing; a format check catches broken values
Summarize a record
Summarize an MRP exception for a planner
Low to moderate
Readable but stable; facts come from the record you give it
Draft text a person edits
A first draft of a customer email about a blocked order
Moderate
Some variety helps; a person reviews before sending
Brainstorm
Ideas for reducing late deliveries
Higher
Variety is the point; nothing goes straight into a process
"Temperature 0 makes the model deterministic." It makes the choice greedy in theory. On hosted models, the serving system can still give different answers to the same request.
"Lower temperature makes answers more accurate." It makes them more uniform. A model that lacks a fact or rule will state the wrong answer more consistently.
"The model plans the whole answer, then types it out." It chooses one token at a time, each based on the text so far.
"Tokens are words." A token is often part of a word. A common rule of thumb for English is about four characters per token, but it varies by language and tokenizer.
"A confident-sounding answer comes from a confident model." The tone is generated like any other text. In the lab, a note that broke the business rule scored almost the same probability as one that followed it.
"Longer answers are better value." Output tokens cost money and time. Ask for the length the process needs.
Pick one answer for each question. The explanation appears after you choose.
1What actually happens when an LLM writes an answer?
Answer: B. An LLM gives every token it knows a probability, a decoding rule picks one, and the loop repeats. This is why answers vary, why they arrive one piece at a time, and why cost grows with length.
2Two clerks ask the same assistant why the same sales order is blocked and get different reasons. What is the most likely cause?
Answer: C. Most decoding settings draw tokens at random, weighted by probability, so identical requests can differ. Even temperature 0 isn't fully repeatable on hosted models. Processes that need consistency need fixed formats, checks and logs.
3A team lowers the temperature to stop wrong quantities appearing in three-way match notes. What should you expect?
Answer: D. Temperature changes how the model picks among the options it has learned, not what it knows. In the lab, the tiny GPT broke the quantity rule at every temperature tested. Checks against purchase order and goods receipt data are what catch wrong values.
4Your team's design depends on setting temperature to 0. What risk should you raise?
Answer: A. Anthropic's documentation says Claude models from version 4.7 accept only the default temperature. A design that relies on one knob can break when the model changes. Consistency should come from formats, checks and logging as well.
5How does the generative AI hub meter LLM use, according to SAP's AI Core guide?
Answer: C. SAP's guide says LLM use is metered in input tokens and output tokens, converted into capacity units at per-model rates. Long prompts and long answers both add cost on every call.
6An extraction job returns half a JSON record now and then. What is the first thing to check?
Answer: D. Every call has a maximum number of output tokens. An answer longer than that stops mid-record. The program should check why each answer stopped and handle a cut-off answer as a failure.
7Which task is the best fit for a higher temperature setting?
Answer: C. Higher randomness adds variety, which helps when ideas are the goal and nothing flows straight into a process. Classification and extraction need the most likely answer, checked against SAP data.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: the model proposes, the decoder decides
Split text generation into two parts:
The model looks at the text so far and returns a score, called a logit, for every token in its vocabulary. Softmax turns the scores into probabilities. This part you built in Build a tiny GPT.
The decoder is a small rule outside the network. It takes those probabilities and chooses one token. Greedy, temperature, top-k and top-p are all decoder rules. None of them changes a single weight.
Then the chosen token is appended and the loop repeats. The Holtzman paper on nucleus sampling makes the point sharply: the decoding strategy alone can change text quality a great deal, with exactly the same model.
Keep this split in mind and the enterprise questions sort themselves. Knowledge, rules and facts live in the model and in the context you give it. Variety, length and stopping live in the decoder. Cost and speed come from how many times the loop runs.
flowchart LR
T[Text so far] --> M[Model]
M --> L[Logits for every token]
L --> D[Decoder rule]
D --> C[One chosen token]
C --> S{Stop?}
S -->|no| T
S -->|end token or length limit| E[Answer]
The loop stops for one of two reasons. Either the model chooses a token that means "end", or the answer reaches the length limit. In the tiny GPT, the end token is a new line. Hosted APIs report the reason too. Anthropic's API, for example, returns a stop_reason such as end_turn for a natural end or max_tokens for a cut-off.
Greedy decoding takes the most likely token every time. It is the default in Hugging Face Transformers, whose documentation recommends it for short outputs where creativity doesn't matter. It also warns that greedy output starts to repeat itself on longer texts.
Greedy has a second weakness. The Hugging Face blog on decoding points out that it can miss a likely sequence hidden behind one unlikely token. It only ever looks one step ahead.
For a business user, greedy has a third, quieter problem: it always gives the same answer to the same prompt, even when several answers are valid. In the lab below, the tiny GPT is asked to complete "Sales order 14212 for customer C-1032 blocked: ". Four reasons appeared about equally often in its training notes. Greedy decoding always picks the one the model rates highest, and never writes "missing export license".
Sampling draws the next token at random, weighted by the probabilities. A token with 30% probability is chosen about 30% of the time.
Temperature reshapes the probabilities before the draw. The decoder divides every logit by the temperature, then applies softmax:
A temperature below 1 makes big scores relatively bigger. Likely tokens get more likely.
A temperature above 1 flattens the distribution. Unlikely tokens get a real chance.
As the temperature approaches 0, sampling becomes greedy decoding, as the Hugging Face blog notes.
The tiny GPT shows it directly. After "Sales order 14212 for customer C-1032 blocked: ", the probability of the most likely first letter is 49% at temperature 0.5, 37% at 1.0 and 31% at 1.5.
Raising the temperature has a cost: very unlikely tokens, which are often garbage, start to appear. Two filters remove the long tail before the draw:
Top-k keeps only the k most likely tokens. The Hugging Face blog notes that GPT-2 used it. Its weakness: k is fixed, whether the model is sure (one good option) or unsure (twenty good options).
Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least p. When the model is sure, the set is small. When it is unsure, the set grows. Holtzman and colleagues proposed it to cut off the "less reliable tail" while keeping variety.
In the lab, top-p 0.9 keeps 3 of 57 characters after "blocked: ", where the model is fairly sure, and 9 of 57 at the start of an invoice number, where any digit is plausible.
Beam search keeps several candidate texts at each step and returns the one with the highest overall probability. The Hugging Face documentation suggests it for input-grounded tasks such as speech recognition. Holtzman and colleagues found that maximizing probability makes open-ended text "bland and strangely repetitive". This topic focuses on sampling, the approach their paper recommends for open-ended text, and doesn't build beam search.
On your laptop, greedy decoding with the same model gives the same text every time. Hosted models are different. In September 2025, Horace He at Thinking Machines Lab sent the same prompt 1,000 times at temperature 0 to the open Qwen3-235B-A22B-Instruct model. The result was 80 different completions.
The cause they identified: a busy server groups many users' requests into one batch, and the batch size changes with load. The GPU routines they examined don't return bit-identical numbers for different batch sizes. Tiny numeric differences flip a close choice between two tokens, and the texts drift apart from there. With "batch-invariant" routines, all 1,000 completions were identical, at a speed cost.
The practical lesson: don't promise "same input, same output" from a hosted LLM. Design for variation, and log what each call returned.
The tiny GPT uses one token per character. Production LLMs use word-piece tokenizers. OpenAI's tiktoken README says a token is about 4 bytes of text on average, which for plain English is about four characters. Other languages, long numbers and codes can split into tokens differently, so count with the tokenizer of the model you use.
Two things follow:
Cost. Providers count input tokens and output tokens. SAP's AI Core guide meters generative AI hub use that way, converting each into capacity units at per-model rates. In the guide's example, the output rate per 1,000 tokens is almost three times the input rate.
Speed. All input tokens can be processed together in one pass. Output tokens come one loop at a time. So a long answer is slower than a long question.
Look at the tiny GPT's generate loop: every new character runs the model over the whole window again. For the earlier positions, that repeats identical work. The Hugging Face documentation on caching says this directly: each prediction depends on the previous tokens, so the model performs the same computations each time.
Production systems keep a KV cache. Attention's keys and values for earlier tokens are stored and reused, so each step computes only the new token. The cost moves to memory: Hugging Face notes the cache can become a bottleneck for long contexts. This is one reason long prompts are expensive to serve, and why providers price and limit context length. The tiny GPT skips the cache to stay short.
Because tokens appear one at a time, an API can send each piece as soon as it exists. That is streaming. The user sees the first words quickly, even if the full answer takes several seconds. SAP's AI Core guide says the orchestration service supports streaming for selected models.
Streaming changes nothing about cost or content. It changes how the wait feels. For a background job that parses the answer, streaming adds complexity for no benefit. For a chat window, it is usually worth it.
Always check why the answer stopped. A cut-off answer that looks complete is a classic production bug.
#Log-probabilities: a weak signal, not a truth test
The probability the model gave each chosen token is a rough measure of how expected it was. Some APIs can return these as log-probabilities. They are tempting as a "confidence score". The lab shows why to be careful: a random order number gets low probability because any digit fits, not because it is wrong. And a note that breaks a business rule can score as high as one that follows it. Use log-probabilities to flag candidates for review, never as proof.
#Build it yourself: turn the decoding knobs on your tiny GPT
You will load the tiny GPT you trained in the previous topic and explore four things: its top guesses at several temperatures, one note written character by character, six decoding settings measured side by side, and a probability score for any note you type. Everything runs offline.
flowchart LR
K[tiny_gpt.pt] --> P[peek: top guesses]
K --> S[stream: one note, live]
K --> C[compare: six settings, measured]
K --> Q[score: how expected is a note]
C --> N[decoding_notes.md]
In VS Code's file list, right-click unit04, choose New File and name it generate_lab.py.
Paste the code below and save with File > Save.
"""Unit 4: how LLMs generate text. Explore decoding with the tiny GPT you trained.
Uses tiny_gpt.pt from the previous topic (Build a tiny GPT) and imports from tiny_gpt.py.
Everything runs offline on your laptop. No accounts, no keys.
How to run (from the unit04 folder, with the course .venv turned on):
python generate_lab.py peek # the model's top guesses for the next character
python generate_lab.py stream # write one note character by character, with timing
python generate_lab.py compare # the same model, six decoding settings, measured
python generate_lab.py score --text "..." # how likely the model finds each character of your text
Optional:
python generate_lab.py peek --prompt "Invoice 51" --temperatures 0.3 1 2
python generate_lab.py stream --temperature 0 --max-chars 40 # greedy, cut off by the length limit
python generate_lab.py compare --count 50 --prompt "Invoice" # faster, invoice notes only
python generate_lab.py compare --grid # a wider temperature sweep for the exercise
"""
import argparse
import math
import re
import sys
import time
import torch
import torch.nn.functional as F
from tiny_gpt import load, pick_device
# ---------------------------------------------------------------- 1. the formats the notes were made from
# These patterns mirror make_note() in tiny_gpt.py. A note "matches a format" if one pattern fits it exactly.
AMT = r"\d{1,2},\d{3}\.(?:00|50) (?:EUR|USD)"
PLANT = r"(?:1010|1020|1710|2010)"
FORMATS = [
r"(?:Sales order \d{5} for customer C-1\d{3} blocked: |Order \d{5} \(C-1\d{3}\) on hold, "
r"|Delivery block on sales order \d{5}, customer C-1\d{3}: )"
rf"(?:credit limit exceeded by {AMT}|overdue items on account|missing export license|price below minimum)\. "
r"(?:Credit team to review before release\.|Sales rep to call the customer\.|Release after payment arrives\."
r"|Check with the credit manager today\.)",
r"(?:Invoice 51\d{4} for PO 45\d{5} held: |Three-way match failed on PO 45\d{5}, "
r"|Payment block on invoice 51\d{4} \(PO 45\d{5}\): )"
rf"(?:invoice quantity \d+ PC, goods receipt \d+ PC|price differs from the purchase order by {AMT}"
r"|no goods receipt posted yet)\. "
r"(?:Buyer to confirm with supplier V-2\d{3}\.|AP clerk to park the invoice\.|Warehouse to check the receipt\."
r"|Ask V-2\d{3} for a credit memo\.)",
rf"(?:MRP exception for material M-4\d{{3}} at plant {PLANT}: |Material M-4\d{{3}}, plant {PLANT}, exception: )"
r"(?:start date in the past|stock below safety stock|reschedule in|reschedule out|opening date in the past)\. "
rf"(?:Planner to reschedule order 1\d{{6}}\.|Expedite with the supplier\.|Check capacity at plant {PLANT}\."
r"|Convert the planned order today\.)",
]
QTY = re.compile(r"invoice quantity (\d+) PC, goods receipt (\d+) PC")
def matches_format(note: str) -> bool:
return any(re.fullmatch(f, note) for f in FORMATS)
def breaks_rule(note: str):
"""None if the note has no quantities; True if the receipt isn't 10 to 40 pieces below the invoice."""
m = QTY.search(note)
if not m:
return None
gap = int(m.group(1)) - int(m.group(2))
return not (10 <= gap <= 40)
# ---------------------------------------------------------------- 2. one decoding step: scores -> one token
def next_token_probs(model, idx, context: int) -> torch.Tensor:
"""Probabilities for the next character, before any decoding setting is applied."""
logits = model(idx[:, -context:])[0, -1, :]
return F.softmax(logits, dim=-1)
def choose(logits: torch.Tensor, temperature: float, top_k: int, top_p: float) -> int:
"""Turn raw scores into one chosen token id. temperature 0 means greedy: always the top score."""
if temperature == 0:
return int(torch.argmax(logits))
logits = logits / temperature
if top_k > 0: # keep only the k highest scores
kth = torch.topk(logits, min(top_k, logits.numel())).values[-1]
logits = logits.masked_fill(logits < kth, float("-inf"))
probs = F.softmax(logits, dim=-1)
if top_p < 1.0: # keep the smallest set whose probabilities reach top_p
sorted_p, order = torch.sort(probs, descending=True)
keep = torch.cumsum(sorted_p, dim=0) - sorted_p < top_p
mask = torch.zeros_like(probs, dtype=torch.bool)
mask[order[keep]] = True
probs = torch.where(mask, probs, torch.zeros_like(probs))
probs = probs / probs.sum()
return int(torch.multinomial(probs, 1))
@torch.no_grad()
def write(model, tok, cfg, prompt: str, temperature=1.0, top_k=0, top_p=1.0, max_chars=200, on_char=None):
"""Generate until a line ends ("end of note") or max_chars is reached ("length limit")."""
idx = torch.tensor([tok.encode(prompt)])
out = []
for _ in range(max_chars):
logits = model(idx[:, -cfg["context"]:])[0, -1, :]
nxt = choose(logits, temperature, top_k, top_p)
ch = tok.chars[nxt]
if ch == "\n":
return "".join(out), "end of note"
out.append(ch)
if on_char:
on_char(ch)
idx = torch.cat([idx, torch.tensor([[nxt]])], dim=1)
return "".join(out), "length limit"
def shown(ch: str) -> str:
return {"\n": "\\n (end of note)", " ": "' ' (space)"}.get(ch, ch)
def check_prompt(tok, prompt: str) -> None:
unknown = sorted(set(prompt) - set(tok.chars))
if unknown:
raise SystemExit(f"The prompt uses characters the model never saw: {unknown}")
# ---------------------------------------------------------------- 3. commands
def peek(args) -> None:
model, tok, cfg, _ = load(args.model, pick_device("cpu"))
model.eval()
check_prompt(tok, args.prompt)
idx = torch.tensor([tok.encode(args.prompt)])
with torch.no_grad():
logits = model(idx[:, -cfg["context"]:])[0, -1, :]
print(f"Prompt: {args.prompt!r}. Top {args.top} guesses for the next character:\n")
temps = args.temperatures
print(f" {'char':<18}" + "".join(f"{'temp ' + format(t, 'g'):>11}" for t in temps))
base = F.softmax(logits, dim=-1)
order = torch.argsort(base, descending=True)[: args.top]
tables = [F.softmax(logits / t, dim=-1) for t in temps]
for i in order.tolist():
print(f" {shown(tok.chars[i]):<18}" + "".join(f"{p[i].item():>11.1%}" for p in tables))
for t, p in zip(temps, tables):
s = torch.sort(p, descending=True).values
nucleus = int((torch.cumsum(s, 0) - s < 0.9).sum())
print(f"\n temp {t:g}: the top choice has {s[0].item():.0%}; "
f"top-p 0.9 keeps {nucleus} of {len(tok.chars)} characters", end="")
print()
def stream(args) -> None:
model, tok, cfg, _ = load(args.model, pick_device("cpu"))
model.eval()
check_prompt(tok, args.prompt)
torch.manual_seed(args.seed)
start = time.time()
first = []
def on_char(ch):
if not first:
first.append(time.time() - start)
sys.stdout.write(ch)
sys.stdout.flush()
time.sleep(args.delay) # slow it down so you can watch it arrive
sys.stdout.write(args.prompt)
note, reason = write(model, tok, cfg, args.prompt, args.temperature, args.top_k, args.top_p,
args.max_chars, on_char)
total = time.time() - start - args.delay * len(note)
print(f"\n\nStopped because: {reason}. Prompt {len(args.prompt)} characters in, "
f"{len(note)} characters out.")
if first:
print(f"First character after {first[0] * 1000:.0f} ms; "
f"{len(note) / max(total, 1e-9):.0f} characters per second of model time.")
def run_setting(model, tok, cfg, prompt, n, seed, temperature, top_k, top_p, seen):
torch.manual_seed(seed)
notes = [write(model, tok, cfg, prompt, temperature, top_k, top_p)[0] for _ in range(n)]
full = [(prompt.strip("\n") + n_) for n_ in notes]
fmt = sum(matches_format(x) for x in full)
qty = [r for r in (breaks_rule(x) for x in full) if r is not None]
return dict(fmt=fmt, distinct=len(set(full)), copies=sum(x in seen for x in full),
qty=len(qty), broke=sum(qty), example=full[0])
def compare(args) -> None:
model, tok, cfg, ckpt = load(args.model, pick_device("cpu"))
model.eval()
check_prompt(tok, args.prompt)
text = open(ckpt["data"], encoding="utf-8").read()
seen = set(text[: int(0.9 * len(text))].splitlines())
if args.grid:
settings = [(f"temperature {t:g}", t, 0, 1.0) for t in (0, 0.3, 0.5, 0.7, 1.0, 1.3, 1.6, 2.0)]
else:
settings = [("greedy (temperature 0)", 0, 0, 1.0), ("temperature 0.5", 0.5, 0, 1.0),
("temperature 1.0", 1.0, 0, 1.0), ("temperature 1.5", 1.5, 0, 1.0),
("temp 1.5 + top-k 5", 1.5, 5, 1.0), ("temp 1.5 + top-p 0.9", 1.5, 0, 0.9)]
n = args.count
print(f"{n} notes per setting, prompt {args.prompt!r}, seed {args.seed}\n")
print(f" {'setting':<24}{'format ok':>10}{'distinct':>10}{'copies':>8}{'qty rule broken':>17}")
rows = []
for name, t, k, p in settings:
r = run_setting(model, tok, cfg, args.prompt, n, args.seed, t, k, p, seen)
rows.append((name, r))
print(f" {name:<24}{r['fmt'] / n:>10.0%}{r['distinct']:>10}{r['copies']:>8}"
f"{str(r['broke']) + ' of ' + str(r['qty']):>17}", flush=True)
print("\nOne example per setting:")
for name, r in rows:
print(f" {name:<24} {r['example'][:110]}")
@torch.no_grad()
def score(args) -> None:
model, tok, cfg, _ = load(args.model, pick_device("cpu"))
model.eval()
text = args.text
check_prompt(tok, text)
ids = tok.encode("\n" + text)
total, rows = 0.0, []
for i in range(1, len(ids)):
p = next_token_probs(model, torch.tensor([ids[:i]]), cfg["context"])[ids[i]].item()
total += math.log(p)
rows.append((tok.chars[ids[i]], p))
print(f"Text: {text}\n")
print(f"Average log-probability per character: {total / len(rows):.2f} "
f"(0 would mean the model was certain of every character)")
low = sorted(range(len(rows)), key=lambda j: rows[j][1])[: args.lowest]
print(f"The {args.lowest} least expected characters:")
for j in sorted(low):
before = "..." + text[max(0, j - 18):j] if j else "(start) "
print(f" {before}[{rows[j][0]}] p = {rows[j][1]:.3f}")
def main() -> None:
p = argparse.ArgumentParser(description="Explore how a language model picks each next token.")
sub = p.add_subparsers(dest="command", required=True)
a = sub.add_parser("peek", help="show the top next-character guesses at several temperatures")
a.add_argument("--prompt", default="Sales order 14212 for customer C-1032 blocked: ")
a.add_argument("--temperatures", type=float, nargs="+", default=[0.5, 1.0, 1.5])
a.add_argument("--top", type=int, default=6)
s = sub.add_parser("stream", help="write one note character by character")
s.add_argument("--prompt", default="Invoice")
s.add_argument("--max-chars", type=int, default=200, help="the length limit, like max_tokens")
s.add_argument("--delay", type=float, default=0.03, help="seconds to pause per character, for watching")
c = sub.add_parser("compare", help="measure several decoding settings on many notes")
c.add_argument("--prompt", default="\n")
c.add_argument("--count", type=int, default=100)
c.add_argument("--grid", action="store_true", help="sweep temperature from 0 to 2 instead")
for sp in (s,):
sp.add_argument("--temperature", type=float, default=1.0, help="0 means greedy")
sp.add_argument("--top-k", type=int, default=0, help="0 means off")
sp.add_argument("--top-p", type=float, default=1.0, help="1.0 means off")
k = sub.add_parser("score", help="how likely the model finds each character of a text")
k.add_argument("--text", default="Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, "
"goods receipt 1460 PC. AP clerk to park the invoice.")
k.add_argument("--lowest", type=int, default=6)
for sp in (a, s, c, k):
sp.add_argument("--model", default="tiny_gpt.pt")
for sp in (s, c):
sp.add_argument("--seed", type=int, default=0)
args = p.parse_args()
{"peek": peek, "stream": stream, "compare": compare, "score": score}[args.command](args)
if __name__ == "__main__":
main()
Prompt: 'Sales order 14212 for customer C-1032 blocked: '. Top 6 guesses for the next character:
char temp 0.5 temp 1 temp 1.5
p 48.6% 37.3% 31.1%
o 25.4% 26.9% 25.0%
c 23.3% 25.8% 24.3%
m 2.7% 8.9% 11.9%
i 0.0% 0.3% 1.2%
r 0.0% 0.2% 0.8%
temp 0.5: the top choice has 49%; top-p 0.9 keeps 3 of 57 characters
temp 1: the top choice has 37%; top-p 0.9 keeps 3 of 57 characters
temp 1.5: the top choice has 31%; top-p 0.9 keeps 4 of 57 characters
How to read it:
The four letters are four block reasons:price below minimum, overdue items, credit limit, missing export license. In sap_notes.txt, each follows "blocked: " about a quarter of the time. Check it yourself: search the file for blocked: m.
The model is not well calibrated. It gives "m" only 9% at temperature 1. A small model learns the patterns roughly, not the exact frequencies.
Low temperature makes it worse. At 0.5, "m" drops to 2.7%. Greedy decoding would never choose it. A low setting doesn't just remove noise; it can hide valid but less likely answers.
Top-p adapts. Here three or four letters cover 90% of the probability, so top-p 0.9 keeps only those.
Now try a place where the model can't know the answer, the digits of an invoice number:
The last lines show the top choice at only 23% even at temperature 0.3, and top-p 0.9 keeping 7 to 10 characters. When every digit is plausible, no setting makes the model sure.
The note appears character by character, slowed down on purpose so you can watch it.
What success looks like:
Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 1460 PC. AP clerk to park the invoice.
Stopped because: end of note. Prompt 7 characters in, 106 characters out.
First character after 4 ms; 394 characters per second of model time.
This is the same note the previous topic showed, with its broken quantity. Same model, same seed (0), same settings: same note. On your laptop, sampling with a fixed seed is repeatable. Your timing numbers will differ.
Now use greedy decoding with a short length limit:
Invoice 516662 for PO 4577991 held: invoice qua
Stopped because: length limit. Prompt 7 characters in, 40 characters out.
The note stopped mid-word because it reached the 40-character limit, just as an API answer stops at max_tokens. The "Stopped because" line is your stop reason. Run the command without --max-chars 40 and greedy decoding finishes the note: ... invoice quantity 170 PC, goods receipt 130 PC. AP clerk to park the invoice. Run it again: you get exactly the same note, every time.
The script writes 100 notes with each of six settings and checks every note in three ways. It takes about 2 minutes.
What success looks like:
100 notes per setting, prompt '\n', seed 0
setting format ok distinct copies qty rule broken
greedy (temperature 0) 100% 1 0 0 of 0
temperature 0.5 86% 100 0 5 of 5
temperature 1.0 64% 100 0 8 of 10
temperature 1.5 4% 100 0 8 of 9
temp 1.5 + top-k 5 16% 100 0 5 of 6
temp 1.5 + top-p 0.9 50% 100 0 13 of 14
The columns:
format ok: the note fits one of the formats tiny_gpt.py made notes from, exactly: right document number lengths, valid amounts, an action that belongs to the same process.
distinct: how many of the 100 notes are different.
copies: notes copied word for word from the training part of the data.
qty rule broken: of the notes with an invoice and a goods receipt quantity, how many break the hidden rule (receipt 10 to 40 pieces below the invoice).
What it shows:
Greedy wrote the same note 100 times. Perfectly formatted and perfectly useless if you wanted 100 different drafts.
Format falls apart as temperature rises. At 1.5, only 4% fit a format. Top-k and top-p win much of it back by cutting the unlikely tail.
The business rule is broken at every setting. Even at 0.5, all 5 quantity notes break it. Decoding can't add a rule the model never learned.
To see what "format not ok" looks like, these are real failures from the model at temperature 1.0:
Delivery block on sales order 10942, customer C-1465: overdue items on account. Convert the planned order today.
Three-way match failed on PO 4567696, price differs from the purchase order by 30,91.00 EUR. Ask V-2301 for a credit memo.
Delivery block on sales order 18121, customer C-1931: credit limit exceeded by 368,,824.00 USD. Check with the credit manager today.
A planned-order action on a sales order, and two broken amounts. Each reads fluently. A format check catches all three. Only a check against system data would catch the broken quantity rule. Unit 5 covers asking a model for a fixed output format.
Text: Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 1460 PC. AP clerk to park the invoice.
Average log-probability per character: -0.31 (0 would mean the model was certain of every character)
The 6 least expected characters:
(start) [I] p = 0.073
...Invoice 51992[1] p = 0.103
... 519921 for PO 457[0] p = 0.077
...19921 for PO 45702[8] p = 0.096
...C, goods receipt 1[4] p = 0.072
..., goods receipt 14[6] p = 0.091
Now score the same note with a receipt that follows the rule:
python generate_lab.py score --text "Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 210 PC. AP clerk to park the invoice."
The average is -0.30: almost the same. The model barely prefers the correct note. Its least expected characters are random digits, which are uncertain because any digit fits, not because they are wrong.
Finally, score the note with the wrong process action:
python generate_lab.py score --text "Delivery block on sales order 10942, customer C-1465: overdue items on account. Convert the planned order today."
The lowest score, p = 0.020, lands on the "o" of "Convert", where the note switches to a plan-to-produce action. Log-probabilities caught this error and missed the quantity one. That is the right expectation: a hint for review, not a check.
As of October 2026, decoding settings in SAP's stack are set per model call through the generative AI hub, mostly via its orchestration service. Unit 5 sets up access and covers the service in depth. Here is how this topic's knobs map to it.
The SDK's reference for the newer orchestration V2 API uses LLMModelDetails with a params dictionary, and names the length limit max_completion_tokens in its example. Note the different names for the same idea. Parameter names and accepted values depend on the model and the SDK version, so check both before you rely on a setting.
A sketch of a classification call with the V2 API:
# Sketch: classify a block reason with low randomness through the orchestration service (not tested)
from gen_ai_hub.orchestration_v2.models.message import SystemMessage, UserMessage
from gen_ai_hub.orchestration_v2.models.template import Template, PromptTemplatingModuleConfig
from gen_ai_hub.orchestration_v2.models.llm_model_details import LLMModelDetails
from gen_ai_hub.orchestration_v2.models.config import ModuleConfig, OrchestrationConfig
from gen_ai_hub.orchestration_v2.service import OrchestrationService
template = Template(template=[
SystemMessage(content="Answer with one of: CREDIT_LIMIT, OVERDUE_ITEMS, EXPORT_LICENSE, PRICE_BELOW_MIN."),
UserMessage(content="Order note: {{?note}}"),
])
llm = LLMModelDetails(name="gpt-4o", params={"max_completion_tokens": 10, "temperature": 0})
config = OrchestrationConfig(modules=ModuleConfig(
prompt_templating=PromptTemplatingModuleConfig(prompt=template, model=llm)))
result = OrchestrationService(config=config).run(
placeholder_values={"note": "Customer C-1032 has open items past due since August."})
label = result.final_result.choices[0].message.content.strip()
if label not in {"CREDIT_LIMIT", "OVERDUE_ITEMS", "EXPORT_LICENSE", "PRICE_BELOW_MIN"}:
raise ValueError(f"Unexpected label: {label!r}") # never pass an unchecked answer on
Three habits from the lab carry over: a short length limit for a short answer, the lowest randomness the model accepts for a fixed-choice task, and a check of the answer against the allowed values.
SAP's AI Core guide (September 2026) says generative AI hub use is metered in input and output tokens. These convert to "GenAI tokens" and then capacity units, at rates that vary by model. Look up the current rates for your model before estimating; don't reuse the guide's example rates as prices.
The same guide says the orchestration service supports response streaming for selected models. Use it for interactive screens; skip it for background jobs that parse the full answer.
The decoding rules are the model's. Calling a model through the generative AI hub doesn't make temperature 0 repeatable or stop a model from inventing values. The orchestration service adds templating, filtering and governed access around the same loop.
Design for variation. Assume the same input can produce a different output. Use fixed answer formats, validate them, and route failures to a person or a retry.
Log enough to explain an answer. Store the prompt or a reference to it, model name and version, decoding settings, the answer, the stop reason and token counts. Mind data protection when prompts contain personal data; Unit 11 covers this.
Check the stop reason on every call. Treat a length cut-off as a failure, not a result.
Validate generated values against SAP. Amounts, quantities and document numbers come from the system, through released APIs, not from the model's text. See Calling your first SAP API.
Budget tokens both ways. Estimate input and output tokens per call, times calls per day, at your model's rates. Long prompts are paid on every call.
Don't hard-wire one knob. Some models drop settings; Anthropic's documentation says Claude models from 4.7 accept only the default temperature. Keep decoding settings in configuration, and test when the model changes.
Use log-probabilities as a hint only. They flag surprising tokens, but random fields look surprising and wrong facts can look normal.
Authorizations apply to inputs. The model can only repeat what you send it or what it learned. Filter the context by the user's SAP authorizations before the call; Unit 7 covers this.
Clean core still applies. Generation runs outside S/4HANA, on BTP or with a provider. Read SAP data through released APIs and write back only through them.
Promising determinism. Temperature 0 on a hosted model is not a guarantee.
Turning temperature down to fix wrong facts. It changes the style of the errors, not their presence.
Using greedy decoding for drafts. You get the same draft every time and lose valid but less likely options.
Using a high temperature without top-p or top-k. The long tail brings in broken tokens.
Ignoring cut-offs. A length limit that is too tight breaks structured answers silently.
Copying settings between models. Parameter names, ranges and defaults differ by model and SDK version.
Reading log-probabilities as confidence in the truth. They measure how expected a token was, given what the model learned.
#Exercise: choose decoding settings for two SAP tasks
You will sweep temperature on invoice notes and choose settings for two tasks: a classifier that must pick one fixed exception category, and a draft writer that suggests varied notes for a person to edit. Your decoding_notes.md feeds Unit 5, where you set the same knobs on a real model.
flowchart LR
G[compare --grid<br/>prompt Invoice] --> T[Table of results]
T --> D1[Settings for a classifier]
T --> D2[Settings for a draft writer]
D1 --> N[decoding_notes.md]
D2 --> N
Before you start: finish Build it yourself above, with the terminal in unit04 and .venv on. The sweep takes about 1 to 2 minutes.
What success looks like (the first lines; examples follow below them):
50 notes per setting, prompt 'Invoice', seed 1
setting format ok distinct copies qty rule broken
temperature 0 100% 1 0 0 of 50
temperature 0.3 96% 50 0 28 of 35
temperature 0.5 82% 50 0 27 of 30
temperature 0.7 82% 50 0 18 of 20
temperature 1 50% 50 0 18 of 18
temperature 1.3 8% 50 0 3 of 3
temperature 1.6 2% 49 0 1 of 1
temperature 2 0% 49 0 0 of 0
Notice the first row: greedy wrote one note 50 times, and that one note happens to follow the rule. Every sampled setting breaks it most of the time.
Test whether top-p rescues a high temperature. Run:
Compare the format ok values of temperature 1.5 and temp 1.5 + top-p 0.9.
In VS Code, create unit04/decoding_notes.md and paste this template:
# Decoding settings for two SAP tasks
## My sweep (prompt "Invoice", seed 1, 50 notes each)
| Temperature | Format ok | Distinct | Qty rule broken |
| --- | --- | --- | --- |
| 0 | | | |
| 0.3 | | | |
| 0.7 | | | |
| 1 | | | |
| 1.3 | | | |
## Task 1: classify a three-way match exception into a fixed category
Settings I would use:
Why:
Check I would add after the model answers:
## Task 2: draft varied notes for an AP clerk to edit
Settings I would use:
Why:
Check I would add after the model answers:
## What no setting fixed
Fill the table from your own output. Then complete both tasks. For each, name a temperature and whether you would add top-p, in one sentence say why, and name one check, such as "the answer must be one of four category codes" or "quantities must match the goods receipt in SAP".
Under What no setting fixed, write two sentences about the quantity rule, and what that means for a real three-way match assistant.
Save your work:
cd ..
git add unit04/decoding_notes.md
git commit -m "Choose decoding settings for two SAP tasks"
If your numbers differ from the sample, that's fine. Record yours, and note your PyTorch version (python -c "import torch; print(torch.__version__)") at the top of the file.
Done when:decoding_notes.md holds your sweep table, settings and a check for each of the two tasks, and two sentences on what no setting fixed, and both generate_lab.py and decoding_notes.md are committed in Git.
Pick one answer for each question. The explanation appears after you choose.
1Which part of text generation do temperature, top-k and top-p change?
Answer: B. The model proposes probabilities; the decoder chooses. Temperature, top-k and top-p are decoder rules applied after the network has run. They change which token is picked, never what the model learned.
2In choose, what does dividing the logits by a temperature of 0.5 do before softmax?
Answer: C. Dividing by a number below 1 makes large logits relatively larger, so softmax concentrates probability on the top choices. In the peek output, the top letter rose from 37% at temperature 1 to 49% at 0.5. As the temperature approaches 0, sampling becomes greedy.
3Why does top-p 0.9 keep 3 characters after "blocked: " but 9 inside an invoice number?
Answer: D. Nucleus sampling adapts to the shape of the distribution. After "blocked: " a few letters cover 90%; for a random digit, probability is spread over many characters. Top-k, by contrast, keeps a fixed number.
4In the compare run, the quantity rule was broken at every sampled temperature. What does that tell a builder?
Answer: A. Decoding chooses among what the model learned and can't add a missing rule. Greedy happened to produce one note that follows it, but wrote that same note every time. In a real three-way match, quantities must be checked against purchase order and goods receipt data.
5You send the same prompt at temperature 0 to a hosted model 1,000 times and get several different answers. What is the most likely explanation?
Answer: C. Thinking Machines Lab traced exactly this to batch sizes changing with load, with routines whose results differ slightly by batch size. A close call between two tokens flips and the texts diverge. Batch-invariant routines made all 1,000 identical, at a speed cost.
6An extraction service returns JSON that sometimes fails to parse. The logs show the stop reason max_tokens on those calls. What do you do?
Answer: B. A max_tokens stop means the answer was cut off at the length limit, as in the stream --max-chars 40 run. The fix is a limit that fits the task, plus code that never passes a cut-off answer on.
7Why are output tokens slower to produce than input tokens?
Answer: D. The input can be processed together in one pass, but each new token depends on the previous one, so the loop runs once per output token. The KV cache makes each step cheaper by reusing earlier keys and values, but the steps still run in order.
8A note that breaks the quantity rule scores almost the same average log-probability as a correct one. What follows?
Answer: C. Log-probabilities measure how expected each token was, given what the model learned. Random digits look surprising and a broken rule can look normal. Use them to flag text for review, never as a truth check.
9In the SAP sketch, which habit protects the process if the model returns an unexpected answer?
Answer: C. A short limit and low temperature help, but neither guarantees a valid answer, and temperature 0 isn't repeatable on hosted models. The check against the four allowed labels is what stops a bad answer from reaching the process.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Generation strategies (Hugging Face Transformers documentation)— greedy search is the default and picks the most likely next token; it starts to repeat itself on longer outputs; sampling with do_sample=True; beam search with num_beams above 1; max_new_tokens
API usage primer for Claude (Claude Platform documentation)— on Claude 4.7 and later models, temperature is deprecated and only its default is accepted; max_tokens; stop_reason values such as end_turn and max_tokens; usage reports input_tokens and output_tokens
Cache strategies (Hugging Face Transformers documentation)— autoregressive models predict one token at a time and would repeat the same key and value computations; a KV cache stores them for reuse; the cache can become a memory bottleneck for long contexts
Orchestration Service V2 API (SAP Cloud SDK for AI, Python, reference)— gen_ai_hub.orchestration_v2 imports; LLMModelDetails with params max_completion_tokens 512 and temperature 0.7; Template with SystemMessage and UserMessage; OrchestrationService(config=...).run(placeholder_values=...); result.final_result.choices[0].message.content
SAP AI Core (SAP Help Portal PDF, version of 2026-09-04)— the generative AI hub is available only in the extended service plan; LLM use is metered in input and output tokens, converted to GenAI tokens and capacity units at per-model rates; response streaming is supported for selected models in the orchestration service