Orchestrate

How LLMs generate text, and what that means for enterprise use

See how a language model picks each next token, what temperature and top-p change, and why that shapes cost, speed, consistency and checks in SAP processes.

Updated Oct 1, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

A large language model (LLM) writes one small piece of text at a time. That piece is a token: a word, part of a word or a punctuation mark. For each next token, the model gives every token it knows a probability. Then a separate rule, the decoding setting, picks one. The chosen token is added to the text, and the model goes again.

That simple loop explains most of what enterprises see when they use LLMs:

  1. The same question can get different answers. Most settings draw the next token at random, weighted by probability. Even the "always pick the top token" setting isn't fully repeatable on large hosted models.
  2. Settings change the style, not the knowledge. A setting called temperature makes the output safer or more varied. No setting makes the model know something it doesn't.
  3. Cost and wait time grow with tokens. Every token in the question and every token in the answer is processed and billed. Answers arrive one token at a time.

In the previous topic, learners trained a tiny GPT on made-up SAP notes. In this one they turn its decoding knobs and measure what changes.

Why it matters to the business

Consistency is a design choice, not a model feature. Picture an AI assistant that suggests why a sales order is blocked. One clerk asks and gets "credit limit exceeded". A colleague asks the same question a minute later and gets "overdue items". Both answers came from the same model. Auditors and users will ask why. A team can lower the randomness, but it can't remove it. Researchers at Thinking Machines Lab sent one prompt 1,000 times to a large open model at temperature 0, the "least random" setting. They got 80 different answers. The cause was the serving system, not the model. So processes that need the same answer every time need more than a setting: fixed answer formats, checks against SAP data, and a log of what was asked and answered.

Lower randomness doesn't mean more correct. In this topic's lab, the tiny GPT broke a simple business rule at every temperature it was tested at. Every training note had a goods receipt slightly below the invoice quantity. The model wrote notes that broke that rule whether it was set to cautious or creative. Turning randomness down made notes more uniform, not more true. In a three-way match process, the check against purchase order and goods receipt data must come from the system, not from the model.

Tokens are the unit of cost and speed. SAP's AI Core guide says use of LLMs in the generative AI hub is metered in input tokens (the prompt) and output tokens (the answer), converted at rates that vary by model. A prompt stuffed with ten pages of order history costs more on every call. A model that writes a long answer takes longer, because each token is generated after the one before. Ask for short, structured answers where a process needs only a decision.

A length limit can cut answers off. Every call sets a maximum number of output tokens. If the answer is longer, it stops mid-sentence, or mid-record. A program that expected a complete result then breaks or, worse, carries on with half of one.

How SAP does it

SAP's main route to LLMs is the generative AI hub in SAP AI Core. As of October 2026, it works like this for decoding:

  • You choose the model and its settings per request. SAP Learning's course on the generative AI hub shows a model set up with a maximum output length and a temperature of 0.2. It notes that "lower temperature means more deterministic outputs". SAP's Python SDK reference shows the same idea in its newer orchestration API.
  • The orchestration service wraps the call. Templates, filters and output formats sit around the model call. Unit 5 covers it in depth.
  • Usage is metered in tokens. SAP's AI Core guide describes input and output tokens converted into capacity units, with conversion rates per model. Output tokens cost more than input tokens in the guide's example.
  • Streaming is available for selected models. The orchestration service can send the answer piece by piece as it is generated, so users see text sooner.

SAP doesn't change how the underlying models generate text. The loop in this topic is the same whether the model is called directly or through SAP. What SAP adds is a governed path: one place for model access, settings, filters and metering.

Choosing decoding settings for SAP tasks

Use this as a starting point, then test with your own data. Not every model accepts every setting. Anthropic's documentation, for example, says its Claude models from version 4.7 accept only the default temperature.

Task Example in SAP Randomness Why
Pick from fixed options Classify a block reason, choose a three-way match exception category As low as the model allows You want the most likely option every time; check it against SAP data anyway
Extract fields Pull PO number, quantity and amount from a supplier email As low as the model allows, plus a fixed output format Variety adds nothing; a format check catches broken values
Summarize a record Summarize an MRP exception for a planner Low to moderate Readable but stable; facts come from the record you give it
Draft text a person edits A first draft of a customer email about a blocked order Moderate Some variety helps; a person reviews before sending
Brainstorm Ideas for reducing late deliveries Higher Variety is the point; nothing goes straight into a process

Questions to ask

  • For this process, does the same input need to give the same output? If yes, how do we get there beyond setting temperature?
  • Which decoding settings does the chosen model accept, and what are the defaults? What happens to our design if the next model version drops a setting?
  • What do we log for each call: prompt, model and version, settings, answer, and token counts?
  • What is the maximum answer length, and what does the system do if an answer is cut off?
  • How many input and output tokens does a typical call use, and what does that mean per month at our volumes?
  • Which generated values (amounts, quantities, document numbers) are checked against SAP before anyone acts on them?
  • Do users need to see the answer as it is written (streaming), or only the final result?

Common misconceptions

  • "Temperature 0 makes the model deterministic." It makes the choice greedy in theory. On hosted models, the serving system can still give different answers to the same request.
  • "Lower temperature makes answers more accurate." It makes them more uniform. A model that lacks a fact or rule will state the wrong answer more consistently.
  • "The model plans the whole answer, then types it out." It chooses one token at a time, each based on the text so far.
  • "Tokens are words." A token is often part of a word. A common rule of thumb for English is about four characters per token, but it varies by language and tokenizer.
  • "A confident-sounding answer comes from a confident model." The tone is generated like any other text. In the lab, a note that broke the business rule scored almost the same probability as one that followed it.
  • "Longer answers are better value." Output tokens cost money and time. Ask for the length the process needs.

Key terms

  • Token: the unit a model reads and writes; often a word piece.
  • Decoding: the rule that turns the model's probabilities into one chosen token.
  • Greedy decoding: always pick the most likely token.
  • Sampling: draw a token at random, weighted by its probability.
  • Temperature: a setting that sharpens (low) or flattens (high) the probabilities before sampling.
  • Top-p (nucleus sampling): sample only from the most likely tokens that together reach a probability p, such as 0.9.
  • Max tokens: the most output tokens a call may produce; the answer stops there.
  • Streaming: sending the answer in pieces as it is generated.
  • Input and output tokens: the prompt and the answer, counted separately for billing.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What actually happens when an LLM writes an answer?

    Answer: B. An LLM gives every token it knows a probability, a decoding rule picks one, and the loop repeats. This is why answers vary, why they arrive one piece at a time, and why cost grows with length.
  2. 2Two clerks ask the same assistant why the same sales order is blocked and get different reasons. What is the most likely cause?

    Answer: C. Most decoding settings draw tokens at random, weighted by probability, so identical requests can differ. Even temperature 0 isn't fully repeatable on hosted models. Processes that need consistency need fixed formats, checks and logs.
  3. 3A team lowers the temperature to stop wrong quantities appearing in three-way match notes. What should you expect?

    Answer: D. Temperature changes how the model picks among the options it has learned, not what it knows. In the lab, the tiny GPT broke the quantity rule at every temperature tested. Checks against purchase order and goods receipt data are what catch wrong values.
  4. 4Your team's design depends on setting temperature to 0. What risk should you raise?

    Answer: A. Anthropic's documentation says Claude models from version 4.7 accept only the default temperature. A design that relies on one knob can break when the model changes. Consistency should come from formats, checks and logging as well.
  5. 5How does the generative AI hub meter LLM use, according to SAP's AI Core guide?

    Answer: C. SAP's guide says LLM use is metered in input tokens and output tokens, converted into capacity units at per-model rates. Long prompts and long answers both add cost on every call.
  6. 6An extraction job returns half a JSON record now and then. What is the first thing to check?

    Answer: D. Every call has a maximum number of output tokens. An answer longer than that stops mid-record. The program should check why each answer stopped and handle a cut-off answer as a failure.
  7. 7Which task is the best fit for a higher temperature setting?

    Answer: C. Higher randomness adds variety, which helps when ideas are the goal and nothing flows straight into a process. Classification and extraction need the most likely answer, checked against SAP data.
Deep layer · 40 min read

Mental model: the model proposes, the decoder decides

Split text generation into two parts:

  • The model looks at the text so far and returns a score, called a logit, for every token in its vocabulary. Softmax turns the scores into probabilities. This part you built in Build a tiny GPT.
  • The decoder is a small rule outside the network. It takes those probabilities and chooses one token. Greedy, temperature, top-k and top-p are all decoder rules. None of them changes a single weight.

Then the chosen token is appended and the loop repeats. The Holtzman paper on nucleus sampling makes the point sharply: the decoding strategy alone can change text quality a great deal, with exactly the same model.

Keep this split in mind and the enterprise questions sort themselves. Knowledge, rules and facts live in the model and in the context you give it. Variety, length and stopping live in the decoder. Cost and speed come from how many times the loop runs.

How it works

The generation loop, one more time

flowchart LR
  T[Text so far] --> M[Model]
  M --> L[Logits for every token]
  L --> D[Decoder rule]
  D --> C[One chosen token]
  C --> S{Stop?}
  S -->|no| T
  S -->|end token or length limit| E[Answer]

The loop stops for one of two reasons. Either the model chooses a token that means "end", or the answer reaches the length limit. In the tiny GPT, the end token is a new line. Hosted APIs report the reason too. Anthropic's API, for example, returns a stop_reason such as end_turn for a natural end or max_tokens for a cut-off.

Greedy decoding

Greedy decoding takes the most likely token every time. It is the default in Hugging Face Transformers, whose documentation recommends it for short outputs where creativity doesn't matter. It also warns that greedy output starts to repeat itself on longer texts.

Greedy has a second weakness. The Hugging Face blog on decoding points out that it can miss a likely sequence hidden behind one unlikely token. It only ever looks one step ahead.

For a business user, greedy has a third, quieter problem: it always gives the same answer to the same prompt, even when several answers are valid. In the lab below, the tiny GPT is asked to complete "Sales order 14212 for customer C-1032 blocked: ". Four reasons appeared about equally often in its training notes. Greedy decoding always picks the one the model rates highest, and never writes "missing export license".

Sampling and temperature

Sampling draws the next token at random, weighted by the probabilities. A token with 30% probability is chosen about 30% of the time.

Temperature reshapes the probabilities before the draw. The decoder divides every logit by the temperature, then applies softmax:

  • A temperature below 1 makes big scores relatively bigger. Likely tokens get more likely.
  • A temperature above 1 flattens the distribution. Unlikely tokens get a real chance.
  • As the temperature approaches 0, sampling becomes greedy decoding, as the Hugging Face blog notes.

The tiny GPT shows it directly. After "Sales order 14212 for customer C-1032 blocked: ", the probability of the most likely first letter is 49% at temperature 0.5, 37% at 1.0 and 31% at 1.5.

Top-k and top-p

Raising the temperature has a cost: very unlikely tokens, which are often garbage, start to appear. Two filters remove the long tail before the draw:

  • Top-k keeps only the k most likely tokens. The Hugging Face blog notes that GPT-2 used it. Its weakness: k is fixed, whether the model is sure (one good option) or unsure (twenty good options).
  • Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least p. When the model is sure, the set is small. When it is unsure, the set grows. Holtzman and colleagues proposed it to cut off the "less reliable tail" while keeping variety.

In the lab, top-p 0.9 keeps 3 of 57 characters after "blocked: ", where the model is fairly sure, and 9 of 57 at the start of an invoice number, where any digit is plausible.

Beam search, briefly

Beam search keeps several candidate texts at each step and returns the one with the highest overall probability. The Hugging Face documentation suggests it for input-grounded tasks such as speech recognition. Holtzman and colleagues found that maximizing probability makes open-ended text "bland and strangely repetitive". This topic focuses on sampling, the approach their paper recommends for open-ended text, and doesn't build beam search.

Why temperature 0 still varies on hosted models

On your laptop, greedy decoding with the same model gives the same text every time. Hosted models are different. In September 2025, Horace He at Thinking Machines Lab sent the same prompt 1,000 times at temperature 0 to the open Qwen3-235B-A22B-Instruct model. The result was 80 different completions.

The cause they identified: a busy server groups many users' requests into one batch, and the batch size changes with load. The GPU routines they examined don't return bit-identical numbers for different batch sizes. Tiny numeric differences flip a close choice between two tokens, and the texts drift apart from there. With "batch-invariant" routines, all 1,000 completions were identical, at a speed cost.

The practical lesson: don't promise "same input, same output" from a hosted LLM. Design for variation, and log what each call returned.

Tokens: the unit of cost and speed

The tiny GPT uses one token per character. Production LLMs use word-piece tokenizers. OpenAI's tiktoken README says a token is about 4 bytes of text on average, which for plain English is about four characters. Other languages, long numbers and codes can split into tokens differently, so count with the tokenizer of the model you use.

Two things follow:

  • Cost. Providers count input tokens and output tokens. SAP's AI Core guide meters generative AI hub use that way, converting each into capacity units at per-model rates. In the guide's example, the output rate per 1,000 tokens is almost three times the input rate.
  • Speed. All input tokens can be processed together in one pass. Output tokens come one loop at a time. So a long answer is slower than a long question.

Why the loop is affordable: the KV cache

Look at the tiny GPT's generate loop: every new character runs the model over the whole window again. For the earlier positions, that repeats identical work. The Hugging Face documentation on caching says this directly: each prediction depends on the previous tokens, so the model performs the same computations each time.

Production systems keep a KV cache. Attention's keys and values for earlier tokens are stored and reused, so each step computes only the new token. The cost moves to memory: Hugging Face notes the cache can become a bottleneck for long contexts. This is one reason long prompts are expensive to serve, and why providers price and limit context length. The tiny GPT skips the cache to stay short.

Streaming and the stop reason

Because tokens appear one at a time, an API can send each piece as soon as it exists. That is streaming. The user sees the first words quickly, even if the full answer takes several seconds. SAP's AI Core guide says the orchestration service supports streaming for selected models.

Streaming changes nothing about cost or content. It changes how the wait feels. For a background job that parses the answer, streaming adds complexity for no benefit. For a chat window, it is usually worth it.

Always check why the answer stopped. A cut-off answer that looks complete is a classic production bug.

Log-probabilities: a weak signal, not a truth test

The probability the model gave each chosen token is a rough measure of how expected it was. Some APIs can return these as log-probabilities. They are tempting as a "confidence score". The lab shows why to be careful: a random order number gets low probability because any digit fits, not because it is wrong. And a note that breaks a business rule can score as high as one that follows it. Use log-probabilities to flag candidates for review, never as proof.

Build it yourself: turn the decoding knobs on your tiny GPT

You will load the tiny GPT you trained in the previous topic and explore four things: its top guesses at several temperatures, one note written character by character, six decoding settings measured side by side, and a probability score for any note you type. Everything runs offline.

Before you start: complete Set up your computer for this course and Set up for Unit 4. You also need three files from the previous topics in your unit04 folder: transformer_block.py from The transformer architecture, and tiny_gpt.py plus the trained tiny_gpt.pt and sap_notes.txt from Build a tiny GPT.

flowchart LR
  K[tiny_gpt.pt] --> P[peek: top guesses]
  K --> S[stream: one note, live]
  K --> C[compare: six settings, measured]
  K --> Q[score: how expected is a note]
  C --> N[decoding_notes.md]

What you need

  • The course folder with its .venv, PyTorch, and the four files above.
  • About 40 minutes. The slowest command, compare, takes about 2 minutes on our two-core test machine.
  • No accounts, no API keys, no cost.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 4 folder:

    cd unit04
  5. Check the files from the previous topics are there:

    • Windows (PowerShell):

      Get-ChildItem transformer_block.py, tiny_gpt.py, tiny_gpt.pt, sap_notes.txt
    • macOS / Linux:

      ls transformer_block.py tiny_gpt.py tiny_gpt.pt sap_notes.txt

    All four names should be listed. If one is missing, see the note under What you need.

Step 2: Save the script

  1. In VS Code's file list, right-click unit04, choose New File and name it generate_lab.py.
  2. Paste the code below and save with File > Save.
"""Unit 4: how LLMs generate text. Explore decoding with the tiny GPT you trained.

Uses tiny_gpt.pt from the previous topic (Build a tiny GPT) and imports from tiny_gpt.py.
Everything runs offline on your laptop. No accounts, no keys.

How to run (from the unit04 folder, with the course .venv turned on):
    python generate_lab.py peek                           # the model's top guesses for the next character
    python generate_lab.py stream                         # write one note character by character, with timing
    python generate_lab.py compare                        # the same model, six decoding settings, measured
    python generate_lab.py score --text "..."             # how likely the model finds each character of your text

Optional:
    python generate_lab.py peek --prompt "Invoice 51" --temperatures 0.3 1 2
    python generate_lab.py stream --temperature 0 --max-chars 40     # greedy, cut off by the length limit
    python generate_lab.py compare --count 50 --prompt "Invoice"      # faster, invoice notes only
    python generate_lab.py compare --grid                             # a wider temperature sweep for the exercise
"""
import argparse
import math
import re
import sys
import time

import torch
import torch.nn.functional as F

from tiny_gpt import load, pick_device

# ---------------------------------------------------------------- 1. the formats the notes were made from
# These patterns mirror make_note() in tiny_gpt.py. A note "matches a format" if one pattern fits it exactly.
AMT = r"\d{1,2},\d{3}\.(?:00|50) (?:EUR|USD)"
PLANT = r"(?:1010|1020|1710|2010)"
FORMATS = [
    r"(?:Sales order \d{5} for customer C-1\d{3} blocked: |Order \d{5} \(C-1\d{3}\) on hold, "
    r"|Delivery block on sales order \d{5}, customer C-1\d{3}: )"
    rf"(?:credit limit exceeded by {AMT}|overdue items on account|missing export license|price below minimum)\. "
    r"(?:Credit team to review before release\.|Sales rep to call the customer\.|Release after payment arrives\."
    r"|Check with the credit manager today\.)",
    r"(?:Invoice 51\d{4} for PO 45\d{5} held: |Three-way match failed on PO 45\d{5}, "
    r"|Payment block on invoice 51\d{4} \(PO 45\d{5}\): )"
    rf"(?:invoice quantity \d+ PC, goods receipt \d+ PC|price differs from the purchase order by {AMT}"
    r"|no goods receipt posted yet)\. "
    r"(?:Buyer to confirm with supplier V-2\d{3}\.|AP clerk to park the invoice\.|Warehouse to check the receipt\."
    r"|Ask V-2\d{3} for a credit memo\.)",
    rf"(?:MRP exception for material M-4\d{{3}} at plant {PLANT}: |Material M-4\d{{3}}, plant {PLANT}, exception: )"
    r"(?:start date in the past|stock below safety stock|reschedule in|reschedule out|opening date in the past)\. "
    rf"(?:Planner to reschedule order 1\d{{6}}\.|Expedite with the supplier\.|Check capacity at plant {PLANT}\."
    r"|Convert the planned order today\.)",
]
QTY = re.compile(r"invoice quantity (\d+) PC, goods receipt (\d+) PC")


def matches_format(note: str) -> bool:
    return any(re.fullmatch(f, note) for f in FORMATS)


def breaks_rule(note: str):
    """None if the note has no quantities; True if the receipt isn't 10 to 40 pieces below the invoice."""
    m = QTY.search(note)
    if not m:
        return None
    gap = int(m.group(1)) - int(m.group(2))
    return not (10 <= gap <= 40)


# ---------------------------------------------------------------- 2. one decoding step: scores -> one token
def next_token_probs(model, idx, context: int) -> torch.Tensor:
    """Probabilities for the next character, before any decoding setting is applied."""
    logits = model(idx[:, -context:])[0, -1, :]
    return F.softmax(logits, dim=-1)


def choose(logits: torch.Tensor, temperature: float, top_k: int, top_p: float) -> int:
    """Turn raw scores into one chosen token id. temperature 0 means greedy: always the top score."""
    if temperature == 0:
        return int(torch.argmax(logits))
    logits = logits / temperature
    if top_k > 0:                                      # keep only the k highest scores
        kth = torch.topk(logits, min(top_k, logits.numel())).values[-1]
        logits = logits.masked_fill(logits < kth, float("-inf"))
    probs = F.softmax(logits, dim=-1)
    if top_p < 1.0:                                    # keep the smallest set whose probabilities reach top_p
        sorted_p, order = torch.sort(probs, descending=True)
        keep = torch.cumsum(sorted_p, dim=0) - sorted_p < top_p
        mask = torch.zeros_like(probs, dtype=torch.bool)
        mask[order[keep]] = True
        probs = torch.where(mask, probs, torch.zeros_like(probs))
        probs = probs / probs.sum()
    return int(torch.multinomial(probs, 1))


@torch.no_grad()
def write(model, tok, cfg, prompt: str, temperature=1.0, top_k=0, top_p=1.0, max_chars=200, on_char=None):
    """Generate until a line ends ("end of note") or max_chars is reached ("length limit")."""
    idx = torch.tensor([tok.encode(prompt)])
    out = []
    for _ in range(max_chars):
        logits = model(idx[:, -cfg["context"]:])[0, -1, :]
        nxt = choose(logits, temperature, top_k, top_p)
        ch = tok.chars[nxt]
        if ch == "\n":
            return "".join(out), "end of note"
        out.append(ch)
        if on_char:
            on_char(ch)
        idx = torch.cat([idx, torch.tensor([[nxt]])], dim=1)
    return "".join(out), "length limit"


def shown(ch: str) -> str:
    return {"\n": "\\n (end of note)", " ": "' ' (space)"}.get(ch, ch)


def check_prompt(tok, prompt: str) -> None:
    unknown = sorted(set(prompt) - set(tok.chars))
    if unknown:
        raise SystemExit(f"The prompt uses characters the model never saw: {unknown}")


# ---------------------------------------------------------------- 3. commands
def peek(args) -> None:
    model, tok, cfg, _ = load(args.model, pick_device("cpu"))
    model.eval()
    check_prompt(tok, args.prompt)
    idx = torch.tensor([tok.encode(args.prompt)])
    with torch.no_grad():
        logits = model(idx[:, -cfg["context"]:])[0, -1, :]
    print(f"Prompt: {args.prompt!r}. Top {args.top} guesses for the next character:\n")
    temps = args.temperatures
    print(f"  {'char':<18}" + "".join(f"{'temp ' + format(t, 'g'):>11}" for t in temps))
    base = F.softmax(logits, dim=-1)
    order = torch.argsort(base, descending=True)[: args.top]
    tables = [F.softmax(logits / t, dim=-1) for t in temps]
    for i in order.tolist():
        print(f"  {shown(tok.chars[i]):<18}" + "".join(f"{p[i].item():>11.1%}" for p in tables))
    for t, p in zip(temps, tables):
        s = torch.sort(p, descending=True).values
        nucleus = int((torch.cumsum(s, 0) - s < 0.9).sum())
        print(f"\n  temp {t:g}: the top choice has {s[0].item():.0%}; "
              f"top-p 0.9 keeps {nucleus} of {len(tok.chars)} characters", end="")
    print()


def stream(args) -> None:
    model, tok, cfg, _ = load(args.model, pick_device("cpu"))
    model.eval()
    check_prompt(tok, args.prompt)
    torch.manual_seed(args.seed)
    start = time.time()
    first = []

    def on_char(ch):
        if not first:
            first.append(time.time() - start)
        sys.stdout.write(ch)
        sys.stdout.flush()
        time.sleep(args.delay)                         # slow it down so you can watch it arrive

    sys.stdout.write(args.prompt)
    note, reason = write(model, tok, cfg, args.prompt, args.temperature, args.top_k, args.top_p,
                         args.max_chars, on_char)
    total = time.time() - start - args.delay * len(note)
    print(f"\n\nStopped because: {reason}. Prompt {len(args.prompt)} characters in, "
          f"{len(note)} characters out.")
    if first:
        print(f"First character after {first[0] * 1000:.0f} ms; "
              f"{len(note) / max(total, 1e-9):.0f} characters per second of model time.")


def run_setting(model, tok, cfg, prompt, n, seed, temperature, top_k, top_p, seen):
    torch.manual_seed(seed)
    notes = [write(model, tok, cfg, prompt, temperature, top_k, top_p)[0] for _ in range(n)]
    full = [(prompt.strip("\n") + n_) for n_ in notes]
    fmt = sum(matches_format(x) for x in full)
    qty = [r for r in (breaks_rule(x) for x in full) if r is not None]
    return dict(fmt=fmt, distinct=len(set(full)), copies=sum(x in seen for x in full),
                qty=len(qty), broke=sum(qty), example=full[0])


def compare(args) -> None:
    model, tok, cfg, ckpt = load(args.model, pick_device("cpu"))
    model.eval()
    check_prompt(tok, args.prompt)
    text = open(ckpt["data"], encoding="utf-8").read()
    seen = set(text[: int(0.9 * len(text))].splitlines())
    if args.grid:
        settings = [(f"temperature {t:g}", t, 0, 1.0) for t in (0, 0.3, 0.5, 0.7, 1.0, 1.3, 1.6, 2.0)]
    else:
        settings = [("greedy (temperature 0)", 0, 0, 1.0), ("temperature 0.5", 0.5, 0, 1.0),
                    ("temperature 1.0", 1.0, 0, 1.0), ("temperature 1.5", 1.5, 0, 1.0),
                    ("temp 1.5 + top-k 5", 1.5, 5, 1.0), ("temp 1.5 + top-p 0.9", 1.5, 0, 0.9)]
    n = args.count
    print(f"{n} notes per setting, prompt {args.prompt!r}, seed {args.seed}\n")
    print(f"  {'setting':<24}{'format ok':>10}{'distinct':>10}{'copies':>8}{'qty rule broken':>17}")
    rows = []
    for name, t, k, p in settings:
        r = run_setting(model, tok, cfg, args.prompt, n, args.seed, t, k, p, seen)
        rows.append((name, r))
        print(f"  {name:<24}{r['fmt'] / n:>10.0%}{r['distinct']:>10}{r['copies']:>8}"
              f"{str(r['broke']) + ' of ' + str(r['qty']):>17}", flush=True)
    print("\nOne example per setting:")
    for name, r in rows:
        print(f"  {name:<24} {r['example'][:110]}")


@torch.no_grad()
def score(args) -> None:
    model, tok, cfg, _ = load(args.model, pick_device("cpu"))
    model.eval()
    text = args.text
    check_prompt(tok, text)
    ids = tok.encode("\n" + text)
    total, rows = 0.0, []
    for i in range(1, len(ids)):
        p = next_token_probs(model, torch.tensor([ids[:i]]), cfg["context"])[ids[i]].item()
        total += math.log(p)
        rows.append((tok.chars[ids[i]], p))
    print(f"Text: {text}\n")
    print(f"Average log-probability per character: {total / len(rows):.2f} "
          f"(0 would mean the model was certain of every character)")
    low = sorted(range(len(rows)), key=lambda j: rows[j][1])[: args.lowest]
    print(f"The {args.lowest} least expected characters:")
    for j in sorted(low):
        before = "..." + text[max(0, j - 18):j] if j else "(start) "
        print(f"  {before}[{rows[j][0]}]   p = {rows[j][1]:.3f}")


def main() -> None:
    p = argparse.ArgumentParser(description="Explore how a language model picks each next token.")
    sub = p.add_subparsers(dest="command", required=True)
    a = sub.add_parser("peek", help="show the top next-character guesses at several temperatures")
    a.add_argument("--prompt", default="Sales order 14212 for customer C-1032 blocked: ")
    a.add_argument("--temperatures", type=float, nargs="+", default=[0.5, 1.0, 1.5])
    a.add_argument("--top", type=int, default=6)
    s = sub.add_parser("stream", help="write one note character by character")
    s.add_argument("--prompt", default="Invoice")
    s.add_argument("--max-chars", type=int, default=200, help="the length limit, like max_tokens")
    s.add_argument("--delay", type=float, default=0.03, help="seconds to pause per character, for watching")
    c = sub.add_parser("compare", help="measure several decoding settings on many notes")
    c.add_argument("--prompt", default="\n")
    c.add_argument("--count", type=int, default=100)
    c.add_argument("--grid", action="store_true", help="sweep temperature from 0 to 2 instead")
    for sp in (s,):
        sp.add_argument("--temperature", type=float, default=1.0, help="0 means greedy")
        sp.add_argument("--top-k", type=int, default=0, help="0 means off")
        sp.add_argument("--top-p", type=float, default=1.0, help="1.0 means off")
    k = sub.add_parser("score", help="how likely the model finds each character of a text")
    k.add_argument("--text", default="Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, "
                                     "goods receipt 1460 PC. AP clerk to park the invoice.")
    k.add_argument("--lowest", type=int, default=6)
    for sp in (a, s, c, k):
        sp.add_argument("--model", default="tiny_gpt.pt")
    for sp in (s, c):
        sp.add_argument("--seed", type=int, default=0)
    args = p.parse_args()
    {"peek": peek, "stream": stream, "compare": compare, "score": score}[args.command](args)


if __name__ == "__main__":
    main()

Step 3: Peek at the model's guesses

python generate_lab.py peek

What success looks like:

Prompt: 'Sales order 14212 for customer C-1032 blocked: '. Top 6 guesses for the next character:

  char                 temp 0.5     temp 1   temp 1.5
  p                       48.6%      37.3%      31.1%
  o                       25.4%      26.9%      25.0%
  c                       23.3%      25.8%      24.3%
  m                        2.7%       8.9%      11.9%
  i                        0.0%       0.3%       1.2%
  r                        0.0%       0.2%       0.8%

  temp 0.5: the top choice has 49%; top-p 0.9 keeps 3 of 57 characters
  temp 1: the top choice has 37%; top-p 0.9 keeps 3 of 57 characters
  temp 1.5: the top choice has 31%; top-p 0.9 keeps 4 of 57 characters

How to read it:

  • The four letters are four block reasons: price below minimum, overdue items, credit limit, missing export license. In sap_notes.txt, each follows "blocked: " about a quarter of the time. Check it yourself: search the file for blocked: m.
  • The model is not well calibrated. It gives "m" only 9% at temperature 1. A small model learns the patterns roughly, not the exact frequencies.
  • Low temperature makes it worse. At 0.5, "m" drops to 2.7%. Greedy decoding would never choose it. A low setting doesn't just remove noise; it can hide valid but less likely answers.
  • Top-p adapts. Here three or four letters cover 90% of the probability, so top-p 0.9 keeps only those.

Now try a place where the model can't know the answer, the digits of an invoice number:

python generate_lab.py peek --prompt "Invoice 51" --temperatures 0.3 1 2

The last lines show the top choice at only 23% even at temperature 0.3, and top-p 0.9 keeping 7 to 10 characters. When every digit is plausible, no setting makes the model sure.

Step 4: Watch a note being written

python generate_lab.py stream

The note appears character by character, slowed down on purpose so you can watch it.

What success looks like:

Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 1460 PC. AP clerk to park the invoice.

Stopped because: end of note. Prompt 7 characters in, 106 characters out.
First character after 4 ms; 394 characters per second of model time.

This is the same note the previous topic showed, with its broken quantity. Same model, same seed (0), same settings: same note. On your laptop, sampling with a fixed seed is repeatable. Your timing numbers will differ.

Now use greedy decoding with a short length limit:

python generate_lab.py stream --temperature 0 --max-chars 40
Invoice 516662 for PO 4577991 held: invoice qua

Stopped because: length limit. Prompt 7 characters in, 40 characters out.

The note stopped mid-word because it reached the 40-character limit, just as an API answer stops at max_tokens. The "Stopped because" line is your stop reason. Run the command without --max-chars 40 and greedy decoding finishes the note: ... invoice quantity 170 PC, goods receipt 130 PC. AP clerk to park the invoice. Run it again: you get exactly the same note, every time.

Step 5: Compare six decoding settings

python generate_lab.py compare

The script writes 100 notes with each of six settings and checks every note in three ways. It takes about 2 minutes.

What success looks like:

100 notes per setting, prompt '\n', seed 0

  setting                  format ok  distinct  copies  qty rule broken
  greedy (temperature 0)        100%         1       0           0 of 0
  temperature 0.5                86%       100       0           5 of 5
  temperature 1.0                64%       100       0          8 of 10
  temperature 1.5                 4%       100       0           8 of 9
  temp 1.5 + top-k 5             16%       100       0           5 of 6
  temp 1.5 + top-p 0.9           50%       100       0         13 of 14

The columns:

  • format ok: the note fits one of the formats tiny_gpt.py made notes from, exactly: right document number lengths, valid amounts, an action that belongs to the same process.
  • distinct: how many of the 100 notes are different.
  • copies: notes copied word for word from the training part of the data.
  • qty rule broken: of the notes with an invoice and a goods receipt quantity, how many break the hidden rule (receipt 10 to 40 pieces below the invoice).

What it shows:

  • Greedy wrote the same note 100 times. Perfectly formatted and perfectly useless if you wanted 100 different drafts.
  • Format falls apart as temperature rises. At 1.5, only 4% fit a format. Top-k and top-p win much of it back by cutting the unlikely tail.
  • The business rule is broken at every setting. Even at 0.5, all 5 quantity notes break it. Decoding can't add a rule the model never learned.

To see what "format not ok" looks like, these are real failures from the model at temperature 1.0:

Delivery block on sales order 10942, customer C-1465: overdue items on account. Convert the planned order today.
Three-way match failed on PO 4567696, price differs from the purchase order by 30,91.00 EUR. Ask V-2301 for a credit memo.
Delivery block on sales order 18121, customer C-1931: credit limit exceeded by 368,,824.00 USD. Check with the credit manager today.

A planned-order action on a sales order, and two broken amounts. Each reads fluently. A format check catches all three. Only a check against system data would catch the broken quantity rule. Unit 5 covers asking a model for a fixed output format.

Step 6: Score a note

python generate_lab.py score

What success looks like:

Text: Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 1460 PC. AP clerk to park the invoice.

Average log-probability per character: -0.31 (0 would mean the model was certain of every character)
The 6 least expected characters:
  (start) [I]   p = 0.073
  ...Invoice 51992[1]   p = 0.103
  ... 519921 for PO 457[0]   p = 0.077
  ...19921 for PO 45702[8]   p = 0.096
  ...C, goods receipt 1[4]   p = 0.072
  ..., goods receipt 14[6]   p = 0.091

Now score the same note with a receipt that follows the rule:

python generate_lab.py score --text "Invoice 519921 for PO 4570287 held: invoice quantity 230 PC, goods receipt 210 PC. AP clerk to park the invoice."

The average is -0.30: almost the same. The model barely prefers the correct note. Its least expected characters are random digits, which are uncertain because any digit fits, not because they are wrong.

Finally, score the note with the wrong process action:

python generate_lab.py score --text "Delivery block on sales order 10942, customer C-1465: overdue items on account. Convert the planned order today."

The lowest score, p = 0.020, lands on the "o" of "Convert", where the note switches to a plan-to-produce action. Log-probabilities caught this error and missed the quantity one. That is the right expectation: a hint for review, not a check.

Step 7: Save your work in Git

From the course folder:

cd ..
git add unit04/generate_lab.py
git commit -m "Explore decoding settings on the tiny GPT"

What each part of the script does

Part What it does
FORMATS, matches_format Patterns that mirror make_note in tiny_gpt.py; a note passes if one pattern fits it exactly
breaks_rule Finds invoice and receipt quantities and checks the receipt is 10 to 40 pieces below the invoice
choose The decoder: greedy at temperature 0, otherwise divide by temperature, apply top-k, apply top-p, then draw
write The generation loop: crop to the context, score, choose, append; stops at a new line or the length limit and says which
peek Prints the top next-character probabilities at several temperatures, and how many characters top-p 0.9 keeps
stream Prints a note as it is generated, with the stop reason and timing
compare Writes many notes per setting and counts format passes, distinct notes, copies and rule breaks
score Adds up the log-probability of each character of a text and lists the least expected ones
load, pick_device Imported from tiny_gpt.py; load uses weights_only=True

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of Set up your computer, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'torch' PyTorch isn't in the Python you're using Check for (.venv) in the prompt. If it's there, follow Set up for Unit 3, Step 2
ModuleNotFoundError: No module named 'tiny_gpt' or 'transformer_block' One of the earlier scripts isn't in the folder you're running from Run from unit04, and save both files there from the earlier topics
tiny_gpt.pt not found The model hasn't been trained, or is in another folder Run python tiny_gpt.py make-data and python tiny_gpt.py train in unit04
FileNotFoundError for sap_notes.txt in compare The training data was deleted or moved Run python tiny_gpt.py make-data; with the default seed it recreates the same file
The prompt uses characters the model never saw The model only knows the 57 characters in its training notes Use plain letters, digits and the punctuation the notes use
Your numbers differ a little from the samples Different PyTorch versions or hardware can change the random draws That's fine. Look for the same pattern: format falls as temperature rises; the rule breaks at every setting
Asked for an API key, or a network or proxy error Nothing in this topic uses the internet or a key Check you are running generate_lab.py

The SAP way

As of October 2026, decoding settings in SAP's stack are set per model call through the generative AI hub, mostly via its orchestration service. Unit 5 sets up access and covers the service in depth. Here is how this topic's knobs map to it.

Settings travel with the model choice

In SAP Cloud SDK for AI (Python), the model and its settings are one object. SAP Learning's course shows:

LLM(name="gpt-4o", version="latest", parameters={"max_tokens": 256, "temperature": 0.2})

The SDK's reference for the newer orchestration V2 API uses LLMModelDetails with a params dictionary, and names the length limit max_completion_tokens in its example. Note the different names for the same idea. Parameter names and accepted values depend on the model and the SDK version, so check both before you rely on a setting.

A sketch of a classification call with the V2 API:

# Sketch: classify a block reason with low randomness through the orchestration service (not tested)
from gen_ai_hub.orchestration_v2.models.message import SystemMessage, UserMessage
from gen_ai_hub.orchestration_v2.models.template import Template, PromptTemplatingModuleConfig
from gen_ai_hub.orchestration_v2.models.llm_model_details import LLMModelDetails
from gen_ai_hub.orchestration_v2.models.config import ModuleConfig, OrchestrationConfig
from gen_ai_hub.orchestration_v2.service import OrchestrationService

template = Template(template=[
    SystemMessage(content="Answer with one of: CREDIT_LIMIT, OVERDUE_ITEMS, EXPORT_LICENSE, PRICE_BELOW_MIN."),
    UserMessage(content="Order note: {{?note}}"),
])
llm = LLMModelDetails(name="gpt-4o", params={"max_completion_tokens": 10, "temperature": 0})
config = OrchestrationConfig(modules=ModuleConfig(
    prompt_templating=PromptTemplatingModuleConfig(prompt=template, model=llm)))

result = OrchestrationService(config=config).run(
    placeholder_values={"note": "Customer C-1032 has open items past due since August."})
label = result.final_result.choices[0].message.content.strip()
if label not in {"CREDIT_LIMIT", "OVERDUE_ITEMS", "EXPORT_LICENSE", "PRICE_BELOW_MIN"}:
    raise ValueError(f"Unexpected label: {label!r}")   # never pass an unchecked answer on

Three habits from the lab carry over: a short length limit for a short answer, the lowest randomness the model accepts for a fixed-choice task, and a check of the answer against the allowed values.

Metering

SAP's AI Core guide (September 2026) says generative AI hub use is metered in input and output tokens. These convert to "GenAI tokens" and then capacity units, at rates that vary by model. Look up the current rates for your model before estimating; don't reuse the guide's example rates as prices.

Streaming

The same guide says the orchestration service supports response streaming for selected models. Use it for interactive screens; skip it for background jobs that parse the full answer.

What SAP doesn't change

The decoding rules are the model's. Calling a model through the generative AI hub doesn't make temperature 0 repeatable or stop a model from inventing values. The orchestration service adds templating, filtering and governed access around the same loop.

Build vs. SAP

Need Your own loop (this topic) Provider API directly Generative AI hub (orchestration)
See and change every decoding rule Yes: choose is yours Only the settings the provider exposes Only the settings the model accepts
Repeatable output Yes, on one machine with a fixed seed Not guaranteed, even at temperature 0 Not guaranteed, even at temperature 0
Quality of the text Toy model Strong models Strong models from several providers
Cost model Your own hardware Provider's per-token prices Tokens converted to SAP capacity units per model
Governance (one access point, filters, metering) None Per provider Built in, the main reason to use it in an SAP landscape
Streaming Yes, as in stream Usually For selected models

Production concerns

  • Design for variation. Assume the same input can produce a different output. Use fixed answer formats, validate them, and route failures to a person or a retry.
  • Log enough to explain an answer. Store the prompt or a reference to it, model name and version, decoding settings, the answer, the stop reason and token counts. Mind data protection when prompts contain personal data; Unit 11 covers this.
  • Check the stop reason on every call. Treat a length cut-off as a failure, not a result.
  • Validate generated values against SAP. Amounts, quantities and document numbers come from the system, through released APIs, not from the model's text. See Calling your first SAP API.
  • Budget tokens both ways. Estimate input and output tokens per call, times calls per day, at your model's rates. Long prompts are paid on every call.
  • Don't hard-wire one knob. Some models drop settings; Anthropic's documentation says Claude models from 4.7 accept only the default temperature. Keep decoding settings in configuration, and test when the model changes.
  • Use log-probabilities as a hint only. They flag surprising tokens, but random fields look surprising and wrong facts can look normal.
  • Authorizations apply to inputs. The model can only repeat what you send it or what it learned. Filter the context by the user's SAP authorizations before the call; Unit 7 covers this.
  • Clean core still applies. Generation runs outside S/4HANA, on BTP or with a provider. Read SAP data through released APIs and write back only through them.

Pitfalls

  • Promising determinism. Temperature 0 on a hosted model is not a guarantee.
  • Turning temperature down to fix wrong facts. It changes the style of the errors, not their presence.
  • Using greedy decoding for drafts. You get the same draft every time and lose valid but less likely options.
  • Using a high temperature without top-p or top-k. The long tail brings in broken tokens.
  • Ignoring cut-offs. A length limit that is too tight breaks structured answers silently.
  • Copying settings between models. Parameter names, ranges and defaults differ by model and SDK version.
  • Reading log-probabilities as confidence in the truth. They measure how expected a token was, given what the model learned.

Exercise: choose decoding settings for two SAP tasks

You will sweep temperature on invoice notes and choose settings for two tasks: a classifier that must pick one fixed exception category, and a draft writer that suggests varied notes for a person to edit. Your decoding_notes.md feeds Unit 5, where you set the same knobs on a real model.

flowchart LR
  G[compare --grid<br/>prompt Invoice] --> T[Table of results]
  T --> D1[Settings for a classifier]
  T --> D2[Settings for a draft writer]
  D1 --> N[decoding_notes.md]
  D2 --> N

Before you start: finish Build it yourself above, with the terminal in unit04 and .venv on. The sweep takes about 1 to 2 minutes.

  1. Run the temperature sweep on invoice notes:

    python generate_lab.py compare --grid --count 50 --prompt "Invoice" --seed 1

    What success looks like (the first lines; examples follow below them):

    50 notes per setting, prompt 'Invoice', seed 1
    
      setting                  format ok  distinct  copies  qty rule broken
      temperature 0                 100%         1       0          0 of 50
      temperature 0.3                96%        50       0         28 of 35
      temperature 0.5                82%        50       0         27 of 30
      temperature 0.7                82%        50       0         18 of 20
      temperature 1                  50%        50       0         18 of 18
      temperature 1.3                 8%        50       0           3 of 3
      temperature 1.6                 2%        49       0           1 of 1
      temperature 2                   0%        49       0           0 of 0

    Notice the first row: greedy wrote one note 50 times, and that one note happens to follow the rule. Every sampled setting breaks it most of the time.

  2. Test whether top-p rescues a high temperature. Run:

    python generate_lab.py compare --count 50 --prompt "Invoice" --seed 1

    Compare the format ok values of temperature 1.5 and temp 1.5 + top-p 0.9.

  3. In VS Code, create unit04/decoding_notes.md and paste this template:

    # Decoding settings for two SAP tasks
    
    ## My sweep (prompt "Invoice", seed 1, 50 notes each)
    
    | Temperature | Format ok | Distinct | Qty rule broken |
    | --- | --- | --- | --- |
    | 0 |  |  |  |
    | 0.3 |  |  |  |
    | 0.7 |  |  |  |
    | 1 |  |  |  |
    | 1.3 |  |  |  |
    
    ## Task 1: classify a three-way match exception into a fixed category
    Settings I would use:
    Why:
    Check I would add after the model answers:
    
    ## Task 2: draft varied notes for an AP clerk to edit
    Settings I would use:
    Why:
    Check I would add after the model answers:
    
    ## What no setting fixed
  4. Fill the table from your own output. Then complete both tasks. For each, name a temperature and whether you would add top-p, in one sentence say why, and name one check, such as "the answer must be one of four category codes" or "quantities must match the goods receipt in SAP".

  5. Under What no setting fixed, write two sentences about the quantity rule, and what that means for a real three-way match assistant.

  6. Save your work:

    cd ..
    git add unit04/decoding_notes.md
    git commit -m "Choose decoding settings for two SAP tasks"

If your numbers differ from the sample, that's fine. Record yours, and note your PyTorch version (python -c "import torch; print(torch.__version__)") at the top of the file.

Done when: decoding_notes.md holds your sweep table, settings and a check for each of the two tasks, and two sentences on what no setting fixed, and both generate_lab.py and decoding_notes.md are committed in Git.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Which part of text generation do temperature, top-k and top-p change?

    Answer: B. The model proposes probabilities; the decoder chooses. Temperature, top-k and top-p are decoder rules applied after the network has run. They change which token is picked, never what the model learned.
  2. 2In choose, what does dividing the logits by a temperature of 0.5 do before softmax?

    Answer: C. Dividing by a number below 1 makes large logits relatively larger, so softmax concentrates probability on the top choices. In the peek output, the top letter rose from 37% at temperature 1 to 49% at 0.5. As the temperature approaches 0, sampling becomes greedy.
  3. 3Why does top-p 0.9 keep 3 characters after "blocked: " but 9 inside an invoice number?

    Answer: D. Nucleus sampling adapts to the shape of the distribution. After "blocked: " a few letters cover 90%; for a random digit, probability is spread over many characters. Top-k, by contrast, keeps a fixed number.
  4. 4In the compare run, the quantity rule was broken at every sampled temperature. What does that tell a builder?

    Answer: A. Decoding chooses among what the model learned and can't add a missing rule. Greedy happened to produce one note that follows it, but wrote that same note every time. In a real three-way match, quantities must be checked against purchase order and goods receipt data.
  5. 5You send the same prompt at temperature 0 to a hosted model 1,000 times and get several different answers. What is the most likely explanation?

    Answer: C. Thinking Machines Lab traced exactly this to batch sizes changing with load, with routines whose results differ slightly by batch size. A close call between two tokens flips and the texts diverge. Batch-invariant routines made all 1,000 identical, at a speed cost.
  6. 6An extraction service returns JSON that sometimes fails to parse. The logs show the stop reason max_tokens on those calls. What do you do?

    Answer: B. A max_tokens stop means the answer was cut off at the length limit, as in the stream --max-chars 40 run. The fix is a limit that fits the task, plus code that never passes a cut-off answer on.
  7. 7Why are output tokens slower to produce than input tokens?

    Answer: D. The input can be processed together in one pass, but each new token depends on the previous one, so the loop runs once per output token. The KV cache makes each step cheaper by reusing earlier keys and values, but the steps still run in order.
  8. 8A note that breaks the quantity rule scores almost the same average log-probability as a correct one. What follows?

    Answer: C. Log-probabilities measure how expected each token was, given what the model learned. Random digits look surprising and a broken rule can look normal. Use them to flag text for review, never as a truth check.
  9. 9In the SAP sketch, which habit protects the process if the model returns an unexpected answer?

    Answer: C. A short limit and low temperature help, but neither guarantees a valid answer, and temperature 0 isn't repeatable on hosted models. The check against the four allowed labels is what stops a bad answer from reaching the process.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in