Orchestrate

Attention, explained

What attention is, why it lets language models link words that sit far apart, what it costs, and how to compute it yourself on an SAP-style sentence.

Updated Oct 1, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

Read this note from a sales clerk: "Order blocked, credit limit exceeded. Please release it." You know at once that "it" means the order, and that the credit limit is why the order is blocked. You linked words that sit apart in the sentence.

Attention is how a language model makes those links. For every word, the model asks "which other words matter for understanding me?" It gives each other word a weight, and the weights add up to 100%. Then it blends in information from the words with the highest weights. "It" gives most of its weight to "order" and picks up the order's meaning.

The model does this for every word, in parallel, many times over, with many different "questions" at once. Nobody writes the questions by hand. The model learns them in training, from huge amounts of text.

Attention is the core of the transformer, the design behind today's large language models (LLMs). It is why these models can follow a long instruction, resolve "it" and "this", and connect a question at the end of a prompt to a fact at the start.

Why it matters to the business

You will never configure attention in an SAP project. But three everyday decisions depend on how it works.

Long prompts cost more than their length suggests. Every word compares itself with every other word. Double the text and that comparison work roughly quadruples, which shows up as computing cost and delay. "Just paste all 400 open sales orders into the prompt" is a cost and speed decision, not a free shortcut.

A bigger window doesn't mean even reading. Models can accept very long inputs now. A widely cited study found models used information best when it sat at the start or the end of a long input, and worse when it sat in the middle. A longer-context version of the same model showed the same pattern. So a design that dumps every credit memo into one prompt and hopes the model finds the right one is fragile. Picking the few relevant documents first (retrieval, Unit 7) usually works better.

The model only links what is in front of it. Attention works over the text in the current request: the instructions, the data you pass, and the conversation so far. It doesn't reach into your SAP system. If the order's credit block reason isn't in the prompt, the model can't attend to it. It may guess instead. Getting the right SAP data into the prompt is the engineering work.

A concrete example from order-to-cash: an assistant drafts a reply to a customer asking why order 4711 hasn't shipped. If the prompt holds the order, its credit block and the customer's last payment, attention connects "why" to the block reason. If the prompt holds ten orders, attention may connect the question to the wrong one. Same model, different result, because of what you put in front of it.

How SAP does it

SAP doesn't ask customers to build attention. You meet it inside models SAP gives you access to.

  • Generative AI hub. As of October 2026, SAP Learning describes the generative AI hub in SAP AI Core as the core component for accessing LLMs. It offers models from several providers through one interface, managed in SAP AI Launchpad. Unit 5 sets it up. Everything in this topic happens inside those models.
  • SAP-RPT-1 for tables. SAP-RPT-1 is SAP's foundation model for structured business data. It predicts values in tables, for classification and regression. SAP Learning lists a small and a large version with different row and column limits. SAP's documentation for the service doesn't describe its internals. But SAP's open-source sibling model, sap-rpt-1-oss, implements a research design from SAP called ConTextTab. That design uses attention in two directions: across the columns of a row, and across rows, so a new row can look at similar example rows.

So attention isn't only for text. Wherever a model needs to decide "which other pieces matter for this one", attention is a common tool.

What attention explains in your project

You hear What attention tells you What to do
"Let's put the whole contract in the prompt." Cost grows faster than length, and the middle of long inputs is used less reliably Retrieve the relevant clauses first (Unit 7); test with the key fact placed in the middle
"The model mixed up two orders." Attention links by meaning, not by record ID; similar records compete for weight Send one case per request where you can; label each record clearly
"It forgot what we told it yesterday." Attention only sees the current request Pass the needed history or data explicitly in each call
"Why did it say that?" Attention weights show where the model looked, not a reason a business user can audit Ask the model to cite the record it used, and check it against SAP
"Can we use a model with a bigger context window?" Bigger windows accept more text but cost more per call and still favor the start and end Decide on test results with your own data, not on the window size alone

Questions to ask

  • How much text goes into each request, and what does that cost per month at our volume?
  • Which SAP data does the prompt contain, and who decided what goes in?
  • When the answer depends on one record among many, where does that record sit in the prompt?
  • Have we tested with the key fact placed in the middle of a long input, not only at the start?
  • If the model uses a record, can it point to which one, so a user can check it?
  • Is any data in the prompt something this user isn't allowed to see in SAP?

Common misconceptions

  • "Attention means the model understands like a person." It is arithmetic: scores, weights and blends. It works well because it was trained on a lot of text, not because it reasons the way a clerk does.
  • "A bigger context window means it reads everything equally well." Research shows models favor the start and end of long inputs. Window size says what fits, not how well it is used.
  • "Attention weights explain the decision." They show where one part of the model looked. A model has many layers and many attention heads, so one weight map is not an audit trail.
  • "Attention was invented for chatbots." It was introduced for machine translation in a paper first posted in 2014, three years before the transformer.
  • "The model looks things up in SAP." Attention only works on the text in the request. Getting SAP data there is your system's job.

Key terms

  • Attention: a step where each word gives weights to other words and blends in their information.
  • Attention weight: how much one word draws on another; one word's weights add up to 100%.
  • Self-attention: attention among the words of the same text.
  • Attention head: one learned "question" a model asks; models run many heads at once.
  • Context window: the most text a model can take in one request.
  • Transformer: the model design built from attention layers; the next topic in this unit.
  • Token: a piece of text, often part of a word, that the model processes as one unit.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In plain words, what does attention let a language model do?

    Answer: B. For every word, attention gives weights to the other words and blends in information from the most relevant ones. That is how "it" picks up the meaning of "order". It doesn't fetch data or keep memories.
  2. 2A team wants to paste all 400 open sales orders into every prompt. What is the main concern?

    Answer: C. Every word compares itself with every other word, so the work grows roughly with the square of the length. Research also shows models use information in the middle of long inputs less reliably. Retrieving the relevant orders first is usually better.
  3. 3An assistant answers a question about order 4711 using details from another order. What is the likely cause?

    Answer: A. Attention links by meaning, not by record ID, so similar records compete for weight. Sending one case per request where possible, and labeling records clearly, reduces mix-ups.
  4. 4Which SAP offering gives a project access to LLMs from several providers?

    Answer: D. SAP Learning describes the generative AI hub as the core component for accessing LLMs, with models from several providers. SAP-RPT-1 predicts values in tables; it is not a general text model.
  5. 5A vendor says their model's large context window means you don't need retrieval. What should you ask?

    Answer: B. Window size says how much text fits, not how well it is used. Research found models favor the start and end of long inputs, so test with the key fact in the middle, on your own data.
  6. 6A manager wants attention weight maps as the audit trail for an AI decision. What is the problem?

    Answer: D. A model has many layers and many heads, and one weight map shows only where one head looked. A better check is to have the model cite the record it used and verify it against SAP.
Deep layer · 40 min read

Mental model: a soft dictionary lookup

A Python dictionary does an exact lookup. You give it a key, it finds the one matching entry and returns that entry's value.

Attention is a soft lookup. Each word brings a query ("what am I looking for?"). Every word offers a key ("what do I contain?") and a value ("what I pass on if you pick me"). The query is compared with every key. Better matches get bigger weights, the weights are made to add up to 1, and the result is the weighted average of the values.

Nothing is an exact match, and nothing is all or nothing. "It" might take 62% of its information from "order" and spread the rest thinly. Because every step is smooth arithmetic, gradient descent can learn the queries, keys and values. You saw that kind of training in Neural networks from scratch.

How it works

Where attention came from

Before attention, translation models read a whole sentence into one fixed-length vector, then wrote the translation from that vector. Bahdanau, Cho and Bengio (ICLR 2015) argued that this single vector is a bottleneck, especially for long sentences. Their fix: while writing each output word, let the model search the source sentence for the relevant parts. It scores each source word, turns the scores into weights with a softmax, and takes a weighted sum. That is attention.

In 2017, "Attention Is All You Need" (Vaswani and colleagues) built a whole model from attention, with no recurrence. That model is the transformer. Its form of attention is the one every LLM uses today, and the one you will compute.

Scaled dot-product attention, step by step

The transformer paper writes it in one line:

Attention(Q, K, V) = softmax( Q Kᵀ / √d_k ) V

Read it right to left, one word at a time:

  1. Project. Each word's vector x is multiplied by three learned matrices: q = x · W_q, k = x · W_k, v = x · W_v. Stack all words and you get the matrices Q, K and V.
  2. Score. Take the dot product of one word's query with every word's key: Q Kᵀ. A high score means "this key matches what I'm looking for".
  3. Scale. Divide by √d_k, the square root of the key length.
  4. Mask (optional). Set the scores a word must not see to minus infinity.
  5. Softmax. Turn each row of scores into weights that add up to 1. Minus infinity becomes a weight of exactly 0.
  6. Blend. Multiply the weights by V: each word's output is a weighted average of the values.
flowchart LR
  X[Word vectors] --> Q[Queries]
  X --> K[Keys]
  X --> V[Values]
  Q --> S[Scores<br/>Q times K]
  K --> S
  S --> C[Divide by<br/>sqrt d_k]
  C --> M[Mask<br/>optional]
  M --> W[Softmax<br/>weights]
  W --> O[Weighted sum<br/>of values]
  V --> O

Why divide by the square root

The paper's reason: if the numbers in queries and keys are random with mean 0 and variance 1, their dot product has variance d_k. Longer vectors give bigger scores. Big scores push softmax to put nearly all the weight on one word. There, its gradients are tiny, so training gets almost no signal about the other words. Dividing by √d_k brings the variance back to 1. Step 5 of the walkthrough measures this.

Masks: who may look at whom

In the original transformer, attention is used in three ways:

Use Who asks Who answers Mask
Encoder self-attention Every input word Every input word None
Decoder self-attention Every output word so far Itself and earlier output words Future positions set to minus infinity
Encoder-decoder attention Every output word Every input word None

GPT-style models, which the rest of this unit builds, use only the second kind. It is called causal attention. While learning to predict the next word, a word must not see the words after it, or it could simply copy the answer. The mask enforces that.

This has a consequence you will see in Step 4. In "order blocked credit limit exceeded", the word "blocked" comes before its cause. Under a causal mask "blocked" can't look ahead to "credit limit". The information must flow to later words instead.

Many heads at once

One head asks one kind of question. The transformer runs several heads side by side, each with its own W_q, W_k and W_v. The outputs are put side by side and mixed back with one more learned matrix, W_o. The base model in the paper used 8 heads of size 64, for a width of 512. In this topic's script, one head asks "which document?" and another asks "what caused this?".

Attention ignores order unless you add it

Look at the formula again. Shuffle the words and each word gets the same weights, just in a different row. Attention on its own doesn't know which word came first. The transformer paper fixes this by adding position information to each word's vector before the first layer. The exercise uses a learned position vector for each slot. The next topic in this unit covers the options.

What it costs

Every query meets every key. For n tokens of width d, the transformer paper puts self-attention at roughly n² · d operations per layer. Double the input and that part of the work roughly quadruples. In exchange, all positions are computed at once instead of one after another, which suits GPUs well.

This is why libraries ship fast, memory-saving versions. As of PyTorch 2.13, torch.nn.functional.scaled_dot_product_attention can dispatch to a FlashAttention-2 kernel, a memory-efficient kernel or a plain C++ version. All of them compute the same formula.

Build it yourself: attention on an SAP-style sentence

You will run attention by hand on a clerk's note, "order blocked credit limit exceeded please release it". Every word gets six readable numbers instead of learned ones. Two hand-built heads decide where each word looks. You will switch the causal mask and the scaling on and off, then check your arithmetic against PyTorch.

Before you start: complete Set up your computer for this course, Set up for Unit 3 and Set up for Unit 4. They give you the orchestrate-course folder with its .venv, numpy, PyTorch and the unit04 folder. This walkthrough doesn't repeat those steps.

flowchart LR
  S[Clerk's note] --> E[Six numbers<br/>per word]
  E --> H1[Head: which document?]
  E --> H2[Head: what cause?]
  H1 --> T[Weight tables]
  H2 --> T
  T --> P[Optional: check<br/>against PyTorch]

What you need

  • The course folder and .venv from the setup topics, with numpy and PyTorch installed.
  • About 40 minutes. No accounts, no API keys, no cost.
  • No internet. Nothing is sent anywhere.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 4 folder. If it doesn't exist yet, create it first.

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit04
      cd unit04
    • macOS / Linux:

      mkdir -p unit04
      cd unit04

Step 2: Save the script

  1. In VS Code's file list, right-click unit04, choose New File and name it attention.py.
  2. Paste the code below and save with File > Save.
"""Unit 4: attention, worked by hand on a short SAP-style sentence.

Every word gets a small vector of six readable features. Two hand-built attention heads
then decide which earlier or later words each word should look at:
  - the "refer" head: words like "it" and "release" look for the business document
  - the "why" head: status words like "blocked" look for the cause
Nothing here is trained. The point is to see the arithmetic of attention, step by step.

How to run (from the unit04 folder, with the course .venv turned on):
    python attention.py                      # both heads on the default sentence
    python attention.py --head refer         # one head only
    python attention.py --causal             # GPT-style: each word sees only earlier words
    python attention.py --no-scale           # skip the divide-by-square-root step
    python attention.py --sentence "release it please"   # your own words from the vocabulary
    python attention.py --scale-demo         # why the square-root scaling exists
    python attention.py --check-torch        # compare with PyTorch's built-in attention
"""
import argparse
import math

import numpy as np

FEATURES = ["document", "status", "cause", "pronoun", "action", "filler"]

# Each word as six numbers, one per feature above. Hand-made, so you can read them.
VOCAB = {
    "order":    [1.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    "invoice":  [1.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    "blocked":  [0.0, 1.0, 0.0, 0.0, 0.0, 0.0],
    "credit":   [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
    "limit":    [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
    "price":    [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
    "exceeded": [0.0, 0.5, 0.5, 0.0, 0.0, 0.0],
    "release":  [0.0, 0.0, 0.0, 0.0, 1.0, 0.0],
    "it":       [0.0, 0.0, 0.0, 1.0, 0.0, 0.0],
    "please":   [0.0, 0.0, 0.0, 0.0, 0.0, 1.0],
    "the":      [0.0, 0.0, 0.0, 0.0, 0.0, 1.0],
}
DEFAULT_SENTENCE = "order blocked credit limit exceeded please release it"
D = len(FEATURES)


def head_weights(name: str):
    """Return the query, key and value matrices (each 6 x 6) for one hand-built head.

    A word's query = its vector @ W_q, "what am I looking for".
    A word's key   = its vector @ W_k, "what do I offer".
    A word's value = its vector @ W_v, "what I pass on if you pick me".
    """
    w_q = np.zeros((D, D))
    w_k = np.zeros((D, D))
    if name == "refer":
        w_q[FEATURES.index("pronoun"), 0] = 3.0   # "it" asks: where is a document?
        w_q[FEATURES.index("action"), 0] = 2.0    # "release" asks that too: release what?
        w_k[FEATURES.index("document"), 0] = 2.0  # documents answer
    elif name == "why":
        w_q[FEATURES.index("status"), 0] = 3.0    # status words ask: what is the cause?
        w_k[FEATURES.index("cause"), 0] = 2.0     # cause words answer
    else:
        raise ValueError(name)
    w_v = np.eye(D)                               # pass the word's features on unchanged
    return w_q, w_k, w_v


def softmax(scores: np.ndarray) -> np.ndarray:
    """Turn each row of scores into weights that are positive and add up to 1."""
    shifted = scores - scores.max(axis=-1, keepdims=True)   # for numerical safety; same result
    e = np.exp(shifted)
    return e / e.sum(axis=-1, keepdims=True)


def attention(x, w_q, w_k, w_v, causal=False, scale=True):
    """Scaled dot-product attention for one head. x has one row per word."""
    q, k, v = x @ w_q, x @ w_k, x @ w_v
    scores = q @ k.T                                   # how well each query matches each key
    if scale:
        scores = scores / math.sqrt(k.shape[-1])       # divide by the square root of the key size
    if causal:
        n = len(x)
        future = np.triu(np.ones((n, n), dtype=bool), k=1)
        scores = np.where(future, -np.inf, scores)     # a word may not look at later words
    weights = softmax(scores)
    return weights @ v, weights


def print_weights(words, weights, title):
    width = max(len(w) for w in words) + 1
    print(f"\n{title}  (row = the word asking, column = the word looked at; each row adds up to 1)")
    print(" " * width + "".join(f"{w[:8]:>9}" for w in words))
    for i, w in enumerate(words):
        cells = "".join(f"{weights[i, j]:9.2f}" for j in range(len(words)))
        print(f"{w:<{width}}{cells}")
    print("Strongest link per word:")
    for i, w in enumerate(words):
        j = int(np.argmax(weights[i]))
        visible = weights[i][weights[i] > 0]
        if len(visible) == 1:
            note = "(only itself is visible)"
        elif np.allclose(visible, visible.max()):
            note = "(spread evenly: no visible word matches what this word asks for)"
        else:
            note = ""
        print(f"  {w:<{width}} -> {words[j]:<10} {weights[i, j]:.2f} {note}")


def run_sentence(args):
    words = args.sentence.lower().split()
    unknown = [w for w in words if w not in VOCAB]
    if unknown:
        raise SystemExit(f"Unknown word(s): {', '.join(unknown)}. Use words from: {', '.join(VOCAB)}")
    x = np.array([VOCAB[w] for w in words])
    heads = ["refer", "why"] if args.head == "both" else [args.head]
    outputs = []
    for name in heads:
        out, weights = attention(x, *head_weights(name), causal=args.causal, scale=not args.no_scale)
        outputs.append(out)
        mode = "causal" if args.causal else "full"
        print_weights(words, weights, f"Head '{name}', {mode} attention, scaling {'off' if args.no_scale else 'on'}")
    # Multi-head: put the heads' outputs side by side, then mix back to 6 numbers with W_o.
    combined = np.concatenate(outputs, axis=-1)
    w_o = np.vstack([np.eye(D)] * len(heads))
    new_x = x + combined @ w_o                         # residual: keep the word, add what it gathered
    print("\nEach word's features before -> after attention (heads combined, original word kept):")
    print(f"{'':<10}" + "".join(f"{f:>10}" for f in FEATURES))
    for i, w in enumerate(words):
        print(f"{w:<10}" + "".join(f" {a:4.2f}>{b:4.2f}" for a, b in zip(x[i], new_x[i])))


def scale_demo():
    rng = np.random.default_rng(0)
    print("Random queries and keys with 8 candidate words, 2,000 trials per size.")
    print(f"{'key size':>9} {'spread of scores':>17} {'top weight, unscaled':>21} {'top weight, scaled':>19}")
    for d in [4, 16, 64, 256, 1024]:
        q = rng.standard_normal((2000, 1, d))
        k = rng.standard_normal((2000, 8, d))
        s = (q @ k.transpose(0, 2, 1))[:, 0, :]
        top_raw = softmax(s).max(axis=-1).mean()
        top_scaled = softmax(s / math.sqrt(d)).max(axis=-1).mean()
        print(f"{d:>9} {s.std():>17.1f} {top_raw:>21.2f} {top_scaled:>19.2f}")
    print("Unscaled, the top weight heads toward 1.00 as vectors grow: one word takes everything,")
    print("and training gets almost no signal about the others. Scaling keeps the spread steady.")


def check_torch(args):
    try:
        import torch
        import torch.nn.functional as F
    except ImportError:
        raise SystemExit("PyTorch is not installed in this environment. See Set up for Unit 3.")
    words = args.sentence.lower().split()
    x = np.array([VOCAB[w] for w in words])
    for name in ["refer", "why"]:
        w_q, w_k, w_v = head_weights(name)
        ours, _ = attention(x, w_q, w_k, w_v, causal=args.causal)
        t = lambda a: torch.tensor(a, dtype=torch.float64)
        theirs = F.scaled_dot_product_attention(t(x @ w_q), t(x @ w_k), t(x @ w_v), is_causal=args.causal)
        diff = float(np.abs(ours - theirs.numpy()).max())
        print(f"Head '{name}': largest difference from PyTorch = {diff:.1e}  {'OK' if diff < 1e-9 else 'MISMATCH'}")


def main():
    p = argparse.ArgumentParser(description="Attention, worked by hand.")
    p.add_argument("--sentence", default=DEFAULT_SENTENCE)
    p.add_argument("--head", choices=["refer", "why", "both"], default="both")
    p.add_argument("--causal", action="store_true", help="each word sees only itself and earlier words")
    p.add_argument("--no-scale", action="store_true", help="skip dividing by the square root of the key size")
    p.add_argument("--scale-demo", action="store_true", help="show why scaling matters")
    p.add_argument("--check-torch", action="store_true", help="compare with PyTorch's attention")
    args = p.parse_args()
    np.set_printoptions(precision=2, suppress=True)
    if args.scale_demo:
        scale_demo()
    elif args.check_torch:
        check_torch(args)
    else:
        run_sentence(args)


if __name__ == "__main__":
    main()

Step 3: Run it and read the "refer" head

Run one head first:

python attention.py --head refer

What success looks like:

Head 'refer', full attention, scaling on  (row = the word asking, column = the word looked at; each row adds up to 1)
             order  blocked   credit    limit exceeded   please  release       it
order         0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
blocked       0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
credit        0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
limit         0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
exceeded      0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
please        0.12     0.12     0.12     0.12     0.12     0.12     0.12     0.12
release       0.42     0.08     0.08     0.08     0.08     0.08     0.08     0.08
it            0.62     0.05     0.05     0.05     0.05     0.05     0.05     0.05
Strongest link per word:
  order     -> order      0.12 (spread evenly: no visible word matches what this word asks for)
  ...
  release   -> order      0.42
  it        -> order      0.62

How to read it:

  • Row "it": 0.62 of its weight goes to "order". The query of "it" (pronoun feature × 3) matches the key of "order" (document feature × 2). The score is 3 × 2 = 6. Divided by √6 it is about 2.45. Softmax turns 2.45 against seven zeros into 0.62.
  • Row "release": a weaker query (2 instead of 3), so a weaker pull, 0.42.
  • Rows at 0.12: these words ask nothing in this head. All scores are 0, so softmax spreads the weight evenly: 1 ÷ 8 ≈ 0.12. That is a valid result, not an error.

Below the table, the "before -> after" block shows what attention did. "It" started with document = 0.00 and ends with 0.62. The word "it" now carries part of the order's meaning.

Step 4: Run both heads, then switch on the causal mask

  1. Run both heads:

    python attention.py

    The "why" head shows "blocked" putting 0.37 on "credit" and 0.37 on "limit". Each status word looks for its cause.

  2. Now run the "why" head as a GPT would, with the causal mask:

    python attention.py --head why --causal

    What success looks like (first rows):

    Head 'why', causal attention, scaling on  (row = the word asking, column = the word looked at; each row adds up to 1)
                 order  blocked   credit    limit exceeded   please  release       it
    order         1.00     0.00     0.00     0.00     0.00     0.00     0.00     0.00
    blocked       0.50     0.50     0.00     0.00     0.00     0.00     0.00     0.00
    credit        0.33     0.33     0.33     0.00     0.00     0.00     0.00     0.00

Everything above the diagonal is now 0.00. "Order" can only see itself. "Blocked" wants a cause but "credit limit" comes later, so it gets nothing useful. Its weight spreads over what it can see. "Exceeded" comes after "credit limit", so it can still find the cause. That is how causal models work: later words gather context from earlier ones.

Step 5: See why the scaling exists

  1. Turn scaling off for one head:

    python attention.py --head refer --no-scale

    The weight of "it" on "order" jumps from 0.62 to about 0.98. Without the division, scores are larger, and softmax is more all-or-nothing.

  2. Run the experiment with random vectors of growing size:

    python attention.py --scale-demo

    What success looks like:

    Random queries and keys with 8 candidate words, 2,000 trials per size.
     key size  spread of scores  top weight, unscaled  top weight, scaled
            4               2.0                  0.52                0.34
           16               4.0                  0.75                0.36
           64               8.0                  0.87                0.36
          256              15.9                  0.94                0.36
         1024              31.8                  0.97                0.36

The spread of scores doubles each time the key size grows four times: it grows with the square root of the size. Unscaled, the top weight creeps toward 1.00, so one word takes everything. Scaled, it stays near 0.36 at every size. That is the transformer paper's argument, measured.

Step 6: Try your own sentence, and check against PyTorch

  1. Use any words from the script's vocabulary (order, invoice, blocked, credit, limit, price, exceeded, release, it, please, the):

    python attention.py --sentence "invoice blocked price exceeded release it"

    On Windows PowerShell the same command works; keep the double quotes.

    In the "before -> after" block, "it" now ends with document = 0.87. With only one document word and fewer words competing, the link is stronger.

  2. Check your attention function against PyTorch's built-in one:

    python attention.py --check-torch
    python attention.py --check-torch --causal

    What success looks like:

    Head 'refer': largest difference from PyTorch = 8.3e-17  OK
    Head 'why': largest difference from PyTorch = 1.1e-16  OK

    Differences around 1e-16 are rounding noise. Your ten lines of numpy compute the same thing as torch.nn.functional.scaled_dot_product_attention.

Step 7: Save your work in Git

From the course folder:

cd ..
git add unit04/attention.py
git commit -m "Compute attention by hand on an SAP-style sentence"

What each part of the script does

Part What it does
FEATURES, VOCAB Six readable numbers per word: document, status, cause, pronoun, action, filler
head_weights Builds W_q, W_k and W_v by hand for the "refer" and "why" heads; W_v passes features on unchanged
softmax Turns each row of scores into weights that add up to 1; subtracting the row maximum first avoids overflow
attention The formula: project, score, divide by the square root of the key size, mask future words if causal, softmax, weighted sum
print_weights Prints the weight table and each word's strongest link
run_sentence Runs the chosen heads, combines them with W_o, and adds the result to each word's original vector
scale_demo Random queries and keys of growing size, with and without scaling
check_torch Runs PyTorch's scaled_dot_product_attention on the same inputs and reports the largest difference

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of Set up your computer, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'numpy' The library isn't in the Python you're using Check for (.venv) in the prompt. If it's there, run pip install -r requirements.txt from the course folder
PyTorch is not installed in this environment PyTorch isn't in this .venv Follow Set up for Unit 3, Step 2. Everything except --check-torch works without it
pip shows ProxyError, SSLError or Could not fetch URL Your network or company proxy blocks the package sites Try another network, or ask IT for access or an internal mirror. The script itself needs no network
Unknown word(s): ... A word in --sentence isn't in the vocabulary Use only the listed words, or add your own word to VOCAB with six numbers
can't open file ... attention.py The terminal isn't in unit04, or the file has another name Run cd unit04 from the course folder, and check the file name
Asked for an API key Nothing in this topic uses a key or account Check you are running attention.py, not another script

Where this shows up in SAP

Unit 4 is a mechanics unit, so this section is short.

Inside the models of the generative AI hub

As of October 2026, SAP Learning describes the generative AI hub as the core component for accessing LLMs in SAP AI Core, with models from several providers and an orchestration service. The attention you computed runs inside those models on every call. You don't configure it. You control what it works on: the prompt, the data and the order you put them in. Unit 5 sets up access to the hub.

Attention over tables: SAP-RPT-1 and ConTextTab

SAP-RPT-1 predicts values in tables from example rows sent with the request. SAP Learning lists sap-rpt-1-small (up to 2,048 rows and 100 columns) and sap-rpt-1-large (up to 65,536 rows and 256 columns). SAP's documentation for the service doesn't describe its architecture.

SAP's open-source sibling, sap-rpt-1-oss in the SAP-samples GitHub organization, implements the ConTextTab paper by SAP researchers. The paper's model alternates two kinds of self-attention layers:

Layer Who attends to whom Mask
Cross-column Each cell attends to the other cells in its row None
Cross-row Each row attends to other rows in the same column Rows may only attend to the provided example rows

The cross-row mask plays the same role as your causal mask: it stops a row from seeing what it must not see. Here, the rows to be predicted must not look at each other, only at the labeled examples. The repository says its checkpoints are for research use only and recommends about 80 GB of GPU memory. It is a reference for how the idea works, not something to run in this course.

Your own attention code

If you ever train your own transformer in SAP AI Core, the attention layer is the same PyTorch function you checked in Step 6. Unit 10 covers running AI workloads on BTP.

Build vs. SAP

Situation Hand-written attention PyTorch's built-in attention A model through SAP
Learning how it works Best: every step is visible Use it to check your work Not visible
Training a small model of your own Too slow and no memory savings Best: fast kernels, same result Not applicable
Text tasks in an SAP process No Only if you train and host a model Best: LLMs through the generative AI hub
Predictions on SAP tables No Possible with a research model and a large GPU SAP-RPT-1, with example rows in the request

Production concerns

  • Cost grows with prompt length. Attention work grows roughly with the square of the length per layer, so long prompts add compute and delay. Measure tokens per request early and test at real volume. Unit 10 covers cost and routing.
  • Position matters. Research on long inputs found the middle is used least reliably, with the same pattern in a longer-context model. Put instructions and the key record where they are easy to use, and test with the key fact moved around.
  • Everything in the prompt can influence everything else. Self-attention lets every token draw on every other token. Text from a customer email, a document or a tool result can therefore steer the output as much as your instructions can. This is the root of prompt injection (Unit 11). Keep untrusted text clearly separated and never let it grant permissions.
  • Only send what the user may see. Attention will happily use any data you include. Filter SAP data by the user's authorizations before it reaches the prompt (Unit 7 and Unit 11).
  • Weights aren't explanations. One head's weight map is a debugging aid, not evidence for an auditor. For business traceability, have the model cite the record it used and check it.
  • Use the library function. For your own models, call scaled_dot_product_attention rather than a hand-written version. It can pick fast, memory-saving kernels and computes the same result.

Pitfalls

  • Forgetting the scaling. Without dividing by √d_k, softmax saturates as vectors grow and training stalls. Step 5 shows it.
  • A mask the wrong way round. In PyTorch's scaled_dot_product_attention, a boolean attn_mask value of True means the position takes part. Other APIs use the opposite convention. Check the documentation of the function you call.
  • Masking after softmax. The mask must set scores to minus infinity before softmax. Zeroing weights afterwards leaves rows that no longer add up to 1.
  • Softmax over the wrong axis. Weights must add up to 1 across the keys (the last axis), one row per query. Summing down the columns gives nonsense that still runs.
  • Leaving out positions. Without position information, "order blocks credit" and "credit blocks order" look the same to attention.
  • Reading one head as the model's reason. Real models have many heads in many layers. Don't build a business explanation on one weight map.

Exercise: train one attention head and watch it learn where to look

So far you built the heads by hand. Now PyTorch learns one. The task: read a row of eight made-up material codes and, at the last position, name the code that sat at a chosen position. Only attention can solve it, because the answer is far from where it is asked for. The Head class you save here is the building block for the tiny GPT later in this unit.

flowchart LR
  R[Row of 8 codes] --> E[Code + position<br/>vectors]
  E --> H[One causal<br/>attention head]
  H --> O[Score for<br/>each code]
  O --> A[Accuracy and<br/>attention map]

Before you start: the same setup as above. PyTorch must be installed (Set up for Unit 3). Training takes a few seconds on a laptop CPU.

  1. In VS Code, create unit04/head.py, paste the code below and save.

    """Unit 4 exercise: one trainable attention head in PyTorch, checked and then trained.
    
    Part 1 checks the head: it must match PyTorch's built-in attention, and in causal mode
    changing a later token must never change an earlier output.
    Part 2 trains a tiny model on a lookup task: read a row of made-up material codes and,
    at the last position, say which code sat at position --target-pos. Only attention can
    solve it, because the answer is far away from where it is asked for.
    
    How to run (from the unit04 folder, with the course .venv turned on):
        python head.py                    # checks, then training with the answer at position 0
        python head.py --target-pos 3     # move the answer; watch the attention follow it
        python head.py --steps 15         # too little training; accuracy stays low
        python head.py --check-only       # only Part 1
    """
    import argparse
    
    import torch
    import torch.nn as nn
    import torch.nn.functional as F
    
    
    class Head(nn.Module):
        """One causal self-attention head: queries, keys and values are learned linear maps."""
    
        def __init__(self, n_embd: int, head_size: int):
            super().__init__()
            self.query = nn.Linear(n_embd, head_size, bias=False)
            self.key = nn.Linear(n_embd, head_size, bias=False)
            self.value = nn.Linear(n_embd, head_size, bias=False)
            self.last_weights = None          # kept so we can print what the head looked at
    
        def forward(self, x: torch.Tensor) -> torch.Tensor:
            # x has shape (batch, positions, n_embd)
            q, k, v = self.query(x), self.key(x), self.value(x)
            scores = q @ k.transpose(-2, -1) / k.shape[-1] ** 0.5          # (batch, positions, positions)
            t = x.shape[1]
            future = torch.triu(torch.ones(t, t, dtype=torch.bool, device=x.device), diagonal=1)
            scores = scores.masked_fill(future, float("-inf"))            # no peeking at later positions
            weights = F.softmax(scores, dim=-1)
            self.last_weights = weights.detach()
            return weights @ v
    
    
    class LookupModel(nn.Module):
        """Token embedding + position embedding -> one attention head -> a score for every code."""
    
        def __init__(self, vocab: int, length: int, n_embd: int = 32, head_size: int = 32):
            super().__init__()
            self.tok = nn.Embedding(vocab, n_embd)
            self.pos = nn.Embedding(length, n_embd)
            self.head = Head(n_embd, head_size)
            self.out = nn.Linear(head_size, vocab)
    
        def forward(self, idx: torch.Tensor) -> torch.Tensor:
            positions = torch.arange(idx.shape[1], device=idx.device)
            x = self.tok(idx) + self.pos(positions)
            return self.out(self.head(x))
    
    
    def run_checks() -> bool:
        torch.manual_seed(0)
        head = Head(n_embd=8, head_size=4).double()
        x = torch.randn(2, 5, 8, dtype=torch.float64)
        ours = head(x)
        q, k, v = head.query(x), head.key(x), head.value(x)
        theirs = F.scaled_dot_product_attention(q, k, v, is_causal=True)
        diff = (ours - theirs).abs().max().item()
        ok1 = diff < 1e-9
        print(f"Check 1, matches PyTorch's attention: largest difference {diff:.1e}  {'OK' if ok1 else 'FAIL'}")
    
        x2 = x.clone()
        x2[:, -1, :] = torch.randn(2, 8, dtype=torch.float64)                # change only the last position
        change = (head(x2)[:, :-1] - ours[:, :-1]).abs().max().item()
        ok2 = change < 1e-12
        print(f"Check 2, later tokens can't change earlier outputs: change {change:.1e}  {'OK' if ok2 else 'FAIL'}")
    
        rows = head.last_weights.sum(dim=-1)
        ok3 = torch.allclose(rows, torch.ones_like(rows))
        print(f"Check 3, every row of weights adds up to 1:  {'OK' if ok3 else 'FAIL'}")
        return ok1 and ok2 and ok3
    
    
    def make_batch(n: int, vocab: int, length: int, target_pos: int):
        idx = torch.randint(0, vocab, (n, length))
        return idx, idx[:, target_pos]
    
    
    def train(args) -> None:
        torch.manual_seed(args.seed)
        model = LookupModel(args.vocab, args.length)
        opt = torch.optim.AdamW(model.parameters(), lr=3e-3)
        print(f"\nTraining: {args.length} codes per row, answer at position {args.target_pos}, {args.steps} steps")
        for step in range(1, args.steps + 1):
            idx, target = make_batch(64, args.vocab, args.length, args.target_pos)
            logits = model(idx)[:, -1, :]                   # only the last position answers
            loss = F.cross_entropy(logits, target)
            opt.zero_grad()
            loss.backward()
            opt.step()
            if step % max(1, args.steps // 5) == 0:
                print(f"  step {step:>4}  loss {loss.item():.3f}")
    
        model.eval()
        with torch.no_grad():
            idx, target = make_batch(1000, args.vocab, args.length, args.target_pos)
            acc = (model(idx)[:, -1, :].argmax(-1) == target).float().mean().item()
            avg = model.head.last_weights[:, -1, :].mean(dim=0)
        print(f"Accuracy on 1,000 new rows: {acc:.2f}  (guessing would give about {1 / args.vocab:.2f})")
        print("Where the last position looks, averaged over those rows:")
        for p, w in enumerate(avg.tolist()):
            print(f"  position {p}: {w:.2f} {'#' * round(w * 40)}")
        peak = int(avg.argmax())
        print(f"Peak attention at position {peak}; the answer was at position {args.target_pos}.")
    
    
    def main() -> None:
        p = argparse.ArgumentParser(description="Check and train one attention head.")
        p.add_argument("--target-pos", type=int, default=0, help="where the answer sits in each row")
        p.add_argument("--length", type=int, default=8, help="codes per row")
        p.add_argument("--vocab", type=int, default=20, help="number of different material codes")
        p.add_argument("--steps", type=int, default=600)
        p.add_argument("--seed", type=int, default=0)
        p.add_argument("--check-only", action="store_true")
        args = p.parse_args()
        if not 0 <= args.target_pos < args.length:
            raise SystemExit(f"--target-pos must be between 0 and {args.length - 1}.")
        passed = run_checks()
        if not passed:
            raise SystemExit("A check failed. Compare your Head class with the published code.")
        if not args.check_only:
            train(args)
    
    
    if __name__ == "__main__":
        main()
  2. In the terminal, inside unit04, run:

    python head.py

    What success looks like (loss values can differ slightly between computers):

    Check 1, matches PyTorch's attention: largest difference 1.1e-16  OK
    Check 2, later tokens can't change earlier outputs: change 0.0e+00  OK
    Check 3, every row of weights adds up to 1:  OK
    
    Training: 8 codes per row, answer at position 0, 600 steps
      step  120  loss 0.012
      ...
    Accuracy on 1,000 new rows: 1.00  (guessing would give about 0.05)
    Where the last position looks, averaged over those rows:
      position 0: 1.00 ########################################
      position 1: 0.00
      ...
    Peak attention at position 0; the answer was at position 0.
  3. Move the answer and run again:

    python head.py --target-pos 3

    The attention bar should jump to position 3. Nobody told the head where to look. It learned it from the loss.

  4. Train far too little:

    python head.py --steps 15

    In our run, accuracy was 0.44 while 0.74 of the attention already sat on the right position. The head learns where to look first; the output layer still has to learn what to say.

  5. Create unit04/attention_notes.md with three short lines: the accuracy and peak position from step 2, the same from step 3, and one sentence on what step 4 showed.

  6. Save your work:

    cd ..
    git add unit04/head.py unit04/attention_notes.md
    git commit -m "Train one attention head on a lookup task"
Part of head.py What it does
Head One causal self-attention head with learned query, key and value matrices; keeps its last weights for printing
LookupModel Code vector plus position vector, one head, then a layer that scores all 20 codes
run_checks Compares with PyTorch's function, proves later tokens can't change earlier outputs, checks rows add up to 1
make_batch Random rows of codes; the answer is the code at --target-pos
train AdamW training on the last position only, then accuracy and the average attention map on 1,000 new rows

If something goes wrong, the table in Build it yourself applies. One more case: --target-pos must be between 0 and 7 means you asked for a position outside the row; use 0 to 7, or change --length.

Done when: python head.py --target-pos 3 prints three OK checks, accuracy of at least 0.95 and peak attention at position 3, and head.py and attention_notes.md are committed in Git.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In scaled dot-product attention, what do the query, key and value each do?

    Answer: B. Each word's query is compared with every key, the scores become weights through softmax, and the output is the weighted average of the values. It is a soft lookup, not an exact match.
  2. 2Why does the formula divide the scores by the square root of the key size?

    Answer: C. For random vectors the dot product's variance grows with the key size, so larger vectors give larger scores. Softmax then saturates and training gets tiny gradients; dividing by the square root keeps the spread steady, as the scale demo showed.
  3. 3In the causal run, why does "blocked" fail to find "credit limit"?

    Answer: A. The causal mask sets every score above the diagonal to minus infinity, so those weights become 0. "Blocked" can only see "order" and itself; "exceeded", which comes later, can still reach the cause.
  4. 4You write your own mask and set the forbidden weights to 0 after the softmax. What goes wrong?

    Answer: D. The mask must set scores to minus infinity before softmax, so the remaining weights are renormalized to add up to 1. Zeroing weights afterwards leaves rows that sum to less than 1.
  5. 5What does attn_mask with the value True mean in PyTorch's scaled_dot_product_attention?

    Answer: B. PyTorch's documentation says a boolean True means the element should take part in attention. Other APIs use the opposite convention, which is a common source of silent bugs.
  6. 6A colleague doubles the length of every prompt to include more SAP documents. What happens to attention's work per layer?

    Answer: C. Self-attention compares each query with every key, about n squared times d operations per layer. Twice the tokens means about four times that work, which shows up as cost and delay.
  7. 7Your trained head reaches 0.44 accuracy with most attention on the right position. What does that tell you?

    Answer: D. In the short run, 0.74 of the attention already sat on the answer's position while accuracy was 0.44. More training lets the output layer turn the gathered information into the right code.
  8. 8Text from a customer email in the prompt makes your assistant ignore its instructions. Why is that possible?

    Answer: A. Attention doesn't know which text is trusted; it blends information across everything in the context. That is the root of prompt injection, so untrusted text must be kept separate and never allowed to grant permissions.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in