Read this note from a sales clerk: "Order blocked, credit limit exceeded. Please release it." You know at once that "it" means the order, and that the credit limit is why the order is blocked. You linked words that sit apart in the sentence.
Attention is how a language model makes those links. For every word, the model asks "which other words matter for understanding me?" It gives each other word a weight, and the weights add up to 100%. Then it blends in information from the words with the highest weights. "It" gives most of its weight to "order" and picks up the order's meaning.
The model does this for every word, in parallel, many times over, with many different "questions" at once. Nobody writes the questions by hand. The model learns them in training, from huge amounts of text.
Attention is the core of the transformer, the design behind today's large language models (LLMs). It is why these models can follow a long instruction, resolve "it" and "this", and connect a question at the end of a prompt to a fact at the start.
You will never configure attention in an SAP project. But three everyday decisions depend on how it works.
Long prompts cost more than their length suggests. Every word compares itself with every other word. Double the text and that comparison work roughly quadruples, which shows up as computing cost and delay. "Just paste all 400 open sales orders into the prompt" is a cost and speed decision, not a free shortcut.
A bigger window doesn't mean even reading. Models can accept very long inputs now. A widely cited study found models used information best when it sat at the start or the end of a long input, and worse when it sat in the middle. A longer-context version of the same model showed the same pattern. So a design that dumps every credit memo into one prompt and hopes the model finds the right one is fragile. Picking the few relevant documents first (retrieval, Unit 7) usually works better.
The model only links what is in front of it. Attention works over the text in the current request: the instructions, the data you pass, and the conversation so far. It doesn't reach into your SAP system. If the order's credit block reason isn't in the prompt, the model can't attend to it. It may guess instead. Getting the right SAP data into the prompt is the engineering work.
A concrete example from order-to-cash: an assistant drafts a reply to a customer asking why order 4711 hasn't shipped. If the prompt holds the order, its credit block and the customer's last payment, attention connects "why" to the block reason. If the prompt holds ten orders, attention may connect the question to the wrong one. Same model, different result, because of what you put in front of it.
SAP doesn't ask customers to build attention. You meet it inside models SAP gives you access to.
Generative AI hub. As of October 2026, SAP Learning describes the generative AI hub in SAP AI Core as the core component for accessing LLMs. It offers models from several providers through one interface, managed in SAP AI Launchpad. Unit 5 sets it up. Everything in this topic happens inside those models.
SAP-RPT-1 for tables. SAP-RPT-1 is SAP's foundation model for structured business data. It predicts values in tables, for classification and regression. SAP Learning lists a small and a large version with different row and column limits. SAP's documentation for the service doesn't describe its internals. But SAP's open-source sibling model, sap-rpt-1-oss, implements a research design from SAP called ConTextTab. That design uses attention in two directions: across the columns of a row, and across rows, so a new row can look at similar example rows.
So attention isn't only for text. Wherever a model needs to decide "which other pieces matter for this one", attention is a common tool.
"Attention means the model understands like a person." It is arithmetic: scores, weights and blends. It works well because it was trained on a lot of text, not because it reasons the way a clerk does.
"A bigger context window means it reads everything equally well." Research shows models favor the start and end of long inputs. Window size says what fits, not how well it is used.
"Attention weights explain the decision." They show where one part of the model looked. A model has many layers and many attention heads, so one weight map is not an audit trail.
"Attention was invented for chatbots." It was introduced for machine translation in a paper first posted in 2014, three years before the transformer.
"The model looks things up in SAP." Attention only works on the text in the request. Getting SAP data there is your system's job.
Pick one answer for each question. The explanation appears after you choose.
1In plain words, what does attention let a language model do?
Answer: B. For every word, attention gives weights to the other words and blends in information from the most relevant ones. That is how "it" picks up the meaning of "order". It doesn't fetch data or keep memories.
2A team wants to paste all 400 open sales orders into every prompt. What is the main concern?
Answer: C. Every word compares itself with every other word, so the work grows roughly with the square of the length. Research also shows models use information in the middle of long inputs less reliably. Retrieving the relevant orders first is usually better.
3An assistant answers a question about order 4711 using details from another order. What is the likely cause?
Answer: A. Attention links by meaning, not by record ID, so similar records compete for weight. Sending one case per request where possible, and labeling records clearly, reduces mix-ups.
4Which SAP offering gives a project access to LLMs from several providers?
Answer: D. SAP Learning describes the generative AI hub as the core component for accessing LLMs, with models from several providers. SAP-RPT-1 predicts values in tables; it is not a general text model.
5A vendor says their model's large context window means you don't need retrieval. What should you ask?
Answer: B. Window size says how much text fits, not how well it is used. Research found models favor the start and end of long inputs, so test with the key fact in the middle, on your own data.
6A manager wants attention weight maps as the audit trail for an AI decision. What is the problem?
Answer: D. A model has many layers and many heads, and one weight map shows only where one head looked. A better check is to have the model cite the record it used and verify it against SAP.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A Python dictionary does an exact lookup. You give it a key, it finds the one matching entry and returns that entry's value.
Attention is a soft lookup. Each word brings a query ("what am I looking for?"). Every word offers a key ("what do I contain?") and a value ("what I pass on if you pick me"). The query is compared with every key. Better matches get bigger weights, the weights are made to add up to 1, and the result is the weighted average of the values.
Nothing is an exact match, and nothing is all or nothing. "It" might take 62% of its information from "order" and spread the rest thinly. Because every step is smooth arithmetic, gradient descent can learn the queries, keys and values. You saw that kind of training in Neural networks from scratch.
Before attention, translation models read a whole sentence into one fixed-length vector, then wrote the translation from that vector. Bahdanau, Cho and Bengio (ICLR 2015) argued that this single vector is a bottleneck, especially for long sentences. Their fix: while writing each output word, let the model search the source sentence for the relevant parts. It scores each source word, turns the scores into weights with a softmax, and takes a weighted sum. That is attention.
In 2017, "Attention Is All You Need" (Vaswani and colleagues) built a whole model from attention, with no recurrence. That model is the transformer. Its form of attention is the one every LLM uses today, and the one you will compute.
Project. Each word's vector x is multiplied by three learned matrices: q = x · W_q, k = x · W_k, v = x · W_v. Stack all words and you get the matrices Q, K and V.
Score. Take the dot product of one word's query with every word's key: Q Kᵀ. A high score means "this key matches what I'm looking for".
Scale. Divide by √d_k, the square root of the key length.
Mask (optional). Set the scores a word must not see to minus infinity.
Softmax. Turn each row of scores into weights that add up to 1. Minus infinity becomes a weight of exactly 0.
Blend. Multiply the weights by V: each word's output is a weighted average of the values.
flowchart LR
X[Word vectors] --> Q[Queries]
X --> K[Keys]
X --> V[Values]
Q --> S[Scores<br/>Q times K]
K --> S
S --> C[Divide by<br/>sqrt d_k]
C --> M[Mask<br/>optional]
M --> W[Softmax<br/>weights]
W --> O[Weighted sum<br/>of values]
V --> O
The paper's reason: if the numbers in queries and keys are random with mean 0 and variance 1, their dot product has variance d_k. Longer vectors give bigger scores. Big scores push softmax to put nearly all the weight on one word. There, its gradients are tiny, so training gets almost no signal about the other words. Dividing by √d_k brings the variance back to 1. Step 5 of the walkthrough measures this.
In the original transformer, attention is used in three ways:
Use
Who asks
Who answers
Mask
Encoder self-attention
Every input word
Every input word
None
Decoder self-attention
Every output word so far
Itself and earlier output words
Future positions set to minus infinity
Encoder-decoder attention
Every output word
Every input word
None
GPT-style models, which the rest of this unit builds, use only the second kind. It is called causal attention. While learning to predict the next word, a word must not see the words after it, or it could simply copy the answer. The mask enforces that.
This has a consequence you will see in Step 4. In "order blocked credit limit exceeded", the word "blocked" comes before its cause. Under a causal mask "blocked" can't look ahead to "credit limit". The information must flow to later words instead.
One head asks one kind of question. The transformer runs several heads side by side, each with its own W_q, W_k and W_v. The outputs are put side by side and mixed back with one more learned matrix, W_o. The base model in the paper used 8 heads of size 64, for a width of 512. In this topic's script, one head asks "which document?" and another asks "what caused this?".
Look at the formula again. Shuffle the words and each word gets the same weights, just in a different row. Attention on its own doesn't know which word came first. The transformer paper fixes this by adding position information to each word's vector before the first layer. The exercise uses a learned position vector for each slot. The next topic in this unit covers the options.
Every query meets every key. For n tokens of width d, the transformer paper puts self-attention at roughly n² · d operations per layer. Double the input and that part of the work roughly quadruples. In exchange, all positions are computed at once instead of one after another, which suits GPUs well.
This is why libraries ship fast, memory-saving versions. As of PyTorch 2.13, torch.nn.functional.scaled_dot_product_attention can dispatch to a FlashAttention-2 kernel, a memory-efficient kernel or a plain C++ version. All of them compute the same formula.
#Build it yourself: attention on an SAP-style sentence
You will run attention by hand on a clerk's note, "order blocked credit limit exceeded please release it". Every word gets six readable numbers instead of learned ones. Two hand-built heads decide where each word looks. You will switch the causal mask and the scaling on and off, then check your arithmetic against PyTorch.
flowchart LR
S[Clerk's note] --> E[Six numbers<br/>per word]
E --> H1[Head: which document?]
E --> H2[Head: what cause?]
H1 --> T[Weight tables]
H2 --> T
T --> P[Optional: check<br/>against PyTorch]
In VS Code's file list, right-click unit04, choose New File and name it attention.py.
Paste the code below and save with File > Save.
"""Unit 4: attention, worked by hand on a short SAP-style sentence.
Every word gets a small vector of six readable features. Two hand-built attention heads
then decide which earlier or later words each word should look at:
- the "refer" head: words like "it" and "release" look for the business document
- the "why" head: status words like "blocked" look for the cause
Nothing here is trained. The point is to see the arithmetic of attention, step by step.
How to run (from the unit04 folder, with the course .venv turned on):
python attention.py # both heads on the default sentence
python attention.py --head refer # one head only
python attention.py --causal # GPT-style: each word sees only earlier words
python attention.py --no-scale # skip the divide-by-square-root step
python attention.py --sentence "release it please" # your own words from the vocabulary
python attention.py --scale-demo # why the square-root scaling exists
python attention.py --check-torch # compare with PyTorch's built-in attention
"""
import argparse
import math
import numpy as np
FEATURES = ["document", "status", "cause", "pronoun", "action", "filler"]
# Each word as six numbers, one per feature above. Hand-made, so you can read them.
VOCAB = {
"order": [1.0, 0.0, 0.0, 0.0, 0.0, 0.0],
"invoice": [1.0, 0.0, 0.0, 0.0, 0.0, 0.0],
"blocked": [0.0, 1.0, 0.0, 0.0, 0.0, 0.0],
"credit": [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
"limit": [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
"price": [0.0, 0.0, 1.0, 0.0, 0.0, 0.0],
"exceeded": [0.0, 0.5, 0.5, 0.0, 0.0, 0.0],
"release": [0.0, 0.0, 0.0, 0.0, 1.0, 0.0],
"it": [0.0, 0.0, 0.0, 1.0, 0.0, 0.0],
"please": [0.0, 0.0, 0.0, 0.0, 0.0, 1.0],
"the": [0.0, 0.0, 0.0, 0.0, 0.0, 1.0],
}
DEFAULT_SENTENCE = "order blocked credit limit exceeded please release it"
D = len(FEATURES)
def head_weights(name: str):
"""Return the query, key and value matrices (each 6 x 6) for one hand-built head.
A word's query = its vector @ W_q, "what am I looking for".
A word's key = its vector @ W_k, "what do I offer".
A word's value = its vector @ W_v, "what I pass on if you pick me".
"""
w_q = np.zeros((D, D))
w_k = np.zeros((D, D))
if name == "refer":
w_q[FEATURES.index("pronoun"), 0] = 3.0 # "it" asks: where is a document?
w_q[FEATURES.index("action"), 0] = 2.0 # "release" asks that too: release what?
w_k[FEATURES.index("document"), 0] = 2.0 # documents answer
elif name == "why":
w_q[FEATURES.index("status"), 0] = 3.0 # status words ask: what is the cause?
w_k[FEATURES.index("cause"), 0] = 2.0 # cause words answer
else:
raise ValueError(name)
w_v = np.eye(D) # pass the word's features on unchanged
return w_q, w_k, w_v
def softmax(scores: np.ndarray) -> np.ndarray:
"""Turn each row of scores into weights that are positive and add up to 1."""
shifted = scores - scores.max(axis=-1, keepdims=True) # for numerical safety; same result
e = np.exp(shifted)
return e / e.sum(axis=-1, keepdims=True)
def attention(x, w_q, w_k, w_v, causal=False, scale=True):
"""Scaled dot-product attention for one head. x has one row per word."""
q, k, v = x @ w_q, x @ w_k, x @ w_v
scores = q @ k.T # how well each query matches each key
if scale:
scores = scores / math.sqrt(k.shape[-1]) # divide by the square root of the key size
if causal:
n = len(x)
future = np.triu(np.ones((n, n), dtype=bool), k=1)
scores = np.where(future, -np.inf, scores) # a word may not look at later words
weights = softmax(scores)
return weights @ v, weights
def print_weights(words, weights, title):
width = max(len(w) for w in words) + 1
print(f"\n{title} (row = the word asking, column = the word looked at; each row adds up to 1)")
print(" " * width + "".join(f"{w[:8]:>9}" for w in words))
for i, w in enumerate(words):
cells = "".join(f"{weights[i, j]:9.2f}" for j in range(len(words)))
print(f"{w:<{width}}{cells}")
print("Strongest link per word:")
for i, w in enumerate(words):
j = int(np.argmax(weights[i]))
visible = weights[i][weights[i] > 0]
if len(visible) == 1:
note = "(only itself is visible)"
elif np.allclose(visible, visible.max()):
note = "(spread evenly: no visible word matches what this word asks for)"
else:
note = ""
print(f" {w:<{width}} -> {words[j]:<10} {weights[i, j]:.2f} {note}")
def run_sentence(args):
words = args.sentence.lower().split()
unknown = [w for w in words if w not in VOCAB]
if unknown:
raise SystemExit(f"Unknown word(s): {', '.join(unknown)}. Use words from: {', '.join(VOCAB)}")
x = np.array([VOCAB[w] for w in words])
heads = ["refer", "why"] if args.head == "both" else [args.head]
outputs = []
for name in heads:
out, weights = attention(x, *head_weights(name), causal=args.causal, scale=not args.no_scale)
outputs.append(out)
mode = "causal" if args.causal else "full"
print_weights(words, weights, f"Head '{name}', {mode} attention, scaling {'off' if args.no_scale else 'on'}")
# Multi-head: put the heads' outputs side by side, then mix back to 6 numbers with W_o.
combined = np.concatenate(outputs, axis=-1)
w_o = np.vstack([np.eye(D)] * len(heads))
new_x = x + combined @ w_o # residual: keep the word, add what it gathered
print("\nEach word's features before -> after attention (heads combined, original word kept):")
print(f"{'':<10}" + "".join(f"{f:>10}" for f in FEATURES))
for i, w in enumerate(words):
print(f"{w:<10}" + "".join(f" {a:4.2f}>{b:4.2f}" for a, b in zip(x[i], new_x[i])))
def scale_demo():
rng = np.random.default_rng(0)
print("Random queries and keys with 8 candidate words, 2,000 trials per size.")
print(f"{'key size':>9} {'spread of scores':>17} {'top weight, unscaled':>21} {'top weight, scaled':>19}")
for d in [4, 16, 64, 256, 1024]:
q = rng.standard_normal((2000, 1, d))
k = rng.standard_normal((2000, 8, d))
s = (q @ k.transpose(0, 2, 1))[:, 0, :]
top_raw = softmax(s).max(axis=-1).mean()
top_scaled = softmax(s / math.sqrt(d)).max(axis=-1).mean()
print(f"{d:>9} {s.std():>17.1f} {top_raw:>21.2f} {top_scaled:>19.2f}")
print("Unscaled, the top weight heads toward 1.00 as vectors grow: one word takes everything,")
print("and training gets almost no signal about the others. Scaling keeps the spread steady.")
def check_torch(args):
try:
import torch
import torch.nn.functional as F
except ImportError:
raise SystemExit("PyTorch is not installed in this environment. See Set up for Unit 3.")
words = args.sentence.lower().split()
x = np.array([VOCAB[w] for w in words])
for name in ["refer", "why"]:
w_q, w_k, w_v = head_weights(name)
ours, _ = attention(x, w_q, w_k, w_v, causal=args.causal)
t = lambda a: torch.tensor(a, dtype=torch.float64)
theirs = F.scaled_dot_product_attention(t(x @ w_q), t(x @ w_k), t(x @ w_v), is_causal=args.causal)
diff = float(np.abs(ours - theirs.numpy()).max())
print(f"Head '{name}': largest difference from PyTorch = {diff:.1e} {'OK' if diff < 1e-9 else 'MISMATCH'}")
def main():
p = argparse.ArgumentParser(description="Attention, worked by hand.")
p.add_argument("--sentence", default=DEFAULT_SENTENCE)
p.add_argument("--head", choices=["refer", "why", "both"], default="both")
p.add_argument("--causal", action="store_true", help="each word sees only itself and earlier words")
p.add_argument("--no-scale", action="store_true", help="skip dividing by the square root of the key size")
p.add_argument("--scale-demo", action="store_true", help="show why scaling matters")
p.add_argument("--check-torch", action="store_true", help="compare with PyTorch's attention")
args = p.parse_args()
np.set_printoptions(precision=2, suppress=True)
if args.scale_demo:
scale_demo()
elif args.check_torch:
check_torch(args)
else:
run_sentence(args)
if __name__ == "__main__":
main()
Head 'refer', full attention, scaling on (row = the word asking, column = the word looked at; each row adds up to 1)
order blocked credit limit exceeded please release it
order 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
blocked 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
credit 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
limit 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
exceeded 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
please 0.12 0.12 0.12 0.12 0.12 0.12 0.12 0.12
release 0.42 0.08 0.08 0.08 0.08 0.08 0.08 0.08
it 0.62 0.05 0.05 0.05 0.05 0.05 0.05 0.05
Strongest link per word:
order -> order 0.12 (spread evenly: no visible word matches what this word asks for)
...
release -> order 0.42
it -> order 0.62
How to read it:
Row "it": 0.62 of its weight goes to "order". The query of "it" (pronoun feature × 3) matches the key of "order" (document feature × 2). The score is 3 × 2 = 6. Divided by √6 it is about 2.45. Softmax turns 2.45 against seven zeros into 0.62.
Row "release": a weaker query (2 instead of 3), so a weaker pull, 0.42.
Rows at 0.12: these words ask nothing in this head. All scores are 0, so softmax spreads the weight evenly: 1 ÷ 8 ≈ 0.12. That is a valid result, not an error.
Below the table, the "before -> after" block shows what attention did. "It" started with document = 0.00 and ends with 0.62. The word "it" now carries part of the order's meaning.
#Step 4: Run both heads, then switch on the causal mask
Run both heads:
python attention.py
The "why" head shows "blocked" putting 0.37 on "credit" and 0.37 on "limit". Each status word looks for its cause.
Now run the "why" head as a GPT would, with the causal mask:
python attention.py --head why --causal
What success looks like (first rows):
Head 'why', causal attention, scaling on (row = the word asking, column = the word looked at; each row adds up to 1)
order blocked credit limit exceeded please release it
order 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
blocked 0.50 0.50 0.00 0.00 0.00 0.00 0.00 0.00
credit 0.33 0.33 0.33 0.00 0.00 0.00 0.00 0.00
Everything above the diagonal is now 0.00. "Order" can only see itself. "Blocked" wants a cause but "credit limit" comes later, so it gets nothing useful. Its weight spreads over what it can see. "Exceeded" comes after "credit limit", so it can still find the cause. That is how causal models work: later words gather context from earlier ones.
The weight of "it" on "order" jumps from 0.62 to about 0.98. Without the division, scores are larger, and softmax is more all-or-nothing.
Run the experiment with random vectors of growing size:
python attention.py --scale-demo
What success looks like:
Random queries and keys with 8 candidate words, 2,000 trials per size.
key size spread of scores top weight, unscaled top weight, scaled
4 2.0 0.52 0.34
16 4.0 0.75 0.36
64 8.0 0.87 0.36
256 15.9 0.94 0.36
1024 31.8 0.97 0.36
The spread of scores doubles each time the key size grows four times: it grows with the square root of the size. Unscaled, the top weight creeps toward 1.00, so one word takes everything. Scaled, it stays near 0.36 at every size. That is the transformer paper's argument, measured.
#Step 6: Try your own sentence, and check against PyTorch
Use any words from the script's vocabulary (order, invoice, blocked, credit, limit, price, exceeded, release, it, please, the):
As of October 2026, SAP Learning describes the generative AI hub as the core component for accessing LLMs in SAP AI Core, with models from several providers and an orchestration service. The attention you computed runs inside those models on every call. You don't configure it. You control what it works on: the prompt, the data and the order you put them in. Unit 5 sets up access to the hub.
SAP-RPT-1 predicts values in tables from example rows sent with the request. SAP Learning lists sap-rpt-1-small (up to 2,048 rows and 100 columns) and sap-rpt-1-large (up to 65,536 rows and 256 columns). SAP's documentation for the service doesn't describe its architecture.
SAP's open-source sibling, sap-rpt-1-oss in the SAP-samples GitHub organization, implements the ConTextTab paper by SAP researchers. The paper's model alternates two kinds of self-attention layers:
Layer
Who attends to whom
Mask
Cross-column
Each cell attends to the other cells in its row
None
Cross-row
Each row attends to other rows in the same column
Rows may only attend to the provided example rows
The cross-row mask plays the same role as your causal mask: it stops a row from seeing what it must not see. Here, the rows to be predicted must not look at each other, only at the labeled examples. The repository says its checkpoints are for research use only and recommends about 80 GB of GPU memory. It is a reference for how the idea works, not something to run in this course.
If you ever train your own transformer in SAP AI Core, the attention layer is the same PyTorch function you checked in Step 6. Unit 10 covers running AI workloads on BTP.
Cost grows with prompt length. Attention work grows roughly with the square of the length per layer, so long prompts add compute and delay. Measure tokens per request early and test at real volume. Unit 10 covers cost and routing.
Position matters. Research on long inputs found the middle is used least reliably, with the same pattern in a longer-context model. Put instructions and the key record where they are easy to use, and test with the key fact moved around.
Everything in the prompt can influence everything else. Self-attention lets every token draw on every other token. Text from a customer email, a document or a tool result can therefore steer the output as much as your instructions can. This is the root of prompt injection (Unit 11). Keep untrusted text clearly separated and never let it grant permissions.
Only send what the user may see. Attention will happily use any data you include. Filter SAP data by the user's authorizations before it reaches the prompt (Unit 7 and Unit 11).
Weights aren't explanations. One head's weight map is a debugging aid, not evidence for an auditor. For business traceability, have the model cite the record it used and check it.
Use the library function. For your own models, call scaled_dot_product_attention rather than a hand-written version. It can pick fast, memory-saving kernels and computes the same result.
Forgetting the scaling. Without dividing by √d_k, softmax saturates as vectors grow and training stalls. Step 5 shows it.
A mask the wrong way round. In PyTorch's scaled_dot_product_attention, a boolean attn_mask value of True means the position takes part. Other APIs use the opposite convention. Check the documentation of the function you call.
Masking after softmax. The mask must set scores to minus infinity before softmax. Zeroing weights afterwards leaves rows that no longer add up to 1.
Softmax over the wrong axis. Weights must add up to 1 across the keys (the last axis), one row per query. Summing down the columns gives nonsense that still runs.
Leaving out positions. Without position information, "order blocks credit" and "credit blocks order" look the same to attention.
Reading one head as the model's reason. Real models have many heads in many layers. Don't build a business explanation on one weight map.
#Exercise: train one attention head and watch it learn where to look
So far you built the heads by hand. Now PyTorch learns one. The task: read a row of eight made-up material codes and, at the last position, name the code that sat at a chosen position. Only attention can solve it, because the answer is far from where it is asked for. The Head class you save here is the building block for the tiny GPT later in this unit.
flowchart LR
R[Row of 8 codes] --> E[Code + position<br/>vectors]
E --> H[One causal<br/>attention head]
H --> O[Score for<br/>each code]
O --> A[Accuracy and<br/>attention map]
Before you start: the same setup as above. PyTorch must be installed (Set up for Unit 3). Training takes a few seconds on a laptop CPU.
In VS Code, create unit04/head.py, paste the code below and save.
"""Unit 4 exercise: one trainable attention head in PyTorch, checked and then trained.
Part 1 checks the head: it must match PyTorch's built-in attention, and in causal mode
changing a later token must never change an earlier output.
Part 2 trains a tiny model on a lookup task: read a row of made-up material codes and,
at the last position, say which code sat at position --target-pos. Only attention can
solve it, because the answer is far away from where it is asked for.
How to run (from the unit04 folder, with the course .venv turned on):
python head.py # checks, then training with the answer at position 0
python head.py --target-pos 3 # move the answer; watch the attention follow it
python head.py --steps 15 # too little training; accuracy stays low
python head.py --check-only # only Part 1
"""
import argparse
import torch
import torch.nn as nn
import torch.nn.functional as F
class Head(nn.Module):
"""One causal self-attention head: queries, keys and values are learned linear maps."""
def __init__(self, n_embd: int, head_size: int):
super().__init__()
self.query = nn.Linear(n_embd, head_size, bias=False)
self.key = nn.Linear(n_embd, head_size, bias=False)
self.value = nn.Linear(n_embd, head_size, bias=False)
self.last_weights = None # kept so we can print what the head looked at
def forward(self, x: torch.Tensor) -> torch.Tensor:
# x has shape (batch, positions, n_embd)
q, k, v = self.query(x), self.key(x), self.value(x)
scores = q @ k.transpose(-2, -1) / k.shape[-1] ** 0.5 # (batch, positions, positions)
t = x.shape[1]
future = torch.triu(torch.ones(t, t, dtype=torch.bool, device=x.device), diagonal=1)
scores = scores.masked_fill(future, float("-inf")) # no peeking at later positions
weights = F.softmax(scores, dim=-1)
self.last_weights = weights.detach()
return weights @ v
class LookupModel(nn.Module):
"""Token embedding + position embedding -> one attention head -> a score for every code."""
def __init__(self, vocab: int, length: int, n_embd: int = 32, head_size: int = 32):
super().__init__()
self.tok = nn.Embedding(vocab, n_embd)
self.pos = nn.Embedding(length, n_embd)
self.head = Head(n_embd, head_size)
self.out = nn.Linear(head_size, vocab)
def forward(self, idx: torch.Tensor) -> torch.Tensor:
positions = torch.arange(idx.shape[1], device=idx.device)
x = self.tok(idx) + self.pos(positions)
return self.out(self.head(x))
def run_checks() -> bool:
torch.manual_seed(0)
head = Head(n_embd=8, head_size=4).double()
x = torch.randn(2, 5, 8, dtype=torch.float64)
ours = head(x)
q, k, v = head.query(x), head.key(x), head.value(x)
theirs = F.scaled_dot_product_attention(q, k, v, is_causal=True)
diff = (ours - theirs).abs().max().item()
ok1 = diff < 1e-9
print(f"Check 1, matches PyTorch's attention: largest difference {diff:.1e} {'OK' if ok1 else 'FAIL'}")
x2 = x.clone()
x2[:, -1, :] = torch.randn(2, 8, dtype=torch.float64) # change only the last position
change = (head(x2)[:, :-1] - ours[:, :-1]).abs().max().item()
ok2 = change < 1e-12
print(f"Check 2, later tokens can't change earlier outputs: change {change:.1e} {'OK' if ok2 else 'FAIL'}")
rows = head.last_weights.sum(dim=-1)
ok3 = torch.allclose(rows, torch.ones_like(rows))
print(f"Check 3, every row of weights adds up to 1: {'OK' if ok3 else 'FAIL'}")
return ok1 and ok2 and ok3
def make_batch(n: int, vocab: int, length: int, target_pos: int):
idx = torch.randint(0, vocab, (n, length))
return idx, idx[:, target_pos]
def train(args) -> None:
torch.manual_seed(args.seed)
model = LookupModel(args.vocab, args.length)
opt = torch.optim.AdamW(model.parameters(), lr=3e-3)
print(f"\nTraining: {args.length} codes per row, answer at position {args.target_pos}, {args.steps} steps")
for step in range(1, args.steps + 1):
idx, target = make_batch(64, args.vocab, args.length, args.target_pos)
logits = model(idx)[:, -1, :] # only the last position answers
loss = F.cross_entropy(logits, target)
opt.zero_grad()
loss.backward()
opt.step()
if step % max(1, args.steps // 5) == 0:
print(f" step {step:>4} loss {loss.item():.3f}")
model.eval()
with torch.no_grad():
idx, target = make_batch(1000, args.vocab, args.length, args.target_pos)
acc = (model(idx)[:, -1, :].argmax(-1) == target).float().mean().item()
avg = model.head.last_weights[:, -1, :].mean(dim=0)
print(f"Accuracy on 1,000 new rows: {acc:.2f} (guessing would give about {1 / args.vocab:.2f})")
print("Where the last position looks, averaged over those rows:")
for p, w in enumerate(avg.tolist()):
print(f" position {p}: {w:.2f} {'#' * round(w * 40)}")
peak = int(avg.argmax())
print(f"Peak attention at position {peak}; the answer was at position {args.target_pos}.")
def main() -> None:
p = argparse.ArgumentParser(description="Check and train one attention head.")
p.add_argument("--target-pos", type=int, default=0, help="where the answer sits in each row")
p.add_argument("--length", type=int, default=8, help="codes per row")
p.add_argument("--vocab", type=int, default=20, help="number of different material codes")
p.add_argument("--steps", type=int, default=600)
p.add_argument("--seed", type=int, default=0)
p.add_argument("--check-only", action="store_true")
args = p.parse_args()
if not 0 <= args.target_pos < args.length:
raise SystemExit(f"--target-pos must be between 0 and {args.length - 1}.")
passed = run_checks()
if not passed:
raise SystemExit("A check failed. Compare your Head class with the published code.")
if not args.check_only:
train(args)
if __name__ == "__main__":
main()
In the terminal, inside unit04, run:
python head.py
What success looks like (loss values can differ slightly between computers):
Check 1, matches PyTorch's attention: largest difference 1.1e-16 OK
Check 2, later tokens can't change earlier outputs: change 0.0e+00 OK
Check 3, every row of weights adds up to 1: OK
Training: 8 codes per row, answer at position 0, 600 steps
step 120 loss 0.012
...
Accuracy on 1,000 new rows: 1.00 (guessing would give about 0.05)
Where the last position looks, averaged over those rows:
position 0: 1.00 ########################################
position 1: 0.00
...
Peak attention at position 0; the answer was at position 0.
Move the answer and run again:
python head.py --target-pos 3
The attention bar should jump to position 3. Nobody told the head where to look. It learned it from the loss.
Train far too little:
python head.py --steps 15
In our run, accuracy was 0.44 while 0.74 of the attention already sat on the right position. The head learns where to look first; the output layer still has to learn what to say.
Create unit04/attention_notes.md with three short lines: the accuracy and peak position from step 2, the same from step 3, and one sentence on what step 4 showed.
Save your work:
cd ..
git add unit04/head.py unit04/attention_notes.md
git commit -m "Train one attention head on a lookup task"
Part of head.py
What it does
Head
One causal self-attention head with learned query, key and value matrices; keeps its last weights for printing
LookupModel
Code vector plus position vector, one head, then a layer that scores all 20 codes
run_checks
Compares with PyTorch's function, proves later tokens can't change earlier outputs, checks rows add up to 1
make_batch
Random rows of codes; the answer is the code at --target-pos
train
AdamW training on the last position only, then accuracy and the average attention map on 1,000 new rows
If something goes wrong, the table in Build it yourself applies. One more case: --target-pos must be between 0 and 7 means you asked for a position outside the row; use 0 to 7, or change --length.
Done when:python head.py --target-pos 3 prints three OK checks, accuracy of at least 0.95 and peak attention at position 3, and head.py and attention_notes.md are committed in Git.
Pick one answer for each question. The explanation appears after you choose.
1In scaled dot-product attention, what do the query, key and value each do?
Answer: B. Each word's query is compared with every key, the scores become weights through softmax, and the output is the weighted average of the values. It is a soft lookup, not an exact match.
2Why does the formula divide the scores by the square root of the key size?
Answer: C. For random vectors the dot product's variance grows with the key size, so larger vectors give larger scores. Softmax then saturates and training gets tiny gradients; dividing by the square root keeps the spread steady, as the scale demo showed.
3In the causal run, why does "blocked" fail to find "credit limit"?
Answer: A. The causal mask sets every score above the diagonal to minus infinity, so those weights become 0. "Blocked" can only see "order" and itself; "exceeded", which comes later, can still reach the cause.
4You write your own mask and set the forbidden weights to 0 after the softmax. What goes wrong?
Answer: D. The mask must set scores to minus infinity before softmax, so the remaining weights are renormalized to add up to 1. Zeroing weights afterwards leaves rows that sum to less than 1.
5What does attn_mask with the value True mean in PyTorch's scaled_dot_product_attention?
Answer: B. PyTorch's documentation says a boolean True means the element should take part in attention. Other APIs use the opposite convention, which is a common source of silent bugs.
6A colleague doubles the length of every prompt to include more SAP documents. What happens to attention's work per layer?
Answer: C. Self-attention compares each query with every key, about n squared times d operations per layer. Twice the tokens means about four times that work, which shows up as cost and delay.
7Your trained head reaches 0.44 accuracy with most attention on the right position. What does that tell you?
Answer: D. In the short run, 0.74 of the attention already sat on the answer's position while accuracy was 0.44. More training lets the output layer turn the gathered information into the right code.
8Text from a customer email in the prompt makes your assistant ignore its instructions. Why is that possible?
Answer: A. Attention doesn't know which text is trusted; it blends information across everything in the context. That is the root of prompt injection, so untrusted text must be kept separate and never allowed to grant permissions.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Attention Is All You Need (Vaswani et al., 2017), arXiv HTML version 7— scaled dot-product attention formula; scaling by 1/sqrt(d_k) because dot products grow with d_k and push softmax into tiny gradients; multi-head attention with h=8 and d_k=64 in the base model; masking future positions to minus infinity; self-attention costs O(n²·d) per layer but O(1) sequential steps; positional encodings
Introduction to SAP-RPT-1 (SAP Learning)— foundation model for relational and structured data; sap-rpt-1-small (up to 2,048 rows, 100 columns) and sap-rpt-1-large (up to 65,536 rows, 256 columns); classification and regression; architecture not described
SAP-samples/sap-rpt-1-oss (GitHub)— open-source model implementing the ConTextTab paper (NeurIPS 2025), formerly named ConTextTab; checkpoints for research purposes only; about 80 GB of GPU memory recommended