Orchestrate

The transformer architecture

How attention, feed-forward layers, shortcuts and normalization stack into the transformer, the three shapes it comes in, and how to build and check one block yourself.

Updated Oct 1, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

The transformer is the design inside today's large language models (LLMs). It looks complicated in diagrams, but it is one simple unit, the block, repeated many times.

Think of a sales order note passing through a shared services team. Each word of the note has its own working file. One block does two things to those files:

  1. A meeting. Every word checks the other words and picks up what matters to it. This is attention, covered in Attention, explained. "It" learns that it means "the order".
  2. Desk work. Each word then works alone on what it gathered, using knowledge the model learned in training. This is the feed-forward step.

Two habits keep the process stable. Each step adds notes to the file instead of rewriting it, so nothing earlier gets lost. And before each step, the file is tidied to a standard scale, so no number grows out of control.

Stack the same block 12, 24 or 48 times. Put a step at the bottom that turns words into numbers and marks their position. Put a step at the top that scores every possible next word. That is the whole architecture.

Why it matters to the business

You will never design a transformer for an SAP project. But knowing its shape helps you make three decisions.

Pick the right shape for the job. Transformers come in three shapes. One reads a whole text at once and is good at understanding and sorting it. One writes text a word at a time and powers chat assistants. One reads one text and writes another, as in translation. Routing 10,000 customer emails to the right order-to-cash queue is a reading job. A chat model can do it, but a smaller reading model may do it faster and cheaper. Ask which shape a proposal uses and why.

Size drives cost. A model's size is mostly how many blocks it stacks and how wide each block is. The GPT-2 paper from 2019 listed four sizes, from 117 million learned numbers (parameters) in 12 blocks to over 1.5 billion in 48 blocks. Every parameter must sit in memory while the model runs. Bigger models need bigger hardware and cost more per request. Whether the quality gain is worth it is a measurement, not a given.

The same design works on tables. SAP's table model, SAP-RPT-1, is named "Relational Pretrained Transformer". It uses transformer blocks on rows and columns of business data, not on sentences. For predictions on SAP tables, such as which sales orders will be delivered late, a table model may fit better than a text model.

A concrete order-to-cash example: blocked sales orders. A reading model sorts incoming customer notes by topic. A writing model drafts the reply to the customer. A table model predicts which open orders are likely to be blocked next week. Three jobs, three shapes, one underlying design.

How SAP does it

SAP doesn't ask customers to build transformers. You meet them inside the models SAP offers, as of October 2026:

  • Generative AI hub. SAP Learning describes the generative AI hub in SAP AI Core as the core component for accessing LLMs. It offers models from several providers through one interface, managed in SAP AI Launchpad. You choose a model, not its architecture. Vendors of the largest chat models publish few internal details. Unit 5 sets up access.
  • SAP-RPT-1. SAP Learning names it "Relational Pretrained Transformer", a foundation model for structured business data. It handles classification and regression. It learns from example rows sent with each request, so there is no separate training step. It comes in a small and a large version with different row and column limits.
  • The research behind it. SAP researchers published the ConTextTab design, whose open-source model is SAP-RPT-1's sibling. Its blocks alternate between looking across the columns of a row and looking across rows. It uses no position markers at all. Column names do that job, because the order of rows and columns in a table carries no meaning.

So "transformer" isn't only a chat technology. In SAP's portfolio it reads text, writes text and predicts values in tables.

Three shapes of transformer

Shape What each word can see Typical job Well-known example SAP-flavored use
Encoder (reader) Every word, before and after it Understand, sort, compare, search BERT Turn product texts into vectors for search; route customer notes to a queue
Decoder (writer) Only the words before it Write text one word at a time GPT family Draft a reply about a blocked order; summarize a dispute case
Encoder-decoder (reader plus writer) The writer sees all of the reader's input, and its own earlier words Turn one text into another The original 2017 transformer, built for translation Translate a supplier's message before it reaches procure-to-pay

The GPT family, which made chat assistants widely known, is decoder-only. For a specific model in the generative AI hub, check the vendor's own documentation rather than assuming.

Questions to ask

  • Is this a reading job, a writing job or a table job? Which model shape does the proposal use, and why that one?
  • Have we tested a smaller model against the large one on our own data, with the quality and cost of each written down?
  • How many parameters does the model have, and what hardware or per-request price does that imply at our volume?
  • For predictions on SAP tables, have we compared a table model such as SAP-RPT-1 with a text model?
  • If we host an open model ourselves, who maintains it, and how do we update it when the vendor releases a new version?
  • Does the vendor publish how the model was built and evaluated, and what does that mean for our audit needs?

Common misconceptions

  • "A transformer is a chatbot." It is a general design. Readers like BERT sort and search text, and SAP-RPT-1 applies it to tables.
  • "More layers always means better answers." Depth adds capacity and cost. Whether it improves your task is something to measure.
  • "The model reads words in order, like a person." Attention looks at all words at once. Order has to be added, usually as position information at the bottom of the stack.
  • "Most of the model is attention." In the standard design, the feed-forward parts hold about two thirds of each block's parameters. Attention is the famous part, not the biggest one.
  • "An LLM is a different technology from the AI in SAP's table model." SAP-RPT-1 is a transformer too. The blocks are similar; the data and the training differ.

Key terms

  • Transformer: a model design made of identical blocks stacked on top of each other.
  • Block (layer): one unit of the stack, made of an attention step and a feed-forward step.
  • Feed-forward network: a small network applied to each word on its own, after attention.
  • Residual connection (shortcut): each step adds its result to the word's vector instead of replacing it.
  • Layer normalization: rescaling a word's numbers to a standard range before each step.
  • Parameter: one learned number in the model; model size is counted in parameters.
  • Encoder / decoder: a block stack that sees the whole text, or one that sees only earlier words.
  • Position information: numbers added at the bottom of the stack so the model knows word order.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In plain words, what is a transformer made of?

    Answer: B. A transformer repeats one block: attention, where words share information, then a feed-forward step, where each word works alone. Stacking 12, 24 or 48 of these blocks, with embeddings below and scoring on top, gives the whole model.
  2. 2A team wants a chat model to sort 10,000 customer emails into order-to-cash queues. What should you ask first?

    Answer: C. Sorting is a reading job, which encoder-style models are built for. A chat model can do it too, but a smaller reader may be faster and cheaper. Ask which shape fits the job and compare on your own data.
  3. 3Why does the number of parameters matter to a project budget?

    Answer: A. Model size, mostly blocks times width, decides how much memory and computing each request needs. That shows up as hardware or per-request cost. Whether a bigger model is worth it is a measurement, not a given.
  4. 4What does the name SAP-RPT-1 tell you about the model?

    Answer: D. SAP Learning spells out RPT as Relational Pretrained Transformer. It applies transformer blocks to rows and columns for classification and regression, learning from example rows sent with each request.
  5. 5A vendor says their model is better because it has more layers. What is the right response?

    Answer: B. More blocks give a model more capacity and also make it more expensive to run. Whether that improves your task is something to test with your own data and compare with the cost.
  6. 6Which statement about the three shapes of transformer is correct?

    Answer: D. An encoder lets every word see every other word, which suits understanding and sorting. A decoder lets each word see only the words before it, which suits writing one word at a time. Encoder-decoder models combine both, as in translation.
Deep layer · 40 min read

Mental model: a shared working file, edited in turns

Give every token a vector, a list of numbers. That vector travels up through the stack, and nothing in the stack ever replaces it. Each step reads it, computes something, and adds its result back.

There are only two kinds of steps, and they alternate:

  • Attention moves information between tokens. "It" reads from "order".
  • Feed-forward works within each token. Every token runs the same small network on its own vector.

Before each step, a layer norm rescales the vector so its numbers stay in a steady range. That's the whole block. Everything else (embeddings at the bottom, scores at the top, encoders versus decoders) is arrangement around it.

This running vector is often called the residual stream. Keep the picture of one file per token, with a meeting and desk work taking turns to add notes. It explains why blocks can be stacked, why depth needs the shortcuts, and where the parameters sit.

How it works

The full stack in one picture

flowchart TB
  T[Token ids] --> E[Token embedding<br/>+ position]
  E --> B1[Block 1]
  B1 --> B2[Block 2]
  B2 --> BN[... Block N]
  BN --> LN[Final layer norm]
  LN --> H[Score for every<br/>vocabulary token]

And inside each block, in the GPT-2 arrangement this topic builds:

flowchart TB
  X[x in] --> N1[Layer norm]
  N1 --> A[Multi-head attention]
  A --> P1((+))
  X --> P1
  P1 --> N2[Layer norm]
  N2 --> F[Feed-forward]
  F --> P2((+))
  P1 --> P2
  P2 --> Y[x out]

The two + circles are the shortcuts. Every block takes a tensor of shape (batch, tokens, width) and returns the same shape. That is what makes stacking possible.

Embeddings and positions

The bottom of the stack turns token ids into vectors with an embedding table: one learned row of numbers per vocabulary entry. You used embeddings in Embeddings and semantic similarity.

Attention by itself ignores order, as the previous topic showed. So the transformer adds position information to each token's vector. The 2017 paper used fixed sine and cosine waves of different frequencies. It also tried learned position vectors, one per slot, and reported nearly identical results. GPT-2 uses learned positions, and so does this topic's code.

Other schemes exist. Rotary position embedding (RoPE), from the RoFormer paper, rotates queries and keys by an angle that depends on position. That way, the attention score carries the relative distance between two tokens. Models differ here, so check a model's own documentation for the scheme it uses.

Two surprises are worth knowing:

  • A causal mask leaks order. A 2022 study found causal language models with no position information at all were still competitive. Its explanation: with a causal mask, each token can infer how many tokens came before it. You will see this in the exercise.
  • Sometimes order should be ignored. SAP's ConTextTab design uses no explicit position embeddings. Column headers do that job, so shuffling the rows or columns of a table doesn't change the prediction.

Multi-head attention inside the block

This is the attention from the previous topic, with every head computed at once. One linear layer produces queries, keys and values for all heads together. The width is split into equal slices, one per head. Each head runs scaled dot-product attention on its slice. The slices are put side by side again, and one more matrix (proj, called W_o before) mixes them.

The width must divide evenly by the number of heads. GPT-2's smallest model and BERT-BASE both use width 768 with 12 heads, so each head works on 64 numbers.

The feed-forward network

After attention, every position runs through the same two-layer network:

FFN(x) = max(0, x·W1 + b1)·W2 + b2

That is the 2017 formula: widen, apply an activation, narrow again. The paper used width 512 inside the stream and 2,048 inside the feed-forward, four times wider. The max(0, ...) is the ReLU activation from Neural networks from scratch. GPT-style models commonly use GELU, a smooth relative of ReLU. This topic's code uses GELU.

"Position-wise" means no information moves between tokens here. Attention gathers, feed-forward processes. Because of the four-times widening, this is where most of a block's parameters sit.

Shortcuts: residual connections

Each step's output is added to its input: x = x + step(x). The 2017 paper took the idea from residual networks for image recognition, which made very deep networks trainable.

Shortcuts help in two ways. First, a block that has learned nothing useful yet adds roughly zero and passes the input through unchanged, so extra depth does little harm at the start. Second, during training the learning signal flows back through the additions without being squeezed by every layer. Step 5 of the walkthrough shows what happens without them. GPT-2 also starts the layers that write into the shortcut path with smaller random values, scaled by 1/√N for N residual layers, to account for the build-up along the path. This topic's code copies that.

Layer normalization, and where it goes

Layer norm rescales each token's vector to an average of 0 and a spread of 1, then applies a learned scale and shift. It works per token, so it behaves the same in training and in use.

Where it sits matters:

Arrangement Formula Used by Trade-off
Post-LN x = norm(x + step(x)) The 2017 paper Xiong et al. showed large gradients near the output at the start, so training needs a learning-rate warm-up
Pre-LN x = x + step(norm(x)) GPT-2 and many later models Better-behaved gradients; Xiong et al. trained without warm-up, faster, with comparable results

GPT-2 moved the norm to the input of each step and added one final norm after the last block. That is the arrangement in this topic's code. PyTorch's nn.TransformerEncoderLayer supports both. Its default is norm_first=False, the post-LN of 2017. You will set norm_first=True to compare it with your block.

The top: final norm, scores and weight tying

After the last block, a final layer norm and one linear layer turn each token's vector into a score for every vocabulary entry. Those scores become next-token probabilities, covered in the last topic of this unit.

The 2017 paper shares one weight matrix between the embedding table and this final layer. This is called weight tying: the same table turns tokens into vectors at the bottom and vectors back into token scores at the top. For GPT-2's smallest model, that one table is about 38.6 million of the 124 million parameters you will count.

Encoder, decoder, encoder-decoder

The block is the same in all three shapes. What changes is the mask and the wiring.

  • Encoder (BERT): no causal mask, so every token sees every token. BERT trains by hiding 15% of input tokens and predicting them. BERT-BASE has 12 layers, width 768, 12 heads and 110 million parameters.
  • Decoder (GPT): a causal mask, so each token sees only itself and earlier tokens. The BERT paper describes this as GPT's constrained self-attention, where every token attends only to its left. That lets the model learn to predict the next token at every position at once.
  • Encoder-decoder (the 2017 transformer): the decoder blocks get a third step between attention and feed-forward. It attends over the encoder's output, often called cross-attention. The paper stacked 6 encoder and 6 decoder layers.

Where the parameters are

For width d and a feed-forward four times wider, one block holds:

Part Parameters GPT-2 smallest, d = 768
Attention: queries, keys, values (3d × d plus biases) and proj (d × d plus biases) about 4d² 2,362,368
Feed-forward: up (d × 4d) and down (4d × d) plus biases about 8d² 4,722,432
Two layer norms (scale and shift each) 4d 3,072

The number of heads doesn't change the count; it only changes how the width is sliced. The feed-forward part is about two thirds of every block. Step 4 counts the whole GPT-2-sized model, layer by layer.

Build it yourself: assemble a transformer block and stack it

You will build a decoder block from the parts above, stack two of them into a tiny model, and push an SAP-style sentence through it. Then you will prove your block matches PyTorch's own layer, count where GPT-2's parameters sit, and see what deep stacks do without shortcuts. Nothing here is trained yet; the exercise trains it.

Before you start: complete Set up your computer for this course, Set up for Unit 3 and Set up for Unit 4. They give you the orchestrate-course folder with its .venv, PyTorch and the unit04 folder. Attention, explained is the topic before this one; its head.py is useful background but not required.

flowchart LR
  S[SAP-style<br/>sentence] --> M[Tiny model:<br/>2 blocks]
  M --> T[Shape trace]
  B[One block] --> C[Check against<br/>PyTorch]
  G[GPT-2 sizes] --> P[Parameter count]
  D[Deep stacks] --> R[With and without<br/>shortcuts]

What you need

  • The course folder and .venv from the setup topics, with PyTorch installed.
  • About 40 minutes. No accounts, no API keys, no cost.
  • No internet. Nothing is sent anywhere.

Step 1: Open the course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Go into the Unit 4 folder, creating it if needed:

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit04
      cd unit04
    • macOS / Linux:

      mkdir -p unit04
      cd unit04

Step 2: Save the script

  1. In VS Code's file list, right-click unit04, choose New File and name it transformer_block.py.
  2. Paste the code below and save with File > Save. The exercise and the next topic import from this file, so keep the name exactly.
"""Unit 4: one transformer block, built from parts you can read, then stacked.

The block is the decoder ("GPT-style") kind with the norm placed first (pre-LN):
    x = x + attention(norm(x))      # tokens share information
    x = x + feed_forward(norm(x))   # each token works on what it gathered
Stack the same block several times, put embeddings below and a scoring layer on top,
and you have the whole architecture of a GPT-style language model.

How to run (from the unit04 folder, with the course .venv turned on):
    python transformer_block.py                      # trace one SAP-style sentence through a tiny model
    python transformer_block.py --check-torch        # compare one block with PyTorch's built-in layer
    python transformer_block.py --count gpt2-small   # where 124 million parameters sit
    python transformer_block.py --count tiny         # the same count for the tiny model
    python transformer_block.py --depth-demo         # why the "x +" shortcuts matter in deep stacks
"""
import argparse

import torch
import torch.nn as nn
import torch.nn.functional as F


class MultiHeadAttention(nn.Module):
    """Several attention heads side by side, then one matrix (proj) that mixes their outputs."""

    def __init__(self, n_embd: int, n_head: int, causal: bool = True):
        super().__init__()
        if n_embd % n_head != 0:
            raise ValueError("n_embd must be divisible by n_head")
        self.n_head, self.causal = n_head, causal
        self.qkv = nn.Linear(n_embd, 3 * n_embd)      # queries, keys and values for all heads at once
        self.proj = nn.Linear(n_embd, n_embd)         # W_o: mixes the heads back together
        self.last_weights = None

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        b, t, c = x.shape
        q, k, v = self.qkv(x).split(c, dim=-1)
        # (b, t, c) -> (b, heads, t, head_size): each head gets its own slice of the width
        q, k, v = (z.view(b, t, self.n_head, c // self.n_head).transpose(1, 2) for z in (q, k, v))
        scores = q @ k.transpose(-2, -1) / k.shape[-1] ** 0.5
        if self.causal:
            future = torch.triu(torch.ones(t, t, dtype=torch.bool, device=x.device), diagonal=1)
            scores = scores.masked_fill(future, float("-inf"))
        weights = F.softmax(scores, dim=-1)
        self.last_weights = weights.detach()
        out = (weights @ v).transpose(1, 2).reshape(b, t, c)   # put the heads side by side again
        return self.proj(out)


class FeedForward(nn.Module):
    """The same small two-layer network applied to every position on its own."""

    def __init__(self, n_embd: int, expand: int = 4):
        super().__init__()
        self.up = nn.Linear(n_embd, expand * n_embd)
        self.act = nn.GELU()
        self.down = nn.Linear(expand * n_embd, n_embd)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        return self.down(self.act(self.up(x)))


class Block(nn.Module):
    """One transformer block: attention, then feed-forward, each with a norm and a shortcut."""

    def __init__(self, n_embd: int, n_head: int, residual: bool = True, norm: bool = True,
                 causal: bool = True):
        super().__init__()
        self.residual = residual
        self.ln1 = nn.LayerNorm(n_embd) if norm else nn.Identity()
        self.attn = MultiHeadAttention(n_embd, n_head, causal)
        self.ln2 = nn.LayerNorm(n_embd) if norm else nn.Identity()
        self.ffn = FeedForward(n_embd)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        if self.residual:
            x = x + self.attn(self.ln1(x))
            x = x + self.ffn(self.ln2(x))
        else:
            x = self.attn(self.ln1(x))
            x = self.ffn(self.ln2(x))
        return x


class TinyTransformer(nn.Module):
    """Token and position embeddings -> n_layer blocks -> final norm -> a score for every token."""

    def __init__(self, vocab: int, context: int, n_embd: int = 32, n_head: int = 4, n_layer: int = 2,
                 residual: bool = True, norm: bool = True, causal: bool = True, positions: bool = True):
        super().__init__()
        self.tok = nn.Embedding(vocab, n_embd)
        self.pos = nn.Embedding(context, n_embd) if positions else None
        self.blocks = nn.ModuleList(Block(n_embd, n_head, residual, norm, causal) for _ in range(n_layer))
        self.ln_f = nn.LayerNorm(n_embd) if norm else nn.Identity()
        self.head = nn.Linear(n_embd, vocab, bias=False)
        self.head.weight = self.tok.weight             # weight tying: one matrix for words in and out
        self.apply(self._init)
        for block in self.blocks:          # GPT-2's tweak: shrink the layers that write into the shortcut path
            for layer in (block.attn.proj, block.ffn.down):
                nn.init.normal_(layer.weight, mean=0.0, std=0.02 / (2 * n_layer) ** 0.5)

    @staticmethod
    def _init(module: nn.Module) -> None:
        # Small random starting values, a common choice for GPT-style models.
        if isinstance(module, (nn.Linear, nn.Embedding)):
            nn.init.normal_(module.weight, mean=0.0, std=0.02)
        if isinstance(module, nn.Linear) and module.bias is not None:
            nn.init.zeros_(module.bias)

    def forward(self, idx: torch.Tensor, trace: bool = False) -> torch.Tensor:
        x = self.tok(idx)
        if self.pos is not None:
            x = x + self.pos(torch.arange(idx.shape[1], device=idx.device))
        if trace:
            print(f"  {'token ids':<34} {tuple(idx.shape)}")
            print(f"  {'after embeddings + positions':<34} {tuple(x.shape)}")
        for i, block in enumerate(self.blocks, start=1):
            x = block(x)
            if trace:
                print(f"  {f'after block {i}':<34} {tuple(x.shape)}")
        x = self.ln_f(x)
        logits = self.head(x)
        if trace:
            print(f"  {'scores for every vocabulary word':<34} {tuple(logits.shape)}")
        return logits


PRESETS = {
    # name: (vocab, context, n_embd, n_head, n_layer)
    "tiny": (14, 16, 32, 4, 2),
    # GPT-2's smallest model: vocabulary 50,257, context 1,024, width 768, 12 blocks.
    # The head count (12 here) slices the width; it doesn't change the parameter count.
    "gpt2-small": (50257, 1024, 768, 12, 12),
}


def count(preset: str) -> None:
    vocab, context, n_embd, n_head, n_layer = PRESETS[preset]
    with torch.device("meta"):                   # shapes only, no memory used for the numbers
        model = TinyTransformer(vocab, context, n_embd, n_head, n_layer)

    def n(module: nn.Module) -> int:
        return sum(p.numel() for p in module.parameters())

    block = model.blocks[0]
    rows = [
        ("token embeddings (shared with output)", n(model.tok)),
        ("position embeddings", n(model.pos)),
        (f"attention, per block x {n_layer}", n(block.attn)),
        (f"feed-forward, per block x {n_layer}", n(block.ffn)),
        (f"layer norms, per block x {n_layer}", n(block.ln1) + n(block.ln2)),
        ("final layer norm", n(model.ln_f)),
    ]
    total = n(model)                              # tied weights are counted once
    print(f"Preset '{preset}': vocab {vocab:,}, context {context:,}, width {n_embd}, "
          f"{n_head} heads, {n_layer} blocks")
    for name, value in rows:
        print(f"  {name:<40} {value:>13,}")
    all_blocks = n_layer * n(block)
    print(f"  {'all blocks together':<40} {all_blocks:>13,}  ({all_blocks / total:.0%} of the total)")
    print(f"  {'total':<40} {total:>13,}")
    ffn_share = n(block.ffn) / n(block)
    print(f"Inside each block, the feed-forward part holds {ffn_share:.0%} of the parameters.")


SENTENCE = "order 4711 blocked credit limit exceeded release it"


def trace() -> None:
    torch.manual_seed(0)
    words = SENTENCE.split()
    vocab = {w: i for i, w in enumerate(sorted(set(words)))}
    idx = torch.tensor([[vocab[w] for w in words]])
    vocab_size, context, n_embd, n_head, n_layer = PRESETS["tiny"]
    model = TinyTransformer(vocab_size, context, n_embd, n_head, n_layer).eval()
    print(f"Sentence: \"{SENTENCE}\"  ({len(words)} tokens)")
    print(f"Tiny model: width {n_embd}, {n_head} heads, {n_layer} blocks. Shapes are (batch, tokens, numbers):")
    with torch.no_grad():
        logits = model(idx, trace=True)
    print("Every block takes and returns the same shape, which is why blocks can be stacked.")
    w = model.blocks[0].attn.last_weights[0]      # (heads, tokens, tokens)
    print(f"Block 1 has {w.shape[0]} heads; each makes its own {w.shape[1]}x{w.shape[2]} weight table.")
    probs = F.softmax(logits[0, -1], dim=-1)
    print(f"Untrained, the model's guess for the word after 'it' is spread thin: "
          f"top probability {probs.max().item():.2f} (even spread would be {1 / vocab_size:.2f}).")


def check_torch() -> bool:
    torch.manual_seed(0)
    n_embd, n_head, t = 16, 4, 6
    ours = Block(n_embd, n_head).double().eval()
    ref = nn.TransformerEncoderLayer(n_embd, n_head, dim_feedforward=4 * n_embd, dropout=0.0,
                                     activation="gelu", batch_first=True, norm_first=True).double().eval()
    with torch.no_grad():                         # copy our weights into PyTorch's layer
        ref.self_attn.in_proj_weight.copy_(ours.attn.qkv.weight)
        ref.self_attn.in_proj_bias.copy_(ours.attn.qkv.bias)
        ref.self_attn.out_proj.weight.copy_(ours.attn.proj.weight)
        ref.self_attn.out_proj.bias.copy_(ours.attn.proj.bias)
        ref.linear1.weight.copy_(ours.ffn.up.weight)
        ref.linear1.bias.copy_(ours.ffn.up.bias)
        ref.linear2.weight.copy_(ours.ffn.down.weight)
        ref.linear2.bias.copy_(ours.ffn.down.bias)
        for a, b in ((ref.norm1, ours.ln1), (ref.norm2, ours.ln2)):
            a.weight.copy_(b.weight)
            a.bias.copy_(b.bias)
        x = torch.randn(2, t, n_embd, dtype=torch.float64)
        mask = nn.Transformer.generate_square_subsequent_mask(t, dtype=torch.float64)
        diff = (ours(x) - ref(x, src_mask=mask, is_causal=True)).abs().max().item()
        x2 = x.clone()
        x2[:, -1] = torch.randn(2, n_embd, dtype=torch.float64)
        leak = (ours(x2)[:, :-1] - ours(x)[:, :-1]).abs().max().item()
    ok1, ok2 = diff < 1e-10, leak < 1e-12
    print(f"Check 1, matches nn.TransformerEncoderLayer (norm_first=True): largest difference {diff:.1e}  "
          f"{'OK' if ok1 else 'FAIL'}")
    print(f"Check 2, a later token can't change earlier outputs: change {leak:.1e}  {'OK' if ok2 else 'FAIL'}")
    return ok1 and ok2


def depth_demo() -> None:
    print("What reaches the top of an untrained stack? 16 tokens, width 32.")
    print("  alike = average similarity of the tokens' vectors (low = still different, 1.00 = all the same)")
    print("  size  = average length of a token's vector (near 0 = the signal has faded away)")
    print(f"{'blocks':>7} | {'with shortcuts':^21} | {'without shortcuts':^21}")
    print(f"{'':>7} | {'alike':>9} {'size':>11} | {'alike':>9} {'size':>11}")
    vocab, context = 14, 16
    for n_layer in (2, 8, 32):
        cells = []
        for residual in (True, False):
            torch.manual_seed(0)
            model = TinyTransformer(vocab, context, n_embd=32, n_head=4, n_layer=n_layer, residual=residual)
            idx = torch.randint(0, vocab, (8, context))
            with torch.no_grad():
                x = model.tok(idx) + model.pos(torch.arange(context))
                for block in model.blocks:
                    x = block(x)
                size = x.norm(dim=-1).mean().item()
                unit = F.normalize(x, dim=-1)                       # length 1, so dot product = cosine
                alike = (unit @ unit.transpose(1, 2)).mean().item()
            cells.append(f"{alike:>9.2f} {size:>11.1e}" if size > 1e-12 else f"{'-':>9} {size:>11.1e}")
        print(f"{n_layer:>7} | {cells[0]} | {cells[1]}")


def main() -> None:
    p = argparse.ArgumentParser(description="Build, check and count a transformer block.")
    p.add_argument("--check-torch", action="store_true", help="compare one block with PyTorch's layer")
    p.add_argument("--count", choices=sorted(PRESETS), help="count parameters for a model size")
    p.add_argument("--depth-demo", action="store_true", help="deep stacks with and without shortcuts")
    args = p.parse_args()
    if args.check_torch:
        if not check_torch():
            raise SystemExit("A check failed. Compare your code with the published script.")
    elif args.count:
        count(args.count)
    elif args.depth_demo:
        depth_demo()
    else:
        trace()


if __name__ == "__main__":
    main()

Step 3: Trace a sentence through the stack

Run the script with no options:

python transformer_block.py

What success looks like:

Sentence: "order 4711 blocked credit limit exceeded release it"  (8 tokens)
Tiny model: width 32, 4 heads, 2 blocks. Shapes are (batch, tokens, numbers):
  token ids                          (1, 8)
  after embeddings + positions       (1, 8, 32)
  after block 1                      (1, 8, 32)
  after block 2                      (1, 8, 32)
  scores for every vocabulary word   (1, 8, 14)
Every block takes and returns the same shape, which is why blocks can be stacked.
Block 1 has 4 heads; each makes its own 8x8 weight table.
Untrained, the model's guess for the word after 'it' is spread thin: top probability 0.11 (even spread would be 0.07).

How to read it:

  • (1, 8): one sentence of eight token ids. Here each word is one token; real models split words into smaller pieces.
  • (1, 8, 32): after the embedding step, every token is 32 numbers. Both blocks keep exactly that shape. You could insert a third block, or a hundredth, without changing anything else.
  • (1, 8, 14): at the top, every position gets one score per vocabulary slot. The tiny model has 14 slots; only 8 are used by this sentence.
  • Top probability 0.11: the model is untrained, so its guess is barely better than an even spread. The small random starting values keep it that way on purpose. Training (the exercise and the next topic) changes the weights, not the shapes.

Step 4: Check your block against PyTorch's

python transformer_block.py --check-torch

What success looks like:

Check 1, matches nn.TransformerEncoderLayer (norm_first=True): largest difference 4.4e-16  OK
Check 2, a later token can't change earlier outputs: change 0.0e+00  OK

The script builds PyTorch's nn.TransformerEncoderLayer with norm_first=True, GELU, no dropout, and a causal mask. It copies your block's weights into it and feeds both the same random input. A difference around 1e-16 is rounding noise: your block computes the same thing as PyTorch's. Check 2 changes only the last token and confirms no earlier output moved. That is the causal mask doing its job.

The name "encoder layer" is PyTorch's. With a causal mask and no cross-attention, it computes exactly a GPT-style decoder block.

Step 5: Count where GPT-2's parameters sit

python transformer_block.py --count gpt2-small

What success looks like:

Preset 'gpt2-small': vocab 50,257, context 1,024, width 768, 12 heads, 12 blocks
  token embeddings (shared with output)       38,597,376
  position embeddings                            786,432
  attention, per block x 12                    2,362,368
  feed-forward, per block x 12                 4,722,432
  layer norms, per block x 12                      3,072
  final layer norm                                 1,536
  all blocks together                         85,054,464  (68% of the total)
  total                                      124,439,808
Inside each block, the feed-forward part holds 67% of the parameters.

Read three things from it:

  1. Feed-forward beats attention two to one in every block, as the table in "How it works" predicted.
  2. The embedding table is big. With 50,257 vocabulary entries, it holds almost a third of this small model. Weight tying saves another table of the same size.
  3. The total is about 124 million. The GPT-2 paper's table lists 117 million for its smallest model with the same layers and width. The paper doesn't show its counting, so treat its table as approximate. Counting the weights of the published configuration, as you just did, gives about 124 million.

Each parameter stored as a 32-bit number takes 4 bytes, so 124 million parameters need about 0.5 GB of memory just to hold the weights. Repeat the arithmetic for a model with billions of parameters and you see why model size is a hardware question.

Then run python transformer_block.py --count tiny for comparison. In the tiny model, the blocks are 96% of everything, because its vocabulary has only 14 entries.

Step 6: See why deep stacks need the shortcuts

python transformer_block.py --depth-demo

What success looks like:

What reaches the top of an untrained stack? 16 tokens, width 32.
  alike = average similarity of the tokens' vectors (low = still different, 1.00 = all the same)
  size  = average length of a token's vector (near 0 = the signal has faded away)
 blocks |    with shortcuts     |   without shortcuts
        |     alike        size |     alike        size
      2 |      0.11     1.8e-01 |      0.91     2.9e-02
      8 |      0.16     1.7e-01 |      1.00     3.0e-03
     32 |      0.13     1.6e-01 |         -     1.1e-23

With shortcuts, the 16 tokens still look different at the top of a 32-block stack, and their vectors keep a steady size. Without shortcuts, after 8 blocks every token points the same way (1.00). After 32 blocks, almost nothing is left at all (1.1e-23; the - means it is too small to compare). A model whose tokens all look identical at the top can't tell "order" from "release". The exercise shows what that does to training.

Step 7: Save your work in Git

From the course folder:

cd ..
git add unit04/transformer_block.py
git commit -m "Build and check a transformer block"

What each part of the script does

Part What it does
MultiHeadAttention One linear layer makes queries, keys and values for all heads; splits the width into heads; scaled dot-product attention with an optional causal mask; proj mixes the heads
FeedForward Widen four times, GELU, narrow again; the same network for every position
Block Pre-LN arrangement: x + attn(norm(x)), then x + ffn(norm(x)); options switch off shortcuts, norms or the mask
TinyTransformer Token and position embeddings, a stack of blocks, a final norm and a tied output layer; small random starting values, with GPT-2's smaller start for the layers that write into the shortcut path
count Builds a model on the meta device and adds up parameters per part
trace Pushes one sentence through the tiny model and prints the shape after each stage
check_torch Copies the block's weights into nn.TransformerEncoderLayer(norm_first=True) and compares outputs; checks the causal mask
depth_demo Untrained stacks of 2, 8 and 32 blocks, with and without shortcuts

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Step 1 of Set up your computer, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'torch' PyTorch isn't in the Python you're using Check for (.venv) in the prompt. If it's there, follow Set up for Unit 3, Step 2
pip shows ProxyError, SSLError or Could not fetch URL Your network or company proxy blocks the package sites Try another network, or ask IT for access or an internal mirror. The script itself needs no network
AttributeError: module 'torch' has no attribute 'device' or a meta error with --count Your PyTorch is very old Run pip install --upgrade torch inside the .venv, as in the Unit 3 setup
n_embd must be divisible by n_head You changed the width or head count to numbers that don't divide Keep the width a multiple of the number of heads, such as 32 and 4
Check 1 says FAIL Your block differs from the published code Compare Block.forward line by line; the order is norm, step, then add
can't open file ... transformer_block.py The terminal isn't in unit04, or the file has another name Run cd unit04 from the course folder, and check the file name
Asked for an API key Nothing in this topic uses a key or account Check you are running transformer_block.py, not another script

Where this shows up in SAP

Unit 4 is a mechanics unit, so this section is short.

Inside the models of the generative AI hub

As of October 2026, SAP Learning describes the generative AI hub as the core component for accessing LLMs in SAP AI Core, with models from several providers managed through SAP AI Launchpad. Every block you built runs inside those models on every call, many times over. You don't configure the architecture. You choose the model, and the model's size and shape decide its cost and speed. Unit 5 sets up access and covers choosing a model.

SAP-RPT-1: transformer blocks on tables

SAP Learning names SAP-RPT-1 "Relational Pretrained Transformer". It is SAP's foundation model for structured data, for classification and regression. It learns from example rows in each request, with no separate training step. It is consumed through SAP AI Core.

SAP's documentation for the service doesn't describe its internals. The ConTextTab paper, by SAP researchers, describes the design behind the open-source sibling model. Its blocks look familiar:

Your block ConTextTab's blocks
Attention across the tokens of a sentence, with a causal mask Attention across the columns of one row (no mask), alternating with attention across rows (masked so rows attend only to the provided example rows)
Self-attention, then a feed-forward network Self-attention, then a feed-forward MLP
Learned position embeddings No position embeddings; column headers act as position information
GPT-2 smallest: 12 blocks, width 768 Base variant: 12 layers, width 768

The same block, rewired for a different kind of data. Unit 12 covers table AI and SAP-RPT-1 in depth.

Your own transformer code

If you train your own model one day in SAP AI Core, the blocks are the PyTorch code you wrote. PyTorch's documentation notes that its nn.Transformer layers implement the original architecture with limited features, and points to building blocks and ecosystem libraries for newer designs. That is one reason this topic builds the block from parts. Unit 10 covers running AI workloads on BTP.

Build vs. SAP

Situation Your own block (this topic) PyTorch's nn.TransformerEncoderLayer A model through SAP
Learning how it works Best: every line is visible Use it to check your work Not visible
Training a small custom model Good: easy to change, as in the exercise Fine for the original design; limited for newer ones Not applicable
Text tasks in an SAP process No No Best: LLMs through the generative AI hub
Predictions on SAP tables No No SAP-RPT-1, with example rows in the request
Embeddings for search No No An open embedding model as in Unit 3; SAP's options are covered in Unit 7

Production concerns

  • Size is cost. Parameters must sit in memory, and every token passes through every block. Measure quality against cost on your own data before choosing the bigger model. Unit 10 covers cost and model routing.
  • Shape is fit. Don't default to a chat model for a sorting or search job. Encoders and table models are often cheaper and easier to evaluate for those jobs.
  • Position limits are real. A model with learned positions, like GPT-2, has a fixed number of position slots (1,024 for GPT-2). Text beyond the context must be cut, split or retrieved. Unit 7 covers retrieval.
  • Architecture details are often private. Vendors of large chat models rarely publish layers, widths or training data. Base your evaluation on measured behavior with your data, and record which model version you tested.
  • Data and authorizations come first. No architecture choice changes what the model is allowed to see. Filter SAP data by the user's authorizations before it reaches any model (Units 7 and 11).
  • Self-hosted models are your responsibility. If you run an open model yourself, you own patching, scaling, monitoring and updates. The generative AI hub moves much of that work to SAP and the model providers.

Pitfalls

  • Forgetting the shortcuts. Without x +, deep stacks lose the signal, as Step 6 showed. Training then stalls.
  • Mixing up the norm placement. Post-LN and pre-LN train differently. PyTorch's layers default to post-LN (norm_first=False). Set it on purpose, and expect post-LN to need a learning-rate warm-up.
  • A width that doesn't divide by the heads. The width is sliced into equal heads. 768 with 12 heads works; 768 with 10 doesn't.
  • Wrong mask for the shape. A decoder without a causal mask can see the answer it should predict, so training looks perfect and generation fails. An encoder with a causal mask loses half its context.
  • Assuming positions are optional. An encoder without position information sees a bag of words. The exercise shows the effect.
  • Counting tied weights twice. With weight tying, the input and output tables are one matrix. Count it once.
  • Treating untrained output as meaningful. A freshly built model produces confident-looking shapes and meaningless numbers. Only training gives the weights meaning.

Exercise: switch parts off and see what breaks

Now you train the tiny transformer and break it on purpose. Each row is a made-up event log for one sales order: eight events such as created, credit_check or picked. Each row holds exactly one goods_issue and one invoice. A made-up rule for this exercise: if the invoice comes before the goods issue, the row is an exception, otherwise ok. The words in a row never tell the answer; only their order does. That makes it a good test of positions, masks and shortcuts.

flowchart LR
  R[Event log<br/>8 events + ?] --> M[Tiny transformer<br/>from Step 2]
  M --> A[exception or ok]
  A --> V[Accuracy on<br/>2,000 new rows]
  V --> N[Notes:<br/>what broke and why]

Before you start: finish Build it yourself above. order_check.py imports the classes from transformer_block.py, so both files must be in unit04. Each run takes 5 to 20 seconds on a laptop CPU.

  1. In VS Code, create unit04/order_check.py, paste the code below and save.

    """Unit 4 exercise: switch parts of a transformer off and see what breaks.
    
    Each row is a made-up event log for one sales order: eight events such as "created",
    "credit_check" and "picked", with exactly one "goods_issue" and one "invoice" somewhere.
    The model reads the row and, at a final "?" token, answers "exception" if the invoice
    came before the goods issue, otherwise "ok". (A rule made up for this exercise.)
    The two words are the same in every row; only their ORDER decides the answer.
    
    How to run (from the unit04 folder, with the course .venv turned on):
        python order_check.py                                # full model: 2 blocks, causal, positions on
        python order_check.py --encoder --no-positions       # sees all words at once, but no order
        python order_check.py --no-positions                 # causal mask on, positions off
        python order_check.py --layers 8 --no-shortcuts      # deep stack without the "x +" shortcuts
        python order_check.py --layers 8                     # the same depth with shortcuts
    """
    import argparse
    import time
    
    import torch
    import torch.nn.functional as F
    
    from transformer_block import TinyTransformer
    
    FILLER = ["created", "changed", "credit_check", "released", "picked", "packed"]
    WORDS = FILLER + ["goods_issue", "invoice", "?", "exception", "ok"]
    ID = {w: i for i, w in enumerate(WORDS)}
    EVENTS = 8                                       # events per row, then the "?" token
    
    
    def make_batch(n: int):
        rows = torch.randint(0, len(FILLER), (n, EVENTS))
        for i in range(n):                           # place one goods_issue and one invoice at random
            a, b = torch.randperm(EVENTS)[:2].tolist()
            rows[i, a], rows[i, b] = ID["goods_issue"], ID["invoice"]
        invoice_first = (rows == ID["invoice"]).int().argmax(1) < (rows == ID["goods_issue"]).int().argmax(1)
        x = torch.cat([rows, torch.full((n, 1), ID["?"])], dim=1)
        y = torch.where(invoice_first, ID["exception"], ID["ok"])
        return x, y
    
    
    def main() -> None:
        p = argparse.ArgumentParser(description="Train a tiny transformer to spot invoice-before-goods-issue.")
        p.add_argument("--layers", type=int, default=2, help="number of transformer blocks")
        p.add_argument("--encoder", action="store_true", help="no causal mask: every word sees every word")
        p.add_argument("--no-positions", action="store_true", help="leave out the position embeddings")
        p.add_argument("--no-shortcuts", action="store_true", help="leave out the residual 'x +' shortcuts")
        p.add_argument("--steps", type=int, default=800)
        p.add_argument("--seed", type=int, default=0)
        args = p.parse_args()
    
        torch.manual_seed(args.seed)
        model = TinyTransformer(len(WORDS), EVENTS + 1, n_embd=32, n_head=4, n_layer=args.layers,
                                residual=not args.no_shortcuts, causal=not args.encoder,
                                positions=not args.no_positions)
        opt = torch.optim.AdamW(model.parameters(), lr=3e-3)
        setup = (f"{args.layers} blocks, {'encoder (no mask)' if args.encoder else 'decoder (causal mask)'}, "
                 f"positions {'off' if args.no_positions else 'on'}, shortcuts {'off' if args.no_shortcuts else 'on'}")
        print(f"Training: {setup}, {args.steps} steps")
        start = time.time()
        for step in range(1, args.steps + 1):
            x, y = make_batch(64)
            loss = F.cross_entropy(model(x)[:, -1], y)    # only the "?" position answers
            opt.zero_grad()
            loss.backward()
            opt.step()
            if step % max(1, args.steps // 4) == 0:
                print(f"  step {step:>4}  loss {loss.item():.3f}")
    
        model.eval()
        with torch.no_grad():
            x, y = make_batch(2000)
            pred = model(x)[:, -1].argmax(-1)
            acc = (pred == y).float().mean().item()
        print(f"Accuracy on 2,000 new rows: {acc:.2f}  (guessing gives about 0.50)  [{time.time() - start:.0f} s]")
        print("Three of those rows:")
        for i in range(3):
            events = " ".join(WORDS[t] for t in x[i, :-1].tolist())
            print(f"  {events}\n    -> model says {WORDS[pred[i]]}, correct is {WORDS[y[i]]}")
    
    
    if __name__ == "__main__":
        main()
  2. In the terminal, inside unit04, train the full model:

    python order_check.py

    What success looks like (loss values and example rows can differ between computers):

    Training: 2 blocks, decoder (causal mask), positions on, shortcuts on, 800 steps
      step  200  loss 0.004
      step  400  loss 0.001
      step  600  loss 0.001
      step  800  loss 0.000
    Accuracy on 2,000 new rows: 1.00  (guessing gives about 0.50)  [5 s]
    Three of those rows:
      packed picked changed invoice picked created credit_check goods_issue
        -> model says exception, correct is exception
      ...
  3. Take away order. Run an encoder (no causal mask) without positions, then the same encoder with positions:

    python order_check.py --encoder --no-positions
    python order_check.py --encoder

    In our runs, the first stayed at 0.48 (guessing) and the second reached 1.00. Without positions and without a mask, the model sees a bag of words. Every row holds one invoice and one goods issue, so it can't do better than a coin toss.

  4. Now keep the causal mask but drop the positions:

    python order_check.py --no-positions

    In our run this reached 1.00. The causal mask leaks order: a token can tell which events came before it. That matches the 2022 finding on causal models without position encodings from "How it works".

  5. Go deep, without and with shortcuts:

    python order_check.py --layers 8 --no-shortcuts
    python order_check.py --layers 8

    In our runs, 8 blocks without shortcuts stayed at 0.52 and the loss never left about 0.69, the value for pure guessing between two answers. With shortcuts, 8 blocks reached 1.00. Step 6 of the walkthrough showed why: without shortcuts, the tokens look the same at the top.

  6. Create unit04/transformer_notes.md with a small table: one row per command from steps 2 to 5, with the accuracy you got. Add one sentence for each of these: why the encoder without positions failed, why the decoder without positions didn't, and why depth without shortcuts failed.

  7. Save your work:

    cd ..
    git add unit04/order_check.py unit04/transformer_notes.md
    git commit -m "Break a tiny transformer on purpose: positions, masks and shortcuts"
Part of order_check.py What it does
FILLER, WORDS, ID The event vocabulary, plus the ? token and the two answers
make_batch Random event logs with one goods_issue and one invoice; the label depends only on their order
Options --encoder, --no-positions, --no-shortcuts, --layers Pass straight into TinyTransformer from transformer_block.py
Training loop AdamW, only the ? position answers, loss printed four times
Evaluation Accuracy on 2,000 new rows and three example rows

If something goes wrong, the table in Build it yourself applies. Two more cases. ModuleNotFoundError: No module named 'transformer_block' means the two files aren't in the same folder, or the terminal isn't in unit04. If the full model stays near 0.50 on your computer, the random start can matter: run python order_check.py --seed 1 and note it in your file.

Keep transformer_block.py. The next topic, Build a tiny GPT, trains these same classes to write text.

Done when: transformer_block.py --check-torch prints two OK lines, order_check.py reaches at least 0.95 accuracy in its default run, transformer_notes.md holds the four-step table with your three explanations, and all three files are committed in Git.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What do the two steps in every transformer block do?

    Answer: B. Attention lets each token gather from others, which is the "meeting". The feed-forward network then runs the same small network on each position on its own, which is the "desk work". Mixing them up leads to wrong expectations about what each part can learn.
  2. 2In the script, why can TinyTransformer stack any number of Block objects without other changes?

    Answer: C. The trace in Step 3 shows (1, 8, 32) before and after every block. Attention and feed-forward both return vectors of the input width, and the shortcuts add them to the input. So blocks plug into each other like identical parts.
  3. 3What is the difference between post-LN and pre-LN?

    Answer: A. The 2017 paper computed norm(x + step(x)). GPT-2 moved the norm to x + step(norm(x)). Xiong et al. showed post-LN has large gradients near the output at the start and needs a warm-up, while pre-LN trained without it.
  4. 4Your count shows GPT-2's smallest model at about 124 million parameters. Where are most of a block's parameters?

    Answer: D. The feed-forward part has about 8d² parameters per block against about 4d² for attention, so it holds roughly two thirds. The mask has no parameters at all, and attention's pairwise cost is computing work, not stored weights.
  5. 5You run the exercise as an encoder without positions and accuracy stays at 0.48. What explains it?

    Answer: B. Without a causal mask, every token sees every token, and without positions nothing marks order. Every row holds one invoice and one goods issue, so the input carries no information about the answer. Adding positions or a causal mask fixes it.
  6. 6A colleague says a causal model must have position embeddings to learn word order. Is that right?

    Answer: C. A 2022 study found causal language models without position encodings were still competitive, and explained it by the mask: each token sees a different number of predecessors. Step 4 of the exercise shows the same effect.
  7. 7You build an 8-block model and forget the x + shortcuts. What do you expect, and why?

    Answer: D. Step 6 showed tokens becoming identical after 8 blocks without shortcuts, and fading to almost nothing after 32. In the exercise, 8 blocks without shortcuts stayed at guessing level while 8 with shortcuts reached 1.00.
  8. 8A business team wants predictions on open sales orders from a table of past orders. Which SAP option matches the architecture idea best?

    Answer: C. SAP Learning names SAP-RPT-1 a Relational Pretrained Transformer for classification and regression on structured data, learning from example rows sent with the request. Its research design uses attention across columns and across rows, which fits tables better than text.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in