Orchestrate

Set up for Unit 4: PyTorch and an optional free GPU

Confirm PyTorch is ready for transformers, measure how long a tiny GPT takes to train on your own computer, and know when a free GPU notebook is worth it.

Updated Oct 1, 2026Foundational 6 minDeep 35 min
Foundational layer · 6 min read

The 60-second version

Unit 4 opens up large language models (LLMs). Learners study attention, the idea that lets a model decide which earlier words matter, then the transformer design built around it. Then they train a tiny GPT on their own computer and watch it learn to write text one character at a time.

This setup adds no new software. Unit 3 already installed PyTorch, the library that does the training. What this setup adds is a measurement: a short script times a tiny GPT-sized model on the learner's own laptop and estimates how long a full training run would take.

That number decides one thing: train on the laptop, or use a free GPU notebook in the browser. A GPU (graphics processing unit) is a chip that does many sums at once, which speeds up neural network training. This course uses Google Colab as the optional free GPU path. As of October 2026 it is free to use, but Google doesn't guarantee a GPU will be available.

Setup takes 20 to 40 minutes, most of it waiting for timing runs.

Why it matters to the business

Leaders hear "we need GPUs" in almost every AI conversation. This setup gives a learner first-hand evidence of what that means. The same small model can take minutes on one machine and hours on another. Model size and hardware drive time, and time drives cost.

That experience carries straight into SAP projects. A team asked to "fine-tune our own model on our sales order notes" should first ask how big the model is, how long training takes and on what hardware. Most enterprise teams never train an LLM from scratch. They call models someone else trained, which Unit 5 covers. Training a tiny one here is about understanding, not about building your own production model.

Two points for a leader:

  • The free GPU is a learning tool, not a platform. Colab's own FAQ says its free resources are "not guaranteed and not unlimited", and limits change without notice.
  • Company data stays out. Everything in Unit 4 uses made-up text. Nothing from your SAP systems belongs in a personal notebook service.

How SAP does it

SAP customers rarely train LLMs themselves. As of October 2026, SAP's generative AI hub, part of SAP AI Core and SAP AI Launchpad, gives access to ready-trained foundation models from several providers. Unit 5 sets that up.

When a team does need to run its own model, SAP AI Core can run it on GPU hardware. A setting called a resource plan chooses the machine size; SAP's tutorial for a GPU model uses the plan infer.s to get a GPU node. That needs a paid BTP setup, so this course shows it as a sketch later, not as a step here.

So Unit 4 sits underneath SAP's offering. It explains what happens inside the models that SAP's services call.

Laptop, local GPU or free notebook

Option What it is Cost Good for
Laptop CPU The normal processor, with the PyTorch you installed in Unit 3 Free The small tiny-GPT size; every other Unit 4 topic
Apple silicon GPU The GPU in M-series Macs, which PyTorch can use directly Free Faster training on a Mac, if PyTorch detects it
NVIDIA GPU on your PC A gaming or workstation graphics card Free if you have one; needs PyTorch's GPU build Faster training; most learners won't have one
Google Colab (free) A notebook in the browser on Google's machines, GPU when available Free, with limits that change The larger tiny-GPT size, or a slow or Intel-based laptop

Most learners need only the first row. The timing script tells each learner whether the others are worth the effort.

Time and money

  • Time: 20 to 40 minutes. The larger timing test can take several minutes on a slow laptop.
  • Money: nothing. Colab's free tier needs a Google account but no payment.
  • Disk space: nothing new. PyTorch is already installed.
  • Honest run times: in our test on a small two-core cloud machine, a full run of the small model was estimated at about 8 minutes. The larger model was estimated at about 9 hours. A recent laptop is usually faster than our test machine, and a GPU much faster again. Each learner measures their own.

Questions to ask IT

  • May learners use Google Colab, and with a personal or work Google account?
  • Does the company network allow colab.research.google.com?
  • Do any team laptops have an NVIDIA GPU or an Apple silicon chip that learners could use?
  • Can learners keep a laptop awake and plugged in for an hour of training?
  • If the company has an SAP AI Core contract, is there a sandbox where GPU resource plans could be shown later?

Common misconceptions

  • "You can't train a GPT without a GPU." You can train a tiny one on a laptop CPU. It is slower, and the model is small, but the mechanics are the same.
  • "A free GPU notebook is a free GPU server." No. Colab's free tier has changing limits, can give you no GPU, times out when idle and deletes its machine afterwards.
  • "A GPU makes everything faster." For very small models the overhead of moving data to the GPU can eat the gain. That is why the setup measures instead of assuming.
  • "Training our own LLM is the normal enterprise path." Most SAP AI work calls existing models through services such as the generative AI hub. Training here is for understanding.

Key terms

  • Attention: the step in a language model that decides which earlier words matter for the next one.
  • Transformer: the model design behind today's LLMs, built from repeated attention layers.
  • GPU: a chip that runs many calculations at once; it speeds up neural network training.
  • CPU: the computer's normal processor.
  • Training step: one round of "predict, measure the error, adjust the weights".
  • Google Colab: a free, browser-based notebook service from Google, with optional GPUs when available.
  • Resource plan: in SAP AI Core, the setting that picks the machine size, including GPU nodes.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does Unit 4's setup add to a learner's computer?

    Answer: B. PyTorch was installed in Unit 3. Unit 4's setup only measures how fast a tiny GPT-sized model trains on the learner's computer, to decide between the laptop and a free GPU notebook.
  2. 2What decision does the timing result drive?

    Answer: C. The script estimates how long a full training run would take. If it is short, the laptop is fine; if it runs into hours, the larger model is better trained in a free GPU notebook such as Colab.
  3. 3A team member plans to rely on free Colab GPUs for a weekly business job. What is the risk?

    Answer: D. Colab's FAQ says free resources are not guaranteed, limits fluctuate, idle runtimes time out and the virtual machines are deleted. It is fine for learning, not for business processes.
  4. 4How do most SAP customers use large language models?

    Answer: A. SAP's generative AI hub gives access to foundation models from several providers. Training a tiny GPT in this course is for understanding what those models do, not a production path.
  5. 5Why does the course measure speed instead of telling everyone to use a GPU?

    Answer: B. The same small model can take minutes on one machine and hours on another. For very small models, moving data to a GPU can cost as much as it saves, so measuring is the honest way to decide.
  6. 6What should a leader ask IT before learners start Unit 4?

    Answer: D. The optional GPU path uses Colab, so IT should confirm it is allowed and reachable, and which Google account to use. Knowing which laptops already have a usable GPU avoids needless spending.
Deep layer · 35 min read

Mental model: measure first, then pick where to train

Unit 4's code runs in the same course folder, the same .venv and with the same PyTorch you installed in Unit 3. The only open question is where training runs. A timing script answers it with a number from your own machine.

flowchart LR
  V[.venv with PyTorch<br/>from Unit 3] --> B[bench_tiny_gpt.py]
  B --> E{Estimated<br/>training time}
  E -->|minutes| L[Train on your laptop]
  E -->|hours| C[Free GPU notebook<br/>Google Colab]
  B --> N[timing_notes.md]
  N --> K[check_unit04.py]

How it works

Where the time goes in training

Training a GPT repeats one training step thousands of times. Each step takes a batch of text snippets, predicts the next character at every position, measures the error and adjusts every weight. The work per step grows with:

  • Layers and width: more and wider layers mean more weights to compute and adjust.
  • Context length: how many characters the model looks back over. Attention compares every position with every earlier one, so doubling the context roughly quadruples that part of the work.
  • Batch size: how many snippets go through at once.

Total time is roughly time per step × number of steps. The timing script measures the first and multiplies by a typical step count. It uses 5,000 steps, the number in a well-known small character-level GPT setup (nanoGPT's Shakespeare example).

CPU, CUDA and MPS

PyTorch calls the hardware it runs on a device:

Device name What it is How PyTorch checks for it
cpu The normal processor; always there Always available
cuda An NVIDIA GPU, with PyTorch's GPU build torch.cuda.is_available()
mps The GPU in Apple silicon Macs, through Apple's Metal torch.backends.mps.is_available()

Code picks a device once, then moves the model and data to it with .to(device). That is the "device-agnostic" pattern in PyTorch's documentation, and every Unit 4 script uses it.

PyTorch's MPS documentation lists macOS 14 or newer and an MPS-capable Mac as requirements. On Windows and Linux, the CPU build from Unit 3 has no cuda support, even if the computer has an NVIDIA card. That is fine: Colab covers the GPU case without reinstalling anything.

Why timing GPUs needs care

GPUs work asynchronously: Python hands them a job and moves on before the job finishes. PyTorch's CUDA documentation warns that time measurements without synchronizing are not accurate. The timing script therefore calls torch.cuda.synchronize() (or torch.mps.synchronize() on a Mac) before reading the clock. It also skips three warm-up steps, because the first steps include one-off setup work.

What Google Colab gives you

As of October 2026, Colab's FAQ describes the free tier like this:

Rule What it means for you
Free of charge, with dynamic usage limits Limits are not published and change over time
GPUs not guaranteed Some days you may get no GPU; use the laptop path then
Notebooks run at most 12 hours Long enough for the larger tiny GPT, with room to spare
Idle runtimes time out Keep the browser tab open while training
Virtual machines are deleted Files you create there disappear; download what you need

PyTorch's own tutorial says PyTorch comes pre-installed in Colab, and shows how to switch on a GPU: Runtime > Change runtime type, then T4 GPU under Hardware accelerator, then Save.

Build it yourself: time a tiny GPT on your computer

You will check that PyTorch is ready, run a timing script that trains a tiny GPT-sized stand-in model for a few steps, write down the result, and optionally run the same script on a free Colab GPU. A check script confirms everything at the end.

Before you start: complete Set up your computer for this course, Set up for Unit 2 and Set up for Unit 3. They install Python, VS Code, Git and PyTorch, and create your orchestrate-course folder with its .venv. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder from Units 1 to 3, with PyTorch installed.
  • About 20 to 40 minutes, and your laptop plugged in.
  • For the optional GPU step only: a Google account. Colab is free; no payment details are needed.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate

Step 2: Confirm PyTorch and see which device it finds

Run this one line (the same on every system):

python -c "import torch; print(torch.__version__, 'cuda:', torch.cuda.is_available(), 'mps:', torch.backends.mps.is_available())"

What success looks like (your version may differ):

2.14.1 cuda: False mps: False
  • cuda: False mps: False is the normal result on most Windows and Linux laptops. You will train on the CPU, which is fine.
  • mps: True means PyTorch can use your Mac's GPU. The timing script will use it automatically.
  • cuda: True means you have an NVIDIA GPU and PyTorch's GPU build. The script will use it automatically.

If you see ModuleNotFoundError: No module named 'torch', go back to Set up for Unit 3, Step 2.

Step 3: Make the Unit 4 folder

From your course folder:

  • Windows (PowerShell):

    New-Item -ItemType Directory -Force unit04
    cd unit04
  • macOS / Linux:

    mkdir -p unit04
    cd unit04

Step 4: Create and run the timing script

The script builds a stand-in model with the same shape and size as the tiny GPT you will build later in this unit. It trains it for 50 steps on random made-up characters and estimates how long 5,000 steps would take. It measures speed only; the model learns nothing useful.

  1. In VS Code's file list, right-click unit04, choose New File and name it bench_tiny_gpt.py.

  2. Paste the code below and save.

  3. In the terminal, inside unit04, run:

    python bench_tiny_gpt.py
"""Unit 4 timing test: how fast can this computer train a tiny GPT-sized model?

It builds a stand-in model the same size as a small character-level GPT, trains it for a
few steps on random made-up characters, and estimates how long a full training run would take.
It measures speed only; the model learns nothing useful. Unit 4 builds the real tiny GPT.

How to run (from the unit04 folder, with the course .venv turned on):
    python bench_tiny_gpt.py                    # small model, best device found
    python bench_tiny_gpt.py --size medium      # bigger model, closer to a common "baby GPT"
    python bench_tiny_gpt.py --device cpu       # force the processor, to compare
    python bench_tiny_gpt.py --steps 20         # fewer timed steps, for a quick look
"""
import argparse
import time

import torch

SIZES = {  # layers, embedding width, attention heads, context length, batch size
    "small": dict(n_layer=4, n_embd=128, n_head=4, block_size=64, batch_size=32),
    "medium": dict(n_layer=6, n_embd=384, n_head=6, block_size=256, batch_size=64),
}
VOCAB = 65          # about the number of different characters in a small text file
FULL_RUN = 5000     # training steps in a typical small character-level GPT run


def pick_device(choice: str) -> torch.device:
    """Use an NVIDIA GPU (cuda) or an Apple silicon GPU (mps) if present, else the CPU."""
    if choice == "cuda" and not torch.cuda.is_available():
        raise SystemExit("No NVIDIA GPU that PyTorch can use here. Run without --device, or use --device cpu.")
    if choice == "mps" and not torch.backends.mps.is_available():
        raise SystemExit("No Apple silicon GPU that PyTorch can use here. Run without --device, or use --device cpu.")
    if choice != "auto":
        return torch.device(choice)
    if torch.cuda.is_available():
        return torch.device("cuda")
    if torch.backends.mps.is_available():
        return torch.device("mps")
    return torch.device("cpu")


def sync(device: torch.device) -> None:
    """GPUs work in the background; wait for them so the timing is honest."""
    if device.type == "cuda":
        torch.cuda.synchronize()
    elif device.type == "mps":
        torch.mps.synchronize()


class StandIn(torch.nn.Module):
    """Same shape and size as a tiny GPT: token embeddings, masked self-attention layers, output layer."""

    def __init__(self, n_layer, n_embd, n_head, block_size):
        super().__init__()
        self.tok = torch.nn.Embedding(VOCAB, n_embd)
        self.pos = torch.nn.Embedding(block_size, n_embd)
        layer = torch.nn.TransformerEncoderLayer(n_embd, n_head, 4 * n_embd, dropout=0.0,
                                                 batch_first=True, norm_first=True)
        self.blocks = torch.nn.TransformerEncoder(layer, n_layer, enable_nested_tensor=False)
        self.head = torch.nn.Linear(n_embd, VOCAB)

    def forward(self, idx):
        t = idx.shape[1]
        x = self.tok(idx) + self.pos(torch.arange(t, device=idx.device))
        mask = torch.nn.Transformer.generate_square_subsequent_mask(t, device=idx.device)
        return self.head(self.blocks(x, mask=mask, is_causal=True))


def main() -> None:
    parser = argparse.ArgumentParser(description="Time a tiny GPT-sized training loop.")
    parser.add_argument("--size", choices=SIZES, default="small")
    parser.add_argument("--device", choices=["auto", "cpu", "cuda", "mps"], default="auto")
    parser.add_argument("--steps", type=int, default=50, help="training steps to time")
    args = parser.parse_args()

    cfg = SIZES[args.size]
    device = pick_device(args.device)
    torch.manual_seed(0)
    model = StandIn(cfg["n_layer"], cfg["n_embd"], cfg["n_head"], cfg["block_size"]).to(device)
    params = sum(p.numel() for p in model.parameters())
    optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3)
    print(f"PyTorch {torch.__version__} on device: {device.type}")
    print(f"Model size '{args.size}': {params / 1e6:.2f} million weights, "
          f"{cfg['n_layer']} layers, context {cfg['block_size']} characters, batch {cfg['batch_size']}")

    def one_step():
        x = torch.randint(0, VOCAB, (cfg["batch_size"], cfg["block_size"]), device=device)
        y = torch.randint(0, VOCAB, (cfg["batch_size"], cfg["block_size"]), device=device)
        loss = torch.nn.functional.cross_entropy(model(x).reshape(-1, VOCAB), y.reshape(-1))
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

    for _ in range(3):          # warm-up steps are slower; don't count them
        one_step()
    sync(device)
    start = time.perf_counter()
    for _ in range(args.steps):
        one_step()
    sync(device)
    per_step = (time.perf_counter() - start) / args.steps

    full = per_step * FULL_RUN
    print(f"Timed {args.steps} steps: {per_step * 1000:.0f} ms per step")
    print(f"Estimated time for a {FULL_RUN}-step training run: {full / 60:.1f} minutes")
    if full > 3600:
        print("Tip: that is over an hour. Use the small size, fewer steps, or a free GPU notebook.")


if __name__ == "__main__":
    main()

What success looks like (from our test on a small two-core cloud machine):

PyTorch 2.14.1 on device: cpu
Model size 'small': 0.82 million weights, 4 layers, context 64 characters, batch 32
Timed 50 steps: 92 ms per step
Estimated time for a 5000-step training run: 7.6 minutes

Your numbers will differ: that is the point. The device line shows cpu, mps or cuda. The version text also differs by system, for example 2.14.1+cpu on Linux. The run itself takes a few seconds to a minute.

Step 5: Time the larger size and write down your results

  1. Run the larger size with fewer timed steps, so it finishes quickly:

    python bench_tiny_gpt.py --size medium --steps 5

    What success looks like (same test machine):

    PyTorch 2.14.1 on device: cpu
    Model size 'medium': 10.80 million weights, 6 layers, context 256 characters, batch 64
    Timed 5 steps: 6458 ms per step
    Estimated time for a 5000-step training run: 538.2 minutes
    Tip: that is over an hour. Use the small size, fewer steps, or a free GPU notebook.

    On our test machine the medium model was about 70 times slower per step than the small one: 13 times more weights, and a context four times longer. About 9 hours is not a realistic laptop run. Your laptop is likely faster; measure it.

  2. If the device line said mps or cuda, also run both sizes with --device cpu to see what the GPU gains you.

  3. In VS Code, create timing_notes.md inside unit04 and fill it in:

    # My Unit 4 timings
    
    - Computer (model, processor):
    - Device the script chose:
    - Small: ms per step / estimated minutes:
    - Medium: ms per step / estimated minutes:
    - Colab GPU, medium: ms per step / estimated minutes (Step 6, optional):
    - My plan for the tiny GPT (laptop small, laptop medium, or Colab):
  4. Save the file.

What each part of the script does:

Part What it does
SIZES Two model sizes. medium matches the shape of nanoGPT's character-level "baby GPT"
pick_device Uses cuda, then mps, then cpu; stops with a clear message if you force a device that isn't there
StandIn Token and position embeddings, masked self-attention layers, and an output layer: a GPT's shape, built from PyTorch's ready-made layers
one_step One training step on random characters: predict, measure the error, backward, step
Warm-up loop Three untimed steps, so one-off setup work doesn't skew the result
sync Waits for the GPU to finish before reading the clock, as PyTorch's docs advise
--size, --device, --steps Optional: choose the model size, force a device, and set how many steps to time

Step 6 (optional): Run the same test on a free Colab GPU

Do this if your medium estimate was more than about an hour, or if you have an Intel Mac where PyTorch won't install. If your company doesn't allow Colab, skip it and use the small size on your laptop.

  1. Open https://colab.research.google.com and sign in with your Google account.

  2. Choose File > New notebook.

  3. Switch on a GPU: open Runtime > Change runtime type, choose T4 GPU under Hardware accelerator, and click Save. If no GPU option is offered, Colab has none for you right now; try again later.

  4. In the first code cell, type the line %%writefile bench_tiny_gpt.py, then paste the whole script from Step 4 below it. Run the cell with the play button to its left, or Shift+Enter. It answers Writing bench_tiny_gpt.py.

  5. Click + Code to add a cell, type this and run it:

    !python bench_tiny_gpt.py --size medium
  6. Copy the result into the Colab line of your timing_notes.md.

What success looks like: the first line says on device: cuda, and the medium time per step is far lower than on most laptops. If it says on device: cpu, the GPU wasn't switched on: repeat item 3, then run both cells again.

%%writefile saves the cell's contents as a file on Colab's machine. The ! at the start of a line runs it as a terminal command. Both disappear when Colab deletes the machine, which is fine: your copy lives in your course folder.

Step 7: Run the Unit 4 check

  1. Go back to your course folder in the terminal:

    cd ..
  2. In VS Code, create a new file in the course folder (not in unit04), paste the script below and save it as check_unit04.py.

  3. Run it:

    python check_unit04.py
"""Check that your computer is ready for Unit 4 (attention, transformers and a tiny GPT).

Run it from your course folder:  python check_unit04.py
It only reads your setup and opens one web page to test your network.
It installs nothing and sends none of your files or keys anywhere.
"""
import importlib.metadata
import importlib.util
import os
import subprocess
import sys
import urllib.error
import urllib.request

problems = 0


def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
    """Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
    global problems
    if ok:
        print(f"  OK       {label}")
    elif optional:
        print(f"  LATER    {label}  ->  {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label}  ->  {fix}")


def reachable(url: str) -> bool:
    """True if the site answers at all (any HTTP status counts; we only test the network)."""
    try:
        urllib.request.urlopen(urllib.request.Request(url, method="HEAD"), timeout=15)
        return True
    except urllib.error.HTTPError:
        return True   # the site answered, even if with an error page
    except Exception:
        return False


print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
       "the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")

print("\n2. PyTorch")
found = importlib.util.find_spec("torch") is not None
version = importlib.metadata.version("torch") if found else ""
report(found, f"torch {version}".strip(), "install it (Set up for Unit 3, Step 2)")

# In a fresh Python: compute a gradient, then ask which accelerator PyTorch can use.
test = (
    "import torch;"
    "w = torch.tensor(3.0, requires_grad=True);"
    "(w * w).backward();"
    "assert abs(w.grad.item() - 6.0) < 1e-6;"
    "print('cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'none')"
)
accelerator = ""
if found:
    try:
        result = subprocess.run([sys.executable, "-c", test], capture_output=True, text=True, timeout=180)
        ok = result.returncode == 0
        detail = (result.stderr.strip().splitlines() or ["unknown error"])[-1] if not ok else ""
        accelerator = result.stdout.strip().splitlines()[-1] if ok else ""
    except subprocess.TimeoutExpired:
        ok, detail = False, "timed out"
    report(ok, "tensor maths and gradients", f"{detail}; reinstall PyTorch (Set up for Unit 3, Step 2)")
if accelerator:
    names = {"cuda": "NVIDIA GPU (cuda)", "mps": "Apple silicon GPU (mps)"}
    report(accelerator != "none", f"local GPU: {names.get(accelerator, 'none found, the CPU will be used')}",
           "not needed; use the small size, or a free GPU notebook for the medium size (Step 6)",
           optional=True)

print("\n3. Network")
report(reachable("https://colab.research.google.com"), "can reach Google Colab (optional free GPU notebook)",
       "only needed for the optional GPU path; try another network or ask IT", optional=True)

print("\n4. Course folder")
report(os.path.isdir("unit04"), "unit04 folder", "create it (Step 3)", optional=True)
report(os.path.exists(os.path.join("unit04", "bench_tiny_gpt.py")), "unit04/bench_tiny_gpt.py",
       "create it (Step 4)", optional=True)
report(os.path.exists(os.path.join("unit04", "timing_notes.md")), "unit04/timing_notes.md",
       "write down your timings (Step 5)", optional=True)

print()
if problems:
    print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
    sys.exit(1)
print("All set. Your computer is ready for Unit 4.")

What success looks like (on a laptop without a GPU):

1. Python
  OK       Python 3.11.15
  OK       virtual environment is active

2. PyTorch
  OK       torch 2.14.1
  OK       tensor maths and gradients
  LATER    local GPU: none found, the CPU will be used  ->  not needed; use the small size, or a free GPU notebook for the medium size (Step 6)

3. Network
  OK       can reach Google Colab (optional free GPU notebook)

4. Course folder
  OK       unit04 folder
  OK       unit04/bench_tiny_gpt.py
  OK       unit04/timing_notes.md

All set. Your computer is ready for Unit 4.

A LATER line doesn't stop you. LATER local GPU is the normal result on most laptops. On an Apple silicon Mac where PyTorch finds the GPU, the line reads OK local GPU: Apple silicon GPU (mps).

What each part of the check does:

Part What it checks
Python Version 3.11 or newer, and that .venv is active
PyTorch That PyTorch is installed, can compute a gradient in a fresh Python, and which GPU it can use, if any
Network Whether Colab answers. It sends no data, only a request for the page header
Course folder That unit04, the timing script and your timing notes exist

Like the earlier checks, it uses only Python's built-in modules and changes nothing.

Step 8: Save your work in Git

From the course folder:

git add check_unit04.py unit04/bench_tiny_gpt.py unit04/timing_notes.md
git commit -m "Set up Unit 4: timing test and training plan"

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'torch' PyTorch isn't installed in the Python you're using Check for (.venv) in the prompt, then repeat Set up for Unit 3, Step 2
No matching distribution found for torch Your system has no current PyTorch build, for example an Intel Mac Use Colab (Step 6) for all of Unit 4's code, or another computer
No NVIDIA GPU that PyTorch can use here You ran --device cuda without a usable NVIDIA GPU and GPU build Run without --device, or with --device cpu
No Apple silicon GPU that PyTorch can use here You ran --device mps on a machine without MPS, or macOS older than 14 Run without --device, or with --device cpu
The medium test seems to hang Each step is slow on your machine Wait, or stop with Ctrl+C and run --size medium --steps 2
The estimate is far slower than expected The laptop is on battery saving, or other programs are busy Plug in, close heavy programs, run again
LATER can reach Google Colab Your network blocked Colab when the check ran Only matters for Step 6. Try another network, or ask IT
Colab offers no GPU, or says usage limits are reached Free GPUs aren't guaranteed Try later, or use the small size on your laptop
Colab shows on device: cpu The GPU wasn't switched on before the cells ran Step 6, item 3, then run both cells again
Colab says the runtime disconnected It timed out while idle, or reached its limits Run the cells again; the %%writefile cell recreates the file
Windows: .venv\Scripts\Activate.ps1 cannot be loaded PowerShell blocks scripts Run Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser, answer Y, and try again

Where this shows up in SAP

What you measured here is the same trade-off SAP customers face at a much larger scale. Bigger models and longer contexts cost more compute per step, and a GPU changes what is practical.

In SAP's offering, most of that cost sits with the model providers. As of October 2026, the generative AI hub in SAP AI Core and SAP AI Launchpad gives access to foundation models from several providers, so a customer pays to call a model rather than to train one. Unit 5 sets that up.

When a team runs its own model in SAP AI Core, the hardware is chosen in the deployment's configuration through a resource plan. SAP's GPU tutorial sets it like this:

# Sketch only: part of an SAP AI Core serving template. Needs a paid SAP AI Core setup.
resourcePlan: infer.s   # SAP's GPU tutorial uses this plan to get a GPU node

Pitfalls

  • Assuming a GPU is always faster. For the small size, the overhead of the GPU can eat most of the gain. Measure both, as Step 5 does.
  • Timing a GPU without synchronizing. Without torch.cuda.synchronize(), a GPU looks impossibly fast because the clock stops before the work does.
  • Counting the first steps. The first steps include setup work. Always warm up before timing.
  • Installing the GPU build "just in case". On a laptop without an NVIDIA card it only costs gigabytes. Colab gives you a GPU without it.
  • Treating Colab as storage. Colab deletes its machines. Keep every file in your course folder and Git.
  • Real data in a personal notebook. Everything in Unit 4 is made up. Keep it that way.

Exercise: decide where you will train your tiny GPT

  1. Run python bench_tiny_gpt.py and python bench_tiny_gpt.py --size medium --steps 5 in unit04.

  2. If you have mps or cuda, also run both with --device cpu.

  3. Optional: run the medium size in Colab (Step 6).

  4. Fill every line of timing_notes.md, including the plan line.

  5. Under a new heading ## Why, write two sentences: which option you chose for the tiny GPT and why, using your own numbers.

  6. Run python check_unit04.py from the course folder.

  7. Save your work in Git:

    git add unit04/timing_notes.md
    git commit -m "Add Unit 4 timing notes and training plan"

Done when: check_unit04.py prints All set, timing_notes.md has your small and medium timings and a plan with two sentences of reasons, and git log shows the commit. The tiny GPT topic later in Unit 4 starts from that plan.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Your laptop has no GPU. What does Step 2's one-line check print, and what does it mean?

    Answer: B. Without an NVIDIA GPU and the GPU build, torch.cuda.is_available() is False; MPS is only on Apple silicon Macs. PyTorch then uses the CPU, which is enough for the small tiny GPT.
  2. 2Why does bench_tiny_gpt.py call sync(device) before reading the clock?

    Answer: C. GPU work is asynchronous; Python moves on before the GPU is done. PyTorch's docs say timings without synchronizing are not accurate, so the script waits with torch.cuda.synchronize() or torch.mps.synchronize().
  3. 3Why are the first three training steps not timed?

    Answer: B. The first steps pay one-off costs such as allocating memory and preparing operations. Timing after a warm-up gives a fair time per step to multiply by 5,000.
  4. 4On our test machine the medium size was about 70 times slower per step than the small one. Which change contributes most to that?

    Answer: D. The medium model has about 13 times more weights and a context of 256 instead of 64. Attention compares each position with every earlier one, so its work grows faster than the context length.
  5. 5Your medium estimate on the laptop is 6 hours. What do you do?

    Answer: C. Hours of training is not a realistic laptop run. The small size trains in minutes locally, and Colab's free GPU can run the medium size, when one is available. The GPU build without an NVIDIA card gains nothing.
  6. 6In Colab, the script prints on device: cpu. What went wrong?

    Answer: A. PyTorch is pre-installed in Colab, but a new notebook has no GPU until you choose T4 GPU under Hardware accelerator and save. Then run both cells again so the script sees the GPU.
  7. 7A colleague wants to train on real sales order notes in Colab to "make it realistic". What do you say?

    Answer: D. Colab here is a personal learning tool with no company agreement behind it. Deleting the machine later doesn't undo the upload, so Unit 4 uses only made-up text.
  8. 8How does a team choose GPU hardware when it runs its own model in SAP AI Core?

    Answer: B. SAP AI Core picks machine sizes through resource plans, and SAP's GPU tutorial uses infer.s to get a GPU node. It needs a paid setup, so the course shows it only as a sketch.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in