Orchestrate

When to fine-tune (and when not to)

Decide whether a task needs fine-tuning by measuring prompts first, then train a small LoRA adapter on SAP-shaped notes and compare the numbers.

Updated Oct 8, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

Fine-tuning means training an existing language model a little more, on your own examples, so it changes how it behaves. You keep the model's general skills and teach it one narrow habit: a format, a tone, a set of labels.

It is the most expensive way to improve an AI feature, and the one teams reach for too early. Most problems are solved more cheaply by a better prompt, by a few worked examples in the prompt, or by giving the model the right documents (retrieval).

Use this rule:

  • If the model doesn't know something, give it the information. Don't fine-tune.
  • If the model knows enough but behaves inconsistently, and better prompts have stopped helping, fine-tuning may be worth it.
  • In every case, measure first. Without an evaluation set, you can't tell whether fine-tuning helped.

Why it matters to the business

Fine-tuning is a project, not a setting. It needs labelled examples, people to check them, compute to train, an evaluation, and someone to repeat all of it when the base model is retired. When it pays off, it pays off in three ways:

  • Consistency. The model returns the same format, labels or tone every time, so downstream automation doesn't break.
  • Cost per call. A tuned model needs a short prompt instead of long instructions and examples. At millions of calls, fewer tokens per call adds up (see token economics).
  • A smaller model. A small tuned model can sometimes match a large general one on one narrow task, which can make running it yourself realistic (see the Unit 12 setup).

Take the running example. Order-to-cash clerks write short notes on blocked sales orders: "customer over the credit limit", "no price for material TG11", "customer asked us to hold delivery". A team wants each note routed automatically to credit management, master data, pricing, logistics or customer service.

That is a good fine-tuning candidate: the labels are fixed, the output format must be exact, and volume is high. But the team should still try a clear prompt and a few examples first. If those reach the target, the fine-tuning project never needs to start.

A bad candidate is "make the model know our pricing conditions". Pricing changes weekly. Knowledge that changes belongs in retrieval or a live API call, not baked into a model.

How SAP does it

As of October 2026, SAP's own AI products point the same way: adapt the input first, and leave model training to specialists.

  • Prompt optimization in the generative AI hub. SAP's tutorial shows the hub rewriting a prompt template automatically against your labelled dataset and a metric you choose, such as exact match on JSON. The improved prompt is stored in the Prompt Registry. This tunes the prompt, not the model, and needs the Extended plan of SAP AI Core.
  • Models that learn from examples at run time. SAP's documentation says SAP-RPT-1, its model for predictions on business tables, works without training or fine-tuning: you send example rows with the request. The tabular AI topic later in this unit covers it.
  • Models SAP has tuned for you. SAP's documentation describes SAP-ABAP-1 as a foundation model fine-tuned on a large amount of ABAP code, available in the generative AI hub. SAP did the fine-tuning; customers call the result.
  • Your own training on SAP AI Core. For teams that do need a custom model, SAP AI Core runs training workflows on GPU resource plans and then serves the trained model.

In the SAP pages opened for this topic, we found no documented option to fine-tune the weights of the partner models in the generative AI hub. If a vendor offers it, ask exactly where training runs and where the tuned model lives.

A decision guide

Work down the ladder. Stop at the first rung that meets your target on your evaluation set.

Rung What you change Typical effort Fixes Doesn't fix
1. Clear prompt The instructions Hours Vague or missing instructions Missing knowledge
2. Examples in the prompt A few worked examples Hours Format, labels, tone Very long or varied tasks
3. Retrieval or tools The data sent with each request Days to weeks Missing or changing knowledge Inconsistent behaviour
4. Fine-tuning The model itself Weeks, then ongoing Consistent behaviour, shorter prompts, smaller models Facts that change
5. Training from scratch Everything Months, large budget Almost never needed for SAP projects

Signs fine-tuning is worth a pilot:

  • The task is narrow and stable, with a fixed set of outputs.
  • Prompts and examples have plateaued below the target, and the errors are about behaviour, not missing facts.
  • You have, or can label, at least dozens of good examples. OpenAI's guide, for example, suggests starting with 50 well-crafted ones.
  • Volume is high enough that shorter prompts or a smaller model save real money.
  • Someone owns the dataset and the retraining after the pilot.

Signs it isn't:

  • The answer depends on data that changes, such as prices, stock or open items.
  • You have no evaluation set yet.
  • The examples contain personal data you aren't allowed to put into a model (see data security and PII).

Vendor terms change too. OpenAI's documentation now says it is winding down its self-serve fine-tuning platform, with no new training jobs for existing customers after 6 January 2027. A tuned model is tied to its base model and its platform; plan for that.

Questions to ask

Your team:

  • What is our evaluation set, and what score did the best prompt reach on it?
  • Are the failures about missing knowledge or about behaviour?
  • Who labels the examples, and who checks the labels?
  • What happens when the base model is retired? Who retrains, and how long does it take?

A vendor or partner proposing fine-tuning:

  • Which prompt and retrieval baselines did you measure, and on which data?
  • Where does training run, and where is our data stored afterwards?
  • Who owns the tuned model or adapter? Can we take it with us?
  • What does one retraining cycle cost, including labelling?

Common misconceptions

  • "Fine-tuning teaches the model our company data." It mostly teaches behaviour. Facts that change belong in retrieval or an API call.
  • "More training data is always better." A few dozen clean, varied examples often beat thousands of noisy ones. Quality decides.
  • "Fine-tuning is a one-off." Base models get retired. The dataset, the training run and the evaluation must be repeatable.
  • "We need to fine-tune to use AI with SAP." Most SAP use cases in this course run on prompts, retrieval and tools.
  • "A fine-tuned model is safe because it's ours." It still makes mistakes and still needs evaluation, guardrails and access control.

Key terms

  • Fine-tuning: further training of an existing model on your own examples, to change its behaviour.
  • Base model: the model you start from.
  • Training example: one input with the ideal output, written or approved by a person.
  • Evaluation set (test set): examples kept out of training and used only to score the model.
  • Few-shot prompt: a prompt that includes a few worked examples.
  • LoRA (low-rank adaptation): a cheap way to fine-tune that trains small add-on matrices instead of the whole model.
  • Adapter: the small file LoRA produces. It works only with the base model it was trained on.
  • Prompt optimization: automatically rewriting a prompt to score better on a labelled dataset, without changing the model.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A sponsor asks to fine-tune a model so it "knows our current pricing conditions". What is the best response?

    Answer: B. Fine-tuning mainly changes behaviour, such as format and labels. Pricing changes often, so it belongs in data fetched with each request. Retraining every time prices change would be slow and costly.
  2. 2What should exist before anyone starts a fine-tuning project?

    Answer: C. Without an evaluation set and a prompt baseline, nobody can show that fine-tuning helped. The baseline also tells you whether cheaper rungs already meet the target.
  3. 3Which task is the strongest candidate for fine-tuning?

    Answer: D. The routing task is narrow, stable, high volume and has a fixed output. The others depend on changing knowledge or happen too rarely to justify training.
  4. 4How does SAP's prompt optimization in the generative AI hub differ from fine-tuning?

    Answer: A. Prompt optimization rewrites the prompt template to score better on a labelled dataset and a chosen metric. The model stays as it is, which makes it a cheaper rung to try first.
  5. 5A partner says "once it's fine-tuned, we're done". What is the main risk they are ignoring?

    Answer: B. A tuned model depends on its base model and platform. OpenAI's wind-down of its fine-tuning platform shows that terms change, so the dataset and training must be repeatable.
  6. 6The best prompt scores 72% on your evaluation set and the target is 90%. The errors are inconsistent labels, not missing facts. What next?

    Answer: C. Behaviour errors that prompts haven't fixed are what fine-tuning addresses. Scoring the pilot on the same evaluation set shows whether it closed the gap. More retrieval wouldn't help, because knowledge isn't the problem.
  7. 7Which question best tests a vendor's fine-tuning proposal?

    Answer: D. A proposal that skipped the cheaper rungs can't show fine-tuning was needed. The baseline numbers are what the tuned model has to beat.
Deep layer · 40 min read

Mental model: fine-tuning changes habits, context supplies facts

A language model has two ways to get something right. It can know it, from training. Or it can be told it, in the prompt. OpenAI's accuracy guide calls these learned memory and in-context memory, and treats them as separate levers that can be combined.

That gives one test for every "should we fine-tune?" question:

  • Is the failure about what the model knows? Change the context: retrieval, tools, a live API call. This is Units 7 and 9.
  • Is the failure about how the model behaves? Format drift, wrong labels, the wrong tone, an instruction it keeps ignoring. First try a clearer prompt and examples. If those plateau, fine-tuning changes the habit itself.

Fine-tuning moves what you'd otherwise put in every prompt (instructions and examples) into the model's weights. You pay once for training instead of on every call, and you take on a training pipeline to maintain.

How it works

The ladder, as a loop

flowchart TD
  E[Evaluation set<br/>and target] --> P[Clear prompt]
  P -->|score| D1{Target met?}
  D1 -->|yes| S[Ship and monitor]
  D1 -->|no| F[Add examples<br/>to the prompt]
  F -->|score| D2{Target met?}
  D2 -->|yes| S
  D2 -->|no| K{Knowledge or<br/>behaviour error?}
  K -->|knowledge| R[Retrieval or tools]
  R --> D1
  K -->|behaviour| T[Fine-tune pilot]
  T -->|score on the same set| D3{Beats best prompt<br/>by enough?}
  D3 -->|yes| S
  D3 -->|no| X[Stop: keep the prompt]

Every arrow labelled "score" uses the same evaluation set, built as in LLM evaluation fundamentals. If the set changes between rungs, the comparison means nothing.

What training does

Supervised fine-tuning shows the model pairs of input and ideal output. For each pair, the model predicts the output one token at a time. The training code measures how far each prediction was from the ideal token (the loss) and nudges the weights to reduce it. After many such nudges, the ideal output becomes the likely one.

Two details matter in practice:

  • Only the answer is scored. The prompt part of each example is masked out, so the model learns to produce answers, not to repeat prompts.
  • Training data and serving format must match. The tuned model expects the same system instruction and chat layout it was trained with.

Full fine-tuning vs. LoRA

Full fine-tuning updates every weight. For a model with billions of parameters that needs a lot of GPU memory, and each tuned copy is as large as the original.

LoRA (low-rank adaptation) freezes the original weights. Next to selected weight matrices it adds two small matrices, A and B, whose product is a low-rank update. Only A and B are trained. The LoRA paper reports, for GPT-3 175B, about 10,000 times fewer trainable parameters and about three times less GPU memory than full fine-tuning, with no added delay when answering.

flowchart LR
  X[input] --> W[frozen weight W]
  X --> A[small matrix A] --> B[small matrix B]
  W --> P((+))
  B --> P
  P --> Y[output]

Hugging Face's PEFT library implements LoRA. Its quicktour shows a 1-billion-parameter model training about 0.04% of its parameters, and an adapter file of 6.3 MB next to a base model of about 700 MB. The adapter is useless without the exact base model it was trained on.

The settings you will meet:

Setting What it controls Value in this topic
r Rank: the size of A and B. Larger can learn more and costs more 8
lora_alpha How strongly the update is scaled 16
lora_dropout Random dropping during training, against overfitting 0.05
target_modules Which weight matrices get adapters the attention projections q_proj, k_proj, v_proj, o_proj
learning rate Size of each nudge 0.0002
steps, batch How many nudges, how many examples per nudge 60 steps of 8

Data: the part that decides the outcome

  • Format. Most tools accept JSONL: one JSON object per line, each holding a messages list of system, user and assistant turns. OpenAI's guide uses this format; the scripts below do too.
  • Amount. OpenAI's guide sets a minimum of 10 examples, says improvements typically show with 50 to 100, and suggests starting with 50 well-crafted ones. If 50 change nothing, revisit the task before adding more.
  • Splits. Keep a validation set to watch training, and a test set the model never sees. Our scripts go one step further: test notes use different phrasings from training notes. That shows whether the model learned the task or memorised the wording.
  • Leakage. A test example that also appears in training inflates the score. Check for it in code.

What changes in production

A tuned model is a new artefact with a life cycle: dataset version, training run, evaluation report, deployment, and retraining when the base model is retired. That is the real cost, more than the GPU hours.

Build it yourself: measure first, then fine-tune a small model

You will create a made-up dataset of clerk notes on blocked sales orders. You'll then score a small open model with a plain prompt and with examples in the prompt, train a LoRA adapter on your laptop, and score again. The script prints one comparison table, which is the evidence for the decision.

Before you start: complete Set up your computer for this course and Set up for Unit 12: local models. This walkthrough doesn't repeat their steps. It uses Hugging Face libraries instead of Ollama, because Ollama runs models but doesn't train them.

flowchart LR
  D[make_ft_data.py<br/>train, val, test] --> B[Score prompts<br/>zero-shot, few-shot]
  B --> T[Train LoRA<br/>adapter]
  T --> E[Score tuned model<br/>same test set]
  E --> R[ft_report.json<br/>decision]

What you need

  • Your course folder with .venv and Python 3.11 or newer, from earlier units.
  • About 4 GB of free disk: the libraries plus the model, which is about 1.5 GB on Hugging Face.
  • 8 GB of memory or more. With less, use --limit and fewer --steps.
  • 30 to 60 minutes. Training on a laptop processor takes several minutes; a GPU makes it faster but isn't needed.
  • Cost: free. No account or key. The model, Qwen3-0.6B, has an Apache-2.0 licence.

Step 1: Install the training libraries

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course. Open a terminal (Terminal > New Terminal) and turn on .venv if the prompt doesn't show (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  2. Linux only: install the processor-only version of PyTorch first. The default Linux install also downloads large GPU packages. The selector on pytorch.org (Linux, Pip, Python, CPU) gives this command:

    pip install torch --index-url https://download.pytorch.org/whl/cpu

    If you have an NVIDIA GPU on Linux, use the command the selector shows for your CUDA version instead.

  3. Open requirements.txt, add these three lines at the end, and save:

    torch
    transformers
    peft
  4. Install (the same on every system):

    pip install -r requirements.txt

What success looks like: a line starting Successfully installed, listing peft, transformers and, on Windows and macOS, torch, with version numbers. This can take a few minutes.

Step 2: Create the dataset

You'll generate 220 made-up notes across five routes. Train and validation notes share four phrasings per route; test notes use two other phrasings that training never sees.

  1. In VS Code's file list, right-click unit12, choose New File, name it make_ft_data.py, paste the code and save.
"""Make a small, made-up dataset for the fine-tuning decision: route blocked-order notes.

Run it from your course folder:
    python unit12/make_ft_data.py

It writes three files in unit12/ft_data/ in the chat format most fine-tuning tools accept:
    train.jsonl  examples the model learns from
    val.jsonl    examples to watch during training (same phrasings as train)
    test.jsonl   examples the model never sees in training, written with DIFFERENT phrasings

Every note is made up. The routes are this course's own labels, not SAP codes.
Built-in Python only; nothing leaves your computer.
"""
import json
import random
from pathlib import Path

SEED = 7                      # same seed, same files, every time
OUT = Path("unit12") / "ft_data"
ROUTES = ["CREDIT", "MASTER_DATA", "PRICING", "STOCK", "CUSTOMER_HOLD"]

# The short instruction every example carries. The fine-tuned model sees only this.
SYSTEM = ("Route the clerk's note about a blocked SAP sales order. "
          'Reply with JSON only: {"route": "<ROUTE>"}. '
          "ROUTE is one of " + ", ".join(ROUTES) + ".")

# Phrasings per route. The first four go to train and val; the last two only to test.
# Holding phrasings back is what tells you whether a model learned the task or the wording.
TEMPLATES = {
    "CREDIT": [
        "Order {so} stuck, customer {cu} is {amt} {cur} over the credit limit.",
        "Credit check failed on {so}. {cu} has invoices overdue since {month}.",
        "{so}: exposure for {cu} too high after last week's orders, needs credit release.",
        "Customer {cu} limit exceeded, {so} worth {amt} {cur} cannot go out.",
        "Finance says {cu} still hasn't paid the {month} statement, so {so} is held.",
        "Risk team flagged {cu}; payment behaviour got worse and {so} waits on them.",
    ],
    "MASTER_DATA": [
        "{so} blocked: ship-to address for {cu} has no postal code.",
        "Tax number missing on customer {cu}, so {so} can't be released.",
        "Incoterms not maintained for {cu}; {so} stays blocked.",
        "{so} cannot be delivered, sales area data for {cu} is incomplete.",
        "Customer record {cu} was created yesterday and half the fields are empty, {so} on hold.",
        "Bank and contact details for {cu} never got set up, which is why {so} won't move.",
    ],
    "PRICING": [
        "No price found for material {mat} on {so}.",
        "{so}: manual discount of {pct}% is above tolerance, needs approval.",
        "Condition record for {mat} expired, {so} shows zero price.",
        "Price on {so} differs from contract for {cu} by {pct}%.",
        "Sales rep typed a different net value on {so} than the agreement with {cu} allows.",
        "Line for {mat} on {so} came through at 0.00 {cur}, someone has to fix the amount.",
    ],
    "STOCK": [
        "Not enough stock of {mat} for {so}, short by {qty} pieces.",
        "{so}: availability check confirms only {qty} of the ordered quantity.",
        "Material {mat} back-ordered, {so} waits for the next receipt.",
        "Plant has no free quantity of {mat}; {so} cannot be scheduled.",
        "Warehouse is out of {mat} until the supplier delivers, {so} sits there.",
        "We can ship only part of {so}, {qty} units of {mat} just aren't on the shelf.",
    ],
    "CUSTOMER_HOLD": [
        "{cu} asked us to hold {so} until further notice.",
        "Customer called: do not ship {so} before {month}.",
        "{so} on hold at {cu}'s request, their warehouse is full.",
        "Buyer at {cu} wants {so} delayed while they check their budget.",
        "Email from {cu}: please park {so}, they are moving sites.",
        "{cu} told the rep to keep {so} back until their project restarts.",
    ],
}
MONTHS = ["May", "June", "July", "August", "September"]


def fill(template: str, rng: random.Random) -> str:
    """Put random, made-up values into one phrasing."""
    return template.format(
        so=str(rng.randint(9000001, 9000999)), cu=str(rng.randint(10100001, 10100999)),
        amt=f"{rng.randint(2, 90) * 250:,}", cur=rng.choice(["EUR", "USD", "GBP"]),
        month=rng.choice(MONTHS), mat=f"TG{rng.randint(10, 99)}",
        pct=rng.randint(6, 25), qty=rng.randint(5, 400))


def example(note: str, route: str) -> dict:
    """One training example in chat format: instruction, note, and the ideal answer."""
    return {"messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": note},
        {"role": "assistant", "content": json.dumps({"route": route})},
    ]}


def make(count_per_route: int, template_slice: slice, rng: random.Random) -> list:
    rows = []
    for route in ROUTES:
        phrasings = TEMPLATES[route][template_slice]
        for i in range(count_per_route):
            rows.append(example(fill(phrasings[i % len(phrasings)], rng), route))
    rng.shuffle(rows)
    return rows


def main() -> None:
    rng = random.Random(SEED)
    splits = {
        "train": make(32, slice(0, 4), rng),   # 160 examples
        "val": make(4, slice(0, 4), rng),      # 20 examples
        "test": make(8, slice(4, 6), rng),     # 40 examples, unseen phrasings
    }
    # Leakage check: no note may appear in both train and test.
    train_notes = {row["messages"][1]["content"] for row in splits["train"]}
    leaked = [r for r in splits["test"] if r["messages"][1]["content"] in train_notes]
    if leaked:
        raise SystemExit(f"{len(leaked)} test notes also appear in train; change SEED.")

    OUT.mkdir(parents=True, exist_ok=True)
    for name, rows in splits.items():
        with open(OUT / f"{name}.jsonl", "w", encoding="utf-8") as handle:
            for row in rows:
                handle.write(json.dumps(row) + "\n")
        print(f"Wrote {len(rows):>3} examples to {OUT / (name + '.jsonl')}")
    print("Leakage check: no test note appears in train.")
    first = splits["train"][0]["messages"]
    print(f"\nOne training example:\n  note:   {first[1]['content']}\n  answer: {first[2]['content']}")


if __name__ == "__main__":
    main()
  1. Run it from the course folder:

    python unit12/make_ft_data.py

What success looks like:

Wrote 160 examples to unit12/ft_data/train.jsonl
Wrote  20 examples to unit12/ft_data/val.jsonl
Wrote  40 examples to unit12/ft_data/test.jsonl
Leakage check: no test note appears in train.

One training example:
  note:   Incoterms not maintained for 10100202; 9000482 stays blocked.
  answer: {"route": "MASTER_DATA"}

On Windows the paths print with backslashes. Open unit12/ft_data/test.jsonl in VS Code: each line is one example with a system instruction, a note and the ideal answer.

Step 3: Create the comparison script

  1. In unit12, create finetune_lora.py, paste the code and save.
"""Should we fine-tune? Measure prompting first, then train a small LoRA adapter and compare.

Run it from your course folder, after make_ft_data.py:
    python unit12/finetune_lora.py --sample          # no model, no download: shows the report format
    python unit12/finetune_lora.py --baseline-only   # score zero-shot and few-shot prompts only
    python unit12/finetune_lora.py                   # baselines, then LoRA training, then the tuned score
    python unit12/finetune_lora.py --steps 20 --limit 10   # a quick, rough run

The first real run downloads Qwen/Qwen3-0.6B from Hugging Face (about 1.5 GB, Apache-2.0).
Everything then runs on this computer. Training on a laptop processor takes several minutes.
"""
import argparse
import json
import random
import sys
import time
from pathlib import Path

DATA = Path("unit12") / "ft_data"
ADAPTER_DIR = Path("unit12") / "ft_adapter"
REPORT = Path("unit12") / "ft_report.json"
ROUTES = ["CREDIT", "MASTER_DATA", "PRICING", "STOCK", "CUSTOMER_HOLD"]

# What --sample prints. Made-up numbers, to show the format only; your run will differ.
SAMPLE_REPORT = {
    "model": "sample (no model was called)",
    "results": [
        {"setup": "zero-shot prompt", "accuracy": 0.55, "valid_json": 0.80, "prompt_tokens": 70},
        {"setup": "few-shot prompt", "accuracy": 0.83, "valid_json": 1.00, "prompt_tokens": 260},
        {"setup": "LoRA fine-tuned", "accuracy": 0.93, "valid_json": 1.00, "prompt_tokens": 70},
    ],
    "train_seconds": 410,
    "trainable_share": 0.0038,
}


def read_jsonl(name: str) -> list:
    path = DATA / f"{name}.jsonl"
    if not path.exists():
        sys.exit(f"{path} not found. Run first: python unit12/make_ft_data.py")
    with open(path, encoding="utf-8") as handle:
        return [json.loads(line) for line in handle if line.strip()]


def few_shot_turns(train: list) -> list:
    """One worked example per route, taken from the training data, as earlier chat turns."""
    turns, seen = [], set()
    for row in train:
        route = json.loads(row["messages"][2]["content"])["route"]
        if route not in seen:
            seen.add(route)
            turns += row["messages"][1:3]
    return turns


def build_prompt(tokenizer, row: dict, shots: list) -> str:
    """System instruction, optional examples, then the note. Qwen3's thinking is switched off."""
    system, note = row["messages"][0], row["messages"][1]
    return tokenizer.apply_chat_template([system, *shots, note], tokenize=False,
                                         add_generation_prompt=True, enable_thinking=False)


def parse_route(text: str):
    """Return (valid, route). Valid means the reply is JSON with a known route."""
    start, end = text.find("{"), text.rfind("}")
    try:
        route = json.loads(text[start:end + 1]).get("route")
    except (ValueError, AttributeError):
        return False, None
    return route in ROUTES, route


def evaluate(model, tokenizer, rows: list, shots: list, label: str, device: str) -> dict:
    """Ask the model about every test note; count right answers and well-formed replies."""
    import torch

    correct = valid = tokens = 0
    model.eval()
    for i, row in enumerate(rows, 1):
        expected = json.loads(row["messages"][2]["content"])["route"]
        inputs = tokenizer(build_prompt(tokenizer, row, shots), return_tensors="pt").to(device)
        tokens += inputs["input_ids"].shape[1]
        with torch.no_grad():
            output = model.generate(**inputs, max_new_tokens=16, do_sample=False,
                                    pad_token_id=tokenizer.pad_token_id or tokenizer.eos_token_id)
        reply = tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
        ok, route = parse_route(reply)
        valid += ok
        correct += route == expected
        print(f"\r  {label}: {i}/{len(rows)} notes scored", end="", flush=True)
    print()
    n = len(rows)
    return {"setup": label, "accuracy": correct / n, "valid_json": valid / n,
            "prompt_tokens": round(tokens / n)}


def encode_for_training(tokenizer, row: dict) -> tuple:
    """Token ids for prompt + answer, and labels that score only the answer tokens."""
    prompt_ids = tokenizer(build_prompt(tokenizer, row, []), add_special_tokens=False)["input_ids"]
    answer = row["messages"][2]["content"] + tokenizer.eos_token
    answer_ids = tokenizer(answer, add_special_tokens=False)["input_ids"]
    return prompt_ids + answer_ids, [-100] * len(prompt_ids) + answer_ids   # -100 = ignore


def batch_loss(model, tokenizer, rows: list, device: str):
    """Pad a batch to one length and return the model's loss on the answer tokens."""
    import torch

    pairs = [encode_for_training(tokenizer, row) for row in rows]
    width = max(len(ids) for ids, _ in pairs)
    pad = tokenizer.pad_token_id if tokenizer.pad_token_id is not None else tokenizer.eos_token_id
    ids = [p[0] + [pad] * (width - len(p[0])) for p in pairs]
    labels = [p[1] + [-100] * (width - len(p[1])) for p in pairs]
    mask = [[1] * len(p[0]) + [0] * (width - len(p[0])) for p in pairs]
    as_tensor = lambda rows_: torch.tensor(rows_, device=device)
    return model(input_ids=as_tensor(ids), attention_mask=as_tensor(mask),
                 labels=as_tensor(labels)).loss


def train_lora(model, tokenizer, train: list, val: list, args, device: str):
    """Freeze the model, add small LoRA matrices, and train only those on the examples."""
    import torch
    from peft import LoraConfig, get_peft_model

    config = LoraConfig(task_type="CAUSAL_LM", r=8, lora_alpha=16, lora_dropout=0.05,
                        target_modules=["q_proj", "k_proj", "v_proj", "o_proj"])
    model = get_peft_model(model, config)
    trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
    total = sum(p.numel() for p in model.parameters())
    print(f"  LoRA trains {trainable:,} of {total:,} parameters ({trainable / total:.2%})")

    optimizer = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=args.lr)
    rng = random.Random(0)
    started = time.time()
    model.train()
    for step in range(1, args.steps + 1):
        loss = batch_loss(model, tokenizer, rng.sample(train, args.batch), device)
        loss.backward()
        optimizer.step()
        optimizer.zero_grad()
        if step % 10 == 0 or step == 1 or step == args.steps:
            print(f"  step {step:>3}/{args.steps}  training loss {loss.item():.3f}"
                  f"  ({time.time() - started:.0f} s)")
    model.eval()
    with torch.no_grad():
        val_loss = batch_loss(model, tokenizer, val, device).item()
    print(f"  validation loss {val_loss:.3f}  (lower is better; compare with the training loss)")
    model.save_pretrained(ADAPTER_DIR)
    print(f"  adapter saved to {ADAPTER_DIR}/")
    return model, time.time() - started, trainable / total


def print_report(report: dict) -> None:
    print(f"\nModel: {report['model']}")
    print(f"{'Setup':<18}{'Accuracy':>10}{'Valid JSON':>12}{'Prompt tokens':>15}")
    for r in report["results"]:
        print(f"{r['setup']:<18}{r['accuracy']:>10.0%}{r['valid_json']:>12.0%}"
              f"{r['prompt_tokens']:>15}")
    if report.get("train_seconds"):
        print(f"Training took {report['train_seconds']:.0f} s and changed "
              f"{report['trainable_share']:.2%} of the parameters.")
    best_prompt = max((r for r in report["results"] if "prompt" in r["setup"]),
                      key=lambda r: r["accuracy"])
    tuned = [r for r in report["results"] if r["setup"] == "LoRA fine-tuned"]
    if tuned:
        gain = tuned[0]["accuracy"] - best_prompt["accuracy"]
        print(f"Fine-tuning vs. the best prompt: {gain:+.0%} accuracy, "
              f"{tuned[0]['prompt_tokens'] - best_prompt['prompt_tokens']:+d} tokens per call.")


def main() -> None:
    parser = argparse.ArgumentParser(description="Compare prompting with LoRA fine-tuning.")
    parser.add_argument("--model", default="Qwen/Qwen3-0.6B", help="Hugging Face model or folder")
    parser.add_argument("--steps", type=int, default=60, help="training steps (default 60)")
    parser.add_argument("--batch", type=int, default=8, help="examples per step (default 8)")
    parser.add_argument("--lr", type=float, default=2e-4, help="learning rate (default 0.0002)")
    parser.add_argument("--limit", type=int, default=0, help="score only the first N test notes")
    parser.add_argument("--baseline-only", action="store_true", help="skip training")
    parser.add_argument("--sample", action="store_true", help="no model: print a sample report")
    args = parser.parse_args()

    if args.sample:
        print_report(SAMPLE_REPORT)
        return

    train, val, test = read_jsonl("train"), read_jsonl("val"), read_jsonl("test")
    if args.limit:
        test = test[:args.limit]
    try:
        import torch
        from transformers import AutoModelForCausalLM, AutoTokenizer
    except ModuleNotFoundError as missing:
        sys.exit(f"Missing library ({missing.name}). Install the Step 1 libraries, then retry.")

    torch.manual_seed(0)
    device = "cuda" if torch.cuda.is_available() else "cpu"
    print(f"Loading {args.model} on {device} (the first run downloads it)...")
    try:
        tokenizer = AutoTokenizer.from_pretrained(args.model)
        model = AutoModelForCausalLM.from_pretrained(args.model, dtype=torch.float32).to(device)
    except OSError as error:
        sys.exit(f"Couldn't load {args.model}: {error}\nCheck the name and your network or proxy.")

    shots = few_shot_turns(train)
    print(f"Scoring {len(test)} test notes the model has never seen:")
    results = [evaluate(model, tokenizer, test, [], "zero-shot prompt", device),
               evaluate(model, tokenizer, test, shots, "few-shot prompt", device)]
    report = {"model": args.model, "results": results}

    if not args.baseline_only:
        print(f"Training a LoRA adapter on {len(train)} examples for {args.steps} steps:")
        model, seconds, share = train_lora(model, tokenizer, train, val, args, device)
        results.append(evaluate(model, tokenizer, test, [], "LoRA fine-tuned", device))
        report.update(train_seconds=seconds, trainable_share=share)

    REPORT.write_text(json.dumps(report, indent=2), encoding="utf-8")
    print_report(report)
    print(f"Report saved to {REPORT}")


if __name__ == "__main__":
    main()
  1. Run it with --sample. This needs no model and no download, so it works even if Step 1 failed:

    python unit12/finetune_lora.py --sample

What success looks like (made-up numbers, to show the format):

Model: sample (no model was called)
Setup               Accuracy  Valid JSON  Prompt tokens
zero-shot prompt         55%         80%             70
few-shot prompt          83%        100%            260
LoRA fine-tuned          93%        100%             70
Training took 410 s and changed 0.38% of the parameters.
Fine-tuning vs. the best prompt: +10% accuracy, -190 tokens per call.

Step 4: Score the prompts first

This run downloads the model once, then scores the 40 test notes with a plain prompt and with five examples in the prompt. It doesn't train anything.

python unit12/finetune_lora.py --baseline-only

What success looks like: a loading line, progress counters such as zero-shot prompt: 40/40 notes scored, then a table with two rows and your own numbers. The download may print progress bars and harmless warnings first.

Write down the few-shot accuracy. This is the number fine-tuning has to beat. If it already meets your target, the honest answer for this task is "don't fine-tune".

Step 5: Train the adapter and compare

python unit12/finetune_lora.py

What success looks like (your numbers and times will differ):

Training a LoRA adapter on 160 examples for 60 steps:
  LoRA trains ... of ... parameters (...%)
  step   1/60  training loss ...
  step  10/60  training loss ...
  ...
  validation loss ...  (lower is better; compare with the training loss)
  adapter saved to unit12/ft_adapter/

Then the table gains a third row, LoRA fine-tuned, and a last line comparing it with the best prompt. The training loss should fall over the steps. If the validation loss is far above the final training loss, the model has memorised the training notes; try fewer --steps.

Look at three columns, not one:

  • Accuracy on unseen phrasings: did the tuned model learn the task?
  • Valid JSON: did it learn the format? Small models often improve here first.
  • Prompt tokens: the tuned model runs with the short instruction only, so each call is cheaper than the few-shot prompt.

Step 6: Save your work

  1. Open .gitignore in the course folder, add this line and save. The adapter is a training output; you can always rebuild it.

    unit12/ft_adapter/
  2. Commit:

    git add requirements.txt .gitignore unit12/make_ft_data.py unit12/finetune_lora.py unit12/ft_data unit12/ft_report.json
    git commit -m "Unit 12: measure prompts, then LoRA fine-tune"

How the code works

Part What it does
TEMPLATES Six phrasings per route; the first four make train and validation data, the last two only test data
leakage check Stops if any test note also appears in training
few_shot_turns Picks one training example per route and adds them as earlier chat turns
build_prompt Uses the model's own chat template with enable_thinking=False, as the Qwen3 model card describes
parse_route Counts a reply as valid only if it holds JSON with a known route
evaluate Greedy decoding (no randomness), then accuracy, valid-JSON rate and average prompt tokens
encode_for_training Marks prompt tokens with -100 so only the answer is scored
train_lora LoraConfig and get_peft_model from PEFT, a plain training loop, validation loss, save_pretrained
ft_report.json The numbers, saved for your decision record and the exercise

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on the path, or .venv is off Turn on .venv (Step 1); on macOS/Linux try python3
Missing library (torch) or (peft) The libraries aren't in this Python Check for (.venv), then repeat Step 1
unit12/ft_data/train.jsonl not found The dataset wasn't created, or you're in the wrong folder Run Step 2 from the course folder
Couldn't load Qwen/Qwen3-0.6B with a connection error The download is blocked by a network or proxy Set HTTPS_PROXY as IT tells you, or ask IT to allow huggingface.co; --sample still works
KeyError: 'qwen3' Your transformers is too old for Qwen3 pip install -U transformers
The run is very slow or the computer freezes Not enough memory, or many apps open Close apps; try --limit 10 --steps 20
You expected a key prompt None is needed This topic uses no account or API key
Fine-tuned accuracy is lower than few-shot Too few steps, too many, or a hard test set That is a valid result. Try --steps 30 and --steps 120 and compare

The SAP way

As of October 2026, SAP's AI stack offers four relevant routes. They map onto the ladder, cheapest first.

Optimize the prompt in the generative AI hub

SAP's Developer Center tutorial on prompt optimization takes a prompt template from the Prompt Registry, a labelled dataset in your object store (records with input and answer fields), a metric to maximise such as json_exact_match, and a target model. The run saves the optimized prompt back to the Prompt Registry and tracks its metrics. It needs the Extended plan of SAP AI Core.

This is rung 1 and 2 done by machine, against your evaluation data. For the routing task above, it is what to try before any training.

Use a model that learns from examples at run time

SAP's AI Core documentation describes SAP-RPT-1 as solving predictive tasks on tables "without requiring any training or fine-tuning", through in-context learning: example rows travel with the request. For tabular predictions, that removes the fine-tuning question entirely. The tabular AI topic later in this unit goes deeper.

Use a model SAP has already tuned

The same documentation describes SAP-ABAP-1 as a foundation model built by SAP and fine-tuned on a large amount of ABAP code, generally available in the generative AI hub since its December 2025 release note. This is the pattern most customers should expect: SAP or a model provider does the fine-tuning, and you consume the result through the hub like any other model. Unit 13 covers AI-assisted SAP engineering.

Train and serve your own model on SAP AI Core

When a custom model is justified, SAP AI Core runs the training itself. SAP Learning's introduction describes the pieces:

Piece Role
Object store Holds the training data and, later, the trained model
Template and executable Define the training pipeline's containers, inputs, outputs and infrastructure
Configuration Binds an executable to a specific dataset and parameters for a run
Execution One training run, with status, logs and metrics in SAP AI Launchpad
Resource plan The machine size: CPUs, GPUs and memory

SAP AI Core's service documentation lists GPU resource plans: Infer-S, Infer-M and Infer-L with a "standard GPU", and Train-L with an "advanced GPU", at least 47 GB of memory and 5 CPU cores. Exact hardware depends on the hyperscaler. After training, SAP Learning says the model is registered in SAP AI Core and ready for deployment; serving then uses an inference resource plan. The Unit 12 setup topic links SAP's tutorial for serving an open model through Ollama on SAP AI Core.

flowchart LR
  D[(Object store<br/>JSONL dataset)] --> X[Execution<br/>Train-L plan]
  X --> M[(Object store<br/>adapter or model)]
  M --> S[Deployment<br/>Infer-S/M/L plan]
  S --> A[Your CAP app<br/>or agent]

This is a sketch of the flow, not something to run now.

Build vs. SAP

Situation Choose Why
Labels or format drift, hub models already used Prompt optimization in the generative AI hub Same evaluation data, no training pipeline to own
Predictions on SAP tables SAP-RPT-1 (Unit 12 tabular topic) In-context learning; no training
ABAP code understanding SAP-ABAP-1 in the hub SAP already tuned it
Narrow, high-volume task; prompts plateaued; data may not leave your boundary LoRA on an open model, trained and served on SAP AI Core or your own infrastructure Short prompts, a small model, data stays inside
Learning and pilots LoRA on a laptop, as in this topic Free, fast feedback on whether tuning helps at all
Knowledge that changes Not fine-tuning: retrieval or tools Weights go stale; data fetched per request doesn't

Production concerns

  • Data protection. A training set built from SAP notes may hold names, customer numbers and free text about people. Apply the same classification and minimisation as in data security and PII. Treat the adapter like the data it learned from: store it, share it and delete it under the same rules.
  • Authorizations. A tuned model has no SAP authorizations of its own. Whatever it routes or drafts, the action still goes through the user's or agent's SAP permissions, as in agent permissions and SAP authorizations.
  • Evaluation. Keep the test set frozen and versioned. Report accuracy per route, not only overall: a model can score well while failing one small team's route. Re-run the evaluation harness on every retrain.
  • Safety regression. Tuning can weaken behaviour you relied on, such as refusing out-of-scope requests. Include such cases in the evaluation set and re-run your red-team tests.
  • Cost. Compare total cost: labelling hours, training runs, a GPU deployment that runs even when idle, and retraining. Against that, count tokens saved per call times volume (see token economics).
  • Life cycle. Version the dataset, the training code, the base model and the adapter together. Plan for the base model's retirement: OpenAI's fine-tuning wind-down, with new jobs ending on 6 January 2027 for active customers, shows platform terms can change under a tuned model.
  • Clean core. Fine-tuning happens outside S/4HANA. The tuned model sits in a side-by-side app or agent and reaches SAP through released APIs, never by modifying SAP itself.

Pitfalls

  • Skipping the baseline. Without prompt scores, a fine-tuned 90% sounds good even if a few-shot prompt reached 92%.
  • Testing on training phrasings. The model looks perfect and fails on real notes. Hold out phrasings, not only rows.
  • Tuning for knowledge. Facts baked into weights go stale and can't be cited.
  • Inconsistent labels. If two clerks would label the same note differently, the model learns the noise. Have a second person check a sample.
  • Mismatched serving format. Training with one system instruction and serving with another quietly lowers accuracy.
  • Over-training. Training loss keeps falling while validation loss rises: the model is memorising. Use fewer steps.
  • Losing the base model. An adapter only works with the exact base model version it was trained on. Record both.

Exercise

Find out whether fine-tuning helps when examples are scarce. The result goes into your Unit 12 decision log and feeds the buy-vs-build topic at the end of the unit.

  1. Open unit12/make_ft_data.py and change make(32, slice(0, 4), rng) in main() to make(10, slice(0, 4), rng). That gives 50 training examples, the starting size OpenAI's guide suggests.
  2. Run python unit12/make_ft_data.py. Check it now says Wrote 50 examples for train.
  3. Run python unit12/finetune_lora.py --steps 30. Training runs on 50 examples instead of 160.
  4. Open unit12/model_log.md from the Unit 12 setup exercise (create it if it doesn't exist) and add a section ## Fine-tuning decision with a table: setup, training examples, accuracy, valid JSON, prompt tokens. Fill in the rows from this run and from Step 5.
  5. Under the table, write two sentences: would you fine-tune for this task, and what would make you change your mind?
  6. Change the number back to 32, run make_ft_data.py again, and commit model_log.md.

Done when model_log.md has a ## Fine-tuning decision table with real numbers for 50 and 160 training examples, and a two-sentence decision that names the best prompt's score.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1An agent keeps quoting last month's credit limits for customers. Which fix fits?

    Answer: B. The failure is knowledge that changes, so the fix is in the context: fetch the value per request. Any fine-tune on credit limits would be stale by next month.
  2. 2In finetune_lora.py, why are the prompt tokens labelled -100 during training?

    Answer: D. -100 tells the loss to ignore those positions. The model learns to produce the answer for a given prompt instead of learning to reproduce prompts.
  3. 3Why do the test notes use phrasings that never appear in training?

    Answer: B. If test notes repeat training phrasings, a model that memorised wording would score well and then fail on real notes. Held-out phrasings measure generalisation.
  4. 4What does LoRA train?

    Answer: C. LoRA freezes the original weights and trains low-rank matrices A and B beside selected layers, here the attention projections. That is why the adapter is small and works only with its base model.
  5. 5Your run shows few-shot at 88% accuracy and 260 prompt tokens, LoRA at 89% and 70 tokens. What is the soundest reading?

    Answer: C. Accuracy is about equal, so the difference is tokens per call. At high volume that may justify the training pipeline; at low volume the few-shot prompt is simpler to own.
  6. 6Training loss keeps falling while validation loss rises. What is happening, and what do you do?

    Answer: A. Diverging losses are the classic sign of overfitting. Fewer steps, or more varied examples, keep the model general.
  7. 7A customer wants a custom model trained inside SAP. Which SAP setup fits?

    Answer: C. SAP AI Core runs training workflows on resource plans, and Train-L is the plan with an advanced GPU. Prompt optimization changes prompts, not weights, and SAP-RPT-1 needs no training.
  8. 8Which production control matters most for an adapter trained on real order notes?

    Answer: D. An adapter is derived from its training data, so it inherits that data's classification. It needs no SAP authorizations of its own; actions still go through the user's or agent's permissions.
  9. 9The base model behind your adapter is announced for retirement. What should already be in place?

    Answer: B. Adapters work only with their exact base model. With the dataset, code and test set versioned, retraining on a new base is a routine run you can score against the old result.

Sources

  • Supervised fine-tuning (OpenAI API documentation) — use cases (classification, specific formats, instruction-following failures); minimum 10 examples, start with 50 well-crafted demonstrations; JSONL chat messages format; set up evals first and hold out data; OpenAI is winding down the fine-tuning platform
  • Deprecations (OpenAI API documentation) — fine-tuning availability: no new training for organizations that never fine-tuned (7 May 2026); job creation removed for organizations without recent fine-tuned inference (2 July 2026); last date for new jobs for active customers 6 January 2027; inference continues until the base model is deprecated
  • Optimizing LLM accuracy (OpenAI API documentation) — context optimization (what the model knows) vs. LLM optimization (how it behaves); start with prompt engineering and an evaluation set; fine-tuning for consistency, format and efficiency; RAG for knowledge; the two are additive
  • LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., arXiv 2106.09685) — freeze pretrained weights, train small low-rank matrices; about 10,000x fewer trainable parameters and 3x less GPU memory than full fine-tuning of GPT-3 175B; no added inference latency
  • PEFT quicktour (Hugging Face documentation) — LoraConfig with task_type CAUSAL_LM, r, lora_alpha, lora_dropout, target_modules; get_peft_model; a 1B model trains about 0.04% of parameters in the example; save_pretrained stores only the adapter (6.3 MB example vs ~700 MB base)
  • Qwen/Qwen3-0.6B (Hugging Face model card) — 0.6B parameters; Apache-2.0 licence; transformers 4.51.0 or newer; enable_thinking=False in apply_chat_template switches thinking off; model.safetensors 1.5 GB
  • PyTorch: Get started — install selector for OS, package and compute platform (choose CPU); macOS command pip3 install torch; stable PyTorch needs Python 3.10 or later
  • SAP AI Core (help.sap.com service documentation, PDF) — resource plans Starter, Basic, Basic.8x, Infer-S/M/L (standard GPU), Train-L (min. 47 GB, 5 CPU cores, advanced GPU); generative AI hub only in the extended plan; SAP-RPT-1 works without training or fine-tuning via in-context learning; SAP-ABAP-1 is fine-tuned on a large amount of ABAP code (GA in generative AI hub, 2025-12-22)
  • Training an ML model (SAP Learning, Introduction to SAP AI Core) — datasets in a hyperscaler object store; templates define executables; configurations bind executables and inputs; executions run the workflow; resource plans bundle CPUs, GPUs and memory; trained model registered in SAP AI Core and ready for deployment
  • Optimize prompts in generative AI hub (SAP Developer Center tutorial) — prompt template from the Prompt Registry, labeled dataset with input and answer fields in the object store, a metric such as json_exact_match, a target model; optimized prompt saved to the Prompt Registry; Extended plan required

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in