Orchestrate

Quantization and local models

Understand how quantization shrinks a model to fit a laptop or a small GPU, measure what it costs in accuracy on an SAP-shaped task, and pick a level with evidence.

Updated Oct 8, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

A language model is a very long list of numbers, called weights. A model with 8 billion weights stored at full precision takes about 16 GB. That is too big for most laptops and for the smallest cloud GPUs.

Quantization stores each weight with fewer bits, the way you might round prices to whole euros. The file gets two to four times smaller and usually faster to run. The model also gets a little less precise.

The skill is knowing how much rounding a task can take. Eight-bit models are usually almost indistinguishable from the original. Four-bit models are the common default for running models locally. Below four bits, quality tends to fall off.

You don't guess. You run the same test on two or three versions and look at the numbers.

Why it matters to the business

Quantization decides where a model can run and what it costs.

  • Hardware cost. In llama.cpp's published table, an 8-billion-parameter model shrinks from 32.1 GB in its original form to 4.9 GB at a common four-bit level. That is the difference between a large GPU server and a single modest GPU or a good laptop.
  • Data stays put. A model small enough to run inside your own boundary means prompts with customer or supplier data never leave it. That is often why teams look at local models in the first place.
  • Speed. Smaller weights mean less data to move per word generated. The same table shows text generation more than twice as fast at four bits as at sixteen.
  • Risk. Too much rounding produces a model that looks fine in a demo and makes more mistakes on real cases. Those mistakes are often quiet: a wrong route, a misread code.

Take the running example from order-to-cash. Clerks write short notes on blocked sales orders: "customer over the credit limit", "no price for material TG11", "customer asked us to hold delivery". A small local model routes each note to credit management, master data, pricing, logistics or customer service.

A four-bit version runs on the clerks' standard laptops. An eight-bit version needs more memory but may route a few more notes correctly. The business question is not "which is better?" but "is the accuracy difference worth the hardware?" Only a measurement on your own notes answers that.

How SAP does it

SAP offers two routes, and quantization matters differently in each. Both statements are as of October 2026.

  • Generative AI hub in SAP AI Core. You call models through an API, and someone else runs them. SAP's Python SDK documentation lists open-weight models such as meta--llama3.1-70b-instruct and mistralai--mistral-small-instruct next to commercial ones. You don't choose a quantization level, and the pages we opened don't say which precision is served. You judge the model by testing its answers.
  • Your own model on SAP AI Core. SAP's developer tutorial shows how to run Ollama, the same local model runner this unit uses, as a custom serving container. It uses the infer.s resource plan and needs an SAP AI Core instance on the Standard or Extended plan. Here you pick the model file, so you pick the quantization level, exactly as on a laptop.

There is no SAP-specific quantization tool in the sources we checked. The quantization itself happens in open-source tools such as llama.cpp before the model reaches SAP AI Core.

A decision guide: how much to shrink

The bit counts and the effects below come from llama.cpp's quantize documentation and Hugging Face's guide to quantization methods. Sizes are for an 8-billion-parameter model.

Level Bits per weight Size (8B model) What to expect Typical use
16-bit (F16, BF16) 16 about 15 GiB The reference quality Evaluation baseline, fine-tuning, large GPUs
8-bit (Q8_0) 8.5 about 8 GiB Very close to 16-bit When memory allows and accuracy matters most
5- to 6-bit (Q5_K_M, Q6_K) 5.7 to 6.6 5.3 to 6.1 GiB A middle ground A small step down from 8-bit
4-bit (Q4_K_M) 4.9 about 4.6 GiB Relatively high accuracy, with a measurable drop The usual default for local models
2- to 3-bit (Q2_K, Q3_K_M) 3.2 to 4.0 3 to 3.7 GiB A noticeable drop, especially at 2-bit Only when nothing else fits, and only after testing

Two rules of thumb follow:

  1. Start at 4-bit, test 8-bit next to it. If 8-bit wins clearly on your task, pay for the memory.
  2. Compare sizes, not just levels. A bigger model at 4 bits can beat a smaller model at 16 bits that takes the same memory. Test both.

Questions to ask

Your team

  • Which quantization level are we running, and against which full-precision baseline did we test it?
  • What is the accuracy difference on our own examples, not on a public benchmark?
  • How much memory does the model need with our longest prompts, not only the file size?
  • Who approved downloading this model file, and from where?

Your vendor or partner

  • For a hosted model: at what precision is it served, and do you change it without notice?
  • For a model you deliver to us: which tool and settings produced the quantized file, and can we reproduce it?
  • What evaluation did you run after quantizing, and on what data?

Your infrastructure team

  • What GPU memory do we have per server, and how many models must share it?
  • Can laptops in the target group hold a 0.5 to 5 GB model in memory alongside their other work?

Common misconceptions

  • "Quantized means low quality." Eight-bit models are usually very close to the original. The drop at four bits is often small for narrow tasks. It is a measured trade-off, not a downgrade by definition.
  • "Smaller is always faster." Generation speed rises as the file shrinks, but reading a long prompt stays about the same in llama.cpp's table. Memory, hardware and context length also matter.
  • "The file size is the memory we need." The runner also keeps the conversation in memory. Long prompts can add a lot on top of the weights.
  • "We'll quantize when we import the model." Ollama's documentation says it does not quantize GGUF models during import. You choose a quantized file or make one with llama.cpp first.
  • "A public benchmark tells us which level to use." Hugging Face's own guide says to benchmark on your task and hardware. A level that is fine for chat can still misroute an SAP document type.

Key terms

  • Weight: one of the learned numbers inside a model. Billions of them make up the model file.
  • Precision: how many bits store each number. More bits mean finer steps between values.
  • Quantization: storing weights with fewer bits, plus a few scale values to map them back.
  • Bits per weight: the average storage per weight, including the scales. Q8_0 averages 8.5, not 8.
  • GGUF: a single model file format used by llama.cpp and Ollama, holding the weights and metadata.
  • Calibration: running sample text through the model while quantizing, so the most important weights are protected.
  • Context (KV cache): the memory a model uses to hold the current conversation.
  • Baseline: the full-precision model you compare quantized versions against.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does quantization change about a model?

    Answer: B. Quantization keeps the same weights but stores them more coarsely. The file gets two to four times smaller. Retraining on your data is fine-tuning, a different technique.
  2. 2A team wants to run a model on clerks' laptops. Which starting point does this topic suggest?

    Answer: C. Four-bit is the usual default for local models, and 8-bit is very close to the original. Testing both on your own examples shows whether the extra memory pays off. Public benchmarks don't measure your routing task.
  3. 3Why is the model file size not the full memory requirement?

    Answer: A. The weights are only part of it. The conversation, or context, also takes memory, and long prompts can add a lot. Plan for the file plus room for the longest prompts you expect.
  4. 4When you call an open-weight model through SAP's generative AI hub, who picks the quantization level?

    Answer: C. In the hub, someone else runs the model and the pages we opened don't state its precision. You can't change it, so you test the answers. If you bring your own model to SAP AI Core, you pick the file and its level.
  5. 5Which question best tests a vendor's claim that their quantized model is "just as good"?

    Answer: B. "Just as good" is only meaningful against a baseline on data like yours. Format and popularity don't tell you how many notes it will misroute. Ask for the measured difference, and how to reproduce it.
  6. 6A bigger model at 4-bit takes the same memory as a smaller model at 16-bit. What should you do?

    Answer: D. Neither rule holds every time. A bigger model with more rounding can beat a smaller one stored precisely, or not. The same test on the same examples decides.
Deep layer · 40 min read

Mental model: rounding with a ruler that moves

Quantization stores each weight as a small integer plus a scale. To read the weight back, you multiply: weight ≈ integer × scale.

Think of measuring parts with a ruler that has only 15 marks (that is 4 bits: the integers −7 to +7). If one ruler must cover every part from a screw to a girder, the marks are far apart and small parts all read as zero. If each small batch of parts gets its own ruler, sized to the largest part in that batch, the marks are close together and the readings are accurate.

That is most of the field:

  • Fewer bits means fewer marks on the ruler, so more rounding error.
  • Smaller blocks means more rulers, so less error but more scales to store.
  • Outliers (a few very large weights) stretch a ruler and ruin precision for everything that shares it.
  • Smarter methods decide which weights matter most and protect them.

How it works

From floats to integers

Models are usually trained and published in 16-bit floating point. The Hugging Face GGUF documentation describes the two common 16-bit types: F16 is IEEE half precision, and BF16 is a shortened version of the 32-bit float that keeps its range.

To quantize a block of weights to b bits, the simplest method works like this:

  1. Find the largest absolute value in the block.
  2. Set the scale so that value maps to the top integer, for example 127 for 8 bits or 7 for 4 bits.
  3. Divide every weight by the scale and round to the nearest integer.
  4. Store the integers and the scale.

This is round-to-nearest quantization. The GGUF documentation describes Q8_0 and Q4_0 exactly this way, with blocks of 32 weights.

Why bits per weight is never a round number

Each block also stores its scale. With a 16-bit scale for every 32 weights, Q8_0 costs 8 + 16 / 32 = 8.5 bits per weight. llama.cpp's table lists Q8_0 at 8.5008 bits per weight.

The newer K-quants group blocks into super-blocks. The GGUF documentation gives Q4_K at 4.5 bits per weight, Q5_K at 5.5, Q6_K at 6.5625 and Q2_K at 2.625. llama.cpp's table lists the Q4_K_M preset at 4.89, higher than Q4_K alone, so that preset doesn't store every tensor at Q4_K.

Not all weights matter equally

The AWQ paper reports that protecting only 1% of the most important weights greatly reduces quantization error. It finds those weights by looking at the activations (the values flowing through the model on real text), not at the size of the weights.

That idea leads to two families of methods. Hugging Face's guide splits them this way:

Family Examples How Trade-off
On the fly, no calibration bitsandbytes, HQQ, torchao Quantize while loading Easy; bitsandbytes is aimed mainly at NVIDIA GPUs
Calibration-based GPTQ, AWQ Run sample text first, then quantize Often the best 4-bit accuracy; GPTQ can overfit its calibration data
GGUF with an importance matrix llama.cpp llama-imatrix Collect importance from calibration text Used for low-bit GGUF files

Hugging Face's guide reports calibration times of about 20 minutes for GPTQ and about 10 minutes for AWQ, for an 8-billion-parameter model on one A100 GPU.

The GGUF pipeline

Ollama and llama.cpp use GGUF files. The Qwen documentation shows the usual pipeline, and Ollama's import page says Ollama doesn't quantize GGUF models during import, so quantizing happens before Ollama sees the file.

flowchart LR
  H[Published model<br/>BF16 weights] --> C[convert to GGUF<br/>BF16]
  C --> Q[llama-quantize<br/>Q8_0, Q4_K_M]
  T[Calibration text<br/>optional] --> I[llama-imatrix]
  I -.-> Q
  Q --> M[Modelfile<br/>FROM file.gguf]
  M --> O[ollama create<br/>then run]

The Qwen documentation also suggests calibration text that represents your target domain. For SAP work that could be a sample of your own notes or document texts, cleared for that use.

Memory: weights plus context

llama.cpp's README says models are fully loaded into memory, so you need RAM (or GPU memory) for the whole file. On top of that, the runner keeps the KV cache: the model's working memory for the current conversation. It grows with the context length.

Ollama uses a 4,096-token context by default, according to its FAQ. The same FAQ describes KV cache quantization through the OLLAMA_KV_CACHE_TYPE setting:

Value Memory compared with f16 Effect described by Ollama
f16 (default) 1x High precision
q8_0 about 1/2 Very small loss; usually no noticeable impact
q4_0 about 1/4 Small to medium loss, more noticeable at long contexts

It only works when Flash Attention is on, which Ollama enables automatically when the hardware supports it. It is a global setting: every model on that Ollama server uses it.

Speed: where the time goes

llama.cpp's table for Llama 3.1 8B shows two speeds:

Level Size (GiB) Reading the prompt (tokens/s) Writing the answer (tokens/s)
F16 14.96 923 29
Q8_0 7.95 865 51
Q4_K_M 4.58 822 72
Q2_K 2.95 784 80

Writing the answer gets much faster as the file shrinks, because each new token has to read every weight. Reading the prompt stays roughly flat. For SAP tasks with long inputs and short outputs, like routing a note, the gain from quantization is smaller than the answer-speed column suggests.

These numbers come from one machine in the README. Yours will differ, which is why you measure.

Build it yourself: measure the trade-off on your machine

You will do two things. First, quantize a made-up layer of weights yourself, so you can see rounding error and the block trick with your own eyes. Then run the same blocked-order routing test on three versions of one small model, at 4, 8 and 16 bits, and compare accuracy, speed and size.

Before you start: complete Set up your computer for this course and Set up for Unit 12: local models. They install Python, Ollama and the ollama library, and pull qwen3:0.6b. This walkthrough doesn't repeat those steps. It uses numpy from Set up for Unit 2.

flowchart LR
  S1[Step 1<br/>open folder] --> S2[Step 2<br/>quantize_demo.py]
  S2 --> S3[Step 3<br/>pull 3 versions]
  S3 --> S4[Step 4<br/>compare_quants.py]
  S4 --> S5[Step 5<br/>read the results]
  S5 --> S6[Step 6<br/>save in Git]

What you need

  • Your course folder orchestrate-course with its .venv, Ollama running, and qwen3:0.6b pulled, from the Unit 12 setup.
  • About 2.5 GB of free disk for two more versions of the model: 832 MB and 1.5 GB, per Ollama's tag list.
  • At least 4 GB of free memory. The largest version is 1.5 GB; Ollama loads one at a time.
  • 30 to 45 minutes. Downloads take most of it on a slow connection.
  • Cost: free. No account or key. The model has an Apache 2.0 licence, per its Ollama page.

Step 1: Open your course folder

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal with Terminal > New Terminal. Turn on the virtual environment if the prompt doesn't start with (.venv):

    Windows (PowerShell):

    .venv\Scripts\Activate.ps1

    macOS or Linux:

    source .venv/bin/activate
  3. Check that the unit12 folder exists. If it doesn't, create it:

    mkdir unit12

Run every command in this topic from the course folder, not from inside unit12.

Step 2: Quantize a layer yourself

This script needs no model and no internet. It makes a grid of about a million random weights shaped like a trained layer, adds 20 unusually large ones, and stores the grid five ways. Then it compares what the layer computes with the original and the rounded weights.

  1. In VS Code's file list, right-click unit12, choose New File, name it quantize_demo.py, paste the code below and save.
"""See what quantization does to numbers, then estimate how big a model file gets.

Run it from your course folder:
    python unit12/quantize_demo.py                  # quantize one made-up weight matrix
    python unit12/quantize_demo.py --params 8       # also size an 8-billion-parameter model
    python unit12/quantize_demo.py --no-outliers    # same test without a few large weights

Part 1 builds a made-up "layer" of weights and stores it four ways: 16-bit floats,
8-bit and 4-bit integers with ONE scale for the whole matrix, and 8-bit and 4-bit
integers with one scale per block of 32 weights (the idea behind GGUF's Q8_0 and Q4_0).
Part 2 turns bits per weight into file sizes. No model, no account, no internet.
"""
import argparse

import numpy as np

ROWS, COLS = 1024, 1024   # one made-up layer: about a million weights
BLOCK = 32                # weights that share one scale in the block methods

# Average bits per weight from llama.cpp's quantize README (Llama 3.1 8B table).
# They include the scales, which is why Q8_0 is 8.5 and not 8.
LLAMA_CPP_BPW = {"F16": 16.0, "Q8_0": 8.5, "Q6_K": 6.56, "Q5_K_M": 5.70, "Q4_K_M": 4.89, "Q3_K_M": 4.00}


def make_weights(rng, outliers: bool) -> np.ndarray:
    """Small random weights, like a trained layer; optionally a few much larger ones."""
    w = rng.normal(0.0, 0.02, size=(ROWS, COLS)).astype(np.float32)
    if outliers:
        spots = rng.choice(w.size, size=20, replace=False)   # 20 of about a million
        w.flat[spots] = rng.choice([-1.0, 1.0], size=20) * 0.5
    return w


def quantize(w: np.ndarray, bits: int, block: int | None) -> tuple:
    """Round weights to signed integers with `bits` bits; return the rebuilt weights and bits per weight.

    block=None: one scale for the whole matrix. block=32: one scale per 32 weights.
    Each scale is stored as a 16-bit float, so it adds 16 bits per block.
    """
    levels = 2 ** (bits - 1) - 1                    # 127 for 8-bit, 7 for 4-bit
    flat = w.reshape(-1, block) if block else w.reshape(1, -1)
    scale = np.abs(flat).max(axis=1, keepdims=True) / levels   # biggest value maps to the top level
    scale[scale == 0] = 1.0
    q = np.clip(np.round(flat / scale), -levels, levels)       # the small integers that get stored
    rebuilt = (q * scale.astype(np.float16).astype(np.float32)).reshape(w.shape)
    scale_bits = 16 * flat.shape[0]
    return rebuilt, bits + scale_bits / w.size


def relative_error(original: np.ndarray, approx: np.ndarray) -> float:
    """Size of the error compared with the size of the original, in percent."""
    return 100 * float(np.linalg.norm(original - approx) / np.linalg.norm(original))


def part_one(seed: int, outliers: bool) -> None:
    rng = np.random.default_rng(seed)
    w = make_weights(rng, outliers)
    x = rng.normal(0.0, 1.0, size=(COLS, 64)).astype(np.float32)   # 64 made-up inputs
    y = w @ x                                                      # what the layer outputs

    methods = [
        ("16-bit float", lambda: (w.astype(np.float16).astype(np.float32), 16.0)),
        ("8-bit, one scale", lambda: quantize(w, 8, None)),
        ("8-bit, blocks of 32", lambda: quantize(w, 8, BLOCK)),
        ("4-bit, one scale", lambda: quantize(w, 4, None)),
        ("4-bit, blocks of 32", lambda: quantize(w, 4, BLOCK)),
    ]
    label = "with 20 large outlier weights" if outliers else "without outliers"
    print(f"Part 1: one made-up layer, {ROWS} x {COLS} weights, {label}\n")
    print(f"{'Method':<22}{'Bits/weight':>12}{'Size':>10}{'Weight error':>14}{'Output error':>14}")
    for name, run in methods:
        rebuilt, bpw = run()
        size_mb = w.size * bpw / 8 / 1e6
        print(f"{name:<22}{bpw:>12.2f}{size_mb:>8.2f} MB"
              f"{relative_error(w, rebuilt):>13.2f}%{relative_error(y, rebuilt @ x):>13.2f}%")
    print("\nOutput error compares what the layer computes with the original weights and with the")
    print("rebuilt ones. Smaller is better. 32-bit floats would take 4.19 MB.")


def part_two(params_billion: list) -> None:
    print("\nPart 2: weight file size = parameters x bits per weight / 8")
    print("(a floor: real files add metadata, and the runner needs memory for the conversation)\n")
    header = f"{'Parameters':<12}" + "".join(f"{name:>10}" for name in LLAMA_CPP_BPW)
    print(header)
    for p in params_billion:
        row = f"{p:>6.2f} B    "
        for bpw in LLAMA_CPP_BPW.values():
            row += f"{p * 1e9 * bpw / 8 / 1e9:>8.2f}GB"
        print(row)


def main() -> None:
    parser = argparse.ArgumentParser(description="See what quantization does to weights and file size.")
    parser.add_argument("--params", type=float, action="append",
                        help="model size in billions of parameters (repeat for several)")
    parser.add_argument("--no-outliers", action="store_true", help="leave out the large weights")
    parser.add_argument("--seed", type=int, default=7, help="change for different random weights")
    args = parser.parse_args()

    part_one(args.seed, outliers=not args.no_outliers)
    part_two(args.params or [0.752, 8.0, 70.0])   # 0.752 B is qwen3:0.6b's parameter count


if __name__ == "__main__":
    main()
  1. Run it (the same on every system):

    python unit12/quantize_demo.py

What success looks like (the numbers are the same on every computer, because the random seed is fixed):

Part 1: one made-up layer, 1024 x 1024 weights, with 20 large outlier weights

Method                 Bits/weight      Size  Weight error  Output error
16-bit float                 16.00    2.10 MB         0.02%         0.02%
8-bit, one scale              8.00    1.05 MB         5.65%         5.64%
8-bit, blocks of 32           8.50    1.11 MB         0.55%         0.55%
4-bit, one scale              4.00    0.52 MB        88.26%        88.38%
4-bit, blocks of 32           4.50    0.59 MB         9.88%         9.84%

Output error compares what the layer computes with the original weights and with the
rebuilt ones. Smaller is better. 32-bit floats would take 4.19 MB.

Part 2: weight file size = parameters x bits per weight / 8
(a floor: real files add metadata, and the runner needs memory for the conversation)

Parameters         F16      Q8_0      Q6_K    Q5_K_M    Q4_K_M    Q3_K_M
  0.75 B        1.50GB    0.80GB    0.62GB    0.54GB    0.46GB    0.38GB
  8.00 B       16.00GB    8.50GB    6.56GB    5.70GB    4.89GB    4.00GB
 70.00 B      140.00GB   74.38GB   57.40GB   49.88GB   42.79GB   35.00GB
  1. Read the table. Three lessons are in it:

    • Outliers ruin a shared scale. With one scale for the whole grid, 20 large weights out of a million stretch the ruler. At 4 bits almost every normal weight rounds to zero, and the output error is 88%.
    • Blocks fix most of it. One scale per 32 weights costs 0.5 extra bits per weight and brings the 4-bit error down to about 10%. This is why GGUF formats use blocks.
    • 8-bit with blocks is almost free. Half the size of 16-bit, with an error near half a percent.
  2. Run it once without the outliers and compare the "one scale" rows:

    python unit12/quantize_demo.py --no-outliers

    The 4-bit one-scale error drops to about 20%, while the block rows barely change. Blocks matter most when weights are uneven, which is the situation AWQ describes in real models.

  3. Check Part 2 against reality. For the 8B model, the Q4_K_M estimate of 4.89 GB matches the 4.9 GB that llama.cpp reports. For qwen3:0.6b (752 million parameters), Ollama's tag list shows 1.5 GB at fp16 and 832 MB at q8_0, close to the estimates. The q4_K_M file is 523 MB, more than the 0.46 GB estimate, because the 4.89 average comes from a much larger model. Treat Part 2 as a first estimate, then check the real file.

Step 3: Pull the three versions of the model

Ollama's library publishes the same model at several quantization levels. For qwen3:0.6b, the tag list shows:

Tag Quantization Download
qwen3:0.6b-q4_K_M Q4_K_M 523 MB (the same file as qwen3:0.6b)
qwen3:0.6b-q8_0 Q8_0 832 MB
qwen3:0.6b-fp16 F16 1.5 GB

The tag list gives qwen3:0.6b and qwen3:0.6b-q4_K_M the same ID, so the model you pulled in the setup is already the 4-bit version.

  1. Pull the three tags, one command at a time (the same on every system):

    ollama pull qwen3:0.6b-q4_K_M
    ollama pull qwen3:0.6b-q8_0
    ollama pull qwen3:0.6b-fp16

    The first should finish quickly, because Ollama already has that file. The other two download.

  2. List what you have:

    ollama list

What success looks like: lines for all three tags, with sizes of about 523 MB, 832 MB and 1.5 GB, plus your earlier qwen3:0.6b.

Step 4: Run the same test on all three

The script sends 20 made-up clerk notes to each version, one at a time, with the same instructions and settings. Each answer must be JSON naming one of five routes. It counts correct routes and valid JSON, and records Ollama's own timings.

  1. In VS Code, right-click unit12, choose New File, name it compare_quants.py, paste the code below and save.
"""Compare the same small model at three quantization levels on one SAP-shaped task.

Run it from your course folder:
    python unit12/compare_quants.py --sample             # no model: shows the output format
    python unit12/compare_quants.py                      # q4_K_M, q8_0 and fp16 of qwen3:0.6b
    python unit12/compare_quants.py --limit 5            # quick try on the first 5 notes
    python unit12/compare_quants.py --models qwen3:1.7b-q4_K_M qwen3:1.7b-q8_0

The task: route a clerk's note about a blocked sales order to one of five teams.
The notes and routes are made up; the routes are this course's labels, not SAP codes.
Everything runs in Ollama on this computer. Results are saved to unit12/quant_results.csv.
"""
import argparse
import csv
import json
import re
import sys
import time
from pathlib import Path

DEFAULT_MODELS = ["qwen3:0.6b-q4_K_M", "qwen3:0.6b-q8_0", "qwen3:0.6b-fp16"]
ROUTES = ["CREDIT", "MASTER_DATA", "PRICING", "STOCK", "CUSTOMER_HOLD"]
RESULTS = Path("unit12") / "quant_results.csv"

SYSTEM = ("Route the clerk's note about a blocked SAP sales order. "
          'Reply with JSON only: {"route": "<ROUTE>"}. ROUTE is one of '
          + ", ".join(ROUTES) + ". "
          "CREDIT: credit limit or unpaid invoices. MASTER_DATA: missing or wrong customer data. "
          "PRICING: wrong or missing price or discount. STOCK: not enough material. "
          "CUSTOMER_HOLD: the customer asked to wait.")

# 20 made-up notes, four per route, with the route a person would choose.
NOTES = [
    ("Order 4711 stuck, customer 10100001 is 6,200 EUR over the credit limit.", "CREDIT"),
    ("Credit check failed on 4712. Invoices overdue since July.", "CREDIT"),
    ("Finance says the customer still hasn't paid the August statement, so 4713 is held.", "CREDIT"),
    ("Risk team flagged 10100007; payment behaviour got worse and 4714 waits on them.", "CREDIT"),
    ("4721 blocked: ship-to address for 10100002 has no postal code.", "MASTER_DATA"),
    ("Tax number missing on customer 10100003, so 4722 can't be released.", "MASTER_DATA"),
    ("Incoterms not maintained for 10100004; 4723 stays blocked.", "MASTER_DATA"),
    ("Customer record was created yesterday and half the fields are empty, 4724 on hold.", "MASTER_DATA"),
    ("No price found for material TG11 on 4731.", "PRICING"),
    ("4732: manual discount of 18% is above tolerance, needs approval.", "PRICING"),
    ("Condition record for TG12 expired, 4733 shows zero price.", "PRICING"),
    ("Line for TG14 on 4734 came through at 0.00 EUR, someone has to fix the amount.", "PRICING"),
    ("Not enough stock of TG11 for 4741, short by 40 pieces.", "STOCK"),
    ("4742: availability check confirms only 12 of the ordered quantity.", "STOCK"),
    ("Material TG13 back-ordered, 4743 waits for the next receipt.", "STOCK"),
    ("Warehouse is out of TG15 until the supplier delivers, 4744 sits there.", "STOCK"),
    ("10100005 asked us to hold 4751 until further notice.", "CUSTOMER_HOLD"),
    ("Customer called: do not ship 4752 before October.", "CUSTOMER_HOLD"),
    ("4753 on hold at the customer's request, their warehouse is full.", "CUSTOMER_HOLD"),
    ("Email from 10100006: please park 4754, they are moving sites.", "CUSTOMER_HOLD"),
]
# Reorder so the routes take turns: then --limit 5 still tests every route once.
NOTES = [NOTES[i + 4 * r] for i in range(4) for r in range(len(ROUTES))]

# What --sample prints. Made-up numbers in a realistic shape, not a measurement.
SAMPLE_ROWS = [
    {"model": "qwen3:0.6b-q4_K_M", "size_mb": 523, "accuracy": 0.70, "valid_json": 0.95,
     "tokens_per_s": 38.2, "load_s": 1.4, "seconds_per_note": 0.6},
    {"model": "qwen3:0.6b-q8_0", "size_mb": 832, "accuracy": 0.75, "valid_json": 1.00,
     "tokens_per_s": 29.5, "load_s": 1.9, "seconds_per_note": 0.8},
    {"model": "qwen3:0.6b-fp16", "size_mb": 1500, "accuracy": 0.75, "valid_json": 1.00,
     "tokens_per_s": 17.1, "load_s": 3.2, "seconds_per_note": 1.3},
]


def read_route(text: str) -> tuple:
    """Return (route or None, whether the reply was valid JSON)."""
    try:
        route = json.loads(text).get("route", "")
        return (route if route in ROUTES else None), True
    except (json.JSONDecodeError, AttributeError):
        # Not clean JSON: still accept a route name if one appears, but count it as invalid JSON.
        found = re.search("|".join(ROUTES), text.upper())
        return (found.group(0) if found else None), False


def model_sizes(client) -> dict:
    """Size on disk of every pulled model, in MB, as Ollama reports it."""
    return {m.model: round(m.size / 1e6) for m in client.list().models}


def run_model(client, model: str, notes: list) -> dict:
    """Ask one model to route every note; return accuracy, speed and load time."""
    correct = valid = answer_tokens = 0
    answer_ns = load_ns = 0
    started = time.perf_counter()
    for i, (note, expected) in enumerate(notes, start=1):
        response = client.chat(
            model=model,
            messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": note}],
            think=False,                                   # answer straight away
            options={"temperature": 0, "num_ctx": 2048},   # same settings for every model
        )
        route, is_json = read_route(response.message.content.strip())
        correct += route == expected
        valid += is_json
        answer_tokens += response.eval_count or 0
        answer_ns += response.eval_duration or 0
        load_ns += response.load_duration or 0
        print(f"  {i:>2}/{len(notes)}  expected {expected:<13} got {route or '(unreadable)'}")
    elapsed = time.perf_counter() - started
    return {
        "model": model,
        "accuracy": round(correct / len(notes), 2),
        "valid_json": round(valid / len(notes), 2),
        "tokens_per_s": round(answer_tokens / (answer_ns / 1e9), 1) if answer_ns else 0.0,
        "load_s": round(load_ns / 1e9, 1),
        "seconds_per_note": round(elapsed / len(notes), 2),
    }


def print_table(rows: list) -> None:
    print(f"\n{'Model':<22}{'Size MB':>9}{'Accuracy':>10}{'Valid JSON':>12}"
          f"{'Tokens/s':>10}{'Load s':>8}{'s/note':>8}")
    for r in rows:
        print(f"{r['model']:<22}{r['size_mb']:>9}{r['accuracy']:>10.0%}{r['valid_json']:>12.0%}"
              f"{r['tokens_per_s']:>10.1f}{r['load_s']:>8.1f}{r['seconds_per_note']:>8.2f}")


def save(rows: list) -> None:
    RESULTS.parent.mkdir(exist_ok=True)
    with RESULTS.open("w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=list(rows[0]))
        writer.writeheader()
        writer.writerows(rows)
    print(f"\nSaved to {RESULTS}")


def main() -> None:
    parser = argparse.ArgumentParser(description="Compare quantization levels of a local model.")
    parser.add_argument("--models", nargs="+", default=DEFAULT_MODELS, help="Ollama tags to compare")
    parser.add_argument("--limit", type=int, default=len(NOTES), help="use only the first N notes")
    parser.add_argument("--sample", action="store_true", help="no model: print sample results")
    args = parser.parse_args()

    if args.sample:
        print("Sample results (made up, no model was called):")
        print_table(SAMPLE_ROWS)
        return

    try:
        from ollama import Client, ResponseError
    except ModuleNotFoundError:
        sys.exit("The ollama library isn't installed. Run: pip install -r requirements.txt")

    client = Client()   # http://127.0.0.1:11434 unless OLLAMA_HOST says otherwise
    notes = NOTES[: max(1, args.limit)]
    try:
        sizes = model_sizes(client)
        missing = [m for m in args.models if m not in sizes]
        if missing:
            sys.exit("Pull these first, one at a time:\n" + "\n".join(f"  ollama pull {m}" for m in missing))
        rows = []
        for model in args.models:
            print(f"\n{model}")
            row = run_model(client, model, notes)
            row["size_mb"] = sizes[model]
            rows.append(row)
    except ConnectionError:
        sys.exit("Can't reach Ollama on this computer. Start the Ollama app "
                 "(Linux: sudo systemctl start ollama), then run this again.")
    except ResponseError as error:
        sys.exit(f"Ollama answered with an error: {error}")

    print_table(rows)
    save(rows)


if __name__ == "__main__":
    main()
  1. Run it with --sample first. This prints the output format without calling any model:

    python unit12/compare_quants.py --sample

What success looks like:

Sample results (made up, no model was called):

Model                   Size MB  Accuracy  Valid JSON  Tokens/s  Load s  s/note
qwen3:0.6b-q4_K_M           523       70%         95%      38.2     1.4    0.60
qwen3:0.6b-q8_0             832       75%        100%      29.5     1.9    0.80
qwen3:0.6b-fp16            1500       75%        100%      17.1     3.2    1.30

These numbers are invented to show the layout. Don't quote them.

  1. Try a quick real run on five notes, one per route:

    python unit12/compare_quants.py --limit 5

    You should see each model's name, then one line per note:

    qwen3:0.6b-q4_K_M
       1/5  expected CREDIT        got CREDIT
       2/5  expected MASTER_DATA   got MASTER_DATA
       3/5  expected PRICING       got PRICING
       4/5  expected STOCK         got STOCK
       5/5  expected CUSTOMER_HOLD got CUSTOMER_HOLD

    Your routes may differ. got (unreadable) means the answer named no route at all; that counts as wrong.

  2. Run the full test:

    python unit12/compare_quants.py

What success looks like: 20 lines per model, the comparison table with your own numbers, and Saved to unit12/quant_results.csv. On a laptop without a GPU this takes a few minutes.

Step 5: Read your results

Open unit12/quant_results.csv in VS Code, or look at the table in the terminal. Ask four questions:

  1. Accuracy. Is the 4-bit version clearly worse than fp16? With 20 notes, one note is 5 percentage points, so a gap of one or two notes is noise. A larger, repeatable gap is a signal.
  2. Valid JSON. Lower precision can break formatting before it breaks understanding. If valid JSON drops, the business process breaks even when the route is right.
  3. Tokens per second. Does the 4-bit version write faster on your machine? Compare with the llama.cpp pattern: big gains in answer speed, small gains in reading the prompt.
  4. Seconds per note. This is what a clerk feels. Routing answers are short, so this is mostly prompt reading and overhead.

A good conclusion names a choice and a reason, for example: "q4_K_M: same accuracy as fp16 on 20 notes, a third of the size, twice the speed. Re-test with 100 real notes before go-live."

Step 6: Save your work in Git

  1. Check what changed:

    git status

    You should see unit12/quantize_demo.py, unit12/compare_quants.py and unit12/quant_results.csv. Models stay in Ollama's own folder and never reach Git.

  2. Save:

    git add unit12/quantize_demo.py unit12/compare_quants.py unit12/quant_results.csv
    git commit -m "Unit 12: measure quantization levels on blocked-order routing"

How the code works

Part What it does
quantize() Finds the largest value per block, sets the scale, rounds to integers, then multiplies back so you can measure the error
scale_bits Adds 16 bits per block for the stored scale, which is why blocks cost 0.5 bits per weight at 32 weights per block
relative_error() Compares the size of the error with the size of the original, as a percentage
y = w @ x Runs 64 made-up inputs through the layer, so you see the error in what the layer computes, not only in the weights
LLAMA_CPP_BPW Bits per weight copied from llama.cpp's table, used for the file-size estimate
NOTES 20 made-up notes, reordered so the routes take turns
read_route() Accepts clean JSON; falls back to finding a route name but counts the answer as invalid JSON
client.chat(... think=False, temperature 0) The same instructions and settings for every model, so only the quantization differs
eval_count, eval_duration, load_duration Ollama's own counters, in nanoseconds, turned into tokens per second and load seconds
model_sizes() Reads the size of each pulled model from Ollama, and tells you which tags still need ollama pull

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on your path, or the terminal isn't in the course folder Reopen the course folder in VS Code; on macOS or Linux try python3; see Set up your computer
ModuleNotFoundError: No module named 'numpy' The virtual environment is off, or numpy isn't installed Turn on .venv (Step 1), then pip install -r requirements.txt
The ollama library isn't installed Same, for the ollama library Turn on .venv, then pip install -r requirements.txt
Can't reach Ollama on this computer The Ollama server isn't running Start the Ollama app; on Linux run sudo systemctl start ollama
Pull these first, one at a time A tag hasn't been downloaded, or was typed differently Run the ollama pull lines the script prints, then ollama list
ollama pull hangs, or fails with a TLS or proxy error Your network or proxy blocks Ollama's registry Try another network, or ask IT to allow ollama.com; the Unit 12 setup explains the proxy settings
The fp16 run is very slow, or the computer freezes Not enough free memory for the 1.5 GB version plus your other apps Close other apps, or run --models qwen3:0.6b-q4_K_M qwen3:0.6b-q8_0 without fp16
No API key is asked for anywhere Correct: local models need no key Nothing to do

The SAP way

As of October 2026, there are two ways to bring this into SAP's platform. Neither has an SAP-specific quantization feature in the sources we opened; quantization happens in open-source tools before the model reaches SAP.

Call a hosted model in the generative AI hub

SAP's Python SDK documentation for the generative AI hub lists open-weight models such as meta--llama3.1-70b-instruct, ibm--granite-13b-chat, mistralai--mistral-small-instruct and mistralai--mistral-large-instruct, next to models from Amazon, Anthropic, Google and OpenAI.

You don't choose the precision these are served at, and the documentation we opened doesn't state it. Treat each model as a black box and run your evaluation set against it, as in the choosing and calling LLMs topic. The hub itself is covered in SAP Generative AI Hub and the orchestration service.

Serve your own quantized model on SAP AI Core

SAP's developer tutorial "Using Custom models on SAP AI Core via Ollama" packages Ollama as a custom serving container. The serving template sets the resource plan with the label ai.sap.com/resourcePlan: infer.s, and the model is pulled into the running Ollama pod through the SAP AI API. The tutorial requires an SAP AI Core instance on the Standard or Extended plan, and it was last dated 13 June 2025.

Because the container runs Ollama, the quantization choice is the same tag choice you made in Step 3. The part of the serving template that matters here looks like this:

# Sketch only: needs an SAP AI Core instance (Standard or Extended plan) and a
# Docker registry; Set up for Unit 5 covers SAP AI Core. Based on SAP's Ollama tutorial.
metadata:
  labels:
    ai.sap.com/resourcePlan: infer.s   # the GPU plan the tutorial uses
# ...the rest of the serving template from SAP's tutorial...
# After deployment, pull the exact tag you tested on your laptop, for example
# qwen3:0.6b-q8_0, not a bare name that could resolve to a different level.

If you need a level that the Ollama library doesn't publish, make a GGUF file with llama.cpp (convert, then llama-quantize, optionally with an importance matrix from your own calibration text), import it with a Modelfile, and serve that. Ollama's import documentation is explicit that it won't quantize during import.

Build vs. SAP

Situation Choose Why
Need the strongest model, data may go to a hosted service Generative AI hub No hardware to size; quantization isn't your decision
Data must stay on a device or a site Quantized model on local hardware A 4-bit file fits laptops and small GPUs
Data must stay in your BTP landscape, open model is good enough Own model on SAP AI Core via Ollama You choose the file and level; SAP AI Core runs the container
Need a level or calibration the library doesn't offer llama.cpp llama-quantize with an importance matrix, then import Full control over the file, at the cost of building it
Fine-tuning a model on limited GPU memory bitsandbytes 4-bit with LoRA (QLoRA) Hugging Face's guide calls it the standard method for QLoRA

Production concerns

  • Evaluate per level, on your data. Hugging Face's guide ends with the advice to benchmark accuracy and speed on your own task and hardware. Keep the full-precision version as the baseline, and rerun the same evaluation set for every level and every model update.
  • Pin the exact file. Use the full tag (qwen3:0.6b-q8_0), not the bare name, and record the model ID from the registry. A bare name like qwen3 points at the library's default, the 8b model, which you didn't choose.
  • Supply chain. A model file is executable behaviour you download from the internet. Pull from a source your security team approves, record where each file came from, and check the licence (Apache 2.0 for Qwen3, per its Ollama page). Treat files you quantize yourself the same way: keep the command, the source file and the calibration text.
  • Calibration data is data. If you build an importance matrix from real SAP texts, those texts need the same approval as any other use of business data.
  • Memory planning. Size for the weights plus the KV cache at your longest real prompt. On a shared Ollama server, OLLAMA_KV_CACHE_TYPE applies to every model, so test all of them after changing it.
  • Authorizations. Quantization doesn't change who may see what. A local model that reads SAP data still needs the same SAP authorization checks in the code that fetches the data, as covered in agent permissions and SAP authorizations.
  • Clean core. Nothing here changes the SAP system. The model runs beside it and reads data through released APIs.
  • Cost. On SAP AI Core you pay for the resource plan while the deployment runs. A smaller quantized model may fit a smaller plan; confirm plan sizes and prices with SAP before committing.

Pitfalls

  • Comparing different models and calling it a quantization test. Change one thing at a time: same model, same prompt, same settings, different level.
  • Too few test cases. With 20 notes, one note is 5 points. Don't decide on a one-note difference.
  • Ignoring format errors. A model that routes correctly but breaks the JSON still breaks the process. Track valid output separately.
  • Trusting the bare model name. qwen3:0.6b is the q4_K_M file today, per the tag list. Write the level into your config explicitly.
  • Expecting Ollama to quantize on import. It doesn't for GGUF files. Quantize with llama.cpp first.
  • Going below 4 bits without calibration. Hugging Face's guide reports a noticeable drop at 2-bit, and the Qwen documentation notes that llama-quantize warns when 1- or 2-bit files are made without an importance matrix.
  • Sizing hardware from the file alone. The KV cache grows with context. Test with your longest realistic prompt.

Exercise

Answer the question this unit keeps coming back to: is a bigger model at lower precision better than a smaller one at higher precision? The result goes into your Unit 12 decision log and feeds the buy-vs-build topic at the end of the unit.

  1. Pull the 4-bit version of the next size up. Ollama's tag list gives it as 1.4 GB, close to the 1.5 GB of qwen3:0.6b-fp16:

    ollama pull qwen3:1.7b-q4_K_M
  2. Run both on the same 20 notes:

    python unit12/compare_quants.py --models qwen3:0.6b-fp16 qwen3:1.7b-q4_K_M
  3. Open unit12/model_log.md (create it if it doesn't exist) and add a section ## Quantization decision with a table: model tag, size MB, accuracy, valid JSON, tokens per second, seconds per note. Copy the rows from this run and from Step 4.

  4. Under the table, write two sentences: which tag would you deploy for the clerks' laptops, and what result would make you change your mind?

  5. Save the results file under a new name so the next run doesn't overwrite it, then commit both files.

    Windows (PowerShell):

    Copy-Item unit12\quant_results.csv unit12\quant_results_size_vs_precision.csv

    macOS or Linux:

    cp unit12/quant_results.csv unit12/quant_results_size_vs_precision.csv

    Then, on every system:

    git add unit12/model_log.md unit12/quant_results_size_vs_precision.csv
    git commit -m "Unit 12: size vs. precision decision"

Done when model_log.md has a ## Quantization decision table with at least four rows of your own measured numbers and a two-sentence choice, and both files are committed.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In quantize_demo.py, why does "4-bit, one scale" lose almost all precision when 20 outliers are present?

    Answer: B. The scale is set by the largest value. With one scale for a million weights, a few large ones make each step so wide that almost all normal weights round to zero. Blocks of 32 give each group its own scale, which is why the block row drops to about 10% error.
  2. 2Why does Q8_0 cost 8.5 bits per weight rather than 8?

    Answer: C. The integers take 8 bits each, and each block of 32 also stores its scale. Sixteen bits spread over 32 weights adds 0.5 bits per weight, matching llama.cpp's 8.5008.
  3. 3Your team needs a Q3_K_M version of a model the Ollama library doesn't publish at that level. What do you do?

    Answer: D. Ollama's import documentation says it doesn't quantize GGUF models during import, so the file must be quantized first. The KV cache setting only affects conversation memory, not weights. The hub doesn't let you pick a level.
  4. 4What does AWQ use to decide which weights to protect?

    Answer: B. The AWQ paper finds important weights from the activations, not from weight size, and it needs no backpropagation. Protecting about 1% of weights greatly reduces the error.
  5. 5In llama.cpp's table, Q4_K_M writes answers more than twice as fast as F16. Why does routing a note speed up much less?

    Answer: C. In the table, generation speed rises steeply as the file shrinks, but prompt processing stays roughly flat. A routing call reads a long instruction and writes a few tokens, so most of its time is in the part quantization helps least.
  6. 6compare_quants.py reports 80% accuracy for q4_K_M and 85% for fp16 on 20 notes. What do you conclude?

    Answer: B. With 20 notes, each note is 5 percentage points, so a one-note gap can come from chance. Run a larger set, ideally real notes cleared for testing, before spending money on memory.
  7. 7You deploy Ollama on SAP AI Core and configure the model as qwen3. What's the risk?

    Answer: A. A bare name points at whatever the library sets as default, which for qwen3 is the 8b model, a 5.2 GB download, per Ollama's library page. Pin the exact tag you evaluated, such as qwen3:0.6b-q8_0, so production runs what you tested.
  8. 8Your security team asks what changes for SAP authorizations when you switch from fp16 to a 4-bit model. What's the right answer?

    Answer: D. Quantization changes how weights are stored, not what data the system may read. Authorization checks belong in the code that fetches SAP data, whichever model level answers afterwards.

Sources

  • quantize README (llama.cpp tools, mirrored on docs.rs) — llama-quantize turns a high-precision GGUF (F32 or BF16) into a quantized one; list of I-quants, K-quants, Q8_0, F16; Llama 3.1 8B table of bits per weight, size and speed (F16 16.0 bpw, 14.96 GiB, 29.17 t/s; Q8_0 8.50, 7.95 GiB, 50.93 t/s; Q4_K_M 4.89, 4.58 GiB, 71.93 t/s; Q2_K 3.16); 8B 32.1 GB to 4.9 GB at Q4_K_M; models fully loaded into memory; --imatrix reduces accuracy loss
  • GGUF (Hugging Face Hub documentation) — binary format for fast loading, stores tensors plus standardized metadata; Q8_0 and Q4_0 are round-to-nearest with 32-weight blocks (legacy); Q4_K 4.5 bpw, Q5_K 5.5, Q6_K 6.5625, Q2_K 2.625; F16 and BF16 definitions
  • Selecting a quantization method (Hugging Face Transformers documentation) — on-the-fly methods (bitsandbytes, HQQ, torchao) vs calibration-based (GPTQ, AWQ); 8-bit about 2x memory saving and very close to bf16; 4-bit about 4x with relatively high accuracy; sub-4-bit noticeable drop; bitsandbytes is the standard for QLoRA; always benchmark on your own task and hardware
  • AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv 2306.00978) — protecting only 1% of salient weights greatly reduces quantization error; salient channels found from activations, not weight size; no backpropagation
  • Quantization with llama.cpp (Qwen documentation) — convert to GGUF in bf16, then llama-quantize to Q8_0, Q5_K_M or Q4_K_M; importance matrix from domain calibration text with llama-imatrix; perplexity comparison with llama-perplexity
  • qwen3 tags (Ollama model library) — qwen3 latest points to 8b (5.2 GB); qwen3:0.6b and qwen3:0.6b-q4_K_M share ID 7df6b6e09427 (523 MB); qwen3:0.6b-q8_0 832 MB; qwen3:0.6b-fp16 1.5 GB; 1.7b: q4_K_M 1.4 GB, q8_0 2.2 GB, fp16 4.1 GB
  • qwen3:0.6b-q8_0 (Ollama model library) — 752M parameters, quantization Q8_0, Apache License 2.0; the fp16 tag page shows 752M parameters, F16, 1.5 GB
  • FAQ (Ollama documentation) — default context 4096 tokens; Flash Attention used automatically when supported; OLLAMA_KV_CACHE_TYPE f16 (default), q8_0 about half the memory, q4_0 about a quarter with small-medium precision loss; global option for all models
  • Importing a model (Ollama documentation) — Ollama does not quantize GGUF models during import; quantize first with llama.cpp's llama-quantize; Modelfile FROM /path/to/file.gguf then ollama create
  • Using Custom models on SAP AI Core via Ollama (SAP Developer Center tutorial) — Ollama as a custom serving template on SAP AI Core; resource plan infer.s; SAP AI Core instance with Standard or Extended plan; model pulled into the Ollama pod through the SAP AI API; dated 13 June 2025

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in