Understand how quantization shrinks a model to fit a laptop or a small GPU, measure what it costs in accuracy on an SAP-shaped task, and pick a level with evidence.
A language model is a very long list of numbers, called weights. A model with 8 billion weights stored at full precision takes about 16 GB. That is too big for most laptops and for the smallest cloud GPUs.
Quantization stores each weight with fewer bits, the way you might round prices to whole euros. The file gets two to four times smaller and usually faster to run. The model also gets a little less precise.
The skill is knowing how much rounding a task can take. Eight-bit models are usually almost indistinguishable from the original. Four-bit models are the common default for running models locally. Below four bits, quality tends to fall off.
You don't guess. You run the same test on two or three versions and look at the numbers.
Quantization decides where a model can run and what it costs.
Hardware cost. In llama.cpp's published table, an 8-billion-parameter model shrinks from 32.1 GB in its original form to 4.9 GB at a common four-bit level. That is the difference between a large GPU server and a single modest GPU or a good laptop.
Data stays put. A model small enough to run inside your own boundary means prompts with customer or supplier data never leave it. That is often why teams look at local models in the first place.
Speed. Smaller weights mean less data to move per word generated. The same table shows text generation more than twice as fast at four bits as at sixteen.
Risk. Too much rounding produces a model that looks fine in a demo and makes more mistakes on real cases. Those mistakes are often quiet: a wrong route, a misread code.
Take the running example from order-to-cash. Clerks write short notes on blocked sales orders: "customer over the credit limit", "no price for material TG11", "customer asked us to hold delivery". A small local model routes each note to credit management, master data, pricing, logistics or customer service.
A four-bit version runs on the clerks' standard laptops. An eight-bit version needs more memory but may route a few more notes correctly. The business question is not "which is better?" but "is the accuracy difference worth the hardware?" Only a measurement on your own notes answers that.
SAP offers two routes, and quantization matters differently in each. Both statements are as of October 2026.
Generative AI hub in SAP AI Core. You call models through an API, and someone else runs them. SAP's Python SDK documentation lists open-weight models such as meta--llama3.1-70b-instruct and mistralai--mistral-small-instruct next to commercial ones. You don't choose a quantization level, and the pages we opened don't say which precision is served. You judge the model by testing its answers.
Your own model on SAP AI Core. SAP's developer tutorial shows how to run Ollama, the same local model runner this unit uses, as a custom serving container. It uses the infer.s resource plan and needs an SAP AI Core instance on the Standard or Extended plan. Here you pick the model file, so you pick the quantization level, exactly as on a laptop.
There is no SAP-specific quantization tool in the sources we checked. The quantization itself happens in open-source tools such as llama.cpp before the model reaches SAP AI Core.
The bit counts and the effects below come from llama.cpp's quantize documentation and Hugging Face's guide to quantization methods. Sizes are for an 8-billion-parameter model.
Level
Bits per weight
Size (8B model)
What to expect
Typical use
16-bit (F16, BF16)
16
about 15 GiB
The reference quality
Evaluation baseline, fine-tuning, large GPUs
8-bit (Q8_0)
8.5
about 8 GiB
Very close to 16-bit
When memory allows and accuracy matters most
5- to 6-bit (Q5_K_M, Q6_K)
5.7 to 6.6
5.3 to 6.1 GiB
A middle ground
A small step down from 8-bit
4-bit (Q4_K_M)
4.9
about 4.6 GiB
Relatively high accuracy, with a measurable drop
The usual default for local models
2- to 3-bit (Q2_K, Q3_K_M)
3.2 to 4.0
3 to 3.7 GiB
A noticeable drop, especially at 2-bit
Only when nothing else fits, and only after testing
Two rules of thumb follow:
Start at 4-bit, test 8-bit next to it. If 8-bit wins clearly on your task, pay for the memory.
Compare sizes, not just levels. A bigger model at 4 bits can beat a smaller model at 16 bits that takes the same memory. Test both.
"Quantized means low quality." Eight-bit models are usually very close to the original. The drop at four bits is often small for narrow tasks. It is a measured trade-off, not a downgrade by definition.
"Smaller is always faster." Generation speed rises as the file shrinks, but reading a long prompt stays about the same in llama.cpp's table. Memory, hardware and context length also matter.
"The file size is the memory we need." The runner also keeps the conversation in memory. Long prompts can add a lot on top of the weights.
"We'll quantize when we import the model." Ollama's documentation says it does not quantize GGUF models during import. You choose a quantized file or make one with llama.cpp first.
"A public benchmark tells us which level to use." Hugging Face's own guide says to benchmark on your task and hardware. A level that is fine for chat can still misroute an SAP document type.
Pick one answer for each question. The explanation appears after you choose.
1What does quantization change about a model?
Answer: B. Quantization keeps the same weights but stores them more coarsely. The file gets two to four times smaller. Retraining on your data is fine-tuning, a different technique.
2A team wants to run a model on clerks' laptops. Which starting point does this topic suggest?
Answer: C. Four-bit is the usual default for local models, and 8-bit is very close to the original. Testing both on your own examples shows whether the extra memory pays off. Public benchmarks don't measure your routing task.
3Why is the model file size not the full memory requirement?
Answer: A. The weights are only part of it. The conversation, or context, also takes memory, and long prompts can add a lot. Plan for the file plus room for the longest prompts you expect.
4When you call an open-weight model through SAP's generative AI hub, who picks the quantization level?
Answer: C. In the hub, someone else runs the model and the pages we opened don't state its precision. You can't change it, so you test the answers. If you bring your own model to SAP AI Core, you pick the file and its level.
5Which question best tests a vendor's claim that their quantized model is "just as good"?
Answer: B. "Just as good" is only meaningful against a baseline on data like yours. Format and popularity don't tell you how many notes it will misroute. Ask for the measured difference, and how to reproduce it.
6A bigger model at 4-bit takes the same memory as a smaller model at 16-bit. What should you do?
Answer: D. Neither rule holds every time. A bigger model with more rounding can beat a smaller one stored precisely, or not. The same test on the same examples decides.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Quantization stores each weight as a small integer plus a scale. To read the weight back, you multiply: weight ≈ integer × scale.
Think of measuring parts with a ruler that has only 15 marks (that is 4 bits: the integers −7 to +7). If one ruler must cover every part from a screw to a girder, the marks are far apart and small parts all read as zero. If each small batch of parts gets its own ruler, sized to the largest part in that batch, the marks are close together and the readings are accurate.
That is most of the field:
Fewer bits means fewer marks on the ruler, so more rounding error.
Smaller blocks means more rulers, so less error but more scales to store.
Outliers (a few very large weights) stretch a ruler and ruin precision for everything that shares it.
Smarter methods decide which weights matter most and protect them.
Models are usually trained and published in 16-bit floating point. The Hugging Face GGUF documentation describes the two common 16-bit types: F16 is IEEE half precision, and BF16 is a shortened version of the 32-bit float that keeps its range.
To quantize a block of weights to b bits, the simplest method works like this:
Find the largest absolute value in the block.
Set the scale so that value maps to the top integer, for example 127 for 8 bits or 7 for 4 bits.
Divide every weight by the scale and round to the nearest integer.
Store the integers and the scale.
This is round-to-nearest quantization. The GGUF documentation describes Q8_0 and Q4_0 exactly this way, with blocks of 32 weights.
Each block also stores its scale. With a 16-bit scale for every 32 weights, Q8_0 costs 8 + 16 / 32 = 8.5 bits per weight. llama.cpp's table lists Q8_0 at 8.5008 bits per weight.
The newer K-quants group blocks into super-blocks. The GGUF documentation gives Q4_K at 4.5 bits per weight, Q5_K at 5.5, Q6_K at 6.5625 and Q2_K at 2.625. llama.cpp's table lists the Q4_K_M preset at 4.89, higher than Q4_K alone, so that preset doesn't store every tensor at Q4_K.
The AWQ paper reports that protecting only 1% of the most important weights greatly reduces quantization error. It finds those weights by looking at the activations (the values flowing through the model on real text), not at the size of the weights.
That idea leads to two families of methods. Hugging Face's guide splits them this way:
Family
Examples
How
Trade-off
On the fly, no calibration
bitsandbytes, HQQ, torchao
Quantize while loading
Easy; bitsandbytes is aimed mainly at NVIDIA GPUs
Calibration-based
GPTQ, AWQ
Run sample text first, then quantize
Often the best 4-bit accuracy; GPTQ can overfit its calibration data
GGUF with an importance matrix
llama.cpp llama-imatrix
Collect importance from calibration text
Used for low-bit GGUF files
Hugging Face's guide reports calibration times of about 20 minutes for GPTQ and about 10 minutes for AWQ, for an 8-billion-parameter model on one A100 GPU.
Ollama and llama.cpp use GGUF files. The Qwen documentation shows the usual pipeline, and Ollama's import page says Ollama doesn't quantize GGUF models during import, so quantizing happens before Ollama sees the file.
flowchart LR
H[Published model<br/>BF16 weights] --> C[convert to GGUF<br/>BF16]
C --> Q[llama-quantize<br/>Q8_0, Q4_K_M]
T[Calibration text<br/>optional] --> I[llama-imatrix]
I -.-> Q
Q --> M[Modelfile<br/>FROM file.gguf]
M --> O[ollama create<br/>then run]
The Qwen documentation also suggests calibration text that represents your target domain. For SAP work that could be a sample of your own notes or document texts, cleared for that use.
llama.cpp's README says models are fully loaded into memory, so you need RAM (or GPU memory) for the whole file. On top of that, the runner keeps the KV cache: the model's working memory for the current conversation. It grows with the context length.
Ollama uses a 4,096-token context by default, according to its FAQ. The same FAQ describes KV cache quantization through the OLLAMA_KV_CACHE_TYPE setting:
Value
Memory compared with f16
Effect described by Ollama
f16 (default)
1x
High precision
q8_0
about 1/2
Very small loss; usually no noticeable impact
q4_0
about 1/4
Small to medium loss, more noticeable at long contexts
It only works when Flash Attention is on, which Ollama enables automatically when the hardware supports it. It is a global setting: every model on that Ollama server uses it.
llama.cpp's table for Llama 3.1 8B shows two speeds:
Level
Size (GiB)
Reading the prompt (tokens/s)
Writing the answer (tokens/s)
F16
14.96
923
29
Q8_0
7.95
865
51
Q4_K_M
4.58
822
72
Q2_K
2.95
784
80
Writing the answer gets much faster as the file shrinks, because each new token has to read every weight. Reading the prompt stays roughly flat. For SAP tasks with long inputs and short outputs, like routing a note, the gain from quantization is smaller than the answer-speed column suggests.
These numbers come from one machine in the README. Yours will differ, which is why you measure.
#Build it yourself: measure the trade-off on your machine
You will do two things. First, quantize a made-up layer of weights yourself, so you can see rounding error and the block trick with your own eyes. Then run the same blocked-order routing test on three versions of one small model, at 4, 8 and 16 bits, and compare accuracy, speed and size.
This script needs no model and no internet. It makes a grid of about a million random weights shaped like a trained layer, adds 20 unusually large ones, and stores the grid five ways. Then it compares what the layer computes with the original and the rounded weights.
In VS Code's file list, right-click unit12, choose New File, name it quantize_demo.py, paste the code below and save.
"""See what quantization does to numbers, then estimate how big a model file gets.
Run it from your course folder:
python unit12/quantize_demo.py # quantize one made-up weight matrix
python unit12/quantize_demo.py --params 8 # also size an 8-billion-parameter model
python unit12/quantize_demo.py --no-outliers # same test without a few large weights
Part 1 builds a made-up "layer" of weights and stores it four ways: 16-bit floats,
8-bit and 4-bit integers with ONE scale for the whole matrix, and 8-bit and 4-bit
integers with one scale per block of 32 weights (the idea behind GGUF's Q8_0 and Q4_0).
Part 2 turns bits per weight into file sizes. No model, no account, no internet.
"""
import argparse
import numpy as np
ROWS, COLS = 1024, 1024 # one made-up layer: about a million weights
BLOCK = 32 # weights that share one scale in the block methods
# Average bits per weight from llama.cpp's quantize README (Llama 3.1 8B table).
# They include the scales, which is why Q8_0 is 8.5 and not 8.
LLAMA_CPP_BPW = {"F16": 16.0, "Q8_0": 8.5, "Q6_K": 6.56, "Q5_K_M": 5.70, "Q4_K_M": 4.89, "Q3_K_M": 4.00}
def make_weights(rng, outliers: bool) -> np.ndarray:
"""Small random weights, like a trained layer; optionally a few much larger ones."""
w = rng.normal(0.0, 0.02, size=(ROWS, COLS)).astype(np.float32)
if outliers:
spots = rng.choice(w.size, size=20, replace=False) # 20 of about a million
w.flat[spots] = rng.choice([-1.0, 1.0], size=20) * 0.5
return w
def quantize(w: np.ndarray, bits: int, block: int | None) -> tuple:
"""Round weights to signed integers with `bits` bits; return the rebuilt weights and bits per weight.
block=None: one scale for the whole matrix. block=32: one scale per 32 weights.
Each scale is stored as a 16-bit float, so it adds 16 bits per block.
"""
levels = 2 ** (bits - 1) - 1 # 127 for 8-bit, 7 for 4-bit
flat = w.reshape(-1, block) if block else w.reshape(1, -1)
scale = np.abs(flat).max(axis=1, keepdims=True) / levels # biggest value maps to the top level
scale[scale == 0] = 1.0
q = np.clip(np.round(flat / scale), -levels, levels) # the small integers that get stored
rebuilt = (q * scale.astype(np.float16).astype(np.float32)).reshape(w.shape)
scale_bits = 16 * flat.shape[0]
return rebuilt, bits + scale_bits / w.size
def relative_error(original: np.ndarray, approx: np.ndarray) -> float:
"""Size of the error compared with the size of the original, in percent."""
return 100 * float(np.linalg.norm(original - approx) / np.linalg.norm(original))
def part_one(seed: int, outliers: bool) -> None:
rng = np.random.default_rng(seed)
w = make_weights(rng, outliers)
x = rng.normal(0.0, 1.0, size=(COLS, 64)).astype(np.float32) # 64 made-up inputs
y = w @ x # what the layer outputs
methods = [
("16-bit float", lambda: (w.astype(np.float16).astype(np.float32), 16.0)),
("8-bit, one scale", lambda: quantize(w, 8, None)),
("8-bit, blocks of 32", lambda: quantize(w, 8, BLOCK)),
("4-bit, one scale", lambda: quantize(w, 4, None)),
("4-bit, blocks of 32", lambda: quantize(w, 4, BLOCK)),
]
label = "with 20 large outlier weights" if outliers else "without outliers"
print(f"Part 1: one made-up layer, {ROWS} x {COLS} weights, {label}\n")
print(f"{'Method':<22}{'Bits/weight':>12}{'Size':>10}{'Weight error':>14}{'Output error':>14}")
for name, run in methods:
rebuilt, bpw = run()
size_mb = w.size * bpw / 8 / 1e6
print(f"{name:<22}{bpw:>12.2f}{size_mb:>8.2f} MB"
f"{relative_error(w, rebuilt):>13.2f}%{relative_error(y, rebuilt @ x):>13.2f}%")
print("\nOutput error compares what the layer computes with the original weights and with the")
print("rebuilt ones. Smaller is better. 32-bit floats would take 4.19 MB.")
def part_two(params_billion: list) -> None:
print("\nPart 2: weight file size = parameters x bits per weight / 8")
print("(a floor: real files add metadata, and the runner needs memory for the conversation)\n")
header = f"{'Parameters':<12}" + "".join(f"{name:>10}" for name in LLAMA_CPP_BPW)
print(header)
for p in params_billion:
row = f"{p:>6.2f} B "
for bpw in LLAMA_CPP_BPW.values():
row += f"{p * 1e9 * bpw / 8 / 1e9:>8.2f}GB"
print(row)
def main() -> None:
parser = argparse.ArgumentParser(description="See what quantization does to weights and file size.")
parser.add_argument("--params", type=float, action="append",
help="model size in billions of parameters (repeat for several)")
parser.add_argument("--no-outliers", action="store_true", help="leave out the large weights")
parser.add_argument("--seed", type=int, default=7, help="change for different random weights")
args = parser.parse_args()
part_one(args.seed, outliers=not args.no_outliers)
part_two(args.params or [0.752, 8.0, 70.0]) # 0.752 B is qwen3:0.6b's parameter count
if __name__ == "__main__":
main()
Run it (the same on every system):
python unit12/quantize_demo.py
What success looks like (the numbers are the same on every computer, because the random seed is fixed):
Part 1: one made-up layer, 1024 x 1024 weights, with 20 large outlier weights
Method Bits/weight Size Weight error Output error
16-bit float 16.00 2.10 MB 0.02% 0.02%
8-bit, one scale 8.00 1.05 MB 5.65% 5.64%
8-bit, blocks of 32 8.50 1.11 MB 0.55% 0.55%
4-bit, one scale 4.00 0.52 MB 88.26% 88.38%
4-bit, blocks of 32 4.50 0.59 MB 9.88% 9.84%
Output error compares what the layer computes with the original weights and with the
rebuilt ones. Smaller is better. 32-bit floats would take 4.19 MB.
Part 2: weight file size = parameters x bits per weight / 8
(a floor: real files add metadata, and the runner needs memory for the conversation)
Parameters F16 Q8_0 Q6_K Q5_K_M Q4_K_M Q3_K_M
0.75 B 1.50GB 0.80GB 0.62GB 0.54GB 0.46GB 0.38GB
8.00 B 16.00GB 8.50GB 6.56GB 5.70GB 4.89GB 4.00GB
70.00 B 140.00GB 74.38GB 57.40GB 49.88GB 42.79GB 35.00GB
Read the table. Three lessons are in it:
Outliers ruin a shared scale. With one scale for the whole grid, 20 large weights out of a million stretch the ruler. At 4 bits almost every normal weight rounds to zero, and the output error is 88%.
Blocks fix most of it. One scale per 32 weights costs 0.5 extra bits per weight and brings the 4-bit error down to about 10%. This is why GGUF formats use blocks.
8-bit with blocks is almost free. Half the size of 16-bit, with an error near half a percent.
Run it once without the outliers and compare the "one scale" rows:
python unit12/quantize_demo.py --no-outliers
The 4-bit one-scale error drops to about 20%, while the block rows barely change. Blocks matter most when weights are uneven, which is the situation AWQ describes in real models.
Check Part 2 against reality. For the 8B model, the Q4_K_M estimate of 4.89 GB matches the 4.9 GB that llama.cpp reports. For qwen3:0.6b (752 million parameters), Ollama's tag list shows 1.5 GB at fp16 and 832 MB at q8_0, close to the estimates. The q4_K_M file is 523 MB, more than the 0.46 GB estimate, because the 4.89 average comes from a much larger model. Treat Part 2 as a first estimate, then check the real file.
The script sends 20 made-up clerk notes to each version, one at a time, with the same instructions and settings. Each answer must be JSON naming one of five routes. It counts correct routes and valid JSON, and records Ollama's own timings.
In VS Code, right-click unit12, choose New File, name it compare_quants.py, paste the code below and save.
"""Compare the same small model at three quantization levels on one SAP-shaped task.
Run it from your course folder:
python unit12/compare_quants.py --sample # no model: shows the output format
python unit12/compare_quants.py # q4_K_M, q8_0 and fp16 of qwen3:0.6b
python unit12/compare_quants.py --limit 5 # quick try on the first 5 notes
python unit12/compare_quants.py --models qwen3:1.7b-q4_K_M qwen3:1.7b-q8_0
The task: route a clerk's note about a blocked sales order to one of five teams.
The notes and routes are made up; the routes are this course's labels, not SAP codes.
Everything runs in Ollama on this computer. Results are saved to unit12/quant_results.csv.
"""
import argparse
import csv
import json
import re
import sys
import time
from pathlib import Path
DEFAULT_MODELS = ["qwen3:0.6b-q4_K_M", "qwen3:0.6b-q8_0", "qwen3:0.6b-fp16"]
ROUTES = ["CREDIT", "MASTER_DATA", "PRICING", "STOCK", "CUSTOMER_HOLD"]
RESULTS = Path("unit12") / "quant_results.csv"
SYSTEM = ("Route the clerk's note about a blocked SAP sales order. "
'Reply with JSON only: {"route": "<ROUTE>"}. ROUTE is one of '
+ ", ".join(ROUTES) + ". "
"CREDIT: credit limit or unpaid invoices. MASTER_DATA: missing or wrong customer data. "
"PRICING: wrong or missing price or discount. STOCK: not enough material. "
"CUSTOMER_HOLD: the customer asked to wait.")
# 20 made-up notes, four per route, with the route a person would choose.
NOTES = [
("Order 4711 stuck, customer 10100001 is 6,200 EUR over the credit limit.", "CREDIT"),
("Credit check failed on 4712. Invoices overdue since July.", "CREDIT"),
("Finance says the customer still hasn't paid the August statement, so 4713 is held.", "CREDIT"),
("Risk team flagged 10100007; payment behaviour got worse and 4714 waits on them.", "CREDIT"),
("4721 blocked: ship-to address for 10100002 has no postal code.", "MASTER_DATA"),
("Tax number missing on customer 10100003, so 4722 can't be released.", "MASTER_DATA"),
("Incoterms not maintained for 10100004; 4723 stays blocked.", "MASTER_DATA"),
("Customer record was created yesterday and half the fields are empty, 4724 on hold.", "MASTER_DATA"),
("No price found for material TG11 on 4731.", "PRICING"),
("4732: manual discount of 18% is above tolerance, needs approval.", "PRICING"),
("Condition record for TG12 expired, 4733 shows zero price.", "PRICING"),
("Line for TG14 on 4734 came through at 0.00 EUR, someone has to fix the amount.", "PRICING"),
("Not enough stock of TG11 for 4741, short by 40 pieces.", "STOCK"),
("4742: availability check confirms only 12 of the ordered quantity.", "STOCK"),
("Material TG13 back-ordered, 4743 waits for the next receipt.", "STOCK"),
("Warehouse is out of TG15 until the supplier delivers, 4744 sits there.", "STOCK"),
("10100005 asked us to hold 4751 until further notice.", "CUSTOMER_HOLD"),
("Customer called: do not ship 4752 before October.", "CUSTOMER_HOLD"),
("4753 on hold at the customer's request, their warehouse is full.", "CUSTOMER_HOLD"),
("Email from 10100006: please park 4754, they are moving sites.", "CUSTOMER_HOLD"),
]
# Reorder so the routes take turns: then --limit 5 still tests every route once.
NOTES = [NOTES[i + 4 * r] for i in range(4) for r in range(len(ROUTES))]
# What --sample prints. Made-up numbers in a realistic shape, not a measurement.
SAMPLE_ROWS = [
{"model": "qwen3:0.6b-q4_K_M", "size_mb": 523, "accuracy": 0.70, "valid_json": 0.95,
"tokens_per_s": 38.2, "load_s": 1.4, "seconds_per_note": 0.6},
{"model": "qwen3:0.6b-q8_0", "size_mb": 832, "accuracy": 0.75, "valid_json": 1.00,
"tokens_per_s": 29.5, "load_s": 1.9, "seconds_per_note": 0.8},
{"model": "qwen3:0.6b-fp16", "size_mb": 1500, "accuracy": 0.75, "valid_json": 1.00,
"tokens_per_s": 17.1, "load_s": 3.2, "seconds_per_note": 1.3},
]
def read_route(text: str) -> tuple:
"""Return (route or None, whether the reply was valid JSON)."""
try:
route = json.loads(text).get("route", "")
return (route if route in ROUTES else None), True
except (json.JSONDecodeError, AttributeError):
# Not clean JSON: still accept a route name if one appears, but count it as invalid JSON.
found = re.search("|".join(ROUTES), text.upper())
return (found.group(0) if found else None), False
def model_sizes(client) -> dict:
"""Size on disk of every pulled model, in MB, as Ollama reports it."""
return {m.model: round(m.size / 1e6) for m in client.list().models}
def run_model(client, model: str, notes: list) -> dict:
"""Ask one model to route every note; return accuracy, speed and load time."""
correct = valid = answer_tokens = 0
answer_ns = load_ns = 0
started = time.perf_counter()
for i, (note, expected) in enumerate(notes, start=1):
response = client.chat(
model=model,
messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": note}],
think=False, # answer straight away
options={"temperature": 0, "num_ctx": 2048}, # same settings for every model
)
route, is_json = read_route(response.message.content.strip())
correct += route == expected
valid += is_json
answer_tokens += response.eval_count or 0
answer_ns += response.eval_duration or 0
load_ns += response.load_duration or 0
print(f" {i:>2}/{len(notes)} expected {expected:<13} got {route or '(unreadable)'}")
elapsed = time.perf_counter() - started
return {
"model": model,
"accuracy": round(correct / len(notes), 2),
"valid_json": round(valid / len(notes), 2),
"tokens_per_s": round(answer_tokens / (answer_ns / 1e9), 1) if answer_ns else 0.0,
"load_s": round(load_ns / 1e9, 1),
"seconds_per_note": round(elapsed / len(notes), 2),
}
def print_table(rows: list) -> None:
print(f"\n{'Model':<22}{'Size MB':>9}{'Accuracy':>10}{'Valid JSON':>12}"
f"{'Tokens/s':>10}{'Load s':>8}{'s/note':>8}")
for r in rows:
print(f"{r['model']:<22}{r['size_mb']:>9}{r['accuracy']:>10.0%}{r['valid_json']:>12.0%}"
f"{r['tokens_per_s']:>10.1f}{r['load_s']:>8.1f}{r['seconds_per_note']:>8.2f}")
def save(rows: list) -> None:
RESULTS.parent.mkdir(exist_ok=True)
with RESULTS.open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=list(rows[0]))
writer.writeheader()
writer.writerows(rows)
print(f"\nSaved to {RESULTS}")
def main() -> None:
parser = argparse.ArgumentParser(description="Compare quantization levels of a local model.")
parser.add_argument("--models", nargs="+", default=DEFAULT_MODELS, help="Ollama tags to compare")
parser.add_argument("--limit", type=int, default=len(NOTES), help="use only the first N notes")
parser.add_argument("--sample", action="store_true", help="no model: print sample results")
args = parser.parse_args()
if args.sample:
print("Sample results (made up, no model was called):")
print_table(SAMPLE_ROWS)
return
try:
from ollama import Client, ResponseError
except ModuleNotFoundError:
sys.exit("The ollama library isn't installed. Run: pip install -r requirements.txt")
client = Client() # http://127.0.0.1:11434 unless OLLAMA_HOST says otherwise
notes = NOTES[: max(1, args.limit)]
try:
sizes = model_sizes(client)
missing = [m for m in args.models if m not in sizes]
if missing:
sys.exit("Pull these first, one at a time:\n" + "\n".join(f" ollama pull {m}" for m in missing))
rows = []
for model in args.models:
print(f"\n{model}")
row = run_model(client, model, notes)
row["size_mb"] = sizes[model]
rows.append(row)
except ConnectionError:
sys.exit("Can't reach Ollama on this computer. Start the Ollama app "
"(Linux: sudo systemctl start ollama), then run this again.")
except ResponseError as error:
sys.exit(f"Ollama answered with an error: {error}")
print_table(rows)
save(rows)
if __name__ == "__main__":
main()
Run it with --sample first. This prints the output format without calling any model:
python unit12/compare_quants.py --sample
What success looks like:
Sample results (made up, no model was called):
Model Size MB Accuracy Valid JSON Tokens/s Load s s/note
qwen3:0.6b-q4_K_M 523 70% 95% 38.2 1.4 0.60
qwen3:0.6b-q8_0 832 75% 100% 29.5 1.9 0.80
qwen3:0.6b-fp16 1500 75% 100% 17.1 3.2 1.30
These numbers are invented to show the layout. Don't quote them.
Try a quick real run on five notes, one per route:
python unit12/compare_quants.py --limit 5
You should see each model's name, then one line per note:
Your routes may differ. got (unreadable) means the answer named no route at all; that counts as wrong.
Run the full test:
python unit12/compare_quants.py
What success looks like: 20 lines per model, the comparison table with your own numbers, and Saved to unit12/quant_results.csv. On a laptop without a GPU this takes a few minutes.
Open unit12/quant_results.csv in VS Code, or look at the table in the terminal. Ask four questions:
Accuracy. Is the 4-bit version clearly worse than fp16? With 20 notes, one note is 5 percentage points, so a gap of one or two notes is noise. A larger, repeatable gap is a signal.
Valid JSON. Lower precision can break formatting before it breaks understanding. If valid JSON drops, the business process breaks even when the route is right.
Tokens per second. Does the 4-bit version write faster on your machine? Compare with the llama.cpp pattern: big gains in answer speed, small gains in reading the prompt.
Seconds per note. This is what a clerk feels. Routing answers are short, so this is mostly prompt reading and overhead.
A good conclusion names a choice and a reason, for example: "q4_K_M: same accuracy as fp16 on 20 notes, a third of the size, twice the speed. Re-test with 100 real notes before go-live."
As of October 2026, there are two ways to bring this into SAP's platform. Neither has an SAP-specific quantization feature in the sources we opened; quantization happens in open-source tools before the model reaches SAP.
SAP's Python SDK documentation for the generative AI hub lists open-weight models such as meta--llama3.1-70b-instruct, ibm--granite-13b-chat, mistralai--mistral-small-instruct and mistralai--mistral-large-instruct, next to models from Amazon, Anthropic, Google and OpenAI.
SAP's developer tutorial "Using Custom models on SAP AI Core via Ollama" packages Ollama as a custom serving container. The serving template sets the resource plan with the label ai.sap.com/resourcePlan: infer.s, and the model is pulled into the running Ollama pod through the SAP AI API. The tutorial requires an SAP AI Core instance on the Standard or Extended plan, and it was last dated 13 June 2025.
Because the container runs Ollama, the quantization choice is the same tag choice you made in Step 3. The part of the serving template that matters here looks like this:
# Sketch only: needs an SAP AI Core instance (Standard or Extended plan) and a
# Docker registry; Set up for Unit 5 covers SAP AI Core. Based on SAP's Ollama tutorial.
metadata:
labels:
ai.sap.com/resourcePlan: infer.s # the GPU plan the tutorial uses
# ...the rest of the serving template from SAP's tutorial...
# After deployment, pull the exact tag you tested on your laptop, for example
# qwen3:0.6b-q8_0, not a bare name that could resolve to a different level.
If you need a level that the Ollama library doesn't publish, make a GGUF file with llama.cpp (convert, then llama-quantize, optionally with an importance matrix from your own calibration text), import it with a Modelfile, and serve that. Ollama's import documentation is explicit that it won't quantize during import.
Evaluate per level, on your data. Hugging Face's guide ends with the advice to benchmark accuracy and speed on your own task and hardware. Keep the full-precision version as the baseline, and rerun the same evaluation set for every level and every model update.
Pin the exact file. Use the full tag (qwen3:0.6b-q8_0), not the bare name, and record the model ID from the registry. A bare name like qwen3 points at the library's default, the 8b model, which you didn't choose.
Supply chain. A model file is executable behaviour you download from the internet. Pull from a source your security team approves, record where each file came from, and check the licence (Apache 2.0 for Qwen3, per its Ollama page). Treat files you quantize yourself the same way: keep the command, the source file and the calibration text.
Calibration data is data. If you build an importance matrix from real SAP texts, those texts need the same approval as any other use of business data.
Memory planning. Size for the weights plus the KV cache at your longest real prompt. On a shared Ollama server, OLLAMA_KV_CACHE_TYPE applies to every model, so test all of them after changing it.
Authorizations. Quantization doesn't change who may see what. A local model that reads SAP data still needs the same SAP authorization checks in the code that fetches the data, as covered in agent permissions and SAP authorizations.
Clean core. Nothing here changes the SAP system. The model runs beside it and reads data through released APIs.
Cost. On SAP AI Core you pay for the resource plan while the deployment runs. A smaller quantized model may fit a smaller plan; confirm plan sizes and prices with SAP before committing.
Comparing different models and calling it a quantization test. Change one thing at a time: same model, same prompt, same settings, different level.
Too few test cases. With 20 notes, one note is 5 points. Don't decide on a one-note difference.
Ignoring format errors. A model that routes correctly but breaks the JSON still breaks the process. Track valid output separately.
Trusting the bare model name.qwen3:0.6b is the q4_K_M file today, per the tag list. Write the level into your config explicitly.
Expecting Ollama to quantize on import. It doesn't for GGUF files. Quantize with llama.cpp first.
Going below 4 bits without calibration. Hugging Face's guide reports a noticeable drop at 2-bit, and the Qwen documentation notes that llama-quantize warns when 1- or 2-bit files are made without an importance matrix.
Sizing hardware from the file alone. The KV cache grows with context. Test with your longest realistic prompt.
Answer the question this unit keeps coming back to: is a bigger model at lower precision better than a smaller one at higher precision? The result goes into your Unit 12 decision log and feeds the buy-vs-build topic at the end of the unit.
Pull the 4-bit version of the next size up. Ollama's tag list gives it as 1.4 GB, close to the 1.5 GB of qwen3:0.6b-fp16:
Open unit12/model_log.md (create it if it doesn't exist) and add a section ## Quantization decision with a table: model tag, size MB, accuracy, valid JSON, tokens per second, seconds per note. Copy the rows from this run and from Step 4.
Under the table, write two sentences: which tag would you deploy for the clerks' laptops, and what result would make you change your mind?
Save the results file under a new name so the next run doesn't overwrite it, then commit both files.
Done whenmodel_log.md has a ## Quantization decision table with at least four rows of your own measured numbers and a two-sentence choice, and both files are committed.
Pick one answer for each question. The explanation appears after you choose.
1In quantize_demo.py, why does "4-bit, one scale" lose almost all precision when 20 outliers are present?
Answer: B. The scale is set by the largest value. With one scale for a million weights, a few large ones make each step so wide that almost all normal weights round to zero. Blocks of 32 give each group its own scale, which is why the block row drops to about 10% error.
2Why does Q8_0 cost 8.5 bits per weight rather than 8?
Answer: C. The integers take 8 bits each, and each block of 32 also stores its scale. Sixteen bits spread over 32 weights adds 0.5 bits per weight, matching llama.cpp's 8.5008.
3Your team needs a Q3_K_M version of a model the Ollama library doesn't publish at that level. What do you do?
Answer: D. Ollama's import documentation says it doesn't quantize GGUF models during import, so the file must be quantized first. The KV cache setting only affects conversation memory, not weights. The hub doesn't let you pick a level.
4What does AWQ use to decide which weights to protect?
Answer: B. The AWQ paper finds important weights from the activations, not from weight size, and it needs no backpropagation. Protecting about 1% of weights greatly reduces the error.
5In llama.cpp's table, Q4_K_M writes answers more than twice as fast as F16. Why does routing a note speed up much less?
Answer: C. In the table, generation speed rises steeply as the file shrinks, but prompt processing stays roughly flat. A routing call reads a long instruction and writes a few tokens, so most of its time is in the part quantization helps least.
6compare_quants.py reports 80% accuracy for q4_K_M and 85% for fp16 on 20 notes. What do you conclude?
Answer: B. With 20 notes, each note is 5 percentage points, so a one-note gap can come from chance. Run a larger set, ideally real notes cleared for testing, before spending money on memory.
7You deploy Ollama on SAP AI Core and configure the model as qwen3. What's the risk?
Answer: A. A bare name points at whatever the library sets as default, which for qwen3 is the 8b model, a 5.2 GB download, per Ollama's library page. Pin the exact tag you evaluated, such as qwen3:0.6b-q8_0, so production runs what you tested.
8Your security team asks what changes for SAP authorizations when you switch from fp16 to a 4-bit model. What's the right answer?
Answer: D. Quantization changes how weights are stored, not what data the system may read. Authorization checks belong in the code that fetches SAP data, whichever model level answers afterwards.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
quantize README (llama.cpp tools, mirrored on docs.rs)— llama-quantize turns a high-precision GGUF (F32 or BF16) into a quantized one; list of I-quants, K-quants, Q8_0, F16; Llama 3.1 8B table of bits per weight, size and speed (F16 16.0 bpw, 14.96 GiB, 29.17 t/s; Q8_0 8.50, 7.95 GiB, 50.93 t/s; Q4_K_M 4.89, 4.58 GiB, 71.93 t/s; Q2_K 3.16); 8B 32.1 GB to 4.9 GB at Q4_K_M; models fully loaded into memory; --imatrix reduces accuracy loss
GGUF (Hugging Face Hub documentation)— binary format for fast loading, stores tensors plus standardized metadata; Q8_0 and Q4_0 are round-to-nearest with 32-weight blocks (legacy); Q4_K 4.5 bpw, Q5_K 5.5, Q6_K 6.5625, Q2_K 2.625; F16 and BF16 definitions
Selecting a quantization method (Hugging Face Transformers documentation)— on-the-fly methods (bitsandbytes, HQQ, torchao) vs calibration-based (GPTQ, AWQ); 8-bit about 2x memory saving and very close to bf16; 4-bit about 4x with relatively high accuracy; sub-4-bit noticeable drop; bitsandbytes is the standard for QLoRA; always benchmark on your own task and hardware
Quantization with llama.cpp (Qwen documentation)— convert to GGUF in bf16, then llama-quantize to Q8_0, Q5_K_M or Q4_K_M; importance matrix from domain calibration text with llama-imatrix; perplexity comparison with llama-perplexity
qwen3 tags (Ollama model library)— qwen3 latest points to 8b (5.2 GB); qwen3:0.6b and qwen3:0.6b-q4_K_M share ID 7df6b6e09427 (523 MB); qwen3:0.6b-q8_0 832 MB; qwen3:0.6b-fp16 1.5 GB; 1.7b: q4_K_M 1.4 GB, q8_0 2.2 GB, fp16 4.1 GB
FAQ (Ollama documentation)— default context 4096 tokens; Flash Attention used automatically when supported; OLLAMA_KV_CACHE_TYPE f16 (default), q8_0 about half the memory, q4_0 about a quarter with small-medium precision loss; global option for all models
Importing a model (Ollama documentation)— Ollama does not quantize GGUF models during import; quantize first with llama.cpp's llama-quantize; Modelfile FROM /path/to/file.gguf then ollama create