So far in this course, every model has run in someone else's data centre. You sent a prompt over the internet and paid per request. Unit 12 asks a different question: when should the model run somewhere you control?
This setup installs one free program, Ollama, that downloads open models and runs them on your own laptop. You then download a small model, about half a gigabyte, and ask it about a blocked sales order.
Three things change when the model is local:
The data stays put. The prompt never leaves the computer.
There is no bill per request. You pay with hardware and electricity instead.
Quality and speed depend on the machine. A small model on a laptop is fast enough to learn with, and clearly weaker than the large hosted models you used before.
Setup takes 30 to 60 minutes and costs nothing. It needs about 5 GB of free disk and works best with 8 GB of memory or more.
Most SAP AI work will keep using hosted models through SAP's generative AI hub. Local and small models still matter, for three business reasons.
Data that must not leave. Some content, such as HR cases or unreleased financials, may be barred from external model providers. A model you host yourself can keep it inside your own boundary.
Cost at volume. A task that runs millions of times, such as tagging every incoming invoice line, can cost less on a small model you host than on a large hosted one. The unit's buy-vs-build topic weighs this properly.
Fit for purpose. Many tasks don't need the biggest model. A small model, sometimes tuned for one job, can be good enough and much faster.
Take the running example: a clerk asks why sales order 9000001 is blocked. The order data includes a customer number and an amount. If policy says that data may not go to an outside provider, a model running inside the company is one way to still offer the assistant.
The trade-off is real. A small model makes more mistakes, and running models yourself means owning updates, security and capacity. This unit teaches you to measure that trade-off rather than guess.
As of October 2026, SAP offers two routes that touch this unit.
Open models in the generative AI hub. SAP's Python SDK reference lists models from Meta, IBM and Mistral AI alongside Amazon, Anthropic, Google and OpenAI. Examples include meta--llama3.1-70b-instruct and mistralai--mistral-small-instruct. You call them like any other hub model, and SAP runs them.
Your own model on SAP AI Core. An SAP Developer Center tutorial packages Ollama, the same runner you install today, as a custom serving template on SAP AI Core. It needs an SAP AI Core instance on the Standard or Extended plan, a Docker image and a GitHub repository.
Neither route is free, so this unit teaches the ideas on your laptop first. AI features built into SAP applications are a third, separate case, covered by the unit's topic on embedded AI.
Time: 30 to 60 minutes. Most of it is downloading.
Money for a learner: nothing. No account, no key, no per-request cost.
Disk: keep at least 5 GB free. Ollama's Windows documentation asks for at least 4 GB for the program alone, and each model adds its own size.
Memory: 8 GB or more is comfortable. With 4 to 8 GB, stay with the smallest model and close other apps.
A graphics card (GPU) is optional. It makes answers faster. Everything in this unit also runs on the processor.
Money for a company: the software is free. Hosting models for real users means servers, often with GPUs, plus people to run them. That cost is the heart of the buy-vs-build decision later in the unit.
"Local means as good as the big models, just free." A half-gigabyte model is far weaker than a large hosted one. It is useful for learning and for narrow tasks.
"Local means secure." The data stays on the machine, which helps. The machine, the runner and the model download still need the same care as any other software.
"You need a gaming GPU." Small models run on an ordinary laptop processor. A GPU makes them faster, not possible.
"Open model means no licence terms." Each model has its own licence, and some restrict commercial use. Check before production.
"Ollama is an SAP product." It is an independent open-source project. SAP's tutorial uses it on SAP AI Core, which is a different thing from SAP supporting it.
Pick one answer for each question. The explanation appears after you choose.
1What is the main reason a team would run a model locally instead of calling a hosted one?
Answer: B. When the model runs on your own machine, the prompt never leaves it. That can unlock data a policy keeps away from outside providers. Accuracy is usually lower, not higher, for small local models.
2A sponsor says "a local model is free, so let's replace every hosted call with one". What is the soundest reply?
Answer: C. There is no bill per request, but someone pays for servers, updates and capacity. A small model also makes more mistakes, so the swap needs a quality comparison before it is decided.
3Which SAP route lets a team run its own model inside SAP's platform?
Answer: D. SAP's Developer Center tutorial packages Ollama as a custom serving template on SAP AI Core, which needs a Standard or Extended plan. The hub's catalogue is the other route, where SAP runs the model for you.
4A learner's laptop has 6 GB of memory and no graphics card. What should they do?
Answer: A. Small models run on the processor; a GPU only makes them faster. With limited memory, the smallest model is the one that fits comfortably.
5Which question should go to IT before anyone installs Ollama?
Answer: B. Ollama installs software and downloads models from the internet, which some company policies and proxies block. It doesn't need SAP licences or an SAP release.
6A vendor says "it's an open model, so there are no licence terms to check". What is wrong?
Answer: D. The runner and the model are licensed separately. Legal should approve the specific model's licence before production use.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 35 min read
#Mental model: a model is a file, a runner is a player
Think of music. A song is a file. A player opens the file and turns it into sound. You can play the same song on different players, and the same player handles many songs.
Local models work the same way:
The model is a file, often a few hundred megabytes to many gigabytes. It holds the learned parameters.
The runner is the player. Ollama loads the file into memory, feeds your prompt through it, and streams the answer back.
Your code is the remote control. It sends a request to the runner and reads the answer. It never touches the model file directly.
That split explains most of this unit. Fine-tuning changes the file. Quantization makes the file smaller. Buy vs. build asks who runs the player and where.
flowchart LR
C[local_chat.py] -->|HTTP request<br/>127.0.0.1:11434| O[Ollama server]
O -->|loads once| M[(qwen3:0.6b<br/>model file)]
M --> O
O -->|answer and timings| C
R[Ollama registry] -.->|ollama pull<br/>one-time download| M
The Ollama server runs in the background on your computer. On Windows and macOS the Ollama app starts it; on Linux it usually runs as a system service. Ollama's FAQ says it listens on 127.0.0.1 port 11434 by default. 127.0.0.1 means "this computer only", so other machines can't reach it.
ollama pull downloads a model from Ollama's registry into a models folder. After that, no internet is needed to use it.
The ollama Python library sends your chat to the server and returns the answer as an object. It is the official client, MIT licensed.
To answer, the runner loads the whole model into memory: GPU memory if it can, otherwise main memory. It also needs room for the conversation so far, called the context. Ollama's FAQ gives a default context of 4,096 tokens.
The registry shows how size grows with parameters for one model family:
Tag
Parameters
Download
qwen3:0.6b
about 0.6 billion
523 MB
qwen3:1.7b
about 1.7 billion
1.4 GB
qwen3:4b
about 4 billion
2.5 GB
qwen3:8b (also qwen3:latest)
about 8 billion
5.2 GB
Our course rule of thumb: the download size plus 1 to 2 GB must fit in free memory. That is why this setup uses qwen3:0.6b and never the bare name qwen3, which would download the 8b model.
Ollama uses a supported GPU when it finds one and falls back to the processor (CPU) when it doesn't. Its hardware page lists NVIDIA cards with compute capability 5.0 or higher and driver 550 or newer, AMD cards through ROCm v7 or Vulkan, and Apple silicon through Metal. Apple silicon Macs share one pool of memory between CPU and GPU.
ollama ps shows each loaded model and, in its Processor column, how much of it sits on the GPU. "100% GPU" means all of it. By default a model stays loaded for 5 minutes after its last request, so the second question is faster than the first.
You will install Ollama, download a small model, check your machine, and ask the model why a made-up sales order is blocked. No account, no key and no per-request cost.
Before you start: complete Set up your computer for this course and Set up for Unit 2. They install Python, VS Code and Git, and create your orchestrate-course folder with its .venv and requirements.txt. This walkthrough doesn't repeat those steps.
In your browser, go to https://ollama.com/download and choose Windows.
Run the downloaded OllamaSetup.exe and follow the installer. Ollama's documentation says it installs in your home folder and doesn't need Administrator rights.
When it finishes, Ollama runs in the background with an icon in the system tray, near the clock.
macOS
Go to https://ollama.com/download and choose macOS.
Open the downloaded .dmg file and drag Ollama into the Applications folder.
Open Ollama from Applications. On first launch it offers to link the ollama command into /usr/local/bin; accept, and enter your Mac password if asked.
Linux
Run Ollama's install script. It asks for your password because it installs system-wide:
curl -fsSL https://ollama.com/install.sh | sh
Check that the service is running:
systemctl status ollama
Look for active (running). Press q to leave the status view. If it isn't running, start it with sudo systemctl start ollama.
All systems: check the install
Close VS Code's terminal and open a new one (Terminal > New Terminal), so it sees the new command. Turn .venv back on as in Step 1.
Run:
ollama -v
What success looks like: one line with the version number, such as ollama version is 0.x.y. Any version is fine.
Pull the model. Always type the tag :0.6b; the bare name qwen3 downloads the 5.2 GB model.
ollama pull qwen3:0.6b
A progress bar runs for a minute or two. The last line says success.
List your models:
ollama list
You should see a line starting with qwen3:0.6b, with its size, about 0.5 GB.
Chat with it in the terminal:
ollama run qwen3:0.6b "In one sentence, what is a sales order?"
The model writes an answer. It may first show its reasoning ("thinking"), because Qwen3 models can reason before they answer. The first run takes longer while the model loads.
See where it is running:
ollama ps
The Processor column says 100% CPU, 100% GPU, or a split. Any of these is fine for this unit.
You will write a short script that sends a made-up blocked order to the model and prints the answer and its speed. The order uses field names from SAP's sales order API, as in earlier units.
In VS Code's file list, right-click unit12, choose New File, and name it local_chat.py.
Paste the code below and save.
"""Ask a small language model on your own computer about a blocked SAP sales order.
Run it from your course folder:
python unit12/local_chat.py --sample # no model: shows the output format
python unit12/local_chat.py # asks qwen3:0.6b through Ollama
python unit12/local_chat.py --model qwen3:1.7b # a bigger model you pulled
python unit12/local_chat.py --question "What should the clerk check first?"
The model runs in Ollama on this computer. The order and the prompt never leave it.
The order is made up, in the shape of SAP's sales order API. The meaning of a block code
is set in each SAP system's configuration; the meanings below belong to this sample only.
"""
import argparse
import json
import sys
# One made-up blocked sales order, with field names from SAP's sales order API.
ORDER = {
"SalesOrder": "9000001",
"SoldToParty": "10100001",
"TotalNetAmount": "18250.00",
"TransactionCurrency": "EUR",
"TotalCreditCheckStatus": "B",
"DeliveryBlockReason": "01",
"HeaderBillingBlockReason": "",
}
# What the codes mean in this made-up system. Your system's configuration decides the real ones.
CODE_MEANINGS = {
"TotalCreditCheckStatus B": "credit check failed",
"DeliveryBlockReason 01": "credit hold",
"empty value": "no block of that kind",
}
SYSTEM = (
"You help order-to-cash clerks understand blocked SAP sales orders. "
"Use only the order data and code meanings given. If the data can't answer, say so. "
"Answer in at most three short sentences."
)
DEFAULT_QUESTION = "Why is this order blocked, and who should act on it?"
# What --sample prints, so you can see the output format without Ollama.
SAMPLE_ANSWER = (
"Sales order 9000001 is blocked because the credit check failed (status B), "
"and delivery block 01 puts it on credit hold. Credit management should review "
"customer 10100001 before the order can be released."
)
SAMPLE_STATS = {"load_duration": 1_210_000_000, "prompt_eval_count": 214,
"eval_count": 46, "eval_duration": 1_150_000_000}
def build_messages(question: str) -> list:
"""Put the instructions, the order and the question into chat messages."""
context = json.dumps({"order": ORDER, "code_meanings": CODE_MEANINGS}, indent=1)
return [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Order data:\n{context}\n\nQuestion: {question}"},
]
def print_stats(stats: dict, model: str) -> None:
"""Show how long the model took. Ollama reports every time in nanoseconds."""
seconds = stats["eval_duration"] / 1e9
speed = stats["eval_count"] / seconds if seconds else 0.0
print(f"\nModel: {model}")
print(f"Load time: {stats['load_duration'] / 1e9:.1f} s")
print(f"Prompt tokens: {stats['prompt_eval_count']}")
print(f"Answer tokens: {stats['eval_count']}")
print(f"Speed: {speed:.1f} tokens per second")
def ask_ollama(model: str, question: str, think: bool) -> tuple:
"""Send the chat to the local Ollama server and return the answer and timing numbers."""
from ollama import Client # imported here so --sample works before the library is installed
client = Client() # http://127.0.0.1:11434 unless the OLLAMA_HOST variable says otherwise
response = client.chat(
model=model,
messages=build_messages(question),
think=think, # qwen3 can "think" first; off by default
options={"temperature": 0, "num_ctx": 4096}, # steady answers, a 4,096-token window
)
stats = {key: getattr(response, key) or 0 for key in SAMPLE_STATS}
return response.message.content.strip(), stats
def main() -> None:
parser = argparse.ArgumentParser(description="Ask a local model about a blocked order.")
parser.add_argument("--model", default="qwen3:0.6b", help="an Ollama model you have pulled")
parser.add_argument("--question", default=DEFAULT_QUESTION, help="your own question")
parser.add_argument("--think", action="store_true", help="let the model reason first (slower)")
parser.add_argument("--sample", action="store_true", help="no model: print a sample answer")
args = parser.parse_args()
print(f"Question: {args.question}\n")
if args.sample:
print(SAMPLE_ANSWER)
print_stats(SAMPLE_STATS, "sample (no model was called)")
return
try:
from ollama import ResponseError
answer, stats = ask_ollama(args.model, args.question, args.think)
except ModuleNotFoundError:
sys.exit("The ollama library isn't installed. Run: pip install -r requirements.txt")
except ConnectionError:
sys.exit("Can't reach Ollama on this computer. Start the Ollama app "
"(Linux: sudo systemctl start ollama), then run this again.")
except ResponseError as error:
if error.status_code == 404:
sys.exit(f"Ollama doesn't have {args.model} yet. Run: ollama pull {args.model}")
sys.exit(f"Ollama answered with an error: {error}")
print(answer)
print_stats(stats, args.model)
if __name__ == "__main__":
main()
First run it with --sample. This prints the output format without calling any model, so it works even if Ollama isn't installed yet:
python unit12/local_chat.py --sample
What success looks like:
Question: Why is this order blocked, and who should act on it?
Sales order 9000001 is blocked because the credit check failed (status B), and delivery block 01 puts it on credit hold. Credit management should review customer 10100001 before the order can be released.
Model: sample (no model was called)
Load time: 1.2 s
Prompt tokens: 214
Answer tokens: 46
Speed: 40.0 tokens per second
Now ask the real local model:
python unit12/local_chat.py
What success looks like: the same layout, with the model's own answer and your machine's numbers. The answer will be worded differently and may be less precise than the sample. A small model can misread a code or add a detail that isn't in the data; spotting that is part of the unit. The first run's load time is the largest; run it again within 5 minutes and the load time drops close to zero, because the model is still in memory.
Ask your own question:
python unit12/local_chat.py --question "What should the clerk check first?"
Speeds vary a lot between machines. Write down your tokens per second; the exercise uses it.
In the course folder (not in unit12), create check_unit12.py, paste the code below and save. It uses only built-in Python, like the earlier checks.
"""Check that your computer is ready for Unit 12 (local models).
Run it from your course folder: python check_unit12.py
It uses built-in Python only. It measures memory and free disk, looks for a GPU, checks the
ollama library, the Ollama program and its local server, and the models this unit uses.
It changes nothing and sends nothing outside your computer.
"""
import ctypes
import importlib.metadata
import json
import os
import platform
import shutil
import subprocess
import sys
import urllib.request
REQUIRED_MODEL = "qwen3:0.6b" # every Unit 12 topic uses it
OPTIONAL_MODEL = "qwen3:1.7b" # used for size and speed comparisons
GB = 1024 ** 3
SERVER = "http://127.0.0.1:11434" # where Ollama listens by default
problems = 0
def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
"""Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
global problems
if ok:
print(f" OK {label}")
elif optional:
print(f" LATER {label} -> {fix}")
else:
problems += 1
print(f" MISSING {label} -> {fix}")
def total_memory() -> int:
"""Installed memory (RAM) in bytes, or 0 if it can't be read."""
if sys.platform == "win32":
class Status(ctypes.Structure):
_fields_ = [("dwLength", ctypes.c_ulong), ("dwMemoryLoad", ctypes.c_ulong),
("ullTotalPhys", ctypes.c_ulonglong), ("ullAvailPhys", ctypes.c_ulonglong),
("ullTotalPageFile", ctypes.c_ulonglong),
("ullAvailPageFile", ctypes.c_ulonglong),
("ullTotalVirtual", ctypes.c_ulonglong),
("ullAvailVirtual", ctypes.c_ulonglong),
("ullAvailExtendedVirtual", ctypes.c_ulonglong)]
status = Status()
status.dwLength = ctypes.sizeof(Status)
if ctypes.windll.kernel32.GlobalMemoryStatusEx(ctypes.byref(status)):
return status.ullTotalPhys
return 0
try:
return os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES")
except (ValueError, OSError, AttributeError):
pass
try: # macOS fallback
out = subprocess.run(["sysctl", "-n", "hw.memsize"], capture_output=True, text=True)
return int(out.stdout.strip())
except (OSError, ValueError):
return 0
def models_folder() -> str:
"""Where Ollama keeps models: OLLAMA_MODELS if set, else the default for this system."""
if os.environ.get("OLLAMA_MODELS"):
return os.environ["OLLAMA_MODELS"]
if sys.platform.startswith("linux") and os.path.isdir("/usr/share/ollama"):
return "/usr/share/ollama/.ollama/models"
return os.path.join(os.path.expanduser("~"), ".ollama", "models")
def free_disk(path: str) -> int:
"""Free bytes on the drive that holds path (or its nearest existing parent folder)."""
while path and not os.path.exists(path):
parent = os.path.dirname(path)
if parent == path:
break
path = parent
try:
return shutil.disk_usage(path or os.path.expanduser("~")).free
except OSError:
return 0
def gpu() -> str:
"""A short description of a GPU Ollama can use, or '' if none was found."""
if shutil.which("nvidia-smi"):
try:
out = subprocess.run(["nvidia-smi", "--query-gpu=name,memory.total",
"--format=csv,noheader"],
capture_output=True, text=True, timeout=20)
first = out.stdout.strip().splitlines()
if first:
return "NVIDIA " + first[0].replace("NVIDIA ", "")
except (OSError, subprocess.SubprocessError):
pass
if sys.platform == "darwin" and platform.machine() == "arm64":
return "Apple silicon GPU (shares the computer's memory)"
return ""
def ollama_get(path: str) -> dict:
"""GET a path from Ollama's server at its default address. Ignores proxies: it's local."""
opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
with opener.open(SERVER + path, timeout=5) as response:
return json.load(response)
print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
"the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")
print("\n2. This computer")
memory = total_memory()
if memory:
report(memory >= 7.5 * GB, f"memory: {memory / GB:.1f} GB",
"under 8 GB: use qwen3:0.6b only and close other apps", optional=memory >= 4 * GB)
else:
report(False, "memory: couldn't be measured", "check it in your system settings",
optional=True)
folder = models_folder()
space = free_disk(folder)
report(space >= 5 * GB, f"free disk where models go: {space / GB:.1f} GB",
"free up at least 5 GB on that drive, or move models with OLLAMA_MODELS")
card = gpu()
report(bool(card), f"GPU: {card}" if card else "GPU: none found",
"fine: this unit's small models run on the processor (CPU)", optional=True)
print("\n3. Ollama")
try:
report(True, f"ollama library {importlib.metadata.version('ollama')}")
except importlib.metadata.PackageNotFoundError:
report(False, "ollama library", "pip install -r requirements.txt (Step 3)")
report(bool(shutil.which("ollama")), "ollama command",
"install Ollama, then open a new terminal (Step 2)")
try:
version = ollama_get("/api/version").get("version", "?")
report(True, f"Ollama server {version} answers on this computer")
names = {m.get("name", "") for m in ollama_get("/api/tags").get("models", [])}
except (OSError, ValueError):
report(False, "Ollama server answers on this computer",
"start the Ollama app (Linux: sudo systemctl start ollama)")
names = None
print("\n4. Models")
if names is None:
report(False, f"model {REQUIRED_MODEL}", "start Ollama first, then run this again")
else:
report(REQUIRED_MODEL in names, f"model {REQUIRED_MODEL}", f"ollama pull {REQUIRED_MODEL} (Step 4)")
report(OPTIONAL_MODEL in names, f"model {OPTIONAL_MODEL}",
f"optional: ollama pull {OPTIONAL_MODEL} (Exercise)", optional=True)
print("\n5. Course folder")
report(os.path.exists(os.path.join("unit12", "local_chat.py")), "unit12/local_chat.py",
"create it (Step 5)")
print()
if problems:
print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
sys.exit(1)
print("All set. Your computer is ready for Unit 12.")
Run it:
python check_unit12.py
What success looks like (your numbers, versions and GPU line will differ):
1. Python
OK Python 3.13.16
OK virtual environment is active
2. This computer
OK memory: 15.8 GB
OK free disk where models go: 112.4 GB
LATER GPU: none found -> fine: this unit's small models run on the processor (CPU)
3. Ollama
OK ollama library 0.6.3
OK ollama command
OK Ollama server 0.x.y answers on this computer
4. Models
OK model qwen3:0.6b
LATER model qwen3:1.7b -> optional: ollama pull qwen3:1.7b (Exercise)
5. Course folder
OK unit12/local_chat.py
All set. Your computer is ready for Unit 12.
LATER lines are fine. MISSING lines must be fixed before the next topic; each names the step that fixes it. A computer with 4 to 8 GB of memory shows a LATER memory line: you can continue with qwen3:0.6b.
You should see requirements.txt, check_unit12.py and unit12/. You must not see .env. Models are stored in Ollama's own folder, outside your course folder, so they never reach Git.
Save:
git add requirements.txt check_unit12.py unit12
git commit -m "Set up Unit 12: Ollama and a local model"
This section is short on purpose: the buy-vs-build and embedded AI topics go deeper.
Hosted open models. As of October 2026, SAP's Python SDK reference for the generative AI hub lists models from Meta, IBM and Mistral AI next to the large commercial providers. SAP runs them; you call them by name, such as mistralai--mistral-small-instruct. Availability changes, so check SAP's current model list before you plan around one.
Ollama inside SAP AI Core. SAP's Developer Center tutorial builds a Docker image with Ollama and nginx, registers it as a custom serving template, and deploys it on the infer.s resource plan. It needs an SAP AI Core instance on the Standard or Extended plan. Treat it as a sketch for now: it needs paid SAP AI Core access, which Set up for Unit 5 discusses.
Same code shape. Your local_chat.py sends a chat to a server and reads back the answer. Swapping the local server for a hub model changes the client, not the idea.
Need
Learn with (this setup)
On SAP
Try an open model
Ollama on your laptop
An open model from the generative AI hub catalogue
Run a model you chose yourself
Ollama on your laptop
Custom serving template on SAP AI Core (Standard or Extended plan)
Keep data inside a boundary
Data never leaves the laptop
Model runs inside your SAP AI Core tenant; check your contract and data policy
Bind to this computer only. Ollama listens on 127.0.0.1 by default. Exposing it to a network gives anyone who can reach it free use of your hardware, and the local server doesn't check who is calling. Put a proper gateway in front before anyone else uses it.
Licences per model. Ollama's MIT licence doesn't cover the models. Record each model's licence and version with every result.
Pin the tag. A tag like latest can change. Record the exact tag you tested, and re-test when you update.
Measure before you swap. A local model that is cheaper and wrong more often may cost more in rework. The fine-tuning and buy-vs-build topics measure this with the evaluation tools from Unit 8.
Capacity. One laptop serves one person. Real users need servers sized for peak load, and GPUs are the expensive part.
Clean core. Nothing here touches S/4HANA. The model runs side by side and reads data that your app passes in.
Pulling qwen3 without a tag. You get the 5.2 GB 8b model. Always type the size.
Judging speed on the first run. The first answer includes load time. Compare speeds on the second run, or use the Speed line, which excludes loading.
Trusting a small model's reading of codes. It may confidently explain 01 wrongly. Give it the meanings, as the script does, and check the answer against them.
Treating the sample answer as the model's.--sample prints fixed text. It shows the format only.
Comparing models on one question. One answer says little. The later topics use evaluation sets.
Add a row: today's date, qwen3:0.6b, the size from ollama list, the processor, the load time and speed from the second run, and yes or no for whether the answer matches CODE_MEANINGS (credit check failed, credit hold, no billing block).
Repeat steps 3 to 5 with --model qwen3:1.7b.
Under the table, write one sentence: which model you would pick for this task on your machine, and why.
Commit:
git add unit12/model_log.md
git commit -m "Unit 12: first local model log"
Done whenmodel_log.md has two rows with real numbers and a one-sentence choice, and python check_unit12.py ends with All set.
Pick one answer for each question. The explanation appears after you choose.
1In the mental model, what does Ollama correspond to?
Answer: B. Ollama is the runner: it loads the model file into memory and turns prompts into answers. The model file holds the parameters, and your script is the remote control that sends requests.
2Why does the setup always pull qwen3:0.6b and never just qwen3?
Answer: C. The registry marks the 8b model as latest, so the bare name pulls 5.2 GB instead of about 0.5 GB. Typing the tag keeps the model small enough for most laptops.
3Ollama returns eval_count 46 and eval_duration 1150000000. What is the speed?
Answer: D. Timings are in nanoseconds, so 1,150,000,000 is 1.15 seconds. 46 tokens divided by 1.15 seconds is 40 tokens per second, which is what the script prints.
4local_chat.py prints "Ollama doesn't have qwen3:0.6b yet". What happened?
Answer: C. The script maps a ResponseError with status 404 to that message. A server that isn't running raises ConnectionError instead, which prints a different message.
5A colleague wants to let the whole team use the Ollama on her desktop by making it listen on the office network. What would you do?
Answer: B. Ollama listens on 127.0.0.1 by default for a reason: the local server doesn't check who is calling. Exposing it lets anyone on the network use the hardware and the models, so put a proper gateway in front first.
6Why does check_unit12.py build its web opener with an empty ProxyHandler?
Answer: D. A proxy setting can send even local requests to a company proxy, which can't reach your laptop's port 11434. Ignoring proxies for a local call avoids that, which is also why the troubleshooting table mentions NO_PROXY.
7Which SAP route matches "run a model we chose ourselves, inside SAP's platform"?
Answer: C. SAP's Developer Center tutorial deploys Ollama as a custom serving template on SAP AI Core, Standard or Extended plan. Hub catalogue models are chosen and run by SAP, which is a different route.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Ollama Windows (Ollama documentation)— Windows 10 22H2 or newer; OllamaSetup.exe installs in the home folder without Administrator; at least 4 GB for the program; models can need tens to hundreds of GB; OLLAMA_MODELS moves them; API at http://localhost:11434
Ollama macOS (Ollama documentation)— macOS Sonoma (v14) or newer; Apple M-series get CPU and GPU, Intel Macs CPU only; drag to Applications; first launch offers to link the ollama command; models in ~/.ollama
Ollama Linux (Ollama documentation)— curl -fsSL https://ollama.com/install.sh | sh; verify with ollama -v; systemd service; systemctl status and journalctl for logs
Ollama FAQ (Ollama documentation)— model folders per system; default context 4096 tokens; ollama ps Processor column; models kept in memory 5 minutes; binds 127.0.0.1:11434; local runs keep prompts on the machine; HTTPS_PROXY for downloads, avoid HTTP_PROXY