Orchestrate

Set up for Unit 12: local models

Install Ollama, download a small open model, check your memory and disk, and ask a model running on your own laptop about a blocked SAP order.

Updated Oct 8, 2026Foundational 8 minDeep 35 min
Foundational layer · 8 min read

The 60-second version

So far in this course, every model has run in someone else's data centre. You sent a prompt over the internet and paid per request. Unit 12 asks a different question: when should the model run somewhere you control?

This setup installs one free program, Ollama, that downloads open models and runs them on your own laptop. You then download a small model, about half a gigabyte, and ask it about a blocked sales order.

Three things change when the model is local:

  • The data stays put. The prompt never leaves the computer.
  • There is no bill per request. You pay with hardware and electricity instead.
  • Quality and speed depend on the machine. A small model on a laptop is fast enough to learn with, and clearly weaker than the large hosted models you used before.

Setup takes 30 to 60 minutes and costs nothing. It needs about 5 GB of free disk and works best with 8 GB of memory or more.

Why it matters to the business

Most SAP AI work will keep using hosted models through SAP's generative AI hub. Local and small models still matter, for three business reasons.

  • Data that must not leave. Some content, such as HR cases or unreleased financials, may be barred from external model providers. A model you host yourself can keep it inside your own boundary.
  • Cost at volume. A task that runs millions of times, such as tagging every incoming invoice line, can cost less on a small model you host than on a large hosted one. The unit's buy-vs-build topic weighs this properly.
  • Fit for purpose. Many tasks don't need the biggest model. A small model, sometimes tuned for one job, can be good enough and much faster.

Take the running example: a clerk asks why sales order 9000001 is blocked. The order data includes a customer number and an amount. If policy says that data may not go to an outside provider, a model running inside the company is one way to still offer the assistant.

The trade-off is real. A small model makes more mistakes, and running models yourself means owning updates, security and capacity. This unit teaches you to measure that trade-off rather than guess.

How SAP does it

As of October 2026, SAP offers two routes that touch this unit.

  • Open models in the generative AI hub. SAP's Python SDK reference lists models from Meta, IBM and Mistral AI alongside Amazon, Anthropic, Google and OpenAI. Examples include meta--llama3.1-70b-instruct and mistralai--mistral-small-instruct. You call them like any other hub model, and SAP runs them.
  • Your own model on SAP AI Core. An SAP Developer Center tutorial packages Ollama, the same runner you install today, as a custom serving template on SAP AI Core. It needs an SAP AI Core instance on the Standard or Extended plan, a Docker image and a GitHub repository.

Neither route is free, so this unit teaches the ideas on your laptop first. AI features built into SAP applications are a third, separate case, covered by the unit's topic on embedded AI.

What this unit adds

Item What it is Cost Used in
Ollama A free program that downloads open models and runs them on your computer Free Every Unit 12 topic
qwen3:0.6b A small open model from the Qwen3 family, about 0.5 GB Free Every Unit 12 topic
qwen3:1.7b (optional) The next size up, about 1.4 GB Free Size and speed comparisons
ollama Python library Lets your Python code talk to Ollama Free (MIT) Every Unit 12 topic
unit12/local_chat.py Asks the local model about a blocked order and times it Free Fine-tuning, quantization, buy vs. build

Later topics in the unit add their own libraries in their own steps, for example for reading documents.

Time and money

  • Time: 30 to 60 minutes. Most of it is downloading.
  • Money for a learner: nothing. No account, no key, no per-request cost.
  • Disk: keep at least 5 GB free. Ollama's Windows documentation asks for at least 4 GB for the program alone, and each model adds its own size.
  • Memory: 8 GB or more is comfortable. With 4 to 8 GB, stay with the smallest model and close other apps.
  • A graphics card (GPU) is optional. It makes answers faster. Everything in this unit also runs on the processor.
  • Money for a company: the software is free. Hosting models for real users means servers, often with GPUs, plus people to run them. That cost is the heart of the buy-vs-build decision later in the unit.

Questions to ask

Ask IT before you install:

  • May we install Ollama on company laptops, and may it download models from the internet?
  • Which open-model licences has legal approved? Each model has its own licence, separate from Ollama's.
  • Does our data policy treat a local model differently from a hosted one? Which data classes would it unlock?
  • Will our proxy allow the model downloads? Ollama fetches models from its own registry.
  • Do our laptops have the memory for this, or should learners use a lab machine?

Ask a vendor or partner who proposes a self-hosted model:

  • Who patches the model runner and the model, and how often?
  • How did you measure quality against the hosted model you are replacing?
  • What does it cost per month at our volume, including idle time?

Common misconceptions

  • "Local means as good as the big models, just free." A half-gigabyte model is far weaker than a large hosted one. It is useful for learning and for narrow tasks.
  • "Local means secure." The data stays on the machine, which helps. The machine, the runner and the model download still need the same care as any other software.
  • "You need a gaming GPU." Small models run on an ordinary laptop processor. A GPU makes them faster, not possible.
  • "Open model means no licence terms." Each model has its own licence, and some restrict commercial use. Check before production.
  • "Ollama is an SAP product." It is an independent open-source project. SAP's tutorial uses it on SAP AI Core, which is a different thing from SAP supporting it.

Key terms

  • Local model: a language model that runs on a computer you control, not in a provider's cloud.
  • Open model (open-weight model): a model whose trained parameters you may download and run, under its licence.
  • Model runner: a program that loads a model file and answers prompts. Ollama is one.
  • Parameters: the numbers a model learned. "0.6b" means about 0.6 billion. More usually means better and slower.
  • GPU: a graphics processor. It does the maths of a model much faster than the main processor.
  • Memory (RAM): the computer's working space. The model must fit in it to run well.
  • Tokens per second: how fast a model writes. It decides whether an answer feels instant or slow.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What is the main reason a team would run a model locally instead of calling a hosted one?

    Answer: B. When the model runs on your own machine, the prompt never leaves it. That can unlock data a policy keeps away from outside providers. Accuracy is usually lower, not higher, for small local models.
  2. 2A sponsor says "a local model is free, so let's replace every hosted call with one". What is the soundest reply?

    Answer: C. There is no bill per request, but someone pays for servers, updates and capacity. A small model also makes more mistakes, so the swap needs a quality comparison before it is decided.
  3. 3Which SAP route lets a team run its own model inside SAP's platform?

    Answer: D. SAP's Developer Center tutorial packages Ollama as a custom serving template on SAP AI Core, which needs a Standard or Extended plan. The hub's catalogue is the other route, where SAP runs the model for you.
  4. 4A learner's laptop has 6 GB of memory and no graphics card. What should they do?

    Answer: A. Small models run on the processor; a GPU only makes them faster. With limited memory, the smallest model is the one that fits comfortably.
  5. 5Which question should go to IT before anyone installs Ollama?

    Answer: B. Ollama installs software and downloads models from the internet, which some company policies and proxies block. It doesn't need SAP licences or an SAP release.
  6. 6A vendor says "it's an open model, so there are no licence terms to check". What is wrong?

    Answer: D. The runner and the model are licensed separately. Legal should approve the specific model's licence before production use.
Deep layer · 35 min read

Mental model: a model is a file, a runner is a player

Think of music. A song is a file. A player opens the file and turns it into sound. You can play the same song on different players, and the same player handles many songs.

Local models work the same way:

  • The model is a file, often a few hundred megabytes to many gigabytes. It holds the learned parameters.
  • The runner is the player. Ollama loads the file into memory, feeds your prompt through it, and streams the answer back.
  • Your code is the remote control. It sends a request to the runner and reads the answer. It never touches the model file directly.

That split explains most of this unit. Fine-tuning changes the file. Quantization makes the file smaller. Buy vs. build asks who runs the player and where.

How it works

flowchart LR
  C[local_chat.py] -->|HTTP request<br/>127.0.0.1:11434| O[Ollama server]
  O -->|loads once| M[(qwen3:0.6b<br/>model file)]
  M --> O
  O -->|answer and timings| C
  R[Ollama registry] -.->|ollama pull<br/>one-time download| M

The pieces

  • The Ollama server runs in the background on your computer. On Windows and macOS the Ollama app starts it; on Linux it usually runs as a system service. Ollama's FAQ says it listens on 127.0.0.1 port 11434 by default. 127.0.0.1 means "this computer only", so other machines can't reach it.
  • ollama pull downloads a model from Ollama's registry into a models folder. After that, no internet is needed to use it.
  • The ollama Python library sends your chat to the server and returns the answer as an object. It is the official client, MIT licensed.

Where the model lives

System Default models folder (Ollama FAQ)
Windows C:\Users\<you>\.ollama\models
macOS ~/.ollama/models
Linux (service install) /usr/share/ollama/.ollama/models

The OLLAMA_MODELS environment variable moves the folder, for example to a larger drive.

Why memory decides what you can run

To answer, the runner loads the whole model into memory: GPU memory if it can, otherwise main memory. It also needs room for the conversation so far, called the context. Ollama's FAQ gives a default context of 4,096 tokens.

The registry shows how size grows with parameters for one model family:

Tag Parameters Download
qwen3:0.6b about 0.6 billion 523 MB
qwen3:1.7b about 1.7 billion 1.4 GB
qwen3:4b about 4 billion 2.5 GB
qwen3:8b (also qwen3:latest) about 8 billion 5.2 GB

Our course rule of thumb: the download size plus 1 to 2 GB must fit in free memory. That is why this setup uses qwen3:0.6b and never the bare name qwen3, which would download the 8b model.

CPU or GPU

Ollama uses a supported GPU when it finds one and falls back to the processor (CPU) when it doesn't. Its hardware page lists NVIDIA cards with compute capability 5.0 or higher and driver 550 or newer, AMD cards through ROCm v7 or Vulkan, and Apple silicon through Metal. Apple silicon Macs share one pool of memory between CPU and GPU.

ollama ps shows each loaded model and, in its Processor column, how much of it sits on the GPU. "100% GPU" means all of it. By default a model stays loaded for 5 minutes after its last request, so the second question is faster than the first.

What the timings mean

Every answer from Ollama carries timing fields. Ollama's API documentation says all times are in nanoseconds (billionths of a second).

Field Meaning
load_duration Time to load the model into memory; near zero if it was already loaded
prompt_eval_count Tokens in your prompt
eval_count Tokens in the answer
eval_duration Time spent writing the answer

Speed in tokens per second is eval_count / (eval_duration / 1,000,000,000). You will use this number all unit to compare models.

Build it yourself: your first local model

You will install Ollama, download a small model, check your machine, and ask the model why a made-up sales order is blocked. No account, no key and no per-request cost.

Before you start: complete Set up your computer for this course and Set up for Unit 2. They install Python, VS Code and Git, and create your orchestrate-course folder with its .venv and requirements.txt. This walkthrough doesn't repeat those steps.

flowchart LR
  S1[Step 1<br/>open folder] --> S2[Step 2<br/>install Ollama]
  S2 --> S3[Step 3<br/>Python library]
  S3 --> S4[Step 4<br/>pull a model]
  S4 --> S5[Step 5<br/>local_chat.py]
  S5 --> S6[Step 6<br/>check_unit12.py]
  S6 --> S7[Step 7<br/>save in Git]

What you need

  • Your course folder from earlier units, with .venv and Python 3.11 or newer.
  • Windows 10 22H2 or newer, macOS Sonoma (14) or newer, or a recent Linux. These are Ollama's stated minimums for Windows and macOS.
  • At least 5 GB of free disk and, ideally, 8 GB of memory.
  • About 30 to 60 minutes and an internet connection for the downloads. After that, the model works offline.
  • Cost: free. No account and no key.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Create the folder for this unit:

    mkdir unit12

Run every command in this topic from the course folder, not from inside unit12.

Step 2: Install Ollama

Windows

  1. In your browser, go to https://ollama.com/download and choose Windows.
  2. Run the downloaded OllamaSetup.exe and follow the installer. Ollama's documentation says it installs in your home folder and doesn't need Administrator rights.
  3. When it finishes, Ollama runs in the background with an icon in the system tray, near the clock.

macOS

  1. Go to https://ollama.com/download and choose macOS.
  2. Open the downloaded .dmg file and drag Ollama into the Applications folder.
  3. Open Ollama from Applications. On first launch it offers to link the ollama command into /usr/local/bin; accept, and enter your Mac password if asked.

Linux

  1. Run Ollama's install script. It asks for your password because it installs system-wide:

    curl -fsSL https://ollama.com/install.sh | sh
  2. Check that the service is running:

    systemctl status ollama

    Look for active (running). Press q to leave the status view. If it isn't running, start it with sudo systemctl start ollama.

All systems: check the install

  1. Close VS Code's terminal and open a new one (Terminal > New Terminal), so it sees the new command. Turn .venv back on as in Step 1.

  2. Run:

    ollama -v

What success looks like: one line with the version number, such as ollama version is 0.x.y. Any version is fine.

Step 3: Add the Python library

  1. Open requirements.txt and add this line at the end, then save:

    ollama
  2. Install (the same on every system):

    pip install -r requirements.txt

What success looks like (your version may be newer):

Successfully installed ollama-0.6.3

If pip says Requirement already satisfied for everything, that is fine too.

Step 4: Download a small model and try it

  1. Pull the model. Always type the tag :0.6b; the bare name qwen3 downloads the 5.2 GB model.

    ollama pull qwen3:0.6b

    A progress bar runs for a minute or two. The last line says success.

  2. List your models:

    ollama list

    You should see a line starting with qwen3:0.6b, with its size, about 0.5 GB.

  3. Chat with it in the terminal:

    ollama run qwen3:0.6b "In one sentence, what is a sales order?"

    The model writes an answer. It may first show its reasoning ("thinking"), because Qwen3 models can reason before they answer. The first run takes longer while the model loads.

  4. See where it is running:

    ollama ps

    The Processor column says 100% CPU, 100% GPU, or a split. Any of these is fine for this unit.

Step 5: Ask the model from Python

You will write a short script that sends a made-up blocked order to the model and prints the answer and its speed. The order uses field names from SAP's sales order API, as in earlier units.

  1. In VS Code's file list, right-click unit12, choose New File, and name it local_chat.py.
  2. Paste the code below and save.
"""Ask a small language model on your own computer about a blocked SAP sales order.

Run it from your course folder:
    python unit12/local_chat.py --sample                 # no model: shows the output format
    python unit12/local_chat.py                          # asks qwen3:0.6b through Ollama
    python unit12/local_chat.py --model qwen3:1.7b       # a bigger model you pulled
    python unit12/local_chat.py --question "What should the clerk check first?"

The model runs in Ollama on this computer. The order and the prompt never leave it.
The order is made up, in the shape of SAP's sales order API. The meaning of a block code
is set in each SAP system's configuration; the meanings below belong to this sample only.
"""
import argparse
import json
import sys

# One made-up blocked sales order, with field names from SAP's sales order API.
ORDER = {
    "SalesOrder": "9000001",
    "SoldToParty": "10100001",
    "TotalNetAmount": "18250.00",
    "TransactionCurrency": "EUR",
    "TotalCreditCheckStatus": "B",
    "DeliveryBlockReason": "01",
    "HeaderBillingBlockReason": "",
}
# What the codes mean in this made-up system. Your system's configuration decides the real ones.
CODE_MEANINGS = {
    "TotalCreditCheckStatus B": "credit check failed",
    "DeliveryBlockReason 01": "credit hold",
    "empty value": "no block of that kind",
}
SYSTEM = (
    "You help order-to-cash clerks understand blocked SAP sales orders. "
    "Use only the order data and code meanings given. If the data can't answer, say so. "
    "Answer in at most three short sentences."
)
DEFAULT_QUESTION = "Why is this order blocked, and who should act on it?"

# What --sample prints, so you can see the output format without Ollama.
SAMPLE_ANSWER = (
    "Sales order 9000001 is blocked because the credit check failed (status B), "
    "and delivery block 01 puts it on credit hold. Credit management should review "
    "customer 10100001 before the order can be released."
)
SAMPLE_STATS = {"load_duration": 1_210_000_000, "prompt_eval_count": 214,
                "eval_count": 46, "eval_duration": 1_150_000_000}


def build_messages(question: str) -> list:
    """Put the instructions, the order and the question into chat messages."""
    context = json.dumps({"order": ORDER, "code_meanings": CODE_MEANINGS}, indent=1)
    return [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": f"Order data:\n{context}\n\nQuestion: {question}"},
    ]


def print_stats(stats: dict, model: str) -> None:
    """Show how long the model took. Ollama reports every time in nanoseconds."""
    seconds = stats["eval_duration"] / 1e9
    speed = stats["eval_count"] / seconds if seconds else 0.0
    print(f"\nModel:          {model}")
    print(f"Load time:      {stats['load_duration'] / 1e9:.1f} s")
    print(f"Prompt tokens:  {stats['prompt_eval_count']}")
    print(f"Answer tokens:  {stats['eval_count']}")
    print(f"Speed:          {speed:.1f} tokens per second")


def ask_ollama(model: str, question: str, think: bool) -> tuple:
    """Send the chat to the local Ollama server and return the answer and timing numbers."""
    from ollama import Client  # imported here so --sample works before the library is installed

    client = Client()  # http://127.0.0.1:11434 unless the OLLAMA_HOST variable says otherwise
    response = client.chat(
        model=model,
        messages=build_messages(question),
        think=think,                                    # qwen3 can "think" first; off by default
        options={"temperature": 0, "num_ctx": 4096},    # steady answers, a 4,096-token window
    )
    stats = {key: getattr(response, key) or 0 for key in SAMPLE_STATS}
    return response.message.content.strip(), stats


def main() -> None:
    parser = argparse.ArgumentParser(description="Ask a local model about a blocked order.")
    parser.add_argument("--model", default="qwen3:0.6b", help="an Ollama model you have pulled")
    parser.add_argument("--question", default=DEFAULT_QUESTION, help="your own question")
    parser.add_argument("--think", action="store_true", help="let the model reason first (slower)")
    parser.add_argument("--sample", action="store_true", help="no model: print a sample answer")
    args = parser.parse_args()

    print(f"Question: {args.question}\n")
    if args.sample:
        print(SAMPLE_ANSWER)
        print_stats(SAMPLE_STATS, "sample (no model was called)")
        return

    try:
        from ollama import ResponseError
        answer, stats = ask_ollama(args.model, args.question, args.think)
    except ModuleNotFoundError:
        sys.exit("The ollama library isn't installed. Run: pip install -r requirements.txt")
    except ConnectionError:
        sys.exit("Can't reach Ollama on this computer. Start the Ollama app "
                 "(Linux: sudo systemctl start ollama), then run this again.")
    except ResponseError as error:
        if error.status_code == 404:
            sys.exit(f"Ollama doesn't have {args.model} yet. Run: ollama pull {args.model}")
        sys.exit(f"Ollama answered with an error: {error}")

    print(answer)
    print_stats(stats, args.model)


if __name__ == "__main__":
    main()
  1. First run it with --sample. This prints the output format without calling any model, so it works even if Ollama isn't installed yet:

    python unit12/local_chat.py --sample

What success looks like:

Question: Why is this order blocked, and who should act on it?

Sales order 9000001 is blocked because the credit check failed (status B), and delivery block 01 puts it on credit hold. Credit management should review customer 10100001 before the order can be released.

Model:          sample (no model was called)
Load time:      1.2 s
Prompt tokens:  214
Answer tokens:  46
Speed:          40.0 tokens per second
  1. Now ask the real local model:

    python unit12/local_chat.py

What success looks like: the same layout, with the model's own answer and your machine's numbers. The answer will be worded differently and may be less precise than the sample. A small model can misread a code or add a detail that isn't in the data; spotting that is part of the unit. The first run's load time is the largest; run it again within 5 minutes and the load time drops close to zero, because the model is still in memory.

  1. Ask your own question:

    python unit12/local_chat.py --question "What should the clerk check first?"

Speeds vary a lot between machines. Write down your tokens per second; the exercise uses it.

Step 6: Run the Unit 12 check

  1. In the course folder (not in unit12), create check_unit12.py, paste the code below and save. It uses only built-in Python, like the earlier checks.
"""Check that your computer is ready for Unit 12 (local models).

Run it from your course folder:  python check_unit12.py
It uses built-in Python only. It measures memory and free disk, looks for a GPU, checks the
ollama library, the Ollama program and its local server, and the models this unit uses.
It changes nothing and sends nothing outside your computer.
"""
import ctypes
import importlib.metadata
import json
import os
import platform
import shutil
import subprocess
import sys
import urllib.request

REQUIRED_MODEL = "qwen3:0.6b"   # every Unit 12 topic uses it
OPTIONAL_MODEL = "qwen3:1.7b"   # used for size and speed comparisons
GB = 1024 ** 3
SERVER = "http://127.0.0.1:11434"   # where Ollama listens by default
problems = 0


def report(ok: bool, label: str, fix: str = "", optional: bool = False) -> None:
    """Print one line: OK, MISSING (must fix) or LATER (optional for now)."""
    global problems
    if ok:
        print(f"  OK       {label}")
    elif optional:
        print(f"  LATER    {label}  ->  {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label}  ->  {fix}")


def total_memory() -> int:
    """Installed memory (RAM) in bytes, or 0 if it can't be read."""
    if sys.platform == "win32":
        class Status(ctypes.Structure):
            _fields_ = [("dwLength", ctypes.c_ulong), ("dwMemoryLoad", ctypes.c_ulong),
                        ("ullTotalPhys", ctypes.c_ulonglong), ("ullAvailPhys", ctypes.c_ulonglong),
                        ("ullTotalPageFile", ctypes.c_ulonglong),
                        ("ullAvailPageFile", ctypes.c_ulonglong),
                        ("ullTotalVirtual", ctypes.c_ulonglong),
                        ("ullAvailVirtual", ctypes.c_ulonglong),
                        ("ullAvailExtendedVirtual", ctypes.c_ulonglong)]
        status = Status()
        status.dwLength = ctypes.sizeof(Status)
        if ctypes.windll.kernel32.GlobalMemoryStatusEx(ctypes.byref(status)):
            return status.ullTotalPhys
        return 0
    try:
        return os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES")
    except (ValueError, OSError, AttributeError):
        pass
    try:  # macOS fallback
        out = subprocess.run(["sysctl", "-n", "hw.memsize"], capture_output=True, text=True)
        return int(out.stdout.strip())
    except (OSError, ValueError):
        return 0


def models_folder() -> str:
    """Where Ollama keeps models: OLLAMA_MODELS if set, else the default for this system."""
    if os.environ.get("OLLAMA_MODELS"):
        return os.environ["OLLAMA_MODELS"]
    if sys.platform.startswith("linux") and os.path.isdir("/usr/share/ollama"):
        return "/usr/share/ollama/.ollama/models"
    return os.path.join(os.path.expanduser("~"), ".ollama", "models")


def free_disk(path: str) -> int:
    """Free bytes on the drive that holds path (or its nearest existing parent folder)."""
    while path and not os.path.exists(path):
        parent = os.path.dirname(path)
        if parent == path:
            break
        path = parent
    try:
        return shutil.disk_usage(path or os.path.expanduser("~")).free
    except OSError:
        return 0


def gpu() -> str:
    """A short description of a GPU Ollama can use, or '' if none was found."""
    if shutil.which("nvidia-smi"):
        try:
            out = subprocess.run(["nvidia-smi", "--query-gpu=name,memory.total",
                                  "--format=csv,noheader"],
                                 capture_output=True, text=True, timeout=20)
            first = out.stdout.strip().splitlines()
            if first:
                return "NVIDIA " + first[0].replace("NVIDIA ", "")
        except (OSError, subprocess.SubprocessError):
            pass
    if sys.platform == "darwin" and platform.machine() == "arm64":
        return "Apple silicon GPU (shares the computer's memory)"
    return ""


def ollama_get(path: str) -> dict:
    """GET a path from Ollama's server at its default address. Ignores proxies: it's local."""
    opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
    with opener.open(SERVER + path, timeout=5) as response:
        return json.load(response)


print("\n1. Python")
v = sys.version_info
report(v >= (3, 11), f"Python {v.major}.{v.minor}.{v.micro}",
       "the course needs Python 3.11 or newer (see Set up for Unit 2, Step 1)")
report(sys.prefix != sys.base_prefix, "virtual environment is active", "activate .venv (Step 1)")

print("\n2. This computer")
memory = total_memory()
if memory:
    report(memory >= 7.5 * GB, f"memory: {memory / GB:.1f} GB",
           "under 8 GB: use qwen3:0.6b only and close other apps", optional=memory >= 4 * GB)
else:
    report(False, "memory: couldn't be measured", "check it in your system settings",
           optional=True)
folder = models_folder()
space = free_disk(folder)
report(space >= 5 * GB, f"free disk where models go: {space / GB:.1f} GB",
       "free up at least 5 GB on that drive, or move models with OLLAMA_MODELS")
card = gpu()
report(bool(card), f"GPU: {card}" if card else "GPU: none found",
       "fine: this unit's small models run on the processor (CPU)", optional=True)

print("\n3. Ollama")
try:
    report(True, f"ollama library {importlib.metadata.version('ollama')}")
except importlib.metadata.PackageNotFoundError:
    report(False, "ollama library", "pip install -r requirements.txt (Step 3)")
report(bool(shutil.which("ollama")), "ollama command",
       "install Ollama, then open a new terminal (Step 2)")
try:
    version = ollama_get("/api/version").get("version", "?")
    report(True, f"Ollama server {version} answers on this computer")
    names = {m.get("name", "") for m in ollama_get("/api/tags").get("models", [])}
except (OSError, ValueError):
    report(False, "Ollama server answers on this computer",
           "start the Ollama app (Linux: sudo systemctl start ollama)")
    names = None

print("\n4. Models")
if names is None:
    report(False, f"model {REQUIRED_MODEL}", "start Ollama first, then run this again")
else:
    report(REQUIRED_MODEL in names, f"model {REQUIRED_MODEL}", f"ollama pull {REQUIRED_MODEL} (Step 4)")
    report(OPTIONAL_MODEL in names, f"model {OPTIONAL_MODEL}",
           f"optional: ollama pull {OPTIONAL_MODEL} (Exercise)", optional=True)

print("\n5. Course folder")
report(os.path.exists(os.path.join("unit12", "local_chat.py")), "unit12/local_chat.py",
       "create it (Step 5)")

print()
if problems:
    print(f"{problems} item(s) to fix. Fix them in order, then run this again.")
    sys.exit(1)
print("All set. Your computer is ready for Unit 12.")
  1. Run it:

    python check_unit12.py

What success looks like (your numbers, versions and GPU line will differ):

1. Python
  OK       Python 3.13.16
  OK       virtual environment is active

2. This computer
  OK       memory: 15.8 GB
  OK       free disk where models go: 112.4 GB
  LATER    GPU: none found  ->  fine: this unit's small models run on the processor (CPU)

3. Ollama
  OK       ollama library 0.6.3
  OK       ollama command
  OK       Ollama server 0.x.y answers on this computer

4. Models
  OK       model qwen3:0.6b
  LATER    model qwen3:1.7b  ->  optional: ollama pull qwen3:1.7b (Exercise)

5. Course folder
  OK       unit12/local_chat.py

All set. Your computer is ready for Unit 12.

LATER lines are fine. MISSING lines must be fixed before the next topic; each names the step that fixes it. A computer with 4 to 8 GB of memory shows a LATER memory line: you can continue with qwen3:0.6b.

Step 7: Save your work in Git

  1. Check what Git sees:

    git status

    You should see requirements.txt, check_unit12.py and unit12/. You must not see .env. Models are stored in Ollama's own folder, outside your course folder, so they never reach Git.

  2. Save:

    git add requirements.txt check_unit12.py unit12
    git commit -m "Set up Unit 12: Ollama and a local model"

How the code works

Part What it does
ORDER A made-up blocked order with SAP API field names
CODE_MEANINGS What the codes mean in this sample; real meanings come from each system's configuration
SYSTEM Instructions: answer only from the data, say so if the data can't answer, keep it short
build_messages() Packs the instructions, the order and the question into chat messages
Client() Connects to Ollama at 127.0.0.1:11434, or at OLLAMA_HOST if you set it
client.chat(...) Sends the chat; think=False skips Qwen3's reasoning step, temperature: 0 keeps answers steady, num_ctx: 4096 sets the context window
print_stats() Turns Ollama's nanosecond timings into seconds and tokens per second
except ConnectionError Ollama isn't running
except ResponseError with 404 The model hasn't been pulled
--sample Prints a fixed answer and timings without any model
total_memory() in the check Reads installed memory with built-in tools on each system
gpu() in the check Looks for nvidia-smi or an Apple silicon Mac
ollama_get() in the check Asks Ollama's /api/version and /api/tags endpoints, ignoring proxies because the server is local

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
ModuleNotFoundError: No module named 'ollama', or the script says the library isn't installed The library isn't in the Python you are using Check for (.venv) in the prompt, then pip install -r requirements.txt
ollama is not recognized, or command not found The terminal started before Ollama was installed, or the macOS link was declined Open a new terminal. On macOS, open the Ollama app again and accept the command-line link
Can't reach Ollama on this computer The Ollama server isn't running Windows/macOS: open the Ollama app. Linux: sudo systemctl start ollama
Ollama doesn't have qwen3:0.6b yet The model wasn't pulled, or the tag was typed differently Run ollama pull qwen3:0.6b and check ollama list
ollama pull hangs, or fails with a TLS or proxy error Your network or proxy blocks Ollama's registry Try another network, or ask IT to allow ollama.com. Ollama's FAQ says to set HTTPS_PROXY for the server, not HTTP_PROXY
The script can't reach Ollama even though ollama ps works An HTTP_PROXY setting is sending local requests to the proxy Add localhost,127.0.0.1 to the NO_PROXY environment variable, or remove HTTP_PROXY
Answers take minutes, or the computer freezes The model doesn't fit in memory Close other apps, stay with qwen3:0.6b, and check ollama ps
The model shows long "thinking" before every answer in ollama run Qwen3 reasons first by default Fine for now; the Python script turns it off with think=False
free disk where models go is MISSING Less than 5 GB free on that drive Free up space, or set OLLAMA_MODELS to a folder on a larger drive and restart Ollama
Linux: MISSING Ollama server right after a reboot The service didn't start sudo systemctl enable --now ollama

Where this shows up in SAP

This section is short on purpose: the buy-vs-build and embedded AI topics go deeper.

  • Hosted open models. As of October 2026, SAP's Python SDK reference for the generative AI hub lists models from Meta, IBM and Mistral AI next to the large commercial providers. SAP runs them; you call them by name, such as mistralai--mistral-small-instruct. Availability changes, so check SAP's current model list before you plan around one.
  • Ollama inside SAP AI Core. SAP's Developer Center tutorial builds a Docker image with Ollama and nginx, registers it as a custom serving template, and deploys it on the infer.s resource plan. It needs an SAP AI Core instance on the Standard or Extended plan. Treat it as a sketch for now: it needs paid SAP AI Core access, which Set up for Unit 5 discusses.
  • Same code shape. Your local_chat.py sends a chat to a server and reads back the answer. Swapping the local server for a hub model changes the client, not the idea.
Need Learn with (this setup) On SAP
Try an open model Ollama on your laptop An open model from the generative AI hub catalogue
Run a model you chose yourself Ollama on your laptop Custom serving template on SAP AI Core (Standard or Extended plan)
Keep data inside a boundary Data never leaves the laptop Model runs inside your SAP AI Core tenant; check your contract and data policy

Production concerns

  • Bind to this computer only. Ollama listens on 127.0.0.1 by default. Exposing it to a network gives anyone who can reach it free use of your hardware, and the local server doesn't check who is calling. Put a proper gateway in front before anyone else uses it.
  • Licences per model. Ollama's MIT licence doesn't cover the models. Record each model's licence and version with every result.
  • Pin the tag. A tag like latest can change. Record the exact tag you tested, and re-test when you update.
  • Measure before you swap. A local model that is cheaper and wrong more often may cost more in rework. The fine-tuning and buy-vs-build topics measure this with the evaluation tools from Unit 8.
  • Capacity. One laptop serves one person. Real users need servers sized for peak load, and GPUs are the expensive part.
  • Clean core. Nothing here touches S/4HANA. The model runs side by side and reads data that your app passes in.

Pitfalls

  • Pulling qwen3 without a tag. You get the 5.2 GB 8b model. Always type the size.
  • Judging speed on the first run. The first answer includes load time. Compare speeds on the second run, or use the Speed line, which excludes loading.
  • Trusting a small model's reading of codes. It may confidently explain 01 wrongly. Give it the meanings, as the script does, and check the answer against them.
  • Treating the sample answer as the model's. --sample prints fixed text. It shows the format only.
  • Comparing models on one question. One answer says little. The later topics use evaluation sets.

Exercise

Start a model log for Unit 12. The quantization and buy-vs-build topics add rows to it.

  1. Pull the next size up:

    ollama pull qwen3:1.7b
  2. In unit12, create model_log.md with this content and save:

    # Unit 12 model log
    
    | Date | Model | Download size | Processor | Load s | Tokens/s | Answer correct? |
    | --- | --- | --- | --- | --- | --- | --- |
  3. Run the small model twice and keep the second run's numbers:

    python unit12/local_chat.py --model qwen3:0.6b
    python unit12/local_chat.py --model qwen3:0.6b
  4. Run ollama ps and note the Processor column.

  5. Add a row: today's date, qwen3:0.6b, the size from ollama list, the processor, the load time and speed from the second run, and yes or no for whether the answer matches CODE_MEANINGS (credit check failed, credit hold, no billing block).

  6. Repeat steps 3 to 5 with --model qwen3:1.7b.

  7. Under the table, write one sentence: which model you would pick for this task on your machine, and why.

  8. Commit:

    git add unit12/model_log.md
    git commit -m "Unit 12: first local model log"

Done when model_log.md has two rows with real numbers and a one-sentence choice, and python check_unit12.py ends with All set.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the mental model, what does Ollama correspond to?

    Answer: B. Ollama is the runner: it loads the model file into memory and turns prompts into answers. The model file holds the parameters, and your script is the remote control that sends requests.
  2. 2Why does the setup always pull qwen3:0.6b and never just qwen3?

    Answer: C. The registry marks the 8b model as latest, so the bare name pulls 5.2 GB instead of about 0.5 GB. Typing the tag keeps the model small enough for most laptops.
  3. 3Ollama returns eval_count 46 and eval_duration 1150000000. What is the speed?

    Answer: D. Timings are in nanoseconds, so 1,150,000,000 is 1.15 seconds. 46 tokens divided by 1.15 seconds is 40 tokens per second, which is what the script prints.
  4. 4local_chat.py prints "Ollama doesn't have qwen3:0.6b yet". What happened?

    Answer: C. The script maps a ResponseError with status 404 to that message. A server that isn't running raises ConnectionError instead, which prints a different message.
  5. 5A colleague wants to let the whole team use the Ollama on her desktop by making it listen on the office network. What would you do?

    Answer: B. Ollama listens on 127.0.0.1 by default for a reason: the local server doesn't check who is calling. Exposing it lets anyone on the network use the hardware and the models, so put a proper gateway in front first.
  6. 6Why does check_unit12.py build its web opener with an empty ProxyHandler?

    Answer: D. A proxy setting can send even local requests to a company proxy, which can't reach your laptop's port 11434. Ignoring proxies for a local call avoids that, which is also why the troubleshooting table mentions NO_PROXY.
  7. 7Which SAP route matches "run a model we chose ourselves, inside SAP's platform"?

    Answer: C. SAP's Developer Center tutorial deploys Ollama as a custom serving template on SAP AI Core, Standard or Extended plan. Hub catalogue models are chosen and run by SAP, which is a different route.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in