Orchestrate

Choosing and calling LLMs

Pick a language model for an SAP task by testing quality, cost, speed and lifecycle on your own cases, then call it so you can switch later.

Updated Oct 1, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

There is no single best large language model (LLM). There is the right model for one task, at one volume, under one budget. A model that writes a careful reply to an angry customer may be far too slow and costly for sorting ten thousand blocked sales orders a day.

Choosing well means testing a few models on your own examples and comparing four things: how often they get it right, what each answer costs, how fast it comes back, and how long the model will stay available. Then you write the choice down, with the numbers.

Calling well means keeping the model name in configuration, not buried in code. Models get new versions and are retired on set dates. A team that can swap a model by changing one setting, and re-run its tests, is protected. A team that can't is stuck.

In SAP's world, the generative AI hub gives one account access to models from several providers. Its orchestration service lets you switch models by name. That is what this topic builds on.

Why it matters to the business

The model choice drives three numbers a leader cares about.

  • Running cost. Models are paid per token, the word pieces they read and write. SAP's own illustration uses a retrieval system handling 25,000 requests a month, each with 3,500 tokens in and 300 out. That is 87.5 million input tokens and 7.5 million output tokens a month. The price per token across models can differ a lot, so the same workload can cost very different amounts.
  • Quality on the task. A cheaper model that mislabels one blocked order in six may send credit cases to the pricing team. The cost of that mistake lands in order-to-cash, not in the AI budget.
  • Continuity. SAP's documentation says model versions have deprecation dates. A process that depends on a retired version stops working on that date unless someone has planned the move.

A concrete case: an order-to-cash team wants an assistant that reads why a sales order is blocked and routes it to credit, pricing or master data. It runs on every blocked order, so volume is high and answers are one word. A small, fast model may be enough. The same team also wants a draft email to the customer for disputed invoices. Volume is low and tone matters, so a larger model may be worth its price. Two tasks, two models, one account.

How SAP does it

As of October 2026, SAP offers model choice through the generative AI hub in SAP AI Core, on the extended service plan. Set up for Unit 5 covers access, including the 30-day trial.

  • One catalog. SAP AI Core lists the models your account can use, with each model's provider, context size, cost figures, and whether it is deprecated or has a retirement date. SAP Note 3437766 holds the token conversion rates, rate limits and deprecation dates.
  • A model library screen. In SAP AI Launchpad, the Model Library has a catalog, a leaderboard of benchmark scores, a chart view and a model card per model. The card shows input types, cost information and your rate limits, with an Increase Quota request.
  • Switch by setting, not by project. SAP says the orchestration service is provider agnostic: you can switch models through a configuration value, without new deployments.
  • Version control. A call can ask for the latest version, which upgrades automatically, or pin a named version, which stays fixed until its deprecation date.
  • Fallbacks. Orchestration (version 2) accepts a list of configurations in order of preference. If the first model isn't available in the region, or fails with certain temporary errors, it tries the next.
  • Approved lists. An administrator can restrict an orchestration deployment to an allow list or a deny list of models, to enforce company standards.

One restriction is worth knowing early: SAP's documentation says SAP AI Core, including the generative AI hub, must not be used to generate synthetic data for training or fine-tuning models, unless an exception is approved and documented.

A decision guide for SAP tasks

Benchmarks and leaderboards help you build a shortlist. They don't tell you how a model does on your blocked orders. Use them to pick two or three candidates, then test.

Task Volume What matters most Where to start the test
Sort blocked sales orders into a few reasons High, every order Accuracy on a fixed label set, cost per call, speed A small, fast model, checked against a larger one
Summarize a three-way match exception for an AP clerk Medium Faithful to the documents, readable A mid-sized model
Draft a customer email about a disputed invoice Low Tone, judgment, few errors A larger model; a person reviews before sending
Explain an MRP exception list to a planner Medium, long input Context size, faithful to the data A model whose context window fits the full list

Two starting strategies are common. OpenAI's agent guide recommends starting with the most capable model to set a quality baseline, then trying smaller models to see if they still meet the target. Teams under tight cost or speed limits sometimes start small and move up only when tests fail. Both work if you measure.

Questions to ask

  • Which two or three models did you test, on how many of our real examples, and what was the score of each?
  • What does one call cost on average, and what will a month cost at our expected volume?
  • How fast is a typical answer, and is that fast enough for the screen or batch job it feeds?
  • Is the model version pinned or set to latest? Who watches deprecation dates, and what is the plan when one arrives?
  • Is there a fallback model if the first one is unavailable? Has the fallback passed the same tests?
  • Is the chosen model on our approved list for this kind of data?
  • How do we switch models later, and how long would it take to re-test?

Common misconceptions

  • "The top of the leaderboard is the right choice." Leaderboards measure general tasks. Your task may be narrow, and a smaller model may score just as well on it at a fraction of the cost.
  • "Choose once and you're done." Versions are deprecated and retired on dates SAP publishes. A model choice needs an owner and a review date.
  • "Cost depends on the number of questions." Cost depends on tokens, in and out. A long prompt with company context can cost more than a short question with a long answer, and the other way round.
  • "Switching models means rebuilding the application." Through SAP's orchestration service, the model is a configuration value. What takes time is re-testing, so keep the tests ready.
  • "A fallback model is free insurance." A fallback only helps if it has passed the same tests. Otherwise it quietly lowers quality when it kicks in.

Key terms

  • LLM (large language model): a model that reads and writes text, called over the internet and paid per use.
  • Token: a word piece; the unit models are metered and priced in.
  • Context window: the largest amount of text, in tokens, a model can take in one call.
  • Latency: how long an answer takes to come back.
  • Model version: a dated release of a model; versions are deprecated and then retired.
  • Fallback: a second model the system uses automatically when the first one fails or isn't available.
  • Allow list: the models an administrator permits; anything else is refused.
  • Model library: the screen in SAP AI Launchpad for exploring models, benchmarks and costs.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1An order-to-cash team asks which LLM is "best". What is the most useful answer?

    Answer: B. There is no single best model, only the right one for a task, volume and budget. Leaderboards help build a shortlist, but only tests on your own examples show which model is good enough for your process.
  2. 2Why does model choice matter so much for a high-volume task like sorting blocked sales orders?

    Answer: D. Models are paid per token, and per-token prices differ across models. At thousands of calls a day, a small difference per call becomes a large monthly figure, so it pays to test whether a cheaper model is good enough.
  3. 3What does SAP's orchestration service change about switching models later?

    Answer: B. SAP describes orchestration as provider agnostic: you switch models through a configuration value. Re-testing is still needed, because a new model can behave differently on your cases.
  4. 4Your assistant pins a model version. What risk must someone own?

    Answer: C. SAP's documentation says a deployment pinned to a model version stops working on that version's deprecation date. Someone needs to watch the dates and plan the move, or use latest and re-test after upgrades.
  5. 5A vendor says their solution has a fallback model "for free resilience". What should you ask?

    Answer: D. A fallback only protects the process if it is good enough at the task. If it was never tested, it can quietly lower quality whenever the primary is unavailable.
  6. 6What does an allow list on an orchestration deployment give a company?

    Answer: A. SAP lets administrators restrict an orchestration deployment to an allow list or deny list of models. It is a control for company standards, for example models approved for certain data.
  7. 7Which use of the generative AI hub does SAP's documentation forbid unless approved and documented?

    Answer: C. SAP's documentation says SAP AI Core, including the generative AI hub, must not be used to generate synthetic data for training or fine-tuning, unless an exception is explicitly approved and documented. The other uses are normal parts of choosing and calling models.
Deep layer · 40 min read

Mental model: a model call is a priced, versioned request

Every LLM call is the same kind of thing: a web request that names a model and version, carries messages and a few parameters, and comes back with text plus a usage count of tokens in and out. Choosing a model means trading four measurable things against each other on your own cases: quality, cost, latency and lifecycle.

flowchart LR
  C[Catalog<br/>what you may call] --> S[Shortlist<br/>2-3 models]
  S --> T[Test on your cases<br/>score, time, tokens]
  T --> D[Decision record<br/>model, version, fallback]
  D --> P[Config value<br/>in your app]
  P -->|new version or retirement| T

The loop at the end is the point. The choice is not made once. It lives in configuration, and your test cases let you re-check it in minutes when a version changes.

How it works

Anatomy of one call

Set up for Unit 5 sent one prompt through SAP's orchestration service. Here is what that call contains, in the version 2 shape SAP documents:

Part Example Why it matters for choice
Model name "name": "gpt-4o" The thing you are choosing; a configuration value
Model version "version": "latest" latest upgrades by itself; a named version is fixed until it is deprecated
Parameters "params": {"max_tokens": 300} Caps and sampling settings; SAP says the possible values depend on the model
Timeout and retries timeout, max_retries How long to wait and how often to retry before giving up
Messages system and user messages Your prompt; its length is input tokens
Response: content the model's text What you score
Response: usage prompt_tokens, completion_tokens What you pay for
Response: model the model that answered Matters when a fallback was used

What drives cost

SAP meters generative AI use in tokens. They are converted into GenAI tokens and then into BTP capacity units, at rates that differ by model (SAP Note 3437766 lists them). SAP's pricing page says output tokens tend to cost slightly more than input tokens.

So cost per call is roughly:

cost per call = input tokens x input price + output tokens x output price

Two levers follow. The shape of the task sets the token ratio: SAP's page notes that summarization can run 3 or 4 input tokens per output token, while drafting text can produce more output than input. And the model sets the price per token. A one-word classifier with a short prompt is cheap on any model; a long system prompt sent on every call is not.

What drives latency

Answers are generated one token at a time, as How LLMs generate text showed. Long answers take longer than short ones, whichever model you use. Beyond that, measure: latency differs by model, by region and by load. Record the median, not one lucky call.

Context window

The catalog lists each version's contextLength. If a task needs a long MRP exception list or a full contract in one call, the model must accept that many tokens plus room for the answer. Filter on it first; it is a hard limit, not a trade-off.

The model catalog

SAP AI Core answers GET {AI_API_URL}/v2/lm/scenarios/foundation-models/models with every model your account can use. The fields that matter for choice:

Field What it tells you
model The name you put in your call
provider Who built it
allowedScenarios Whether you can call it through orchestration, through a foundation-models deployment, or both
versions[].capabilities For example text-generation or image-recognition
versions[].contextLength Maximum context, in tokens
versions[].cost A list with inputCost and outputCost values
versions[].deprecated, retirementDate Lifecycle warnings
versions[].isLatest Whether this is the version latest points to
versions[].metadata Benchmark figures where SAP provides them

Versions and lifecycle

SAP's Model Lifecycle page gives two options. Auto upgrade: use latest, and your calls move to new versions as SAP adds them. Manual upgrade: name a version, and it stays until you change it. If you pin a version, calls stop working on its deprecation date. latest is the default when you name no version.

Neither is free. latest can change behaviour under you; a pinned version needs someone watching dates. Either way, keep your test cases and re-run them when the version changes.

Fallbacks, timeouts and retries

In orchestration version 2, modules can be a list of configurations in order of preference. SAP documents when orchestration moves to the next one:

  • for every request, if the model isn't supported in the deployed region or environment;
  • for non-streaming requests, also on 408 Request Timeout, 429 Too Many Requests and any 5xx server error.

Other errors fail the request. A successful response lists skipped attempts under intermediate_failures. The final_result.model field tells you which model answered.

Per model, SAP documents a timeout from 1 to 600 seconds (default 600) and max_retries from 0 to 5 (default 2). For a user waiting on screen, 600 seconds is far too long; set a timeout that matches the screen or job.

sequenceDiagram
  participant App as Your script
  participant O as Orchestration
  participant A as Model A
  participant B as Model B
  App->>O: config with modules [A, B]
  O->>A: call
  A-->>O: 429 or not in region
  O->>B: same prompt
  B-->>O: answer
  O-->>App: final_result (model B) + intermediate_failures

Build it yourself: shortlist, compare and add a fallback

You will build one script, choose_model.py, with three commands. catalog reads SAP's model list and filters it. compare sends six made-up blocked sales orders to each model you name, scores the one-word answers, and records time, tokens and a cost index. fallback calls a primary model with a backup behind it and shows which one answered.

flowchart LR
  E[.env<br/>AICORE_ lines] --> CAT[catalog<br/>filter models]
  CAT --> CMP[compare<br/>6 test cases per model]
  CMP --> CSV[model_comparison.csv]
  CMP --> FB[fallback<br/>primary + backup]

Before you start: complete Set up your computer for this course and Set up for Unit 5. They create your orchestrate-course folder, its .venv, the AICORE_ lines in .env, and install sap-ai-sdk-gen. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder from Unit 5 setup, with sap-ai-sdk-gen and python-dotenv installed.
  • About 45 minutes.
  • For real calls: SAP AI Core access with the generative AI hub (the trial or a company account). A comparison of two models is 12 short calls, a small per-request charge on a paid account and no charge during the trial.
  • No account? Every command has a --sample option with made-up data.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Confirm the SDK is there (same on every system):

    python -c "import gen_ai_hub.orchestration_v2; print('ready')"

    You should see ready. No new libraries are needed for this topic.

Step 2: Save the script

  1. In VS Code's file list, right-click unit05, choose New File and name it choose_model.py.
  2. Paste the code below and save.
"""Unit 5: shortlist models from SAP's catalog, then compare them on your own SAP-shaped test cases.

Three commands (run from your course folder, with .venv turned on):
    python unit05/choose_model.py catalog --sample                 # no account: a made-up catalog
    python unit05/choose_model.py catalog --min-context 100000     # real catalog from SAP AI Core
    python unit05/choose_model.py compare --sample                 # no account: made-up answers
    python unit05/choose_model.py compare --models MODEL_A,MODEL_B # real calls through orchestration
    python unit05/choose_model.py fallback --sample                # no account: see a fallback happen
    python unit05/choose_model.py fallback --primary MODEL_A --backup MODEL_B

The real paths read the AICORE_ lines in .env (see "Set up for Unit 5").
"compare" writes unit05/model_comparison.csv, which you reuse in the evaluation unit.
"""
import argparse
import base64
import csv
import json
import os
import statistics
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
from pathlib import Path

HERE = Path(__file__).resolve().parent
LABELS = ["CREDIT", "PRICING", "INCOMPLETE", "EXPORT"]

# Six blocked sales orders with the label a person gave each one. All made up.
CASES = [
    ("Order 4711: customer 10023 has open items of 52,000 EUR against a credit limit of 50,000 EUR.", "CREDIT"),
    ("Order 4712: no price found for material M-200 in sales org 1010; net value is zero.", "PRICING"),
    ("Order 4713: payment terms and Incoterms are missing in the header.", "INCOMPLETE"),
    ("Order 4714: ship-to party is in a country on the embargo list; trade compliance check pending.", "EXPORT"),
    ("Order 4715: credit check failed after the customer's rating dropped to high risk.", "CREDIT"),
    ("Order 4716: manual price change of 40 percent exceeds the allowed tolerance.", "PRICING"),
]
SYSTEM = ("You classify why an SAP sales order is blocked. Answer with exactly one word from this list: "
          + ", ".join(LABELS) + ". No other text.")

# A made-up catalog in the shape SAP documents for GET /v2/lm/scenarios/foundation-models/models.
SAMPLE_CATALOG = {"count": 4, "resources": [
    {"model": "sample--large", "provider": "SampleAI", "executableId": "sample",
     "allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
     "versions": [{"name": "2026-06-01", "isLatest": True, "deprecated": False, "retirementDate": "",
                   "contextLength": 400000, "capabilities": ["text-generation"],
                   "cost": [{"inputCost": "0.0050"}, {"outputCost": "0.0250"}]}]},
    {"model": "sample--small", "provider": "SampleAI", "executableId": "sample",
     "allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
     "versions": [{"name": "2026-06-01", "isLatest": True, "deprecated": False, "retirementDate": "",
                   "contextLength": 128000, "capabilities": ["text-generation"],
                   "cost": [{"inputCost": "0.0004"}, {"outputCost": "0.0016"}]}]},
    {"model": "sample--old", "provider": "SampleAI", "executableId": "sample",
     "allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
     "versions": [{"name": "2024-05-13", "isLatest": True, "deprecated": True, "retirementDate": "2026-11-30",
                   "contextLength": 128000, "capabilities": ["text-generation"],
                   "cost": [{"inputCost": "0.0030"}, {"outputCost": "0.0090"}]}]},
    {"model": "sample--embedder", "provider": "SampleAI", "executableId": "sample",
     "allowedScenarios": [{"scenarioId": "foundation-models", "executableId": "sample"}],
     "versions": [{"name": "1", "isLatest": True, "deprecated": False, "retirementDate": "",
                   "contextLength": 8000, "capabilities": ["embeddings"],
                   "cost": [{"inputCost": "0.0001"}]}]},
]}
# Made-up answers for --sample: the small model gets one case wrong.
SAMPLE_ANSWERS = {
    "sample--large": (["CREDIT", "PRICING", "INCOMPLETE", "EXPORT", "CREDIT", "PRICING"], 1.9),
    "sample--small": (["CREDIT", "PRICING", "INCOMPLETE", "EXPORT", "CREDIT", "INCOMPLETE"], 0.6),
}


# ---------- the catalog ----------

def env_or_exit() -> dict:
    """Read the five AICORE_ settings from .env, or stop with a clear message."""
    from dotenv import load_dotenv
    load_dotenv()
    names = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
             "AICORE_RESOURCE_GROUP"]
    missing = [n for n in names if not os.environ.get(n)]
    if missing:
        sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5. "
                 "Or add --sample to try without an account.")
    return {n: os.environ[n] for n in names}


def fetch_catalog(env: dict) -> dict:
    """Get a token with the service key details, then ask SAP AI Core for its model list."""
    basic = base64.b64encode(f"{env['AICORE_CLIENT_ID']}:{env['AICORE_CLIENT_SECRET']}".encode()).decode()
    body = urllib.parse.urlencode({"grant_type": "client_credentials"}).encode()
    request = urllib.request.Request(env["AICORE_AUTH_URL"], data=body, headers={
        "Authorization": f"Basic {basic}", "Content-Type": "application/x-www-form-urlencoded"})
    try:
        with urllib.request.urlopen(request, timeout=30) as reply:
            token = json.loads(reply.read())["access_token"]
        url = env["AICORE_BASE_URL"].rstrip("/") + "/lm/scenarios/foundation-models/models"
        request = urllib.request.Request(url, headers={
            "Authorization": f"Bearer {token}", "AI-Resource-Group": env["AICORE_RESOURCE_GROUP"]})
        with urllib.request.urlopen(request, timeout=30) as reply:
            return json.loads(reply.read())
    except urllib.error.HTTPError as error:
        sys.exit(f"SAP AI Core answered HTTP {error.code} {error.reason}. Run check_unit05.py to find the cause.")
    except (urllib.error.URLError, KeyError) as error:
        sys.exit(f"Could not reach SAP AI Core ({error}). Check your network and the AICORE_ lines in .env.")


def cost_of(version: dict, key: str) -> float:
    """The catalog lists cost as [{"inputCost": "..."}, {"outputCost": "..."}]."""
    for item in version.get("cost", []):
        if key in item:
            try:
                return float(item[key])
            except (TypeError, ValueError):
                return 0.0
    return 0.0


def shortlist(catalog: dict, min_context: int) -> list:
    """Keep chat models you can call through orchestration, with enough context, newest version only."""
    rows = []
    for model in catalog.get("resources", []):
        if not any(s.get("scenarioId") == "orchestration" for s in model.get("allowedScenarios", [])):
            continue
        versions = model.get("versions", [])
        version = next((v for v in versions if v.get("isLatest")), versions[0] if versions else None)
        if not version or "text-generation" not in version.get("capabilities", []):
            continue
        if (version.get("contextLength") or 0) < min_context:
            continue
        rows.append({"model": model["model"], "provider": model.get("provider", ""),
                     "context": version.get("contextLength") or 0,
                     "in_cost": cost_of(version, "inputCost"), "out_cost": cost_of(version, "outputCost"),
                     "deprecated": bool(version.get("deprecated")),
                     "retires": version.get("retirementDate") or ""})
    return sorted(rows, key=lambda r: (r["deprecated"], r["in_cost"] + r["out_cost"]))


def cmd_catalog(args) -> None:
    catalog = SAMPLE_CATALOG if args.sample else fetch_catalog(env_or_exit())
    rows = shortlist(catalog, args.min_context)
    print(f"{catalog.get('count', len(catalog.get('resources', [])))} models in the catalog; "
          f"{len(rows)} of them are chat models usable through orchestration with at least "
          f"{args.min_context:,} tokens of context.\n")
    print(f"{'model':<34}{'provider':<14}{'context':>10}{'in cost':>10}{'out cost':>10}  status")
    for r in rows:
        status = "ok"
        if r["deprecated"] or r["retires"]:
            status = "DEPRECATED" if r["deprecated"] else "retiring"
            status += f", retires {r['retires']}" if r["retires"] else ""
        print(f"{r['model']:<34}{r['provider'][:13]:<14}{r['context']:>10,}"
              f"{r['in_cost']:>10.4f}{r['out_cost']:>10.4f}  {status}")
    if not rows:
        print("(none) Lower --min-context, or ask your administrator which models your account offers.")


# ---------- calling models ----------

def make_config(model: str, max_out: int = 0, backup: str = ""):
    """Build an orchestration v2 configuration: one template, one model (plus an optional fallback)."""
    from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
                                             PromptTemplatingModuleConfig, SystemMessage, Template,
                                             UserMessage)

    def module(name: str):
        params = {}
        if max_out:   # Anthropic models take max_tokens; the newer OpenAI models take max_completion_tokens
            params["max_tokens" if name.startswith("anthropic--") else "max_completion_tokens"] = max_out
        template = Template(template=[SystemMessage(content=SYSTEM), UserMessage(content="{{?order}}")])
        return ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
            prompt=template, model=LLMModelDetails(name=name, params=params or None, timeout=60, max_retries=1)))

    return OrchestrationConfig(modules=[module(model), module(backup)] if backup else module(model))


def call(service, text: str):
    """One call. Returns (answer, input tokens, output tokens, seconds, model that answered, failures)."""
    start = time.perf_counter()
    result = service.run(placeholder_values={"order": text})
    seconds = time.perf_counter() - start
    final = result.final_result
    failures = [f"{f.code} {f.message}" for f in (result.intermediate_failures or [])]
    return (final.choices[0].message.content or "", final.usage.prompt_tokens, final.usage.completion_tokens,
            seconds, final.model, failures)


def clean(answer: str) -> str:
    """Models sometimes add a full stop or spaces. Keep the first word, in capitals."""
    words = answer.strip().split()
    return words[0].strip(".,:;!*\"'").upper() if words else ""


def cmd_compare(args) -> None:
    models = list(SAMPLE_ANSWERS) if args.sample else [m.strip() for m in args.models.split(",") if m.strip()]
    if not models:
        sys.exit("Name the models to compare, for example --models MODEL_A,MODEL_B (catalog lists them).")
    catalog = SAMPLE_CATALOG if args.sample else fetch_catalog(env_or_exit())
    prices = {r["model"]: r for r in shortlist(catalog, 0)}

    rows, summary = [], []
    for model in models:
        print(f"\n{model}")
        correct, seconds, tokens_in, tokens_out = 0, [], 0, 0
        service = None
        if not args.sample:
            from gen_ai_hub.orchestration_v2 import OrchestrationService
            service = OrchestrationService(config=make_config(model, args.max_out))
        try:
            for i, (text, expected) in enumerate(CASES):
                if args.sample:
                    answers, delay = SAMPLE_ANSWERS[model]
                    answer = answers[i] if i < len(answers) else expected   # your own added cases: made-up correct
                    t_in, t_out, took = 70, 2, delay + 0.1 * (i % 3)
                else:
                    try:
                        answer, t_in, t_out, took, _, _ = call(service, text)
                    except Exception as error:   # keep going: one failure shouldn't hide the rest
                        print(f"  case {i + 1}: call failed: {type(error).__name__}: {str(error)[:200]}")
                        continue
                got = clean(answer)
                ok = got == expected
                correct += ok
                seconds.append(took)
                tokens_in += t_in
                tokens_out += t_out
                print(f"  case {i + 1}: expected {expected:<10} got {got:<10} {'ok' if ok else 'WRONG'}"
                      f"  {took:.1f}s")
                rows.append({"model": model, "case": i + 1, "expected": expected, "got": got,
                             "correct": ok, "seconds": round(took, 2), "tokens_in": t_in, "tokens_out": t_out})
        finally:
            if service is not None:
                service.close_http_connection()
        price = prices.get(model, {})
        cost = (tokens_in * price.get("in_cost", 0) + tokens_out * price.get("out_cost", 0)) / 1000
        summary.append((model, correct, statistics.median(seconds) if seconds else 0.0,
                        tokens_in, tokens_out, cost, bool(price)))

    print(f"\n{'model':<34}{'correct':>9}{'median s':>10}{'tokens in':>11}{'tokens out':>11}{'cost index':>12}")
    for model, correct, median, t_in, t_out, cost, priced in summary:
        print(f"{model:<34}{f'{correct}/{len(CASES)}':>9}{median:>10.1f}{t_in:>11}{t_out:>11}"
              f"{(f'{cost:.4f}' if priced else 'n/a'):>12}")
    out = HERE / "model_comparison.csv"
    with open(out, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=list(rows[0]) if rows else ["model"])
        writer.writeheader()
        writer.writerows(rows)
    print(f"\nSaved {len(rows)} rows to {out}")
    if args.sample:
        print("[sample] Made-up answers and timings; no model was called.")


def cmd_fallback(args) -> None:
    text, expected = CASES[0]
    if args.sample:
        print(f"[sample] Primary {args.primary} is 'not supported' in this made-up region; "
              f"orchestration used {args.backup}.")
        print(f"Answer: {expected}   answered by: {args.backup}")
        print(f"Skipped: 400 Model {args.primary} not supported.")
        return
    env_or_exit()
    from gen_ai_hub.orchestration_v2 import OrchestrationService
    service = OrchestrationService(config=make_config(args.primary, args.max_out, backup=args.backup))
    try:
        answer, _, _, took, answered_by, failures = call(service, text)
    except Exception as error:
        sys.exit(f"Both models failed: {type(error).__name__}: {str(error)[:400]}")
    finally:
        service.close_http_connection()
    print(f"Answer: {clean(answer)}   answered by: {answered_by}   ({took:.1f}s)")
    print("Skipped: " + ("; ".join(failures) if failures else "nothing, the primary model answered"))


def main() -> None:
    parser = argparse.ArgumentParser(description="Shortlist and compare LLMs in SAP's generative AI hub.")
    sub = parser.add_subparsers(dest="command", required=True)
    p = sub.add_parser("catalog", help="list chat models usable through orchestration")
    p.add_argument("--min-context", type=int, default=0, help="smallest context window you need, in tokens")
    p.add_argument("--sample", action="store_true", help="use a made-up catalog (no account)")
    p = sub.add_parser("compare", help="run the six test cases on each model")
    p.add_argument("--models", default="", help="comma-separated model names from the catalog")
    p.add_argument("--max-out", type=int, default=0, help="optional cap on output tokens per answer")
    p.add_argument("--sample", action="store_true", help="made-up answers (no account)")
    p = sub.add_parser("fallback", help="call a primary model with a backup model behind it")
    p.add_argument("--primary", default="sample--not-in-region")
    p.add_argument("--backup", default="sample--small")
    p.add_argument("--max-out", type=int, default=0, help="optional cap on output tokens per answer")
    p.add_argument("--sample", action="store_true", help="show a made-up fallback (no account)")
    args = parser.parse_args()

    {"catalog": cmd_catalog, "compare": cmd_compare, "fallback": cmd_fallback}[args.command](args)


if __name__ == "__main__":
    main()

Step 3: Read the catalog

  1. No account:

    python unit05/choose_model.py catalog --sample
  2. With your key: list chat models you can call through orchestration. Add --min-context with the context size your task needs, in tokens:

    python unit05/choose_model.py catalog --min-context 100000

What success looks like (with --sample; the names and numbers are made up):

4 models in the catalog; 3 of them are chat models usable through orchestration with at least 0 tokens of context.

model                             provider         context   in cost  out cost  status
sample--small                     SampleAI         128,000    0.0004    0.0016  ok
sample--large                     SampleAI         400,000    0.0050    0.0250  ok
sample--old                       SampleAI         128,000    0.0030    0.0090  DEPRECATED, retires 2026-11-30

The list is sorted with current models first, cheapest first. The embedding model in the sample catalog is left out: it can't be called through orchestration and doesn't generate text. With your key, the names, providers, costs and dates are your account's real ones and will differ from this.

If the table says (none), nothing matched your filter. Lower --min-context, or run python check_unit05.py to confirm your access.

  1. Pick two models from your list for the next step: a larger one and a smaller, cheaper one. Avoid any marked DEPRECATED or retiring.

Step 4: Compare models on your test cases

  1. No account:

    python unit05/choose_model.py compare --sample
  2. With your key: replace the two names with the ones you picked. No spaces around the comma.

    python unit05/choose_model.py compare --models MODEL_A,MODEL_B

What success looks like (with --sample):

sample--large
  case 1: expected CREDIT     got CREDIT     ok  1.9s
  case 2: expected PRICING    got PRICING    ok  2.0s
  case 3: expected INCOMPLETE got INCOMPLETE ok  2.1s
  case 4: expected EXPORT     got EXPORT     ok  1.9s
  case 5: expected CREDIT     got CREDIT     ok  2.0s
  case 6: expected PRICING    got PRICING    ok  2.1s

sample--small
  case 1: expected CREDIT     got CREDIT     ok  0.6s
  case 2: expected PRICING    got PRICING    ok  0.7s
  case 3: expected INCOMPLETE got INCOMPLETE ok  0.8s
  case 4: expected EXPORT     got EXPORT     ok  0.6s
  case 5: expected CREDIT     got CREDIT     ok  0.7s
  case 6: expected PRICING    got INCOMPLETE WRONG  0.8s

model                               correct  median s  tokens in tokens out  cost index
sample--large                           6/6       2.0        420         12      0.0024
sample--small                           5/6       0.7        420         12      0.0002

Saved 12 rows to /Users/you/orchestrate-course/unit05/model_comparison.csv
[sample] Made-up answers and timings; no model was called.

Read the summary as a decision. In the sample, the small model is about three times faster and about a tenth of the cost index, but it sent a pricing case to the wrong team. Is one in six wrong acceptable? Not for routing work automatically. It might be for a suggestion a person confirms. That judgment is the point of the exercise.

With real models, expect different timings on every run, and possibly different answers; that is normal. Six cases are enough to learn the method, not to make a production decision. The Exercise grows the set.

  1. Optional: cap the answer length with --max-out, for example --max-out 20. If answers come back empty with a cap, remove it: some models spend part of the allowance on internal reasoning before the visible answer, which SAP's parameter list says max_completion_tokens includes.

Step 5: Add a fallback

  1. No account:

    python unit05/choose_model.py fallback --sample
  2. With your key: put a name that is not in your catalog as the primary, so you can watch orchestration skip it. Use your smaller model as the backup:

    python unit05/choose_model.py fallback --primary not-a-real-model --backup MODEL_B
  3. Then try two real models. The primary should answer, and nothing is skipped:

    python unit05/choose_model.py fallback --primary MODEL_A --backup MODEL_B

What success looks like (with --sample):

[sample] Primary sample--not-in-region is 'not supported' in this made-up region; orchestration used sample--small.
Answer: CREDIT   answered by: sample--small
Skipped: 400 Model sample--not-in-region not supported.

With your key and a made-up primary, you should see your backup model under answered by and an error about the primary under Skipped. The exact wording of the error comes from SAP and may differ. With two real models, Skipped reads nothing, the primary model answered.

Step 6: Save your work in Git

  1. Check what Git sees:

    git status

    You should see unit05/choose_model.py and unit05/model_comparison.csv. You must not see .env.

  2. Save:

    git add unit05/choose_model.py unit05/model_comparison.csv
    git commit -m "Unit 5: shortlist and compare models"

What each part of the script does

Part What it does
CASES Six made-up blocked sales orders, each with the label a person would give
SYSTEM The instruction: answer with exactly one of four labels
SAMPLE_CATALOG, SAMPLE_ANSWERS Made-up data in SAP's documented catalog shape, for --sample
env_or_exit Loads .env with load_dotenv() and checks the five AICORE_ settings
fetch_catalog Swaps the client ID and secret for a token, then calls the model discovery endpoint
shortlist Keeps chat models allowed in orchestration, reads the latest version's context, cost and lifecycle, sorts current and cheap first
make_config Builds an orchestration v2 configuration: template, model, optional output cap, a 60-second timeout and one retry; with a backup, modules becomes a list
call Runs one request and returns the answer, token counts, seconds, the model that answered and any skipped attempts
clean Keeps the first word in capitals, so credit. still counts as CREDIT
cmd_compare Runs every case on every model, prints a summary and writes model_comparison.csv
cost index Tokens divided by 1,000, times the catalog's cost values; for comparing models, not a bill

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on your path, or the terminal opened before you installed it Close and reopen VS Code; see Set up your computer
ModuleNotFoundError: No module named 'gen_ai_hub' or 'dotenv' The virtual environment is off, or the libraries are missing Turn on .venv (Step 1), then pip install -r requirements.txt
Missing in .env: AICORE_... The service key details aren't in .env Run unit05/key_to_env.py from Set up for Unit 5, Step 5, or use --sample
SAP AI Core answered HTTP 401 Wrong or old client ID or secret Create a new service key and run key_to_env.py again
SAP AI Core answered HTTP 404 AICORE_BASE_URL doesn't end in /v2, or points at the wrong host Re-run key_to_env.py; it adds /v2
Could not reach SAP AI Core Network, proxy or firewall blocks the call Try another network; ask IT whether BTP and SAP AI Core addresses are allowed
A case shows call failed: ... not supported That model isn't available in your region or account Pick another from catalog
A case shows call failed: ... 429 Rate limit reached Wait a minute and run again; see the model card's rate limits in SAP AI Launchpad
got is empty Output cap too small for that model Run without --max-out
Every model gets the same case wrong The case or its label may be ambiguous Read the case again; fix the label or the wording before blaming the models

The SAP way

As of October 2026, here is how SAP's stack implements each idea in this topic.

Two ways to reach a model

SAP AI Core offers models under two scenarios:

orchestration foundation-models
How you call One orchestration deployment; the model is a name in the request A deployment per model, created with its executableId
Switch model Change a configuration value Create or patch a deployment
Extras Templating, filtering, masking, grounding, fallbacks The model's own API, for example chat completions for Azure OpenAI models
When to use Default: provider agnostic, simpler to test and compare When you need something only the model's own API offers

SAP's documentation names a limit too: orchestration supports LangChain, but not every open-source framework. This course uses orchestration through the SAP Cloud SDK for AI.

Discovering and comparing models

  • API: GET /v2/lm/scenarios/foundation-models/models, as in the script.
  • SAP AI Launchpad, Model Library: under Generative AI Hub > Model Library. Catalog lists models with filters and search. Leaderboard ranks them by benchmark. Chart plots two measures against each other. Each model card shows input types, cost information, metrics where available, deprecation notices and your rate limits. SAP lists the roles that can open it: genai_experimenter, genai_manager, genai_administrator or orchestration_executor.
  • SAP Note 3437766: token conversion rates, rate limits and deprecation dates. It needs an SAP login.

Enforcing an approved list

An administrator can restrict an orchestration deployment with two parameter bindings, modelFilterList and modelFilterListType (allow or deny). This is a sketch of the configuration body, based on SAP's sample; the model names are placeholders:

{
  "name": "orchestration-approved-models",
  "executableId": "orchestration",
  "scenarioId": "orchestration",
  "versionId": "0.0.1",
  "parameterBindings": [
    {"key": "modelFilterList",
     "value": "[{\"modelName\": \"APPROVED_MODEL_1\"}, {\"modelName\": \"APPROVED_MODEL_2\"}]"},
    {"key": "modelFilterListType", "value": "allow"}
  ]
}

Model parameters across providers

SAP's harmonized API accepts OpenAI-style parameters such as max_tokens, max_completion_tokens, temperature (0.0 to 2.0), top_p and stop. SAP notes that possible values depend on the model. One rule is documented explicitly: Anthropic models require max_tokens, and orchestration sets it to the model's maximum if you leave it out. The script sends no parameters by default for this reason, and picks the cap name by provider only when you ask for one.

Licensing and cost

The generative AI hub is available only in the extended plan of SAP AI Core; Set up for Unit 5 covers the trial route. Usage is metered in tokens and billed in capacity units, at rates that differ by model.

Build vs. SAP

Situation Call the provider directly SAP generative AI hub, orchestration SAP generative AI hub, foundation-models deployment
Quick personal prototype, no SAP data Simple, if you have an account Works; needs SAP access More setup than needed
Several providers under one contract One contract per provider Yes Yes
Compare and switch models often Code changes per provider API Change a name in configuration New deployment per model
Company-approved model list Your own controls Allow or deny list on the deployment Control which deployments exist
Fallback to another provider You build it Built in (v2), as a list You build it
Need a provider-only feature first Yes Only once orchestration supports it Native API shape
Process runs on SAP data, side by side on BTP Extra contract and data review Fits the BTP security model Fits the BTP security model

For most SAP processes in this course, orchestration is the default. Go direct, or to a foundation-models deployment, when you need something orchestration doesn't expose yet.

Production concerns

  • Security and SAP authorizations. The service key is a secret; it stays in .env or a secret store, never in code or Git. Restrict models with an allow list. Remember that the model sees whatever you put in the prompt: a blocked-order text may hold customer names and amounts. The orchestration topic later in this unit covers data masking and filtering; Unit 11 covers AI security.
  • Evaluation. Keep your test cases in version control with their expected answers. Re-run them when you change model, version, prompt or fallback. Unit 8 turns this into a full evaluation harness.
  • Cost. Estimate monthly cost from tokens per call times volume, using the rates in SAP Note 3437766. Watch long system prompts sent on every call, and long answers you don't need.
  • Operations. Set a timeout that fits the user's wait and a small max_retries. Log the model that actually answered and any intermediate_failures, so a silent fallback shows up in monitoring. Check rate limits on the model card, and request more with Increase Quota before go-live.
  • Lifecycle. Decide per use case between latest and a pinned version, and write down who checks deprecation dates. Put the next review date in the decision record.
  • Clean core. The model call runs side by side, outside the ERP. When a later version reads real blocked orders, read them through released APIs, as in Calling your first SAP API, not by changing S/4HANA.

Pitfalls

  • Choosing from a leaderboard alone. General benchmarks don't measure your labels. Test on your cases.
  • Testing on three examples. A model that gets three easy cases right tells you little. Include hard and ambiguous ones.
  • Comparing models with different prompts. Change one thing at a time, or you can't tell what made the difference.
  • One timing per model. Latency varies call to call. Use the median of several.
  • Hard-coding the model name in many files. Keep it in one configuration value.
  • An untested fallback. If the backup was never scored, a fallback quietly lowers quality.
  • Ignoring intermediate_failures. If you don't log which model answered, you won't notice the primary is failing every day.
  • Forgetting retirement dates. A pinned version stops working on its date. Track it like a certificate expiry.
  • Reading the cost index as a price. It only ranks models. Use SAP Note 3437766 and your contract for money.

Exercise: write a model decision record

You will grow the test set, re-run the comparison and write a one-page decision. Unit 8 reuses your test cases and your CSV as the start of an evaluation harness.

  1. Open unit05/choose_model.py and find CASES.

  2. Add four more blocked orders below the existing six, each on its own line in the same format. Make at least two of them hard, for example an order blocked for both credit and missing data. Give each the label you think is right.

  3. Save the file.

  4. Run the comparison on your two models. With no account, use compare --sample; your new cases then get made-up correct answers, so the point is the method, not the numbers.

    python unit05/choose_model.py compare --models MODEL_A,MODEL_B
  5. Run it a second time. Note whether any answers changed between runs.

  6. In unit05, create model_choice.md and fill in these headings:

    • Task: one sentence.
    • Volume: calls per day you expect.
    • Models tested: names and versions.
    • Results: correct, median seconds and cost index for each, from both runs.
    • Choice: the model, and latest or a pinned version, with one sentence why.
    • Fallback: which model, and whether it passed the same cases.
    • Review date: when to re-run, and who owns deprecation dates.
  7. Save your work:

    git add unit05/choose_model.py unit05/model_comparison.csv unit05/model_choice.md
    git commit -m "Unit 5: model decision record"

Done when: model_comparison.csv has rows for 10 cases per model, and model_choice.md names a model, a version policy, a fallback and a review date, each backed by numbers from your runs.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does the response's usage block give you when choosing a model?

    Answer: C. usage reports prompt_tokens and completion_tokens. Cost is metered in tokens, so these counts, times each model's rates, are how you compare what calls cost.
  2. 2In shortlist, why does the script check allowedScenarios for orchestration?

    Answer: A. The catalog lists every model the account can use, including some only reachable through their own foundation-models deployment. The script calls models through orchestration, so it keeps only models allowed in that scenario.
  3. 3What happens to a call that pins a model version when that version reaches its deprecation date?

    Answer: B. SAP's Model Lifecycle page says a deployment that names a version stops working on that version's deprecation date. Only latest upgrades automatically, and that can change behaviour, so re-run your tests either way.
  4. 4In make_config, what turns a single model into a fallback setup?

    Answer: D. Orchestration version 2 accepts a list of module configurations and tries them in order. Retries repeat the same model; they don't switch to another one.
  5. 5Your primary model starts returning 429 Too Many Requests on a non-streaming call with a fallback set. What does SAP document will happen?

    Answer: C. SAP lists 408, 429 and 5xx errors as reasons to move to the next configuration for non-streaming requests. The skipped attempt shows up in intermediate_failures, which is why you should log it.
  6. 6The small model is a tenth of the cost index but got one of six cases wrong. What would you do next?

    Answer: B. Six cases are too few for a production decision. Grow the test set and compare against a target the business agreed, such as acceptable for suggestions but not for automatic routing.
  7. 7Why does the script send no output cap unless you pass --max-out?

    Answer: D. SAP says possible parameter values depend on the model, and Anthropic models need max_tokens, which orchestration fills in if missing. Sending nothing by default is safe; the option picks the cap name by provider only when asked.
  8. 8A colleague wants to read the cost index as next month's AI bill. What should you tell them?

    Answer: D. The catalog's cost fields have no stated unit on SAP's documentation page, so the index only compares models with each other. Real rates, converted to capacity units, are in SAP Note 3437766 and the company's agreement.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in