There is no single best large language model (LLM). There is the right model for one task, at one volume, under one budget. A model that writes a careful reply to an angry customer may be far too slow and costly for sorting ten thousand blocked sales orders a day.
Choosing well means testing a few models on your own examples and comparing four things: how often they get it right, what each answer costs, how fast it comes back, and how long the model will stay available. Then you write the choice down, with the numbers.
Calling well means keeping the model name in configuration, not buried in code. Models get new versions and are retired on set dates. A team that can swap a model by changing one setting, and re-run its tests, is protected. A team that can't is stuck.
In SAP's world, the generative AI hub gives one account access to models from several providers. Its orchestration service lets you switch models by name. That is what this topic builds on.
The model choice drives three numbers a leader cares about.
Running cost. Models are paid per token, the word pieces they read and write. SAP's own illustration uses a retrieval system handling 25,000 requests a month, each with 3,500 tokens in and 300 out. That is 87.5 million input tokens and 7.5 million output tokens a month. The price per token across models can differ a lot, so the same workload can cost very different amounts.
Quality on the task. A cheaper model that mislabels one blocked order in six may send credit cases to the pricing team. The cost of that mistake lands in order-to-cash, not in the AI budget.
Continuity. SAP's documentation says model versions have deprecation dates. A process that depends on a retired version stops working on that date unless someone has planned the move.
A concrete case: an order-to-cash team wants an assistant that reads why a sales order is blocked and routes it to credit, pricing or master data. It runs on every blocked order, so volume is high and answers are one word. A small, fast model may be enough. The same team also wants a draft email to the customer for disputed invoices. Volume is low and tone matters, so a larger model may be worth its price. Two tasks, two models, one account.
As of October 2026, SAP offers model choice through the generative AI hub in SAP AI Core, on the extended service plan. Set up for Unit 5 covers access, including the 30-day trial.
One catalog. SAP AI Core lists the models your account can use, with each model's provider, context size, cost figures, and whether it is deprecated or has a retirement date. SAP Note 3437766 holds the token conversion rates, rate limits and deprecation dates.
A model library screen. In SAP AI Launchpad, the Model Library has a catalog, a leaderboard of benchmark scores, a chart view and a model card per model. The card shows input types, cost information and your rate limits, with an Increase Quota request.
Switch by setting, not by project. SAP says the orchestration service is provider agnostic: you can switch models through a configuration value, without new deployments.
Version control. A call can ask for the latest version, which upgrades automatically, or pin a named version, which stays fixed until its deprecation date.
Fallbacks. Orchestration (version 2) accepts a list of configurations in order of preference. If the first model isn't available in the region, or fails with certain temporary errors, it tries the next.
Approved lists. An administrator can restrict an orchestration deployment to an allow list or a deny list of models, to enforce company standards.
One restriction is worth knowing early: SAP's documentation says SAP AI Core, including the generative AI hub, must not be used to generate synthetic data for training or fine-tuning models, unless an exception is approved and documented.
Benchmarks and leaderboards help you build a shortlist. They don't tell you how a model does on your blocked orders. Use them to pick two or three candidates, then test.
Task
Volume
What matters most
Where to start the test
Sort blocked sales orders into a few reasons
High, every order
Accuracy on a fixed label set, cost per call, speed
A small, fast model, checked against a larger one
Summarize a three-way match exception for an AP clerk
Medium
Faithful to the documents, readable
A mid-sized model
Draft a customer email about a disputed invoice
Low
Tone, judgment, few errors
A larger model; a person reviews before sending
Explain an MRP exception list to a planner
Medium, long input
Context size, faithful to the data
A model whose context window fits the full list
Two starting strategies are common. OpenAI's agent guide recommends starting with the most capable model to set a quality baseline, then trying smaller models to see if they still meet the target. Teams under tight cost or speed limits sometimes start small and move up only when tests fail. Both work if you measure.
"The top of the leaderboard is the right choice." Leaderboards measure general tasks. Your task may be narrow, and a smaller model may score just as well on it at a fraction of the cost.
"Choose once and you're done." Versions are deprecated and retired on dates SAP publishes. A model choice needs an owner and a review date.
"Cost depends on the number of questions." Cost depends on tokens, in and out. A long prompt with company context can cost more than a short question with a long answer, and the other way round.
"Switching models means rebuilding the application." Through SAP's orchestration service, the model is a configuration value. What takes time is re-testing, so keep the tests ready.
"A fallback model is free insurance." A fallback only helps if it has passed the same tests. Otherwise it quietly lowers quality when it kicks in.
Pick one answer for each question. The explanation appears after you choose.
1An order-to-cash team asks which LLM is "best". What is the most useful answer?
Answer: B. There is no single best model, only the right one for a task, volume and budget. Leaderboards help build a shortlist, but only tests on your own examples show which model is good enough for your process.
2Why does model choice matter so much for a high-volume task like sorting blocked sales orders?
Answer: D. Models are paid per token, and per-token prices differ across models. At thousands of calls a day, a small difference per call becomes a large monthly figure, so it pays to test whether a cheaper model is good enough.
3What does SAP's orchestration service change about switching models later?
Answer: B. SAP describes orchestration as provider agnostic: you switch models through a configuration value. Re-testing is still needed, because a new model can behave differently on your cases.
4Your assistant pins a model version. What risk must someone own?
Answer: C. SAP's documentation says a deployment pinned to a model version stops working on that version's deprecation date. Someone needs to watch the dates and plan the move, or use latest and re-test after upgrades.
5A vendor says their solution has a fallback model "for free resilience". What should you ask?
Answer: D. A fallback only protects the process if it is good enough at the task. If it was never tested, it can quietly lower quality whenever the primary is unavailable.
6What does an allow list on an orchestration deployment give a company?
Answer: A. SAP lets administrators restrict an orchestration deployment to an allow list or deny list of models. It is a control for company standards, for example models approved for certain data.
7Which use of the generative AI hub does SAP's documentation forbid unless approved and documented?
Answer: C. SAP's documentation says SAP AI Core, including the generative AI hub, must not be used to generate synthetic data for training or fine-tuning, unless an exception is explicitly approved and documented. The other uses are normal parts of choosing and calling models.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: a model call is a priced, versioned request
Every LLM call is the same kind of thing: a web request that names a model and version, carries messages and a few parameters, and comes back with text plus a usage count of tokens in and out. Choosing a model means trading four measurable things against each other on your own cases: quality, cost, latency and lifecycle.
flowchart LR
C[Catalog<br/>what you may call] --> S[Shortlist<br/>2-3 models]
S --> T[Test on your cases<br/>score, time, tokens]
T --> D[Decision record<br/>model, version, fallback]
D --> P[Config value<br/>in your app]
P -->|new version or retirement| T
The loop at the end is the point. The choice is not made once. It lives in configuration, and your test cases let you re-check it in minutes when a version changes.
SAP meters generative AI use in tokens. They are converted into GenAI tokens and then into BTP capacity units, at rates that differ by model (SAP Note 3437766 lists them). SAP's pricing page says output tokens tend to cost slightly more than input tokens.
So cost per call is roughly:
cost per call = input tokens x input price + output tokens x output price
Two levers follow. The shape of the task sets the token ratio: SAP's page notes that summarization can run 3 or 4 input tokens per output token, while drafting text can produce more output than input. And the model sets the price per token. A one-word classifier with a short prompt is cheap on any model; a long system prompt sent on every call is not.
Answers are generated one token at a time, as How LLMs generate text showed. Long answers take longer than short ones, whichever model you use. Beyond that, measure: latency differs by model, by region and by load. Record the median, not one lucky call.
The catalog lists each version's contextLength. If a task needs a long MRP exception list or a full contract in one call, the model must accept that many tokens plus room for the answer. Filter on it first; it is a hard limit, not a trade-off.
SAP AI Core answers GET {AI_API_URL}/v2/lm/scenarios/foundation-models/models with every model your account can use. The fields that matter for choice:
Field
What it tells you
model
The name you put in your call
provider
Who built it
allowedScenarios
Whether you can call it through orchestration, through a foundation-models deployment, or both
SAP's Model Lifecycle page gives two options. Auto upgrade: use latest, and your calls move to new versions as SAP adds them. Manual upgrade: name a version, and it stays until you change it. If you pin a version, calls stop working on its deprecation date. latest is the default when you name no version.
Neither is free. latest can change behaviour under you; a pinned version needs someone watching dates. Either way, keep your test cases and re-run them when the version changes.
In orchestration version 2, modules can be a list of configurations in order of preference. SAP documents when orchestration moves to the next one:
for every request, if the model isn't supported in the deployed region or environment;
for non-streaming requests, also on 408 Request Timeout, 429 Too Many Requests and any 5xx server error.
Other errors fail the request. A successful response lists skipped attempts under intermediate_failures. The final_result.model field tells you which model answered.
Per model, SAP documents a timeout from 1 to 600 seconds (default 600) and max_retries from 0 to 5 (default 2). For a user waiting on screen, 600 seconds is far too long; set a timeout that matches the screen or job.
sequenceDiagram
participant App as Your script
participant O as Orchestration
participant A as Model A
participant B as Model B
App->>O: config with modules [A, B]
O->>A: call
A-->>O: 429 or not in region
O->>B: same prompt
B-->>O: answer
O-->>App: final_result (model B) + intermediate_failures
#Build it yourself: shortlist, compare and add a fallback
You will build one script, choose_model.py, with three commands. catalog reads SAP's model list and filters it. compare sends six made-up blocked sales orders to each model you name, scores the one-word answers, and records time, tokens and a cost index. fallback calls a primary model with a backup behind it and shows which one answered.
flowchart LR
E[.env<br/>AICORE_ lines] --> CAT[catalog<br/>filter models]
CAT --> CMP[compare<br/>6 test cases per model]
CMP --> CSV[model_comparison.csv]
CMP --> FB[fallback<br/>primary + backup]
Before you start: complete Set up your computer for this course and Set up for Unit 5. They create your orchestrate-course folder, its .venv, the AICORE_ lines in .env, and install sap-ai-sdk-gen. This walkthrough doesn't repeat those steps.
Your course folder from Unit 5 setup, with sap-ai-sdk-gen and python-dotenv installed.
About 45 minutes.
For real calls: SAP AI Core access with the generative AI hub (the trial or a company account). A comparison of two models is 12 short calls, a small per-request charge on a paid account and no charge during the trial.
No account? Every command has a --sample option with made-up data.
#Step 1: Open your course folder and turn on the virtual environment
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
If the prompt doesn't start with (.venv), turn it on:
In VS Code's file list, right-click unit05, choose New File and name it choose_model.py.
Paste the code below and save.
"""Unit 5: shortlist models from SAP's catalog, then compare them on your own SAP-shaped test cases.
Three commands (run from your course folder, with .venv turned on):
python unit05/choose_model.py catalog --sample # no account: a made-up catalog
python unit05/choose_model.py catalog --min-context 100000 # real catalog from SAP AI Core
python unit05/choose_model.py compare --sample # no account: made-up answers
python unit05/choose_model.py compare --models MODEL_A,MODEL_B # real calls through orchestration
python unit05/choose_model.py fallback --sample # no account: see a fallback happen
python unit05/choose_model.py fallback --primary MODEL_A --backup MODEL_B
The real paths read the AICORE_ lines in .env (see "Set up for Unit 5").
"compare" writes unit05/model_comparison.csv, which you reuse in the evaluation unit.
"""
import argparse
import base64
import csv
import json
import os
import statistics
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
from pathlib import Path
HERE = Path(__file__).resolve().parent
LABELS = ["CREDIT", "PRICING", "INCOMPLETE", "EXPORT"]
# Six blocked sales orders with the label a person gave each one. All made up.
CASES = [
("Order 4711: customer 10023 has open items of 52,000 EUR against a credit limit of 50,000 EUR.", "CREDIT"),
("Order 4712: no price found for material M-200 in sales org 1010; net value is zero.", "PRICING"),
("Order 4713: payment terms and Incoterms are missing in the header.", "INCOMPLETE"),
("Order 4714: ship-to party is in a country on the embargo list; trade compliance check pending.", "EXPORT"),
("Order 4715: credit check failed after the customer's rating dropped to high risk.", "CREDIT"),
("Order 4716: manual price change of 40 percent exceeds the allowed tolerance.", "PRICING"),
]
SYSTEM = ("You classify why an SAP sales order is blocked. Answer with exactly one word from this list: "
+ ", ".join(LABELS) + ". No other text.")
# A made-up catalog in the shape SAP documents for GET /v2/lm/scenarios/foundation-models/models.
SAMPLE_CATALOG = {"count": 4, "resources": [
{"model": "sample--large", "provider": "SampleAI", "executableId": "sample",
"allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
"versions": [{"name": "2026-06-01", "isLatest": True, "deprecated": False, "retirementDate": "",
"contextLength": 400000, "capabilities": ["text-generation"],
"cost": [{"inputCost": "0.0050"}, {"outputCost": "0.0250"}]}]},
{"model": "sample--small", "provider": "SampleAI", "executableId": "sample",
"allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
"versions": [{"name": "2026-06-01", "isLatest": True, "deprecated": False, "retirementDate": "",
"contextLength": 128000, "capabilities": ["text-generation"],
"cost": [{"inputCost": "0.0004"}, {"outputCost": "0.0016"}]}]},
{"model": "sample--old", "provider": "SampleAI", "executableId": "sample",
"allowedScenarios": [{"scenarioId": "orchestration", "executableId": "orchestration"}],
"versions": [{"name": "2024-05-13", "isLatest": True, "deprecated": True, "retirementDate": "2026-11-30",
"contextLength": 128000, "capabilities": ["text-generation"],
"cost": [{"inputCost": "0.0030"}, {"outputCost": "0.0090"}]}]},
{"model": "sample--embedder", "provider": "SampleAI", "executableId": "sample",
"allowedScenarios": [{"scenarioId": "foundation-models", "executableId": "sample"}],
"versions": [{"name": "1", "isLatest": True, "deprecated": False, "retirementDate": "",
"contextLength": 8000, "capabilities": ["embeddings"],
"cost": [{"inputCost": "0.0001"}]}]},
]}
# Made-up answers for --sample: the small model gets one case wrong.
SAMPLE_ANSWERS = {
"sample--large": (["CREDIT", "PRICING", "INCOMPLETE", "EXPORT", "CREDIT", "PRICING"], 1.9),
"sample--small": (["CREDIT", "PRICING", "INCOMPLETE", "EXPORT", "CREDIT", "INCOMPLETE"], 0.6),
}
# ---------- the catalog ----------
def env_or_exit() -> dict:
"""Read the five AICORE_ settings from .env, or stop with a clear message."""
from dotenv import load_dotenv
load_dotenv()
names = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
"AICORE_RESOURCE_GROUP"]
missing = [n for n in names if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5. "
"Or add --sample to try without an account.")
return {n: os.environ[n] for n in names}
def fetch_catalog(env: dict) -> dict:
"""Get a token with the service key details, then ask SAP AI Core for its model list."""
basic = base64.b64encode(f"{env['AICORE_CLIENT_ID']}:{env['AICORE_CLIENT_SECRET']}".encode()).decode()
body = urllib.parse.urlencode({"grant_type": "client_credentials"}).encode()
request = urllib.request.Request(env["AICORE_AUTH_URL"], data=body, headers={
"Authorization": f"Basic {basic}", "Content-Type": "application/x-www-form-urlencoded"})
try:
with urllib.request.urlopen(request, timeout=30) as reply:
token = json.loads(reply.read())["access_token"]
url = env["AICORE_BASE_URL"].rstrip("/") + "/lm/scenarios/foundation-models/models"
request = urllib.request.Request(url, headers={
"Authorization": f"Bearer {token}", "AI-Resource-Group": env["AICORE_RESOURCE_GROUP"]})
with urllib.request.urlopen(request, timeout=30) as reply:
return json.loads(reply.read())
except urllib.error.HTTPError as error:
sys.exit(f"SAP AI Core answered HTTP {error.code} {error.reason}. Run check_unit05.py to find the cause.")
except (urllib.error.URLError, KeyError) as error:
sys.exit(f"Could not reach SAP AI Core ({error}). Check your network and the AICORE_ lines in .env.")
def cost_of(version: dict, key: str) -> float:
"""The catalog lists cost as [{"inputCost": "..."}, {"outputCost": "..."}]."""
for item in version.get("cost", []):
if key in item:
try:
return float(item[key])
except (TypeError, ValueError):
return 0.0
return 0.0
def shortlist(catalog: dict, min_context: int) -> list:
"""Keep chat models you can call through orchestration, with enough context, newest version only."""
rows = []
for model in catalog.get("resources", []):
if not any(s.get("scenarioId") == "orchestration" for s in model.get("allowedScenarios", [])):
continue
versions = model.get("versions", [])
version = next((v for v in versions if v.get("isLatest")), versions[0] if versions else None)
if not version or "text-generation" not in version.get("capabilities", []):
continue
if (version.get("contextLength") or 0) < min_context:
continue
rows.append({"model": model["model"], "provider": model.get("provider", ""),
"context": version.get("contextLength") or 0,
"in_cost": cost_of(version, "inputCost"), "out_cost": cost_of(version, "outputCost"),
"deprecated": bool(version.get("deprecated")),
"retires": version.get("retirementDate") or ""})
return sorted(rows, key=lambda r: (r["deprecated"], r["in_cost"] + r["out_cost"]))
def cmd_catalog(args) -> None:
catalog = SAMPLE_CATALOG if args.sample else fetch_catalog(env_or_exit())
rows = shortlist(catalog, args.min_context)
print(f"{catalog.get('count', len(catalog.get('resources', [])))} models in the catalog; "
f"{len(rows)} of them are chat models usable through orchestration with at least "
f"{args.min_context:,} tokens of context.\n")
print(f"{'model':<34}{'provider':<14}{'context':>10}{'in cost':>10}{'out cost':>10} status")
for r in rows:
status = "ok"
if r["deprecated"] or r["retires"]:
status = "DEPRECATED" if r["deprecated"] else "retiring"
status += f", retires {r['retires']}" if r["retires"] else ""
print(f"{r['model']:<34}{r['provider'][:13]:<14}{r['context']:>10,}"
f"{r['in_cost']:>10.4f}{r['out_cost']:>10.4f} {status}")
if not rows:
print("(none) Lower --min-context, or ask your administrator which models your account offers.")
# ---------- calling models ----------
def make_config(model: str, max_out: int = 0, backup: str = ""):
"""Build an orchestration v2 configuration: one template, one model (plus an optional fallback)."""
from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
PromptTemplatingModuleConfig, SystemMessage, Template,
UserMessage)
def module(name: str):
params = {}
if max_out: # Anthropic models take max_tokens; the newer OpenAI models take max_completion_tokens
params["max_tokens" if name.startswith("anthropic--") else "max_completion_tokens"] = max_out
template = Template(template=[SystemMessage(content=SYSTEM), UserMessage(content="{{?order}}")])
return ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=name, params=params or None, timeout=60, max_retries=1)))
return OrchestrationConfig(modules=[module(model), module(backup)] if backup else module(model))
def call(service, text: str):
"""One call. Returns (answer, input tokens, output tokens, seconds, model that answered, failures)."""
start = time.perf_counter()
result = service.run(placeholder_values={"order": text})
seconds = time.perf_counter() - start
final = result.final_result
failures = [f"{f.code} {f.message}" for f in (result.intermediate_failures or [])]
return (final.choices[0].message.content or "", final.usage.prompt_tokens, final.usage.completion_tokens,
seconds, final.model, failures)
def clean(answer: str) -> str:
"""Models sometimes add a full stop or spaces. Keep the first word, in capitals."""
words = answer.strip().split()
return words[0].strip(".,:;!*\"'").upper() if words else ""
def cmd_compare(args) -> None:
models = list(SAMPLE_ANSWERS) if args.sample else [m.strip() for m in args.models.split(",") if m.strip()]
if not models:
sys.exit("Name the models to compare, for example --models MODEL_A,MODEL_B (catalog lists them).")
catalog = SAMPLE_CATALOG if args.sample else fetch_catalog(env_or_exit())
prices = {r["model"]: r for r in shortlist(catalog, 0)}
rows, summary = [], []
for model in models:
print(f"\n{model}")
correct, seconds, tokens_in, tokens_out = 0, [], 0, 0
service = None
if not args.sample:
from gen_ai_hub.orchestration_v2 import OrchestrationService
service = OrchestrationService(config=make_config(model, args.max_out))
try:
for i, (text, expected) in enumerate(CASES):
if args.sample:
answers, delay = SAMPLE_ANSWERS[model]
answer = answers[i] if i < len(answers) else expected # your own added cases: made-up correct
t_in, t_out, took = 70, 2, delay + 0.1 * (i % 3)
else:
try:
answer, t_in, t_out, took, _, _ = call(service, text)
except Exception as error: # keep going: one failure shouldn't hide the rest
print(f" case {i + 1}: call failed: {type(error).__name__}: {str(error)[:200]}")
continue
got = clean(answer)
ok = got == expected
correct += ok
seconds.append(took)
tokens_in += t_in
tokens_out += t_out
print(f" case {i + 1}: expected {expected:<10} got {got:<10} {'ok' if ok else 'WRONG'}"
f" {took:.1f}s")
rows.append({"model": model, "case": i + 1, "expected": expected, "got": got,
"correct": ok, "seconds": round(took, 2), "tokens_in": t_in, "tokens_out": t_out})
finally:
if service is not None:
service.close_http_connection()
price = prices.get(model, {})
cost = (tokens_in * price.get("in_cost", 0) + tokens_out * price.get("out_cost", 0)) / 1000
summary.append((model, correct, statistics.median(seconds) if seconds else 0.0,
tokens_in, tokens_out, cost, bool(price)))
print(f"\n{'model':<34}{'correct':>9}{'median s':>10}{'tokens in':>11}{'tokens out':>11}{'cost index':>12}")
for model, correct, median, t_in, t_out, cost, priced in summary:
print(f"{model:<34}{f'{correct}/{len(CASES)}':>9}{median:>10.1f}{t_in:>11}{t_out:>11}"
f"{(f'{cost:.4f}' if priced else 'n/a'):>12}")
out = HERE / "model_comparison.csv"
with open(out, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=list(rows[0]) if rows else ["model"])
writer.writeheader()
writer.writerows(rows)
print(f"\nSaved {len(rows)} rows to {out}")
if args.sample:
print("[sample] Made-up answers and timings; no model was called.")
def cmd_fallback(args) -> None:
text, expected = CASES[0]
if args.sample:
print(f"[sample] Primary {args.primary} is 'not supported' in this made-up region; "
f"orchestration used {args.backup}.")
print(f"Answer: {expected} answered by: {args.backup}")
print(f"Skipped: 400 Model {args.primary} not supported.")
return
env_or_exit()
from gen_ai_hub.orchestration_v2 import OrchestrationService
service = OrchestrationService(config=make_config(args.primary, args.max_out, backup=args.backup))
try:
answer, _, _, took, answered_by, failures = call(service, text)
except Exception as error:
sys.exit(f"Both models failed: {type(error).__name__}: {str(error)[:400]}")
finally:
service.close_http_connection()
print(f"Answer: {clean(answer)} answered by: {answered_by} ({took:.1f}s)")
print("Skipped: " + ("; ".join(failures) if failures else "nothing, the primary model answered"))
def main() -> None:
parser = argparse.ArgumentParser(description="Shortlist and compare LLMs in SAP's generative AI hub.")
sub = parser.add_subparsers(dest="command", required=True)
p = sub.add_parser("catalog", help="list chat models usable through orchestration")
p.add_argument("--min-context", type=int, default=0, help="smallest context window you need, in tokens")
p.add_argument("--sample", action="store_true", help="use a made-up catalog (no account)")
p = sub.add_parser("compare", help="run the six test cases on each model")
p.add_argument("--models", default="", help="comma-separated model names from the catalog")
p.add_argument("--max-out", type=int, default=0, help="optional cap on output tokens per answer")
p.add_argument("--sample", action="store_true", help="made-up answers (no account)")
p = sub.add_parser("fallback", help="call a primary model with a backup model behind it")
p.add_argument("--primary", default="sample--not-in-region")
p.add_argument("--backup", default="sample--small")
p.add_argument("--max-out", type=int, default=0, help="optional cap on output tokens per answer")
p.add_argument("--sample", action="store_true", help="show a made-up fallback (no account)")
args = parser.parse_args()
{"catalog": cmd_catalog, "compare": cmd_compare, "fallback": cmd_fallback}[args.command](args)
if __name__ == "__main__":
main()
What success looks like (with --sample; the names and numbers are made up):
4 models in the catalog; 3 of them are chat models usable through orchestration with at least 0 tokens of context.
model provider context in cost out cost status
sample--small SampleAI 128,000 0.0004 0.0016 ok
sample--large SampleAI 400,000 0.0050 0.0250 ok
sample--old SampleAI 128,000 0.0030 0.0090 DEPRECATED, retires 2026-11-30
The list is sorted with current models first, cheapest first. The embedding model in the sample catalog is left out: it can't be called through orchestration and doesn't generate text. With your key, the names, providers, costs and dates are your account's real ones and will differ from this.
If the table says (none), nothing matched your filter. Lower --min-context, or run python check_unit05.py to confirm your access.
Pick two models from your list for the next step: a larger one and a smaller, cheaper one. Avoid any marked DEPRECATED or retiring.
sample--large
case 1: expected CREDIT got CREDIT ok 1.9s
case 2: expected PRICING got PRICING ok 2.0s
case 3: expected INCOMPLETE got INCOMPLETE ok 2.1s
case 4: expected EXPORT got EXPORT ok 1.9s
case 5: expected CREDIT got CREDIT ok 2.0s
case 6: expected PRICING got PRICING ok 2.1s
sample--small
case 1: expected CREDIT got CREDIT ok 0.6s
case 2: expected PRICING got PRICING ok 0.7s
case 3: expected INCOMPLETE got INCOMPLETE ok 0.8s
case 4: expected EXPORT got EXPORT ok 0.6s
case 5: expected CREDIT got CREDIT ok 0.7s
case 6: expected PRICING got INCOMPLETE WRONG 0.8s
model correct median s tokens in tokens out cost index
sample--large 6/6 2.0 420 12 0.0024
sample--small 5/6 0.7 420 12 0.0002
Saved 12 rows to /Users/you/orchestrate-course/unit05/model_comparison.csv
[sample] Made-up answers and timings; no model was called.
Read the summary as a decision. In the sample, the small model is about three times faster and about a tenth of the cost index, but it sent a pricing case to the wrong team. Is one in six wrong acceptable? Not for routing work automatically. It might be for a suggestion a person confirms. That judgment is the point of the exercise.
With real models, expect different timings on every run, and possibly different answers; that is normal. Six cases are enough to learn the method, not to make a production decision. The Exercise grows the set.
Optional: cap the answer length with --max-out, for example --max-out 20. If answers come back empty with a cap, remove it: some models spend part of the allowance on internal reasoning before the visible answer, which SAP's parameter list says max_completion_tokens includes.
[sample] Primary sample--not-in-region is 'not supported' in this made-up region; orchestration used sample--small.
Answer: CREDIT answered by: sample--small
Skipped: 400 Model sample--not-in-region not supported.
With your key and a made-up primary, you should see your backup model under answered by and an error about the primary under Skipped. The exact wording of the error comes from SAP and may differ. With two real models, Skipped reads nothing, the primary model answered.
Six made-up blocked sales orders, each with the label a person would give
SYSTEM
The instruction: answer with exactly one of four labels
SAMPLE_CATALOG, SAMPLE_ANSWERS
Made-up data in SAP's documented catalog shape, for --sample
env_or_exit
Loads .env with load_dotenv() and checks the five AICORE_ settings
fetch_catalog
Swaps the client ID and secret for a token, then calls the model discovery endpoint
shortlist
Keeps chat models allowed in orchestration, reads the latest version's context, cost and lifecycle, sorts current and cheap first
make_config
Builds an orchestration v2 configuration: template, model, optional output cap, a 60-second timeout and one retry; with a backup, modules becomes a list
call
Runs one request and returns the answer, token counts, seconds, the model that answered and any skipped attempts
clean
Keeps the first word in capitals, so credit. still counts as CREDIT
cmd_compare
Runs every case on every model, prints a summary and writes model_comparison.csv
cost index
Tokens divided by 1,000, times the catalog's cost values; for comparing models, not a bill
The model's own API, for example chat completions for Azure OpenAI models
When to use
Default: provider agnostic, simpler to test and compare
When you need something only the model's own API offers
SAP's documentation names a limit too: orchestration supports LangChain, but not every open-source framework. This course uses orchestration through the SAP Cloud SDK for AI.
API:GET /v2/lm/scenarios/foundation-models/models, as in the script.
SAP AI Launchpad, Model Library: under Generative AI Hub > Model Library. Catalog lists models with filters and search. Leaderboard ranks them by benchmark. Chart plots two measures against each other. Each model card shows input types, cost information, metrics where available, deprecation notices and your rate limits. SAP lists the roles that can open it: genai_experimenter, genai_manager, genai_administrator or orchestration_executor.
SAP Note 3437766: token conversion rates, rate limits and deprecation dates. It needs an SAP login.
An administrator can restrict an orchestration deployment with two parameter bindings, modelFilterList and modelFilterListType (allow or deny). This is a sketch of the configuration body, based on SAP's sample; the model names are placeholders:
SAP's harmonized API accepts OpenAI-style parameters such as max_tokens, max_completion_tokens, temperature (0.0 to 2.0), top_p and stop. SAP notes that possible values depend on the model. One rule is documented explicitly: Anthropic models require max_tokens, and orchestration sets it to the model's maximum if you leave it out. The script sends no parameters by default for this reason, and picks the cap name by provider only when you ask for one.
The generative AI hub is available only in the extended plan of SAP AI Core; Set up for Unit 5 covers the trial route. Usage is metered in tokens and billed in capacity units, at rates that differ by model.
SAP generative AI hub, foundation-models deployment
Quick personal prototype, no SAP data
Simple, if you have an account
Works; needs SAP access
More setup than needed
Several providers under one contract
One contract per provider
Yes
Yes
Compare and switch models often
Code changes per provider API
Change a name in configuration
New deployment per model
Company-approved model list
Your own controls
Allow or deny list on the deployment
Control which deployments exist
Fallback to another provider
You build it
Built in (v2), as a list
You build it
Need a provider-only feature first
Yes
Only once orchestration supports it
Native API shape
Process runs on SAP data, side by side on BTP
Extra contract and data review
Fits the BTP security model
Fits the BTP security model
For most SAP processes in this course, orchestration is the default. Go direct, or to a foundation-models deployment, when you need something orchestration doesn't expose yet.
Security and SAP authorizations. The service key is a secret; it stays in .env or a secret store, never in code or Git. Restrict models with an allow list. Remember that the model sees whatever you put in the prompt: a blocked-order text may hold customer names and amounts. The orchestration topic later in this unit covers data masking and filtering; Unit 11 covers AI security.
Evaluation. Keep your test cases in version control with their expected answers. Re-run them when you change model, version, prompt or fallback. Unit 8 turns this into a full evaluation harness.
Cost. Estimate monthly cost from tokens per call times volume, using the rates in SAP Note 3437766. Watch long system prompts sent on every call, and long answers you don't need.
Operations. Set a timeout that fits the user's wait and a small max_retries. Log the model that actually answered and any intermediate_failures, so a silent fallback shows up in monitoring. Check rate limits on the model card, and request more with Increase Quota before go-live.
Lifecycle. Decide per use case between latest and a pinned version, and write down who checks deprecation dates. Put the next review date in the decision record.
Clean core. The model call runs side by side, outside the ERP. When a later version reads real blocked orders, read them through released APIs, as in Calling your first SAP API, not by changing S/4HANA.
You will grow the test set, re-run the comparison and write a one-page decision. Unit 8 reuses your test cases and your CSV as the start of an evaluation harness.
Open unit05/choose_model.py and find CASES.
Add four more blocked orders below the existing six, each on its own line in the same format. Make at least two of them hard, for example an order blocked for both credit and missing data. Give each the label you think is right.
Save the file.
Run the comparison on your two models. With no account, use compare --sample; your new cases then get made-up correct answers, so the point is the method, not the numbers.
Done when:model_comparison.csv has rows for 10 cases per model, and model_choice.md names a model, a version policy, a fallback and a review date, each backed by numbers from your runs.
Pick one answer for each question. The explanation appears after you choose.
1What does the response's usage block give you when choosing a model?
Answer: C. usage reports prompt_tokens and completion_tokens. Cost is metered in tokens, so these counts, times each model's rates, are how you compare what calls cost.
2In shortlist, why does the script check allowedScenarios for orchestration?
Answer: A. The catalog lists every model the account can use, including some only reachable through their own foundation-models deployment. The script calls models through orchestration, so it keeps only models allowed in that scenario.
3What happens to a call that pins a model version when that version reaches its deprecation date?
Answer: B. SAP's Model Lifecycle page says a deployment that names a version stops working on that version's deprecation date. Only latest upgrades automatically, and that can change behaviour, so re-run your tests either way.
4In make_config, what turns a single model into a fallback setup?
Answer: D. Orchestration version 2 accepts a list of module configurations and tries them in order. Retries repeat the same model; they don't switch to another one.
5Your primary model starts returning 429 Too Many Requests on a non-streaming call with a fallback set. What does SAP document will happen?
Answer: C. SAP lists 408, 429 and 5xx errors as reasons to move to the next configuration for non-streaming requests. The skipped attempt shows up in intermediate_failures, which is why you should log it.
6The small model is a tenth of the cost index but got one of six cases wrong. What would you do next?
Answer: B. Six cases are too few for a production decision. Grow the test set and compare against a target the business agreed, such as acceptable for suggestions but not for automatic routing.
7Why does the script send no output cap unless you pass --max-out?
Answer: D. SAP says possible parameter values depend on the model, and Anthropic models need max_tokens, which orchestration fills in if missing. Sending nothing by default is safe; the option picks the cap name by provider only when asked.
8A colleague wants to read the cost index as next month's AI bill. What should you tell them?
Answer: D. The catalog's cost fields have no stated unit on SAP's documentation page, so the index only compares models with each other. Real rates, converted to capacity units, are in SAP Note 3437766 and the company's agreement.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Supported Models (SAP AI Core documentation, SAP-docs on GitHub)— model discovery endpoint GET /v2/lm/scenarios/foundation-models/models; fields allowedScenarios, provider, versions[].contextLength, cost (inputCost, outputCost), deprecated, isLatest, retirementDate, streamingSupported, metadata; SAP Note 3437766 for token conversion rates, rate limits and deprecation dates
Model Lifecycle (SAP AI Core documentation, SAP-docs on GitHub)— model versions have deprecation dates; a deployment pinned to a version stops working on that date; auto upgrade with modelVersion latest or manual upgrade with a named version; latest is the default
Models (SAP AI Core documentation, SAP-docs on GitHub)— SAP-hosted and remote models; scenarios foundation-models and orchestration; orchestration is provider agnostic and lets you switch models with a config parameter without new deployments; must not be used to generate synthetic data for training or fine-tuning models
Orchestration with Fallbacks (SAP AI Core documentation, SAP-docs on GitHub)— a list of module configurations in preference order; orchestration V2 only; switches when a model isn't supported in the region and, for non-streaming requests, on 408, 429 and 5xx errors; intermediate_failures in the response
Model Configuration (SAP AI Core documentation, SAP-docs on GitHub)— model name required; params depend on the model; Anthropic models need max_tokens and orchestration sets it to the model maximum if missing; version defaults to latest; timeout 1 to 600 seconds (default 600); max_retries 0 to 5 (default 2)
A practical guide to building agents (OpenAI, PDF)— set up evals to establish a baseline, meet the accuracy target with the best models, then replace larger models with smaller ones where results stay acceptable