Most AI applications send every request to one model. That is simple, but you pay the price of your hardest request on every easy one.
Model selection means choosing the model for each task from evidence: you score a few models on your own labelled examples, then pick the cheapest one that meets your quality and speed targets.
Routing goes one step further and chooses per request. Easy requests go to a small, cheap, fast model. Hard ones go to a large model. A router can decide from what the request looks like, or let the small model try first and pass the request on when it is unsure. That second pattern is called a cascade.
Fallbacks keep the service running when a model is overloaded or down: the request moves to the next model on a list.
All three rest on one thing: an evaluation set, a few hundred real examples with the right answer, that tells you whether a cheaper path is still good enough.
Take the running example of this unit: an assistant that reads each blocked sales order and routes it to the right team: credit, pricing, master data or export.
Most blocked orders are routine. One clear block reason, no surprises. A small model handles them well. A few are messy: two block reasons at once, or a note from the sales rep that changes the picture. Those need a stronger model.
In this topic's hands-on replay, with made-up prices and models:
Approach
Correct
Cost per month (6,000 orders)
Small model for everything
74%
0.86 USD
Large model for everything
97%
43.20 USD
Cascade: small first, large when unsure
93%
14.86 USD
The cascade costs about a third of the large model and is right on 93 in 100 orders. Whether that is good enough is a business decision. A misrouted order costs a clerk a few minutes and may delay a delivery. If the team needs 97%, the answer is the large model. The point is that the choice is made with numbers, not by habit.
Two more things change once you route:
Resilience. SAP's documentation says model calls can be refused with a "too many requests" error when a rate limit or the provider's capacity is reached. A fallback model keeps orders flowing. In the replay, a 150-request outage of the large model failed every one of those requests until a fallback was added.
Hidden quality drops. A fallback or router that was never tested can quietly lower quality. In the replay, falling back to the small model kept the service up but cut accuracy during the outage to 77%.
Published research points the same way. The FrugalGPT study (Chen, Zaharia and Zou, 2023) found that provider prices could differ by two orders of magnitude, and that a cascade matched the strongest model on their test tasks at a fraction of its cost. Your savings depend on your data, so measure them.
As of October 2026, from the SAP sources opened for this topic:
Choosing. The generative AI hub in SAP AI Core lists the models your account can use, with costs and lifecycle dates. In SAP AI Launchpad, the Model Library adds a leaderboard and a chart of benchmark scores. Benchmarks help you shortlist; they don't score your blocked orders.
Testing.Evaluations, added in December 2025, benchmarks models and prompts as orchestration configurations on your own dataset, with predefined or custom metrics. This is where an SAP team can compare candidate models on its evaluation set.
Fallbacks. The orchestration service (version 2) accepts a list of configurations in order of preference. If the first model isn't available in the region, or fails with certain temporary errors, orchestration tries the next. The response says which model answered and what failed along the way.
Rate limits. Limits are set per model, in requests per minute. Teams can check them and request increases through an API or the model card.
Routing per request. No managed router that picks a model per request was found in the SAP sources opened for this topic. Your team builds that logic in its application, and sends each request to orchestration with the chosen model name.
The first orchestration endpoint is scheduled for decommissioning on 31 October 2026. Fallbacks need version 2, so a migration is due anyway.
SAP-delivered AI, such as Joule, chooses its own models. This topic applies to AI you build.
Evaluation picks the cheapest model that meets the target
Most teams, as the first step
Paying large-model prices on easy requests
Rules router
Visible features: number of block reasons, a note, order value
Cases where "hard" is easy to spot
Rules that looked right in testing and miss hard cases in production
Cascade
A small model answers; low confidence sends it to a large model
Mostly easy traffic with a hard tail
Confidently wrong small answers; slower escalated requests
Fallback chain
Next model on the list when one fails or is overloaded
Every production system
An untested fallback lowering quality without anyone noticing
A practical order: pick one model per task with an evaluation first. Add a tested fallback before go-live. Add routing only when volume makes the saving worth the extra testing.
"The best model is the one at the top of the leaderboard." The best model is the cheapest one that meets your target on your data.
"Routing always saves money." In a cascade you pay for the small call even when you escalate. If most requests escalate, the cascade can cost more than the large model alone.
"A cascade is as fast as the small model." Escalated requests wait for both models. In the replay, the slowest 5% of cascade requests waited longer than with the large model alone.
"A fallback is free insurance." A fallback that was never evaluated keeps the service up while quietly answering worse.
"The model's confidence tells us when it is wrong." It helps, but some wrong answers come back confident. Test the threshold on labelled data.
"We chose the model once." New models arrive in SAP's catalog every few weeks and versions retire. Selection is a repeatable test, not a one-off decision.
Pick one answer for each question. The explanation appears after you choose.
1The team proposes the large model for blocked-order triage "because it scores best". What should you ask first?
Answer: C. Model selection picks the cheapest model that meets an agreed target on your own examples. "Scores best" says nothing about whether a cheaper model would also be good enough for this task.
2In the replay, the cascade is right on 93% of orders at about a third of the large model's cost. Who decides whether 93% is enough?
Answer: B. The target is a business trade-off: what a wrong route costs order-to-cash against what accuracy costs in model fees. Engineers measure the options; the process owner sets the bar.
3A vendor says routing "always saves money". When can a cascade cost more than the large model alone?
Answer: D. In a cascade every escalated request pays for the small call and the large call. If the threshold sends most traffic onward, the total can pass the cost of using the large model for everything.
4The large model is overloaded for an hour. What keeps blocked orders flowing without a hidden quality drop?
Answer: A. A fallback keeps requests answered, but only a tested fallback keeps quality known. In the replay, falling back to the small model, which misses the target, cut accuracy during the outage to 77%.
5As of October 2026, what does SAP's generative AI hub offer for routing between models?
Answer: C. Orchestration version 2 takes a list of configurations as fallbacks. No managed per-request router was found in the SAP sources for this topic, so the routing logic lives in your application.
6Why should the team re-run its evaluation regularly after choosing a model?
Answer: C. SAP's catalog adds models every few weeks and versions have retirement dates. A ready evaluation lets the team switch in hours instead of rebuilding its case.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A model choice is a point on a quality-cost-latency map, and an evaluation set is the only way to plot it.
Each model is one point: its accuracy on your cases, its cost per request, its p95 wait. A router doesn't add a new model. It adds new points, mixes of models, and some of them sit where no single model can: almost the quality of the large model at a fraction of the cost.
Everything else follows from that:
Selection is reading the map: the cheapest point that meets your target.
Routing is drawing new points, and every new point must be measured, because routes fail in ways single models don't.
Fallbacks are the points you land on when your chosen one is unavailable. They belong on the map too, measured in advance.
flowchart LR
E[Evaluation set<br/>labelled cases] --> S[Score each model<br/>and each route]
S --> P[Pick cheapest that<br/>meets the target]
P --> R[Router + fallback chain<br/>in configuration]
R --> L[Logs: model used,<br/>escalations, failures]
L -->|new model, drift,<br/>outage| S
A choice needs a target first, agreed with the process owner: "route at least 90% of blocked orders correctly, with a p95 under 6 seconds". Defining success metrics with the business shows how to agree one.
Then score every candidate on the same evaluation set, with the same prompt. Record three numbers per model:
Anthropic's model guide describes two ways to start: with a fast, low-cost model, moving up only if it falls short, or with the most capable model, then testing whether cheaper ones keep up. Both work if you measure. The guide also notes that some models take an effort setting that trades intelligence for latency and cost within one model, so "which model" can become "which model at which setting".
A small evaluation set is a trap. Hard cases are rare by definition. In this topic's 150-case set, only 19 are hard, so one more mistake on them moves the hard-case accuracy by five points. Make sure the hard tail is represented, and keep adding production cases that went wrong. Building an evaluation harness has the tooling.
Rules routers decide before any model is called, from features you can see: the number of block reasons, whether a sales rep left a note, the order value or the customer's risk class. They are cheap, fast and easy to explain. Their weakness: "looks simple" is not "is simple". In the replay, one in four hard orders shows only one signal, and the rules send it to the small model.
Cascades let the cheap model try first and escalate when its answer is doubtful. FrugalGPT (Chen, Zaharia and Zou, 2023) studied this pattern, scoring each answer and stopping at the first model whose answer passes. Two cheap escalation signals are common:
Invalid output. The answer isn't one of the allowed queues, or the JSON doesn't match the schema. Structured outputs make this check exact.
Low confidence. The model returns a confidence with its answer, and the cascade escalates below a threshold. Treat that number as a signal to test, not as a probability: in the replay, one in five wrong small-model answers still reports high confidence.
A cascade has two costs people forget. You pay for the small call on every escalated request, and the user waits for both calls.
Learned routers predict, per request, whether the small model's answer will be good enough. RouteLLM (Ong and others, 2024) trained routers on human preference data from Chatbot Arena and reported cost reductions of more than two times in some cases without lowering quality. Its routers kept working when routing between model pairs not seen in training. A learned router is itself a model that needs evaluation and monitoring, so most enterprise teams start with rules or a cascade.
Routing is not only between models. The same patterns can choose a prompt, a reasoning effort setting, or whether to answer from a cache first (Caching for LLM applications).
Failures are routine at volume. SAP documents that SAP AI Core limits requests per minute per model, and that a 429 Too Many Requests can mean either your limit or temporary back-end saturation at the provider. Three mechanisms handle failures, and they fit together:
Mechanism
What it does
Use when
Cost
Retry with backoff
Waits and tries the same model again; honours Retry-After
Short spikes, a single 429
Extra wait; can add load if done without jitter
Fallback
Tries the next model on the list
The model is down, overloaded or not in the region
Quality of the fallback; needs evaluation
Circuit breaker
After repeated failures, stops calling the model for a while, then tries one call
An outage that lasts minutes
A healthy model may be skipped briefly after recovery
Martin Fowler describes the circuit breaker as a wrapper that monitors failures, trips after a threshold so calls fail fast, and after a reset timeout lets a trial call through. For a model call, "fail fast" means going straight to the fallback instead of waiting for a timeout on every request.
Order matters. Retry a short spike. Fall back when retries would make the user wait too long. Trip the breaker when the same model keeps failing, so that thousands of requests don't each wait out a timeout.
#Build it yourself: select, route and survive an outage
You will score three stand-in models on 150 labelled blocked orders, choose one, compare routing strategies, confirm the best on a month of new traffic, and then switch off the large model to see fallbacks and a circuit breaker at work. One script, model_router.py, does all of it with built-in Python.
Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. This script needs nothing else: no library, no account, no network.
flowchart LR
EV[150 labelled orders] --> A[evaluate<br/>score 3 models]
EV --> B[sweep<br/>compare routes]
B --> C[route<br/>500 new orders]
C --> D[fallback<br/>large model down]
In VS Code, right-click the unit10 folder, choose New File, name it model_router.py, paste the code below and save.
"""Unit 10: choose a model per task with your own evaluation data, route easy requests to a
small model, and keep answering when a model fails.
The task is the blocked-orders triage from earlier units: read a blocked sales order and route
it to one queue (credit, pricing, master_data or export). Three stand-in "models" answer it.
They are functions with made-up accuracy, prices and response times, so everything runs
offline with built-in Python. No model is called; no account; no key.
python unit10/model_router.py evaluate score each model on 150 labelled cases
python unit10/model_router.py evaluate --target 0.95 pick the cheapest model that meets 95%
python unit10/model_router.py route --strategy small send all traffic to one model
python unit10/model_router.py route --strategy rules route by what the order looks like
python unit10/model_router.py route --strategy cascade small model first, escalate when unsure
python unit10/model_router.py route --strategy cascade --threshold 0.9 replay new traffic
python unit10/model_router.py sweep compare strategies on the evaluation set
python unit10/model_router.py fallback an outage of the large model, no fallback
python unit10/model_router.py fallback --chain large,mid --breaker
"""
import argparse
import random
import statistics
import sys
# ---------------------------------------------------------------------------------------------
# EXAMPLE models, prices and timings, made up for this course. Replace them with your own
# measurements. Prices are USD per million tokens; seconds are typical response times.
# ---------------------------------------------------------------------------------------------
MODELS = {
"small": {"in": 0.10, "out": 0.40, "seconds": 0.6},
"mid": {"in": 1.00, "out": 4.00, "seconds": 1.4},
"large": {"in": 5.00, "out": 20.00, "seconds": 3.2},
}
# How often each stand-in model gets a case right, by how hard the case really is.
SKILL = {
"small": {"easy": 0.97, "medium": 0.55, "hard": 0.20},
"mid": {"easy": 0.99, "medium": 0.90, "hard": 0.55},
"large": {"easy": 1.00, "medium": 0.97, "hard": 0.88},
}
TOKENS_IN, TOKENS_OUT = 1320, 30 # one triage call: order data in, a short JSON answer out
ORDERS_PER_MONTH = 6000 # the volume used in the token economics topic
QUEUES = ["credit", "pricing", "master_data", "export"]
# ---------------------------------------------------------------------------------------------
# Made-up, SAP-shaped blocked orders. Each case has what the app can SEE (the signals SAP
# shows and whether a sales rep left a note) and what it can't: how hard the case really is.
# ---------------------------------------------------------------------------------------------
def make_cases(count: int, seed: int) -> list:
rng = random.Random(seed)
cases = []
for i in range(count):
difficulty = rng.choices(["easy", "medium", "hard"], weights=[60, 28, 12])[0]
label = rng.choice(QUEUES)
if difficulty == "easy": # one clear signal; sometimes an unimportant note
signals, note = 1, rng.random() < 0.15
elif difficulty == "medium": # one signal, and a note that changes the answer
signals, note = 1, rng.random() < 0.80
else: # several signals, often a note; one in four looks simple
signals, note = (1 if rng.random() < 0.25 else rng.choice([2, 3])), rng.random() < 0.6
cases.append({"order": str(9000200 + i + seed * 1000), "label": label, "difficulty": difficulty,
"signals": signals, "note": note})
return cases
EVAL_SET = make_cases(150, seed=1) # labelled by people: your evaluation data
TRAFFIC = make_cases(500, seed=2) # next month's live requests, labelled afterwards by reviewers
def call(model: str, case: dict) -> dict:
"""The stand-in model. Same model and same order always give the same result."""
rng = random.Random(f"{model}|{case['order']}")
correct = rng.random() < SKILL[model][case["difficulty"]]
answer = case["label"] if correct else rng.choice([q for q in QUEUES if q != case["label"]])
# Confidence the model reports with its answer. It is usually lower when the model is
# wrong, but not always: some wrong answers come back confident.
if correct:
confidence = rng.uniform(0.85, 0.99) if case["difficulty"] == "easy" else rng.uniform(0.6, 0.97)
else:
confidence = rng.uniform(0.88, 0.97) if rng.random() < 0.2 else rng.uniform(0.4, 0.8)
seconds = MODELS[model]["seconds"] * rng.uniform(0.7, 1.6)
cost = (TOKENS_IN * MODELS[model]["in"] + TOKENS_OUT * MODELS[model]["out"]) / 1e6
return {"model": model, "answer": answer, "confidence": round(confidence, 2),
"seconds": seconds, "cost": cost}
def p95(values: list) -> float:
ordered = sorted(values)
return ordered[max(0, int(round(0.95 * len(ordered))) - 1)]
# ---------------------------------------------------------------------------------------------
# evaluate: score every model on the evaluation set, then choose
# ---------------------------------------------------------------------------------------------
def evaluate(args) -> None:
print(f"Evaluation set: {len(EVAL_SET)} labelled blocked orders "
f"({sum(c['difficulty'] == 'hard' for c in EVAL_SET)} hard)\n")
print(f"{'model':<7}{'correct':>9}{'accuracy':>10}{'p50 s':>8}{'p95 s':>8}{'USD / 1,000':>13}")
rows = []
for model in MODELS:
results = [call(model, c) for c in EVAL_SET]
right = sum(r["answer"] == c["label"] for r, c in zip(results, EVAL_SET))
seconds = [r["seconds"] for r in results]
row = {"model": model, "accuracy": right / len(EVAL_SET), "p50": statistics.median(seconds),
"p95": p95(seconds), "per_1000": 1000 * results[0]["cost"]}
rows.append(row)
print(f"{model:<7}{right:>6}/{len(EVAL_SET)}{row['accuracy']:>10.0%}{row['p50']:>8.2f}"
f"{row['p95']:>8.2f}{row['per_1000']:>13.3f}")
print(f"\nTarget: accuracy >= {args.target:.0%}, p95 <= {args.p95_budget:.1f} s")
passing = [r for r in rows if r["accuracy"] >= args.target and r["p95"] <= args.p95_budget]
if not passing:
print("No single model meets both targets. Try routing, a better prompt, or relax a target.")
return
best = min(passing, key=lambda r: r["per_1000"])
monthly = best["per_1000"] * ORDERS_PER_MONTH / 1000
print(f"Cheapest model that passes: {best['model']} "
f"({monthly:.2f} USD a month at {ORDERS_PER_MONTH:,} orders, EXAMPLE prices)")
# ---------------------------------------------------------------------------------------------
# route: three ways to decide which model answers each request
# ---------------------------------------------------------------------------------------------
def route_rules(case: dict) -> list:
"""Decide from what the order looks like, before any model is called."""
if case["signals"] >= 2:
return [call("large", case)]
if case["note"]:
return [call("mid", case)]
return [call("small", case)]
def route_cascade(case: dict, threshold: float) -> list:
"""Ask the small model first; escalate to the large model when it is unsure."""
first = call("small", case)
if first["confidence"] >= threshold and first["answer"] in QUEUES:
return [first]
return [first, call("large", case)]
def run_traffic(strategy: str, threshold: float, cases: list) -> dict:
stats = {"right": 0, "cost": 0.0, "seconds": [], "answered_by": {m: 0 for m in MODELS},
"escalated": 0, "hard_wrong": 0}
for case in cases:
if strategy in MODELS:
calls = [call(strategy, case)]
elif strategy == "rules":
calls = route_rules(case)
else:
calls = route_cascade(case, threshold)
final = calls[-1]
stats["right"] += final["answer"] == case["label"]
stats["hard_wrong"] += final["answer"] != case["label"] and case["difficulty"] == "hard"
stats["cost"] += sum(c["cost"] for c in calls) # you pay for every call, also the first
stats["seconds"].append(sum(c["seconds"] for c in calls)) # and wait for every call
stats["answered_by"][final["model"]] += 1
stats["escalated"] += len(calls) > 1
return stats
def route(args) -> None:
s = run_traffic(args.strategy, args.threshold, TRAFFIC)
n = len(TRAFFIC)
label = f"cascade, threshold {args.threshold}" if args.strategy == "cascade" else args.strategy
print(f"Strategy: {label} ({n} requests)\n")
print(f"accuracy {s['right'] / n:>6.1%}")
shares = ", ".join(f"{m} {v / n:.0%}" for m, v in s["answered_by"].items() if v)
print(f"answered by {shares}")
if args.strategy == "cascade":
print(f"escalated {s['escalated'] / n:>6.0%}")
print(f"p50 / p95 wait {statistics.median(s['seconds']):.2f} / {p95(s['seconds']):.2f} s")
print(f"cost per 1,000 {1000 * s['cost'] / n:>6.3f} USD")
print(f"cost per month {ORDERS_PER_MONTH * s['cost'] / n:>6.2f} USD at {ORDERS_PER_MONTH:,} orders")
print(f"wrong on hard cases {s['hard_wrong']:>6}")
print("\nPrices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.")
def sweep(args) -> None:
n = len(EVAL_SET)
print(f"Evaluation set: {n} labelled blocked orders\n")
print(f"{'strategy':<18}{'accuracy':>9}{'escalated':>11}{'p95 s':>8}{'USD / 1,000':>13}")
rows = [("small", None), ("large", None), ("rules", None)] + [("cascade", t) for t in args.thresholds]
for strategy, threshold in rows:
s = run_traffic(strategy, threshold or 0.0, EVAL_SET)
name = f"cascade {threshold}" if threshold else strategy
esc = f"{s['escalated'] / n:.0%}" if threshold else "-"
print(f"{name:<18}{s['right'] / n:>9.1%}{esc:>11}{p95(s['seconds']):>8.2f}{1000 * s['cost'] / n:>13.3f}")
print("\nPick the cheapest row that meets your target, then confirm it on new traffic with route.")
# ---------------------------------------------------------------------------------------------
# fallback: requests 150 to 299 arrive while the large model times out
# ---------------------------------------------------------------------------------------------
TIMEOUT = 8.0 # seconds the app waits before it gives up on one call
OUTAGE = range(150, 300)
BREAKER_FAILURES = 5 # failures in a row that open the circuit breaker
BREAKER_SKIP = 40 # requests that skip the failing model before it is tried again
def fallback(args) -> None:
chain = args.chain.split(",")
for model in chain:
if model not in MODELS:
sys.exit(f"Unknown model '{model}'. Use names from: {', '.join(MODELS)}")
stats = {"right": 0, "failed": 0, "seconds": [], "answered_by": {m: 0 for m in MODELS},
"timeouts": 0, "outage_right": 0}
in_a_row, skip_until = 0, -1
for i, case in enumerate(TRAFFIC):
waited, final = 0.0, None
for model in chain:
if args.breaker and model == chain[0] and i < skip_until:
continue # breaker open: don't even try the failing model
if model == "large" and i in OUTAGE:
waited += TIMEOUT # the call hangs until our timeout
stats["timeouts"] += 1
if model == chain[0]:
in_a_row += 1
if args.breaker and in_a_row >= BREAKER_FAILURES:
skip_until, in_a_row = i + BREAKER_SKIP, 0 # open, then try again later
continue
if model == chain[0]:
in_a_row = 0
final = call(model, case)
waited += final["seconds"]
break
stats["seconds"].append(waited)
if final is None:
stats["failed"] += 1
continue
stats["answered_by"][final["model"]] += 1
ok = final["answer"] == case["label"]
stats["right"] += ok
stats["outage_right"] += ok and i in OUTAGE
n = len(TRAFFIC)
print(f"Chain: {' -> '.join(chain)}{' circuit breaker on' if args.breaker else ''}")
print(f"Outage: the large model times out for requests {OUTAGE.start}-{OUTAGE.stop - 1} "
f"(timeout {TIMEOUT:.0f} s)\n")
print(f"requests failed {stats['failed']:>6}")
shares = ", ".join(f"{m} {v}" for m, v in stats["answered_by"].items() if v)
print(f"answered by {shares}")
print(f"timed-out calls {stats['timeouts']:>6}")
print(f"accuracy, all {stats['right'] / n:>6.1%}")
print(f"accuracy, outage {stats['outage_right'] / len(OUTAGE):>6.1%}")
print(f"p50 / p95 wait {statistics.median(stats['seconds']):.2f} / {p95(stats['seconds']):.2f} s")
def main() -> None:
parser = argparse.ArgumentParser(description="Model selection, routing and fallbacks, offline.")
sub = parser.add_subparsers(dest="command", required=True)
ev = sub.add_parser("evaluate", help="score each model on the evaluation set")
ev.add_argument("--target", type=float, default=0.90, help="accuracy needed (0 to 1)")
ev.add_argument("--p95-budget", type=float, default=6.0, help="slowest acceptable p95 in seconds")
ro = sub.add_parser("route", help="replay the sample traffic with one routing strategy")
ro.add_argument("--strategy", choices=list(MODELS) + ["rules", "cascade"], default="cascade")
ro.add_argument("--threshold", type=float, default=0.8, help="cascade: confidence needed to keep the small answer")
sw = sub.add_parser("sweep", help="compare strategies and cascade thresholds")
sw.add_argument("--thresholds", type=float, nargs="+", default=[0.7, 0.8, 0.85, 0.9, 0.95])
fb = sub.add_parser("fallback", help="replay traffic during an outage of the large model")
fb.add_argument("--chain", default="large", help="models to try in order, e.g. large,mid")
fb.add_argument("--breaker", action="store_true", help="skip the first model after repeated failures")
args = parser.parse_args()
for name in ("target", "threshold"):
if hasattr(args, name) and not 0 < getattr(args, name) <= 1:
sys.exit(f"--{name} must be above 0 and at most 1.")
if hasattr(args, "thresholds") and not all(0 < t <= 1 for t in args.thresholds):
sys.exit("--thresholds must each be above 0 and at most 1.")
{"evaluate": evaluate, "route": route, "sweep": sweep, "fallback": fallback}[args.command](args)
if __name__ == "__main__":
main()
Check that the file is in the right place:
Windows (PowerShell):
dir unit10\model_router.py
macOS / Linux:
ls unit10/model_router.py
You should see the file name. An error means the file is in another folder or has another name.
Evaluation set: 150 labelled blocked orders (19 hard)
model correct accuracy p50 s p95 s USD / 1,000
small 114/150 76% 0.66 0.94 0.144
mid 131/150 87% 1.59 2.17 1.440
large 149/150 99% 3.86 5.00 7.200
Target: accuracy >= 90%, p95 <= 6.0 s
Cheapest model that passes: large (43.20 USD a month at 6,000 orders, EXAMPLE prices)
Read the table row by row. The small model is fast and costs 0.144 USD per 1,000 orders, but routes only 76% correctly. The large model gets 149 of 150 right and costs 50 times more. The mid model sits between them and misses the 90% target.
Strategy: rules (500 requests)
accuracy 88.8%
answered by small 54%, mid 36%, large 11%
p50 / p95 wait 0.92 / 3.85 s
cost per 1,000 1.353 USD
cost per month 8.12 USD at 6,000 orders
wrong on hard cases 21
Prices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.
Accuracy fell from 94% on the evaluation set to 88.8% on new traffic, below the target. The rules were tuned on what hard orders look like, and some hard orders in the new month look simple.
Strategy: cascade, threshold 0.8 (500 requests)
accuracy 92.6%
answered by small 68%, large 32%
escalated 32%
p50 / p95 wait 0.83 / 5.31 s
cost per 1,000 2.477 USD
cost per month 14.86 USD at 6,000 orders
wrong on hard cases 16
Prices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.
The cascade holds the target at 92.6%, for about a third of the large model's monthly cost. Its weakness is visible too: p95 rises above five seconds, and 16 hard orders are still wrong, mostly ones where the small model answered wrongly with high confidence.
For comparison, run the large model alone:
python unit10/model_router.py route --strategy large
It reaches 97.0% at 43.20 USD a month. Write down the three options with their numbers. Choosing between them is the decision the process owner makes.
The script simulates an outage: requests 150 to 299 reach a large model that hangs until the app's 8-second timeout.
Run with no fallback:
python unit10/model_router.py fallback
Chain: large
Outage: the large model times out for requests 150-299 (timeout 8 s)
requests failed 150
answered by large 350
timed-out calls 150
accuracy, all 68.2%
accuracy, outage 0.0%
p50 / p95 wait 4.29 / 8.00 s
150 requests failed. Every one of them also made the user wait the full 8 seconds.
Chain: large -> small
Outage: the large model times out for requests 150-299 (timeout 8 s)
requests failed 0
answered by small 150, large 350
timed-out calls 150
accuracy, all 91.2%
accuracy, outage 76.7%
p50 / p95 wait 4.29 / 8.85 s
Nothing failed, but accuracy during the outage dropped to 76.7%. Nobody would notice unless the logs show which model answered.
Chain: large -> mid
Outage: the large model times out for requests 150-299 (timeout 8 s)
requests failed 0
answered by mid 150, large 350
timed-out calls 150
accuracy, all 93.8%
accuracy, outage 85.3%
p50 / p95 wait 4.29 / 10.06 s
Better quality during the outage, but p95 is now above 10 seconds: every request in the outage waits 8 seconds for the timeout before the fallback even starts.
Chain: large -> mid circuit breaker on
Outage: the large model times out for requests 150-299 (timeout 8 s)
requests failed 0
answered by mid 176, large 324
timed-out calls 20
accuracy, all 93.6%
accuracy, outage 85.3%
p50 / p95 wait 3.11 / 5.07 s
Only 20 calls timed out instead of 150. After five timeouts in a row, the breaker skips the large model for 40 requests, then lets one trial call through. p95 is back near normal. The price: 26 requests after the outage still went to the mid model while the breaker was open.
Model catalog API.GET /v2/lm/scenarios/foundation-models/models lists models with cost figures, context length and lifecycle fields. Choosing and calling LLMs reads it in code.
Model Library in SAP AI Launchpad: catalog mode with filters, leaderboard mode ranking models by benchmark, and chart mode plotting two measures against each other. Each model card shows cost information, deprecation notices and your rate limits, with Increase Quota.
Keep watching. SAP's What's New page for SAP AI Core added new models in most months of 2026, each pointing to SAP Note 3437766 for conversion rates, rate limits and deprecation dates. A ready evaluation set turns each new model into a one-hour test.
Evaluations (added 8 December 2025) benchmarks models and prompts as orchestration configurations on your own dataset, with predefined metrics or custom LLM-as-a-judge metrics. For model selection, that means one configuration per candidate model, the same prompt, and your labelled blocked orders. Building an evaluation harness covers how it fits with your own release gate.
In orchestration version 2, modules can be a list of module configurations, tried in order. SAP documents when orchestration moves on:
all requests: the model isn't supported in the deployed region or environment;
non-streaming requests only:408 Request Timeout, 429 Too Many Requests, or any 5xx server error.
Any other error fails the request. A successful response carries intermediate_failures with the skipped attempts; if every configuration fails, the error lists each attempt in order. Each entry is a full configuration, so the fallback can carry its own prompt, parameters and modules. SAP's own example drops the filtering module in the fallback. Check that your fallback keeps the filtering and masking your process needs.
In the Python SDK (sap-ai-sdk-gen 7.4.1, read for this topic), OrchestrationConfig.modules accepts one ModuleConfig or a list. This is a sketch: it needs SAP AI Core with the generative AI hub, set up in Set up for Unit 5, and the model names are placeholders for models in your catalog.
# Sketch: needs SAP AI Core with the generative AI hub (Unit 5). Model names are placeholders.
from gen_ai_hub.orchestration_v2.models.config import ModuleConfig, OrchestrationConfig
from gen_ai_hub.orchestration_v2.models.llm_model_details import LLMModelDetails
from gen_ai_hub.orchestration_v2.models.message import SystemMessage, UserMessage
from gen_ai_hub.orchestration_v2.models.template import PromptTemplatingModuleConfig, Template
from gen_ai_hub.orchestration_v2.service import OrchestrationService
PROMPT = Template(template=[
SystemMessage(content="Route the blocked sales order to one queue: credit, pricing, master_data or export."),
UserMessage(content="{{?order}}"),
])
def module(model_name: str) -> ModuleConfig:
# A short timeout per model so the fallback starts before the user gives up.
return ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=PROMPT, model=LLMModelDetails(name=model_name, timeout=8, max_retries=1)))
config = OrchestrationConfig(modules=[module("PRIMARY_MODEL"), module("TESTED_FALLBACK_MODEL")])
response = OrchestrationService().run(config=config, placeholder_values={"order": "..."})
print(response.final_result.model, response.intermediate_failures)
The SDK documents timeout from 1 to 600 seconds (default 600) and max_retries from 0 to 5 (default 2) per model, and notes both are currently ignored for Vertex AI models. Its run_with_retries method retries 429 and server errors with exponential backoff and jitter, and uses Retry-After when present.
SAP AI Core limits requests per minute per model, shared by all versions of that model, across the tenant by default. You can set separate limits per resource group, which lets you isolate a batch job from an interactive assistant. GET /v2/admin/quota/model shows your limits; POST /v2/admin/quota/requests asks for more and returns a requestId for an approval process. SAP lists model fallbacks, caching and workload isolation among the mitigations for 429s.
Routing changes your quota needs: a cascade sends every request to the small model and a share to the large one. Size each model's limit from the routing shares, plus headroom for fallback traffic during an outage.
No managed per-request router was found in the SAP AI Core sources opened for this topic. Build it in your application on SAP BTP: compute the route, then call orchestration with the chosen model name (or a fallback list headed by it). Because the model is a configuration value in orchestration, routing needs no extra deployments. An allow list on the orchestration deployment, described in Choosing and calling LLMs, keeps every route inside the approved models.
For work that can wait, SAP added batch consumption for models called through foundation-models deployments on 8 May 2026, described as a way to reduce cost. An overnight re-triage of all open orders is a candidate. Check the Batch Consumption page for supported models before planning on it.
Authorizations and data. Every model on a route sees the prompt. Each must be approved for the data, including fallbacks. Enforce this with the orchestration allow list rather than trusting the router code. Keep masking and filtering on every fallback configuration.
Evaluation. Score each route, not just each model, and confirm on data the threshold wasn't tuned on. Re-run when a model, version, prompt, threshold or fallback changes. Keep hard cases in the set on purpose.
Observability. Log on every request the route taken, the model that answered, the small model's confidence, whether it escalated, and any intermediate_failures. Alert on the escalation rate and the fallback rate: both drifting up is an early warning. Observability for AI systems shows where these attributes go.
Cost. Count every call in a cascade, not just the one that answered. Recompute monthly cost from real routing shares, since traffic mix changes at month-end.
Latency. Escalated and fallback requests are your tail. Set per-model timeouts to fit the latency budget, not the 600-second default.
Quotas. Size rate limits per model from routing shares, and isolate batch and interactive workloads in separate resource groups.
Lifecycle. Pin or track versions for every model on every route. A retired fallback fails at the worst moment. Migrate off the first orchestration endpoint before its 31 October 2026 decommissioning; fallbacks need version 2.
Clean core. Routing and fallbacks live in your side-by-side app on BTP. S/4HANA is only read through released APIs.
Done whenrouting-decision.md holds results from at least six runs, names a route that meets 90% on new traffic with the month-end mix, and names a fallback chain that was evaluated, with its accuracy during the outage.
Pick one answer for each question. The explanation appears after you choose.
1Why does the topic call an evaluation set the only way to "plot the map" of models and routes?
Answer: B. Each model or route is a point of accuracy, cost and p95. Only scoring all of them on the same labelled cases puts the points on one map, so you can pick the cheapest that meets the target.
2In Step 4, the 0.95 cascade costs almost as much as the large model alone. Why?
Answer: D. A higher threshold keeps fewer small-model answers, so 79% of requests go on to the large model. Each of those also paid for the small call, so the total approaches the cost of the large model alone.
3The rules router scored 94% on the evaluation set and 88.8% on new traffic. What is the lesson?
Answer: B. The rules were chosen from how hard orders looked in the evaluation set. In new traffic some hard orders look simple and go to the small model. Only a check on fresh data exposed it.
4Why does the cascade's p95 exceed the large model's p95 in the replay?
Answer: C. For an escalated order, the user waits for both calls in sequence. Those orders form the slow tail, so the cascade's p95 is the sum of two calls rather than one.
5Your orchestration fallback list is set up, but during an outage every request still takes over 8 seconds. What would you do?
Answer: D. Each request waits out the primary's timeout before orchestration tries the next model. A breaker in your app stops calling the failing model after repeated failures, so requests go straight to the fallback. It isn't part of the documented orchestration fallback.
6Which errors make SAP's orchestration (v2) move to the next configuration for a non-streaming request?
Answer: A. SAP documents the region case for all requests and, for non-streaming requests, 408, 429 and any 5xx. Other errors fail the request and are returned to you.
7In SAP AI Core, how are rate limits for generative models set by default?
Answer: C. SAP documents per-model requests-per-minute limits shared by all versions and all resource groups by default, with optional resource-group limits and a quota request API.
8A fallback configuration copied from SAP's example has no filtering module. What is the risk?
Answer: D. Each entry in the list is a full configuration, so modules can differ between preferences. If the fallback omits filtering or masking, requests that fall back run without it.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Orchestration with Fallbacks (SAP AI Core documentation, SAP-docs on GitHub)— modules as a list of configurations in preference order; orchestration V2 only; switches when the model isn't supported in the region (streaming and non-streaming) and, for non-streaming requests, on 408, 429 and 5xx; other errors fail the request; intermediate_failures in a successful response; error list in preference order when all fail; configurations may differ per preference
Rate Limit Management (SAP AI Core documentation, SAP-docs on GitHub)— requests-per-minute limits per model per tenant, shared by all versions of a model; optional resource-group limits; 429 can mean a rate limit or temporary back-end saturation; Retry-After header; exponential backoff with jitter; use model fallbacks, cache responses, isolate workloads; GET /v2/admin/quota/model and POST /v2/admin/quota/requests
What's New for SAP AI Core (SAP-docs on GitHub, read 7 October 2026)— timeout and max tries on model configurations (2025-09-01); quota management API (2025-10-06); orchestration fallbacks (2025-11-17); Evaluations added to Optimizations (2025-12-08); batch consumption for foundation models (2026-05-08); first orchestration endpoint to be decommissioned on 31 October 2026; frequent new-model entries pointing to SAP Note 3437766
SAP Cloud SDK for AI, generative AI package (sap-ai-sdk-gen 7.4.1 on PyPI)— installed and read in this run: orchestration_v2 OrchestrationConfig.modules accepts one ModuleConfig or a list tried in order; LLMModelDetails name, version (default latest), params, timeout 1 to 600 (default 600), max_retries 0 to 5 (default 2), timeout and retries currently ignored for Vertex AI models; CompletionPostResponse.intermediate_failures; OrchestrationService.run_with_retries retries 429 and server errors with exponential backoff and jitter, honouring Retry-After
RouteLLM: Learning to Route LLMs with Preference Data (Ong et al.; arXiv 2406.18665)— routers choose between a strong and a weak model per query with a cost threshold; trained on Chatbot Arena preference data, with augmentation; over 2x cost reduction in some cases without lowering quality; routers kept working on model pairs not seen in training
Choosing the right model (Claude Platform documentation, Anthropic)— criteria capabilities, speed and cost; two starting approaches, a fast low-cost model first or the most capable model first; an effort parameter trades intelligence for latency and cost within one model
CircuitBreaker (Martin Fowler, 6 March 2014)— wrap a remote call in a breaker that monitors failures; after a threshold it trips and fails fast; after a reset timeout it tries a trial call (half-open)