Orchestrate

Model selection and routing

Choose a model per task from your own evaluation data, send easy requests to a small model, and keep answering when a model fails.

Updated Oct 7, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

Most AI applications send every request to one model. That is simple, but you pay the price of your hardest request on every easy one.

Model selection means choosing the model for each task from evidence: you score a few models on your own labelled examples, then pick the cheapest one that meets your quality and speed targets.

Routing goes one step further and chooses per request. Easy requests go to a small, cheap, fast model. Hard ones go to a large model. A router can decide from what the request looks like, or let the small model try first and pass the request on when it is unsure. That second pattern is called a cascade.

Fallbacks keep the service running when a model is overloaded or down: the request moves to the next model on a list.

All three rest on one thing: an evaluation set, a few hundred real examples with the right answer, that tells you whether a cheaper path is still good enough.

Why it matters to the business

Take the running example of this unit: an assistant that reads each blocked sales order and routes it to the right team: credit, pricing, master data or export.

Most blocked orders are routine. One clear block reason, no surprises. A small model handles them well. A few are messy: two block reasons at once, or a note from the sales rep that changes the picture. Those need a stronger model.

In this topic's hands-on replay, with made-up prices and models:

Approach Correct Cost per month (6,000 orders)
Small model for everything 74% 0.86 USD
Large model for everything 97% 43.20 USD
Cascade: small first, large when unsure 93% 14.86 USD

The cascade costs about a third of the large model and is right on 93 in 100 orders. Whether that is good enough is a business decision. A misrouted order costs a clerk a few minutes and may delay a delivery. If the team needs 97%, the answer is the large model. The point is that the choice is made with numbers, not by habit.

Two more things change once you route:

  • Resilience. SAP's documentation says model calls can be refused with a "too many requests" error when a rate limit or the provider's capacity is reached. A fallback model keeps orders flowing. In the replay, a 150-request outage of the large model failed every one of those requests until a fallback was added.
  • Hidden quality drops. A fallback or router that was never tested can quietly lower quality. In the replay, falling back to the small model kept the service up but cut accuracy during the outage to 77%.

Published research points the same way. The FrugalGPT study (Chen, Zaharia and Zou, 2023) found that provider prices could differ by two orders of magnitude, and that a cascade matched the strongest model on their test tasks at a fraction of its cost. Your savings depend on your data, so measure them.

How SAP does it

As of October 2026, from the SAP sources opened for this topic:

  • Choosing. The generative AI hub in SAP AI Core lists the models your account can use, with costs and lifecycle dates. In SAP AI Launchpad, the Model Library adds a leaderboard and a chart of benchmark scores. Benchmarks help you shortlist; they don't score your blocked orders.
  • Testing. Evaluations, added in December 2025, benchmarks models and prompts as orchestration configurations on your own dataset, with predefined or custom metrics. This is where an SAP team can compare candidate models on its evaluation set.
  • Fallbacks. The orchestration service (version 2) accepts a list of configurations in order of preference. If the first model isn't available in the region, or fails with certain temporary errors, orchestration tries the next. The response says which model answered and what failed along the way.
  • Rate limits. Limits are set per model, in requests per minute. Teams can check them and request increases through an API or the model card.
  • Routing per request. No managed router that picks a model per request was found in the SAP sources opened for this topic. Your team builds that logic in its application, and sends each request to orchestration with the chosen model name.

The first orchestration endpoint is scheduled for decommissioning on 31 October 2026. Fallbacks need version 2, so a migration is due anyway.

SAP-delivered AI, such as Joule, chooses its own models. This topic applies to AI you build.

Three ways to put models to work

Pattern How it decides Good for Watch out for
One model per task Evaluation picks the cheapest model that meets the target Most teams, as the first step Paying large-model prices on easy requests
Rules router Visible features: number of block reasons, a note, order value Cases where "hard" is easy to spot Rules that looked right in testing and miss hard cases in production
Cascade A small model answers; low confidence sends it to a large model Mostly easy traffic with a hard tail Confidently wrong small answers; slower escalated requests
Fallback chain Next model on the list when one fails or is overloaded Every production system An untested fallback lowering quality without anyone noticing

A practical order: pick one model per task with an evaluation first. Add a tested fallback before go-live. Add routing only when volume makes the saving worth the extra testing.

Questions to ask

  • Which models did you score, on how many of our own examples, and how many of them were hard cases?
  • What accuracy and response time are we targeting, and who agreed them with the business?
  • For a router or cascade: what share of requests goes to each model, and what does that cost per month?
  • How was the routing threshold chosen, and was it confirmed on data the team didn't tune it on?
  • What happens when the main model is rate-limited or down? Has the fallback model passed the same evaluation?
  • How will we notice if traffic moves to the fallback or to the small model and quality drops?
  • When a new model or version appears in the catalog, how quickly can we re-run the evaluation and switch?

Common misconceptions

  • "The best model is the one at the top of the leaderboard." The best model is the cheapest one that meets your target on your data.
  • "Routing always saves money." In a cascade you pay for the small call even when you escalate. If most requests escalate, the cascade can cost more than the large model alone.
  • "A cascade is as fast as the small model." Escalated requests wait for both models. In the replay, the slowest 5% of cascade requests waited longer than with the large model alone.
  • "A fallback is free insurance." A fallback that was never evaluated keeps the service up while quietly answering worse.
  • "The model's confidence tells us when it is wrong." It helps, but some wrong answers come back confident. Test the threshold on labelled data.
  • "We chose the model once." New models arrive in SAP's catalog every few weeks and versions retire. Selection is a repeatable test, not a one-off decision.

Key terms

  • Model selection: choosing the model for a task from evaluation results against a quality, speed and cost target.
  • Evaluation set: real examples with known right answers, used to score models and routes.
  • Router: logic that chooses a model for each request.
  • Cascade: a router that asks a cheap model first and escalates to a stronger one when needed.
  • Confidence threshold: the confidence below which a cascade escalates.
  • Escalation rate: the share of requests passed to the stronger model.
  • Fallback: the next model tried when one fails or is overloaded.
  • Rate limit: the most requests per minute a model accepts for your account.
  • Circuit breaker: a switch that stops calling a failing model for a while, so users don't wait on it.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1The team proposes the large model for blocked-order triage "because it scores best". What should you ask first?

    Answer: C. Model selection picks the cheapest model that meets an agreed target on your own examples. "Scores best" says nothing about whether a cheaper model would also be good enough for this task.
  2. 2In the replay, the cascade is right on 93% of orders at about a third of the large model's cost. Who decides whether 93% is enough?

    Answer: B. The target is a business trade-off: what a wrong route costs order-to-cash against what accuracy costs in model fees. Engineers measure the options; the process owner sets the bar.
  3. 3A vendor says routing "always saves money". When can a cascade cost more than the large model alone?

    Answer: D. In a cascade every escalated request pays for the small call and the large call. If the threshold sends most traffic onward, the total can pass the cost of using the large model for everything.
  4. 4The large model is overloaded for an hour. What keeps blocked orders flowing without a hidden quality drop?

    Answer: A. A fallback keeps requests answered, but only a tested fallback keeps quality known. In the replay, falling back to the small model, which misses the target, cut accuracy during the outage to 77%.
  5. 5As of October 2026, what does SAP's generative AI hub offer for routing between models?

    Answer: C. Orchestration version 2 takes a list of configurations as fallbacks. No managed per-request router was found in the SAP sources for this topic, so the routing logic lives in your application.
  6. 6Why should the team re-run its evaluation regularly after choosing a model?

    Answer: C. SAP's catalog adds models every few weeks and versions have retirement dates. A ready evaluation lets the team switch in hours instead of rebuilding its case.
Deep layer · 40 min read

Mental model

A model choice is a point on a quality-cost-latency map, and an evaluation set is the only way to plot it.

Each model is one point: its accuracy on your cases, its cost per request, its p95 wait. A router doesn't add a new model. It adds new points, mixes of models, and some of them sit where no single model can: almost the quality of the large model at a fraction of the cost.

Everything else follows from that:

  • Selection is reading the map: the cheapest point that meets your target.
  • Routing is drawing new points, and every new point must be measured, because routes fail in ways single models don't.
  • Fallbacks are the points you land on when your chosen one is unavailable. They belong on the map too, measured in advance.
flowchart LR
  E[Evaluation set<br/>labelled cases] --> S[Score each model<br/>and each route]
  S --> P[Pick cheapest that<br/>meets the target]
  P --> R[Router + fallback chain<br/>in configuration]
  R --> L[Logs: model used,<br/>escalations, failures]
  L -->|new model, drift,<br/>outage| S

How it works

Step 1: a target, then a scoreboard

A choice needs a target first, agreed with the process owner: "route at least 90% of blocked orders correctly, with a p95 under 6 seconds". Defining success metrics with the business shows how to agree one.

Then score every candidate on the same evaluation set, with the same prompt. Record three numbers per model:

Measure How Why
Accuracy (or your task metric) Compare answers with the labels Does it meet the target?
Cost per 1,000 requests Tokens in and out times the price What does it cost at volume? (Token economics)
p50 and p95 latency Time every call Does it fit the latency budget? (Latency budgets)

Anthropic's model guide describes two ways to start: with a fast, low-cost model, moving up only if it falls short, or with the most capable model, then testing whether cheaper ones keep up. Both work if you measure. The guide also notes that some models take an effort setting that trades intelligence for latency and cost within one model, so "which model" can become "which model at which setting".

A small evaluation set is a trap. Hard cases are rare by definition. In this topic's 150-case set, only 19 are hard, so one more mistake on them moves the hard-case accuracy by five points. Make sure the hard tail is represented, and keep adding production cases that went wrong. Building an evaluation harness has the tooling.

Step 2: routing patterns

When no single model is cheap enough at the quality you need, route.

flowchart TD
  Q[Blocked order] --> RT{Router}
  RT -->|rules: 1 signal, no note| S[Small model]
  RT -->|rules: note| M[Mid model]
  RT -->|rules: 2+ signals| L[Large model]
  Q --> C1[Cascade: small model first]
  C1 -->|confident enough| A[Answer]
  C1 -->|unsure or<br/>invalid output| L2[Large model] --> A

Rules routers decide before any model is called, from features you can see: the number of block reasons, whether a sales rep left a note, the order value or the customer's risk class. They are cheap, fast and easy to explain. Their weakness: "looks simple" is not "is simple". In the replay, one in four hard orders shows only one signal, and the rules send it to the small model.

Cascades let the cheap model try first and escalate when its answer is doubtful. FrugalGPT (Chen, Zaharia and Zou, 2023) studied this pattern, scoring each answer and stopping at the first model whose answer passes. Two cheap escalation signals are common:

  1. Invalid output. The answer isn't one of the allowed queues, or the JSON doesn't match the schema. Structured outputs make this check exact.
  2. Low confidence. The model returns a confidence with its answer, and the cascade escalates below a threshold. Treat that number as a signal to test, not as a probability: in the replay, one in five wrong small-model answers still reports high confidence.

A cascade has two costs people forget. You pay for the small call on every escalated request, and the user waits for both calls.

Learned routers predict, per request, whether the small model's answer will be good enough. RouteLLM (Ong and others, 2024) trained routers on human preference data from Chatbot Arena and reported cost reductions of more than two times in some cases without lowering quality. Its routers kept working when routing between model pairs not seen in training. A learned router is itself a model that needs evaluation and monitoring, so most enterprise teams start with rules or a cascade.

Routing is not only between models. The same patterns can choose a prompt, a reasoning effort setting, or whether to answer from a cache first (Caching for LLM applications).

Step 3: fallbacks, retries and circuit breakers

Failures are routine at volume. SAP documents that SAP AI Core limits requests per minute per model, and that a 429 Too Many Requests can mean either your limit or temporary back-end saturation at the provider. Three mechanisms handle failures, and they fit together:

Mechanism What it does Use when Cost
Retry with backoff Waits and tries the same model again; honours Retry-After Short spikes, a single 429 Extra wait; can add load if done without jitter
Fallback Tries the next model on the list The model is down, overloaded or not in the region Quality of the fallback; needs evaluation
Circuit breaker After repeated failures, stops calling the model for a while, then tries one call An outage that lasts minutes A healthy model may be skipped briefly after recovery

Martin Fowler describes the circuit breaker as a wrapper that monitors failures, trips after a threshold so calls fail fast, and after a reset timeout lets a trial call through. For a model call, "fail fast" means going straight to the fallback instead of waiting for a timeout on every request.

Order matters. Retry a short spike. Fall back when retries would make the user wait too long. Trip the breaker when the same model keeps failing, so that thousands of requests don't each wait out a timeout.

Build it yourself: select, route and survive an outage

You will score three stand-in models on 150 labelled blocked orders, choose one, compare routing strategies, confirm the best on a month of new traffic, and then switch off the large model to see fallbacks and a circuit breaker at work. One script, model_router.py, does all of it with built-in Python.

Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. This script needs nothing else: no library, no account, no network.

flowchart LR
  EV[150 labelled orders] --> A[evaluate<br/>score 3 models]
  EV --> B[sweep<br/>compare routes]
  B --> C[route<br/>500 new orders]
  C --> D[fallback<br/>large model down]

What you need

  • Your course folder with .venv from the earlier setup topics.
  • About 45 minutes.
  • Cost: free. No account, no key, no network.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check your Python version (the same command on every system):

    python --version

    Any version from 3.10 up works.

Step 2: Create the script

  1. In VS Code, right-click the unit10 folder, choose New File, name it model_router.py, paste the code below and save.
"""Unit 10: choose a model per task with your own evaluation data, route easy requests to a
small model, and keep answering when a model fails.

The task is the blocked-orders triage from earlier units: read a blocked sales order and route
it to one queue (credit, pricing, master_data or export). Three stand-in "models" answer it.
They are functions with made-up accuracy, prices and response times, so everything runs
offline with built-in Python. No model is called; no account; no key.

  python unit10/model_router.py evaluate                     score each model on 150 labelled cases
  python unit10/model_router.py evaluate --target 0.95       pick the cheapest model that meets 95%
  python unit10/model_router.py route --strategy small       send all traffic to one model
  python unit10/model_router.py route --strategy rules       route by what the order looks like
  python unit10/model_router.py route --strategy cascade     small model first, escalate when unsure
  python unit10/model_router.py route --strategy cascade --threshold 0.9   replay new traffic
  python unit10/model_router.py sweep                        compare strategies on the evaluation set
  python unit10/model_router.py fallback                     an outage of the large model, no fallback
  python unit10/model_router.py fallback --chain large,mid --breaker
"""
import argparse
import random
import statistics
import sys

# ---------------------------------------------------------------------------------------------
# EXAMPLE models, prices and timings, made up for this course. Replace them with your own
# measurements. Prices are USD per million tokens; seconds are typical response times.
# ---------------------------------------------------------------------------------------------
MODELS = {
    "small": {"in": 0.10, "out": 0.40, "seconds": 0.6},
    "mid": {"in": 1.00, "out": 4.00, "seconds": 1.4},
    "large": {"in": 5.00, "out": 20.00, "seconds": 3.2},
}
# How often each stand-in model gets a case right, by how hard the case really is.
SKILL = {
    "small": {"easy": 0.97, "medium": 0.55, "hard": 0.20},
    "mid": {"easy": 0.99, "medium": 0.90, "hard": 0.55},
    "large": {"easy": 1.00, "medium": 0.97, "hard": 0.88},
}
TOKENS_IN, TOKENS_OUT = 1320, 30   # one triage call: order data in, a short JSON answer out
ORDERS_PER_MONTH = 6000            # the volume used in the token economics topic
QUEUES = ["credit", "pricing", "master_data", "export"]


# ---------------------------------------------------------------------------------------------
# Made-up, SAP-shaped blocked orders. Each case has what the app can SEE (the signals SAP
# shows and whether a sales rep left a note) and what it can't: how hard the case really is.
# ---------------------------------------------------------------------------------------------
def make_cases(count: int, seed: int) -> list:
    rng = random.Random(seed)
    cases = []
    for i in range(count):
        difficulty = rng.choices(["easy", "medium", "hard"], weights=[60, 28, 12])[0]
        label = rng.choice(QUEUES)
        if difficulty == "easy":      # one clear signal; sometimes an unimportant note
            signals, note = 1, rng.random() < 0.15
        elif difficulty == "medium":  # one signal, and a note that changes the answer
            signals, note = 1, rng.random() < 0.80
        else:                         # several signals, often a note; one in four looks simple
            signals, note = (1 if rng.random() < 0.25 else rng.choice([2, 3])), rng.random() < 0.6
        cases.append({"order": str(9000200 + i + seed * 1000), "label": label, "difficulty": difficulty,
                      "signals": signals, "note": note})
    return cases


EVAL_SET = make_cases(150, seed=1)  # labelled by people: your evaluation data
TRAFFIC = make_cases(500, seed=2)   # next month's live requests, labelled afterwards by reviewers


def call(model: str, case: dict) -> dict:
    """The stand-in model. Same model and same order always give the same result."""
    rng = random.Random(f"{model}|{case['order']}")
    correct = rng.random() < SKILL[model][case["difficulty"]]
    answer = case["label"] if correct else rng.choice([q for q in QUEUES if q != case["label"]])
    # Confidence the model reports with its answer. It is usually lower when the model is
    # wrong, but not always: some wrong answers come back confident.
    if correct:
        confidence = rng.uniform(0.85, 0.99) if case["difficulty"] == "easy" else rng.uniform(0.6, 0.97)
    else:
        confidence = rng.uniform(0.88, 0.97) if rng.random() < 0.2 else rng.uniform(0.4, 0.8)
    seconds = MODELS[model]["seconds"] * rng.uniform(0.7, 1.6)
    cost = (TOKENS_IN * MODELS[model]["in"] + TOKENS_OUT * MODELS[model]["out"]) / 1e6
    return {"model": model, "answer": answer, "confidence": round(confidence, 2),
            "seconds": seconds, "cost": cost}


def p95(values: list) -> float:
    ordered = sorted(values)
    return ordered[max(0, int(round(0.95 * len(ordered))) - 1)]


# ---------------------------------------------------------------------------------------------
# evaluate: score every model on the evaluation set, then choose
# ---------------------------------------------------------------------------------------------
def evaluate(args) -> None:
    print(f"Evaluation set: {len(EVAL_SET)} labelled blocked orders "
          f"({sum(c['difficulty'] == 'hard' for c in EVAL_SET)} hard)\n")
    print(f"{'model':<7}{'correct':>9}{'accuracy':>10}{'p50 s':>8}{'p95 s':>8}{'USD / 1,000':>13}")
    rows = []
    for model in MODELS:
        results = [call(model, c) for c in EVAL_SET]
        right = sum(r["answer"] == c["label"] for r, c in zip(results, EVAL_SET))
        seconds = [r["seconds"] for r in results]
        row = {"model": model, "accuracy": right / len(EVAL_SET), "p50": statistics.median(seconds),
               "p95": p95(seconds), "per_1000": 1000 * results[0]["cost"]}
        rows.append(row)
        print(f"{model:<7}{right:>6}/{len(EVAL_SET)}{row['accuracy']:>10.0%}{row['p50']:>8.2f}"
              f"{row['p95']:>8.2f}{row['per_1000']:>13.3f}")
    print(f"\nTarget: accuracy >= {args.target:.0%}, p95 <= {args.p95_budget:.1f} s")
    passing = [r for r in rows if r["accuracy"] >= args.target and r["p95"] <= args.p95_budget]
    if not passing:
        print("No single model meets both targets. Try routing, a better prompt, or relax a target.")
        return
    best = min(passing, key=lambda r: r["per_1000"])
    monthly = best["per_1000"] * ORDERS_PER_MONTH / 1000
    print(f"Cheapest model that passes: {best['model']} "
          f"({monthly:.2f} USD a month at {ORDERS_PER_MONTH:,} orders, EXAMPLE prices)")


# ---------------------------------------------------------------------------------------------
# route: three ways to decide which model answers each request
# ---------------------------------------------------------------------------------------------
def route_rules(case: dict) -> list:
    """Decide from what the order looks like, before any model is called."""
    if case["signals"] >= 2:
        return [call("large", case)]
    if case["note"]:
        return [call("mid", case)]
    return [call("small", case)]


def route_cascade(case: dict, threshold: float) -> list:
    """Ask the small model first; escalate to the large model when it is unsure."""
    first = call("small", case)
    if first["confidence"] >= threshold and first["answer"] in QUEUES:
        return [first]
    return [first, call("large", case)]


def run_traffic(strategy: str, threshold: float, cases: list) -> dict:
    stats = {"right": 0, "cost": 0.0, "seconds": [], "answered_by": {m: 0 for m in MODELS},
             "escalated": 0, "hard_wrong": 0}
    for case in cases:
        if strategy in MODELS:
            calls = [call(strategy, case)]
        elif strategy == "rules":
            calls = route_rules(case)
        else:
            calls = route_cascade(case, threshold)
        final = calls[-1]
        stats["right"] += final["answer"] == case["label"]
        stats["hard_wrong"] += final["answer"] != case["label"] and case["difficulty"] == "hard"
        stats["cost"] += sum(c["cost"] for c in calls)       # you pay for every call, also the first
        stats["seconds"].append(sum(c["seconds"] for c in calls))  # and wait for every call
        stats["answered_by"][final["model"]] += 1
        stats["escalated"] += len(calls) > 1
    return stats


def route(args) -> None:
    s = run_traffic(args.strategy, args.threshold, TRAFFIC)
    n = len(TRAFFIC)
    label = f"cascade, threshold {args.threshold}" if args.strategy == "cascade" else args.strategy
    print(f"Strategy: {label}   ({n} requests)\n")
    print(f"accuracy            {s['right'] / n:>6.1%}")
    shares = ", ".join(f"{m} {v / n:.0%}" for m, v in s["answered_by"].items() if v)
    print(f"answered by         {shares}")
    if args.strategy == "cascade":
        print(f"escalated           {s['escalated'] / n:>6.0%}")
    print(f"p50 / p95 wait      {statistics.median(s['seconds']):.2f} / {p95(s['seconds']):.2f} s")
    print(f"cost per 1,000      {1000 * s['cost'] / n:>6.3f} USD")
    print(f"cost per month      {ORDERS_PER_MONTH * s['cost'] / n:>6.2f} USD at {ORDERS_PER_MONTH:,} orders")
    print(f"wrong on hard cases {s['hard_wrong']:>6}")
    print("\nPrices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.")


def sweep(args) -> None:
    n = len(EVAL_SET)
    print(f"Evaluation set: {n} labelled blocked orders\n")
    print(f"{'strategy':<18}{'accuracy':>9}{'escalated':>11}{'p95 s':>8}{'USD / 1,000':>13}")
    rows = [("small", None), ("large", None), ("rules", None)] + [("cascade", t) for t in args.thresholds]
    for strategy, threshold in rows:
        s = run_traffic(strategy, threshold or 0.0, EVAL_SET)
        name = f"cascade {threshold}" if threshold else strategy
        esc = f"{s['escalated'] / n:.0%}" if threshold else "-"
        print(f"{name:<18}{s['right'] / n:>9.1%}{esc:>11}{p95(s['seconds']):>8.2f}{1000 * s['cost'] / n:>13.3f}")
    print("\nPick the cheapest row that meets your target, then confirm it on new traffic with route.")


# ---------------------------------------------------------------------------------------------
# fallback: requests 150 to 299 arrive while the large model times out
# ---------------------------------------------------------------------------------------------
TIMEOUT = 8.0          # seconds the app waits before it gives up on one call
OUTAGE = range(150, 300)
BREAKER_FAILURES = 5   # failures in a row that open the circuit breaker
BREAKER_SKIP = 40      # requests that skip the failing model before it is tried again


def fallback(args) -> None:
    chain = args.chain.split(",")
    for model in chain:
        if model not in MODELS:
            sys.exit(f"Unknown model '{model}'. Use names from: {', '.join(MODELS)}")
    stats = {"right": 0, "failed": 0, "seconds": [], "answered_by": {m: 0 for m in MODELS},
             "timeouts": 0, "outage_right": 0}
    in_a_row, skip_until = 0, -1
    for i, case in enumerate(TRAFFIC):
        waited, final = 0.0, None
        for model in chain:
            if args.breaker and model == chain[0] and i < skip_until:
                continue  # breaker open: don't even try the failing model
            if model == "large" and i in OUTAGE:
                waited += TIMEOUT  # the call hangs until our timeout
                stats["timeouts"] += 1
                if model == chain[0]:
                    in_a_row += 1
                    if args.breaker and in_a_row >= BREAKER_FAILURES:
                        skip_until, in_a_row = i + BREAKER_SKIP, 0  # open, then try again later
                continue
            if model == chain[0]:
                in_a_row = 0
            final = call(model, case)
            waited += final["seconds"]
            break
        stats["seconds"].append(waited)
        if final is None:
            stats["failed"] += 1
            continue
        stats["answered_by"][final["model"]] += 1
        ok = final["answer"] == case["label"]
        stats["right"] += ok
        stats["outage_right"] += ok and i in OUTAGE
    n = len(TRAFFIC)
    print(f"Chain: {' -> '.join(chain)}{'   circuit breaker on' if args.breaker else ''}")
    print(f"Outage: the large model times out for requests {OUTAGE.start}-{OUTAGE.stop - 1} "
          f"(timeout {TIMEOUT:.0f} s)\n")
    print(f"requests failed     {stats['failed']:>6}")
    shares = ", ".join(f"{m} {v}" for m, v in stats["answered_by"].items() if v)
    print(f"answered by         {shares}")
    print(f"timed-out calls     {stats['timeouts']:>6}")
    print(f"accuracy, all       {stats['right'] / n:>6.1%}")
    print(f"accuracy, outage    {stats['outage_right'] / len(OUTAGE):>6.1%}")
    print(f"p50 / p95 wait      {statistics.median(stats['seconds']):.2f} / {p95(stats['seconds']):.2f} s")


def main() -> None:
    parser = argparse.ArgumentParser(description="Model selection, routing and fallbacks, offline.")
    sub = parser.add_subparsers(dest="command", required=True)
    ev = sub.add_parser("evaluate", help="score each model on the evaluation set")
    ev.add_argument("--target", type=float, default=0.90, help="accuracy needed (0 to 1)")
    ev.add_argument("--p95-budget", type=float, default=6.0, help="slowest acceptable p95 in seconds")
    ro = sub.add_parser("route", help="replay the sample traffic with one routing strategy")
    ro.add_argument("--strategy", choices=list(MODELS) + ["rules", "cascade"], default="cascade")
    ro.add_argument("--threshold", type=float, default=0.8, help="cascade: confidence needed to keep the small answer")
    sw = sub.add_parser("sweep", help="compare strategies and cascade thresholds")
    sw.add_argument("--thresholds", type=float, nargs="+", default=[0.7, 0.8, 0.85, 0.9, 0.95])
    fb = sub.add_parser("fallback", help="replay traffic during an outage of the large model")
    fb.add_argument("--chain", default="large", help="models to try in order, e.g. large,mid")
    fb.add_argument("--breaker", action="store_true", help="skip the first model after repeated failures")
    args = parser.parse_args()
    for name in ("target", "threshold"):
        if hasattr(args, name) and not 0 < getattr(args, name) <= 1:
            sys.exit(f"--{name} must be above 0 and at most 1.")
    if hasattr(args, "thresholds") and not all(0 < t <= 1 for t in args.thresholds):
        sys.exit("--thresholds must each be above 0 and at most 1.")
    {"evaluate": evaluate, "route": route, "sweep": sweep, "fallback": fallback}[args.command](args)


if __name__ == "__main__":
    main()
  1. Check that the file is in the right place:

    • Windows (PowerShell):

      dir unit10\model_router.py
    • macOS / Linux:

      ls unit10/model_router.py

    You should see the file name. An error means the file is in another folder or has another name.

Step 3: Score the three models

  1. Run:

    python unit10/model_router.py evaluate
  2. You should see:

    Evaluation set: 150 labelled blocked orders (19 hard)
    
    model    correct  accuracy   p50 s   p95 s  USD / 1,000
    small     114/150       76%    0.66    0.94        0.144
    mid       131/150       87%    1.59    2.17        1.440
    large     149/150       99%    3.86    5.00        7.200
    
    Target: accuracy >= 90%, p95 <= 6.0 s
    Cheapest model that passes: large (43.20 USD a month at 6,000 orders, EXAMPLE prices)
  3. Read the table row by row. The small model is fast and costs 0.144 USD per 1,000 orders, but routes only 76% correctly. The large model gets 149 of 150 right and costs 50 times more. The mid model sits between them and misses the 90% target.

  4. Lower the target to see the choice change:

    python unit10/model_router.py evaluate --target 0.85
    Target: accuracy >= 85%, p95 <= 6.0 s
    Cheapest model that passes: mid (8.64 USD a month at 6,000 orders, EXAMPLE prices)

    The target decides the model. That is why the target is agreed with the business first.

Step 4: Compare routing strategies on the evaluation set

  1. Run:

    python unit10/model_router.py sweep
  2. You should see:

    Evaluation set: 150 labelled blocked orders
    
    strategy           accuracy  escalated   p95 s  USD / 1,000
    small                 76.0%          -    0.94        0.144
    large                 99.3%          -    5.00        7.200
    rules                 94.0%          -    4.07        1.260
    cascade 0.7           90.7%        20%    5.24        1.584
    cascade 0.8           95.3%        29%    5.46        2.208
    cascade 0.85          96.7%        32%    5.46        2.448
    cascade 0.9           96.7%        54%    5.64        4.032
    cascade 0.95          98.0%        79%    5.65        5.808
    
    Pick the cheapest row that meets your target, then confirm it on new traffic with route.
  3. What the rows mean:

    • rules reaches 94% at 1.26 USD per 1,000, the best-looking row.
    • cascade 0.8 reaches 95.3% at 2.21 USD, sending 29% of orders to the large model.
    • Higher thresholds escalate more and cost more. At 0.95, the cascade costs almost as much as the large model alone.
    • Every cascade row has a higher p95 than the large model alone, because escalated orders wait for two calls.
  4. Both rules and the 0.8 cascade pass the 90% target on this set. Before choosing, check them on data you didn't tune on.

Step 5: Confirm on new traffic

  1. Replay 500 new orders, labelled afterwards by reviewers, with the rules router:

    python unit10/model_router.py route --strategy rules
    Strategy: rules   (500 requests)
    
    accuracy             88.8%
    answered by         small 54%, mid 36%, large 11%
    p50 / p95 wait      0.92 / 3.85 s
    cost per 1,000       1.353 USD
    cost per month        8.12 USD at 6,000 orders
    wrong on hard cases     21
    
    Prices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.
  2. Accuracy fell from 94% on the evaluation set to 88.8% on new traffic, below the target. The rules were tuned on what hard orders look like, and some hard orders in the new month look simple.

  3. Now the cascade:

    python unit10/model_router.py route --strategy cascade --threshold 0.8
    Strategy: cascade, threshold 0.8   (500 requests)
    
    accuracy             92.6%
    answered by         small 68%, large 32%
    escalated              32%
    p50 / p95 wait      0.83 / 5.31 s
    cost per 1,000       2.477 USD
    cost per month       14.86 USD at 6,000 orders
    wrong on hard cases     16
    
    Prices and timings are EXAMPLES. Accuracy uses the reviewers' labels for this traffic.
  4. The cascade holds the target at 92.6%, for about a third of the large model's monthly cost. Its weakness is visible too: p95 rises above five seconds, and 16 hard orders are still wrong, mostly ones where the small model answered wrongly with high confidence.

  5. For comparison, run the large model alone:

    python unit10/model_router.py route --strategy large

    It reaches 97.0% at 43.20 USD a month. Write down the three options with their numbers. Choosing between them is the decision the process owner makes.

Step 6: Switch off the large model

The script simulates an outage: requests 150 to 299 reach a large model that hangs until the app's 8-second timeout.

  1. Run with no fallback:

    python unit10/model_router.py fallback
    Chain: large
    Outage: the large model times out for requests 150-299 (timeout 8 s)
    
    requests failed        150
    answered by         large 350
    timed-out calls        150
    accuracy, all        68.2%
    accuracy, outage      0.0%
    p50 / p95 wait      4.29 / 8.00 s

    150 requests failed. Every one of them also made the user wait the full 8 seconds.

  2. Add the small model as a fallback:

    python unit10/model_router.py fallback --chain large,small
    Chain: large -> small
    Outage: the large model times out for requests 150-299 (timeout 8 s)
    
    requests failed          0
    answered by         small 150, large 350
    timed-out calls        150
    accuracy, all        91.2%
    accuracy, outage     76.7%
    p50 / p95 wait      4.29 / 8.85 s

    Nothing failed, but accuracy during the outage dropped to 76.7%. Nobody would notice unless the logs show which model answered.

  3. Use the mid model instead:

    python unit10/model_router.py fallback --chain large,mid
    Chain: large -> mid
    Outage: the large model times out for requests 150-299 (timeout 8 s)
    
    requests failed          0
    answered by         mid 150, large 350
    timed-out calls        150
    accuracy, all        93.8%
    accuracy, outage     85.3%
    p50 / p95 wait      4.29 / 10.06 s

    Better quality during the outage, but p95 is now above 10 seconds: every request in the outage waits 8 seconds for the timeout before the fallback even starts.

  4. Add the circuit breaker:

    python unit10/model_router.py fallback --chain large,mid --breaker
    Chain: large -> mid   circuit breaker on
    Outage: the large model times out for requests 150-299 (timeout 8 s)
    
    requests failed          0
    answered by         mid 176, large 324
    timed-out calls         20
    accuracy, all        93.6%
    accuracy, outage     85.3%
    p50 / p95 wait      3.11 / 5.07 s

    Only 20 calls timed out instead of 150. After five timeouts in a row, the breaker skips the large model for 40 requests, then lets one trial call through. p95 is back near normal. The price: 26 requests after the outage still went to the mid model while the breaker was open.

Step 7: Save your work

  1. Save the script with Git:

    git add unit10/model_router.py
    git commit -m "Unit 10: model selection, routing and fallbacks"

How the code works

Part What it does
MODELS, SKILL Made-up price, speed and accuracy per model and per difficulty
make_cases Builds blocked orders with what the app can see (signals, note) and what it can't (difficulty); 12% are hard
EVAL_SET, TRAFFIC 150 labelled orders to choose with, and 500 new orders to confirm with
call The stand-in model: same model and order always give the same answer, confidence, time and cost
evaluate Scores each model and picks the cheapest that meets --target and --p95-budget
route_rules Picks a model from visible features, before any call
route_cascade Calls the small model; escalates on an invalid answer or confidence below the threshold
run_traffic Replays cases through a strategy, adding the cost and time of every call made
sweep Runs all strategies and several thresholds on the evaluation set
fallback Replays traffic while the large model times out, tries the chain in order, and opens a circuit breaker after 5 failures in a row

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
can't open file ... model_router.py You are not in the course folder, or the file has another name Run cd to orchestrate-course; check the file is in unit10
error: the following arguments are required: command No subcommand after the script name Add one: evaluate, sweep, route or fallback
error: argument --strategy: invalid choice A strategy name is misspelled Use small, mid, large, rules or cascade
Unknown model 'xl' A name in --chain doesn't exist Use names from MODELS, separated by commas with no spaces
SyntaxError or IndentationError Part of the code was not pasted, or the indentation changed Select all in the file, delete, and paste the whole block again
ModuleNotFoundError A line was changed to import a library this script doesn't use Paste the code again; the script needs only built-in Python
Your numbers differ from the ones shown You changed MODELS, SKILL or the case counts Expected after edits. Paste the code again to get the published numbers
A proxy or network error Not possible here: the script makes no network calls If you see one, you are running a different file

The SAP way

As of October 2026, from the SAP sources opened for this topic.

Discover and shortlist

  • Model catalog API. GET /v2/lm/scenarios/foundation-models/models lists models with cost figures, context length and lifecycle fields. Choosing and calling LLMs reads it in code.
  • Model Library in SAP AI Launchpad: catalog mode with filters, leaderboard mode ranking models by benchmark, and chart mode plotting two measures against each other. Each model card shows cost information, deprecation notices and your rate limits, with Increase Quota.
  • Keep watching. SAP's What's New page for SAP AI Core added new models in most months of 2026, each pointing to SAP Note 3437766 for conversion rates, rate limits and deprecation dates. A ready evaluation set turns each new model into a one-hour test.

Evaluate on your data

Evaluations (added 8 December 2025) benchmarks models and prompts as orchestration configurations on your own dataset, with predefined metrics or custom LLM-as-a-judge metrics. For model selection, that means one configuration per candidate model, the same prompt, and your labelled blocked orders. Building an evaluation harness covers how it fits with your own release gate.

Fallbacks in orchestration

In orchestration version 2, modules can be a list of module configurations, tried in order. SAP documents when orchestration moves on:

  • all requests: the model isn't supported in the deployed region or environment;
  • non-streaming requests only: 408 Request Timeout, 429 Too Many Requests, or any 5xx server error.

Any other error fails the request. A successful response carries intermediate_failures with the skipped attempts; if every configuration fails, the error lists each attempt in order. Each entry is a full configuration, so the fallback can carry its own prompt, parameters and modules. SAP's own example drops the filtering module in the fallback. Check that your fallback keeps the filtering and masking your process needs.

In the Python SDK (sap-ai-sdk-gen 7.4.1, read for this topic), OrchestrationConfig.modules accepts one ModuleConfig or a list. This is a sketch: it needs SAP AI Core with the generative AI hub, set up in Set up for Unit 5, and the model names are placeholders for models in your catalog.

# Sketch: needs SAP AI Core with the generative AI hub (Unit 5). Model names are placeholders.
from gen_ai_hub.orchestration_v2.models.config import ModuleConfig, OrchestrationConfig
from gen_ai_hub.orchestration_v2.models.llm_model_details import LLMModelDetails
from gen_ai_hub.orchestration_v2.models.message import SystemMessage, UserMessage
from gen_ai_hub.orchestration_v2.models.template import PromptTemplatingModuleConfig, Template
from gen_ai_hub.orchestration_v2.service import OrchestrationService

PROMPT = Template(template=[
    SystemMessage(content="Route the blocked sales order to one queue: credit, pricing, master_data or export."),
    UserMessage(content="{{?order}}"),
])

def module(model_name: str) -> ModuleConfig:
    # A short timeout per model so the fallback starts before the user gives up.
    return ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
        prompt=PROMPT, model=LLMModelDetails(name=model_name, timeout=8, max_retries=1)))

config = OrchestrationConfig(modules=[module("PRIMARY_MODEL"), module("TESTED_FALLBACK_MODEL")])
response = OrchestrationService().run(config=config, placeholder_values={"order": "..."})
print(response.final_result.model, response.intermediate_failures)

The SDK documents timeout from 1 to 600 seconds (default 600) and max_retries from 0 to 5 (default 2) per model, and notes both are currently ignored for Vertex AI models. Its run_with_retries method retries 429 and server errors with exponential backoff and jitter, and uses Retry-After when present.

Rate limits and quotas

SAP AI Core limits requests per minute per model, shared by all versions of that model, across the tenant by default. You can set separate limits per resource group, which lets you isolate a batch job from an interactive assistant. GET /v2/admin/quota/model shows your limits; POST /v2/admin/quota/requests asks for more and returns a requestId for an approval process. SAP lists model fallbacks, caching and workload isolation among the mitigations for 429s.

Routing changes your quota needs: a cascade sends every request to the small model and a share to the large one. Size each model's limit from the routing shares, plus headroom for fallback traffic during an outage.

Routing per request

No managed per-request router was found in the SAP AI Core sources opened for this topic. Build it in your application on SAP BTP: compute the route, then call orchestration with the chosen model name (or a fallback list headed by it). Because the model is a configuration value in orchestration, routing needs no extra deployments. An allow list on the orchestration deployment, described in Choosing and calling LLMs, keeps every route inside the approved models.

Batch work

For work that can wait, SAP added batch consumption for models called through foundation-models deployments on 8 May 2026, described as a way to reduce cost. An overnight re-triage of all open orders is a candidate. Check the Batch Consumption page for supported models before planning on it.

Build vs. SAP

Need Use Why
Shortlist models Model Library leaderboard and catalog API Quick view of options, costs and lifecycle
Score candidates on your data SAP Evaluations, or your own harness Configurations tested the way they will run
One model per task Orchestration with the model name in configuration Switch by changing a value, no redeployment
Survive outages and 429s Orchestration fallback list (v2) plus SDK retries Built in; response shows what failed
Avoid waiting out timeouts Your own circuit breaker in the app Not part of the documented fallback
Cheap-first routing Your own rules router or cascade on BTP No managed router found in SAP sources
Learned router Your own, trained on your logs Only when volume justifies a second model to evaluate
Joule and SAP-delivered AI Nothing to build SAP chooses the models

Production concerns

  • Authorizations and data. Every model on a route sees the prompt. Each must be approved for the data, including fallbacks. Enforce this with the orchestration allow list rather than trusting the router code. Keep masking and filtering on every fallback configuration.
  • Evaluation. Score each route, not just each model, and confirm on data the threshold wasn't tuned on. Re-run when a model, version, prompt, threshold or fallback changes. Keep hard cases in the set on purpose.
  • Observability. Log on every request the route taken, the model that answered, the small model's confidence, whether it escalated, and any intermediate_failures. Alert on the escalation rate and the fallback rate: both drifting up is an early warning. Observability for AI systems shows where these attributes go.
  • Cost. Count every call in a cascade, not just the one that answered. Recompute monthly cost from real routing shares, since traffic mix changes at month-end.
  • Latency. Escalated and fallback requests are your tail. Set per-model timeouts to fit the latency budget, not the 600-second default.
  • Quotas. Size rate limits per model from routing shares, and isolate batch and interactive workloads in separate resource groups.
  • Lifecycle. Pin or track versions for every model on every route. A retired fallback fails at the worst moment. Migrate off the first orchestration endpoint before its 31 October 2026 decommissioning; fallbacks need version 2.
  • Clean core. Routing and fallbacks live in your side-by-side app on BTP. S/4HANA is only read through released APIs.

Pitfalls

  • Choosing by leaderboard. General benchmarks don't measure your queues.
  • Tuning and confirming on the same data. The rules router looked best on the evaluation set and missed the target on new traffic.
  • Trusting confidence. Some wrong small-model answers are confident; the cascade can't catch them.
  • Counting only the final call. Cascades pay for the small call on every escalation.
  • Forgetting the tail. A cascade's p95 can be worse than the large model's alone.
  • An untested fallback. It keeps the lights on and lowers quality silently.
  • The 600-second default timeout. Users give up long before a fallback starts.
  • No breaker. Every request in an outage waits out the timeout.
  • A fallback without the same filters. Each fallback configuration needs its own masking and filtering.
  • Too few hard cases. A set with a handful of hard cases can't tell two routes apart where it matters.

Exercise

Write the routing decision for the blocked-orders triage. The AI business case topic later in this unit reuses its monthly cost and accuracy figures.

  1. Open unit10/model_router.py.

  2. Make hard orders more common, as at month-end: in make_cases, change weights=[60, 28, 12] to weights=[50, 28, 22]. Save.

  3. Run python unit10/model_router.py evaluate and python unit10/model_router.py sweep. Note which strategies still meet 90%.

  4. Run route on the new traffic for large, rules and cascade at the threshold you prefer. Note accuracy, p95 and monthly cost for each.

  5. Run fallback --chain large,mid --breaker and one other chain of your choice.

  6. Create unit10/routing-decision.md with:

    • a table of every run: strategy, accuracy, p95, cost per month;
    • the route you choose for normal days and for month-end, with one sentence why;
    • the fallback chain and the breaker settings;
    • three log fields and two alerts you would add in production;
    • one line on which numbers came from stand-ins and must be re-measured with real models.
  7. Change the weights back to [60, 28, 12], then commit:

    git add unit10/model_router.py unit10/routing-decision.md
    git commit -m "Unit 10: routing decision for blocked-order triage"

Done when routing-decision.md holds results from at least six runs, names a route that meets 90% on new traffic with the month-end mix, and names a fallback chain that was evaluated, with its accuracy during the outage.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the topic call an evaluation set the only way to "plot the map" of models and routes?

    Answer: B. Each model or route is a point of accuracy, cost and p95. Only scoring all of them on the same labelled cases puts the points on one map, so you can pick the cheapest that meets the target.
  2. 2In Step 4, the 0.95 cascade costs almost as much as the large model alone. Why?

    Answer: D. A higher threshold keeps fewer small-model answers, so 79% of requests go on to the large model. Each of those also paid for the small call, so the total approaches the cost of the large model alone.
  3. 3The rules router scored 94% on the evaluation set and 88.8% on new traffic. What is the lesson?

    Answer: B. The rules were chosen from how hard orders looked in the evaluation set. In new traffic some hard orders look simple and go to the small model. Only a check on fresh data exposed it.
  4. 4Why does the cascade's p95 exceed the large model's p95 in the replay?

    Answer: C. For an escalated order, the user waits for both calls in sequence. Those orders form the slow tail, so the cascade's p95 is the sum of two calls rather than one.
  5. 5Your orchestration fallback list is set up, but during an outage every request still takes over 8 seconds. What would you do?

    Answer: D. Each request waits out the primary's timeout before orchestration tries the next model. A breaker in your app stops calling the failing model after repeated failures, so requests go straight to the fallback. It isn't part of the documented orchestration fallback.
  6. 6Which errors make SAP's orchestration (v2) move to the next configuration for a non-streaming request?

    Answer: A. SAP documents the region case for all requests and, for non-streaming requests, 408, 429 and any 5xx. Other errors fail the request and are returned to you.
  7. 7In SAP AI Core, how are rate limits for generative models set by default?

    Answer: C. SAP documents per-model requests-per-minute limits shared by all versions and all resource groups by default, with optional resource-group limits and a quota request API.
  8. 8A fallback configuration copied from SAP's example has no filtering module. What is the risk?

    Answer: D. Each entry in the list is a full configuration, so modules can differ between preferences. If the fallback omits filtering or masking, requests that fall back run without it.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in