Orchestrate

Latency budgets and performance

Find where the seconds go in an AI request, set a latency budget per use case, and meet it with streaming, parallel calls, shorter answers and timeouts.

Updated Oct 7, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

Every AI request spends its time in a few places. The network carries the question. Your app reads business data, for example a sales order from SAP. It searches documents. Then the model reads the prompt and writes the answer, one word piece at a time. Writing is usually the slowest part.

A latency budget is the time a use case is allowed to take, agreed with the business, and split across those steps. "95 in 100 clerks see the first words within 3 seconds" is a budget. "It should be fast" is not.

Four design moves do most of the work:

  • Stream the answer, so people see words while the rest is written.
  • Run independent steps at the same time, such as reading SAP and searching notes.
  • Ask for shorter answers. Fewer words written means less waiting.
  • Put a time limit on slow steps, and have a plan for when one runs out.

Speed is measured with percentiles, not averages. The slowest 1 in 20 requests is what users remember.

Why it matters to the business

Take the running example: an assistant that explains blocked sales orders to order-to-cash clerks. A clerk works through a queue of 60 blocked orders a day. If each answer takes 9 seconds, that is 9 minutes of waiting a day per clerk. Worse, Jakob Nielsen's research says that at around 10 seconds people start thinking about other things. The clerk opens another window, and the assistant becomes "the slow thing".

Latency is a business decision for three reasons:

  • Adoption. A correct answer that arrives too late is not used. People go back to calling the credit team.
  • Trade-offs with quality and completeness. Every speed-up costs something. A shorter answer may leave out a detail. A time limit on the SAP call means some answers come without live order data. Someone with process knowledge must decide which is acceptable.
  • Cost. Many speed-ups also save money: fewer output tokens are both faster and cheaper. Others cost more, such as a bigger machine or sending a request twice to beat a slow server.

The budget also differs by use case. A clerk in a live conversation needs the first words fast. An overnight job that classifies 5,000 orders only needs to finish before the morning shift. Treating both the same wastes money on one and frustrates users of the other.

How SAP does it

As of October 2026, SAP's tools give builders the same levers that any AI platform does:

  • Streaming in the orchestration service. The generative AI hub's orchestration service (see SAP generative AI hub and orchestration) can stream answers. SAP's Python and JavaScript SDKs both offer a stream method, and the JavaScript SDK lets an app cancel a stream the user no longer needs.
  • Streaming and filtering together. When an output filter is on, it checks the streamed text in chunks before passing it on. That keeps unsafe text off the screen, but it adds a little delay before the first words appear.
  • Time limits and retries per model call. Each model in an orchestration configuration has a timeout and a number of retries. In the current Python SDK, the defaults are generous: up to 600 seconds and 2 retries. Those defaults suit a batch job, not a clerk waiting at a screen. Ask your team what they set.
  • Pushing text to the screen. SAP has shown how a CAP application can pass streamed text to a browser as it arrives.

For AI features SAP builds into its own applications, such as Joule, SAP runs the service and decides how it performs. Your latency budget applies to the apps and agents your team builds.

Budgets by use case

The table gives example starting points. They follow Nielsen's three limits: 0.1 second feels instant, 1 second keeps the user's flow of thought, and around 10 seconds loses their attention. Agree the real numbers with the process owner, and write them down.

Use case What the user is doing First visible result (p95) Complete (p95) Main lever
Inline suggestion in a Fiori form, such as a proposed block reason Typing, mid-task Under 1 s Under 1 s A small model or no model; short output
Chat assistant for blocked orders Waiting for an explanation Under 3 s Under 6 to 10 s Streaming; shorter answers
Agent that checks several systems before acting Watching progress A status line under 1 s Under 30 to 60 s Show each step; run steps in parallel
Overnight classification of open orders Not waiting Not relevant Before the morning shift Throughput and cost, not latency

Two lines in every budget: first visible result and complete. Streaming moves the first number. Only shorter answers, faster models or less work move the second.

Questions to ask

  • What is the budget for this use case, as a percentile, and who agreed it?
  • Do we measure the time to first visible words separately from the time to a complete answer?
  • Where does the time go today, step by step? Which step is the largest at the 95th percentile?
  • Which steps could run at the same time, and why don't they?
  • What happens when SAP or the model is slow? Is there a time limit, and what does the user see when it runs out?
  • What timeout and retry settings do our model calls use? Do they fit the budget?
  • If we stream, how does the output filter work with streaming, and could unsafe text reach the screen?
  • Would a shorter answer serve the clerk just as well?

Common misconceptions

  • "A bigger server makes the AI faster." Most of the time is the model writing tokens and the calls to other systems. Your app's own server is rarely the slow part.
  • "Shorter prompts are the key to speed." OpenAI's guidance says halving the prompt may gain only 1 to 5%, while halving the answer may save about half the time.
  • "Streaming makes it faster." Streaming changes when people see the first words, not when the answer is complete. It feels faster, which matters, but the total stays the same.
  • "The average response time is fine, so we're fine." An average hides the slow tail. Google's SRE book notes that users prefer a slightly slower system to one that varies a lot.
  • "A time limit is just a technical setting." When a time limit runs out, someone gets a partial answer or an error. What they see is a business decision.
  • "Retries make it reliable." Retries can turn one slow request into three slow ones. They need to fit inside the budget.

Key terms

  • Latency: the time from asking to getting a result.
  • Latency budget: the time a use case may take, as a percentile, split across its steps.
  • Time to first token (TTFT): how long until the model produces the first piece of its answer.
  • Streaming: sending the answer piece by piece while it is written.
  • Output tokens: the word pieces the model writes. Each one takes time, so they drive both latency and cost.
  • Percentile (p50, p95): p95 is the time that 95 in 100 requests beat. p50 is the median.
  • Tail latency: the slowest few requests, the ones above p95 or p99.
  • Timeout: the longest a step may run before the app gives up on it.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1The blocked-orders assistant takes 9 seconds to answer, and clerks have started switching windows while they wait. Which change shows them words soonest, without changing the answer?

    Answer: B. Streaming shows the first words as soon as they are written, so the wait feels much shorter even though the complete answer takes as long. A larger server rarely helps, because most of the time is spent writing tokens and calling other systems.
  2. 2Which budget statement can a team actually test and report on?

    Answer: D. A budget names a percentile, a time and what is measured. An average hides the slow tail that users remember, and "fast" or "as quick as possible" cannot be checked.
  3. 3Your team wants to speed up the assistant. According to the guidance in this topic, which change is likely to save the most time?

    Answer: C. Writing tokens is almost always the slowest step, so halving the answer can save close to half the time. Halving the prompt may gain only a few percent.
  4. 4To meet the budget, the team sets a 1-second limit on the SAP order read. What must the business decide?

    Answer: A. A time limit means some answers arrive without the live order data. The process owner must decide whether that is acceptable and how the answer says so, for example "order details unavailable, based on notes only".
  5. 5An overnight job classifies 5,000 open orders before the morning shift. How should its latency budget differ from the chat assistant's?

    Answer: D. Nobody waits on a single answer in a batch job, so time to first words is irrelevant. The budget is the whole run finishing before the shift, which shifts attention to throughput and cost.
  6. 6The model calls in your team's app use the SDK's default settings. What should a leader ask about them?

    Answer: B. In the current Python SDK, a model call may wait up to 600 seconds and retry twice by default. That suits a batch job, but a clerk at a screen needs limits that fit an interactive budget.
Deep layer · 40 min read

Mental model

A request is a path of steps, and the clerk waits for the longest chain through it, the critical path. A latency budget is a sum along that path. Every second you want back has to come from a step on it.

Two ideas make the rest click:

  1. There are two clocks. One stops when the first words appear. The other stops when the answer is complete. Streaming moves only the first. Writing fewer tokens, choosing a faster model, or removing steps moves both.
  2. Budgets live in the tail. The median request is rarely the problem. The 95th percentile is set by the occasional slow SAP call or long answer, so that is where you measure and where you design.
flowchart LR
  G[Gateway] --> S[SAP order read]
  G --> N[Notes search]
  S --> M1[Model: first token]
  N --> M1
  M1 --> M2[Model: write tokens]
  M2 --> F[Output filter]
  F --> U[Clerk sees text]

Reading SAP and searching notes don't depend on each other, so they can run side by side. The model can't start until both are done. After that, everything is a straight line.

How it works

Where the time goes

Stage What happens What drives it Typical lever
Gateway and network Request arrives; app signs in to the AI service; data travels Distance, connection reuse, token caching for sign-in Keep connections and access tokens alive
Business data OData call to SAP, such as reading a sales order SAP system load, query size, network to the system Select only needed fields; time limit; run in parallel
Retrieval Search notes or documents (see RAG fundamentals) Index size, reranking Smaller top-k; run in parallel
Model: time to first token The model reads the whole prompt and produces the first token Model size, prompt length, provider load Smaller model; shorter prompt helps a little
Model: generation The model writes each output token in turn Number of output tokens × time per token Shorter answers; max_tokens; faster model
Post-processing Output filter, masking, JSON validation Filter type, chunk size when streaming Stream with filtering; validate once

The Claude documentation separates baseline latency, how fast a model works in general, from time to first token (TTFT), the time until the first piece of output when streaming. OpenTelemetry's AI conventions have matching metrics: gen_ai.server.time_to_first_token and gen_ai.server.time_per_output_token from the server side, and gen_ai.client.operation.duration from the client. They are still marked incubating, so names can change.

A rough model for one model call:

model time ≈ time to first token + output tokens × time per output token

With the made-up numbers in this topic's script, a 220-token answer at 50 tokens per second spends about 4.4 seconds writing, after about 0.6 seconds to the first token. That is why OpenAI's latency guide calls token generation "almost always the highest latency step" and says cutting half the output tokens may cut about half the latency, while halving the prompt may gain only 1 to 5%.

Streaming: the first clock

Without streaming, the clerk sees nothing until the last token is written and checked. With streaming, the server sends each piece as soon as it exists, and the app shows it. The total time is the same. The time to first visible text drops to roughly TTFT plus the first chunk.

Streaming has costs:

  • Filters work on chunks. An output filter can't judge text it hasn't seen. SAP's orchestration service checks streamed output in chunks; in version 7.4.1 of SAP's Python SDK, chunk_size defaults to 100 and is described as the minimum number of characters per chunk that post-LLM modules work on. The first words therefore wait for the first 100 or so characters plus the check.
  • Structured output arrives in pieces. A JSON answer is not valid until it is complete. If the app needs the parsed fields, streaming doesn't help the user, but a status line does.
  • The UI needs a channel. The browser needs server-sent events or WebSockets. SAP's example uses CAP with WebSockets.

Parallel calls: shortening the path

Two independent steps in sequence cost a + b. In parallel they cost max(a, b). In Python, asyncio.gather runs awaitables concurrently and returns their results in order. Parallelism only helps for steps that don't need each other's output. The model call needs both the order and the notes, so it stays after them.

OpenAI's guide makes the same point with two shirts drying at once. It also lists the reverse move, making fewer requests: if two model calls run one after the other, combining them into one prompt removes a round trip.

Tail latency and fan-out

Averages hide the tail. Google's SRE book recommends percentiles because "a simple average can obscure these tail latencies", and notes that users prefer a slightly slower system to one with high variance.

The tail gets worse when one request waits for many calls. Dean and Barroso's example in The Tail at Scale: if 1 in 100 calls to a server is slow and a request must wait for 100 such calls, 63% of requests are slow. An agent that calls ten SAP APIs per question has the same problem at a smaller scale.

Two remedies from that paper:

  • Hedged requests. Send a second copy of a request after the first has taken longer than its 95th-percentile time, and use whichever answers first. The paper reports this adds only about 5% load while cutting the tail. For model calls, the second copy is paid too, so weigh the cost.
  • Time limits with a fallback. Give up on a slow, optional step and answer without it. This is what the script in this topic does with the SAP read.

Timeouts and retries are part of the budget

A timeout bounds the worst case. A retry adds a second attempt inside it. Worst case for one step is roughly:

worst case ≈ (retries + 1) × timeout

In SAP's Python SDK 7.4.1, a model in an orchestration configuration defaults to a timeout of 600 seconds and max_retries of 2. For a clerk at a screen, set them to fit the budget, and decide what the clerk sees when the limit is hit.

Measuring p50 and p95 properly

  • Measure where the user is. Server-side timing misses the network and the browser. Start the clock when the request leaves the client.
  • Record each stage as a span (see Observability for AI systems). A p95 for the whole request tells you that you are over budget; stage percentiles tell you why.
  • Use enough requests. With 5 requests, p95 is just the slowest one. With 60, it is the 57th slowest. Hundreds give stable numbers.
  • Don't add stage percentiles. The p95 of the whole request is not the sum of the stage p95s, because the slow requests in each stage are usually different ones. Measure end to end separately.
  • Warm and cold differ. The first call after a deployment or an idle period often includes sign-in and connection setup.

Build it yourself: find the slow step and fix it

You will run a stopwatch over 60 made-up requests of the blocked-orders assistant and see where each second goes. Then you will change the design one lever at a time: streaming, parallel calls, a time limit on SAP, and shorter answers. You will check each design against a budget of 3 seconds to first words and 6 seconds to a complete answer, both at p95. An optional last step times real streamed calls to a model in SAP AI Core.

Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. The main script uses only built-in Python. The optional Step 8 also needs the SAP AI Core trial and sap-ai-sdk-gen from Set up for Unit 5.

flowchart LR
  A[latency_budget.py<br/>60 made-up requests] --> T[Stage table<br/>p50 and p95]
  T --> B{Budget check}
  B -->|FAIL| O[Change one option]
  O --> A
  B -->|PASS| C[latency_runs.csv]

What you need

  • Your course folder with .venv from the earlier setup topics.
  • About 40 to 60 minutes.
  • Cost: free. No model is called and no account is needed, except in the optional Step 8, which uses your SAP AI Core trial (no charge during the trial; small per-request charges after it).

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check your Python version (the same on every system):

    python --version
    Python 3.14.4

    Any version from 3.10 up works.

Step 2: Create the script

The script plays the assistant's five stages for 60 requests: the gateway, the SAP order read, the notes search, the model and the output filter. About 1 SAP call in 20 is very slow, as real systems sometimes are. Each request has its own stopwatch. Times are sped up 20 times so the run takes a few seconds, then scaled back, so every number is what a clerk would feel.

  1. In VS Code, right-click unit10, choose New File, name it latency_budget.py, paste the code below and save.
"""Unit 10: where does the time go in one AI request, and does it fit the budget?

The blocked-orders assistant answers a clerk's question in five stages:

  gateway    your app receives the request and signs in to the AI service
  sap        read the sales order from SAP (an OData call)
  notes      search the team's notes for the block reason (retrieval)
  model      the model reads the prompt, then writes the answer token by token
  filter     the output filter checks the answer before the clerk sees it

With no options, the script plays 60 made-up requests (no model, no SAP system, nothing costs
money), measures each stage with a stopwatch, and checks the results against a latency budget.
Options change the design so you can see what each one buys:

    python unit10/latency_budget.py                       the first design: one step after another
    python unit10/latency_budget.py --parallel            read SAP and search notes at the same time
    python unit10/latency_budget.py --stream              show the answer as it is written
    python unit10/latency_budget.py --short               ask for a shorter answer (fewer output tokens)
    python unit10/latency_budget.py --sap-timeout 1.0     give up on a slow SAP call after 1 second
    python unit10/latency_budget.py --parallel --stream --sap-timeout 1.0 --short   all of them

    python unit10/latency_budget.py --llm --model MODEL_NAME   time real streamed calls to a model
                                                               in SAP AI Core (needs the Unit 5 setup)

Times are simulated, then sped up 20 times so a run takes a few seconds. Every number printed
is in "assistant seconds", the time a clerk would feel.
"""
import argparse
import asyncio
import csv
import os
import random
import sys
import time
from pathlib import Path

HERE = Path(__file__).parent
SPEED = 20  # run 20 times faster than real time; all reported numbers are scaled back

# The latency budget, agreed with the order-to-cash process owner (example values).
BUDGET = {
    "first_text_p95": 3.0,   # 95 in 100 clerks see the first words within 3 seconds
    "complete_p95": 6.0,     # 95 in 100 answers are complete within 6 seconds
}

# Made-up open orders, shaped like the SAP sales order fields used in earlier units.
ORDERS = [
    {"SalesOrder": str(9000001 + i), "DeliveryBlockReason": random.Random(i).choice(["01", "02", ""])}
    for i in range(60)
]

TOKENS_PER_SECOND = 50      # how fast the example model writes, after the first token
CHARS_PER_TOKEN = 4         # a rough rule for English text
FILTER_CHUNK_CHARS = 100    # the output filter checks text in chunks of at least this many characters


def percentile(values: list[float], pct: float) -> float:
    """Nearest-rank percentile: sort, then take the value at the pct position."""
    ordered = sorted(values)
    rank = max(1, round(pct / 100 * len(ordered)))
    return ordered[rank - 1]


async def wait(seconds: float) -> None:
    """Sleep for a simulated duration, sped up by SPEED."""
    await asyncio.sleep(seconds / SPEED)


class Stopwatch:
    """Measures elapsed assistant seconds since the request started."""

    def __init__(self):
        self.start = time.perf_counter()

    def now(self) -> float:
        return (time.perf_counter() - self.start) * SPEED


# ---------- the five stages (simulated) ----------

async def gateway(rng: random.Random) -> None:
    await wait(rng.uniform(0.05, 0.10))


async def read_sales_order(rng: random.Random, order: dict) -> dict:
    """Stand-in for GET .../A_SalesOrder('...'). Usually quick; about 1 call in 20 is very slow."""
    slow = rng.random() < 0.05
    await wait(rng.uniform(2.5, 4.0) if slow else rng.uniform(0.25, 0.55))
    return order


async def search_notes(rng: random.Random) -> list[str]:
    await wait(rng.uniform(0.2, 0.45))
    return ["note-17", "note-42"]


async def model_answer(rng: random.Random, output_tokens: int, stream: bool, clock: Stopwatch) -> tuple[float, int]:
    """The model reads the prompt (time to first token), then writes output_tokens tokens.

    Returns (when the clerk saw the first words, tokens written). Without streaming, the clerk
    sees nothing until the whole answer and the output filter are done.
    """
    await wait(rng.uniform(0.4, 0.8))                        # time to first token
    first_token_at = clock.now()
    tokens = max(10, int(rng.gauss(output_tokens, output_tokens * 0.15)))
    seconds_per_token = 1 / TOKENS_PER_SECOND
    if stream:
        # The output filter needs a chunk of FILTER_CHUNK_CHARS before it can pass text on.
        first_chunk_tokens = FILTER_CHUNK_CHARS // CHARS_PER_TOKEN
        await wait(first_chunk_tokens * seconds_per_token)
        await wait(0.05)                                     # filter checks the first chunk
        first_text_at = clock.now()
        await wait((tokens - first_chunk_tokens) * seconds_per_token)
        await wait(0.05)                                     # filter checks the last chunk
        return first_text_at, tokens
    await wait(tokens * seconds_per_token)
    await wait(0.15)                                         # filter checks the whole answer
    return clock.now(), tokens


# ---------- one request ----------

async def handle_request(i: int, args) -> dict:
    rng = random.Random(args.seed * 1000 + i)
    order = ORDERS[i % len(ORDERS)]
    clock = Stopwatch()
    marks = {}
    degraded = False

    await gateway(rng)
    marks["gateway"] = clock.now()

    async def sap_step():
        nonlocal degraded
        if args.sap_timeout:
            try:   # wait_for cancels the call if it takes longer than the timeout
                return await asyncio.wait_for(read_sales_order(rng, order), args.sap_timeout / SPEED)
            except asyncio.TimeoutError:
                degraded = True          # answer from the notes alone, and say so to the clerk
                return None
        return await read_sales_order(rng, order)

    async def timed(name, step):
        """Run one step and record how long it took on its own."""
        before = clock.now()
        result = await step
        marks[name] = clock.now() - before
        return result

    before = clock.now()
    if args.parallel:
        # Start both steps at once; wait until both are done.
        await asyncio.gather(timed("sap", sap_step()), timed("notes", search_notes(rng)))
    else:
        await timed("sap", sap_step())
        await timed("notes", search_notes(rng))
    sap_and_notes = clock.now() - before

    before_model = clock.now()
    output_tokens = 90 if args.short else 220
    first_text, tokens = await model_answer(rng, output_tokens, args.stream, clock)
    complete = clock.now()
    return {
        "request": i + 1,
        "sales_order": order["SalesOrder"],
        "gateway": marks["gateway"],
        "sap": marks["sap"],
        "notes": marks["notes"],
        "sap_and_notes": sap_and_notes,
        "model_and_filter": complete - before_model,
        "first_text": first_text,
        "complete": complete,
        "output_tokens": tokens,
        "degraded": degraded,
    }


async def run_all(args) -> list[dict]:
    # Requests are independent, so play them all at once; each keeps its own stopwatch.
    return await asyncio.gather(*(handle_request(i, args) for i in range(args.requests)))


def design_name(args) -> str:
    parts = [p for p, on in [("parallel", args.parallel), ("stream", args.stream), ("short", args.short),
                             (f"sap-timeout {args.sap_timeout}", bool(args.sap_timeout))] if on]
    return " + ".join(parts) if parts else "first design (one step after another)"


def report(rows: list[dict], args) -> int:
    print(f"Design: {design_name(args)}")
    print(f"Requests: {len(rows)}   (times in seconds, as the clerk feels them)\n")
    print(f"{'stage':<20}{'p50':>7}{'p95':>7}")
    stages = [("gateway", "gateway"), ("SAP order read", "sap"), ("notes search", "notes"),
              ("SAP + notes window", "sap_and_notes"), ("model + filter", "model_and_filter")]
    for label, key in stages:
        values = [r[key] for r in rows]
        print(f"{label:<20}{percentile(values, 50):>7.2f}{percentile(values, 95):>7.2f}")
    first = [r["first_text"] for r in rows]
    complete = [r["complete"] for r in rows]
    print(f"{'first words seen':<20}{percentile(first, 50):>7.2f}{percentile(first, 95):>7.2f}")
    print(f"{'answer complete':<20}{percentile(complete, 50):>7.2f}{percentile(complete, 95):>7.2f}")
    tokens = sum(r["output_tokens"] for r in rows) / len(rows)
    print(f"\nAverage output tokens per answer: {tokens:.0f}")
    degraded = sum(r["degraded"] for r in rows)
    if args.sap_timeout:
        print(f"Answers without live SAP data (SAP call gave up): {degraded} of {len(rows)}")

    print("\nBudget check")
    checks = [
        ("first words p95", percentile(first, 95), BUDGET["first_text_p95"]),
        ("complete answer p95", percentile(complete, 95), BUDGET["complete_p95"]),
    ]
    failed = 0
    for name, value, limit in checks:
        ok = value <= limit
        failed += not ok
        print(f"  {'PASS' if ok else 'FAIL'}  {name} {value:.2f} s (budget {limit:.1f} s)")
    return failed


def write_csv(rows: list[dict], path: Path, args) -> None:
    new = not path.exists()
    with open(path, "a", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        if new:
            writer.writerow(["design", "requests", "first_text_p95", "complete_p95", "avg_output_tokens",
                             "degraded"])
        first = [r["first_text"] for r in rows]
        complete = [r["complete"] for r in rows]
        writer.writerow([design_name(args), len(rows), f"{percentile(first, 95):.2f}",
                         f"{percentile(complete, 95):.2f}",
                         f"{sum(r['output_tokens'] for r in rows) / len(rows):.0f}",
                         sum(r["degraded"] for r in rows)])
    print(f"\nAdded one line to {path.name}")


# ---------- optional: time real streamed calls in SAP AI Core ----------

def time_real_model(args) -> None:
    """Stream a few real answers through SAP's orchestration service and time them."""
    from dotenv import load_dotenv
    load_dotenv()
    needed = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
              "AICORE_RESOURCE_GROUP"]
    missing = [n for n in needed if not os.environ.get(n)]
    if missing:
        sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5, "
                 "or leave out --llm to use the simulation.")
    if not args.model:
        sys.exit("Name a model with --model, for example one from 'python unit05/choose_model.py catalog'.")
    from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
                                             OrchestrationService, PromptTemplatingModuleConfig,
                                             SystemMessage, Template, UserMessage)
    from gen_ai_hub.orchestration_v2.models.streaming import GlobalStreamOptions

    words = 60 if args.short else 150
    template = Template(template=[
        SystemMessage(content="You explain blocked SAP sales orders to order-to-cash clerks in plain words."),
        UserMessage(content=f"In about {words} words, explain what delivery block reason {{{{?reason}}}} "
                            "usually means and what the clerk should check next.")])
    config = OrchestrationConfig(
        modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
            prompt=template, model=LLMModelDetails(name=args.model, timeout=30, max_retries=0))),
        stream=GlobalStreamOptions(enabled=True))
    try:
        service = OrchestrationService(config=config)
    except Exception as error:
        sys.exit(f"Could not connect to SAP AI Core ({type(error).__name__}). "
                 "Run python check_unit05.py to find the cause, or leave out --llm.")

    first_times, totals = [], []
    print(f"Model: {args.model}   streamed calls: {args.runs}\n")
    print(f"{'call':<6}{'first text':>12}{'complete':>10}{'chars':>8}")
    for n in range(1, args.runs + 1):
        start = time.perf_counter()
        first_at, text = None, ""
        try:
            for chunk in service.stream(placeholder_values={"reason": "01"}, timeout=30):
                result = chunk.final_result
                if result and result.choices and result.choices[0].delta.content:
                    if first_at is None:
                        first_at = time.perf_counter() - start
                    text += result.choices[0].delta.content
        except Exception as error:
            print(f"{n:<6}failed: {type(error).__name__}: {error}")
            continue
        total = time.perf_counter() - start
        first_times.append(first_at or total)
        totals.append(total)
        print(f"{n:<6}{first_at or total:>11.2f}s{total:>9.2f}s{len(text):>8}")
    if totals:
        print(f"\nfirst text   p50 {percentile(first_times, 50):.2f} s   p95 {percentile(first_times, 95):.2f} s")
        print(f"complete     p50 {percentile(totals, 50):.2f} s   p95 {percentile(totals, 95):.2f} s")
        print("With only a few calls, p95 is the slowest call. Use 20 or more for a real estimate.")


def main() -> None:
    parser = argparse.ArgumentParser(description="Measure where time goes in the blocked-orders assistant.")
    parser.add_argument("--parallel", action="store_true", help="read SAP and search notes at the same time")
    parser.add_argument("--stream", action="store_true", help="show the answer while it is written")
    parser.add_argument("--short", action="store_true", help="ask for a shorter answer")
    parser.add_argument("--sap-timeout", type=float, default=0.0,
                        help="give up on the SAP read after this many seconds (0 = wait forever)")
    parser.add_argument("--requests", type=int, default=60, help="how many requests to play (default 60)")
    parser.add_argument("--seed", type=int, default=8, help="change to play a different day")
    parser.add_argument("--csv", action="store_true", help="append the summary to unit10/latency_runs.csv")
    parser.add_argument("--llm", action="store_true", help="time real streamed calls in SAP AI Core")
    parser.add_argument("--model", default="", help="model name for --llm")
    parser.add_argument("--runs", type=int, default=5, help="number of real calls for --llm (default 5)")
    args = parser.parse_args()

    if args.llm:
        time_real_model(args)
        return
    rows = asyncio.run(run_all(args))
    failed = report(rows, args)
    if args.csv:
        write_csv(rows, HERE / "latency_runs.csv", args)
    if not failed:
        print("\nThe design fits the budget.")


if __name__ == "__main__":
    main()

Step 3: Measure the first design

The first design does one step after another and shows the answer only when it is complete.

  1. Run:

    python unit10/latency_budget.py
  2. You should see:

    Design: first design (one step after another)
    Requests: 60   (times in seconds, as the clerk feels them)
    
    stage                   p50    p95
    gateway                0.09   0.11
    SAP order read         0.40   2.74
    notes search           0.34   0.44
    SAP + notes window     0.74   2.97
    model + filter         5.35   6.21
    first words seen       6.13   8.83
    answer complete        6.13   8.83
    
    Average output tokens per answer: 225
    
    Budget check
      FAIL  first words p95 8.83 s (budget 3.0 s)
      FAIL  complete answer p95 8.83 s (budget 6.0 s)
  3. Read the table from the top. The SAP read's p50 is 0.40 s, but its p95 is 2.74 s: the occasional slow call. The model and filter take 5.35 s at p50, most of the request. Because nothing is streamed, the clerk sees the first words only when the answer is complete, so both lines are the same.

Your numbers may differ from these by a few hundredths of a second. The stopwatch is real, and computers vary.

Step 4: Stream the answer

  1. Run:

    python unit10/latency_budget.py --stream
  2. You should see:

    Design: stream
    Requests: 60   (times in seconds, as the clerk feels them)
    
    stage                   p50    p95
    gateway                0.08   0.11
    SAP order read         0.40   2.74
    notes search           0.35   0.45
    SAP + notes window     0.73   2.97
    model + filter         5.33   6.19
    first words seen       2.04   4.40
    answer complete        6.11   8.80
    
    Average output tokens per answer: 225
    
    Budget check
      FAIL  first words p95 4.40 s (budget 3.0 s)
      FAIL  complete answer p95 8.80 s (budget 6.0 s)
  3. First words dropped from 8.83 s to 4.40 s at p95. The complete answer didn't change: streaming moves only the first clock. The first words still fail the budget. Look at what remains before the model starts: the SAP read and the notes search, one after the other.

Step 5: Read SAP and search notes at the same time

  1. Run:

    python unit10/latency_budget.py --parallel --stream
  2. You should see:

    Design: parallel + stream
    Requests: 60   (times in seconds, as the clerk feels them)
    
    stage                   p50    p95
    gateway                0.09   0.11
    SAP order read         0.39   2.74
    notes search           0.34   0.44
    SAP + notes window     0.43   2.74
    model + filter         5.31   6.22
    first words seen       1.71   4.18
    answer complete        5.82   8.57
    
    Average output tokens per answer: 225
    
    Budget check
      FAIL  first words p95 4.18 s (budget 3.0 s)
      FAIL  complete answer p95 8.57 s (budget 6.0 s)
  3. The SAP + notes window is now as long as the slower of the two, 0.43 s at p50 instead of 0.73 s. The median clerk saves about 0.3 s. The p95 barely moved, because the tail is the slow SAP call, and running it in parallel doesn't make it faster.

Step 6: Put a time limit on the SAP call

If the order read takes longer than 1 second, the assistant gives up on it and answers from the notes alone. In a real app, the answer must say so.

  1. Run:

    python unit10/latency_budget.py --parallel --stream --sap-timeout 1.0
  2. You should see:

    Design: parallel + stream + sap-timeout 1.0
    Requests: 60   (times in seconds, as the clerk feels them)
    
    stage                   p50    p95
    gateway                0.08   0.11
    SAP order read         0.39   1.01
    notes search           0.34   0.44
    SAP + notes window     0.43   1.01
    model + filter         5.33   6.21
    first words seen       1.70   2.24
    answer complete        5.84   6.95
    
    Average output tokens per answer: 225
    Answers without live SAP data (SAP call gave up): 4 of 60
    
    Budget check
      PASS  first words p95 2.24 s (budget 3.0 s)
      FAIL  complete answer p95 6.95 s (budget 6.0 s)
  3. First words now pass the budget. The price is on the line above: 4 of 60 answers came without live SAP data. Whether that is acceptable is a question for the process owner, not for the code. The complete answer still fails, because the model writes about 225 tokens every time.

Step 7: Ask for shorter answers

  1. Run:

    python unit10/latency_budget.py --parallel --stream --sap-timeout 1.0 --short
  2. You should see:

    Design: parallel + stream + short + sap-timeout 1.0
    Requests: 60   (times in seconds, as the clerk feels them)
    
    stage                   p50    p95
    gateway                0.08   0.11
    SAP order read         0.40   1.00
    notes search           0.35   0.44
    SAP + notes window     0.43   1.01
    model + filter         2.64   2.98
    first words seen       1.70   2.25
    answer complete        3.17   3.76
    
    Average output tokens per answer: 92
    Answers without live SAP data (SAP call gave up): 4 of 60
    
    Budget check
      PASS  first words p95 2.25 s (budget 3.0 s)
      PASS  complete answer p95 3.76 s (budget 6.0 s)
    
    The design fits the budget.
  3. Fewer output tokens cut the model's time from 5.33 s to 2.64 s at p50, and the design now fits. In a real app, --short stands for a prompt that asks for about 60 words and a max_tokens limit. Check with your evaluation set from Building an evaluation harness that shorter answers are still good enough.

  4. Save the two designs side by side for later:

    python unit10/latency_budget.py --csv
    python unit10/latency_budget.py --parallel --stream --sap-timeout 1.0 --short --csv

    Each run ends with Added one line to latency_runs.csv.

Step 8 (optional): Time real streamed calls in SAP AI Core

This step calls a real model, so it needs your SAP AI Core trial and the AICORE_ lines in .env from Set up for Unit 5. During the trial there is no charge; after it, each call has a small per-request charge.

  1. Check that Unit 5 still works:

    python check_unit05.py
  2. Pick a model name your account offers. check_unit05.py prints some, or run python unit05/choose_model.py catalog from the model topic.

  3. Run five streamed calls, putting your model's name in place of MODEL_NAME:

    python unit10/latency_budget.py --llm --model MODEL_NAME
  4. You should see one line per call, then the percentiles. Your times depend on the model, region and load; this shape is what matters:

    Model: MODEL_NAME   streamed calls: 5
    
    call    first text  complete   chars
    1            0.93s     3.71s     812
    ...
    
    first text   p50 0.95 s   p95 1.20 s
    complete     p50 3.62 s   p95 4.10 s
    With only a few calls, p95 is the slowest call. Use 20 or more for a real estimate.
  5. Add --short and compare the complete times. Add --runs 20 for steadier percentiles.

The sample numbers above are illustrative. This step was not run against a live SAP AI Core account when this topic was written.

Step 9: Save your work

  1. The CSV is output, not code. Keep it out of Git by adding one line to .gitignore in the course folder:

    unit10/latency_runs.csv
  2. Save:

    git add .gitignore unit10/latency_budget.py
    git commit -m "Unit 10: latency budget for the blocked-orders assistant"

How the code works

Part What it does
BUDGET The two targets, in one place: first words and complete answer, both at p95
wait() and SPEED Simulated waits, 20 times faster than real; Stopwatch.now() scales the measured time back
read_sales_order, search_notes Stand-ins for the SAP OData read and the retrieval step; about 1 SAP call in 20 is slow
model_answer Time to first token, then one wait per output token. With --stream, the first text appears after the first filter chunk of 100 characters
asyncio.gather Runs the SAP read and notes search at the same time with --parallel
asyncio.wait_for Gives up on the SAP read after --sap-timeout seconds and marks the answer as degraded
timed() Records each stage's own duration, even when two run at once
percentile() Nearest-rank percentile: sort, then take the value at the 50% or 95% position
run_all Plays all requests at once with gather; each has its own stopwatch, so they don't affect each other
write_csv With --csv, appends one summary line per design to latency_runs.csv
time_real_model With --llm, streams real answers through SAP's orchestration service and times the first text and the end

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
can't open file ... latency_budget.py You are not in the course folder, or the file is named differently Run cd to orchestrate-course; check the file is in unit10
SyntaxError near the top of the file Part of the code was not pasted Select all in the file, delete, and paste the whole block again
Your numbers differ from the ones shown by a few hundredths The stopwatch is real; computers vary Expected. Large differences mean another program is busy; close it and rerun
Numbers differ a lot, or PASS/FAIL changes You changed --seed or --requests, or edited the code Run without those options to get the published numbers
Missing in .env: AICORE_... Step 8 can't find your SAP AI Core details Run unit05/key_to_env.py from the Unit 5 setup, or skip Step 8
ModuleNotFoundError: No module named 'gen_ai_hub' or 'dotenv' The Unit 5 libraries aren't in this Python Check for (.venv) in the prompt, then pip install -r requirements.txt
Could not connect to SAP AI Core Wrong key, expired trial, or no orchestration deployment Run python check_unit05.py, which names the cause
failed: ... on a call in Step 8 The model name is wrong for your account, or the network or proxy blocks the call Check the name in the catalog; try another network or ask IT to allow SAP AI Core

The SAP way

As of October 2026, the same levers exist in SAP's stack. What changes is where you set them.

Streaming through the orchestration service

SAP's Python SDK documentation for orchestration V2 shows streaming in three parts:

  • Turn it on in the configuration: OrchestrationConfig(modules=..., stream=GlobalStreamOptions(enabled=True)). The SDK refuses to mix them: a streaming configuration must be used with stream(), and a non-streaming one with run().
  • Call stream() (or astream() in async code) and read each chunk's text from chunk.final_result.choices[0].delta.content. Streamed chunks carry a delta instead of a message.
  • Tune chunking with GlobalStreamOptions(chunk_size=...), and give the output filter extra context with FilteringStreamOptions(overlap=...).

Step 8's time_real_model function follows this pattern. Tool call arguments can also arrive split across chunks, so an agent that streams must join them before parsing.

SAP's JavaScript SDK offers the same through stream() over server-sent events, with toContentStream() for the text. After the stream ends, getFinishReason() and getTokenUsage() return the finish reason and token counts. A stream can be cancelled with an AbortController, which is worth doing when a clerk closes the panel: generation you don't show is time and tokens you still pay for.

Getting the text to the user

The browser needs a channel for streamed text. An SAP Community post by an SAP employee shows a CAP backend that streams from the orchestration client and emits each chunk to the browser over WebSockets. Server-sent events are the other common choice. Either way, the CAP or Python service sits on BTP, not in S/4HANA, which keeps the core clean.

Timeouts and retries per model

In sap-ai-sdk-gen 7.4.1, each model in an orchestration configuration is described by LLMModelDetails, with:

Field Range Default What to set for a clerk at a screen
timeout 1 to 600 seconds 600 Inside the complete-answer budget
max_retries 0 to 5 2 Low, so that (retries + 1) × timeout still fits

stream() and run() also accept a timeout per request. The SDK notes that both model fields are currently ignored for Vertex AI models, so check what applies to the model you choose.

Measuring it

Record the stages as spans and the model call's timing as the OpenTelemetry metrics above, as shown in Observability for AI systems. The orchestration response's request_id links your measurement to SAP's side if you need to raise a ticket about a slow call.

Build vs. SAP

Situation Use Why
Learning the levers, comparing designs This topic's simulation Free, repeatable, shows each stage
A chat assistant on BTP using SAP-approved models Orchestration stream() with output filtering and chunk tuning Filters and masking still apply while streaming
A structured result the app needs in full, such as JSON fields run() without streaming, plus a progress message Partial JSON can't be used; status text keeps the user informed
A clerk-facing app with a CAP backend CAP plus WebSockets or server-sent events SAP has shown the CAP and WebSockets pattern
An overnight batch over thousands of orders Non-streaming calls, run concurrently within rate limits Throughput and cost matter; nobody watches
Joule and other SAP-built AI features SAP's own operations SAP runs them; your budget covers what you build

Production concerns

  • Streaming and safety. Text on screen can't be taken back. Keep output filtering on while streaming, and accept the chunk delay. Don't stream around a filter to save half a second.
  • Degraded answers must be visible. When a time limit skips the SAP read, the answer must say it was based on notes only. A silent partial answer looks like a wrong one.
  • SAP authorizations stay on the critical path. The OData read runs with the user's rights (see grounding on SAP data without breaking authorizations). Don't move to a faster technical user to save time; that trades a latency problem for a security one.
  • Retries, hedging and cost. Every retry or hedged model call is paid. Count them in your cost model, which the token economics topic later in this unit builds.
  • Rate limits. Running calls in parallel, or many users at once, raises the request rate. A rate-limit error costs more time than the parallel call saved.
  • Load and the tail. Latency grows with load. Measure p95 at expected peak, such as month-end close, not on a quiet afternoon.
  • Cancel what nobody reads. If the user closes the panel, cancel the stream and the pending SAP calls.
  • Budgets in the SLO. Put the two latency targets in the service level objectives from the observability topic, alert on them over a window, and split them by prompt and model version.
  • Clean core. All of this lives in the side-by-side app on BTP. Nothing in S/4HANA changes to make the assistant faster.

Pitfalls

  • Optimizing the prompt length first. The output is usually where the time is.
  • One number for "latency". Track first visible text and complete answer separately.
  • Averages in the budget. A good average can hide a painful tail.
  • Adding stage p95s to predict the total. The slow requests in each stage are usually different ones. Measure end to end.
  • Default timeouts. A 600-second limit with two retries is not a budget.
  • Parallelizing dependent steps. If the second step needs the first one's output, running them together only adds errors.
  • Measuring five calls. With five requests, p95 is just the slowest one.
  • Streaming JSON to the user. Half a JSON object is not an answer. Stream text, or show progress instead.

Exercise

Turn the time limit into a choice the process owner can make. Your results feed the token economics topic later in this unit, which puts a price on each design.

  1. Delete unit10/latency_runs.csv if it exists, so you start fresh.

  2. Run the full design with three different SAP time limits, saving each:

    python unit10/latency_budget.py --parallel --stream --short --sap-timeout 0.5 --csv
    python unit10/latency_budget.py --parallel --stream --short --sap-timeout 1.0 --csv
    python unit10/latency_budget.py --parallel --stream --short --sap-timeout 2.0 --csv
  3. Open unit10/latency_runs.csv in VS Code. For each row, note first_text_p95 and degraded.

  4. Run the same three commands with --requests 300 added, and compare. More requests give steadier percentiles.

  5. Create unit10/latency-budget.md with a three-row table: time limit, first-words p95, answers without SAP data out of 300. Below it, write two sentences: which time limit you would propose and why, and what the clerk should see when the limit is hit.

  6. Commit the note.

Done when latency-budget.md shows three time limits from the 300-request runs, each with its p95 and degraded count, and a recommendation that names the trade-off between speed and answers without live SAP data.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the first design, the SAP read's p50 is 0.40 s but its p95 is 2.74 s. What does that gap tell you?

    Answer: B. A low median with a high p95 means most calls are quick and a few are very slow. The budget is set at p95, so those few calls decide whether the design passes. That is why the time limit, not parallelism, fixed the first-words line.
  2. 2You add --stream and the complete-answer p95 doesn't change at all. Why?

    Answer: C. The model still writes the same number of tokens at the same speed. Streaming lets the clerk see the first chunk early, which moves the first-words clock only. To move the complete-answer clock you need fewer tokens, a faster model or less work.
  3. 3Running the SAP read and notes search with asyncio.gather saved about 0.3 s at p50 but barely changed p95. What explains this?

    Answer: A. Parallel steps cost max(a, b) instead of a + b. That helps the typical request, but when the SAP call is very slow, the max is still that slow call. Fixing the tail needs a time limit or a faster SAP query.
  4. 4Why does the streamed design show first words only after about 100 characters plus a check, rather than after the first token?

    Answer: D. An output filter can only judge text it has seen. In SAP's Python SDK 7.4.1, the streaming chunk size defaults to 100 characters for post-LLM modules, so the first words wait for that chunk and its check.
  5. 5A colleague sets timeout=600 and max_retries=2 on the model in an interactive assistant. What is the problem?

    Answer: B. The worst case is roughly (retries + 1) × timeout, which here is many minutes. These are the SDK's defaults, which suit batch work. An interactive budget needs a timeout and retry count that fit inside the complete-answer target.
  6. 6The blocked-orders assistant must return its answer as JSON fields the UI reads. What should you do about streaming?

    Answer: C. A JSON object isn't usable until it is complete, so streaming the raw text gives the clerk nothing useful. A status message covers the wait. Switching off the filter trades safety for speed, which the topic advises against.
  7. 7Two of the 60 requests miss the budget, and you want to know why. Which data helps most?

    Answer: C. Spans show which stage took the time in the exact requests that were slow. Averages and day totals hide them, and stage p95s can't be added, because different requests are slow in different stages.
  8. 8After a time limit on the SAP call, 4 of 60 answers come from notes only. What must happen before this goes live?

    Answer: D. A time limit trades completeness for speed, which is a business decision. The answer must tell the clerk it lacks live order data. A technical user would bypass the clerk's SAP authorizations, trading a latency problem for a security one.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in