Find where the seconds go in an AI request, set a latency budget per use case, and meet it with streaming, parallel calls, shorter answers and timeouts.
Every AI request spends its time in a few places. The network carries the question. Your app reads business data, for example a sales order from SAP. It searches documents. Then the model reads the prompt and writes the answer, one word piece at a time. Writing is usually the slowest part.
A latency budget is the time a use case is allowed to take, agreed with the business, and split across those steps. "95 in 100 clerks see the first words within 3 seconds" is a budget. "It should be fast" is not.
Four design moves do most of the work:
Stream the answer, so people see words while the rest is written.
Run independent steps at the same time, such as reading SAP and searching notes.
Ask for shorter answers. Fewer words written means less waiting.
Put a time limit on slow steps, and have a plan for when one runs out.
Speed is measured with percentiles, not averages. The slowest 1 in 20 requests is what users remember.
Take the running example: an assistant that explains blocked sales orders to order-to-cash clerks. A clerk works through a queue of 60 blocked orders a day. If each answer takes 9 seconds, that is 9 minutes of waiting a day per clerk. Worse, Jakob Nielsen's research says that at around 10 seconds people start thinking about other things. The clerk opens another window, and the assistant becomes "the slow thing".
Latency is a business decision for three reasons:
Adoption. A correct answer that arrives too late is not used. People go back to calling the credit team.
Trade-offs with quality and completeness. Every speed-up costs something. A shorter answer may leave out a detail. A time limit on the SAP call means some answers come without live order data. Someone with process knowledge must decide which is acceptable.
Cost. Many speed-ups also save money: fewer output tokens are both faster and cheaper. Others cost more, such as a bigger machine or sending a request twice to beat a slow server.
The budget also differs by use case. A clerk in a live conversation needs the first words fast. An overnight job that classifies 5,000 orders only needs to finish before the morning shift. Treating both the same wastes money on one and frustrates users of the other.
As of October 2026, SAP's tools give builders the same levers that any AI platform does:
Streaming in the orchestration service. The generative AI hub's orchestration service (see SAP generative AI hub and orchestration) can stream answers. SAP's Python and JavaScript SDKs both offer a stream method, and the JavaScript SDK lets an app cancel a stream the user no longer needs.
Streaming and filtering together. When an output filter is on, it checks the streamed text in chunks before passing it on. That keeps unsafe text off the screen, but it adds a little delay before the first words appear.
Time limits and retries per model call. Each model in an orchestration configuration has a timeout and a number of retries. In the current Python SDK, the defaults are generous: up to 600 seconds and 2 retries. Those defaults suit a batch job, not a clerk waiting at a screen. Ask your team what they set.
Pushing text to the screen. SAP has shown how a CAP application can pass streamed text to a browser as it arrives.
For AI features SAP builds into its own applications, such as Joule, SAP runs the service and decides how it performs. Your latency budget applies to the apps and agents your team builds.
The table gives example starting points. They follow Nielsen's three limits: 0.1 second feels instant, 1 second keeps the user's flow of thought, and around 10 seconds loses their attention. Agree the real numbers with the process owner, and write them down.
Use case
What the user is doing
First visible result (p95)
Complete (p95)
Main lever
Inline suggestion in a Fiori form, such as a proposed block reason
Typing, mid-task
Under 1 s
Under 1 s
A small model or no model; short output
Chat assistant for blocked orders
Waiting for an explanation
Under 3 s
Under 6 to 10 s
Streaming; shorter answers
Agent that checks several systems before acting
Watching progress
A status line under 1 s
Under 30 to 60 s
Show each step; run steps in parallel
Overnight classification of open orders
Not waiting
Not relevant
Before the morning shift
Throughput and cost, not latency
Two lines in every budget: first visible result and complete. Streaming moves the first number. Only shorter answers, faster models or less work move the second.
"A bigger server makes the AI faster." Most of the time is the model writing tokens and the calls to other systems. Your app's own server is rarely the slow part.
"Shorter prompts are the key to speed." OpenAI's guidance says halving the prompt may gain only 1 to 5%, while halving the answer may save about half the time.
"Streaming makes it faster." Streaming changes when people see the first words, not when the answer is complete. It feels faster, which matters, but the total stays the same.
"The average response time is fine, so we're fine." An average hides the slow tail. Google's SRE book notes that users prefer a slightly slower system to one that varies a lot.
"A time limit is just a technical setting." When a time limit runs out, someone gets a partial answer or an error. What they see is a business decision.
"Retries make it reliable." Retries can turn one slow request into three slow ones. They need to fit inside the budget.
Pick one answer for each question. The explanation appears after you choose.
1The blocked-orders assistant takes 9 seconds to answer, and clerks have started switching windows while they wait. Which change shows them words soonest, without changing the answer?
Answer: B. Streaming shows the first words as soon as they are written, so the wait feels much shorter even though the complete answer takes as long. A larger server rarely helps, because most of the time is spent writing tokens and calling other systems.
2Which budget statement can a team actually test and report on?
Answer: D. A budget names a percentile, a time and what is measured. An average hides the slow tail that users remember, and "fast" or "as quick as possible" cannot be checked.
3Your team wants to speed up the assistant. According to the guidance in this topic, which change is likely to save the most time?
Answer: C. Writing tokens is almost always the slowest step, so halving the answer can save close to half the time. Halving the prompt may gain only a few percent.
4To meet the budget, the team sets a 1-second limit on the SAP order read. What must the business decide?
Answer: A. A time limit means some answers arrive without the live order data. The process owner must decide whether that is acceptable and how the answer says so, for example "order details unavailable, based on notes only".
5An overnight job classifies 5,000 open orders before the morning shift. How should its latency budget differ from the chat assistant's?
Answer: D. Nobody waits on a single answer in a batch job, so time to first words is irrelevant. The budget is the whole run finishing before the shift, which shifts attention to throughput and cost.
6The model calls in your team's app use the SDK's default settings. What should a leader ask about them?
Answer: B. In the current Python SDK, a model call may wait up to 600 seconds and retry twice by default. That suits a batch job, but a clerk at a screen needs limits that fit an interactive budget.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A request is a path of steps, and the clerk waits for the longest chain through it, the critical path. A latency budget is a sum along that path. Every second you want back has to come from a step on it.
Two ideas make the rest click:
There are two clocks. One stops when the first words appear. The other stops when the answer is complete. Streaming moves only the first. Writing fewer tokens, choosing a faster model, or removing steps moves both.
Budgets live in the tail. The median request is rarely the problem. The 95th percentile is set by the occasional slow SAP call or long answer, so that is where you measure and where you design.
flowchart LR
G[Gateway] --> S[SAP order read]
G --> N[Notes search]
S --> M1[Model: first token]
N --> M1
M1 --> M2[Model: write tokens]
M2 --> F[Output filter]
F --> U[Clerk sees text]
Reading SAP and searching notes don't depend on each other, so they can run side by side. The model can't start until both are done. After that, everything is a straight line.
The model reads the whole prompt and produces the first token
Model size, prompt length, provider load
Smaller model; shorter prompt helps a little
Model: generation
The model writes each output token in turn
Number of output tokens × time per token
Shorter answers; max_tokens; faster model
Post-processing
Output filter, masking, JSON validation
Filter type, chunk size when streaming
Stream with filtering; validate once
The Claude documentation separates baseline latency, how fast a model works in general, from time to first token (TTFT), the time until the first piece of output when streaming. OpenTelemetry's AI conventions have matching metrics: gen_ai.server.time_to_first_token and gen_ai.server.time_per_output_token from the server side, and gen_ai.client.operation.duration from the client. They are still marked incubating, so names can change.
A rough model for one model call:
model time ≈ time to first token + output tokens × time per output token
With the made-up numbers in this topic's script, a 220-token answer at 50 tokens per second spends about 4.4 seconds writing, after about 0.6 seconds to the first token. That is why OpenAI's latency guide calls token generation "almost always the highest latency step" and says cutting half the output tokens may cut about half the latency, while halving the prompt may gain only 1 to 5%.
Without streaming, the clerk sees nothing until the last token is written and checked. With streaming, the server sends each piece as soon as it exists, and the app shows it. The total time is the same. The time to first visible text drops to roughly TTFT plus the first chunk.
Streaming has costs:
Filters work on chunks. An output filter can't judge text it hasn't seen. SAP's orchestration service checks streamed output in chunks; in version 7.4.1 of SAP's Python SDK, chunk_size defaults to 100 and is described as the minimum number of characters per chunk that post-LLM modules work on. The first words therefore wait for the first 100 or so characters plus the check.
Structured output arrives in pieces. A JSON answer is not valid until it is complete. If the app needs the parsed fields, streaming doesn't help the user, but a status line does.
The UI needs a channel. The browser needs server-sent events or WebSockets. SAP's example uses CAP with WebSockets.
Two independent steps in sequence cost a + b. In parallel they cost max(a, b). In Python, asyncio.gather runs awaitables concurrently and returns their results in order. Parallelism only helps for steps that don't need each other's output. The model call needs both the order and the notes, so it stays after them.
OpenAI's guide makes the same point with two shirts drying at once. It also lists the reverse move, making fewer requests: if two model calls run one after the other, combining them into one prompt removes a round trip.
Averages hide the tail. Google's SRE book recommends percentiles because "a simple average can obscure these tail latencies", and notes that users prefer a slightly slower system to one with high variance.
The tail gets worse when one request waits for many calls. Dean and Barroso's example in The Tail at Scale: if 1 in 100 calls to a server is slow and a request must wait for 100 such calls, 63% of requests are slow. An agent that calls ten SAP APIs per question has the same problem at a smaller scale.
Two remedies from that paper:
Hedged requests. Send a second copy of a request after the first has taken longer than its 95th-percentile time, and use whichever answers first. The paper reports this adds only about 5% load while cutting the tail. For model calls, the second copy is paid too, so weigh the cost.
Time limits with a fallback. Give up on a slow, optional step and answer without it. This is what the script in this topic does with the SAP read.
A timeout bounds the worst case. A retry adds a second attempt inside it. Worst case for one step is roughly:
worst case ≈ (retries + 1) × timeout
In SAP's Python SDK 7.4.1, a model in an orchestration configuration defaults to a timeout of 600 seconds and max_retries of 2. For a clerk at a screen, set them to fit the budget, and decide what the clerk sees when the limit is hit.
Measure where the user is. Server-side timing misses the network and the browser. Start the clock when the request leaves the client.
Record each stage as a span (see Observability for AI systems). A p95 for the whole request tells you that you are over budget; stage percentiles tell you why.
Use enough requests. With 5 requests, p95 is just the slowest one. With 60, it is the 57th slowest. Hundreds give stable numbers.
Don't add stage percentiles. The p95 of the whole request is not the sum of the stage p95s, because the slow requests in each stage are usually different ones. Measure end to end separately.
Warm and cold differ. The first call after a deployment or an idle period often includes sign-in and connection setup.
You will run a stopwatch over 60 made-up requests of the blocked-orders assistant and see where each second goes. Then you will change the design one lever at a time: streaming, parallel calls, a time limit on SAP, and shorter answers. You will check each design against a budget of 3 seconds to first words and 6 seconds to a complete answer, both at p95. An optional last step times real streamed calls to a model in SAP AI Core.
Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. The main script uses only built-in Python. The optional Step 8 also needs the SAP AI Core trial and sap-ai-sdk-gen from Set up for Unit 5.
flowchart LR
A[latency_budget.py<br/>60 made-up requests] --> T[Stage table<br/>p50 and p95]
T --> B{Budget check}
B -->|FAIL| O[Change one option]
O --> A
B -->|PASS| C[latency_runs.csv]
Your course folder with .venv from the earlier setup topics.
About 40 to 60 minutes.
Cost: free. No model is called and no account is needed, except in the optional Step 8, which uses your SAP AI Core trial (no charge during the trial; small per-request charges after it).
#Step 1: Open your course folder and turn on the virtual environment
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
If the prompt doesn't start with (.venv), turn it on:
Windows (PowerShell):
.venv\Scripts\Activate.ps1
macOS / Linux:
source .venv/bin/activate
Check your Python version (the same on every system):
The script plays the assistant's five stages for 60 requests: the gateway, the SAP order read, the notes search, the model and the output filter. About 1 SAP call in 20 is very slow, as real systems sometimes are. Each request has its own stopwatch. Times are sped up 20 times so the run takes a few seconds, then scaled back, so every number is what a clerk would feel.
In VS Code, right-click unit10, choose New File, name it latency_budget.py, paste the code below and save.
"""Unit 10: where does the time go in one AI request, and does it fit the budget?
The blocked-orders assistant answers a clerk's question in five stages:
gateway your app receives the request and signs in to the AI service
sap read the sales order from SAP (an OData call)
notes search the team's notes for the block reason (retrieval)
model the model reads the prompt, then writes the answer token by token
filter the output filter checks the answer before the clerk sees it
With no options, the script plays 60 made-up requests (no model, no SAP system, nothing costs
money), measures each stage with a stopwatch, and checks the results against a latency budget.
Options change the design so you can see what each one buys:
python unit10/latency_budget.py the first design: one step after another
python unit10/latency_budget.py --parallel read SAP and search notes at the same time
python unit10/latency_budget.py --stream show the answer as it is written
python unit10/latency_budget.py --short ask for a shorter answer (fewer output tokens)
python unit10/latency_budget.py --sap-timeout 1.0 give up on a slow SAP call after 1 second
python unit10/latency_budget.py --parallel --stream --sap-timeout 1.0 --short all of them
python unit10/latency_budget.py --llm --model MODEL_NAME time real streamed calls to a model
in SAP AI Core (needs the Unit 5 setup)
Times are simulated, then sped up 20 times so a run takes a few seconds. Every number printed
is in "assistant seconds", the time a clerk would feel.
"""
import argparse
import asyncio
import csv
import os
import random
import sys
import time
from pathlib import Path
HERE = Path(__file__).parent
SPEED = 20 # run 20 times faster than real time; all reported numbers are scaled back
# The latency budget, agreed with the order-to-cash process owner (example values).
BUDGET = {
"first_text_p95": 3.0, # 95 in 100 clerks see the first words within 3 seconds
"complete_p95": 6.0, # 95 in 100 answers are complete within 6 seconds
}
# Made-up open orders, shaped like the SAP sales order fields used in earlier units.
ORDERS = [
{"SalesOrder": str(9000001 + i), "DeliveryBlockReason": random.Random(i).choice(["01", "02", ""])}
for i in range(60)
]
TOKENS_PER_SECOND = 50 # how fast the example model writes, after the first token
CHARS_PER_TOKEN = 4 # a rough rule for English text
FILTER_CHUNK_CHARS = 100 # the output filter checks text in chunks of at least this many characters
def percentile(values: list[float], pct: float) -> float:
"""Nearest-rank percentile: sort, then take the value at the pct position."""
ordered = sorted(values)
rank = max(1, round(pct / 100 * len(ordered)))
return ordered[rank - 1]
async def wait(seconds: float) -> None:
"""Sleep for a simulated duration, sped up by SPEED."""
await asyncio.sleep(seconds / SPEED)
class Stopwatch:
"""Measures elapsed assistant seconds since the request started."""
def __init__(self):
self.start = time.perf_counter()
def now(self) -> float:
return (time.perf_counter() - self.start) * SPEED
# ---------- the five stages (simulated) ----------
async def gateway(rng: random.Random) -> None:
await wait(rng.uniform(0.05, 0.10))
async def read_sales_order(rng: random.Random, order: dict) -> dict:
"""Stand-in for GET .../A_SalesOrder('...'). Usually quick; about 1 call in 20 is very slow."""
slow = rng.random() < 0.05
await wait(rng.uniform(2.5, 4.0) if slow else rng.uniform(0.25, 0.55))
return order
async def search_notes(rng: random.Random) -> list[str]:
await wait(rng.uniform(0.2, 0.45))
return ["note-17", "note-42"]
async def model_answer(rng: random.Random, output_tokens: int, stream: bool, clock: Stopwatch) -> tuple[float, int]:
"""The model reads the prompt (time to first token), then writes output_tokens tokens.
Returns (when the clerk saw the first words, tokens written). Without streaming, the clerk
sees nothing until the whole answer and the output filter are done.
"""
await wait(rng.uniform(0.4, 0.8)) # time to first token
first_token_at = clock.now()
tokens = max(10, int(rng.gauss(output_tokens, output_tokens * 0.15)))
seconds_per_token = 1 / TOKENS_PER_SECOND
if stream:
# The output filter needs a chunk of FILTER_CHUNK_CHARS before it can pass text on.
first_chunk_tokens = FILTER_CHUNK_CHARS // CHARS_PER_TOKEN
await wait(first_chunk_tokens * seconds_per_token)
await wait(0.05) # filter checks the first chunk
first_text_at = clock.now()
await wait((tokens - first_chunk_tokens) * seconds_per_token)
await wait(0.05) # filter checks the last chunk
return first_text_at, tokens
await wait(tokens * seconds_per_token)
await wait(0.15) # filter checks the whole answer
return clock.now(), tokens
# ---------- one request ----------
async def handle_request(i: int, args) -> dict:
rng = random.Random(args.seed * 1000 + i)
order = ORDERS[i % len(ORDERS)]
clock = Stopwatch()
marks = {}
degraded = False
await gateway(rng)
marks["gateway"] = clock.now()
async def sap_step():
nonlocal degraded
if args.sap_timeout:
try: # wait_for cancels the call if it takes longer than the timeout
return await asyncio.wait_for(read_sales_order(rng, order), args.sap_timeout / SPEED)
except asyncio.TimeoutError:
degraded = True # answer from the notes alone, and say so to the clerk
return None
return await read_sales_order(rng, order)
async def timed(name, step):
"""Run one step and record how long it took on its own."""
before = clock.now()
result = await step
marks[name] = clock.now() - before
return result
before = clock.now()
if args.parallel:
# Start both steps at once; wait until both are done.
await asyncio.gather(timed("sap", sap_step()), timed("notes", search_notes(rng)))
else:
await timed("sap", sap_step())
await timed("notes", search_notes(rng))
sap_and_notes = clock.now() - before
before_model = clock.now()
output_tokens = 90 if args.short else 220
first_text, tokens = await model_answer(rng, output_tokens, args.stream, clock)
complete = clock.now()
return {
"request": i + 1,
"sales_order": order["SalesOrder"],
"gateway": marks["gateway"],
"sap": marks["sap"],
"notes": marks["notes"],
"sap_and_notes": sap_and_notes,
"model_and_filter": complete - before_model,
"first_text": first_text,
"complete": complete,
"output_tokens": tokens,
"degraded": degraded,
}
async def run_all(args) -> list[dict]:
# Requests are independent, so play them all at once; each keeps its own stopwatch.
return await asyncio.gather(*(handle_request(i, args) for i in range(args.requests)))
def design_name(args) -> str:
parts = [p for p, on in [("parallel", args.parallel), ("stream", args.stream), ("short", args.short),
(f"sap-timeout {args.sap_timeout}", bool(args.sap_timeout))] if on]
return " + ".join(parts) if parts else "first design (one step after another)"
def report(rows: list[dict], args) -> int:
print(f"Design: {design_name(args)}")
print(f"Requests: {len(rows)} (times in seconds, as the clerk feels them)\n")
print(f"{'stage':<20}{'p50':>7}{'p95':>7}")
stages = [("gateway", "gateway"), ("SAP order read", "sap"), ("notes search", "notes"),
("SAP + notes window", "sap_and_notes"), ("model + filter", "model_and_filter")]
for label, key in stages:
values = [r[key] for r in rows]
print(f"{label:<20}{percentile(values, 50):>7.2f}{percentile(values, 95):>7.2f}")
first = [r["first_text"] for r in rows]
complete = [r["complete"] for r in rows]
print(f"{'first words seen':<20}{percentile(first, 50):>7.2f}{percentile(first, 95):>7.2f}")
print(f"{'answer complete':<20}{percentile(complete, 50):>7.2f}{percentile(complete, 95):>7.2f}")
tokens = sum(r["output_tokens"] for r in rows) / len(rows)
print(f"\nAverage output tokens per answer: {tokens:.0f}")
degraded = sum(r["degraded"] for r in rows)
if args.sap_timeout:
print(f"Answers without live SAP data (SAP call gave up): {degraded} of {len(rows)}")
print("\nBudget check")
checks = [
("first words p95", percentile(first, 95), BUDGET["first_text_p95"]),
("complete answer p95", percentile(complete, 95), BUDGET["complete_p95"]),
]
failed = 0
for name, value, limit in checks:
ok = value <= limit
failed += not ok
print(f" {'PASS' if ok else 'FAIL'} {name} {value:.2f} s (budget {limit:.1f} s)")
return failed
def write_csv(rows: list[dict], path: Path, args) -> None:
new = not path.exists()
with open(path, "a", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
if new:
writer.writerow(["design", "requests", "first_text_p95", "complete_p95", "avg_output_tokens",
"degraded"])
first = [r["first_text"] for r in rows]
complete = [r["complete"] for r in rows]
writer.writerow([design_name(args), len(rows), f"{percentile(first, 95):.2f}",
f"{percentile(complete, 95):.2f}",
f"{sum(r['output_tokens'] for r in rows) / len(rows):.0f}",
sum(r["degraded"] for r in rows)])
print(f"\nAdded one line to {path.name}")
# ---------- optional: time real streamed calls in SAP AI Core ----------
def time_real_model(args) -> None:
"""Stream a few real answers through SAP's orchestration service and time them."""
from dotenv import load_dotenv
load_dotenv()
needed = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
"AICORE_RESOURCE_GROUP"]
missing = [n for n in needed if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5, "
"or leave out --llm to use the simulation.")
if not args.model:
sys.exit("Name a model with --model, for example one from 'python unit05/choose_model.py catalog'.")
from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
OrchestrationService, PromptTemplatingModuleConfig,
SystemMessage, Template, UserMessage)
from gen_ai_hub.orchestration_v2.models.streaming import GlobalStreamOptions
words = 60 if args.short else 150
template = Template(template=[
SystemMessage(content="You explain blocked SAP sales orders to order-to-cash clerks in plain words."),
UserMessage(content=f"In about {words} words, explain what delivery block reason {{{{?reason}}}} "
"usually means and what the clerk should check next.")])
config = OrchestrationConfig(
modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=args.model, timeout=30, max_retries=0))),
stream=GlobalStreamOptions(enabled=True))
try:
service = OrchestrationService(config=config)
except Exception as error:
sys.exit(f"Could not connect to SAP AI Core ({type(error).__name__}). "
"Run python check_unit05.py to find the cause, or leave out --llm.")
first_times, totals = [], []
print(f"Model: {args.model} streamed calls: {args.runs}\n")
print(f"{'call':<6}{'first text':>12}{'complete':>10}{'chars':>8}")
for n in range(1, args.runs + 1):
start = time.perf_counter()
first_at, text = None, ""
try:
for chunk in service.stream(placeholder_values={"reason": "01"}, timeout=30):
result = chunk.final_result
if result and result.choices and result.choices[0].delta.content:
if first_at is None:
first_at = time.perf_counter() - start
text += result.choices[0].delta.content
except Exception as error:
print(f"{n:<6}failed: {type(error).__name__}: {error}")
continue
total = time.perf_counter() - start
first_times.append(first_at or total)
totals.append(total)
print(f"{n:<6}{first_at or total:>11.2f}s{total:>9.2f}s{len(text):>8}")
if totals:
print(f"\nfirst text p50 {percentile(first_times, 50):.2f} s p95 {percentile(first_times, 95):.2f} s")
print(f"complete p50 {percentile(totals, 50):.2f} s p95 {percentile(totals, 95):.2f} s")
print("With only a few calls, p95 is the slowest call. Use 20 or more for a real estimate.")
def main() -> None:
parser = argparse.ArgumentParser(description="Measure where time goes in the blocked-orders assistant.")
parser.add_argument("--parallel", action="store_true", help="read SAP and search notes at the same time")
parser.add_argument("--stream", action="store_true", help="show the answer while it is written")
parser.add_argument("--short", action="store_true", help="ask for a shorter answer")
parser.add_argument("--sap-timeout", type=float, default=0.0,
help="give up on the SAP read after this many seconds (0 = wait forever)")
parser.add_argument("--requests", type=int, default=60, help="how many requests to play (default 60)")
parser.add_argument("--seed", type=int, default=8, help="change to play a different day")
parser.add_argument("--csv", action="store_true", help="append the summary to unit10/latency_runs.csv")
parser.add_argument("--llm", action="store_true", help="time real streamed calls in SAP AI Core")
parser.add_argument("--model", default="", help="model name for --llm")
parser.add_argument("--runs", type=int, default=5, help="number of real calls for --llm (default 5)")
args = parser.parse_args()
if args.llm:
time_real_model(args)
return
rows = asyncio.run(run_all(args))
failed = report(rows, args)
if args.csv:
write_csv(rows, HERE / "latency_runs.csv", args)
if not failed:
print("\nThe design fits the budget.")
if __name__ == "__main__":
main()
The first design does one step after another and shows the answer only when it is complete.
Run:
python unit10/latency_budget.py
You should see:
Design: first design (one step after another)
Requests: 60 (times in seconds, as the clerk feels them)
stage p50 p95
gateway 0.09 0.11
SAP order read 0.40 2.74
notes search 0.34 0.44
SAP + notes window 0.74 2.97
model + filter 5.35 6.21
first words seen 6.13 8.83
answer complete 6.13 8.83
Average output tokens per answer: 225
Budget check
FAIL first words p95 8.83 s (budget 3.0 s)
FAIL complete answer p95 8.83 s (budget 6.0 s)
Read the table from the top. The SAP read's p50 is 0.40 s, but its p95 is 2.74 s: the occasional slow call. The model and filter take 5.35 s at p50, most of the request. Because nothing is streamed, the clerk sees the first words only when the answer is complete, so both lines are the same.
Your numbers may differ from these by a few hundredths of a second. The stopwatch is real, and computers vary.
Design: stream
Requests: 60 (times in seconds, as the clerk feels them)
stage p50 p95
gateway 0.08 0.11
SAP order read 0.40 2.74
notes search 0.35 0.45
SAP + notes window 0.73 2.97
model + filter 5.33 6.19
first words seen 2.04 4.40
answer complete 6.11 8.80
Average output tokens per answer: 225
Budget check
FAIL first words p95 4.40 s (budget 3.0 s)
FAIL complete answer p95 8.80 s (budget 6.0 s)
First words dropped from 8.83 s to 4.40 s at p95. The complete answer didn't change: streaming moves only the first clock. The first words still fail the budget. Look at what remains before the model starts: the SAP read and the notes search, one after the other.
#Step 5: Read SAP and search notes at the same time
Design: parallel + stream
Requests: 60 (times in seconds, as the clerk feels them)
stage p50 p95
gateway 0.09 0.11
SAP order read 0.39 2.74
notes search 0.34 0.44
SAP + notes window 0.43 2.74
model + filter 5.31 6.22
first words seen 1.71 4.18
answer complete 5.82 8.57
Average output tokens per answer: 225
Budget check
FAIL first words p95 4.18 s (budget 3.0 s)
FAIL complete answer p95 8.57 s (budget 6.0 s)
The SAP + notes window is now as long as the slower of the two, 0.43 s at p50 instead of 0.73 s. The median clerk saves about 0.3 s. The p95 barely moved, because the tail is the slow SAP call, and running it in parallel doesn't make it faster.
Design: parallel + stream + sap-timeout 1.0
Requests: 60 (times in seconds, as the clerk feels them)
stage p50 p95
gateway 0.08 0.11
SAP order read 0.39 1.01
notes search 0.34 0.44
SAP + notes window 0.43 1.01
model + filter 5.33 6.21
first words seen 1.70 2.24
answer complete 5.84 6.95
Average output tokens per answer: 225
Answers without live SAP data (SAP call gave up): 4 of 60
Budget check
PASS first words p95 2.24 s (budget 3.0 s)
FAIL complete answer p95 6.95 s (budget 6.0 s)
First words now pass the budget. The price is on the line above: 4 of 60 answers came without live SAP data. Whether that is acceptable is a question for the process owner, not for the code. The complete answer still fails, because the model writes about 225 tokens every time.
Design: parallel + stream + short + sap-timeout 1.0
Requests: 60 (times in seconds, as the clerk feels them)
stage p50 p95
gateway 0.08 0.11
SAP order read 0.40 1.00
notes search 0.35 0.44
SAP + notes window 0.43 1.01
model + filter 2.64 2.98
first words seen 1.70 2.25
answer complete 3.17 3.76
Average output tokens per answer: 92
Answers without live SAP data (SAP call gave up): 4 of 60
Budget check
PASS first words p95 2.25 s (budget 3.0 s)
PASS complete answer p95 3.76 s (budget 6.0 s)
The design fits the budget.
Fewer output tokens cut the model's time from 5.33 s to 2.64 s at p50, and the design now fits. In a real app, --short stands for a prompt that asks for about 60 words and a max_tokens limit. Check with your evaluation set from Building an evaluation harness that shorter answers are still good enough.
Each run ends with Added one line to latency_runs.csv.
#Step 8 (optional): Time real streamed calls in SAP AI Core
This step calls a real model, so it needs your SAP AI Core trial and the AICORE_ lines in .env from Set up for Unit 5. During the trial there is no charge; after it, each call has a small per-request charge.
Check that Unit 5 still works:
python check_unit05.py
Pick a model name your account offers. check_unit05.py prints some, or run python unit05/choose_model.py catalog from the model topic.
Run five streamed calls, putting your model's name in place of MODEL_NAME:
You should see one line per call, then the percentiles. Your times depend on the model, region and load; this shape is what matters:
Model: MODEL_NAME streamed calls: 5
call first text complete chars
1 0.93s 3.71s 812
...
first text p50 0.95 s p95 1.20 s
complete p50 3.62 s p95 4.10 s
With only a few calls, p95 is the slowest call. Use 20 or more for a real estimate.
Add --short and compare the complete times. Add --runs 20 for steadier percentiles.
The sample numbers above are illustrative. This step was not run against a live SAP AI Core account when this topic was written.
SAP's Python SDK documentation for orchestration V2 shows streaming in three parts:
Turn it on in the configuration: OrchestrationConfig(modules=..., stream=GlobalStreamOptions(enabled=True)). The SDK refuses to mix them: a streaming configuration must be used with stream(), and a non-streaming one with run().
Call stream() (or astream() in async code) and read each chunk's text from chunk.final_result.choices[0].delta.content. Streamed chunks carry a delta instead of a message.
Tune chunking with GlobalStreamOptions(chunk_size=...), and give the output filter extra context with FilteringStreamOptions(overlap=...).
Step 8's time_real_model function follows this pattern. Tool call arguments can also arrive split across chunks, so an agent that streams must join them before parsing.
SAP's JavaScript SDK offers the same through stream() over server-sent events, with toContentStream() for the text. After the stream ends, getFinishReason() and getTokenUsage() return the finish reason and token counts. A stream can be cancelled with an AbortController, which is worth doing when a clerk closes the panel: generation you don't show is time and tokens you still pay for.
The browser needs a channel for streamed text. An SAP Community post by an SAP employee shows a CAP backend that streams from the orchestration client and emits each chunk to the browser over WebSockets. Server-sent events are the other common choice. Either way, the CAP or Python service sits on BTP, not in S/4HANA, which keeps the core clean.
In sap-ai-sdk-gen 7.4.1, each model in an orchestration configuration is described by LLMModelDetails, with:
Field
Range
Default
What to set for a clerk at a screen
timeout
1 to 600 seconds
600
Inside the complete-answer budget
max_retries
0 to 5
2
Low, so that (retries + 1) × timeout still fits
stream() and run() also accept a timeout per request. The SDK notes that both model fields are currently ignored for Vertex AI models, so check what applies to the model you choose.
Record the stages as spans and the model call's timing as the OpenTelemetry metrics above, as shown in Observability for AI systems. The orchestration response's request_id links your measurement to SAP's side if you need to raise a ticket about a slow call.
Streaming and safety. Text on screen can't be taken back. Keep output filtering on while streaming, and accept the chunk delay. Don't stream around a filter to save half a second.
Degraded answers must be visible. When a time limit skips the SAP read, the answer must say it was based on notes only. A silent partial answer looks like a wrong one.
SAP authorizations stay on the critical path. The OData read runs with the user's rights (see grounding on SAP data without breaking authorizations). Don't move to a faster technical user to save time; that trades a latency problem for a security one.
Retries, hedging and cost. Every retry or hedged model call is paid. Count them in your cost model, which the token economics topic later in this unit builds.
Rate limits. Running calls in parallel, or many users at once, raises the request rate. A rate-limit error costs more time than the parallel call saved.
Load and the tail. Latency grows with load. Measure p95 at expected peak, such as month-end close, not on a quiet afternoon.
Cancel what nobody reads. If the user closes the panel, cancel the stream and the pending SAP calls.
Budgets in the SLO. Put the two latency targets in the service level objectives from the observability topic, alert on them over a window, and split them by prompt and model version.
Clean core. All of this lives in the side-by-side app on BTP. Nothing in S/4HANA changes to make the assistant faster.
Turn the time limit into a choice the process owner can make. Your results feed the token economics topic later in this unit, which puts a price on each design.
Delete unit10/latency_runs.csv if it exists, so you start fresh.
Run the full design with three different SAP time limits, saving each:
Open unit10/latency_runs.csv in VS Code. For each row, note first_text_p95 and degraded.
Run the same three commands with --requests 300 added, and compare. More requests give steadier percentiles.
Create unit10/latency-budget.md with a three-row table: time limit, first-words p95, answers without SAP data out of 300. Below it, write two sentences: which time limit you would propose and why, and what the clerk should see when the limit is hit.
Commit the note.
Done whenlatency-budget.md shows three time limits from the 300-request runs, each with its p95 and degraded count, and a recommendation that names the trade-off between speed and answers without live SAP data.
Pick one answer for each question. The explanation appears after you choose.
1In the first design, the SAP read's p50 is 0.40 s but its p95 is 2.74 s. What does that gap tell you?
Answer: B. A low median with a high p95 means most calls are quick and a few are very slow. The budget is set at p95, so those few calls decide whether the design passes. That is why the time limit, not parallelism, fixed the first-words line.
2You add --stream and the complete-answer p95 doesn't change at all. Why?
Answer: C. The model still writes the same number of tokens at the same speed. Streaming lets the clerk see the first chunk early, which moves the first-words clock only. To move the complete-answer clock you need fewer tokens, a faster model or less work.
3Running the SAP read and notes search with asyncio.gather saved about 0.3 s at p50 but barely changed p95. What explains this?
Answer: A. Parallel steps cost max(a, b) instead of a + b. That helps the typical request, but when the SAP call is very slow, the max is still that slow call. Fixing the tail needs a time limit or a faster SAP query.
4Why does the streamed design show first words only after about 100 characters plus a check, rather than after the first token?
Answer: D. An output filter can only judge text it has seen. In SAP's Python SDK 7.4.1, the streaming chunk size defaults to 100 characters for post-LLM modules, so the first words wait for that chunk and its check.
5A colleague sets timeout=600 and max_retries=2 on the model in an interactive assistant. What is the problem?
Answer: B. The worst case is roughly (retries + 1) × timeout, which here is many minutes. These are the SDK's defaults, which suit batch work. An interactive budget needs a timeout and retry count that fit inside the complete-answer target.
6The blocked-orders assistant must return its answer as JSON fields the UI reads. What should you do about streaming?
Answer: C. A JSON object isn't usable until it is complete, so streaming the raw text gives the clerk nothing useful. A status message covers the wait. Switching off the filter trades safety for speed, which the topic advises against.
7Two of the 60 requests miss the budget, and you want to know why. Which data helps most?
Answer: C. Spans show which stage took the time in the exact requests that were slow. Averages and day totals hide them, and stage p95s can't be added, because different requests are slow in different stages.
8After a time limit on the SAP call, 4 of 60 answers come from notes only. What must happen before this goes live?
Answer: D. A time limit trades completeness for speed, which is a business decision. The answer must tell the clerk it lacks live order data. A technical user would bypass the clerk's SAP authorizations, trading a latency problem for a security one.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Latency optimization (OpenAI API documentation)— seven principles; generating tokens is almost always the highest-latency step; cutting 50% of output tokens may cut about 50% of latency, while cutting 50% of the prompt may give only 1 to 5%; make fewer requests; parallelize; stream; don't default to an LLM
Reducing latency (Claude documentation)— definitions of baseline latency and time to first token (TTFT); choose a faster model, shorten prompt and output, set max_tokens, stream responses to improve perceived responsiveness
The Tail at Scale, a summary (The Morning Paper, 2015)— fan-out to 100 servers where 1 in 100 calls is slow makes 63% of user requests slow; hedged requests sent after the 95th-percentile latency add about 5% load and shorten the tail
Orchestration Service V2 API (SAP Cloud SDK for AI, Python, help.sap.com)— stream() and astream(); GlobalStreamOptions(enabled=True) in OrchestrationConfig; chunk_size; FilteringStreamOptions(overlap=...) for output filtering; streamed chunks carry delta instead of message. Version 7.4.1 of sap-ai-sdk-gen was installed and read in this run: chunk_size defaults to 100, described as the minimum characters per chunk for post-LLM modules; LLMModelDetails timeout 1 to 600 s (default 600) and max_retries 0 to 5 (default 2); stream() accepts a per-request timeout