Record traces, metrics and user feedback for an AI assistant, set service level objectives, and catch the quiet failures where nothing errors but answers get worse.
An AI assistant can fail in two ways. The loud way: it errors, times out or goes down. Normal IT monitoring catches that. The quiet way: it keeps answering, quickly and without errors, but the answers get worse, or the cost per question doubles. Nothing turns red. Users just stop trusting it.
Observability means the system records enough about each request that your team can answer new questions about it, including ones nobody thought to ask in advance. For AI, it rests on four kinds of data:
Traces: the step-by-step story of each request. Which notes were searched, which SAP data was read, which model answered, how long each took.
Take the running example: an assistant that explains blocked sales orders to order-to-cash clerks. A developer changes the prompt to include each order's history. Every test still passes. In production, three things happen at once:
Tokens per question roughly double, so the monthly bill climbs.
A document cleanup the same week removed the notes for one block reason. Questions about it now get vague answers.
Clerks quietly go back to calling the credit team.
Error rate: unchanged. Response time: a little slower, within target. A dashboard built only on IT signals shows green all month. With quality signals, the team sees the share of grounded answers fall and thumbs-down rise on the day of the change, and can trace both to one prompt version.
Observability pays for itself in three ways:
Faster fixes. A trace shows which step failed, so a complaint becomes a ticket with a cause.
Cost control. Token use is visible per request and per change, before the invoice arrives.
Evidence for the business case. The same data shows adoption and quality, which is what a sponsor asks for at the next budget round.
The risk runs the other way too. Traces can hold prompts and answers, which can hold customer names, prices or employee data. What you record is a data protection decision, not just a technical one.
As of October 2026, three SAP services cover operations for apps you build on SAP BTP. SAP's learning material describes how they fit together:
SAP Cloud ALM is the central view. Support teams use it to spot problems across BTP apps early.
SAP Cloud Logging is for in-depth analysis. An SAP session from March 2025 describes it as logs, metrics and traces on OpenSearch, with OpenTelemetry ingestion, dashboards and alerts to a webhook or Slack. Development plans are for evaluation only; production needs a standard or large plan.
SAP Alert Notification service turns technical events, such as an app crash, into alerts, including in SAP Cloud ALM.
For the AI part of an app, two SAP developer tools matter:
SAP's orchestration service (in the generative AI hub) returns a request ID, the token counts, and the result of each step it ran, such as filtering or masking. Recording these in your traces links your view to SAP's.
The SAP Cloud SDK for Python includes a telemetry module that traces AI calls with OpenTelemetry, the open standard. It is versioned 0.x, so expect it to change.
AI features SAP builds into its own applications, such as Joule, are run by SAP. You observe the apps and agents your team builds.
A good AI dashboard has one row per question the business asks. If a row is missing, that kind of failure is invisible.
Question
Signal
Example target (SLO)
Catches
Is it up?
Error rate
At most 5 in 100 requests fail
Outages, SAP timeouts
Is it fast enough?
95th-percentile response time
95 in 100 answers within 3 seconds
Slow model, slow SAP calls
What does it cost?
Tokens and cost per request
Within the budget per 1,000 questions
Prompt changes, runaway agents
Is it right?
Share of answers grounded in sources
At least 85 in 100
Missing documents, retrieval bugs
Do users trust it?
Thumbs-down rate
At most 1 in 4 rated answers
Everything the other rows miss
Is it used?
Requests per day, repeat users
Agreed with the process owner
Silent abandonment
The last three rows are what make this AI observability rather than ordinary IT monitoring. Every number should also be split by version: of the prompt, the model and the app. Most quiet failures start with a change.
"If there are no errors, it's working." AI systems mostly fail quietly. A confident wrong answer returns status 200.
"Our IT monitoring covers it." It covers up, fast and errors. It does not cover cost per question or answer quality.
"Log everything, just in case." Recording every prompt and answer creates a store of sensitive data. Record facts by default; capture content only with a policy.
"Averages are enough." An average of one second can hide one user in twenty waiting ten. Track the 95th percentile.
"Observability is a tool you buy." The tool stores data. Deciding what to record, the targets and who acts on alerts is design work your team owns.
"User feedback is too sparse to matter." Only a minority click a thumb, but a sudden change in the thumbs-down rate is one of the earliest quality warnings.
Pick one answer for each question. The explanation appears after you choose.
1The blocked-orders assistant shows no errors and normal response times, but clerks have stopped using it. What is the most likely gap in monitoring?
Answer: B. AI systems often fail quietly: answers get worse while errors and speed look normal. Only quality signals, such as grounded-answer rate and thumbs-down rate, reveal this. A lower error threshold would not help, because there are no errors.
2A developer changed the prompt, and the monthly token bill rose sharply. Which practice would have shown the cause on the first day?
Answer: C. When each request records its prompt version and token counts, the dashboard can compare versions side by side, and the jump appears the day the change ships. The invoice arrives weeks later and doesn't say which change caused it.
3Why does the dashboard track the 95th-percentile response time and not only the average?
Answer: D. A good average can sit on top of a slow tail, such as one request in twenty timing out on an SAP call. The 95th percentile shows what the slowest users experience, which is what drives complaints.
4Your team wants to store every prompt and answer in traces "to debug faster". What should a leader ask first?
Answer: A. Prompts and answers can contain customer, pricing or employee data, so recording them is a data protection decision. The safe default is to record facts such as model, tokens and timing, and capture content only under an agreed policy.
5Which SAP service is described as the central view where support teams spot problems across BTP apps early?
Answer: C. SAP describes Cloud ALM as the central monitoring view, Cloud Logging as the place for in-depth root-cause analysis, and Alert Notification as the service that turns technical events into alerts. They work together rather than replacing each other.
6Only about 3 in 10 users click thumbs up or down. Is the feedback still useful?
Answer: B. Feedback is sparse, but its rate is stable when nothing changes, so a jump after a release is a strong signal. It works best alongside automatic checks such as grounded-answer rate, not instead of them.
7An alert fires every night at 2 a.m. for a few slow requests, and nobody ever does anything about it. What should change?
Answer: D. An alert nobody acts on trains people to ignore alerts, including the important ones. Targets and thresholds should reflect what users need, and each alert should lead to a clear action.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Treat every request as a unit of evidence. When it finishes, it should leave behind enough to answer three questions without rerunning it: what happened (a trace), how it compares (metrics), and was it any good (quality signals attached to it).
The second idea: AI failures are mostly quiet. Classic monitoring watches the four golden signals that Google's SRE book names: latency, traffic, errors and saturation. Keep them. Then add the signals that only AI systems need: tokens and cost per request, grounding, and user feedback. All of them are split by version, because most quiet failures begin with a change.
flowchart LR
R[One request] --> T[Trace<br/>spans per step]
R --> M[Metrics<br/>duration, tokens]
U[User thumb] -->|span ID| T
T --> D[Health report<br/>by version]
M --> D
D --> S{SLO broken?}
S -->|yes| A[Alert a person]
S -->|no| Q[Stay quiet]
Numbers aggregated in time buckets, such as a histogram of durations
Dashboards and alerts over all requests
Low; independent of traffic
Log
A timestamped event with fields
Detail inside a step, audit trails
Medium; grows with verbosity
They join on IDs. Every span carries a trace ID, and a well-set-up logger stamps the same trace ID on each log line written during the request. Feedback that arrives later attaches to a span ID. If your app calls SAP's orchestration service, record its request ID on your span too. Then a ticket to SAP support can name the exact call.
OpenTelemetry's generative AI conventions name the common facts. As of October 2026 they are in use and still under active development, and they now live in their own OpenTelemetry repository.
Fact
Attribute
Notes
Kind of step
gen_ai.operation.name
invoke_agent, chat, retrieval, execute_tool
Model asked
gen_ai.request.model
Also record the model that answered if it can differ
Prompt version, documents returned, grounded flag, order ID if policy allows
Two metrics from the same conventions are worth emitting from day one: gen_ai.client.operation.duration (a histogram in seconds) and gen_ai.client.token.usage (a histogram in tokens, split by gen_ai.token.type, input or output).
Content is opt-in. The OpenTelemetry project states that, by default, its AI conventions capture no prompt content or tool arguments, only metadata. Not every library follows that default. Traceloop's OpenLLMetry, for example, documents that it records prompts, completions and embeddings on spans unless you set TRACELOOP_TRACE_CONTENT=false. Check the default of every instrumentation library you install.
Errors and latency come for free. Quality needs deliberate design. Three sources, cheapest first:
Facts the app already knows. Did retrieval return any documents? Did the model's structured output validate? Did the agent hit its step limit? Record each as a span attribute.
User feedback. A thumb or a short reason, sent with the span ID so it attaches to the request. Phoenix, for example, accepts it through POST /v1/span_annotations, with an annotator kind of HUMAN, LLM or CODE.
Automatic evaluation of a sample. Run the groundedness and correctness checks from Unit 8 on a share of production traffic, and attach the scores the same way.
Averages hide the tail. The SRE book's example: a service with an average latency of 100 ms at 1,000 requests per second can easily have 1% of requests taking 5 seconds. So targets use percentiles: p95 within 3 seconds, not "average within 1 second".
An SLO is a target plus a window: "95 in 100 answers within 3 seconds, measured over a day". An alert fires when an SLO is broken, not when one request is slow. The SRE book's rule for alerts that wake a person is that every one must be actionable. A page that needs no action trains people to ignore pages.
At volume, keeping every trace is expensive. OpenTelemetry describes two approaches:
Head sampling decides when a request starts, for example "keep 10% at random". It is cheap and simple, but it may drop the one trace you need.
Tail sampling decides after the whole trace is complete, so it can keep every error and every slow request. It runs in the OpenTelemetry Collector and costs more to operate.
Metrics are not sampled. That is why the duration and token histograms above matter: they count every request even when only a share of traces is kept. Feedback is the exception that needs care: a thumb attached to a span that was sampled away has nothing to attach to. Keep traces that receive feedback, or record the score as an attribute your metrics can count.
You will play a day of the blocked-orders assistant, record it as traces, metrics and user feedback, and write a health report that checks four SLOs. Then you will replay the day with a bad change halfway through and watch the report catch a failure that has no errors in it.
Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv, install the OpenTelemetry libraries, and start Phoenix. This walkthrough doesn't repeat those steps. Phoenix is optional here: everything also runs without Docker.
flowchart LR
A[observe_assistant.py<br/>200 made-up requests] --> F1[traces.jsonl]
A --> F2[feedback.jsonl]
A -. --phoenix .-> P[Phoenix<br/>localhost:6006]
F1 --> R[ai_health_report.py]
F2 --> R
R --> O[Report + ALERT lines]
This script is the "app". It plays 200 requests of the assistant, one every 30 seconds of simulated time, so a whole day runs in a second. Each request becomes a trace with three child spans: retrieval, the SAP tool call and the model call. About 3 in 100 SAP calls time out. About 3 in 10 users click a thumb. Everything is written to two files; with --phoenix it is also sent to Phoenix.
With --incident, request 101 onward runs a new prompt version, v2. It adds the order history to the prompt, and the same week, the notes for block reason 02 went missing from the search index.
In VS Code, right-click unit10, choose New File, name it observe_assistant.py, paste the code below and save.
"""Unit 10: record a day of the blocked-orders assistant as traces, metrics and user feedback.
Observability means you can answer new questions about a running system from the data it
records. This script plays 200 made-up requests of the assistant (no model, no SAP system,
nothing costs money) and records each one the way a production app would:
traces one trace per request, with spans for retrieval, the SAP tool call and the model
metrics OpenTelemetry's standard AI histograms: operation duration and token usage
feedback a thumbs up or down from some users, linked to the request by its span ID
Everything is written to files in unit10/ so the report script can read it without Docker.
How to run (from your course folder, with .venv turned on):
python unit10/observe_assistant.py a normal day
python unit10/observe_assistant.py --incident a day where a change goes wrong halfway
python unit10/observe_assistant.py --phoenix also send traces and feedback to Phoenix
"""
import argparse
import json
import os
import random
import sys
import time
import urllib.error
import urllib.request
from pathlib import Path
from opentelemetry import trace
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, SpanExporter, SpanExportResult
HERE = Path(__file__).parent
TRACES_FILE = HERE / "traces.jsonl"
FEEDBACK_FILE = HERE / "feedback.jsonl"
PHOENIX = os.environ.get("PHOENIX_URL", "http://localhost:6006")
MODEL = "example-small-model" # a label only: this script calls no model
# Made-up open orders, shaped like the SAP sales order fields used in earlier units.
ORDERS = [
{"SalesOrder": str(9000001 + i), "DeliveryBlockReason": random.Random(i).choice(["01", "02", ""]),
"TotalCreditCheckStatus": random.Random(i + 99).choice(["A", "B"])}
for i in range(40)
]
class JsonLinesExporter(SpanExporter):
"""Write each finished span as one line of JSON. A tiny stand-in for a trace backend."""
def __init__(self, path: Path):
self.file = open(path, "w", encoding="utf-8")
def export(self, spans):
for s in spans:
self.file.write(json.dumps({
"trace_id": f"{s.context.trace_id:032x}",
"span_id": f"{s.context.span_id:016x}",
"parent_id": f"{s.parent.span_id:016x}" if s.parent else None,
"name": s.name,
"start_ns": s.start_time,
"duration_ms": round((s.end_time - s.start_time) / 1_000_000, 1),
"status": s.status.status_code.name, # UNSET, OK or ERROR
"attributes": dict(s.attributes),
}) + "\n")
return SpanExportResult.SUCCESS
def shutdown(self):
self.file.close()
def setup(send_to_phoenix: bool):
"""Create the tracer and meter providers, and say where the data goes."""
resource = Resource.create({
"service.name": "blocked-orders-assistant",
"service.version": "1.4.0",
"openinference.project.name": "orchestrate-unit10-observability", # Phoenix project
})
tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(JsonLinesExporter(TRACES_FILE)))
if send_to_phoenix:
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint=f"{PHOENIX}/v1/traces")))
reader = InMemoryMetricReader() # keeps metrics in memory; production would export them
meter_provider = MeterProvider(resource=resource, metric_readers=[reader])
return tracer_provider, meter_provider, reader
def one_request(tracer, meter_instruments, rng, order, start_ns, prompt_version, notes_missing):
"""Play one assistant request with explicit timestamps, so 200 requests take no real time.
Returns (root span ID, whether the answer was grounded in notes)."""
duration_hist, token_hist = meter_instruments
t = start_ns
ms = 1_000_000 # nanoseconds per millisecond
root = tracer.start_span("invoke_agent blocked-orders-assistant", start_time=t)
root.set_attribute("gen_ai.operation.name", "invoke_agent")
root.set_attribute("app.prompt.version", prompt_version)
root.set_attribute("app.sales_order", order["SalesOrder"]) # an ID your policy allows; no names or text
ctx = trace.set_span_in_context(root)
# 1. Retrieval: search the block-reason notes.
span = tracer.start_span("retrieval block-reason-notes", context=ctx, start_time=t)
span.set_attribute("gen_ai.operation.name", "retrieval")
found = 0 if (notes_missing and order["DeliveryBlockReason"] == "02") else rng.randint(2, 4)
span.set_attribute("app.documents.returned", found)
t += int(rng.uniform(60, 140) * ms)
span.end(end_time=t)
# 2. Tool: read the order from SAP. About 3 in 100 calls time out.
span = tracer.start_span("execute_tool get_sales_order", context=ctx, start_time=t)
span.set_attribute("gen_ai.operation.name", "execute_tool")
span.set_attribute("gen_ai.tool.name", "get_sales_order")
if rng.random() < 0.03:
t += 5000 * ms # the client gives up after 5 seconds
span.set_attribute("error.type", "timeout")
span.set_status(trace.Status(trace.StatusCode.ERROR, "SAP API did not answer within 5 s"))
span.end(end_time=t)
root.set_attribute("error.type", "tool_timeout")
root.set_status(trace.Status(trace.StatusCode.ERROR, "tool call failed"))
root.end(end_time=t)
return root.get_span_context().span_id, False
t += int(rng.uniform(40, 120) * ms)
span.end(end_time=t)
# 3. Model call. Prompt v2 adds the full order history to the prompt: more tokens, more time.
input_tokens = int(rng.gauss(820, 60)) * (2 if prompt_version == "v2" else 1)
output_tokens = int(rng.gauss(95, 20)) if found else int(rng.gauss(40, 8))
span = tracer.start_span(f"chat {MODEL}", context=ctx, start_time=t, kind=trace.SpanKind.CLIENT)
attrs = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "example",
"gen_ai.request.model": MODEL}
for k, v in attrs.items():
span.set_attribute(k, v)
span.set_attribute("gen_ai.usage.input_tokens", input_tokens)
span.set_attribute("gen_ai.usage.output_tokens", output_tokens)
model_ms = 150 + input_tokens * 0.15 + output_tokens * 4 * rng.uniform(0.8, 1.3)
t += int(model_ms * ms)
span.end(end_time=t)
# The same facts as metrics: cheap to keep for every request, even when traces are sampled.
duration_hist.record(model_ms / 1000, attributes=attrs)
token_hist.record(input_tokens, attributes={**attrs, "gen_ai.token.type": "input"})
token_hist.record(output_tokens, attributes={**attrs, "gen_ai.token.type": "output"})
root.set_attribute("app.answer.grounded", bool(found))
root.end(end_time=t)
return root.get_span_context().span_id, bool(found)
def send_feedback_to_phoenix(rows):
"""Attach each thumbs up or down to its trace in Phoenix (span annotations API)."""
body = {"data": [{"span_id": r["span_id"], "name": "user feedback", "annotator_kind": "HUMAN",
"result": {"label": r["label"], "score": r["score"]}} for r in rows]}
req = urllib.request.Request(f"{PHOENIX}/v1/span_annotations?sync=false", method="POST",
data=json.dumps(body).encode(), headers={"Content-Type": "application/json"})
try:
urllib.request.urlopen(req, timeout=10)
print(f"Sent {len(rows)} feedback annotations to Phoenix.")
except (urllib.error.URLError, OSError) as e:
print(f"Could not send feedback to Phoenix ({e}). The traces and local files are fine.")
def main():
parser = argparse.ArgumentParser(description="Record a day of assistant requests as telemetry.")
parser.add_argument("--requests", type=int, default=200, help="how many requests to play (default 200)")
parser.add_argument("--incident", action="store_true", help="halfway through, deploy a change that goes wrong")
parser.add_argument("--phoenix", action="store_true", help="also send traces and feedback to Phoenix")
parser.add_argument("--seed", type=int, default=7, help="change it to get a different day")
args = parser.parse_args()
if args.phoenix:
try:
urllib.request.urlopen(PHOENIX, timeout=3)
except urllib.error.HTTPError:
pass # it answered: fine
except (urllib.error.URLError, OSError):
print(f"Phoenix did not answer at {PHOENIX}. Start it (Unit 10 setup, Step 4) or run without --phoenix.")
sys.exit(1)
rng = random.Random(args.seed)
tracer_provider, meter_provider, reader = setup(args.phoenix)
tracer = tracer_provider.get_tracer("orchestrate.unit10")
meter = meter_provider.get_meter("orchestrate.unit10")
# Metric names and units from OpenTelemetry's generative AI conventions.
duration_hist = meter.create_histogram("gen_ai.client.operation.duration", unit="s",
description="GenAI operation duration")
token_hist = meter.create_histogram("gen_ai.client.token.usage", unit="{token}",
description="Number of input and output tokens used")
start_of_day = time.time_ns() - args.requests * 30 * 1_000_000_000 # one request every 30 s, ending now
feedback = []
for i in range(args.requests):
changed = args.incident and i >= args.requests // 2
span_id, grounded = one_request(
tracer, (duration_hist, token_hist), rng, rng.choice(ORDERS),
start_of_day + i * 30 * 1_000_000_000,
prompt_version="v2" if changed else "v1", notes_missing=changed)
if rng.random() < 0.3: # about 3 in 10 users click a thumb
happy = rng.random() < (0.9 if grounded else 0.25)
feedback.append({"span_id": f"{span_id:016x}", "label": "thumbs-up" if happy else "thumbs-down",
"score": 1 if happy else 0})
tracer_provider.shutdown() # flush every span to the files (and to Phoenix)
with open(FEEDBACK_FILE, "w", encoding="utf-8") as f:
for row in feedback:
f.write(json.dumps(row) + "\n")
print(f"Recorded {args.requests} requests{' (with an incident halfway)' if args.incident else ''}.")
print(f" traces -> {TRACES_FILE.relative_to(HERE.parent)}")
print(f" feedback -> {FEEDBACK_FILE.relative_to(HERE.parent)} ({len(feedback)} thumbs)")
print(" metrics (kept in memory, summarised here):")
for rm in reader.get_metrics_data().resource_metrics:
for sm in rm.scope_metrics:
for m in sm.metrics:
for p in m.data.data_points:
kind = p.attributes.get("gen_ai.token.type", "")
label = f"{m.name}{' (' + kind + ')' if kind else ''}"
print(f" {label:<45} count {p.count:>4} average {p.sum / p.count:8.2f} {m.unit}")
if args.phoenix:
send_feedback_to_phoenix(feedback)
print(f"Open {PHOENIX} and choose project orchestrate-unit10-observability.")
print("Next: python unit10/ai_health_report.py")
if __name__ == "__main__":
main()
Run a normal day (the same on every system):
python unit10/observe_assistant.py
What success looks like (from our test; the numbers are the same for you because of the fixed seed):
Recorded 200 requests.
traces -> unit10/traces.jsonl
feedback -> unit10/feedback.jsonl (60 thumbs)
metrics (kept in memory, summarised here):
gen_ai.client.operation.duration count 196 average 0.67 s
gen_ai.client.token.usage (input) count 196 average 822.47 {token}
gen_ai.client.token.usage (output) count 196 average 94.47 {token}
Next: python unit10/ai_health_report.py
The metrics count 196, not 200: four requests failed on the SAP call before reaching the model. On Windows the arrows show the path with \ instead of /; that is normal.
Open unit10/traces.jsonl in VS Code. Each line is one span. Find a line whose "parent_id" is null: that is a request's root span, with app.prompt.version and app.answer.grounded.
This script is the "dashboard". It rebuilds each request from its spans, groups requests by prompt version, computes the numbers from the foundational layer's table, and prints an ALERT line for each broken SLO. It uses built-in Python only.
Right-click unit10, choose New File, name it ai_health_report.py, paste and save.
"""Unit 10: turn recorded traces and feedback into a health report with alerts.
Reads unit10/traces.jsonl and unit10/feedback.jsonl (written by observe_assistant.py) and
answers the questions an operations team asks of an AI assistant every day:
Is it up? error rate
Is it fast enough? p50 and p95 latency, and which step the time goes to
What does it cost? tokens and an example cost per request
Is it any good? share of answers grounded in notes, and the thumbs-down rate
It groups requests by prompt version, so you can see what a change did, and checks each
group against service level objectives (SLOs). Built-in Python only.
How to run (from your course folder):
python unit10/ai_health_report.py
"""
import json
import statistics
import sys
from collections import defaultdict
from pathlib import Path
HERE = Path(__file__).parent
# Example prices per million tokens, made up for this exercise. Use your contract's real prices.
PRICE_PER_MILLION = {"input": 0.50, "output": 2.00}
# Service level objectives: the line between "fine" and "someone should look".
SLO = {
"error_rate_max": 0.05, # at most 5 in 100 requests fail
"p95_ms_max": 3000, # 95 in 100 answers within 3 seconds
"grounded_rate_min": 0.85, # at least 85 in 100 answers use retrieved notes
"thumbs_down_rate_max": 0.25, # at most 1 in 4 rated answers get a thumbs down
}
def percentile(values, pct):
"""The value below which pct percent of the values fall (nearest-rank method)."""
ordered = sorted(values)
rank = max(1, round(pct / 100 * len(ordered)))
return ordered[rank - 1]
def load(name):
path = HERE / name
if not path.exists():
print(f"{path} not found. Run: python unit10/observe_assistant.py")
sys.exit(1)
with open(path, encoding="utf-8") as f:
return [json.loads(line) for line in f if line.strip()]
def main():
spans = load("traces.jsonl")
feedback = {row["span_id"]: row["score"] for row in load("feedback.jsonl")}
# Rebuild each request from its spans: the root span plus its children, by trace ID.
by_trace = defaultdict(list)
for s in spans:
by_trace[s["trace_id"]].append(s)
groups = defaultdict(list)
for trace_spans in by_trace.values():
root = next(s for s in trace_spans if s["parent_id"] is None)
chat = next((s for s in trace_spans if s["attributes"].get("gen_ai.operation.name") == "chat"), None)
steps = {s["attributes"].get("gen_ai.operation.name"): s["duration_ms"]
for s in trace_spans if s["parent_id"] is not None}
groups[root["attributes"].get("app.prompt.version", "unknown")].append({
"ms": root["duration_ms"],
"error": root["status"] == "ERROR",
"grounded": root["attributes"].get("app.answer.grounded", False),
"tokens_in": chat["attributes"]["gen_ai.usage.input_tokens"] if chat else 0,
"tokens_out": chat["attributes"]["gen_ai.usage.output_tokens"] if chat else 0,
"steps": steps,
"score": feedback.get(root["span_id"]),
})
alerts = []
for version in sorted(groups):
reqs = groups[version]
ok = [r for r in reqs if not r["error"]]
rated = [r["score"] for r in reqs if r["score"] is not None]
error_rate = 1 - len(ok) / len(reqs)
p50, p95 = percentile([r["ms"] for r in reqs], 50), percentile([r["ms"] for r in reqs], 95)
grounded = sum(r["grounded"] for r in ok) / len(ok) if ok else 0
down = rated.count(0) / len(rated) if rated else 0
tin = statistics.mean(r["tokens_in"] for r in ok) if ok else 0
tout = statistics.mean(r["tokens_out"] for r in ok) if ok else 0
cost = (tin * PRICE_PER_MILLION["input"] + tout * PRICE_PER_MILLION["output"]) / 1_000_000
step_ms = defaultdict(list)
for r in ok:
for step, ms in r["steps"].items():
step_ms[step].append(ms)
share = {step: sum(v) / sum(r["ms"] for r in ok) for step, v in step_ms.items()}
print(f"\nPrompt version {version}: {len(reqs)} requests")
print(f" Errors {error_rate:6.1%}")
print(f" Latency p50 {p50:6.0f} ms p95 {p95:6.0f} ms")
print(" Time by step " + " ".join(f"{k} {v:.0%}" for k, v in sorted(share.items(), key=lambda x: -x[1])))
print(f" Tokens/request in {tin:6.0f} out {tout:4.0f} example cost ${cost * 1000:.2f} per 1,000 requests")
print(f" Grounded {grounded:6.1%} of successful answers used retrieved notes")
print(f" Thumbs down {down:6.1%} of {len(rated)} rated answers")
checks = [
(error_rate > SLO["error_rate_max"], f"error rate {error_rate:.1%} is above {SLO['error_rate_max']:.0%}"),
(p95 > SLO["p95_ms_max"], f"p95 latency {p95:.0f} ms is above {SLO['p95_ms_max']} ms"),
(grounded < SLO["grounded_rate_min"], f"grounded answers {grounded:.1%} is below {SLO['grounded_rate_min']:.0%}"),
(down > SLO["thumbs_down_rate_max"], f"thumbs-down rate {down:.1%} is above {SLO['thumbs_down_rate_max']:.0%}"),
]
alerts += [f"ALERT [{version}] {text}" for broken, text in checks if broken]
print()
print("\n".join(alerts) if alerts else "All service level objectives met.")
if __name__ == "__main__":
main()
Run it:
python unit10/ai_health_report.py
What success looks like:
Prompt version v1: 200 requests
Errors 2.0%
Latency p50 845 ms p95 1054 ms
Time by step chat 79% retrieval 12% execute_tool 9%
Tokens/request in 822 out 94 example cost $0.60 per 1,000 requests
Grounded 100.0% of successful answers used retrieved notes
Thumbs down 10.0% of 60 rated answers
All service level objectives met.
Read it like an operator. Four SAP timeouts gave a 2% error rate, under the 5% target. The p95 of about one second is far under three. The model call takes about four fifths of the time, so that is where a latency project would start. A valid "quiet" result is the last line: no alerts means nobody gets woken up.
Notice what the p95 hides: the four timed-out requests took over five seconds each. They are 2% of traffic, so they sit above the 95th percentile. A p99 target, or an alert on error.type = timeout, would catch them.
What success looks like (the first run's lines are shortened here):
Recorded 200 requests (with an incident halfway).
...
Prompt version v1: 100 requests
Errors 2.0%
Latency p50 842 ms p95 1027 ms
Time by step chat 79% retrieval 12% execute_tool 9%
Tokens/request in 819 out 94 example cost $0.60 per 1,000 requests
Grounded 100.0% of successful answers used retrieved notes
Thumbs down 10.3% of 29 rated answers
Prompt version v2: 100 requests
Errors 1.0%
Latency p50 917 ms p95 1168 ms
Time by step chat 81% retrieval 11% execute_tool 8%
Tokens/request in 1638 out 83 example cost $0.98 per 1,000 requests
Grounded 73.7% of successful answers used retrieved notes
Thumbs down 27.3% of 33 rated answers
ALERT [v2] grounded answers 73.7% is below 85%
ALERT [v2] thumbs-down rate 27.3% is above 25%
Compare the two blocks. Version v2 has fewer errors and stays well inside the latency target. An IT dashboard would call it healthy. Yet one answer in four is no longer grounded in notes, users noticed, and each request costs about 60% more. No ALERT covers cost yet: you add that in the exercise.
Sent 62 feedback annotations to Phoenix.
Open http://localhost:6006 and choose project orchestrate-unit10-observability.
Next: python unit10/ai_health_report.py
app.prompt.version == 'v2' and annotations['user feedback'].label == 'thumbs-down'
9
The same, after the change
Click one row from the third filter. In the span's detail panel, the annotations area shows user feedback with label thumbs-down. That is the thumb, attached to the exact request it rated.
If the project doesn't appear, refresh the page; spans are sent in batches. If the filter bar shows a syntax error, check that quotes are straight ' quotes.
SAP's learning material describes a division of labour:
Service
Role
Typical user
SAP Cloud ALM
Central, high-level monitoring of BTP apps, so problems are spotted early
Central support team
SAP Cloud Logging
In-depth monitoring and root-cause analysis
The DevOps team that owns the app
SAP Alert Notification service
Turns technical events, such as an app crash, into real-time events and alerts, including in Cloud ALM
Both
An SAP user group session from March 2025 describes SAP Cloud Logging in more detail: logs, metrics and traces stored on OpenSearch, ingestion over OpenTelemetry (gRPC) among other methods, OpenSearch Dashboards for queries and visualisations, and alerts on any ingested data to destinations such as a webhook or Slack. It lists dev and build-code plans for evaluation only, and standard and large plans for production. On Kyma, the Telemetry module ships OTLP to an SAP Cloud Logging instance, as covered in the Unit 10 setup.
Because the protocol is OpenTelemetry, your spans and attribute names don't change. Only the exporter's configuration does: the SAP Cloud Logging endpoint and the credentials from its service binding, and the gRPC variant of the OTLP exporter if your instance accepts only gRPC. Check the instance's binding for the exact endpoint and protocol.
SAP publishes the SAP Cloud SDK for Python on PyPI (Apache 2.0, Python 3.11 or later). Its telemetry module, as described in the user guide shipped with version 0.58.1 (6 October 2026), works in two layers:
auto_instrument() traces AI library calls and common HTTP and web frameworks automatically. It exports to OTEL_EXPORTER_OTLP_ENDPOINT over gRPC by default (OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf switches to HTTP), or prints with OTEL_TRACES_EXPORTER=console. Without either variable it logs a warning and does nothing.
invoke_agent_span, chat_span and execute_tool_span create spans with the OpenTelemetry AI conventions, for the business context automation can't see: which agent, which tenant, which operation.
A sketch of how the assistant would look with it:
# Sketch: needs `pip install sap-cloud-sdk` (Python 3.11+) and an OTLP endpoint.
from sap_cloud_sdk.core.telemetry import auto_instrument, invoke_agent_span, execute_tool_span
auto_instrument() # call before importing AI libraries
with invoke_agent_span(provider="example", agent_name="blocked-orders-assistant", # provider: a label
attributes={"app.prompt.version": "v2"}):
with execute_tool_span(tool_name="get_sales_order", tool_type="function"):
order = read_sales_order("9000001") # your SAP API call, traced as a child span
answer = ask_model(order) # an instrumented AI client adds the chat span
The generative AI hub's orchestration service (see SAP generative AI hub and orchestration) returns more than an answer. The response schema in SAP's JavaScript SDK (@sap-ai-sdk/orchestration 2.16.0) includes:
request_id: record it on your chat span. It is the link between your trace and SAP's side when you raise a ticket.
final_result with token usage: prompt_tokens, completion_tokens and total_tokens. Map them to gen_ai.usage.input_tokens and gen_ai.usage.output_tokens.
intermediate_results: the output of each module that ran, such as templating, grounding, input_masking, input_filtering and output_filtering. Record facts about them (for example, "output filter triggered"), not their content.
intermediate_failures: module configurations that failed before a fallback succeeded. A rising count is an early warning even when every request succeeds.
The SDK's response object exposes these as getRequestId(), getTokenUsage(), getIntermediateResults() and getIntermediateFailures().
Data protection first. Treat traces as a data store. Default to facts (IDs your policy allows, lengths, counts, versions). If you capture content, mask personal data first, restrict access, and set retention in the backend. The guardrails and data masking topics in Unit 11 go further.
SAP authorizations. A trace viewer can show data from SAP the viewer has no authorization to see in SAP itself. Either keep SAP field values out of spans, or restrict the trace backend to people with matching SAP roles.
Cardinality. Metric attributes such as gen_ai.request.model are fine. A user ID or order number as a metric attribute creates a series per value and breaks the budget. Keep high-cardinality IDs on spans.
Sampling with intent. Use tail sampling to keep every error, every slow trace and every trace with feedback; sample the rest. Keep metrics unsampled.
Alert design. Alert on SLOs over a window, route each alert to an owner, and delete alerts nobody acts on. Quality alerts (grounding, thumbs-down) need a business owner as well as a technical one.
Cost of observability. Traces with content can exceed the size of the requests themselves. Watch the ingestion volume of your backend like any other bill.
Version everything. Prompt, model, retrieval index and app versions on every root span. The incident in this topic is only findable because of app.prompt.version.
Clean core. Instrumentation lives in your side-by-side app on BTP. Nothing is installed in S/4HANA to trace an AI app that reads it.
Dashboards of only IT signals. Green on errors and latency says nothing about answer quality.
Feedback without an ID. A thumb stored without the span ID can be counted but never investigated.
Averages in SLOs. A good average can sit on top of a painful tail.
Inventing attribute names per team.model, llm_model and gen_ai.request.model in three services make one query impossible. Agree on the conventions and keep your names under app..
Turning on an instrumentation library without checking its content default. Some record full prompts unless told not to.
Alerts on single requests. One slow answer is noise; a broken SLO over an hour is a signal.
Forgetting the flush. A short-lived job that exits before its batch is sent loses its traces. Call shutdown() on the provider.
Add the missing cost alert, and find which orders the incident hurt. The cost numbers feed the token economics topic later in this unit.
Open unit10/ai_health_report.py. In the SLO dictionary, below the thumbs_down_rate_max line, add:
"cost_per_1000_max": 0.80, # at most $0.80 per 1,000 requests, at the example prices
In the checks list, below the thumbs-down line, add (same indentation):
(cost * 1000 > SLO["cost_per_1000_max"], f"cost ${cost * 1000:.2f} per 1,000 requests is above ${SLO['cost_per_1000_max']:.2f}"),
Open unit10/observe_assistant.py. Below the line root.set_attribute("app.sales_order", ...), add:
root.set_attribute("app.order.block_reason", order["DeliveryBlockReason"] or "none")
Create unit10/by_reason.py with this code and save it:
"""Count ungrounded answers by delivery block reason, from unit10/traces.jsonl."""
import json
from collections import Counter
from pathlib import Path
total, ungrounded = Counter(), Counter()
for line in open(Path(__file__).parent / "traces.jsonl", encoding="utf-8"):
span = json.loads(line)
attrs = span["attributes"]
if span["parent_id"] is None and span["status"] != "ERROR": # successful requests only
reason = attrs.get("app.order.block_reason", "not recorded")
total[reason] += 1
ungrounded[reason] += not attrs.get("app.answer.grounded", False)
for reason in sorted(total):
print(f"block reason {reason:<5} {ungrounded[reason]:>3} of {total[reason]:>3} answers not grounded")
ALERT [v2] grounded answers 73.7% is below 85%
ALERT [v2] thumbs-down rate 27.3% is above 25%
ALERT [v2] cost $0.98 per 1,000 requests is above $0.80
And the breakdown points at one block reason:
block reason 01 0 of 65 answers not grounded
block reason 02 26 of 50 answers not grounded
block reason none 0 of 82 answers not grounded
With Phoenix running, add --phoenix to the first command and try the filter app.order.block_reason == '02' and app.answer.grounded == False.
Write two sentences in unit10/incident-note.md: what broke, and what you would check first. Commit the three scripts and the note.
Done when the normal day prints All service level objectives met., the incident day prints three ALERT lines including cost, and by_reason.py shows every ungrounded answer under block reason 02.
Pick one answer for each question. The explanation appears after you choose.
1In the incident replay, version v2 has fewer errors and a p95 well inside target. Why does the report still raise alerts?
Answer: C. The change removed notes for one block reason, so answers stopped being grounded and users rated them down, while every request still succeeded. Only the quality SLOs saw it. That is the quiet failure the topic is built around.
2Why does the script emit gen_ai.client.token.usage as a metric when the same token counts are already on each chat span?
Answer: B. Traces are often sampled at volume, so totals computed from traces would undercount. Histograms aggregate every request cheaply, which makes them the right source for cost dashboards and alerts.
3How does a user's thumbs-down end up attached to the exact request it rates?
Answer: D. Each thumb is stored with the root span ID of the request it rates; in a live app, that ID travels with the answer to the user interface. Phoenix's span annotations API and the report both join on it. Without it, feedback can be counted but not investigated.
4Your team switches on an AI instrumentation library and wants to know whether prompts are now stored in traces. What is the right move?
Answer: B. The OpenTelemetry conventions capture no content by default, but libraries differ; OpenLLMetry records prompts unless TRACELOOP_TRACE_CONTENT=false is set. Checking the setting and looking at a real trace is the only way to be sure.
5Which of these belongs on a span but not as a metric attribute?
Answer: D. Each distinct metric attribute value creates a new time series, so an order number would create thousands. Low-cardinality values such as model, token type and operation are fine as metric attributes; IDs belong on spans.
6The report's p95 for a normal day is about one second, yet four requests took over five seconds. What explains this?
Answer: C. The p95 is the value 95% of requests beat, so the slowest 5% don't move it. Four timeouts in 200 requests are 2%, invisible to a p95 target. A p99 target or an alert on the timeout error type would catch them.
7Your app calls SAP's orchestration service and a user reports a wrong answer. Which value from the response should already be on your chat span, so you can raise a precise ticket?
Answer: A. The request ID identifies the call on SAP's side, so support can find it. Token counts help with cost, not tracing a single call, and storing templating output or the message would put content in traces.
8You plan to move the assistant's traces from Phoenix to SAP Cloud Logging. What changes in the instrumented code?
Answer: D. SAP Cloud Logging ingests OpenTelemetry, so the spans and attribute names stay as they are. You change the exporter's endpoint, credentials and, if needed, protocol, using the details from the instance's service binding.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sampling (OpenTelemetry documentation)— head sampling decides when a trace starts; tail sampling decides after the whole trace is seen and can keep errors and slow traces; tail sampling runs in the Collector and costs more to operate
SAP Cloud SDK for Python (PyPI)— published by SAP SE under Apache 2.0, Python 3.11 or later, includes a Telemetry and Observability module. Version 0.58.1 (6 October 2026) was read for its telemetry user guide and source: auto_instrument(), invoke_agent_span, chat_span, execute_tool_span, OTEL_EXPORTER_OTLP_ENDPOINT, gRPC by default, built on Traceloop
OrchestrationResponse type definitions (@sap-ai-sdk/orchestration on unpkg)— getRequestId, getTokenUsage, getFinishReason, getIntermediateResults and getIntermediateFailures. Version 2.16.0 schema read in this run: request_id, intermediate_results (templating, grounding, input_masking, input_filtering, output_filtering, llm), final_result with prompt_tokens, completion_tokens and total_tokens