Every program that talks to SAP or to an AI model will meet failures. The network drops. The model provider says "too many requests". A key expires. A field name is wrong. Someone types bad data.
Good AI code does three things about this:
Handles errors on purpose. Some failures go away if you wait and try again. Others never will. Code should tell them apart: retry the first kind a few times, and stop with a clear message on the second kind.
Keeps a log. A log is the program's diary: what it did, when, and what went wrong. When a nightly job fails at 3 a.m., the log is how anyone finds out why. It must never contain passwords, keys or personal data.
Is debugged, not guessed at. When the result is wrong, a developer pauses the program with a debugger and looks at the actual values, instead of changing code and hoping.
None of this is AI-specific. But AI systems call more outside services than most programs, and they fail in more ways, so these habits matter more.
Take the running example from Unit 1: a job reads blocked sales orders from S/4HANA every night and asks a model to summarize them for the credit team. Here is how weak error handling hurts:
Silent failure. The SAP call times out, the script catches the error and carries on with zero orders. The morning report says "no blocked orders". Credit managers relax while 40 orders wait. A wrong answer that looks right is worse than a crash.
Retry storms. The model provider is overloaded. Every job retries at once, as fast as it can, and makes the overload worse. If the provider charges per request, failed retries also cost money.
Duplicate business documents. A program retries a call that creates something, such as a sales order or a payment proposal. The first call actually worked; the answer just got lost. Now there are two.
Leaked secrets. A developer logs a whole request "to debug it", including the API key. Logs are copied to support tickets and monitoring tools. The key is now in five places.
Slow fixes. Without a log that says which call failed, with which status and which SAP transaction ID, a two-minute fix becomes a two-day investigation across three teams.
The business decision is simple to state: for every AI job that touches SAP, someone should be able to say what happens when each outside call fails, and where the evidence ends up.
SAP systems and services already come with error and log tools. Your team's code should fit into them rather than invent its own. As of October 2026:
SAP Gateway error logs (S/4HANA and other ABAP systems). When an OData call fails, SAP Gateway can record it. SAP Learning describes two logs: the SAP Gateway error log (/IWFND/ERROR_LOG) for requests that fail when the request is analyzed, and the backend error log (/IWBEP/ERROR_LOG) for failures inside the service's implementation. The error details include the call stack and the full request and response, and a Replay option to repeat the call. That is why an AI app should log any transaction ID SAP returns: it lets an SAP administrator find the matching entry.
SAP Cloud Logging (SAP BTP). For apps you build and run on SAP BTP, SAP Learning describes SAP Cloud Logging as the service to ingest, store and analyze application logs, with prebuilt dashboards and support for the OpenTelemetry standard. It links to SAP Cloud ALM, SAP's central monitoring tool, so operations teams can jump from an alert into the detailed logs.
Model calls. Model providers' own libraries already retry some failures. Anthropic's Python library, for example, retries connection errors, rate limits and server errors twice by default. SAP's orchestration service can fall back to a second model configuration when the first fails, as covered in SAP Generative AI Hub and the orchestration service.
What SAP does not do for you: decide which failures your job should retry, what it should do with a half-finished batch, and what must never appear in your logs. That is your team's design.
"Fail loudly" means stop at once with a clear message and a non-zero result, so a person or a scheduler notices. "Fail gracefully" means record the problem, skip the affected piece and finish the rest. Neither is always right. Use this guide:
Situation
Behavior
Why
API key missing, wrong or refused
Stop at once, no retries
Retrying cannot fix it, and repeated refused logins can lock accounts
SAP or the model says "too many requests" (429) or "unavailable" (503)
Wait, retry a few times, then stop
Usually temporary; but give up after a limit so the job doesn't hang
One sales order out of 200 can't be read
Record it, skip it, finish the rest, and report "199 of 200"
One bad record shouldn't block the credit team's whole list
Most records fail
Stop and alert
A broken job that "finishes" hides a real outage
A call that creates or changes SAP data fails with no clear answer
Don't retry blindly; check whether it worked first
A blind retry can create a duplicate document
The model's summary call fails, but the data was read
Save the data, skip the summary, flag it
The order list is still useful without the prose
The model returns text that isn't in the expected format
Treat as an error: retry once or route to a person
"Catching every error makes the program robust." Catching everything and carrying on is how silent failures happen. Catch only the errors you know how to handle; let the rest stop the program loudly.
"Retrying always helps." Retrying a wrong key, a wrong field name or a bad request just fails again, slower. Retrying a call that creates data can make duplicates.
"More logging is always better." Logs full of noise hide the line that matters, cost money to store, and are the most common place secrets leak.
"The AI provider handles errors for us." Provider libraries retry a few temporary failures. They don't know your business rules: what to skip, what to report, and what a half-finished run means.
"Debugging means adding print statements." Prints help, but a debugger shows every value at the moment of the problem without changing the code.
Pick one answer for each question. The explanation appears after you choose.
1A nightly job's SAP call times out. The code catches the error and reports "0 blocked orders". What is the real problem?
Answer: B. Catching an error and carrying on as if nothing happened produces a wrong answer that looks right. The job should either retry and succeed, or stop and say that SAP could not be read.
2Which failure should the job retry after a short wait?
Answer: C. A 429 says "not now", so waiting and trying again can work. A refused key or a wrong field fails the same way every time; retrying only delays the clear message someone needs.
3A job creates sales orders in SAP. One call fails with no clear answer. What should happen next?
Answer: B. A call that creates data is not safe to repeat blindly. If the first call actually worked and only the answer was lost, a retry makes a duplicate document.
4One sales order out of 200 can't be read; the other 199 are fine. What is the best behavior?
Answer: C. One bad record shouldn't block the credit team's whole list, but the gap must be visible. Reporting "199 of 200, here is the one that failed" is graceful and honest.
5What should you ask to see before an AI job goes live?
Answer: B. Logs are copied into tickets and monitoring tools, so a leaked key or customer email spreads quickly. Seeing a real log line is the fastest way to check what the job records.
6Why should the AI app log the transaction ID that an SAP OData error returns?
Answer: D. SAP Gateway keeps error logs with the call stack and full request and response. The ID in your log connects your side of the failure to SAP's side, which turns a long investigation into a quick lookup.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: sort every failure into one of three bins
Every failure your code meets belongs in one of three bins, and each bin has one right response:
Worth waiting for (transient). Timeouts, dropped connections, HTTP 429, 502, 503, 504. Response: wait, retry a limited number of times, then give up loudly.
Worth telling someone (permanent). A refused key (401, 403), a wrong address (404), a bad request (400), a missing setting. Response: stop that piece of work at once, with a message that says what to fix.
A bug in your own code. A KeyError, a TypeError, a wrong answer with no error at all. Response: don't catch it. Let it crash during development, then reproduce it and find it with a debugger.
Around those bins sit two tools. Logs are the flight recorder: they say which bin each failure went into, when, and with which IDs. The debugger is the pause button you use on bin 3.
The rest of this topic builds code that sorts failures into these bins automatically, writes a safe log of what happened, and then uses the debugger on a real bug.
When something goes wrong while Python runs, it raises an exception: an object that describes the problem, such as KeyError or TimeoutError. If no code handles it, Python stops and prints a traceback.
A traceback lists the chain of function calls that led to the error. Python's tutorial calls the order "most recent call last", so:
Read from the bottom up. The last line is the error type and message. The lines just above it show where it happened.
Find your own file. Lines inside libraries (paths containing site-packages) usually show where the error surfaced, not where it started. The last line that points into your file is usually where to look.
Since Python 3.11, tracebacks underline the exact part of the line with ~ and ^ characters.
Traceback (most recent call last):
File ".../unit01/buggy_totals.py", line 34, in <module>
main()
File ".../unit01/buggy_totals.py", line 28, in main
if is_blocked(order) and needs_review(order):
~~~~~~~~~~^^^^^^^
File ".../unit01/buggy_totals.py", line 18, in is_blocked
return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])
~~~~~^^^^^^^^^^^^^^^^^
KeyError: 'DeliveryBlock'
Bottom line first: KeyError: 'DeliveryBlock', a dictionary has no such key. One line up: it happened in is_blocked, line 18. The field in SAP's Sales Order API is DeliveryBlockReason. You will fix exactly this bug later.
try:
reply = call_sap() # code that might fail
except TimeoutError as err: # runs only if that error happened
log.warning("SAP timed out: %s", err)
else: # runs only if nothing failed
process(reply)
finally: # always runs, error or not
close_connection()
The Python tutorial gives the rules: one matching except block runs; else runs only when the try block raised nothing; finally always runs last, which makes it the place for clean-up. Putting the success path in else rather than inside try keeps you from accidentally catching errors from code you didn't mean to protect.
Three habits follow:
Catch narrowly.except KeyError: says what you expect. A bare except: or a broad except Exception: in the middle of your logic hides bugs from bin 3 together with the failures you meant to handle.
Catch where you can decide. A low-level function that sends a request can't know whether one failed order should stop the job. Let the error travel up to the code that can decide, usually the main loop.
Raise your own exception types. A custom class, such as class TransientError(Exception), lets the rest of the code react to meaning ("retry this") instead of to a library's details. Python's tutorial recommends deriving from Exception and ending the name in Error.
When you turn a library's error into your own, chain them with raise ... from err. The traceback then shows both: your clear message and the original cause. (raise ... from None hides the cause; use it rarely.) Python 3.11 also added err.add_note("...") to attach context, such as which sales order was being read, without changing the error type.
#Transient or permanent: deciding from the HTTP answer
For calls over HTTP, the status code is the first clue to the bin:
Answer
Typical meaning
Bin
Timeout or dropped connection
Network, proxy or an overloaded server
Transient
429 Too Many Requests
You are being rate-limited
Transient: wait, ideally as long as the server says
502, 503, 504
A gateway or the service is temporarily unavailable
Transient
400 Bad Request
Something in the request is wrong, such as a field name
Permanent
401 Unauthorized, 403 Forbidden
Key missing, wrong, or not allowed
Permanent
404 Not Found
Wrong address or a record that doesn't exist
Permanent
500 Internal Server Error
The server failed; the cause could be either
Depends: retrying once is common, but log it as an error
Treat this as a starting point, not a law. Check each API's own documentation. Anthropic's error reference, for example, uses its own 529 status for "overloaded", and its Python library retries 408, 409, 429 and 5xx answers by default.
An SAP OData V2 service returns error details in the response body as JSON. The shape is an error object with a code, a message whose text is in message.value, and often an innererror with more detail. SAP systems commonly put a transactionid there. The exact content varies by system and service, so read it defensively: if the body isn't the shape you expect, keep the first part of the raw text instead of crashing while handling an error.
What to retry. Only transient errors, and only calls that are safe to repeat.
How long to wait. Waits grow each time: exponential backoff.
How to avoid a crowd. If 50 jobs all fail at 02:00 and all wait exactly 1, 2, 4 seconds, they all come back together. AWS's architecture blog calls these "clusters of calls". Adding jitter, a random wait, spreads them out. The "full jitter" version picks a wait anywhere between zero and the current backoff limit: sleep = random_between(0, min(cap, base * 2 ** attempt)).
When to stop. A fixed number of attempts and a cap on each wait, so a job never hangs for hours.
If the server sends a Retry-After header, it is telling you how long to wait. Use it, but still cap it.
flowchart TD
A["Call SAP"] --> B{"Answer?"}
B -->|"2xx"| OK["Use the data"]
B -->|"429, 503, timeout"| C{"Attempts left?"}
C -->|"Yes"| W["Wait: Retry-After<br/>or backoff + jitter"]
W --> A
C -->|"No"| T["Give up: log ERROR,<br/>stop or skip"]
B -->|"400, 401, 404"| P["No retry: log,<br/>stop or skip"]
Is it safe to repeat? RFC 9110, the HTTP standard, defines idempotent methods: repeating them has the same effect as sending them once. GET (reading) and the other safe methods are idempotent, and so are PUT and DELETE. POST, which SAP's OData APIs use to create entities such as sales orders, is not. The standard says a client should not automatically retry a non-idempotent request unless it knows the request is safe to repeat, or knows the first one didn't complete. For writes to SAP, check first (read the order back, or look for your own reference number) before you try again.
Retries also cost money and quota. Every retried model call may be billed, and a retry storm can use up a whole team's rate limit. Keep attempts low (three or four in total is typical) and log every retry, so you can see when they start to climb.
Python's logging module is built into the language. Its HOWTO gives a simple rule for what goes where:
You want to...
Use
Show normal output to the person running the script
print()
Record what happened during normal operation
log.info(), or log.debug() for detail
Warn about something unexpected that you handled
log.warning()
Report an error the code can't handle here
Raise an exception
Record an error you handled without raising
log.error(), or log.exception() inside an except block to include the traceback
The five levels, from least to most serious, are DEBUG, INFO, WARNING, ERROR and CRITICAL. By default only WARNING and above are shown, which is why a first log.info() often seems to do nothing.
Three parts work together:
A logger is what your code calls. Create one per file with log = logging.getLogger(__name__), so each entry says which module wrote it.
A handler decides where entries go: the screen, a file, or a log service. Each handler can have its own level, so the screen shows INFO while the file keeps DEBUG.
A formatter decides what each line looks like. A filter can change or drop entries before they are written; you will use one to add a run ID and remove secrets.
Pass values as arguments, log.warning("order %s failed", number), rather than building the string yourself. The logging module only formats the message if the entry is actually written.
A line like WARNING order 9000004 skipped is fine for a person. A log tool searching thousands of runs needs fields. Structured logging writes each entry as data, commonly one JSON object per line (JSON Lines):
Now you can ask "show every 429 last week" or "everything from run 70265a47". The run ID (a correlation ID) ties all entries from one run together. In a bigger system you pass the same ID along to every service you call, so one ID finds the story everywhere.
OWASP's Logging Cheat Sheet lists data that should usually be removed, masked, hashed or encrypted rather than logged. For AI apps on SAP, the ones that matter most are:
Access tokens, API keys, passwords, encryption keys and connection strings. That includes Authorization and APIKey headers.
Sensitive personal data. Customer and employee names, emails, phone numbers, bank details. In AI apps this includes the prompt and the model's answer when they contain such data.
Session identifiers and payment card data.
Two practical rules follow. Never log whole requests, headers or prompts by default; log IDs, counts, status codes and timings instead. And add a safety net: a filter that masks known secrets and obvious patterns (emails, apikey=...) in every entry, including tracebacks. OWASP also warns about log injection: if user text containing line breaks reaches your log, it can forge fake entries. Replacing carriage returns and line feeds in messages prevents that.
Debugging is a loop: reproduce, observe, explain, fix, check.
Reproduce the problem with the smallest input you can, ideally sample data rather than a live system.
Observe. Read the traceback if there is one. If the answer is just wrong, pause the program where the wrong value appears.
Explain the cause in one sentence before changing anything ("the amount is text, so the comparison is alphabetical").
Fix that cause, not the symptom.
Check with the same input, and keep that input as a test.
VS Code's debugger, which comes with the Python extension you installed in Set up your computer, does step 2 well. You set a breakpoint by clicking in the margin left of a line number (or pressing F9). When the program reaches it, it pauses. The Run and Debug view then shows Variables (every value right now), Watch (expressions you choose), and Call Stack (how you got here). The Debug Console lets you type Python expressions and see the result in the paused program. The toolbar moves you on: Continue (F5), Step Over (F10, run this line), Step Into (F11, go inside the function on this line) and Step Out (finish this function).
Two more tools save time with loops over many sales orders. A conditional breakpoint pauses only when an expression is true, such as order["SalesOrder"] == "9000002". A logpoint prints a message with values, such as amount={order["TotalNetAmount"]}, without pausing and without editing your code.
Without VS Code, Python's built-in breakpoint() does the same job in the terminal: put it on a line, run the script, and you get a (Pdb) prompt where p expression prints a value, n runs the next line and c continues.
#Build it yourself: a sturdy SAP reader and a debugging session
You will build two things:
sturdy_calls.py reads blocked sales orders and their items. It sorts every failure into the three bins, retries what is worth retrying, skips one bad order without losing the others, and writes a safe JSON log. Its --sample mode acts out failures on purpose (a 503, a 429, a timeout, a missing order, a wrong key) so you can watch each behavior on your own computer.
buggy_totals.py has two bugs on purpose. You will find one from its traceback and the other with the VS Code debugger.
flowchart LR
S["SAP API<br/>(sample or sandbox)"] --> R["with_retries<br/>+ check"]
R -->|"data"| P["sturdy_result.json"]
R -->|"events"| L["logs/sturdy_calls.jsonl"]
R -->|"screen"| T["Terminal"]
P -. "--llm" .-> M["Claude summary"]
Before you start: complete Set up your computer for this course. It creates your orchestrate-course folder with its .venv, installs requests, anthropic and python-dotenv, stores SAP_API_KEY in .env, and installs VS Code with the Python extension. Calling your first SAP API explains the Sales Order API used here.
In VS Code's file list, right-click unit01, choose New File and name it sturdy_calls.py.
Paste the whole script below and save with Ctrl+S (Windows, Linux) or Cmd+S (macOS).
"""Errors, logging and debugging: read blocked sales orders without falling over.
How to run (from the orchestrate-course folder, with .venv turned on):
python unit01/sturdy_calls.py --sample made-up data; everything works
python unit01/sturdy_calls.py --sample --scenario flaky 503, 429 and a timeout, then success
python unit01/sturdy_calls.py --sample --scenario partial one order's items are missing (404)
python unit01/sturdy_calls.py --sample --scenario badkey wrong key: stops at once, no retries
python unit01/sturdy_calls.py --sample --scenario down service down: retries, then stops
python unit01/sturdy_calls.py SAP's sandbox (needs SAP_API_KEY in .env)
python unit01/sturdy_calls.py --sample --llm also ask Claude for a summary (needs ANTHROPIC_API_KEY)
Options:
--log-level DEBUG show more on screen (the log file always keeps DEBUG)
--max-attempts 4 how many times to try one call before giving up
Exit codes: 0 all good, 1 finished but some orders failed,
2 stopped by a problem retrying can't fix, 3 stopped because SAP stayed unavailable.
"""
from __future__ import annotations
import argparse
import json
import logging
import os
import random
import re
import sys
import time
import uuid
from dataclasses import dataclass, field
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
SERVICE = "/sap/opu/odata/sap/API_SALES_ORDER_SRV"
SANDBOX = "https://sandbox.api.sap.com/s4hanacloud" + SERVICE
ORDER_FIELDS = ["SalesOrder", "SoldToParty", "TotalNetAmount", "TransactionCurrency",
"DeliveryBlockReason", "HeaderBillingBlockReason"]
ITEM_FIELDS = ["SalesOrderItem", "Material", "RequestedQuantity", "NetAmount"]
TRANSIENT_STATUS = {429, 502, 503, 504} # "try again later" answers
MAX_WAIT_S = 30.0 # never sleep longer than this between tries
log = logging.getLogger("sturdy_calls")
# ------------------------------------------------------------------ our own exceptions
class SAPCallError(Exception):
"""A call to SAP failed. Carries the details support will ask for."""
def __init__(self, message: str, *, status: int | None = None, sap_code: str | None = None,
transaction_id: str | None = None, path: str = ""):
super().__init__(message)
self.status = status
self.sap_code = sap_code
self.transaction_id = transaction_id
self.path = path
class TransientError(SAPCallError):
"""Might work later: 429, 502, 503, 504, a timeout or a dropped connection."""
def __init__(self, message: str, *, retry_after: float | None = None, **details):
super().__init__(message, **details)
self.retry_after = retry_after
class PermanentError(SAPCallError):
"""Will fail the same way every time: wrong key, wrong address, bad request."""
class SetupError(Exception):
"""Something on our side is wrong, such as a missing or refused key. Fix it, then rerun."""
@dataclass
class Reply:
status: int
text: str
headers: dict = field(default_factory=dict)
# ------------------------------------------------------------------ reading SAP's answers
def parse_sap_error(text: str) -> tuple[str | None, str, str | None]:
"""Pull code, message and transaction ID out of an OData error body, if it is one."""
try:
err = json.loads(text)["error"]
message = err.get("message", "")
if isinstance(message, dict): # OData V2 puts the text in message.value
message = message.get("value", "")
inner = err.get("innererror") or {}
return err.get("code"), str(message), inner.get("transactionid")
except (ValueError, KeyError, TypeError, AttributeError):
return None, (text or "").strip()[:200] or "(empty body)", None
def parse_retry_after(value: str | None) -> float | None:
"""Retry-After is either a number of seconds or an HTTP date."""
if not value:
return None
if value.strip().isdigit():
return float(value.strip())
try:
when = parsedate_to_datetime(value)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError):
return None
def check(reply: Reply, path: str) -> Reply:
"""Return the reply if it succeeded; otherwise raise the right kind of error."""
if 200 <= reply.status < 300:
return reply
code, message, txn = parse_sap_error(reply.text)
details = {"status": reply.status, "sap_code": code, "transaction_id": txn, "path": path}
if reply.status in TRANSIENT_STATUS:
raise TransientError(f"HTTP {reply.status}: {message}",
retry_after=parse_retry_after(reply.headers.get("Retry-After")), **details)
raise PermanentError(f"HTTP {reply.status}: {message}", **details)
# ------------------------------------------------------------------ retries
STATS = {"calls": 0, "retries": 0, "waited_s": 0.0}
def with_retries(call, what: str, *, max_attempts: int, rng: random.Random,
base: float = 0.5, cap: float = 8.0):
"""Run call(); on a TransientError wait and try again, up to max_attempts in total.
Only wrap calls that are safe to repeat (GET). PermanentError is never retried."""
for attempt in range(1, max_attempts + 1):
STATS["calls"] += 1
try:
result = call()
except TransientError as err:
if attempt == max_attempts:
log.error("%s: gave up after %d attempts (%s)", what, attempt, err,
extra={"status": err.status, "attempt": attempt})
raise
if err.retry_after is not None:
wait = min(err.retry_after, MAX_WAIT_S) # the server told us
else:
wait = rng.uniform(0, min(cap, base * 2 ** attempt)) # backoff with full jitter
log.warning("%s: %s; retry %d of %d in %.1f s", what, err, attempt, max_attempts - 1,
wait, extra={"status": err.status, "attempt": attempt, "wait_s": round(wait, 2)})
STATS["retries"] += 1
STATS["waited_s"] += wait
time.sleep(wait)
else:
if attempt > 1:
log.info("%s: succeeded on attempt %d", what, attempt, extra={"attempt": attempt})
return result
# ------------------------------------------------------------------ talking to SAP
class LiveSAP:
"""Real GET requests to SAP's sandbox, with timeouts."""
def __init__(self, key: str):
import requests # imported here so --sample works without the library
self.requests = requests
self.session = requests.Session()
self.session.headers.update({"APIKey": key, "Accept": "application/json"})
def get(self, path: str, params: dict | None = None) -> Reply:
exc = self.requests.exceptions
try:
resp = self.session.get(SANDBOX + path, params=params, timeout=(5, 30))
except (exc.Timeout, exc.ConnectionError) as err:
raise TransientError(f"could not reach SAP ({type(err).__name__})", path=path) from err
except exc.RequestException as err:
raise PermanentError(f"request could not be sent ({type(err).__name__})", path=path) from err
return Reply(resp.status_code, resp.text, dict(resp.headers))
SAMPLE_ORDERS = [
{"SalesOrder": "9000001", "SoldToParty": "CUST-A", "TotalNetAmount": "18250.00",
"TransactionCurrency": "USD", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000002", "SoldToParty": "CUST-B", "TotalNetAmount": "9500.00",
"TransactionCurrency": "USD", "DeliveryBlockReason": "", "HeaderBillingBlockReason": "02"},
{"SalesOrder": "9000003", "SoldToParty": "CUST-C", "TotalNetAmount": "4200.00",
"TransactionCurrency": "EUR", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000004", "SoldToParty": "CUST-A", "TotalNetAmount": "31000.00",
"TransactionCurrency": "USD", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000005", "SoldToParty": "CUST-D", "TotalNetAmount": "760.00",
"TransactionCurrency": "EUR", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
]
SAMPLE_ITEMS = {
"9000001": [{"SalesOrderItem": "10", "Material": "PUMP-100", "RequestedQuantity": "10", "NetAmount": "15000.00"},
{"SalesOrderItem": "20", "Material": "SEAL-7", "RequestedQuantity": "50", "NetAmount": "3250.00"}],
"9000002": [{"SalesOrderItem": "10", "Material": "VALVE-3", "RequestedQuantity": "19", "NetAmount": "9500.00"}],
"9000004": [{"SalesOrderItem": "10", "Material": "MOTOR-9", "RequestedQuantity": "4", "NetAmount": "31000.00"}],
}
def odata_error(status: int, message: str, n: int) -> str:
"""An error body shaped like an OData V2 error. Code and transaction ID are made up."""
return json.dumps({"error": {"code": f"SAMPLE/{status}",
"message": {"lang": "en", "value": message},
"innererror": {"transactionid": f"SAMPLE-TXN-{status}-{n}"}}})
class SampleSAP:
"""Answers like the sandbox, from made-up data, and fails on purpose per scenario."""
def __init__(self, scenario: str):
self.scenario = scenario
self.seen: dict[str, int] = {}
def get(self, path: str, params: dict | None = None) -> Reply:
n = self.seen[path] = self.seen.get(path, 0) + 1 # how often we've been asked this
s = self.scenario
if path == "/A_SalesOrder":
if s == "badkey":
return Reply(401, "Invalid ApiKey")
if s == "down" or (s == "flaky" and n == 1):
return Reply(503, odata_error(503, "Service temporarily unavailable", n))
if s == "flaky" and n == 2:
return Reply(429, "Too many requests", {"Retry-After": "1"})
top = int((params or {}).get("$top", 50))
return Reply(200, json.dumps({"d": {"results": SAMPLE_ORDERS[:top]}}))
match = re.fullmatch(r"/A_SalesOrder\('(\d+)'\)/to_Item", path)
if match:
order = match.group(1)
if s == "flaky" and order == "9000002" and n == 1:
raise TransientError("could not reach SAP (ReadTimeout)", path=path)
if s == "partial" and order == "9000004":
return Reply(404, odata_error(404, f"Resource not found for order {order}", n))
return Reply(200, json.dumps({"d": {"results": SAMPLE_ITEMS.get(order, [])}}))
return Reply(404, odata_error(404, "Unknown path", n))
# ------------------------------------------------------------------ logging setup
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
KEY_IN_TEXT = re.compile(r"(?i)\b(apikey|api_key|authorization|x-api-key)(\s*[:=]\s*)(\S+)")
class RunContext(logging.Filter):
"""Adds the run ID to every record and removes secrets and personal data from it."""
def __init__(self, run_id: str, secrets: list[str]):
super().__init__()
self.run_id = run_id
self.secrets = [s for s in secrets if s and len(s) >= 8]
def mask(self, text: str) -> str:
for secret in self.secrets:
text = text.replace(secret, "***")
text = KEY_IN_TEXT.sub(r"\1\2***", text)
return EMAIL.sub("<email>", text)
def filter(self, record: logging.LogRecord) -> bool:
record.run_id = self.run_id
message = self.mask(record.getMessage())
record.msg, record.args = message.replace("\r", " ").replace("\n", " "), None # one event, one line
return True
class MaskedText(logging.Formatter):
"""Plain text for the screen, with secrets masked in tracebacks too."""
def __init__(self, fmt: str, context: RunContext):
super().__init__(fmt)
self.context = context
def formatException(self, ei) -> str:
return self.context.mask(super().formatException(ei))
class JsonLines(logging.Formatter):
"""One JSON object per line: easy for people to grep and for log tools to search."""
EXTRA = ("status", "sap_code", "transaction_id", "attempt", "wait_s", "order", "path", "request_id")
def __init__(self, context: RunContext):
super().__init__()
self.context = context
def format(self, record: logging.LogRecord) -> str:
entry = {"time": datetime.fromtimestamp(record.created, timezone.utc).isoformat(timespec="milliseconds"),
"level": record.levelname, "logger": record.name,
"run_id": getattr(record, "run_id", None), "message": record.getMessage()}
for name in self.EXTRA:
if getattr(record, name, None) is not None:
entry[name] = getattr(record, name)
if record.exc_info:
entry["exception"] = self.context.mask(self.formatException(record.exc_info))
return json.dumps(entry, ensure_ascii=False)
def setup_logging(level: str, run_id: str, secrets: list[str], log_dir: str = "logs") -> str:
os.makedirs(log_dir, exist_ok=True)
path = os.path.join(log_dir, "sturdy_calls.jsonl")
context = RunContext(run_id, secrets)
screen = logging.StreamHandler() # writes to stderr
screen.setLevel(level)
screen.setFormatter(MaskedText("%(levelname)-7s %(message)s", context))
screen.addFilter(context)
file = logging.FileHandler(path, encoding="utf-8") # appends; one line per event
file.setLevel(logging.DEBUG)
file.setFormatter(JsonLines(context))
file.addFilter(context)
# Libraries (requests, anthropic, ...) log too. Keep only their warnings; keep all of ours.
logging.basicConfig(level=logging.WARNING, handlers=[screen, file], force=True)
log.setLevel(logging.DEBUG)
return path
# ------------------------------------------------------------------ the work
def is_blocked(order: dict) -> bool:
return bool(order.get("DeliveryBlockReason") or order.get("HeaderBillingBlockReason"))
def read_orders(api, top: int, attempts: int, rng: random.Random) -> list[dict]:
path = "/A_SalesOrder"
params = {"$top": str(top), "$select": ",".join(ORDER_FIELDS), "$format": "json"}
reply = with_retries(lambda: check(api.get(path, params), path), "read orders",
max_attempts=attempts, rng=rng)
return json.loads(reply.text)["d"]["results"]
def read_items(api, order: str, attempts: int, rng: random.Random) -> list[dict]:
path = f"/A_SalesOrder('{order}')/to_Item"
params = {"$select": ",".join(ITEM_FIELDS), "$format": "json"}
reply = with_retries(lambda: check(api.get(path, params), path), f"items of {order}",
max_attempts=attempts, rng=rng)
return json.loads(reply.text)["d"]["results"]
def summarize_with_claude(result: dict) -> str | None:
"""Optional: ask a model for two sentences. Fails gracefully: the data file matters more."""
import anthropic # imported here so the rest runs without it
client = anthropic.Anthropic(max_retries=3, timeout=60.0) # the SDK retries 429 and 5xx itself
facts = {"blocked": [{k: o[k] for k in ("sales_order", "customer", "net_amount", "currency")}
for o in result["orders"]],
"could_not_read": [f["sales_order"] for f in result["failed"]]}
try:
message = client.messages.create(
model=os.environ.get("LLM_MODEL", "claude-opus-5-5"),
max_tokens=200,
messages=[{"role": "user", "content":
"In two sentences for a credit manager, summarize these blocked sales orders. "
"Use only these facts.\n" + json.dumps(facts)}])
except anthropic.AuthenticationError as err: # a wrong key never fixes itself: fail loudly
raise SetupError("Claude refused the key: check ANTHROPIC_API_KEY in .env") from err
except (anthropic.RateLimitError, anthropic.APIConnectionError, anthropic.InternalServerError) as err:
log.warning("summary skipped after the SDK's retries: %s", type(err).__name__)
return None # the order data is still saved
log.info("summary received", extra={"request_id": message._request_id})
return "".join(block.text for block in message.content if block.type == "text")
def main() -> int:
parser = argparse.ArgumentParser(description="Read blocked sales orders, with retries and logs.")
parser.add_argument("--sample", action="store_true", help="made-up data; no key or internet")
parser.add_argument("--scenario", default="ok", choices=["ok", "flaky", "partial", "badkey", "down"],
help="which failure the sample data acts out")
parser.add_argument("--llm", action="store_true", help="also ask Claude for a summary")
parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
parser.add_argument("--max-attempts", type=int, default=4, help="tries per call, including the first")
parser.add_argument("--top", type=int, default=50, help="read at most this many orders")
args = parser.parse_args()
try:
from dotenv import load_dotenv
load_dotenv()
except ImportError:
if not args.sample:
print("python-dotenv is missing: run pip install -r requirements.txt", file=sys.stderr)
return 2
# Check the settings first: a missing key should stop the run before any work is done.
if args.llm and not os.environ.get("ANTHROPIC_API_KEY"):
print("--llm needs ANTHROPIC_API_KEY in .env (or leave out --llm)", file=sys.stderr)
return 2
sap_key = os.environ.get("SAP_API_KEY", "")
run_id = uuid.uuid4().hex[:8]
log_path = setup_logging(args.log_level, run_id, [sap_key, os.environ.get("ANTHROPIC_API_KEY", "")])
rng = random.Random(7) if args.sample else random.Random() # fixed seed: same waits every sample run
if args.sample:
api, source = SampleSAP(args.scenario), f"sample ({args.scenario})"
else:
if not sap_key:
log.error("SAP_API_KEY is not set. Add it to .env, or run with --sample")
return 2
api, source = LiveSAP(sap_key), "sandbox"
print(f"Run {run_id}: reading from {source}. Log file: {log_path}")
log.info("run started: source=%s", source)
log.debug("settings: max_attempts=%d top=%d", args.max_attempts, args.top)
# 1. Read the orders. Without them there is nothing to do, so failures here stop the run.
try:
orders = read_orders(api, args.top, args.max_attempts, rng)
except PermanentError as err:
log.error("stopped: %s", err, extra={"status": err.status, "sap_code": err.sap_code,
"transaction_id": err.transaction_id})
if err.status in (401, 403):
print("The key was refused. Check SAP_API_KEY in .env (Show API Key on api.sap.com).")
return 2
except TransientError:
print("SAP stayed unavailable. Nothing was written; run again later.")
return 3
blocked = [o for o in orders if is_blocked(o)]
print(f"Read {len(orders)} orders; {len(blocked)} have a delivery or billing block")
if not blocked:
blocked = orders[:2]
print("No blocked orders in this data, so reading items for the first two anyway")
# 2. Read items per order. One bad order must not sink the others: record it and go on.
result = {"run_id": run_id, "source": source, "orders": [], "failed": []}
for order in blocked:
number = order["SalesOrder"]
try:
items = read_items(api, number, args.max_attempts, rng)
except SAPCallError as err:
log.warning("order %s skipped: %s", number, err,
extra={"order": number, "status": err.status, "sap_code": err.sap_code,
"transaction_id": err.transaction_id})
result["failed"].append({"sales_order": number, "status": err.status,
"reason": str(err), "transaction_id": err.transaction_id})
continue
log.debug("order %s: %d item(s)", number, len(items), extra={"order": number})
result["orders"].append({"sales_order": number, "customer": order["SoldToParty"],
"net_amount": order["TotalNetAmount"],
"currency": order["TransactionCurrency"], "items": len(items)})
with open("sturdy_result.json", "w", encoding="utf-8") as out:
json.dump(result, out, indent=2)
print(f"Saved {len(result['orders'])} orders, {len(result['failed'])} failed, to sturdy_result.json")
print(f"Calls: {STATS['calls']}, retries: {STATS['retries']}, waited {STATS['waited_s']:.1f} s")
if args.llm:
try:
summary = summarize_with_claude(result)
except SetupError as err:
log.error("stopped: %s", err)
return 2
print("Summary:", summary or "(skipped; see the log)")
log.info("run finished: %d saved, %d failed", len(result["orders"]), len(result["failed"]))
return 1 if result["failed"] else 0
if __name__ == "__main__":
sys.exit(main())
On macOS or Linux, use python3 if python isn't found.
You should see this. Your run ID, the eight characters after Run, will be different every time:
Run 24073cd4: reading from sample (ok). Log file: logs/sturdy_calls.jsonl
INFO run started: source=sample (ok)
Read 5 orders; 3 have a delivery or billing block
Saved 3 orders, 0 failed, to sturdy_result.json
Calls: 4, retries: 0, waited 0.0 s
INFO run finished: 3 saved, 0 failed
Notice the two kinds of lines. Lines starting with INFO come from the logger; the others come from print(). Results are for the person at the terminal; log lines are the record of what happened.
Four calls: one for the order list, then one for the items of each of the three blocked orders.
Run the flaky scenario. The sample service answers the first order call with 503, the second with 429 and Retry-After: 1, and lets one item call time out once:
Run 55548a7a: reading from sample (flaky). Log file: logs/sturdy_calls.jsonl
INFO run started: source=sample (flaky)
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 1 of 3 in 0.3 s
WARNING read orders: HTTP 429: Too many requests; retry 2 of 3 in 1.0 s
INFO read orders: succeeded on attempt 3
Read 5 orders; 3 have a delivery or billing block
WARNING items of 9000002: could not reach SAP (ReadTimeout); retry 1 of 3 in 0.2 s
INFO items of 9000002: succeeded on attempt 2
Saved 3 orders, 0 failed, to sturdy_result.json
Calls: 7, retries: 3, waited 1.5 s
INFO run finished: 3 saved, 0 failed
Read the waits. After the 503 the script picked a random wait with jitter (0.3 s). After the 429 it waited exactly 1.0 s, because the answer's Retry-After header said so. In sample mode the random numbers are seeded, so you get the same waits every time; against a real service they differ on every run.
The end result is the same as Step 4. Three transient failures cost three extra calls and 1.5 seconds, and nobody had to do anything.
#Step 6: Watch one bad order being skipped, not the whole run
Run the partial scenario. The items of order 9000004 come back as 404:
Run 70265a47: reading from sample (partial). Log file: logs/sturdy_calls.jsonl
INFO run started: source=sample (partial)
Read 5 orders; 3 have a delivery or billing block
WARNING order 9000004 skipped: HTTP 404: Resource not found for order 9000004
Saved 2 orders, 1 failed, to sturdy_result.json
Calls: 4, retries: 0, waited 0.0 s
INFO run finished: 2 saved, 1 failed
No retries: a 404 is permanent, so trying again would only fail again. The order was recorded and the other two were saved.
Look at the exit code, the number a scheduler would see:
Windows (PowerShell):
$LASTEXITCODE
macOS / Linux:
echo $?
You should see 1: "finished, but some orders failed". A scheduler or pipeline can alert on that, even though the run didn't crash.
Open sturdy_result.json in VS Code. Under "failed" you see order 9000004 with its status, reason and "transaction_id": "SAMPLE-TXN-404-1". In a real system that ID is what an SAP administrator would look up.
Run 9c76bf35: reading from sample (badkey). Log file: logs/sturdy_calls.jsonl
INFO run started: source=sample (badkey)
ERROR stopped: HTTP 401: Invalid ApiKey
The key was refused. Check SAP_API_KEY in .env (Show API Key on api.sap.com).
No retries, a clear instruction, and exit code 2.
Run the down scenario. The service answers 503 every time:
python unit01/sturdy_calls.py --sample --scenario down
You should see three retries with growing random waits, then a clear stop with exit code 3. It takes about three seconds:
Run 1e4c9f32: reading from sample (down). Log file: logs/sturdy_calls.jsonl
INFO run started: source=sample (down)
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 1 of 3 in 0.3 s
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 2 of 3 in 0.3 s
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 3 of 3 in 2.6 s
ERROR read orders: gave up after 4 attempts (HTTP 503: Service temporarily unavailable)
SAP stayed unavailable. Nothing was written; run again later.
Notice that the second wait (0.3 s) is shorter than you might expect. That is full jitter: each wait is random between zero and a limit that doubles (1, 2, 4 seconds), so waits spread out instead of lining up.
Try --max-attempts 2 with the same scenario. The script gives up after one retry.
In VS Code's file list, open logs/sturdy_calls.jsonl. Every run so far has added lines. Each line is one JSON object with time, level, run_id and message, plus fields such as status, attempt, wait_s, order and transaction_id where they apply.
The file also has DEBUG lines the screen didn't show, such as settings: max_attempts=4 top=50. The screen handler shows INFO and above; the file handler keeps everything.
Check that .env has your SAP_API_KEY. If not, follow the key steps in the setup topic.
Run without --sample:
python unit01/sturdy_calls.py
You should see reading from sandbox and the same kind of summary with SAP's demo data. Valid differences:
The number of orders is up to 50, the default --top.
0 have a delivery or billing block is a valid answer about demo data. The script then reads items for the first two orders so you still see item calls.
If your network blocks the sandbox, you see could not reach SAP (ConnectionError) or (ProxyError) with retries, then exit code 3. That is the script doing its job: the failure is transient from its point of view.
#Step 10 (optional): Add a model summary, and see a graceful skip
If you haven't yet, add ANTHROPIC_API_KEY="..." to .env, as in the setup topic.
Run:
python unit01/sturdy_calls.py --sample --llm
After the usual lines you see INFO summary received and a line starting with Summary:. The log file entry for summary received carries a request_id; that is what a provider's support team asks for.
Look at summarize_with_claude in the script. The Anthropic library is created with max_retries=3 and timeout=60.0, so the library itself retries rate limits and server errors. The function only decides what happens after that: a refused key stops the run (exit code 2); a rate limit or outage that outlasts the retries skips the summary but keeps the saved data.
Without a key, --llm stops before any work with --llm needs ANTHROPIC_API_KEY in .env (or leave out --llm). Checking settings first means you never wait for a whole run only to fail at the end.
Now the debugging half. buggy_totals.py should list the blocked orders above 10,000 that need a credit manager's review. The right answer is ['9000001', '9000004'].
Create unit01/buggy_totals.py and paste this. Don't fix anything yet:
"""Practice file for debugging: it has two bugs on purpose.
Run (from the orchestrate-course folder): python unit01/buggy_totals.py
The right answer is: Blocked orders above 10,000: ['9000001', '9000004']
"""
ORDERS = [ # shaped like A_SalesOrder from SAP's Sales Order API; amounts arrive as text
{"SalesOrder": "9000001", "TotalNetAmount": "18250.00", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000002", "TotalNetAmount": "9500.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": "02"},
{"SalesOrder": "9000003", "TotalNetAmount": "4200.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000004", "TotalNetAmount": "31000.00", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
{"SalesOrder": "9000005", "TotalNetAmount": "760.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
]
LIMIT = "10000.00" # blocked orders above this value go to a credit manager
def is_blocked(order):
return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])
def needs_review(order):
return order["TotalNetAmount"] > LIMIT
def main():
review = []
for order in ORDERS:
if is_blocked(order) and needs_review(order):
review.append(order["SalesOrder"])
print("Blocked orders above 10,000:", review)
if __name__ == "__main__":
main()
Run it:
python unit01/buggy_totals.py
You should see a traceback ending in:
File ".../unit01/buggy_totals.py", line 18, in is_blocked
return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])
~~~~~^^^^^^^^^^^^^^^^^
KeyError: 'DeliveryBlock'
Read it bottom up: a KeyError for 'DeliveryBlock' on line 18, in is_blocked. The orders have DeliveryBlockReason, the field name in SAP's Sales Order API.
On line 18, change order["DeliveryBlock"] to order["DeliveryBlockReason"]. Save and run again.
In buggy_totals.py, click in the margin just left of the line number of return order["TotalNetAmount"] > LIMIT (line 22). A red dot appears: that is a breakpoint.
Open the Run and Debug view: the icon with a play button and a bug on the left, or Ctrl+Shift+D (Cmd+Shift+D on macOS).
Click Run and Debug. If VS Code asks what to debug, choose Python Debugger, then Python File. You can also press F5 with the file open.
The program stops at the red dot, and the line is highlighted. In the Variables section, expand order. It is 9000001, and TotalNetAmount shows '18250.00', in quotes.
Click the Debug Console tab at the bottom and type each line, pressing Enter after each:
You should see <class 'str'>, True, and True. The amount is text, and text is compared character by character: "9" comes after "1", so "9500.00" counts as larger than "10000.00".
Press F5 (Continue) twice. The second stop is order 9000002; look at its amount in Variables. Then press F5 until the program ends, or click the red square (Stop).
Fix the cause. At the top of the file, under the docstring, add:
from decimal import Decimal
Then change the limit and the comparison so both are numbers:
LIMIT = Decimal("10000.00") # blocked orders above this value go to a credit manager
Bins, not cases.check decides transient or permanent in one place. Every other function just reacts to the exception type.
Retries only around reads. Every call wrapped in with_retries is a GET. A write would need its own check-before-retry logic.
Two outputs, two audiences.print() for the person, the logger for the record. The JSON file is for tools, the screen for people.
Exit codes mean something. 0, 1, 2 and 3 let a scheduler tell "all good", "partly done", "fix the setup" and "try later" apart without reading text.
Sample and live share code. Only the class that sends requests changes, so the failure handling you watched in sample mode is the same code that runs against SAP.
When your OData call to an S/4HANA or other ABAP-based system fails, SAP keeps its own record. SAP Learning's Gateway course describes:
SAP Gateway error log (/IWFND/ERROR_LOG): errors found while the OData request is analyzed, for example a request that doesn't fit the service's data model.
SAP Gateway backend error log (/IWBEP/ERROR_LOG): errors while the service implementation processes a valid request, such as an invalid key value or a coding error.
Error details with an overview by error ID and time, the error context, the call stack, the full request and response, and a Replay button.
SAP Gateway Client (/IWFND/GW_CLIENT) to send test requests and jump to the error log.
Two consequences for your code. Log the transaction ID and timestamp from SAP's error body, because they are how an administrator finds the matching entry. And plan who can look: these logs live in the customer's system, so agree with the SAP basis team early how errors get analyzed during the project. The course source covers SAP Gateway in ABAP-based systems; for SAP S/4HANA Cloud Public Edition, ask SAP or your partner which monitoring tools apply.
#On SAP BTP: application logs and SAP Cloud Logging
An app deployed to the SAP BTP Cloud Foundry runtime (Unit 6, Deploying AI apps on SAP BTP) writes its logs to standard output and standard error, the screen of a program with no screen. The platform collects them and cf logs <app> --recent shows the latest lines. That is why sturdy_calls.py logs to the terminal and why one-JSON-object-per-line matters: the platform passes lines through as they are.
For storage, search and dashboards, SAP Learning describes SAP Cloud Logging: it ingests, stores and analyzes application logs from custom apps on SAP BTP, accepts data over the OpenTelemetry standard, comes with prebuilt dashboards, and links from SAP Cloud ALM, so an operator can go from a central alert to the detailed log lines. SAP also publishes an open-source Python library, sap_cf_logging, that writes JSON logs with a correlation ID for Cloud Foundry apps; check its release activity before you depend on it.
When your model calls go through SAP's generative AI hub instead of directly to a provider, the orchestration service can run a list of configurations as fallback: SAP's SDK documentation says it moves to the next one when the model is unavailable, on a timeout, on too many requests or on internal service errors, and not on content filter hits or invalid input. That is the same transient-versus-permanent split as this topic, done by the platform. SAP Generative AI Hub and the orchestration service shows how to set it up, and how to read the failed attempts it reports.
Secrets and personal data. Mask at the source, as RunContext does, and also check what libraries log at DEBUG level; keep third-party loggers at WARNING in production. Never log prompts or model answers by default when they can contain customer or employee data. If you must keep them for evaluation, store them separately, with access control and a retention period agreed with the customer's data protection officer.
Who can read logs. Logs are data. Decide with the customer who may read the logs in each place they end up (your log platform, SAP Cloud Logging, SAP's own error logs) before go-live, not after the first incident.
Writes to SAP. Never put a create or change call inside a blind retry loop. Before retrying, read back to check whether the first call worked, or send your own unique reference with each create and look for it. For AI agents that propose SAP changes, a person approves first (Unit 11 covers agent permissions).
Retry budgets and cost. Cap attempts per call and total wait per run. Every retried model call may be billed and counts toward rate limits. Log retries as WARNING and track how many you see per day; a rising count is an early outage signal.
Timeouts everywhere. The requests library waits without limit unless you pass a timeout, and Anthropic's library defaults to 10 minutes. Set timeouts that match the job: an interactive app needs seconds, a nightly batch can afford more.
Alerting. A log nobody reads is a diary, not monitoring. Alert on exit codes other than 0, on ERROR entries, and on a share of failed records above a threshold. Unit 10 builds this into full observability.
Correlation across systems. Carry one ID from the user's request through your service, the model call and the SAP call. Log SAP's transaction ID and the model provider's request ID next to it.
Clean core. Don't change SAP code to add logging or error handling for your AI app. Use SAP's standard logs and APIs, and do your handling in the side-by-side app.
except Exception: pass. The single most expensive line in enterprise Python. It turns every bin-3 bug into a silent wrong answer.
No timeout. A call that hangs forever holds a batch job and its schedule slot. Always pass timeout=.
Retrying permanent errors. Retrying a 401 delays the message someone needs, and repeated refused logins can trigger lockouts.
Retrying without jitter. Many jobs retrying on the same schedule hit the recovering service together.
Logging the whole request "for now". Debug code that logs headers or prompts ends up in production. Log IDs and counts instead.
Treating text amounts as numbers. SAP's OData V2 APIs send amounts as text. Comparing or adding them as text gives wrong answers without errors, as Step 11 showed. Convert with Decimal once, at the edge.
logging.basicConfig doing nothing. It has no effect if logging was already configured, for example by a library. Pass force=True, as the script does, or configure logging once at the start of main.
Debugging live data first. Reproduce on sample data. Pausing a program in the middle of a live SAP session can leave half-done work and hold locks.
#Exercise: add a run summary and prove the logs are clean
The Unit 10 topics on observability use run summaries like this one to build dashboards and alerts. You will add one, then write a small check that no key ever reaches the log file.
Open unit01/sturdy_calls.py. In main, find the line with open("sturdy_result.json", "w", encoding="utf-8") as out:.
Just above it, add these lines, indented like the code around them:
Run --scenario flaky and check that retries is 3 and failed_by_status is {}.
Create unit01/check_logs.py with this code. It reads every value in .env whose name ends in _API_KEY and searches the log file for it:
"""Fail if any API key from .env appears in the log file."""
import sys
from pathlib import Path
from dotenv import dotenv_values
keys = [v for k, v in dotenv_values(".env").items() if k.endswith("_API_KEY") and v]
log_file = Path("logs/sturdy_calls.jsonl")
text = log_file.read_text(encoding="utf-8") if log_file.exists() else ""
leaks = sum(text.count(key) for key in keys)
print(f"Checked {len(keys)} key(s) against {log_file}: {leaks} leak(s)")
sys.exit(1 if leaks else 0)
Run it:
python unit01/check_logs.py
You should see Checked 2 key(s) against logs/sturdy_calls.jsonl: 0 leak(s) (the number of keys depends on your .env).
Commit both files with a message such as Unit 1: run summary and log check.
Done when:sturdy_result.json from a partial run ends with a run_summary showing "failed_by_status": {"404": 1}, and check_logs.py prints 0 leak(s) and exits with code 0.
Pick one answer for each question. The explanation appears after you choose.
1In sturdy_calls.py, why is a 404 for one order's items not retried?
Answer: B. check sorts 404 into PermanentError, and with_retries only retries TransientError. The main loop then records the order and moves on, which is the graceful part.
2The order list answers 429 with Retry-After: 1. How long does with_retries wait before the next try?
Answer: C. When Retry-After is present, the script uses it instead of its own jittered backoff, but never waits longer than MAX_WAIT_S. Step 5 shows exactly this: retry 2 of 3 in 1.0 s.
3Why does the retry loop use full jitter instead of waiting exactly 1, 2 and 4 seconds?
Answer: A. Plain backoff keeps failed clients in step, so they hit the recovering service together. A random wait between zero and the limit spreads them out. The fixed seed in sample mode is only there so the walkthrough output is repeatable.
4You extend the script to create sales orders with POST. A create call times out. What should the code do?
Answer: D. POST is not idempotent. The first call may have created the order even though the answer never arrived, so a blind retry risks a duplicate. Read back, or search for your own reference, before retrying.
5What is the job of the RunContext filter?
Answer: B. The filter runs on every record before a handler writes it. It attaches run_id and cleans the message; both formatters use the same masking for tracebacks. Removing line breaks from messages also prevents forged log lines.
6buggy_totals.py says order 9000002, worth 9,500, is above 10,000. What does the debugger show as the cause?
Answer: C. In the Debug Console, type(order["TotalNetAmount"]) shows str, and "9500.00" > "10000.00" is True because "9" sorts after "1". Converting both sides to Decimal fixes the cause.
7A batch job fails at 3 a.m. with an HTTP 500 from an S/4HANA OData service. What in your log helps the SAP administrator most?
Answer: D. SAP Gateway's error logs list errors by ID and time, with the call stack and the full request and response. Your log should carry the matching ID, and never the key.
8The Anthropic client in the script is created with max_retries=3. What does that mean for summarize_with_claude?
Answer: C. Anthropic's library retries connection errors, 429 and 5xx answers itself. The function only maps the outcome to the business rule: a refused key stops the run; an outage that outlasts the retries skips the summary but keeps the data.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Logging HOWTO (Python documentation)— when to print vs. log vs. raise, the five levels and their meaning, WARNING as default, getLogger(__name__), logger.exception
Debug code with Visual Studio Code (VS Code documentation)— breakpoints in the editor margin or F9, conditional breakpoints, logpoints, Continue F5, Step Over F10, Step Into F11, Step Out, Variables, Watch, Call Stack, Debug Console
Logging Cheat Sheet (OWASP)— data to exclude, mask or hash in logs (access tokens, passwords, keys, connection strings, sensitive personal data); remove CR, LF and delimiters to prevent log injection
RFC 9110: HTTP Semantics— idempotent methods (PUT, DELETE and the safe methods such as GET); clients should not automatically retry non-idempotent requests; Retry-After with 503
Python SDK (Claude Platform documentation)— typed errors (AuthenticationError, RateLimitError, APIConnectionError, APIStatusError), 2 automatic retries by default for connection errors, 408, 409, 429 and 5xx, max_retries, 10-minute default timeout, _request_id
Managing an SAP Gateway Service (SAP Learning)— SAP Gateway error log /IWFND/ERROR_LOG and backend error log /IWBEP/ERROR_LOG, error context, call stack, full request and response, Replay; SAP Gateway Client /IWFND/GW_CLIENT