Orchestrate

Errors, logging and debugging

Make AI code that calls SAP and model APIs retry what is worth retrying, stop on what isn't, leave a safe log trail, and find bugs with a debugger.

Updated Oct 3, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

Every program that talks to SAP or to an AI model will meet failures. The network drops. The model provider says "too many requests". A key expires. A field name is wrong. Someone types bad data.

Good AI code does three things about this:

  • Handles errors on purpose. Some failures go away if you wait and try again. Others never will. Code should tell them apart: retry the first kind a few times, and stop with a clear message on the second kind.
  • Keeps a log. A log is the program's diary: what it did, when, and what went wrong. When a nightly job fails at 3 a.m., the log is how anyone finds out why. It must never contain passwords, keys or personal data.
  • Is debugged, not guessed at. When the result is wrong, a developer pauses the program with a debugger and looks at the actual values, instead of changing code and hoping.

None of this is AI-specific. But AI systems call more outside services than most programs, and they fail in more ways, so these habits matter more.

Why it matters to the business

Take the running example from Unit 1: a job reads blocked sales orders from S/4HANA every night and asks a model to summarize them for the credit team. Here is how weak error handling hurts:

  • Silent failure. The SAP call times out, the script catches the error and carries on with zero orders. The morning report says "no blocked orders". Credit managers relax while 40 orders wait. A wrong answer that looks right is worse than a crash.
  • Retry storms. The model provider is overloaded. Every job retries at once, as fast as it can, and makes the overload worse. If the provider charges per request, failed retries also cost money.
  • Duplicate business documents. A program retries a call that creates something, such as a sales order or a payment proposal. The first call actually worked; the answer just got lost. Now there are two.
  • Leaked secrets. A developer logs a whole request "to debug it", including the API key. Logs are copied to support tickets and monitoring tools. The key is now in five places.
  • Slow fixes. Without a log that says which call failed, with which status and which SAP transaction ID, a two-minute fix becomes a two-day investigation across three teams.

The business decision is simple to state: for every AI job that touches SAP, someone should be able to say what happens when each outside call fails, and where the evidence ends up.

How SAP does it

SAP systems and services already come with error and log tools. Your team's code should fit into them rather than invent its own. As of October 2026:

  • SAP Gateway error logs (S/4HANA and other ABAP systems). When an OData call fails, SAP Gateway can record it. SAP Learning describes two logs: the SAP Gateway error log (/IWFND/ERROR_LOG) for requests that fail when the request is analyzed, and the backend error log (/IWBEP/ERROR_LOG) for failures inside the service's implementation. The error details include the call stack and the full request and response, and a Replay option to repeat the call. That is why an AI app should log any transaction ID SAP returns: it lets an SAP administrator find the matching entry.
  • SAP Cloud Logging (SAP BTP). For apps you build and run on SAP BTP, SAP Learning describes SAP Cloud Logging as the service to ingest, store and analyze application logs, with prebuilt dashboards and support for the OpenTelemetry standard. It links to SAP Cloud ALM, SAP's central monitoring tool, so operations teams can jump from an alert into the detailed logs.
  • Model calls. Model providers' own libraries already retry some failures. Anthropic's Python library, for example, retries connection errors, rate limits and server errors twice by default. SAP's orchestration service can fall back to a second model configuration when the first fails, as covered in SAP Generative AI Hub and the orchestration service.

What SAP does not do for you: decide which failures your job should retry, what it should do with a half-finished batch, and what must never appear in your logs. That is your team's design.

Fail loudly or fail gracefully: a decision guide

"Fail loudly" means stop at once with a clear message and a non-zero result, so a person or a scheduler notices. "Fail gracefully" means record the problem, skip the affected piece and finish the rest. Neither is always right. Use this guide:

Situation Behavior Why
API key missing, wrong or refused Stop at once, no retries Retrying cannot fix it, and repeated refused logins can lock accounts
SAP or the model says "too many requests" (429) or "unavailable" (503) Wait, retry a few times, then stop Usually temporary; but give up after a limit so the job doesn't hang
One sales order out of 200 can't be read Record it, skip it, finish the rest, and report "199 of 200" One bad record shouldn't block the credit team's whole list
Most records fail Stop and alert A broken job that "finishes" hides a real outage
A call that creates or changes SAP data fails with no clear answer Don't retry blindly; check whether it worked first A blind retry can create a duplicate document
The model's summary call fails, but the data was read Save the data, skip the summary, flag it The order list is still useful without the prose
The model returns text that isn't in the expected format Treat as an error: retry once or route to a person Never pass a broken answer on as if it were fine

Questions to ask

Ask your team, vendor or partner:

  • For each outside call (SAP, model, other services): what happens if it times out? Is there a timeout at all?
  • Which failures are retried, how many times, and how long can the job wait in total?
  • Does the job ever retry a call that changes data in SAP? How do you prevent duplicates?
  • When a nightly job partly fails, how do we find out before the business does? Who gets the alert?
  • What exactly goes into the logs? Show me a real log line. Could it contain keys, tokens, customer names or emails?
  • Where are the logs kept, for how long, and who can read them?
  • If SAP returns an error, do we log its transaction ID so an SAP administrator can find it in the system's error log?
  • Can a developer reproduce a production failure on test data, and step through it with a debugger?

Common misconceptions

  • "Catching every error makes the program robust." Catching everything and carrying on is how silent failures happen. Catch only the errors you know how to handle; let the rest stop the program loudly.
  • "Retrying always helps." Retrying a wrong key, a wrong field name or a bad request just fails again, slower. Retrying a call that creates data can make duplicates.
  • "More logging is always better." Logs full of noise hide the line that matters, cost money to store, and are the most common place secrets leak.
  • "The AI provider handles errors for us." Provider libraries retry a few temporary failures. They don't know your business rules: what to skip, what to report, and what a half-finished run means.
  • "Debugging means adding print statements." Prints help, but a debugger shows every value at the moment of the problem without changing the code.

Key terms

  • Exception: Python's signal that something went wrong while the program ran. Unhandled, it stops the program and prints a traceback.
  • Traceback: the report Python prints when it stops: which lines were running, ending with the error type and message.
  • Transient error: a failure that may go away on its own, such as a timeout, a 429 "too many requests" or a 503 "service unavailable".
  • Permanent error: a failure that will repeat until someone changes something, such as a refused key or a wrong address.
  • Retry with backoff: trying again after a wait that grows each time. Jitter adds randomness so many clients don't retry at the same moment.
  • Idempotent: a call that has the same effect whether it runs once or several times. Reading data is idempotent; creating a sales order usually isn't.
  • Log level: how serious a log entry is: DEBUG, INFO, WARNING, ERROR or CRITICAL.
  • Structured log: log entries written as data (for example one JSON object per line), so tools can search by field.
  • Debugger: a tool that pauses a running program at a breakpoint so you can inspect its values and step through it line by line.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A nightly job's SAP call times out. The code catches the error and reports "0 blocked orders". What is the real problem?

    Answer: B. Catching an error and carrying on as if nothing happened produces a wrong answer that looks right. The job should either retry and succeed, or stop and say that SAP could not be read.
  2. 2Which failure should the job retry after a short wait?

    Answer: C. A 429 says "not now", so waiting and trying again can work. A refused key or a wrong field fails the same way every time; retrying only delays the clear message someone needs.
  3. 3A job creates sales orders in SAP. One call fails with no clear answer. What should happen next?

    Answer: B. A call that creates data is not safe to repeat blindly. If the first call actually worked and only the answer was lost, a retry makes a duplicate document.
  4. 4One sales order out of 200 can't be read; the other 199 are fine. What is the best behavior?

    Answer: C. One bad record shouldn't block the credit team's whole list, but the gap must be visible. Reporting "199 of 200, here is the one that failed" is graceful and honest.
  5. 5What should you ask to see before an AI job goes live?

    Answer: B. Logs are copied into tickets and monitoring tools, so a leaked key or customer email spreads quickly. Seeing a real log line is the fastest way to check what the job records.
  6. 6Why should the AI app log the transaction ID that an SAP OData error returns?

    Answer: D. SAP Gateway keeps error logs with the call stack and full request and response. The ID in your log connects your side of the failure to SAP's side, which turns a long investigation into a quick lookup.
Deep layer · 40 min read

Mental model: sort every failure into one of three bins

Every failure your code meets belongs in one of three bins, and each bin has one right response:

  1. Worth waiting for (transient). Timeouts, dropped connections, HTTP 429, 502, 503, 504. Response: wait, retry a limited number of times, then give up loudly.
  2. Worth telling someone (permanent). A refused key (401, 403), a wrong address (404), a bad request (400), a missing setting. Response: stop that piece of work at once, with a message that says what to fix.
  3. A bug in your own code. A KeyError, a TypeError, a wrong answer with no error at all. Response: don't catch it. Let it crash during development, then reproduce it and find it with a debugger.

Around those bins sit two tools. Logs are the flight recorder: they say which bin each failure went into, when, and with which IDs. The debugger is the pause button you use on bin 3.

The rest of this topic builds code that sorts failures into these bins automatically, writes a safe log of what happened, and then uses the debugger on a real bug.

How it works

Exceptions and tracebacks

When something goes wrong while Python runs, it raises an exception: an object that describes the problem, such as KeyError or TimeoutError. If no code handles it, Python stops and prints a traceback.

A traceback lists the chain of function calls that led to the error. Python's tutorial calls the order "most recent call last", so:

  • Read from the bottom up. The last line is the error type and message. The lines just above it show where it happened.
  • Find your own file. Lines inside libraries (paths containing site-packages) usually show where the error surfaced, not where it started. The last line that points into your file is usually where to look.
  • Since Python 3.11, tracebacks underline the exact part of the line with ~ and ^ characters.
Traceback (most recent call last):
  File ".../unit01/buggy_totals.py", line 34, in <module>
    main()
  File ".../unit01/buggy_totals.py", line 28, in main
    if is_blocked(order) and needs_review(order):
       ~~~~~~~~~~^^^^^^^
  File ".../unit01/buggy_totals.py", line 18, in is_blocked
    return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])
                ~~~~~^^^^^^^^^^^^^^^^^
KeyError: 'DeliveryBlock'

Bottom line first: KeyError: 'DeliveryBlock', a dictionary has no such key. One line up: it happened in is_blocked, line 18. The field in SAP's Sales Order API is DeliveryBlockReason. You will fix exactly this bug later.

Handling exceptions: try, except, else, finally

try:
    reply = call_sap()            # code that might fail
except TimeoutError as err:       # runs only if that error happened
    log.warning("SAP timed out: %s", err)
else:                             # runs only if nothing failed
    process(reply)
finally:                          # always runs, error or not
    close_connection()

The Python tutorial gives the rules: one matching except block runs; else runs only when the try block raised nothing; finally always runs last, which makes it the place for clean-up. Putting the success path in else rather than inside try keeps you from accidentally catching errors from code you didn't mean to protect.

Three habits follow:

  • Catch narrowly. except KeyError: says what you expect. A bare except: or a broad except Exception: in the middle of your logic hides bugs from bin 3 together with the failures you meant to handle.
  • Catch where you can decide. A low-level function that sends a request can't know whether one failed order should stop the job. Let the error travel up to the code that can decide, usually the main loop.
  • Raise your own exception types. A custom class, such as class TransientError(Exception), lets the rest of the code react to meaning ("retry this") instead of to a library's details. Python's tutorial recommends deriving from Exception and ending the name in Error.

When you turn a library's error into your own, chain them with raise ... from err. The traceback then shows both: your clear message and the original cause. (raise ... from None hides the cause; use it rarely.) Python 3.11 also added err.add_note("...") to attach context, such as which sales order was being read, without changing the error type.

Transient or permanent: deciding from the HTTP answer

For calls over HTTP, the status code is the first clue to the bin:

Answer Typical meaning Bin
Timeout or dropped connection Network, proxy or an overloaded server Transient
429 Too Many Requests You are being rate-limited Transient: wait, ideally as long as the server says
502, 503, 504 A gateway or the service is temporarily unavailable Transient
400 Bad Request Something in the request is wrong, such as a field name Permanent
401 Unauthorized, 403 Forbidden Key missing, wrong, or not allowed Permanent
404 Not Found Wrong address or a record that doesn't exist Permanent
500 Internal Server Error The server failed; the cause could be either Depends: retrying once is common, but log it as an error

Treat this as a starting point, not a law. Check each API's own documentation. Anthropic's error reference, for example, uses its own 529 status for "overloaded", and its Python library retries 408, 409, 429 and 5xx answers by default.

An SAP OData V2 service returns error details in the response body as JSON. The shape is an error object with a code, a message whose text is in message.value, and often an innererror with more detail. SAP systems commonly put a transactionid there. The exact content varies by system and service, so read it defensively: if the body isn't the shape you expect, keep the first part of the raw text instead of crashing while handling an error.

Retries that don't make things worse

A retry loop needs four decisions:

  1. What to retry. Only transient errors, and only calls that are safe to repeat.
  2. How long to wait. Waits grow each time: exponential backoff.
  3. How to avoid a crowd. If 50 jobs all fail at 02:00 and all wait exactly 1, 2, 4 seconds, they all come back together. AWS's architecture blog calls these "clusters of calls". Adding jitter, a random wait, spreads them out. The "full jitter" version picks a wait anywhere between zero and the current backoff limit: sleep = random_between(0, min(cap, base * 2 ** attempt)).
  4. When to stop. A fixed number of attempts and a cap on each wait, so a job never hangs for hours.

If the server sends a Retry-After header, it is telling you how long to wait. Use it, but still cap it.

flowchart TD
  A["Call SAP"] --> B{"Answer?"}
  B -->|"2xx"| OK["Use the data"]
  B -->|"429, 503, timeout"| C{"Attempts left?"}
  C -->|"Yes"| W["Wait: Retry-After<br/>or backoff + jitter"]
  W --> A
  C -->|"No"| T["Give up: log ERROR,<br/>stop or skip"]
  B -->|"400, 401, 404"| P["No retry: log,<br/>stop or skip"]

Is it safe to repeat? RFC 9110, the HTTP standard, defines idempotent methods: repeating them has the same effect as sending them once. GET (reading) and the other safe methods are idempotent, and so are PUT and DELETE. POST, which SAP's OData APIs use to create entities such as sales orders, is not. The standard says a client should not automatically retry a non-idempotent request unless it knows the request is safe to repeat, or knows the first one didn't complete. For writes to SAP, check first (read the order back, or look for your own reference number) before you try again.

Retries also cost money and quota. Every retried model call may be billed, and a retry storm can use up a whole team's rate limit. Keep attempts low (three or four in total is typical) and log every retry, so you can see when they start to climb.

Logging: levels, loggers, handlers

Python's logging module is built into the language. Its HOWTO gives a simple rule for what goes where:

You want to... Use
Show normal output to the person running the script print()
Record what happened during normal operation log.info(), or log.debug() for detail
Warn about something unexpected that you handled log.warning()
Report an error the code can't handle here Raise an exception
Record an error you handled without raising log.error(), or log.exception() inside an except block to include the traceback

The five levels, from least to most serious, are DEBUG, INFO, WARNING, ERROR and CRITICAL. By default only WARNING and above are shown, which is why a first log.info() often seems to do nothing.

Three parts work together:

  • A logger is what your code calls. Create one per file with log = logging.getLogger(__name__), so each entry says which module wrote it.
  • A handler decides where entries go: the screen, a file, or a log service. Each handler can have its own level, so the screen shows INFO while the file keeps DEBUG.
  • A formatter decides what each line looks like. A filter can change or drop entries before they are written; you will use one to add a run ID and remove secrets.

Pass values as arguments, log.warning("order %s failed", number), rather than building the string yourself. The logging module only formats the message if the entry is actually written.

Structured logs and correlation IDs

A line like WARNING order 9000004 skipped is fine for a person. A log tool searching thousands of runs needs fields. Structured logging writes each entry as data, commonly one JSON object per line (JSON Lines):

{"time": "2026-10-03T06:29:56.379+00:00", "level": "WARNING", "run_id": "70265a47", "message": "order 9000004 skipped: HTTP 404: ...", "status": 404, "transaction_id": "SAMPLE-TXN-404-1", "order": "9000004"}

Now you can ask "show every 429 last week" or "everything from run 70265a47". The run ID (a correlation ID) ties all entries from one run together. In a bigger system you pass the same ID along to every service you call, so one ID finds the story everywhere.

What never goes in a log

OWASP's Logging Cheat Sheet lists data that should usually be removed, masked, hashed or encrypted rather than logged. For AI apps on SAP, the ones that matter most are:

  • Access tokens, API keys, passwords, encryption keys and connection strings. That includes Authorization and APIKey headers.
  • Sensitive personal data. Customer and employee names, emails, phone numbers, bank details. In AI apps this includes the prompt and the model's answer when they contain such data.
  • Session identifiers and payment card data.

Two practical rules follow. Never log whole requests, headers or prompts by default; log IDs, counts, status codes and timings instead. And add a safety net: a filter that masks known secrets and obvious patterns (emails, apikey=...) in every entry, including tracebacks. OWASP also warns about log injection: if user text containing line breaks reaches your log, it can forge fake entries. Replacing carriage returns and line feeds in messages prevents that.

Debugging: from symptom to cause

Debugging is a loop: reproduce, observe, explain, fix, check.

  1. Reproduce the problem with the smallest input you can, ideally sample data rather than a live system.
  2. Observe. Read the traceback if there is one. If the answer is just wrong, pause the program where the wrong value appears.
  3. Explain the cause in one sentence before changing anything ("the amount is text, so the comparison is alphabetical").
  4. Fix that cause, not the symptom.
  5. Check with the same input, and keep that input as a test.

VS Code's debugger, which comes with the Python extension you installed in Set up your computer, does step 2 well. You set a breakpoint by clicking in the margin left of a line number (or pressing F9). When the program reaches it, it pauses. The Run and Debug view then shows Variables (every value right now), Watch (expressions you choose), and Call Stack (how you got here). The Debug Console lets you type Python expressions and see the result in the paused program. The toolbar moves you on: Continue (F5), Step Over (F10, run this line), Step Into (F11, go inside the function on this line) and Step Out (finish this function).

Two more tools save time with loops over many sales orders. A conditional breakpoint pauses only when an expression is true, such as order["SalesOrder"] == "9000002". A logpoint prints a message with values, such as amount={order["TotalNetAmount"]}, without pausing and without editing your code.

Without VS Code, Python's built-in breakpoint() does the same job in the terminal: put it on a line, run the script, and you get a (Pdb) prompt where p expression prints a value, n runs the next line and c continues.

Build it yourself: a sturdy SAP reader and a debugging session

You will build two things:

  1. sturdy_calls.py reads blocked sales orders and their items. It sorts every failure into the three bins, retries what is worth retrying, skips one bad order without losing the others, and writes a safe JSON log. Its --sample mode acts out failures on purpose (a 503, a 429, a timeout, a missing order, a wrong key) so you can watch each behavior on your own computer.
  2. buggy_totals.py has two bugs on purpose. You will find one from its traceback and the other with the VS Code debugger.
flowchart LR
  S["SAP API<br/>(sample or sandbox)"] --> R["with_retries<br/>+ check"]
  R -->|"data"| P["sturdy_result.json"]
  R -->|"events"| L["logs/sturdy_calls.jsonl"]
  R -->|"screen"| T["Terminal"]
  P -. "--llm" .-> M["Claude summary"]

Before you start: complete Set up your computer for this course. It creates your orchestrate-course folder with its .venv, installs requests, anthropic and python-dotenv, stores SAP_API_KEY in .env, and installs VS Code with the Python extension. Calling your first SAP API explains the Sales Order API used here.

What you need

  • The course setup above. Nothing new to install.
  • For the sample runs and the debugging session: nothing else. Free, and they work offline.
  • For the optional sandbox run: your free SAP Business Accelerator Hub key in .env as SAP_API_KEY. Free.
  • For the optional summary: ANTHROPIC_API_KEY in .env. Each run makes one small request: a small per-request charge.
  • About 60 to 90 minutes.

Step 1: Open your course folder and turn on the environment

  1. Open VS Code, choose File > Open Folder and open your orchestrate-course folder.

  2. Open a terminal with Terminal > New Terminal.

  3. Turn on the virtual environment:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that the prompt starts with (.venv).

  5. Make sure the unit01 folder exists. Stay in the course folder: every command in this walkthrough runs from there.

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit01
    • macOS / Linux:

      mkdir -p unit01

Step 2: Keep logs out of Git

The script writes its log to a logs folder. Logs belong to one computer and one run; they don't belong in your repository, even when they are clean.

  1. Add a line to .gitignore:

    • Windows (PowerShell):

      Add-Content .gitignore "logs/"
    • macOS / Linux:

      echo "logs/" >> .gitignore
  2. Open .gitignore in VS Code and check it now ends with logs/, and still lists .env.

Step 3: Save the script

  1. In VS Code's file list, right-click unit01, choose New File and name it sturdy_calls.py.
  2. Paste the whole script below and save with Ctrl+S (Windows, Linux) or Cmd+S (macOS).
"""Errors, logging and debugging: read blocked sales orders without falling over.

How to run (from the orchestrate-course folder, with .venv turned on):
  python unit01/sturdy_calls.py --sample                     made-up data; everything works
  python unit01/sturdy_calls.py --sample --scenario flaky    503, 429 and a timeout, then success
  python unit01/sturdy_calls.py --sample --scenario partial  one order's items are missing (404)
  python unit01/sturdy_calls.py --sample --scenario badkey   wrong key: stops at once, no retries
  python unit01/sturdy_calls.py --sample --scenario down     service down: retries, then stops
  python unit01/sturdy_calls.py                              SAP's sandbox (needs SAP_API_KEY in .env)
  python unit01/sturdy_calls.py --sample --llm               also ask Claude for a summary (needs ANTHROPIC_API_KEY)

Options:
  --log-level DEBUG   show more on screen (the log file always keeps DEBUG)
  --max-attempts 4    how many times to try one call before giving up

Exit codes: 0 all good, 1 finished but some orders failed,
            2 stopped by a problem retrying can't fix, 3 stopped because SAP stayed unavailable.
"""
from __future__ import annotations

import argparse
import json
import logging
import os
import random
import re
import sys
import time
import uuid
from dataclasses import dataclass, field
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

SERVICE = "/sap/opu/odata/sap/API_SALES_ORDER_SRV"
SANDBOX = "https://sandbox.api.sap.com/s4hanacloud" + SERVICE
ORDER_FIELDS = ["SalesOrder", "SoldToParty", "TotalNetAmount", "TransactionCurrency",
                "DeliveryBlockReason", "HeaderBillingBlockReason"]
ITEM_FIELDS = ["SalesOrderItem", "Material", "RequestedQuantity", "NetAmount"]
TRANSIENT_STATUS = {429, 502, 503, 504}  # "try again later" answers
MAX_WAIT_S = 30.0                         # never sleep longer than this between tries

log = logging.getLogger("sturdy_calls")


# ------------------------------------------------------------------ our own exceptions

class SAPCallError(Exception):
    """A call to SAP failed. Carries the details support will ask for."""

    def __init__(self, message: str, *, status: int | None = None, sap_code: str | None = None,
                 transaction_id: str | None = None, path: str = ""):
        super().__init__(message)
        self.status = status
        self.sap_code = sap_code
        self.transaction_id = transaction_id
        self.path = path


class TransientError(SAPCallError):
    """Might work later: 429, 502, 503, 504, a timeout or a dropped connection."""

    def __init__(self, message: str, *, retry_after: float | None = None, **details):
        super().__init__(message, **details)
        self.retry_after = retry_after


class PermanentError(SAPCallError):
    """Will fail the same way every time: wrong key, wrong address, bad request."""


class SetupError(Exception):
    """Something on our side is wrong, such as a missing or refused key. Fix it, then rerun."""


@dataclass
class Reply:
    status: int
    text: str
    headers: dict = field(default_factory=dict)


# ------------------------------------------------------------------ reading SAP's answers

def parse_sap_error(text: str) -> tuple[str | None, str, str | None]:
    """Pull code, message and transaction ID out of an OData error body, if it is one."""
    try:
        err = json.loads(text)["error"]
        message = err.get("message", "")
        if isinstance(message, dict):  # OData V2 puts the text in message.value
            message = message.get("value", "")
        inner = err.get("innererror") or {}
        return err.get("code"), str(message), inner.get("transactionid")
    except (ValueError, KeyError, TypeError, AttributeError):
        return None, (text or "").strip()[:200] or "(empty body)", None


def parse_retry_after(value: str | None) -> float | None:
    """Retry-After is either a number of seconds or an HTTP date."""
    if not value:
        return None
    if value.strip().isdigit():
        return float(value.strip())
    try:
        when = parsedate_to_datetime(value)
        return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
    except (TypeError, ValueError):
        return None


def check(reply: Reply, path: str) -> Reply:
    """Return the reply if it succeeded; otherwise raise the right kind of error."""
    if 200 <= reply.status < 300:
        return reply
    code, message, txn = parse_sap_error(reply.text)
    details = {"status": reply.status, "sap_code": code, "transaction_id": txn, "path": path}
    if reply.status in TRANSIENT_STATUS:
        raise TransientError(f"HTTP {reply.status}: {message}",
                             retry_after=parse_retry_after(reply.headers.get("Retry-After")), **details)
    raise PermanentError(f"HTTP {reply.status}: {message}", **details)


# ------------------------------------------------------------------ retries

STATS = {"calls": 0, "retries": 0, "waited_s": 0.0}


def with_retries(call, what: str, *, max_attempts: int, rng: random.Random,
                 base: float = 0.5, cap: float = 8.0):
    """Run call(); on a TransientError wait and try again, up to max_attempts in total.
    Only wrap calls that are safe to repeat (GET). PermanentError is never retried."""
    for attempt in range(1, max_attempts + 1):
        STATS["calls"] += 1
        try:
            result = call()
        except TransientError as err:
            if attempt == max_attempts:
                log.error("%s: gave up after %d attempts (%s)", what, attempt, err,
                          extra={"status": err.status, "attempt": attempt})
                raise
            if err.retry_after is not None:
                wait = min(err.retry_after, MAX_WAIT_S)           # the server told us
            else:
                wait = rng.uniform(0, min(cap, base * 2 ** attempt))  # backoff with full jitter
            log.warning("%s: %s; retry %d of %d in %.1f s", what, err, attempt, max_attempts - 1,
                        wait, extra={"status": err.status, "attempt": attempt, "wait_s": round(wait, 2)})
            STATS["retries"] += 1
            STATS["waited_s"] += wait
            time.sleep(wait)
        else:
            if attempt > 1:
                log.info("%s: succeeded on attempt %d", what, attempt, extra={"attempt": attempt})
            return result


# ------------------------------------------------------------------ talking to SAP

class LiveSAP:
    """Real GET requests to SAP's sandbox, with timeouts."""

    def __init__(self, key: str):
        import requests  # imported here so --sample works without the library
        self.requests = requests
        self.session = requests.Session()
        self.session.headers.update({"APIKey": key, "Accept": "application/json"})

    def get(self, path: str, params: dict | None = None) -> Reply:
        exc = self.requests.exceptions
        try:
            resp = self.session.get(SANDBOX + path, params=params, timeout=(5, 30))
        except (exc.Timeout, exc.ConnectionError) as err:
            raise TransientError(f"could not reach SAP ({type(err).__name__})", path=path) from err
        except exc.RequestException as err:
            raise PermanentError(f"request could not be sent ({type(err).__name__})", path=path) from err
        return Reply(resp.status_code, resp.text, dict(resp.headers))


SAMPLE_ORDERS = [
    {"SalesOrder": "9000001", "SoldToParty": "CUST-A", "TotalNetAmount": "18250.00",
     "TransactionCurrency": "USD", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000002", "SoldToParty": "CUST-B", "TotalNetAmount": "9500.00",
     "TransactionCurrency": "USD", "DeliveryBlockReason": "", "HeaderBillingBlockReason": "02"},
    {"SalesOrder": "9000003", "SoldToParty": "CUST-C", "TotalNetAmount": "4200.00",
     "TransactionCurrency": "EUR", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000004", "SoldToParty": "CUST-A", "TotalNetAmount": "31000.00",
     "TransactionCurrency": "USD", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000005", "SoldToParty": "CUST-D", "TotalNetAmount": "760.00",
     "TransactionCurrency": "EUR", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
]
SAMPLE_ITEMS = {
    "9000001": [{"SalesOrderItem": "10", "Material": "PUMP-100", "RequestedQuantity": "10", "NetAmount": "15000.00"},
                {"SalesOrderItem": "20", "Material": "SEAL-7", "RequestedQuantity": "50", "NetAmount": "3250.00"}],
    "9000002": [{"SalesOrderItem": "10", "Material": "VALVE-3", "RequestedQuantity": "19", "NetAmount": "9500.00"}],
    "9000004": [{"SalesOrderItem": "10", "Material": "MOTOR-9", "RequestedQuantity": "4", "NetAmount": "31000.00"}],
}


def odata_error(status: int, message: str, n: int) -> str:
    """An error body shaped like an OData V2 error. Code and transaction ID are made up."""
    return json.dumps({"error": {"code": f"SAMPLE/{status}",
                                 "message": {"lang": "en", "value": message},
                                 "innererror": {"transactionid": f"SAMPLE-TXN-{status}-{n}"}}})


class SampleSAP:
    """Answers like the sandbox, from made-up data, and fails on purpose per scenario."""

    def __init__(self, scenario: str):
        self.scenario = scenario
        self.seen: dict[str, int] = {}

    def get(self, path: str, params: dict | None = None) -> Reply:
        n = self.seen[path] = self.seen.get(path, 0) + 1   # how often we've been asked this
        s = self.scenario
        if path == "/A_SalesOrder":
            if s == "badkey":
                return Reply(401, "Invalid ApiKey")
            if s == "down" or (s == "flaky" and n == 1):
                return Reply(503, odata_error(503, "Service temporarily unavailable", n))
            if s == "flaky" and n == 2:
                return Reply(429, "Too many requests", {"Retry-After": "1"})
            top = int((params or {}).get("$top", 50))
            return Reply(200, json.dumps({"d": {"results": SAMPLE_ORDERS[:top]}}))
        match = re.fullmatch(r"/A_SalesOrder\('(\d+)'\)/to_Item", path)
        if match:
            order = match.group(1)
            if s == "flaky" and order == "9000002" and n == 1:
                raise TransientError("could not reach SAP (ReadTimeout)", path=path)
            if s == "partial" and order == "9000004":
                return Reply(404, odata_error(404, f"Resource not found for order {order}", n))
            return Reply(200, json.dumps({"d": {"results": SAMPLE_ITEMS.get(order, [])}}))
        return Reply(404, odata_error(404, "Unknown path", n))


# ------------------------------------------------------------------ logging setup

EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
KEY_IN_TEXT = re.compile(r"(?i)\b(apikey|api_key|authorization|x-api-key)(\s*[:=]\s*)(\S+)")


class RunContext(logging.Filter):
    """Adds the run ID to every record and removes secrets and personal data from it."""

    def __init__(self, run_id: str, secrets: list[str]):
        super().__init__()
        self.run_id = run_id
        self.secrets = [s for s in secrets if s and len(s) >= 8]

    def mask(self, text: str) -> str:
        for secret in self.secrets:
            text = text.replace(secret, "***")
        text = KEY_IN_TEXT.sub(r"\1\2***", text)
        return EMAIL.sub("<email>", text)

    def filter(self, record: logging.LogRecord) -> bool:
        record.run_id = self.run_id
        message = self.mask(record.getMessage())
        record.msg, record.args = message.replace("\r", " ").replace("\n", " "), None  # one event, one line
        return True


class MaskedText(logging.Formatter):
    """Plain text for the screen, with secrets masked in tracebacks too."""

    def __init__(self, fmt: str, context: RunContext):
        super().__init__(fmt)
        self.context = context

    def formatException(self, ei) -> str:
        return self.context.mask(super().formatException(ei))


class JsonLines(logging.Formatter):
    """One JSON object per line: easy for people to grep and for log tools to search."""
    EXTRA = ("status", "sap_code", "transaction_id", "attempt", "wait_s", "order", "path", "request_id")

    def __init__(self, context: RunContext):
        super().__init__()
        self.context = context

    def format(self, record: logging.LogRecord) -> str:
        entry = {"time": datetime.fromtimestamp(record.created, timezone.utc).isoformat(timespec="milliseconds"),
                 "level": record.levelname, "logger": record.name,
                 "run_id": getattr(record, "run_id", None), "message": record.getMessage()}
        for name in self.EXTRA:
            if getattr(record, name, None) is not None:
                entry[name] = getattr(record, name)
        if record.exc_info:
            entry["exception"] = self.context.mask(self.formatException(record.exc_info))
        return json.dumps(entry, ensure_ascii=False)


def setup_logging(level: str, run_id: str, secrets: list[str], log_dir: str = "logs") -> str:
    os.makedirs(log_dir, exist_ok=True)
    path = os.path.join(log_dir, "sturdy_calls.jsonl")
    context = RunContext(run_id, secrets)

    screen = logging.StreamHandler()                   # writes to stderr
    screen.setLevel(level)
    screen.setFormatter(MaskedText("%(levelname)-7s %(message)s", context))
    screen.addFilter(context)

    file = logging.FileHandler(path, encoding="utf-8")  # appends; one line per event
    file.setLevel(logging.DEBUG)
    file.setFormatter(JsonLines(context))
    file.addFilter(context)

    # Libraries (requests, anthropic, ...) log too. Keep only their warnings; keep all of ours.
    logging.basicConfig(level=logging.WARNING, handlers=[screen, file], force=True)
    log.setLevel(logging.DEBUG)
    return path


# ------------------------------------------------------------------ the work

def is_blocked(order: dict) -> bool:
    return bool(order.get("DeliveryBlockReason") or order.get("HeaderBillingBlockReason"))


def read_orders(api, top: int, attempts: int, rng: random.Random) -> list[dict]:
    path = "/A_SalesOrder"
    params = {"$top": str(top), "$select": ",".join(ORDER_FIELDS), "$format": "json"}
    reply = with_retries(lambda: check(api.get(path, params), path), "read orders",
                         max_attempts=attempts, rng=rng)
    return json.loads(reply.text)["d"]["results"]


def read_items(api, order: str, attempts: int, rng: random.Random) -> list[dict]:
    path = f"/A_SalesOrder('{order}')/to_Item"
    params = {"$select": ",".join(ITEM_FIELDS), "$format": "json"}
    reply = with_retries(lambda: check(api.get(path, params), path), f"items of {order}",
                         max_attempts=attempts, rng=rng)
    return json.loads(reply.text)["d"]["results"]


def summarize_with_claude(result: dict) -> str | None:
    """Optional: ask a model for two sentences. Fails gracefully: the data file matters more."""
    import anthropic  # imported here so the rest runs without it

    client = anthropic.Anthropic(max_retries=3, timeout=60.0)  # the SDK retries 429 and 5xx itself
    facts = {"blocked": [{k: o[k] for k in ("sales_order", "customer", "net_amount", "currency")}
                         for o in result["orders"]],
             "could_not_read": [f["sales_order"] for f in result["failed"]]}
    try:
        message = client.messages.create(
            model=os.environ.get("LLM_MODEL", "claude-opus-5-5"),
            max_tokens=200,
            messages=[{"role": "user", "content":
                       "In two sentences for a credit manager, summarize these blocked sales orders. "
                       "Use only these facts.\n" + json.dumps(facts)}])
    except anthropic.AuthenticationError as err:  # a wrong key never fixes itself: fail loudly
        raise SetupError("Claude refused the key: check ANTHROPIC_API_KEY in .env") from err
    except (anthropic.RateLimitError, anthropic.APIConnectionError, anthropic.InternalServerError) as err:
        log.warning("summary skipped after the SDK's retries: %s", type(err).__name__)
        return None                            # the order data is still saved
    log.info("summary received", extra={"request_id": message._request_id})
    return "".join(block.text for block in message.content if block.type == "text")


def main() -> int:
    parser = argparse.ArgumentParser(description="Read blocked sales orders, with retries and logs.")
    parser.add_argument("--sample", action="store_true", help="made-up data; no key or internet")
    parser.add_argument("--scenario", default="ok", choices=["ok", "flaky", "partial", "badkey", "down"],
                        help="which failure the sample data acts out")
    parser.add_argument("--llm", action="store_true", help="also ask Claude for a summary")
    parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
    parser.add_argument("--max-attempts", type=int, default=4, help="tries per call, including the first")
    parser.add_argument("--top", type=int, default=50, help="read at most this many orders")
    args = parser.parse_args()

    try:
        from dotenv import load_dotenv
        load_dotenv()
    except ImportError:
        if not args.sample:
            print("python-dotenv is missing: run  pip install -r requirements.txt", file=sys.stderr)
            return 2

    # Check the settings first: a missing key should stop the run before any work is done.
    if args.llm and not os.environ.get("ANTHROPIC_API_KEY"):
        print("--llm needs ANTHROPIC_API_KEY in .env (or leave out --llm)", file=sys.stderr)
        return 2
    sap_key = os.environ.get("SAP_API_KEY", "")
    run_id = uuid.uuid4().hex[:8]
    log_path = setup_logging(args.log_level, run_id, [sap_key, os.environ.get("ANTHROPIC_API_KEY", "")])
    rng = random.Random(7) if args.sample else random.Random()  # fixed seed: same waits every sample run

    if args.sample:
        api, source = SampleSAP(args.scenario), f"sample ({args.scenario})"
    else:
        if not sap_key:
            log.error("SAP_API_KEY is not set. Add it to .env, or run with --sample")
            return 2
        api, source = LiveSAP(sap_key), "sandbox"
    print(f"Run {run_id}: reading from {source}. Log file: {log_path}")
    log.info("run started: source=%s", source)
    log.debug("settings: max_attempts=%d top=%d", args.max_attempts, args.top)

    # 1. Read the orders. Without them there is nothing to do, so failures here stop the run.
    try:
        orders = read_orders(api, args.top, args.max_attempts, rng)
    except PermanentError as err:
        log.error("stopped: %s", err, extra={"status": err.status, "sap_code": err.sap_code,
                                             "transaction_id": err.transaction_id})
        if err.status in (401, 403):
            print("The key was refused. Check SAP_API_KEY in .env (Show API Key on api.sap.com).")
        return 2
    except TransientError:
        print("SAP stayed unavailable. Nothing was written; run again later.")
        return 3

    blocked = [o for o in orders if is_blocked(o)]
    print(f"Read {len(orders)} orders; {len(blocked)} have a delivery or billing block")
    if not blocked:
        blocked = orders[:2]
        print("No blocked orders in this data, so reading items for the first two anyway")

    # 2. Read items per order. One bad order must not sink the others: record it and go on.
    result = {"run_id": run_id, "source": source, "orders": [], "failed": []}
    for order in blocked:
        number = order["SalesOrder"]
        try:
            items = read_items(api, number, args.max_attempts, rng)
        except SAPCallError as err:
            log.warning("order %s skipped: %s", number, err,
                        extra={"order": number, "status": err.status, "sap_code": err.sap_code,
                               "transaction_id": err.transaction_id})
            result["failed"].append({"sales_order": number, "status": err.status,
                                     "reason": str(err), "transaction_id": err.transaction_id})
            continue
        log.debug("order %s: %d item(s)", number, len(items), extra={"order": number})
        result["orders"].append({"sales_order": number, "customer": order["SoldToParty"],
                                 "net_amount": order["TotalNetAmount"],
                                 "currency": order["TransactionCurrency"], "items": len(items)})

    with open("sturdy_result.json", "w", encoding="utf-8") as out:
        json.dump(result, out, indent=2)
    print(f"Saved {len(result['orders'])} orders, {len(result['failed'])} failed, to sturdy_result.json")
    print(f"Calls: {STATS['calls']}, retries: {STATS['retries']}, waited {STATS['waited_s']:.1f} s")

    if args.llm:
        try:
            summary = summarize_with_claude(result)
        except SetupError as err:
            log.error("stopped: %s", err)
            return 2
        print("Summary:", summary or "(skipped; see the log)")

    log.info("run finished: %d saved, %d failed", len(result["orders"]), len(result["failed"]))
    return 1 if result["failed"] else 0


if __name__ == "__main__":
    sys.exit(main())

Step 4: Run it when everything works

  1. Run:

    python unit01/sturdy_calls.py --sample

    On macOS or Linux, use python3 if python isn't found.

  2. You should see this. Your run ID, the eight characters after Run, will be different every time:

Run 24073cd4: reading from sample (ok). Log file: logs/sturdy_calls.jsonl
INFO    run started: source=sample (ok)
Read 5 orders; 3 have a delivery or billing block
Saved 3 orders, 0 failed, to sturdy_result.json
Calls: 4, retries: 0, waited 0.0 s
INFO    run finished: 3 saved, 0 failed
  1. Notice the two kinds of lines. Lines starting with INFO come from the logger; the others come from print(). Results are for the person at the terminal; log lines are the record of what happened.
  2. Four calls: one for the order list, then one for the items of each of the three blocked orders.

Step 5: Watch transient failures being retried

  1. Run the flaky scenario. The sample service answers the first order call with 503, the second with 429 and Retry-After: 1, and lets one item call time out once:

    python unit01/sturdy_calls.py --sample --scenario flaky
  2. You should see:

Run 55548a7a: reading from sample (flaky). Log file: logs/sturdy_calls.jsonl
INFO    run started: source=sample (flaky)
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 1 of 3 in 0.3 s
WARNING read orders: HTTP 429: Too many requests; retry 2 of 3 in 1.0 s
INFO    read orders: succeeded on attempt 3
Read 5 orders; 3 have a delivery or billing block
WARNING items of 9000002: could not reach SAP (ReadTimeout); retry 1 of 3 in 0.2 s
INFO    items of 9000002: succeeded on attempt 2
Saved 3 orders, 0 failed, to sturdy_result.json
Calls: 7, retries: 3, waited 1.5 s
INFO    run finished: 3 saved, 0 failed
  1. Read the waits. After the 503 the script picked a random wait with jitter (0.3 s). After the 429 it waited exactly 1.0 s, because the answer's Retry-After header said so. In sample mode the random numbers are seeded, so you get the same waits every time; against a real service they differ on every run.
  2. The end result is the same as Step 4. Three transient failures cost three extra calls and 1.5 seconds, and nobody had to do anything.

Step 6: Watch one bad order being skipped, not the whole run

  1. Run the partial scenario. The items of order 9000004 come back as 404:

    python unit01/sturdy_calls.py --sample --scenario partial
  2. You should see:

Run 70265a47: reading from sample (partial). Log file: logs/sturdy_calls.jsonl
INFO    run started: source=sample (partial)
Read 5 orders; 3 have a delivery or billing block
WARNING order 9000004 skipped: HTTP 404: Resource not found for order 9000004
Saved 2 orders, 1 failed, to sturdy_result.json
Calls: 4, retries: 0, waited 0.0 s
INFO    run finished: 2 saved, 1 failed
  1. No retries: a 404 is permanent, so trying again would only fail again. The order was recorded and the other two were saved.

  2. Look at the exit code, the number a scheduler would see:

    • Windows (PowerShell):

      $LASTEXITCODE
    • macOS / Linux:

      echo $?

    You should see 1: "finished, but some orders failed". A scheduler or pipeline can alert on that, even though the run didn't crash.

  3. Open sturdy_result.json in VS Code. Under "failed" you see order 9000004 with its status, reason and "transaction_id": "SAMPLE-TXN-404-1". In a real system that ID is what an SAP administrator would look up.

Step 7: Watch permanent failures stop the run

  1. Run the badkey scenario:

    python unit01/sturdy_calls.py --sample --scenario badkey
  2. You should see:

Run 9c76bf35: reading from sample (badkey). Log file: logs/sturdy_calls.jsonl
INFO    run started: source=sample (badkey)
ERROR   stopped: HTTP 401: Invalid ApiKey
The key was refused. Check SAP_API_KEY in .env (Show API Key on api.sap.com).

No retries, a clear instruction, and exit code 2.

  1. Run the down scenario. The service answers 503 every time:

    python unit01/sturdy_calls.py --sample --scenario down
  2. You should see three retries with growing random waits, then a clear stop with exit code 3. It takes about three seconds:

Run 1e4c9f32: reading from sample (down). Log file: logs/sturdy_calls.jsonl
INFO    run started: source=sample (down)
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 1 of 3 in 0.3 s
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 2 of 3 in 0.3 s
WARNING read orders: HTTP 503: Service temporarily unavailable; retry 3 of 3 in 2.6 s
ERROR   read orders: gave up after 4 attempts (HTTP 503: Service temporarily unavailable)
SAP stayed unavailable. Nothing was written; run again later.
  1. Notice that the second wait (0.3 s) is shorter than you might expect. That is full jitter: each wait is random between zero and a limit that doubles (1, 2, 4 seconds), so waits spread out instead of lining up.
  2. Try --max-attempts 2 with the same scenario. The script gives up after one retry.

Step 8: Read the log file

  1. In VS Code's file list, open logs/sturdy_calls.jsonl. Every run so far has added lines. Each line is one JSON object with time, level, run_id and message, plus fields such as status, attempt, wait_s, order and transaction_id where they apply.

  2. The file also has DEBUG lines the screen didn't show, such as settings: max_attempts=4 top=50. The screen handler shows INFO and above; the file handler keeps everything.

  3. Find every warning from the command line:

    • Windows (PowerShell):

      Select-String -Path logs\sturdy_calls.jsonl -Pattern '"WARNING"'
    • macOS / Linux:

      grep '"WARNING"' logs/sturdy_calls.jsonl
  4. To see DEBUG lines on screen as well, run:

    python unit01/sturdy_calls.py --sample --scenario partial --log-level DEBUG

    You now also see DEBUG order 9000001: 2 item(s) and similar lines.

Step 9 (optional): Run it against SAP's sandbox

  1. Check that .env has your SAP_API_KEY. If not, follow the key steps in the setup topic.

  2. Run without --sample:

    python unit01/sturdy_calls.py
  3. You should see reading from sandbox and the same kind of summary with SAP's demo data. Valid differences:

    • The number of orders is up to 50, the default --top.
    • 0 have a delivery or billing block is a valid answer about demo data. The script then reads items for the first two orders so you still see item calls.
    • If your network blocks the sandbox, you see could not reach SAP (ConnectionError) or (ProxyError) with retries, then exit code 3. That is the script doing its job: the failure is transient from its point of view.

Step 10 (optional): Add a model summary, and see a graceful skip

  1. If you haven't yet, add ANTHROPIC_API_KEY="..." to .env, as in the setup topic.

  2. Run:

    python unit01/sturdy_calls.py --sample --llm
  3. After the usual lines you see INFO summary received and a line starting with Summary:. The log file entry for summary received carries a request_id; that is what a provider's support team asks for.

  4. Look at summarize_with_claude in the script. The Anthropic library is created with max_retries=3 and timeout=60.0, so the library itself retries rate limits and server errors. The function only decides what happens after that: a refused key stops the run (exit code 2); a rate limit or outage that outlasts the retries skips the summary but keeps the saved data.

  5. Without a key, --llm stops before any work with --llm needs ANTHROPIC_API_KEY in .env (or leave out --llm). Checking settings first means you never wait for a whole run only to fail at the end.

Step 11: Find a bug from its traceback

Now the debugging half. buggy_totals.py should list the blocked orders above 10,000 that need a credit manager's review. The right answer is ['9000001', '9000004'].

  1. Create unit01/buggy_totals.py and paste this. Don't fix anything yet:
"""Practice file for debugging: it has two bugs on purpose.

Run (from the orchestrate-course folder):  python unit01/buggy_totals.py
The right answer is: Blocked orders above 10,000: ['9000001', '9000004']
"""

ORDERS = [  # shaped like A_SalesOrder from SAP's Sales Order API; amounts arrive as text
    {"SalesOrder": "9000001", "TotalNetAmount": "18250.00", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000002", "TotalNetAmount": "9500.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": "02"},
    {"SalesOrder": "9000003", "TotalNetAmount": "4200.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000004", "TotalNetAmount": "31000.00", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
    {"SalesOrder": "9000005", "TotalNetAmount": "760.00", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
]
LIMIT = "10000.00"  # blocked orders above this value go to a credit manager


def is_blocked(order):
    return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])


def needs_review(order):
    return order["TotalNetAmount"] > LIMIT


def main():
    review = []
    for order in ORDERS:
        if is_blocked(order) and needs_review(order):
            review.append(order["SalesOrder"])
    print("Blocked orders above 10,000:", review)


if __name__ == "__main__":
    main()
  1. Run it:

    python unit01/buggy_totals.py
  2. You should see a traceback ending in:

  File ".../unit01/buggy_totals.py", line 18, in is_blocked
    return bool(order["DeliveryBlock"] or order["HeaderBillingBlockReason"])
                ~~~~~^^^^^^^^^^^^^^^^^
KeyError: 'DeliveryBlock'
  1. Read it bottom up: a KeyError for 'DeliveryBlock' on line 18, in is_blocked. The orders have DeliveryBlockReason, the field name in SAP's Sales Order API.
  2. On line 18, change order["DeliveryBlock"] to order["DeliveryBlockReason"]. Save and run again.
  3. You should now see:
Blocked orders above 10,000: ['9000001', '9000002', '9000004']

No crash, but the answer is wrong: 9000002 is worth 9,500. This is bin 3 without a traceback, the hardest kind of bug.

Step 12: Find the second bug with the debugger

  1. In buggy_totals.py, click in the margin just left of the line number of return order["TotalNetAmount"] > LIMIT (line 22). A red dot appears: that is a breakpoint.

  2. Open the Run and Debug view: the icon with a play button and a bug on the left, or Ctrl+Shift+D (Cmd+Shift+D on macOS).

  3. Click Run and Debug. If VS Code asks what to debug, choose Python Debugger, then Python File. You can also press F5 with the file open.

  4. The program stops at the red dot, and the line is highlighted. In the Variables section, expand order. It is 9000001, and TotalNetAmount shows '18250.00', in quotes.

  5. Click the Debug Console tab at the bottom and type each line, pressing Enter after each:

    type(order["TotalNetAmount"])
    order["TotalNetAmount"] > LIMIT
    "9500.00" > "10000.00"

    You should see <class 'str'>, True, and True. The amount is text, and text is compared character by character: "9" comes after "1", so "9500.00" counts as larger than "10000.00".

  6. Press F5 (Continue) twice. The second stop is order 9000002; look at its amount in Variables. Then press F5 until the program ends, or click the red square (Stop).

  7. Fix the cause. At the top of the file, under the docstring, add:

    from decimal import Decimal

    Then change the limit and the comparison so both are numbers:

    LIMIT = Decimal("10000.00")  # blocked orders above this value go to a credit manager
    def needs_review(order):
        return Decimal(order["TotalNetAmount"]) > LIMIT
  8. Click the red dot to remove the breakpoint, save, and run:

    python unit01/buggy_totals.py
  9. You should see the right answer:

Blocked orders above 10,000: ['9000001', '9000004']

Decimal is the right type for money: it keeps exact decimal values, as Calling your first SAP API explained for SAP amounts.

Step 13: Save your work in Git

  1. Check what Git sees. The logs folder must not appear:

    git status
  2. Commit the two scripts:

    git add .gitignore unit01/sturdy_calls.py unit01/buggy_totals.py
    git commit -m "Unit 1: retries, structured logs and a debugging session"

What each part of the script does

Part of the script What it does
SAPCallError, TransientError, PermanentError, SetupError Our own exception types. The rest of the code reacts to their meaning, not to library details
Reply A plain answer: status, text and headers, the same for sample and live
parse_sap_error Reads code, message.value and innererror.transactionid from an OData error body; falls back to raw text
parse_retry_after Reads Retry-After as seconds or as an HTTP date
check Sorts an answer into success, transient (429, 502, 503, 504) or permanent (everything else)
with_retries Calls, and on a transient error waits (Retry-After, or backoff with full jitter) and tries again, up to --max-attempts; counts calls and waits
LiveSAP Real GET requests with a timeout of 5 s to connect and 30 s to read; turns network errors into TransientError with raise ... from
SampleSAP, odata_error Made-up orders and items, and scripted failures for each --scenario
RunContext A logging filter: adds the run ID, masks keys, apikey=... and emails, and puts every message on one line
MaskedText The screen formatter: plain text, with secrets masked in tracebacks too
JsonLines The file formatter: one JSON object per entry, with extra fields and a masked traceback
setup_logging Screen at your chosen level, file at DEBUG; libraries only show warnings
read_orders, read_items The two GET calls, each wrapped in with_retries
summarize_with_claude Optional model call with the library's own retries; a refused key stops, an outage skips the summary
main Checks settings first, stops loudly if orders can't be read, skips single failed orders, saves the result, returns an exit code

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed or the terminal can't find it macOS/Linux: use python3. Windows: repeat the Python step in the setup topic and open a new terminal
can't open file '...unit01/sturdy_calls.py' The terminal isn't in the course folder, or the file has another name Run cd to your orchestrate-course folder, and check the file name in VS Code
ModuleNotFoundError: No module named 'requests' (or anthropic, dotenv) The library isn't installed in the active environment Turn on .venv (Step 1), then pip install -r requirements.txt. --sample runs need neither requests nor anthropic
TypeError: unsupported operand type(s) for | Your Python is older than 3.10 Install a current Python, as in the setup topic
SAP_API_KEY is not set The script can't find your key Check .env is in the course folder with SAP_API_KEY="...", or use --sample
stopped: HTTP 401 without --sample The sandbox refused your key Copy it again with Show API Key on api.sap.com and update .env
could not reach SAP (ConnectionError) or (ProxyError) Your network or a company proxy blocks sandbox.api.sap.com Try another network, or ask IT to allow the site. --sample still works
HTTP 429 warnings on every run You are calling too often Wait a few minutes; lower --top while experimenting
--llm needs ANTHROPIC_API_KEY No model key in .env Add it, or leave out --llm
stopped: Claude refused the key The model key is wrong or revoked Create a new key in the Claude Console and update .env
PermissionError on logs/sturdy_calls.jsonl The log file is locked, often by another program on Windows Close programs that have the file open, or delete the logs folder
The debugger doesn't stop at the red dot The line never ran, or you started without debugging Use Run and Debug or F5, not the plain Run button; check the breakpoint is on a line that runs
The debugger asks for a configuration every time No launch.json yet Choose Python Debugger, then Python File; or use create a launch.json file in the Run and Debug view

Why the script is built this way

  • Bins, not cases. check decides transient or permanent in one place. Every other function just reacts to the exception type.
  • Retries only around reads. Every call wrapped in with_retries is a GET. A write would need its own check-before-retry logic.
  • Two outputs, two audiences. print() for the person, the logger for the record. The JSON file is for tools, the screen for people.
  • Exit codes mean something. 0, 1, 2 and 3 let a scheduler tell "all good", "partly done", "fix the setup" and "try later" apart without reading text.
  • Sample and live share code. Only the class that sends requests changes, so the failure handling you watched in sample mode is the same code that runs against SAP.

The SAP way

In the SAP system: Gateway error logs

When your OData call to an S/4HANA or other ABAP-based system fails, SAP keeps its own record. SAP Learning's Gateway course describes:

  • SAP Gateway error log (/IWFND/ERROR_LOG): errors found while the OData request is analyzed, for example a request that doesn't fit the service's data model.
  • SAP Gateway backend error log (/IWBEP/ERROR_LOG): errors while the service implementation processes a valid request, such as an invalid key value or a coding error.
  • Error details with an overview by error ID and time, the error context, the call stack, the full request and response, and a Replay button.
  • SAP Gateway Client (/IWFND/GW_CLIENT) to send test requests and jump to the error log.

Two consequences for your code. Log the transaction ID and timestamp from SAP's error body, because they are how an administrator finds the matching entry. And plan who can look: these logs live in the customer's system, so agree with the SAP basis team early how errors get analyzed during the project. The course source covers SAP Gateway in ABAP-based systems; for SAP S/4HANA Cloud Public Edition, ask SAP or your partner which monitoring tools apply.

On SAP BTP: application logs and SAP Cloud Logging

An app deployed to the SAP BTP Cloud Foundry runtime (Unit 6, Deploying AI apps on SAP BTP) writes its logs to standard output and standard error, the screen of a program with no screen. The platform collects them and cf logs <app> --recent shows the latest lines. That is why sturdy_calls.py logs to the terminal and why one-JSON-object-per-line matters: the platform passes lines through as they are.

For storage, search and dashboards, SAP Learning describes SAP Cloud Logging: it ingests, stores and analyzes application logs from custom apps on SAP BTP, accepts data over the OpenTelemetry standard, comes with prebuilt dashboards, and links from SAP Cloud ALM, so an operator can go from a central alert to the detailed log lines. SAP also publishes an open-source Python library, sap_cf_logging, that writes JSON logs with a correlation ID for Cloud Foundry apps; check its release activity before you depend on it.

Model calls through SAP

When your model calls go through SAP's generative AI hub instead of directly to a provider, the orchestration service can run a list of configurations as fallback: SAP's SDK documentation says it moves to the next one when the model is unavailable, on a timeout, on too many requests or on internal service errors, and not on content filter hits or invalid input. That is the same transient-versus-permanent split as this topic, done by the platform. SAP Generative AI Hub and the orchestration service shows how to set it up, and how to read the failed attempts it reports.

Build vs. SAP

Need Build it in your code Use SAP's tools Recommendation
Decide what to retry and what to skip Yes: only your code knows the business rule No Always yours
Retry model calls on 429 and 5xx Possible Provider SDK retries; orchestration fallback Use the SDK or orchestration; add your own limits around them
Find why an OData call failed inside SAP No: you only see the answer Gateway error logs, with call stack and replay Log SAP's transaction ID, then use SAP's logs
Store and search logs for a BTP app A log file works on a laptop only SAP Cloud Logging, with SAP Cloud ALM SAP Cloud Logging, or the customer's existing log platform
Remove secrets and personal data from logs Yes: a filter like RunContext Libraries may redact some fields Do it in your code; never rely on the tool downstream
Debug a wrong result VS Code debugger on sample data SAP Gateway Client to replay a request Both: reproduce in SAP, then debug your code locally

Production concerns

  • Secrets and personal data. Mask at the source, as RunContext does, and also check what libraries log at DEBUG level; keep third-party loggers at WARNING in production. Never log prompts or model answers by default when they can contain customer or employee data. If you must keep them for evaluation, store them separately, with access control and a retention period agreed with the customer's data protection officer.
  • Who can read logs. Logs are data. Decide with the customer who may read the logs in each place they end up (your log platform, SAP Cloud Logging, SAP's own error logs) before go-live, not after the first incident.
  • Writes to SAP. Never put a create or change call inside a blind retry loop. Before retrying, read back to check whether the first call worked, or send your own unique reference with each create and look for it. For AI agents that propose SAP changes, a person approves first (Unit 11 covers agent permissions).
  • Retry budgets and cost. Cap attempts per call and total wait per run. Every retried model call may be billed and counts toward rate limits. Log retries as WARNING and track how many you see per day; a rising count is an early outage signal.
  • Timeouts everywhere. The requests library waits without limit unless you pass a timeout, and Anthropic's library defaults to 10 minutes. Set timeouts that match the job: an interactive app needs seconds, a nightly batch can afford more.
  • Alerting. A log nobody reads is a diary, not monitoring. Alert on exit codes other than 0, on ERROR entries, and on a share of failed records above a threshold. Unit 10 builds this into full observability.
  • Correlation across systems. Carry one ID from the user's request through your service, the model call and the SAP call. Log SAP's transaction ID and the model provider's request ID next to it.
  • Clean core. Don't change SAP code to add logging or error handling for your AI app. Use SAP's standard logs and APIs, and do your handling in the side-by-side app.

Pitfalls

  • except Exception: pass. The single most expensive line in enterprise Python. It turns every bin-3 bug into a silent wrong answer.
  • No timeout. A call that hangs forever holds a batch job and its schedule slot. Always pass timeout=.
  • Retrying permanent errors. Retrying a 401 delays the message someone needs, and repeated refused logins can trigger lockouts.
  • Retrying without jitter. Many jobs retrying on the same schedule hit the recovering service together.
  • Logging the whole request "for now". Debug code that logs headers or prompts ends up in production. Log IDs and counts instead.
  • Treating text amounts as numbers. SAP's OData V2 APIs send amounts as text. Comparing or adding them as text gives wrong answers without errors, as Step 11 showed. Convert with Decimal once, at the edge.
  • logging.basicConfig doing nothing. It has no effect if logging was already configured, for example by a library. Pass force=True, as the script does, or configure logging once at the start of main.
  • Debugging live data first. Reproduce on sample data. Pausing a program in the middle of a live SAP session can leave half-done work and hold locks.

Exercise: add a run summary and prove the logs are clean

The Unit 10 topics on observability use run summaries like this one to build dashboards and alerts. You will add one, then write a small check that no key ever reaches the log file.

  1. Open unit01/sturdy_calls.py. In main, find the line with open("sturdy_result.json", "w", encoding="utf-8") as out:.

  2. Just above it, add these lines, indented like the code around them:

    by_status = {}
    for failure in result["failed"]:
        key = str(failure["status"])
        by_status[key] = by_status.get(key, 0) + 1
    result["run_summary"] = {"calls": STATS["calls"], "retries": STATS["retries"],
                             "waited_s": round(STATS["waited_s"], 2), "failed_by_status": by_status}
  3. Save, then run:

    python unit01/sturdy_calls.py --sample --scenario partial
  4. Open sturdy_result.json. At the end you should see:

    "run_summary": {
      "calls": 4,
      "retries": 0,
      "waited_s": 0.0,
      "failed_by_status": {
        "404": 1
      }
    }
  5. Run --scenario flaky and check that retries is 3 and failed_by_status is {}.

  6. Create unit01/check_logs.py with this code. It reads every value in .env whose name ends in _API_KEY and searches the log file for it:

    """Fail if any API key from .env appears in the log file."""
    import sys
    from pathlib import Path
    
    from dotenv import dotenv_values
    
    keys = [v for k, v in dotenv_values(".env").items() if k.endswith("_API_KEY") and v]
    log_file = Path("logs/sturdy_calls.jsonl")
    text = log_file.read_text(encoding="utf-8") if log_file.exists() else ""
    leaks = sum(text.count(key) for key in keys)
    print(f"Checked {len(keys)} key(s) against {log_file}: {leaks} leak(s)")
    sys.exit(1 if leaks else 0)
  7. Run it:

    python unit01/check_logs.py

    You should see Checked 2 key(s) against logs/sturdy_calls.jsonl: 0 leak(s) (the number of keys depends on your .env).

  8. Commit both files with a message such as Unit 1: run summary and log check.

Done when: sturdy_result.json from a partial run ends with a run_summary showing "failed_by_status": {"404": 1}, and check_logs.py prints 0 leak(s) and exits with code 0.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In sturdy_calls.py, why is a 404 for one order's items not retried?

    Answer: B. check sorts 404 into PermanentError, and with_retries only retries TransientError. The main loop then records the order and moves on, which is the graceful part.
  2. 2The order list answers 429 with Retry-After: 1. How long does with_retries wait before the next try?

    Answer: C. When Retry-After is present, the script uses it instead of its own jittered backoff, but never waits longer than MAX_WAIT_S. Step 5 shows exactly this: retry 2 of 3 in 1.0 s.
  3. 3Why does the retry loop use full jitter instead of waiting exactly 1, 2 and 4 seconds?

    Answer: A. Plain backoff keeps failed clients in step, so they hit the recovering service together. A random wait between zero and the limit spreads them out. The fixed seed in sample mode is only there so the walkthrough output is repeatable.
  4. 4You extend the script to create sales orders with POST. A create call times out. What should the code do?

    Answer: D. POST is not idempotent. The first call may have created the order even though the answer never arrived, so a blind retry risks a duplicate. Read back, or search for your own reference, before retrying.
  5. 5What is the job of the RunContext filter?

    Answer: B. The filter runs on every record before a handler writes it. It attaches run_id and cleans the message; both formatters use the same masking for tracebacks. Removing line breaks from messages also prevents forged log lines.
  6. 6buggy_totals.py says order 9000002, worth 9,500, is above 10,000. What does the debugger show as the cause?

    Answer: C. In the Debug Console, type(order["TotalNetAmount"]) shows str, and "9500.00" > "10000.00" is True because "9" sorts after "1". Converting both sides to Decimal fixes the cause.
  7. 7A batch job fails at 3 a.m. with an HTTP 500 from an S/4HANA OData service. What in your log helps the SAP administrator most?

    Answer: D. SAP Gateway's error logs list errors by ID and time, with the call stack and the full request and response. Your log should carry the matching ID, and never the key.
  8. 8The Anthropic client in the script is created with max_retries=3. What does that mean for summarize_with_claude?

    Answer: C. Anthropic's library retries connection errors, 429 and 5xx answers itself. The function only maps the outcome to the business rule: a refused key stops the run; an outage that outlasts the retries skips the summary but keeps the data.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in