Orchestrate

Building an AI API

Turn a model call into a dependable web service with a clear contract, keys, limits, a fallback and tests, and see how SAP BTP does the same.

Updated Oct 2, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An AI API is a model call wrapped in a service that other systems can call. A sales app, an SAP Fiori screen or a workflow sends a small, well-defined request. The service checks it, asks the model, checks the answer, and sends back a small, well-defined response.

The model is the least stable part of that chain. It can be slow, down, expensive or simply wrong. A good AI API hides that from its callers. It promises a fixed contract (what you send, what you get back, what the errors mean), and it keeps that promise even when the model misbehaves.

In practice that means seven things around the model call: who may call it (keys), what they may send (input checks), what the model may return (a schema), how long it may take (timeouts), what happens when it fails (a fallback), how much it may cost (rate limits and a cache), and how you trace a single call later (a request ID). This topic builds all seven.

Why it matters to the business

Most AI pilots start as one person's script. The value appears when many systems and users can rely on it, and that is where pilots stall. An API is the step that turns "it worked in the demo" into "the order desk uses it every day".

Take the course's running example: blocked sales orders in order-to-cash. A sales rep opens an order that is on credit hold and wants to know why, in plain words, and who should act. Without an API, every app that wants this explanation builds its own prompt, its own model connection and its own error handling. With one, they all call a single service:

  • One place to govern. The prompt, the model choice and the safety rules live in one service. Changing the model is a configuration change, not a change in every app.
  • Predictable cost. Rate limits and a cache cap how often the paid model is called. OWASP lists unlimited resource use as a top API risk, and recommends spending limits or billing alerts for paid services.
  • Graceful failure. When the model is down, the rep still gets a short, rule-based answer instead of an error screen. The business process doesn't stop because a model provider had a bad hour.
  • Evidence. Every answer carries a request ID, the model name and the prompt version. When someone asks "why did the system say that?", you can find out.

The risk runs the other way too. An AI API with no key, no limits and no checks is an open, paid door into your systems.

How SAP does it

As of October 2026, SAP's guidance for generative AI applications on SAP Business Technology Platform (BTP) centers on a few building blocks. SAP's Architecture Center describes them in its "AI golden path":

  • SAP Cloud Application Programming Model (CAP) as the backend framework, running on Cloud Foundry or Kyma in BTP.
  • SAP Cloud SDK for AI, available for Java, JavaScript and Python, to call models through the generative AI hub in SAP AI Core.
  • The orchestration service for templating, masking, filtering and model fallback around each call (covered in SAP Generative AI Hub and the orchestration service).
  • The prompt registry to manage prompts instead of hard-coding them.
  • SAP Cloud Identity Services for sign-in, and destinations for secure connections to SAP back ends.

Around the API, API Management in SAP Integration Suite can publish it to other teams. SAP's learning material lists what it adds: API key and OAuth checks, rate limiting and spike arrest, caching, usage analytics and a developer portal where other developers find and subscribe to APIs.

So SAP gives you the parts. Deciding the contract, the limits and the fallback is still your team's job.

A decision guide: what your AI API must promise

Use this table when you review an AI API design, your own or a partner's. Each row is one promise to the callers.

Promise What it means in plain words What goes wrong without it
A fixed contract Callers know exactly what to send and what comes back Every model change breaks the apps that call it
Only known callers Each caller has its own key or sign-in Anyone who finds the address can run up the bill
Checked input Bad or oversized requests are refused before the model sees them Costs and errors from requests that never made sense
Checked output The model's answer must fit a schema before anyone sees it The model "releases" an order in a sentence nobody checked
A time limit Each model call stops after a set number of seconds One slow call blocks the sales rep's screen
A fallback A safe, simpler answer when the model fails The order desk sees error pages during an outage
Limits and reuse A cap per caller, and a cache for repeated questions A loop in one app spends the month's budget in an hour
Traceability A request ID, model name and prompt version on every answer Nobody can explain a wrong answer after the fact

Questions to ask

  • What exactly does the API accept and return? Can you show me the contract, and who must approve a change to it?
  • What happens when the model is slow or down? What does the user see?
  • How is the model's answer checked before it reaches a user or another system?
  • Who can call the API, and how do we revoke one caller without affecting the others?
  • What are the rate limits and the monthly spending cap? Who gets the alert?
  • Does the API log customer data? If it logs anything, where, and for how long?
  • Which prompt and model version produced a given answer, and can we find it by request ID?
  • Are there automated tests for the error cases, not just the happy path?

Common misconceptions

  • "The API is just a thin wrapper around the model." The model call is a few lines. The contract, checks, limits and fallback are most of the work, and most of the value.
  • "If the model is good, we don't need to check its answers." Even strong models sometimes return answers that break the agreed format or overstep. The API checks every answer.
  • "An error is the honest response when the model fails." Sometimes. For many business tasks, a simpler rule-based answer, clearly labeled, is more useful than an error.
  • "Internal APIs don't need keys." Internal callers loop, retry and leak too. Every caller should be known and limitable.
  • "Caching AI answers is cheating." If the question and the prompt version are the same, the same answer is often fine, and much cheaper. The cache key must include the prompt version.

Key terms

  • API (application programming interface): a defined way for one program to ask another for something.
  • Contract: the agreed request fields, response fields and error codes of an API.
  • Endpoint: one address of an API, such as /v1/explain.
  • Status code: a three-digit number in every web response. 200 means success; 401 means "who are you?"; 404 means "not found"; 422 means "your request is malformed"; 429 means "too many requests".
  • API key: a long secret string that identifies a calling program.
  • Rate limit: the maximum number of calls one caller may make in a time window.
  • Fallback: a safe, simpler answer the API gives when the model fails.
  • Request ID: a unique label on each call, used to find it in logs later.
  • Prompt version: a name for the exact prompt in use, so answers can be traced to it.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What is the main job of an AI API around a model call?

    Answer: B. The API promises fixed requests, responses and errors, and keeps that promise when the model is slow, down or wrong. It doesn't make the model smarter; it makes the service around it dependable.
  2. 2The model provider has an outage for an hour. What should the sales rep see in a well-built blocked-order API?

    Answer: C. A fallback keeps the business process moving with a safe, simpler answer. Labeling it as a fallback keeps it honest, so nobody mistakes it for the model's analysis.
  3. 3Why does a well-built AI API need rate limits and a cache?

    Answer: A. Every model call is paid. OWASP lists unlimited resource use as a top API risk and recommends spending limits or alerts. A cache avoids paying twice for the same question.
  4. 4A partner says the model is strong, so checking its answers is unnecessary. What is the best reply?

    Answer: D. The API checks every answer against a schema before anyone sees it. That catches the rare answer that breaks the format or oversteps, such as claiming an order was released.
  5. 5As of October 2026, which combination does SAP's Architecture Center describe for generative AI applications on BTP?

    Answer: B. SAP's AI golden path names CAP as the backend framework on Cloud Foundry or Kyma, with SAP Cloud SDK for AI for model access. The orchestration service and prompt registry sit around those calls.
  6. 6Someone asks why the API said order 4711 needed a credit review. What makes the question answerable?

    Answer: C. With the request ID you find the call in the logs, and the model name and prompt version tell you exactly what produced it. Without them the answer can't be traced.
  7. 7What is the most useful question to ask a vendor about their AI API?

    Answer: D. Outages and slow calls will happen. Knowing the fallback behavior tells you whether your business process keeps running, which matters more than the language or the newest model.
Deep layer · 40 min read

Mental model: a stable contract around an unstable part

Treat the model as an unreliable supplier inside your service. It is usually good, sometimes slow, occasionally down, and now and then returns something that breaks the agreement. Your API is the buyer who checks every delivery.

The contract faces outward and never changes by accident. The model faces inward and can be swapped, retried or bypassed. Everything in this topic is one of three moves: check what comes in, check what comes out, and decide what to do when the middle fails.

flowchart LR
  C[Caller<br/>app or workflow] -->|request| K{Key and<br/>rate limit}
  K -->|ok| V{Input<br/>valid?}
  V -->|yes| CA{In cache?}
  CA -->|no| M[Model call<br/>with timeout]
  M -->|answer| S{Fits the<br/>schema?}
  S -->|yes| R[Response<br/>source: model]
  S -->|no| F[Fallback<br/>source: fallback]
  M -->|error or timeout| F
  CA -->|yes| RC[Response<br/>source: cache]
  K -->|no| E1[401 or 429]
  V -->|no| E2[422 or 404]

How it works

The contract: request, response, errors

The contract has three parts, and each should be written down before any code.

The request. Accept the smallest input that does the job. For the blocked-order explainer, that is an order number and an optional length. Validate both: the order number is digits only, and the length sits within a range. Never accept a free-text prompt from the caller if the task doesn't need one; a free-text field is an open door for prompt injection (covered in Unit 11).

The response. Return fields, not prose. The explanation is text, but around it sit fields a program can act on: the suggested next_step from a fixed list, the source (model, cache or fallback), the model name, the prompt version and a request ID. The caller's code branches on fields, never on parsing sentences.

The errors. Use standard HTTP status codes so every caller understands them:

Code Meaning in this API Who fixes it
200 An answer, from the model, the cache or the fallback Nobody
401 No API key, or the wrong one The caller
404 The order isn't there The caller
422 The request body doesn't match the contract The caller
429 Too many requests; a Retry-After header says when to try again The caller waits
500 A bug in the API itself You

FastAPI produces several of these for you. Its documentation explains that when a request doesn't match the declared model, FastAPI raises a RequestValidationError and answers 422. For your own errors you raise HTTPException(status_code=..., detail=...), and you can attach headers such as Retry-After.

Notice what is missing from the table: a 5xx code for "the model failed". In this API, a model failure becomes a 200 with source: "fallback". That is a design choice. For an explanation, a safe rule-based answer is more useful than an error. For a task where a wrong answer is worse than none, such as posting a document, you would return 503 instead and let the caller decide.

Authentication: who may call

Every caller needs an identity. On your laptop, the simplest form is an API key: a long random string the caller sends in a header. FastAPI's APIKeyHeader reads the header and describes it in the API's documentation, so the /docs page can send it for you. With auto_error=False it returns nothing instead of failing, so your own function can decide what to do. The script compares keys with secrets.compare_digest, which takes the same time whether the first or last character differs, so an attacker can't guess the key by timing.

API keys are fine for a course and for machine-to-machine calls behind a gateway. On SAP BTP, apps normally use OAuth tokens instead. SAP's Python tutorial for Cloud Foundry puts the application router (a Node.js entry point) in front of the Python app and binds the XSUAA service. The Python app then validates each token with the sap-xssec library and checks a scope with check_scope. The tutorial shows a direct call without a token being refused with 403. The deployment topic later in this unit sets that up.

Authentication says who is calling. Authorization says what they may see. OWASP's top API risk is broken object-level authorization: a caller changes an ID in the request and reads data that isn't theirs. Here, that would be a rep asking about an order outside their sales area. The API must check access to each object it reads, not just that the caller is signed in. Unit 7 connects this to SAP authorizations.

Validating the model's answer

The model gets a strict JSON schema, as in Structured outputs and function calling. The schema lists two fields: explanation and next_step, where next_step must be one of four values. SAP's SDK reference describes the strict flag on JSONResponseSchema as enabling strict schema adherence.

A schema sent to the model is a request, not a guarantee. So the API validates the answer again with a Pydantic model, ModelAnswer, which also limits the length. If validation fails, the answer is thrown away and the fallback is used. The bad answer never reaches the caller. The model proposes; your code disposes.

Timeouts, retries and the fallback

A model call can hang. Set a timeout on every call. In the SAP Cloud SDK for AI, LLMModelDetails takes a timeout in seconds and a max_retries count; the reference notes both are ignored for Vertex AI models. Keep retries low (one is plenty): each retry adds waiting time for the user and cost for you.

When the call still fails, the SDK raises OrchestrationError, whose message says what went wrong. Network problems and expired credentials raise other exceptions. The script turns all of them into one internal ModelUnavailable error and serves the fallback. The fallback is boring on purpose: it builds a sentence from the order's reason field with a fixed rule. It is always available, costs nothing and never claims more than the data says.

Unit 5's orchestration topic showed a second kind of fallback: a list of model configurations that the orchestration service tries in order. Use both. The service-level fallback tries another model; your API-level fallback covers the case where orchestration itself is unreachable.

Rate limits and caching

The script keeps a list of recent call times per key. If a key made 10 calls in the last 60 seconds, the next call gets 429 and a Retry-After header. This is a sliding window limit, held in memory. It is enough for one process on one machine. With several instances, each keeps its own count, so production limits belong in a shared store or a gateway such as SAP API Management.

The cache stores good model answers for five minutes. Its key is the order number, the length, the model name and the prompt version. That last part matters: if you change the prompt and forget to change the version, old answers keep coming back. Fallback answers are never cached, so the next call tries the model again.

Request IDs and logging

Every response carries an X-Request-ID header, and the same ID appears in the response body and in the log line. If a caller sends its own ID (8 to 64 letters, digits or dashes), the API reuses it, so one ID can follow a request across systems. Anything else is replaced, so a caller can't inject odd text into your logs.

The log line has the ID, method, path, status and time. It has no request body, no customer names and no keys. The model call may need customer data; your logs don't. Unit 10 builds on this with tracing.

Sync, async and streaming

The script uses a plain def endpoint and the SDK's blocking run(), which is the simplest pattern to read. The SDK also offers arun() for async def endpoints, and stream(), which yields partial answers as the model writes them. Streaming suits chat screens. It suits this API less: you can't validate a JSON answer against a schema until it is complete.

Build it yourself: an AI API for blocked orders

You will turn last topic's hello_api.py into a real AI service. A sales app sends an order number; the API answers with a plain-words explanation and a suggested next step. Around the model call you will add a key, input checks, a schema check, a timeout, a fallback, a cache, a rate limit and request IDs. Then you will prove every promise with automated tests.

flowchart LR
  T[test_ai_api.py<br/>13 checks] -->|in memory| A[ai_api.py]
  B[Browser /docs<br/>or curl] -->|X-API-Key| A
  A -->|default| S[Sample model<br/>made-up answers]
  A -->|--llm| O[SAP orchestration<br/>service]

Before you start: complete Set up your computer for this course and Set up for Unit 6. They create your orchestrate-course folder with its .venv, and install FastAPI and Uvicorn. For the optional real model (Step 7) you also need Set up for Unit 5, which puts the AICORE_ lines in .env and installs sap-ai-sdk-gen. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder with the Unit 6 setup done (python check_unit06.py ends with All set).
  • About 60 to 90 minutes.
  • Cost: free with the sample model. Step 7 makes real model calls, a small per-request charge on a paid SAP AI Core account; on a trial, check what your plan allows.
  • No SAP account? Steps 1 to 6 and the exercise work without one.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that FastAPI is there:

    python -c "import fastapi, httpx; print('ready', fastapi.__version__)"

    You should see ready and a version number. The fastapi[standard] install from the Unit 6 setup includes httpx, which the tests need.

Run every command in this topic from the course folder.

Step 2: Create a key for your API

Your API will refuse callers without a key. You create the key once and keep it in .env, never in code.

  1. Print a new random key:

    python -c "import secrets; print(secrets.token_urlsafe(32))"

    You see a line of about 43 random letters, digits, dashes and underscores. Each run gives a different one.

  2. Open .env in VS Code (it is in the course folder; earlier units created it). Add this line at the end, with your key between the quotes, and save:

    ORCHESTRATE_API_KEY="paste-the-key-here"
  3. Check that Git ignores .env:

    git check-ignore .env

    It should print .env. If it prints nothing, add a line .env to .gitignore before you go on.

Step 3: Save the API

  1. In VS Code's file list, right-click unit06, choose New File, name it ai_api.py, paste the code below and save.
"""Unit 6: an AI API for blocked sales orders, built with FastAPI.

It puts a model behind POST /v1/explain, with the parts a real AI service needs: an API key,
input checks, a schema for the model's answer, a timeout, a fallback, a cache, a rate limit
and a request ID on every response.

How to run (from your course folder, with .venv turned on):
    python unit06/ai_api.py                            # sample model: made-up answers, no SAP account
    python unit06/ai_api.py --llm                      # real model through SAP's orchestration service
    python unit06/ai_api.py --llm --model MODEL_NAME   # another model from your catalog
    python unit06/ai_api.py --port 8080                # another port if 8000 is taken
It needs ORCHESTRATE_API_KEY in .env. --llm also needs the AICORE_ lines from "Set up for Unit 5".
Stop it with Ctrl+C.
"""
import argparse
import json
import logging
import os
import re
import secrets
import sys
import time
import uuid
from collections import defaultdict, deque
from typing import Literal, Optional

from fastapi import Depends, FastAPI, HTTPException, Request
from fastapi.security import APIKeyHeader
from pydantic import BaseModel, Field, ValidationError

PROMPT_VERSION = "explain-v1"   # change this whenever you change the prompt; it is part of the cache key

BLOCKED_ORDERS = {   # made-up data in the shape of the course's running example
    "4711": {"customer": "Made-up Retail GmbH", "reason": "Credit limit exceeded", "net_value": 12500.00, "currency": "EUR"},
    "4712": {"customer": "Example Foods Ltd", "reason": "Missing export documents", "net_value": 8300.50, "currency": "GBP"},
    "4713": {"customer": "Sample Tools Inc", "reason": "Credit limit exceeded", "net_value": 21000.00, "currency": "USD"},
    "4714": {"customer": "Demo Garden AG", "reason": "Incomplete delivery address", "net_value": 640.00, "currency": "CHF"},
}

NextStep = Literal["credit_review", "complete_documents", "fix_master_data", "contact_customer"]
NEXT_STEP_BY_REASON = {   # used by the fallback when the model can't answer
    "Credit limit exceeded": "credit_review",
    "Missing export documents": "complete_documents",
    "Incomplete delivery address": "fix_master_data",
}

# ---------- the contract: what callers send, what the model must return, what callers get ----------

class ExplainRequest(BaseModel):
    """What a caller sends to /v1/explain."""
    sales_order: str = Field(pattern=r"^[0-9]{1,10}$", description="SAP sales order number, digits only",
                             examples=["4711"])
    max_words: int = Field(default=60, ge=20, le=120, description="Longest explanation you want, in words")


class ModelAnswer(BaseModel):
    """What the model must return. Your code checks it; the model only proposes."""
    explanation: str = Field(min_length=1, max_length=800)
    next_step: NextStep


class ExplainResponse(BaseModel):
    """What the caller gets back, whatever happened inside."""
    sales_order: str
    explanation: str
    next_step: NextStep
    source: Literal["model", "cache", "fallback"]
    fallback_reason: Optional[str] = None
    model: str
    prompt_version: str
    tokens: int
    request_id: str


SCHEMA = {   # the JSON Schema sent to the model (strict mode needs every field required, no extras)
    "type": "object",
    "properties": {
        "explanation": {"type": "string", "description": "Why the order is blocked, in plain words"},
        "next_step": {"type": "string", "enum": list(NextStep.__args__),
                      "description": "Who should act next"},
    },
    "required": ["explanation", "next_step"],
    "additionalProperties": False,
}

SYSTEM = ("You explain blocked SAP sales orders to sales representatives in plain words. "
          "Use only the order data you are given. Never say that an order is released or will ship: "
          "only the responsible team can remove a block.")
USER = "Explain this blocked order in at most {{?max_words}} words.\n<order>{{?order}}</order>"


class ModelUnavailable(Exception):
    """The model didn't answer in time or the service returned an error."""


# ---------- two interchangeable models: a sample one and the real one ----------

class SampleModel:
    """Made-up answers, so the API runs without an SAP account. Two orders misbehave on purpose:
    4713 times out, and 4714 gets an answer that breaks the schema."""
    name = "sample-model"

    def __init__(self, delay: float = 0.4):
        self.delay = delay   # pretend the model takes a moment

    def answer(self, sales_order: str, order: dict, max_words: int) -> tuple[str, int]:
        time.sleep(self.delay)
        if sales_order == "4713":
            raise ModelUnavailable("sample model: simulated timeout")
        if sales_order == "4714":
            return json.dumps({"explanation": "The order is fine, I released it.", "next_step": "release_order"}), 95
        text = (f"Order {sales_order} for {order['customer']} is on hold. Reason in SAP: {order['reason']}. "
                f"It is worth {order['net_value']:,.2f} {order['currency']}. The responsible team must review "
                "it before it can ship.")
        return json.dumps({"explanation": text, "next_step": NEXT_STEP_BY_REASON[order["reason"]]}), 120


class OrchestrationModel:
    """The real model, called through SAP's orchestration service (version 2) with a strict JSON schema."""

    def __init__(self, model_name: str, timeout_seconds: int = 30):
        from dotenv import load_dotenv
        load_dotenv()
        needed = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
                  "AICORE_RESOURCE_GROUP"]
        missing = [n for n in needed if not os.environ.get(n)]
        if missing:
            sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5, "
                     "or leave out --llm to use the sample model.")
        from gen_ai_hub.orchestration_v2 import (JSONResponseSchema, LLMModelDetails, ModuleConfig,
                                                 OrchestrationConfig, OrchestrationService,
                                                 PromptTemplatingModuleConfig, ResponseFormatJsonSchema,
                                                 SystemMessage, Template, UserMessage)
        self.name = model_name
        template = Template(template=[SystemMessage(content=SYSTEM), UserMessage(content=USER)],
                            response_format=ResponseFormatJsonSchema(json_schema=JSONResponseSchema(
                                name="order_explanation", description="Why a sales order is blocked",
                                schema=SCHEMA, strict=True)))
        config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
            prompt=template, model=LLMModelDetails(name=model_name, params={"max_tokens": 400},
                                                   timeout=timeout_seconds, max_retries=1))))
        try:   # the SDK signs in and finds the orchestration deployment here, once, at startup
            self.service = OrchestrationService(config=config)
        except Exception as error:
            sys.exit(f"Could not connect to SAP AI Core ({type(error).__name__}). "
                     "Run python check_unit05.py to find the cause, or leave out --llm.")

    def answer(self, sales_order: str, order: dict, max_words: int) -> tuple[str, int]:
        from gen_ai_hub.orchestration_v2 import OrchestrationError
        try:
            response = self.service.run(placeholder_values={
                "order": json.dumps({"sales_order": sales_order, **order}), "max_words": str(max_words)})
        except OrchestrationError as error:
            raise ModelUnavailable(f"orchestration error {error.code}: {error.message}") from error
        except Exception as error:   # network errors, timeouts, expired credentials
            raise ModelUnavailable(f"{type(error).__name__}: {error}") from error
        result = response.final_result
        tokens = result.usage.total_tokens if result.usage else 0
        return result.choices[0].message.content or "", tokens


# ---------- the app ----------

def fallback_answer(sales_order: str, order: dict) -> dict:
    """A safe, rule-based answer for when the model fails. Boring on purpose."""
    return {"explanation": (f"Order {sales_order} for {order['customer']} is blocked: {order['reason'].lower()}. "
                            "The responsible team must review it before it can ship."),
            "next_step": NEXT_STEP_BY_REASON.get(order["reason"], "contact_customer")}


def create_app(model, api_key: str, rate_limit_per_minute: int = 10, cache_seconds: int = 300) -> FastAPI:
    """Build the API around any object with a .name and an .answer() method."""
    log = logging.getLogger("ai_api")
    app = FastAPI(title="Orchestrate blocked-order explainer", version="1.0.0",
                  description="Explains blocked SAP sales orders in plain words. Course example with made-up data.")
    cache: dict = {}                     # (order, words, model, prompt version) -> (time stored, answer)
    calls: dict = defaultdict(deque)     # API key -> times of recent calls

    @app.middleware("http")
    async def request_id_and_log(request: Request, call_next):
        incoming = request.headers.get("X-Request-ID", "")
        request_id = incoming if re.fullmatch(r"[A-Za-z0-9-]{8,64}", incoming) else uuid.uuid4().hex
        request.state.request_id = request_id
        started = time.perf_counter()
        response = await call_next(request)
        response.headers["X-Request-ID"] = request_id
        # One line per request: no body, no customer data, no keys.
        log.info("%s %s %s %s %.0fms", request_id, request.method, request.url.path, response.status_code,
                 (time.perf_counter() - started) * 1000)
        return response

    key_header = APIKeyHeader(name="X-API-Key", auto_error=False, description="Your ORCHESTRATE_API_KEY")

    def require_key(key: Optional[str] = Depends(key_header)) -> str:
        if not key or not secrets.compare_digest(key.encode(), api_key.encode()):
            raise HTTPException(status_code=401, detail="Missing or wrong API key",
                                headers={"WWW-Authenticate": "APIKey"})
        return key

    def rate_limit(key: str = Depends(require_key)) -> None:
        now = time.monotonic()
        recent = calls[key]
        while recent and now - recent[0] > 60:
            recent.popleft()
        if len(recent) >= rate_limit_per_minute:
            wait = int(60 - (now - recent[0])) + 1
            raise HTTPException(status_code=429, detail=f"Too many requests; try again in {wait} seconds",
                                headers={"Retry-After": str(wait)})
        recent.append(now)

    @app.get("/health")
    def health() -> dict:
        """Is the app running? No key needed, no model call."""
        return {"status": "ok", "model": model.name, "prompt_version": PROMPT_VERSION}

    @app.get("/v1/blocked-orders", dependencies=[Depends(require_key)])
    def blocked_orders() -> list:
        """List the made-up blocked orders."""
        return [{"sales_order": number, **order} for number, order in BLOCKED_ORDERS.items()]

    @app.post("/v1/explain", response_model=ExplainResponse, dependencies=[Depends(rate_limit)])
    def explain(body: ExplainRequest, request: Request) -> dict:
        """Explain why one order is blocked. Uses the model, a cached answer, or a safe fallback."""
        request_id = request.state.request_id
        order = BLOCKED_ORDERS.get(body.sales_order)
        if order is None:
            raise HTTPException(status_code=404, detail=f"No blocked order {body.sales_order} in the sample data")

        cache_key = (body.sales_order, body.max_words, model.name, PROMPT_VERSION)
        stored = cache.get(cache_key)
        if stored and time.monotonic() - stored[0] < cache_seconds:
            return {**stored[1], "source": "cache", "tokens": 0, "request_id": request_id}

        base = {"sales_order": body.sales_order, "model": model.name, "prompt_version": PROMPT_VERSION,
                "request_id": request_id}
        try:
            raw, tokens = model.answer(body.sales_order, order, body.max_words)
            checked = ModelAnswer.model_validate_json(raw)    # the model proposes, your code disposes
        except ModelUnavailable as error:
            log.warning("%s model unavailable: %s", request_id, error)
            return {**base, **fallback_answer(body.sales_order, order), "source": "fallback",
                    "fallback_reason": "model unavailable", "tokens": 0}
        except ValidationError:
            log.warning("%s model answer rejected: it did not match the schema", request_id)
            return {**base, **fallback_answer(body.sales_order, order), "source": "fallback",
                    "fallback_reason": "model answer rejected", "tokens": tokens}

        result = {**base, **checked.model_dump(), "source": "model", "tokens": tokens}
        cache[cache_key] = (time.monotonic(), result)   # only good model answers are cached
        return result

    return app


def main() -> None:
    import uvicorn
    from dotenv import load_dotenv

    parser = argparse.ArgumentParser(description="Run the Unit 6 AI API on your computer.")
    parser.add_argument("--llm", action="store_true", help="call a real model through SAP's orchestration service")
    parser.add_argument("--model", default="gpt-4o-mini", help="model name from your catalog (with --llm)")
    parser.add_argument("--port", type=int, default=8000, help="port to listen on (default 8000)")
    parser.add_argument("--rate-limit", type=int, default=10, help="requests per minute per key (default 10)")
    args = parser.parse_args()

    load_dotenv()   # reads ORCHESTRATE_API_KEY (and the AICORE_ lines) from .env in your course folder
    api_key = os.environ.get("ORCHESTRATE_API_KEY", "")
    if len(api_key) < 20:
        sys.exit("ORCHESTRATE_API_KEY is missing or shorter than 20 characters in .env. See Step 2.")

    logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
    model = OrchestrationModel(args.model) if args.llm else SampleModel()
    app = create_app(model, api_key, rate_limit_per_minute=args.rate_limit)
    print(f"Model: {model.name}. Open http://127.0.0.1:{args.port}/docs in your browser. Press Ctrl+C to stop.")
    uvicorn.run(app, host="127.0.0.1", port=args.port, access_log=False)


if __name__ == "__main__":
    main()

Step 4: Run the contract tests (no server, no account)

Before you start the server, prove the API keeps its promises. FastAPI's TestClient sends requests to the app directly in memory, without a running server, as its testing guide explains.

  1. In unit06, create test_ai_api.py, paste the code below and save.
"""Unit 6: contract tests for the AI API. No server, no SAP account, no model costs.

FastAPI's TestClient sends requests to the app in memory, with the sample model behind it.

How to run (from your course folder, with .venv turned on):
    python unit06/test_ai_api.py
"""
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).parent))   # so Python finds ai_api.py next to this file

from fastapi.testclient import TestClient   # noqa: E402

from ai_api import SampleModel, create_app   # noqa: E402

KEY = "test-key-0123456789abcdef"
GOOD = {"X-API-Key": KEY}
client = TestClient(create_app(SampleModel(delay=0), KEY, rate_limit_per_minute=100))


def check(name: str, condition: bool, detail: str = "") -> bool:
    print(f"  {'PASS' if condition else 'FAIL'}  {name}" + (f"  ({detail})" if detail and not condition else ""))
    return condition


def explain(body: dict, headers: dict = GOOD):
    return client.post("/v1/explain", json=body, headers=headers)


results = []
print("Contract tests for unit06/ai_api.py")

r = client.get("/health")
results.append(check("health answers without a key", r.status_code == 200, r.text))

r = explain({"sales_order": "4711"}, headers={})
results.append(check("no key -> 401", r.status_code == 401, r.text))

r = explain({"sales_order": "4711"}, headers={"X-API-Key": "wrong-key-0123456789"})
results.append(check("wrong key -> 401", r.status_code == 401, r.text))

r = explain({"sales_order": "47-11"})
results.append(check("malformed order number -> 422", r.status_code == 422, r.text))

r = explain({"sales_order": "4711", "max_words": 500})
results.append(check("max_words out of range -> 422", r.status_code == 422, r.text))

r = explain({"sales_order": "9999"})
results.append(check("unknown order -> 404", r.status_code == 404, r.text))

r = explain({"sales_order": "4711"})
body = r.json()
results.append(check("4711 answered by the model", r.status_code == 200 and body.get("source") == "model", r.text))
results.append(check("next_step is credit_review", body.get("next_step") == "credit_review", r.text))
results.append(check("response carries X-Request-ID", bool(r.headers.get("X-Request-ID")), str(r.headers)))

r = explain({"sales_order": "4711"})
results.append(check("same question again -> cache", r.json().get("source") == "cache", r.text))

r = explain({"sales_order": "4713"})
body = r.json()
results.append(check("model timeout -> 200 with fallback",
                     r.status_code == 200 and body.get("fallback_reason") == "model unavailable", r.text))

r = explain({"sales_order": "4714"})
body = r.json()
results.append(check("answer that breaks the schema -> fallback",
                     body.get("fallback_reason") == "model answer rejected" and "released" not in body["explanation"],
                     r.text))

limited = TestClient(create_app(SampleModel(delay=0), KEY, rate_limit_per_minute=3))
codes = [limited.post("/v1/explain", json={"sales_order": "4712"}, headers=GOOD).status_code for _ in range(4)]
last = limited.post("/v1/explain", json={"sales_order": "4712"}, headers=GOOD)
results.append(check("4th call in a minute -> 429 with Retry-After",
                     codes == [200, 200, 200, 429] and "Retry-After" in last.headers, str(codes)))

passed = sum(results)
print(f"{passed} of {len(results)} passed")
sys.exit(0 if passed == len(results) else 1)
  1. Run it:

    python unit06/test_ai_api.py

What success looks like:

Contract tests for unit06/ai_api.py
  PASS  health answers without a key
  PASS  no key -> 401
  PASS  wrong key -> 401
  PASS  malformed order number -> 422
  PASS  max_words out of range -> 422
  PASS  unknown order -> 404
  PASS  4711 answered by the model
  PASS  next_step is credit_review
  PASS  response carries X-Request-ID
  PASS  same question again -> cache
f53ed1db343645b9ad66f4f59daaf47b model unavailable: sample model: simulated timeout
  PASS  model timeout -> 200 with fallback
6698aaf08eb14769b4dbc0ae91075760 model answer rejected: it did not match the schema
  PASS  answer that breaks the schema -> fallback
  PASS  4th call in a minute -> 429 with Retry-After
13 of 13 passed

The two lines without PASS are the API's own warnings, printed because two tests make the model fail on purpose. Your IDs will differ. Any FAIL line shows what came back in brackets; compare it with your code.

Step 5: Start the API and call it

  1. Start it with the sample model:

    python unit06/ai_api.py

What success looks like:

Model: sample-model. Open http://127.0.0.1:8000/docs in your browser. Press Ctrl+C to stop.
INFO:     Started server process [463]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
  1. Open http://127.0.0.1:8000/docs. Click Authorize near the top right. In the box for APIKeyHeader, paste your key from .env (without the quotes), click Authorize, then Close.

  2. Click POST /v1/explain, then Try it out. Set the body to {"sales_order": "4711"} and click Execute. You should see status 200 and "source": "model". Click Execute again: now "source": "cache" and "tokens": 0.

  3. Or call it from a second terminal (Terminal > New Terminal; the first one is busy). Turn on .venv there too (Step 1). First load your key into the terminal:

    • Windows (PowerShell):

      $key = (Select-String -Path .env -Pattern '^ORCHESTRATE_API_KEY="(.*)"').Matches[0].Groups[1].Value
      Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/v1/explain -ContentType "application/json" -Headers @{"X-API-Key" = $key} -Body '{"sales_order": "4711"}'
    • macOS / Linux:

      key=$(grep '^ORCHESTRATE_API_KEY=' .env | cut -d'"' -f2)
      curl -X POST http://127.0.0.1:8000/v1/explain -H "Content-Type: application/json" -H "X-API-Key: $key" -d '{"sales_order": "4711"}'

What success looks like (macOS/Linux; PowerShell shows the same values as a list):

{"sales_order":"4711","explanation":"Order 4711 for Made-up Retail GmbH is on hold. Reason in SAP: Credit limit exceeded. It is worth 12,500.00 EUR. The responsible team must review it before it can ship.","next_step":"credit_review","source":"model","fallback_reason":null,"model":"sample-model","prompt_version":"explain-v1","tokens":120,"request_id":"b1bdc2aad5104f299c4a857fc6854a69"}

If you already called 4711 in the browser, source is cache instead. Both are correct.

  1. Look at the first terminal. Each call left one line:
2026-10-02 09:25:11,494 INFO b1bdc2aad5104f299c4a857fc6854a69 POST /v1/explain 200 403ms
2026-10-02 09:25:11,503 INFO 564fab14225e4932a55d3e7c9799cfda POST /v1/explain 200 1ms

The first call took about 400 ms (the sample model's pretend delay); the cached one took 1 ms. No customer name and no key appears in the log.

Step 6: Make it fail on purpose

Run these from the second terminal. Each one tests a promise from the contract.

  • macOS / Linux:

    curl -X POST http://127.0.0.1:8000/v1/explain -H "Content-Type: application/json" -H "X-API-Key: $key" -d '{"sales_order": "4713"}'
    curl -X POST http://127.0.0.1:8000/v1/explain -H "Content-Type: application/json" -d '{"sales_order": "4711"}'
    curl -X POST http://127.0.0.1:8000/v1/explain -H "Content-Type: application/json" -H "X-API-Key: $key" -d '{"sales_order": "47-11"}'
  • Windows (PowerShell): Invoke-RestMethod stops with a red error for status codes 400 and above. The message still contains the API's answer.

    Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/v1/explain -ContentType "application/json" -Headers @{"X-API-Key" = $key} -Body '{"sales_order": "4713"}'
    Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/v1/explain -ContentType "application/json" -Body '{"sales_order": "4711"}'
    Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8000/v1/explain -ContentType "application/json" -Headers @{"X-API-Key" = $key} -Body '{"sales_order": "47-11"}'

What success looks like (shortened):

{"sales_order":"4713","explanation":"Order 4713 for Sample Tools Inc is blocked: credit limit exceeded. The responsible team must review it before it can ship.","next_step":"credit_review","source":"fallback","fallback_reason":"model unavailable",...}
{"detail":"Missing or wrong API key"}
{"detail":[{"type":"string_pattern_mismatch","loc":["body","sales_order"],"msg":"String should match pattern '^[0-9]{1,10}$'",...}]}

The sample model times out on 4713 on purpose, so you see the fallback. Try 4714 too: the sample model "releases" the order with a next_step that isn't on the list, the API rejects it, and you get the fallback with "fallback_reason": "model answer rejected". Press Ctrl+C in the first terminal when you are done.

Step 7 (optional): Put a real model behind it

This step needs the AICORE_ lines in .env from Set up for Unit 5. Each call is a small per-request charge on a paid account.

  1. Start the API with the real model:

    python unit06/ai_api.py --llm

    The first line now says Model: gpt-4o-mini. If that model isn't offered in your account, run python check_unit05.py to list usable models and add --model MODEL_NAME.

  2. Repeat Step 5. The explanation is now the model's own wording, tokens shows the real token count, and next_step still comes from the fixed list. Ask for 4713: the real model answers it normally, because only the sample model fails on purpose.

  3. Stop with Ctrl+C.

What each part of the script does:

Part What it does
ExplainRequest The request contract: digits-only order number, max_words between 20 and 120; anything else gets 422
ModelAnswer and SCHEMA What the model must return; the schema goes to the model, ModelAnswer checks the answer again
ExplainResponse The response contract, including source, prompt_version and request_id
SampleModel Made-up answers with no account; fails on 4713 and 4714 on purpose
OrchestrationModel The real model through SAP's orchestration service, with a strict schema, a 30-second timeout and one retry; built once at startup
fallback_answer A fixed-rule answer for when the model fails
request_id_and_log Adds X-Request-ID to every response and writes one log line without customer data
require_key Reads X-API-Key and compares it safely with your key; 401 if missing or wrong
rate_limit At most --rate-limit calls per key per minute; 429 with Retry-After beyond that
explain Looks up the order, tries the cache, calls the model, checks the answer, falls back if needed
create_app Builds the API around any model object, so tests can use the sample model
--llm, --model, --port, --rate-limit Optional: real model, model name, port, and limit

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on PATH, or .venv is off Turn on .venv (Step 1); see Set up your computer if Python itself is missing
ModuleNotFoundError: No module named 'fastapi' (or httpx) The library isn't in this .venv Check for (.venv) in the prompt, then pip install -r requirements.txt
ModuleNotFoundError: No module named 'gen_ai_hub' with --llm The SAP Cloud SDK for AI isn't installed Follow Step 2 of Set up for Unit 5, or leave out --llm
ORCHESTRATE_API_KEY is missing or shorter than 20 characters The key line isn't in .env, or you ran the script from another folder Do Step 2; run from the course folder
{"detail":"Missing or wrong API key"} The header is missing or the key differs from .env Load the key again (Step 5); restart the API after editing .env
Missing in .env: AICORE_... or Could not connect to SAP AI Core The Unit 5 settings are missing, wrong or expired Run python check_unit05.py; or leave out --llm
Every answer has "fallback_reason": "model unavailable" with --llm The model name isn't offered, or the network or a proxy blocks SAP's servers Read the warning line in the first terminal; try --model with a name from check_unit05.py; ask IT about the proxy
address already in use Port 8000 is taken, often by hello_api.py Stop the other program with Ctrl+C, or add --port 8080 and use that port in your calls
{"detail":"Too many requests; try again in 41 seconds"} You hit the rate limit Wait, or start with --rate-limit 30 while experimenting
A test shows FAIL Your file differs from the published code Copy the code again; the text in brackets shows what came back

The SAP way

The script runs on your laptop with an API key. On SAP BTP, the same ideas map to managed services. As of October 2026:

Concern In this topic On SAP BTP
Framework FastAPI CAP, or a Python app on Cloud Foundry or Kyma; SAP's Architecture Center names CAP as the backend framework for generative AI apps
Model access OrchestrationService from SAP Cloud SDK for AI (Python) The same SDK, also available for Java and JavaScript
Prompt A constant in the code, with PROMPT_VERSION The prompt registry, which SAP recommends over hard-coded prompts
Model fallback Your rule-based fallback_answer Orchestration's model fallback, plus your own fallback for when orchestration is unreachable
Caller identity X-API-Key XSUAA or SAP Cloud Identity Services; Python apps validate tokens with sap-xssec
Rate limits, publishing An in-memory sliding window API Management in SAP Integration Suite: API key and OAuth checks, rate limiting, spike arrest, caching, analytics, developer portal
Connections to S/4HANA Made-up data in the script Destinations to SAP back ends

Calling the model. The SDK reference shows the pattern the script uses: build an OrchestrationService with a configuration, call run(placeholder_values=...), and read final_result.choices[0].message.content. Errors arrive as OrchestrationError. The response also has a request_id from SAP's side; log it next to your own ID when you need to raise a support ticket.

Securing a Python app. SAP's tutorial for Python on Cloud Foundry lists sap-xssec in requirements.txt and uses xssec.create_security_context(...) and check_scope(...) on the token. It reads the bound XSUAA credentials with the cfenv library and refuses requests without an authorization header with 403. In the tutorial, the application router handles sign-in and passes the token on. This is a sketch of what replaces require_key, adapted to FastAPI:

# Sketch: needs a BTP app bound to an XSUAA instance (set up in "Deploying AI apps on SAP BTP", later in Unit 6).
from cfenv import AppEnv
from fastapi import HTTPException, Request
from sap import xssec

env = AppEnv()
uaa_service = env.get_service(name="my-xsuaa").credentials   # the XSUAA instance bound to this app
REQUIRED_SCOPE = "openid"   # the scope SAP's tutorial checks; your app defines its own in the deployment topic


def require_token(request: Request):
    header = request.headers.get("authorization", "")
    if not header.startswith("Bearer "):
        raise HTTPException(status_code=403, detail="No token")
    security_context = xssec.create_security_context(header[len("Bearer "):], uaa_service)
    if not security_context.check_scope(REQUIRED_SCOPE):
        raise HTTPException(status_code=403, detail="Missing scope")
    return security_context

Licensing notes. SAP's Architecture Center lists the SAP AI Core extended plan for model access through the generative AI hub; what each call costs is covered in Choosing and calling LLMs. API Management is part of SAP Integration Suite, a separate service. Check your company's entitlements before you plan on either.

Build vs. SAP

Situation Build it yourself (FastAPI and your own checks) Use SAP services
Learning, a prototype, a proof of concept Fastest; everything is visible Overhead you don't need yet
One team, one caller, low volume Fine, with a key and the tests from this topic Optional
Several teams or partners call the API Hard to manage keys and limits by hand API Management for keys, limits, analytics and a portal
Users sign in with company accounts You would build token handling yourself XSUAA or SAP Cloud Identity Services with sap-xssec
The answer must respect S/4HANA authorizations Easy to get wrong Destinations to SAP back ends, with authorizations covered in Unit 7
Business data model, Fiori front end, SAP-style services Extra work CAP, calling the Python AI service or the SDK directly
Prompts change often, by non-developers Code releases for every change The prompt registry

A common project shape uses both: a CAP service owns the business data and the user's identity, and calls a small Python AI service like this one for the model step.

Production concerns

  • Secrets. Keys live in .env on your laptop and in a secret store or service binding on BTP. Never log them, never return them, and rotate them when people leave.
  • Authorization per object. Being signed in isn't enough. Check that the caller may see this order before you read it, as OWASP's top API risk warns.
  • Cost caps. OWASP recommends spending limits for paid services, or billing alerts where limits aren't possible. Combine a per-caller rate limit, a cache and a monthly alert on SAP AI Core usage.
  • Timeouts end to end. The caller, the API and the model call each need a timeout, and the inner ones must be shorter. Otherwise the caller gives up while the API still pays for the answer.
  • Versioning. The path starts with /v1. A change that breaks callers goes in /v2, with /v1 kept for a while. PROMPT_VERSION changes with every prompt edit.
  • Evaluation. The contract tests prove the plumbing. They don't prove the explanations are good. Unit 8 builds an evaluation harness for that.
  • Observability. Log the request ID, status, time, model, prompt version, tokens and source. A rising fallback rate is your first sign of a model problem. Unit 10 covers tracing.
  • Data protection. The model sees order data; logs don't need it. If prompts include personal data, use orchestration's masking module.
  • Clean core. The API reads SAP data through released APIs and runs outside S/4HANA, on BTP. It doesn't modify the core system.
  • Scaling. The in-memory cache and rate limit work for one instance only. With several instances, move both to a shared store or a gateway.

Pitfalls

  • Accepting a free-text prompt from callers. Accept the fields the task needs and build the prompt yourself.
  • Trusting the schema you sent. Validate the model's answer again in your code.
  • Retrying without limits. Three retries on a 30-second timeout can mean a 2-minute wait, paid four times.
  • Caching without the prompt version. Old answers survive the prompt change you just made.
  • Caching fallbacks. One outage then serves fallback answers for the whole cache period.
  • Logging request bodies "for debugging". Customer data ends up in a log that many people can read.
  • Returning the model's raw error. Callers see provider details they can't act on. Log the detail; return a clean answer or code.
  • Building the SDK client on every request. It signs in each time. Build it once at startup, as the script does.
  • Hiding fallbacks. If source isn't in the response, nobody can tell a model answer from a rule.

Exercise: add a second endpoint and prove it with tests

You will add a POST /v1/explain-batch endpoint that explains up to five orders in one call. This is the shape the Unit 9 order-exception agent will call.

  1. Open unit06/ai_api.py. Below ExplainRequest, add a request model:

    class BatchRequest(BaseModel):
        """Up to five orders in one call."""
        sales_orders: list[str] = Field(min_length=1, max_length=5)
  2. Inside create_app, below the explain function, add the endpoint. It reuses explain for each order and skips orders that aren't found:

    @app.post("/v1/explain-batch", dependencies=[Depends(rate_limit)])
    def explain_batch(body: BatchRequest, request: Request) -> dict:
        """Explain several orders. Unknown orders are listed, not fatal."""
        answers, not_found = [], []
        for number in body.sales_orders:
            try:
                answers.append(explain(ExplainRequest(sales_order=number), request))
            except (HTTPException, ValidationError):
                not_found.append(number)
        return {"answers": answers, "not_found": not_found, "request_id": request.state.request_id}
  3. Open unit06/test_ai_api.py. Above the line passed = sum(results), add two tests:

    r = client.post("/v1/explain-batch", json={"sales_orders": ["4711", "4712", "9999"]}, headers=GOOD)
    results.append(check("batch explains two and lists one unknown",
                         len(r.json().get("answers", [])) == 2 and r.json().get("not_found") == ["9999"], r.text))
    r = client.post("/v1/explain-batch", json={"sales_orders": ["4711"] * 6}, headers=GOOD)
    results.append(check("batch of six -> 422", r.status_code == 422, r.text))
  4. Run the tests:

    python unit06/test_ai_api.py

    You should see 15 of 15 passed.

  5. Start the API and open http://127.0.0.1:8000/docs. Find POST /v1/explain-batch and try it with {"sales_orders": ["4711", "4714"]}. Note which answer came from the model and which from the fallback.

  6. In unit06, create api_contract.md. In plain words, write the request fields, the response fields, and one line for each status code your API can return. This is the start of the solution design document later in this unit.

  7. Commit your work:

    git add unit06/ai_api.py unit06/test_ai_api.py unit06/api_contract.md
    git commit -m "Add the blocked-order AI API with contract tests"

    Don't add .env; git status should not list it.

Done when: python unit06/test_ai_api.py prints 15 of 15 passed, the batch endpoint answers in /docs, and api_contract.md is committed with every status code explained.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In this API, what happens when the model returns {"next_step": "release_order", ...}?

    Answer: C. A schema sent to the model is a request, not a guarantee. The API validates the answer again with ModelAnswer; release_order isn't in the allowed list, so the answer is thrown away and the fallback is served with fallback_reason: "model answer rejected".
  2. 2Why does the API answer 200 with source: "fallback" instead of 503 when the model is down?

    Answer: B. This is a design choice for this task. An explanation from a fixed rule keeps the sales rep working, and source tells the caller what happened. For a task where a wrong answer is worse than none, 503 would be the better contract.
  3. 3Which cache key avoids serving stale answers after you edit the prompt?

    Answer: D. The script's cache key includes PROMPT_VERSION and the model name. Change either and old answers no longer match. Keying on the order number alone keeps serving answers from the old prompt.
  4. 4Why does require_key use secrets.compare_digest instead of ==?

    Answer: C. A plain comparison can stop at the first differing character, which leaks timing information. compare_digest compares in constant time, so an attacker can't guess the key one character at a time.
  5. 5You run three copies of this API behind a load balancer, each with --rate-limit 10. What limit does one caller really get?

    Answer: B. The sliding window lives in each process's memory, so each copy keeps its own count. Production limits belong in a shared store or a gateway such as SAP API Management.
  6. 6How does the script call a model through SAP's orchestration service?

    Answer: D. Building the service signs in and finds the deployment, so the script does it once. Each request calls run(placeholder_values=...) and reads final_result.choices[0].message.content. Streaming doesn't fit, because a JSON answer can't be validated until it is complete.
  7. 7A caller sends order 4712 but is not responsible for that customer. Signed-in callers all pass require_key. What is missing?

    Answer: C. Authentication says who is calling; it doesn't say which orders they may see. OWASP's top API risk is exactly this gap. The API must check access to each object it reads, which Unit 7 ties to SAP authorizations.
  8. 8When you deploy this API to SAP BTP Cloud Foundry, what replaces the X-API-Key check, following SAP's Python tutorial?

    Answer: A. SAP's tutorial binds XSUAA, puts the application router in front for sign-in, and validates each token with sap-xssec and check_scope. Keys in manifest.yml would end up in source control, and the orchestration credentials identify your app, not its callers.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in