Orchestrate

REST API design for AI services

Design an AI API others can rely on: resources, errors, versions, pages, safe retries, background jobs, streaming and an OpenAPI contract.

Updated Oct 3, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An API is how one program asks another for something. When your team wraps an AI model in an API, other teams, partners and workflows start to depend on it. API design is the set of rules that makes that dependence safe.

Good design answers a handful of plain questions. What are the things the API offers? How does it say "that failed, and here is why"? What happens when a caller retries? How does it change without breaking old callers? What does it do when the work takes two minutes, not two seconds?

AI services make these questions sharper. A model call is slow, costs money per request and gives a slightly different answer each time. So an AI API needs three things a plain data API can often skip: safe retries, background jobs for slow work, and streaming for answers people read as they appear.

The rules are written down in one file, the OpenAPI specification. It is the contract. Other teams build against it, and tools check it automatically.

Why it matters to the business

Take the running example: the blocked-order explainer from Building an AI API. A sales app asks it why order 4711 is blocked and gets a plain-words answer with a suggested next step.

Now picture three bad days.

  • The double task. The network drops just as the API answers. The sales app retries. The API runs the model twice and opens two credit-review tasks for the same order. Finance works the same case twice. A design rule called idempotency would have made the retry return the first answer.
  • The timeout at month end. A team sends 400 blocked orders in one call. The model needs a few seconds each. The call times out after a minute, and nobody knows which orders were done. A background job would have accepted the work, returned at once, and let the caller check progress.
  • The silent break. A developer renames a field from next_step to nextAction. Every caller that reads next_step breaks overnight. Versioning and a published contract would have caught it before release.

Each of these costs real money: duplicate work, failed batch runs and incident calls. None needs new technology. They need decisions made early and written down.

Design also decides adoption. An API with clear errors and a published contract is one that partner teams can use without a meeting. That is how an AI pilot turns into a service many processes use.

How SAP does it

SAP's own APIs show these rules in practice. As of October 2026:

  • Published contracts. The SAP Business Accelerator Hub (api.sap.com) lists SAP's APIs with their specification files. After logging in you can download EDMX files for OData services and JSON or YAML files for OpenAPI services. Teams generate client code from them.
  • OData conventions. Many SAP business APIs use OData, an open standard with its own answers to the same questions. The standard pages with $top and $skip, or a server-provided next link. Updates can carry an If-Match header so two people can't overwrite each other's changes. Slow requests can be handed off with Prefer: respond-async and a 202 Accepted answer.
  • Extra security on changes. SAP's Cloud SDK fetches a CSRF token by default before requests that change data, such as POST or PATCH, and sends it along. Your own code must do the same when the SAP system requires it.
  • Streaming for AI. The SAP Cloud SDK for AI can stream model answers from the orchestration service in chunks, using the server-sent events standard, so a chat screen shows words as they arrive.
  • Long-running AI resources. In SAP AI Core, a model deployment is a resource with a status such as RUNNING or STOPPED. You ask for a change of status and check back, rather than waiting on one long call.

The lesson for your own AI services: follow the same habits SAP's APIs follow, so your service feels familiar to the teams who already call SAP. Gateways such as SAP API Management add keys and rate limits on top (covered in Building an AI API); they do not fix a poorly designed contract.

A decision guide: eight choices to make before the first caller

Choice The plain question What goes wrong if you skip it
Resources What "things" does the API offer, such as orders or explanations? Endpoints named after actions pile up and nobody can guess the next one
Status codes How does the caller know it worked, or whose fault a failure is? Callers read error text to guess, and break when the wording changes
Error format Do all errors look the same? Each team writes its own error handling for each endpoint
Versioning How will it change without breaking callers? A rename breaks production integrations overnight
Pagination How does a caller get 10,000 items? Slow, huge responses, or silently missing items
Idempotency What happens when a caller retries? Duplicate tasks, duplicate charges, duplicate documents
Background jobs What happens when work takes minutes? Timeouts with no record of what was finished
Streaming Should people see the answer as it is written? Chat screens that look frozen for ten seconds

You don't need all eight on day one. You need a written answer to each, even if the answer is "not yet, and here is why".

Questions to ask

  • Where is the OpenAPI specification, and who approves changes to it?
  • What happens if a caller sends the same request twice? Show me.
  • How will you change a field name a year from now? What is the version policy, and how long do old versions live?
  • Which requests can take more than 30 seconds, and how does the caller learn the result?
  • Does every error say whose fault it is (caller or service) in a standard format?
  • How do partner teams get access to the contract, a test key and sample data?
  • If the model provider is down, which status code does the caller see, and is it documented?

Common misconceptions

  • "REST means JSON over HTTP." JSON is just the format. REST design is about resources, standard methods and status codes that every caller understands the same way.
  • "Retries are the caller's problem." Retries happen on every network. If the API can't recognize a retry, it will do the work twice. That is the service's design problem.
  • "We'll add versioning when we need it." The first breaking change is when you need it, and by then callers already depend on the unversioned paths.
  • "Streaming makes the model faster." The total time is about the same. Streaming makes the wait feel shorter because people read while the model writes.
  • "The documentation is the contract." A wiki page drifts. A machine-readable OpenAPI file, checked automatically on every change, is the contract.

Key terms

  • Resource: a thing the API offers, such as a blocked order or an explanation, with its own address.
  • HTTP method: the verb of a request: GET reads, POST creates, PUT replaces, PATCH changes, DELETE removes.
  • Status code: a three-digit number in every answer: 2xx worked, 4xx the caller must fix something, 5xx the service failed.
  • Problem details: a standard JSON shape for errors (RFC 9457), so all errors look alike.
  • Pagination: splitting a long list into pages the caller fetches one by one.
  • Idempotency key: a unique value the caller sends so a retried request isn't processed twice.
  • Background job: work the API accepts now and finishes later; the caller checks its status.
  • Server-sent events (SSE): a standard way for a server to send a stream of small messages over one HTTP response.
  • OpenAPI specification: a machine-readable file describing every endpoint, field and error of an API.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A sales app retries a request after a network drop, and the AI service opens two credit-review tasks for one order. Which design rule was missing?

    Answer: B. Idempotency lets the service recognize a retried request and return the first result instead of doing the work again. Pagination, streaming and versioning solve other problems and would not stop the duplicate task.
  2. 2A team plans to send 400 orders to the AI service in one call. What is the better design?

    Answer: C. A background job returns at once with a way to check progress, so a slow batch never ends in a timeout with an unknown outcome. Longer timeouts only move the failure, and streaming suits one answer read by a person, not a batch.
  3. 3Why does a published OpenAPI specification matter to the business?

    Answer: A. The OpenAPI file is the contract in machine-readable form. Partner teams build against it and tools check every change, which is how a pilot becomes a service many processes use. It doesn't change model quality or replace security.
  4. 4A developer wants to rename next_step to nextAction next week. What should you ask first?

    Answer: D. A rename breaks every caller that reads the old field. The version policy says how breaking changes ship, for example under a new major version, and how long the old one stays available.
  5. 5Which statement about SAP's own APIs is accurate, as described in this topic?

    Answer: C. The SAP Business Accelerator Hub offers EDMX files for OData services and JSON or YAML files for OpenAPI services. OData has paging options, and the SAP Cloud SDK for AI can stream answers.
  6. 6A vendor says streaming will make their AI service faster. What is the accurate view?

    Answer: B. Streaming sends the answer in pieces as the model writes it, so the wait feels shorter. The model still takes about the same total time, and errors still need handling.
Deep layer · 40 min read

Mental model: the API is a promise to strangers

Your API will be called by code you never see, written by people you never meet, retried by networks you don't control. Design is deciding, in advance, what you promise those strangers in every situation: success, their mistake, your failure, a retry, a slow request.

An AI service adds three facts that shape those promises:

  1. It is slow. Seconds per answer, minutes per batch. So some work must move to background jobs, and some answers should stream.
  2. It costs money per call. So a retry must not mean a second model call.
  3. It is not deterministic. So the API's shape must be fixed even when the text changes. Callers branch on fields and status codes, never on wording.

The OpenAPI specification is the promise in writing. Write it, export it, check it automatically, and change it only on purpose.

flowchart LR
  D[Design rules] --> C[Code<br/>FastAPI app]
  C -->|export| S[openapi.json<br/>the contract]
  S --> L[Lint<br/>rules check]
  S --> G[Client code<br/>other teams]
  S --> T[Tests<br/>Unit 6 next topic]

How it works

Resources and verbs

REST design names resources (nouns) and uses the HTTP methods (verbs) on them. Instead of /getBlockedOrders and /explainOrder, you have:

Request Meaning
GET /v1/blocked-orders List blocked orders
GET /v1/blocked-orders/4711 Read one
POST /v1/explanations Create an explanation
GET /v1/explanations/exp_72f0c532e21c Read that explanation again
POST /v1/explanation-jobs Start a background job
GET /v1/explanation-jobs/job_31e3163166d3 Check the job

Two choices here are AI-specific. First, an explanation is a resource with an ID, not a fire-and-forget answer. The caller can fetch it again, link to it from a ticket and audit it later. Second, a slow batch is a job resource, so its progress has an address.

RFC 9110, the HTTP standard, sorts methods into two useful groups. Safe methods (GET, HEAD) only read. Idempotent methods (GET, PUT, DELETE and others) have the same intended effect whether sent once or many times. POST and PATCH are neither. That gap is why POST needs extra care, below.

Status codes: whose move is it?

A status code tells the caller's code what to do next without reading any text.

Code Meaning (RFC 9110 unless noted) In this API
200 OK Here is what you asked for Reads, job status
201 Created A new resource exists; Location says where New explanation
202 Accepted Accepted, not finished yet Background job started
401 Unauthorized Missing or wrong credentials No API key
404 Not Found No such resource Unknown order, explanation or job
409 Conflict Clashes with the resource's current state Same idempotency key still running
412 Precondition Failed A condition in the request headers was false OData updates with a stale If-Match (SAP section)
422 Unprocessable Content Understood, but the content is wrong Bad order number, reused idempotency key
429 Too Many Requests Slow down; see Building an AI API Rate limit
503 Service Unavailable The service can't answer right now Model down, when no fallback is safe

Rule of thumb: 4xx means the caller must change something; 5xx means you must. A model failure is never a 4xx.

One error format: problem details

If every endpoint invents its own error shape, every caller writes custom parsing. RFC 9457 defines one shape, served with the media type application/problem+json:

{
  "type": "https://example.com/problems/order-not-found",
  "title": "Blocked order not found",
  "status": 404,
  "detail": "There is no blocked order 9999 in the sample data.",
  "instance": "/v1/blocked-orders/9999",
  "request_id": "bef0da40d5654c2882032fa75985e76b"
}
  • type is a URI naming the kind of problem. Callers branch on it. When it is absent, RFC 9457 says it means about:blank.
  • title is a short summary that stays the same for that type.
  • status repeats the HTTP status code, and must match it.
  • detail explains this occurrence for a human. RFC 9457 says consumers should not parse it.
  • instance points at this occurrence; here, the path.
  • request_id is an extension member: RFC 9457 lets you add your own fields. It ties the error to your log line.

FastAPI's default errors look different ({"detail": ...}). The script below replaces its handlers so every error, including validation errors and unknown paths, comes back as problem details.

Versioning

Callers depend on field names, types and meanings. Any change that can break a caller is a breaking change: removing or renaming a field, changing a type, making an optional field required, changing what a status code means. Adding an optional field or a new endpoint is usually safe, if callers ignore fields they don't know (say so in your docs).

The simplest policy, used by this course:

  • Put the major version in the path: /v1/.... A breaking change ships as /v2/..., side by side with v1 for an announced period.
  • Put the full version in the contract: info.version in the OpenAPI file (1.1.0 here) changes with every release, so callers can see what moved.
  • Never change v1's behavior in place, even to "fix" it, if callers might depend on the old behavior.

For an AI service, the model and prompt version are not the API version. Swapping a model changes wording, not shape, so it doesn't need /v2. Report it in a field (model here; Building an AI API also returned prompt_version) so callers and auditors can see which one answered.

Pagination

Never return an unbounded list. Two common designs:

Design Request Strength Weakness
Offset ?limit=50&offset=100 Simple; jump to any page Items shift if data changes between calls, so rows get skipped or repeated; deep offsets get slow
Cursor ?limit=50&cursor=eyJh... Stable while data changes; fast Only "next page"; no jumping

The script uses a cursor: an opaque token the server hands out as next_cursor. Inside it is just {"after": "4713"}, base64-encoded, but callers must treat it as a black box. That lets you change how it works later without a new version. next_cursor is null on the last page.

OData, which many SAP APIs use, offers both styles: $top and $skip for client-driven paging, and server-driven paging where the response carries a next link with an opaque $skiptoken.

Idempotency: safe retries for POST

POST isn't idempotent, so a retried POST normally creates a second resource and, for an AI API, a second model call. The fix is an idempotency key. The caller generates a random value (a UUID) per action and sends it in a header. The server remembers the key with the result. A retry with the same key gets the stored result instead of new work.

The IETF is standardizing this as the Idempotency-Key header. As of its draft 07 (October 2025) it is still an Internet-Draft, not an RFC, but it already describes the behavior:

Situation Server answer
New key Do the work, store the result with the key
Same key, same body, first request finished Return the stored result, success or error
Same key, different body 422: the key was misused
Same key while the first request is still running 409 Conflict: retry later

The server needs to compare bodies, so it stores a fingerprint (a hash) of the request body with the key. Stored keys should expire after a documented time. One subtlety in the script: if the work fails, it forgets the key, so the caller can retry the same action after fixing the cause.

sequenceDiagram
  participant A as Sales app
  participant S as AI service
  A->>S: POST /v1/explanations<br/>Idempotency-Key: 3f2a...
  S->>S: new key: run model, store result
  S--xA: 201 Created (network drops)
  A->>S: same request, same key
  S->>A: 201 with stored result<br/>no second model call

Long-running work: 202 and a job to poll

Any request that might take longer than a caller's timeout becomes a job:

  1. POST /v1/explanation-jobs with the list of orders.
  2. The server answers at once with 202 Accepted, a Location header pointing at the job, and Retry-After: 1 (seconds to wait before checking).
  3. The caller polls GET on that location. The job shows status (queued, running, succeeded, failed), progress (done of total), results and per-item errors.

Per-item errors matter. One bad order number in a batch of 400 should not fail the other 399. The job records {"sales_order": "9999", "type": ".../order-not-found"} and carries on.

Polling is simple and works through every proxy. Its alternative is a callback (webhook): the caller registers a URL and the server calls it when done. Callbacks save polling traffic but need the caller to expose an endpoint, which many SAP landscapes won't allow. Start with polling.

OData has the same pattern: a client sends Prefer: respond-async, and a service may answer 202 Accepted with a Location to a status monitor.

Streaming: server-sent events

For one answer a person reads, streaming sends pieces as the model writes them. The standard for this over plain HTTP is server-sent events (SSE), media type text/event-stream. Each event is a few text lines:

event: delta
data: {"text": "Order "}
id: 0

event: done
data: {"id": "exp_c056d1f1c04f", "sales_order": "4714", "explanation": "...", "next_step": "fix_master_data", ...}

FastAPI supports SSE directly from version 0.135.0: declare response_class=EventSourceResponse and yield events. FastAPI also sends keep-alive pings and sets Cache-Control: no-cache for you, per its documentation.

Two design rules for AI streams:

  • Errors before the stream, not inside it. Once the first event is sent, the status is already 200. So check everything that can fail (key, input, order exists) before streaming starts. The script does this in a FastAPI dependency.
  • End with one complete, validated object. Callers that need fields, not text, wait for the final done event. As Building an AI API explained, you can't validate a JSON answer against its schema until it is complete; the done event is where that check has happened.

OpenAPI 3.2 added itemSchema to describe each item in streams such as text/event-stream. The FastAPI version tested here already uses it in the exported spec.

The contract: OpenAPI

FastAPI builds the OpenAPI document from your code: paths, parameters, Pydantic models and the responses= you declare. Three habits make it a good contract:

  • Give every operation an operationId and a summary. Code generators turn operationId into method names such as createExplanation.
  • Declare every error response, with its media type, not just the success.
  • Export it to a file and check it into Git, so every change shows up in a pull request diff.

Then check it with a linter: a script that reads the spec and enforces your rules. The walkthrough builds a small one.

Build it yourself: redesign the blocked-orders API

You will rebuild the blocked-orders API from Building an AI API around the rules above: versioned resources, problem-details errors, cursor pages, idempotent creates, a background job and a stream. Then you will export its OpenAPI contract and check it with your own linter. Everything runs on your laptop with made-up data and a sample model, so it is free.

flowchart LR
  D[designed_api.py<br/>--demo] -->|in memory| A[The API]
  B[Browser /docs<br/>or curl] -->|X-API-Key| A
  A -->|--export-openapi| S[openapi.json]
  S --> L[lint_api.py<br/>7 rules]

Before you start: complete Set up your computer for this course and Set up for Unit 6, and do Steps 1 and 2 of Building an AI API so .env has your ORCHESTRATE_API_KEY. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder with the Unit 6 setup done (python check_unit06.py ends with All set).
  • ORCHESTRATE_API_KEY in .env (from Building an AI API, Step 2).
  • About 60 minutes.
  • Cost: free. No SAP account and no model account needed.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that your FastAPI version has server-sent events (0.135.0 or newer):

    python -c "import fastapi; print(fastapi.__version__)"

    If the number is lower, run pip install --upgrade "fastapi[standard]".

Step 2: Save the API

  1. In VS Code's file list, right-click unit06, choose New File, name it designed_api.py, paste the code below and save.
"""Unit 6: the blocked-orders AI API, redesigned with REST rules.

It shows the design choices an AI service needs, each on a real endpoint:
  - resources and verbs:   /v1/blocked-orders, /v1/explanations, /v1/explanation-jobs
  - versioning:            the major version in the path (/v1)
  - pagination:            ?limit= and an opaque ?cursor= on the list
  - idempotency:           an Idempotency-Key header on POST /v1/explanations
  - one error format:      application/problem+json (RFC 9457) for every error
  - long-running jobs:     202 Accepted, a Location header and a job you poll
  - streaming:             server-sent events on POST /v1/explanations/stream
  - a contract:            the OpenAPI spec, exported to a file

How to run (from your course folder, with .venv turned on):
    python unit06/designed_api.py --demo                         # walk through every rule, no server
    python unit06/designed_api.py --export-openapi unit06/openapi.json
    python unit06/designed_api.py                                # start the server on port 8000
It needs ORCHESTRATE_API_KEY in .env (from "Building an AI API"). The model is a made-up sample model.
Stop the server with Ctrl+C.
"""
import argparse
import base64
import hashlib
import json
import os
import re
import secrets
import sys
import threading
import time
import uuid
from collections.abc import Iterable
from typing import Literal, Optional

from fastapi import Depends, FastAPI, Header, Query, Request, Response
from fastapi.exceptions import RequestValidationError
from fastapi.responses import JSONResponse
from fastapi.security import APIKeyHeader
from fastapi.sse import EventSourceResponse, ServerSentEvent
from pydantic import BaseModel, Field
from starlette.exceptions import HTTPException as StarletteHTTPException

API_VERSION = "1.1.0"           # the contract's version; the path only carries the major version (v1)
PROBLEM_BASE = "https://example.com/problems/"   # replace with a page that documents your error types

BLOCKED_ORDERS = {   # made-up data in the shape of the course's running example
    "4711": {"customer": "Made-up Retail GmbH", "reason": "Credit limit exceeded", "net_value": 12500.00, "currency": "EUR"},
    "4712": {"customer": "Example Foods Ltd", "reason": "Missing export documents", "net_value": 8300.50, "currency": "GBP"},
    "4713": {"customer": "Sample Tools Inc", "reason": "Credit limit exceeded", "net_value": 21000.00, "currency": "USD"},
    "4714": {"customer": "Demo Garden AG", "reason": "Incomplete delivery address", "net_value": 640.00, "currency": "CHF"},
    "4715": {"customer": "Fictional Motors SA", "reason": "Credit limit exceeded", "net_value": 47200.00, "currency": "EUR"},
    "4716": {"customer": "Placeholder Pharma BV", "reason": "Missing export documents", "net_value": 3150.75, "currency": "EUR"},
    "4717": {"customer": "Invented Textiles Srl", "reason": "Incomplete delivery address", "net_value": 980.00, "currency": "EUR"},
}
NEXT_STEP_BY_REASON = {"Credit limit exceeded": "credit_review", "Missing export documents": "complete_documents",
                       "Incomplete delivery address": "fix_master_data"}
NextStep = Literal["credit_review", "complete_documents", "fix_master_data", "contact_customer"]
SalesOrder = Field(pattern=r"^[0-9]{1,10}$", description="SAP sales order number, digits only", examples=["4711"])


# ---------- the contract: request and response shapes ----------

class BlockedOrder(BaseModel):
    sales_order: str
    customer: str
    reason: str
    net_value: float
    currency: str


class BlockedOrderPage(BaseModel):
    """One page of a list. next_cursor is null on the last page."""
    items: list[BlockedOrder]
    next_cursor: Optional[str] = Field(description="Pass this as ?cursor= to get the next page; null when done")


class ExplanationRequest(BaseModel):
    sales_order: str = SalesOrder
    max_words: int = Field(default=60, ge=20, le=120, description="Longest explanation you want, in words")


class Explanation(BaseModel):
    id: str
    sales_order: str
    explanation: str
    next_step: NextStep
    model: str
    created_at: str


class JobRequest(BaseModel):
    sales_orders: list[str] = Field(min_length=1, max_length=50, description="1 to 50 order numbers")


class Job(BaseModel):
    id: str
    status: Literal["queued", "running", "succeeded", "failed"]
    total: int
    done: int
    results: list[Explanation] = []
    errors: list[dict] = []


class Problem(BaseModel):
    """RFC 9457 problem details. Clients branch on type and status, never on the detail text."""
    type: str
    title: str
    status: int
    detail: str
    instance: str
    request_id: str


PROBLEM_RESPONSES = {   # documented in the OpenAPI spec for every operation that can fail this way
    401: {"model": Problem, "content": {"application/problem+json": {}}, "description": "Missing or wrong API key"},
    404: {"model": Problem, "content": {"application/problem+json": {}}, "description": "No such resource"},
    422: {"model": Problem, "content": {"application/problem+json": {}},
          "description": "Request does not match the contract, or an Idempotency-Key was reused with another body"},
}


class ApiProblem(Exception):
    def __init__(self, status: int, slug: str, title: str, detail: str, headers: Optional[dict] = None):
        self.status, self.slug, self.title, self.detail, self.headers = status, slug, title, detail, headers or {}


# ---------- the sample model ----------

class SampleModel:
    """Made-up answers so everything runs without an account. Swap in the real model from 'Building an AI API'."""
    name = "sample-model"

    def explain(self, sales_order: str, order: dict, max_words: int) -> dict:
        text = (f"Order {sales_order} for {order['customer']} is on hold because of: {order['reason'].lower()}. "
                f"It is worth {order['net_value']:,.2f} {order['currency']}. The responsible team must review it "
                "before it can ship.")
        words = text.split()[:max_words]
        return {"explanation": " ".join(words), "next_step": NEXT_STEP_BY_REASON.get(order["reason"], "contact_customer")}


# ---------- the app ----------

def create_app(api_key: str, model=None, job_seconds_per_order: float = 0.3) -> FastAPI:
    model = model or SampleModel()
    app = FastAPI(title="Orchestrate blocked-order explainer", version=API_VERSION,
                  description="Explains blocked SAP sales orders. Course example with made-up data. "
                              "Every error is application/problem+json (RFC 9457).")
    explanations: dict[str, dict] = {}          # explanation id -> explanation
    idempotency: dict[str, tuple[str, dict]] = {}   # Idempotency-Key -> (request fingerprint, stored result)
    jobs: dict[str, dict] = {}
    lock = threading.Lock()

    # --- one error format for everything ---
    def problem(request: Request, status: int, slug: str, title: str, detail: str, headers=None) -> JSONResponse:
        body = {"type": PROBLEM_BASE + slug, "title": title, "status": status, "detail": detail,
                "instance": request.url.path, "request_id": getattr(request.state, "request_id", "")}
        return JSONResponse(body, status_code=status, headers=headers, media_type="application/problem+json")

    @app.exception_handler(ApiProblem)
    async def handle_api_problem(request: Request, error: ApiProblem):
        return problem(request, error.status, error.slug, error.title, error.detail, error.headers)

    @app.exception_handler(RequestValidationError)
    async def handle_validation(request: Request, error: RequestValidationError):
        first = error.errors()[0]
        where = ".".join(str(part) for part in first["loc"])
        return problem(request, 422, "invalid-request", "Request does not match the contract", f"{where}: {first['msg']}")

    @app.exception_handler(StarletteHTTPException)
    async def handle_http(request: Request, error: StarletteHTTPException):   # e.g. unknown paths
        return problem(request, error.status_code, "http-error", "HTTP error", str(error.detail))

    @app.middleware("http")
    async def request_id(request: Request, call_next):
        incoming = request.headers.get("X-Request-ID", "")
        request.state.request_id = incoming if re.fullmatch(r"[A-Za-z0-9-]{8,64}", incoming) else uuid.uuid4().hex
        response = await call_next(request)
        response.headers["X-Request-ID"] = request.state.request_id
        return response

    key_header = APIKeyHeader(name="X-API-Key", auto_error=False, description="Your ORCHESTRATE_API_KEY")

    def require_key(key: Optional[str] = Depends(key_header)) -> None:
        if not key or not secrets.compare_digest(key.encode(), api_key.encode()):
            raise ApiProblem(401, "unauthorized", "Missing or wrong API key", "Send your key in the X-API-Key header.",
                             {"WWW-Authenticate": "APIKey"})

    def find_order(sales_order: str) -> dict:
        order = BLOCKED_ORDERS.get(sales_order)
        if order is None:
            raise ApiProblem(404, "order-not-found", "Blocked order not found",
                             f"There is no blocked order {sales_order} in the sample data.")
        return order

    def make_explanation(sales_order: str, max_words: int) -> dict:
        answer = model.explain(sales_order, find_order(sales_order), max_words)
        return {"id": "exp_" + uuid.uuid4().hex[:12], "sales_order": sales_order, **answer, "model": model.name,
                "created_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())}

    secured = [Depends(require_key)]

    @app.get("/health", tags=["service"], operation_id="getHealth", summary="Is the service up?")
    def health() -> dict:
        return {"status": "ok", "version": API_VERSION}

    # --- pagination: limit plus an opaque cursor ---
    @app.get("/v1/blocked-orders", response_model=BlockedOrderPage, dependencies=secured, tags=["blocked orders"],
             operation_id="listBlockedOrders", summary="List blocked orders, one page at a time",
             responses={401: PROBLEM_RESPONSES[401], 422: PROBLEM_RESPONSES[422]})
    def list_blocked_orders(limit: int = Query(default=3, ge=1, le=100, description="Items per page"),
                            cursor: Optional[str] = Query(default=None, description="next_cursor from the previous page"),
                            reason: Optional[str] = Query(default=None, description="Filter by block reason")):
        numbers = sorted(n for n, o in BLOCKED_ORDERS.items() if reason is None or o["reason"] == reason)
        after = ""
        if cursor:
            try:
                after = json.loads(base64.urlsafe_b64decode(cursor.encode()))["after"]
            except Exception:
                raise ApiProblem(422, "invalid-cursor", "Invalid cursor", "Use the next_cursor value exactly as returned.")
        remaining = [n for n in numbers if n > after]
        page = remaining[:limit]
        next_cursor = None
        if len(remaining) > limit:   # the cursor hides the position, so you can change how it works later
            next_cursor = base64.urlsafe_b64encode(json.dumps({"after": page[-1]}).encode()).decode()
        return {"items": [{"sales_order": n, **BLOCKED_ORDERS[n]} for n in page], "next_cursor": next_cursor}

    @app.get("/v1/blocked-orders/{sales_order}", response_model=BlockedOrder, dependencies=secured,
             tags=["blocked orders"], operation_id="getBlockedOrder", summary="Read one blocked order",
             responses={401: PROBLEM_RESPONSES[401], 404: PROBLEM_RESPONSES[404], 422: PROBLEM_RESPONSES[422]})
    def get_blocked_order(sales_order: str):
        return {"sales_order": sales_order, **find_order(sales_order)}

    # --- create a resource, safely retryable with Idempotency-Key ---
    @app.post("/v1/explanations", response_model=Explanation, status_code=201, dependencies=secured,
              tags=["explanations"], operation_id="createExplanation", summary="Explain one blocked order",
              responses={**PROBLEM_RESPONSES, 409: {"model": Problem, "content": {"application/problem+json": {}},
                                                     "description": "A request with this Idempotency-Key is still running"}})
    def create_explanation(body: ExplanationRequest, request: Request, response: Response,
                           idempotency_key: Optional[str] = Header(default=None, alias="Idempotency-Key",
                                                                   description="A new random value (UUID) per action. "
                                                                   "Retries with the same key and body return the first result.",
                                                                   pattern=r"^[A-Za-z0-9-]{8,64}$")):
        fingerprint = hashlib.sha256(body.model_dump_json().encode()).hexdigest()
        if idempotency_key:
            with lock:   # check and reserve the key in one step, so two parallel retries can't both run
                stored = idempotency.get(idempotency_key)
                if stored is None:
                    idempotency[idempotency_key] = (fingerprint, None)   # None: still in progress
            if stored:
                if stored[0] != fingerprint:
                    raise ApiProblem(422, "idempotency-key-reused", "Idempotency key already used",
                                     "This Idempotency-Key was used with a different request body. Use a new key.")
                if stored[1] is None:
                    raise ApiProblem(409, "request-in-progress", "Request still in progress",
                                     "A request with this Idempotency-Key is still running. Retry shortly.",
                                     {"Retry-After": "1"})
                response.headers["Location"] = f"/v1/explanations/{stored[1]['id']}"
                response.headers["Idempotent-Replay"] = "true"   # our own header: this is the stored result
                return stored[1]
        try:
            result = make_explanation(body.sales_order, body.max_words)
        except ApiProblem:
            if idempotency_key:
                with lock:
                    idempotency.pop(idempotency_key, None)   # a failed attempt may be retried with the same key
            raise
        with lock:
            explanations[result["id"]] = result
            if idempotency_key:
                idempotency[idempotency_key] = (fingerprint, result)
        response.headers["Location"] = f"/v1/explanations/{result['id']}"
        return result

    @app.get("/v1/explanations/{explanation_id}", response_model=Explanation, dependencies=secured,
             tags=["explanations"], operation_id="getExplanation", summary="Read an explanation again",
             responses={401: PROBLEM_RESPONSES[401], 404: PROBLEM_RESPONSES[404], 422: PROBLEM_RESPONSES[422]})
    def get_explanation(explanation_id: str):
        if explanation_id not in explanations:
            raise ApiProblem(404, "explanation-not-found", "Explanation not found", f"No explanation {explanation_id}.")
        return explanations[explanation_id]

    # --- streaming: server-sent events, then one final validated object ---
    def prepare_explanation(body: ExplanationRequest) -> dict:
        """Runs before the stream starts, so a missing order is still a normal 404 problem."""
        result = make_explanation(body.sales_order, body.max_words)
        explanations[result["id"]] = result
        return result

    @app.post("/v1/explanations/stream", response_class=EventSourceResponse, dependencies=secured,
              tags=["explanations"], operation_id="streamExplanation",
              summary="Explain one order, sending words as they are ready",
              responses={401: PROBLEM_RESPONSES[401], 404: PROBLEM_RESPONSES[404], 422: PROBLEM_RESPONSES[422]})
    def stream_explanation(result: dict = Depends(prepare_explanation)) -> Iterable[ServerSentEvent]:
        for number, word in enumerate(result["explanation"].split()):
            time.sleep(0.05)   # pretend the model writes word by word
            yield ServerSentEvent(event="delta", id=str(number), data={"text": word + " "})
        yield ServerSentEvent(event="done", data=result)   # the full, schema-checked object comes last

    # --- long-running work: 202 Accepted, Location, poll the job ---
    def run_job(job_id: str) -> None:
        job = jobs[job_id]
        job["status"] = "running"
        for number in job["sales_orders"]:
            time.sleep(job_seconds_per_order)   # pretend each explanation takes a while
            try:
                job["results"].append(make_explanation(number, 60))
            except ApiProblem as error:
                job["errors"].append({"sales_order": number, "type": PROBLEM_BASE + error.slug, "detail": error.detail})
            job["done"] += 1
        job["status"] = "succeeded" if job["results"] else "failed"

    @app.post("/v1/explanation-jobs", response_model=Job, status_code=202, dependencies=secured,
              tags=["jobs"], operation_id="createExplanationJob", summary="Explain many orders in the background",
              responses={401: PROBLEM_RESPONSES[401], 422: PROBLEM_RESPONSES[422]})
    def create_job(body: JobRequest, response: Response):
        job_id = "job_" + uuid.uuid4().hex[:12]
        jobs[job_id] = {"id": job_id, "status": "queued", "total": len(body.sales_orders), "done": 0,
                        "results": [], "errors": [], "sales_orders": body.sales_orders}
        accepted = dict(jobs[job_id])   # what the caller sees now: queued, nothing done yet
        threading.Thread(target=run_job, args=(job_id,), daemon=True).start()
        response.headers["Location"] = f"/v1/explanation-jobs/{job_id}"
        response.headers["Retry-After"] = "1"   # seconds the caller should wait before the first poll
        return accepted

    @app.get("/v1/explanation-jobs/{job_id}", response_model=Job, dependencies=secured, tags=["jobs"],
             operation_id="getExplanationJob", summary="Check a background job",
             responses={401: PROBLEM_RESPONSES[401], 404: PROBLEM_RESPONSES[404], 422: PROBLEM_RESPONSES[422]})
    def get_job(job_id: str, response: Response):
        if job_id not in jobs:
            raise ApiProblem(404, "job-not-found", "Job not found", f"No job {job_id}.")
        if jobs[job_id]["status"] in ("queued", "running"):
            response.headers["Retry-After"] = "1"
        return jobs[job_id]

    def openapi_with_problems() -> dict:
        """FastAPI files error schemas under the success media type; move them to application/problem+json."""
        if app.openapi_schema is None:
            from fastapi.openapi.utils import get_openapi
            spec = get_openapi(title=app.title, version=app.version, description=app.description, routes=app.routes)
            for path in spec["paths"].values():
                for operation in path.values():
                    for code, answer in operation.get("responses", {}).items():
                        if "application/problem+json" in answer.get("content", {}):
                            answer["content"] = {"application/problem+json": {
                                "schema": {"$ref": "#/components/schemas/Problem"}}}
            app.openapi_schema = spec
        return app.openapi_schema

    app.openapi = openapi_with_problems
    return app


# ---------- demo: every rule, in memory, no server ----------

def demo(api_key: str) -> None:
    from fastapi.testclient import TestClient
    client = TestClient(create_app(api_key, job_seconds_per_order=0.2))
    key = {"X-API-Key": api_key}

    def show(title: str, response, fields=None) -> None:
        body = response.json() if "json" in response.headers.get("content-type", "") else response.text
        if fields and isinstance(body, dict):
            body = {f: body.get(f) for f in fields}
        print(f"\n== {title}\n{response.status_code} {response.headers.get('content-type', '')}")
        for header in ("Location", "Retry-After", "Idempotent-Replay"):
            if header in response.headers:
                print(f"{header}: {response.headers[header]}")
        print(json.dumps(body, indent=2)[:600])

    page1 = client.get("/v1/blocked-orders?limit=3", headers=key)
    show("1. First page of blocked orders", page1, ["next_cursor"])
    print("orders:", [o["sales_order"] for o in page1.json()["items"]])
    page2 = client.get("/v1/blocked-orders", params={"limit": 3, "cursor": page1.json()["next_cursor"]}, headers=key)
    print("next page orders:", [o["sales_order"] for o in page2.json()["items"]])

    show("2. An error, in one format for every error (RFC 9457)", client.get("/v1/blocked-orders/9999", headers=key))

    idem = {**key, "Idempotency-Key": str(uuid.uuid4())}
    first = client.post("/v1/explanations", json={"sales_order": "4711"}, headers=idem)
    show("3. Create an explanation (201 Created)", first, ["id", "next_step", "model"])
    again = client.post("/v1/explanations", json={"sales_order": "4711"}, headers=idem)
    show("4. Same request retried with the same Idempotency-Key", again, ["id"])
    print("same id as before:", again.json()["id"] == first.json()["id"])
    show("5. Same key, different body", client.post("/v1/explanations", json={"sales_order": "4712"}, headers=idem),
         ["status", "title"])

    created = client.post("/v1/explanation-jobs", json={"sales_orders": ["4711", "4712", "9999"]}, headers=key)
    show("6. Start a background job (202 Accepted)", created, ["id", "status", "total", "done"])
    location = created.headers["Location"]
    while True:
        time.sleep(float(created.headers.get("Retry-After", "1")) / 2)
        status = client.get(location, headers=key).json()
        print(f"   poll: {status['status']} {status['done']}/{status['total']}")
        if status["status"] in ("succeeded", "failed"):
            break
    print(f"   results: {len(status['results'])}, errors: {[e['sales_order'] for e in status['errors']]}")

    print("\n== 7. Stream an explanation (server-sent events)")
    with client.stream("POST", "/v1/explanations/stream", json={"sales_order": "4714"},
                       headers=key) as stream:
        lines = [line for line in stream.iter_lines() if line.startswith(("event:", "data:"))]
    print("\n".join(lines[:4]) + "\n...\n" + "\n".join(lines[-2:])[:300])

    show("8. No key", client.get("/v1/blocked-orders"), ["status", "title"])
    print("\nDemo finished. Next: export the contract with --export-openapi unit06/openapi.json")


def main() -> None:
    parser = argparse.ArgumentParser(description="The Unit 6 blocked-orders API, designed with REST rules.")
    parser.add_argument("--demo", action="store_true", help="walk through every rule in memory, no server")
    parser.add_argument("--export-openapi", metavar="FILE", help="write the OpenAPI spec to FILE and stop")
    parser.add_argument("--port", type=int, default=8000, help="port to listen on (default 8000)")
    args = parser.parse_args()

    from dotenv import load_dotenv
    load_dotenv()
    api_key = os.environ.get("ORCHESTRATE_API_KEY", "")
    if args.export_openapi:
        spec = create_app(api_key or "not-needed-for-export").openapi()
        with open(args.export_openapi, "w", encoding="utf-8") as file:
            json.dump(spec, file, indent=2)
        print(f"Wrote {args.export_openapi}: OpenAPI {spec['openapi']}, {len(spec['paths'])} paths, "
              f"API version {spec['info']['version']}")
        return
    if len(api_key) < 20:
        sys.exit("ORCHESTRATE_API_KEY is missing or shorter than 20 characters in .env. "
                 "See 'Building an AI API', Step 2.")
    if args.demo:
        demo(api_key)
        return
    import uvicorn
    print(f"Open http://127.0.0.1:{args.port}/docs in your browser. Press Ctrl+C to stop.")
    uvicorn.run(create_app(api_key), host="127.0.0.1", port=args.port)


if __name__ == "__main__":
    main()

Step 3: Run the demo (no server)

The demo sends requests to the API in memory, with FastAPI's TestClient, and prints what comes back for each design rule.

  1. Run it:

    python unit06/designed_api.py --demo

What success looks like (shortened; your IDs differ, and the poll lines may show other counts):

== 1. First page of blocked orders
200 application/json
{
  "next_cursor": "eyJhZnRlciI6ICI0NzEzIn0="
}
orders: ['4711', '4712', '4713']
next page orders: ['4714', '4715', '4716']

== 2. An error, in one format for every error (RFC 9457)
404 application/problem+json
{
  "type": "https://example.com/problems/order-not-found",
  "title": "Blocked order not found",
  "status": 404,
  "detail": "There is no blocked order 9999 in the sample data.",
  "instance": "/v1/blocked-orders/9999",
  "request_id": "bef0da40d5654c2882032fa75985e76b"
}

== 3. Create an explanation (201 Created)
201 application/json
Location: /v1/explanations/exp_72f0c532e21c
...
== 4. Same request retried with the same Idempotency-Key
201 application/json
Location: /v1/explanations/exp_72f0c532e21c
Idempotent-Replay: true
...
same id as before: True

== 5. Same key, different body
422 application/problem+json
...
== 6. Start a background job (202 Accepted)
202 application/json
Location: /v1/explanation-jobs/job_31e3163166d3
Retry-After: 1
...
   poll: running 2/3
   poll: succeeded 3/3
   results: 2, errors: ['9999']

== 7. Stream an explanation (server-sent events)
event: delta
data: {"text": "Order "}
...
event: done
data: {"id": "exp_c056d1f1c04f", "sales_order": "4714", ...}

== 8. No key
401 application/problem+json
...
Demo finished. Next: export the contract with --export-openapi unit06/openapi.json
  1. Read it against the rules. Section 4 is the important one: the retry got the same explanation ID, so the model ran once. In section 6, the job finished with two results and one per-item error, without failing the batch.

Step 4: Export the contract

  1. Write the OpenAPI file:

    python unit06/designed_api.py --export-openapi unit06/openapi.json

What success looks like:

Wrote unit06/openapi.json: OpenAPI 3.1.0, 8 paths, API version 1.1.0
  1. Open unit06/openapi.json in VS Code. Find "/v1/explanations", then its "post" section. You will see the Idempotency-Key header parameter, the 201 response and the error responses, each with application/problem+json. Find "/v1/explanations/stream": its 200 response uses text/event-stream with an itemSchema.

Step 5: Check the contract with a linter

  1. In unit06, create lint_api.py, paste the code below and save. It uses only modules built into Python.
"""Unit 6: check an OpenAPI spec against the course's API design rules.

Built-in modules only. It reads the JSON spec that designed_api.py exports and prints one line per rule:
OK when every operation follows it, FAIL with the operations that don't.

How to run (from your course folder, with .venv turned on):
    python unit06/lint_api.py unit06/openapi.json
"""
import json
import re
import sys

METHODS = {"get", "put", "post", "patch", "delete"}
UNVERSIONED = {"/health"}   # paths that may live outside /v1


def operations(spec: dict):
    for path, item in spec.get("paths", {}).items():
        for method, operation in item.items():
            if method in METHODS:
                yield path, method, operation


def rule_versioned_paths(spec):
    """Every business path starts with a major version, such as /v1/."""
    return [p for p in spec["paths"] if p not in UNVERSIONED and not re.match(r"^/v[0-9]+/", p)]


def rule_resource_names(spec):
    """Path segments are lowercase nouns with hyphens; no verbs such as /getOrders."""
    bad = []
    for path in spec["paths"]:
        for segment in path.strip("/").split("/"):
            if segment.startswith("{"):
                continue
            if not re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", segment) or re.match(r"^(get|create|update|delete|do)", segment):
                bad.append(f"{path} ({segment})")
    return bad


def rule_operation_ids(spec):
    """Every operation has a unique operationId and a summary, so generated clients get good names."""
    seen, bad = set(), []
    for path, method, op in operations(spec):
        op_id = op.get("operationId")
        if not op_id or op_id in seen or not op.get("summary"):
            bad.append(f"{method.upper()} {path}")
        seen.add(op_id)
    return bad


def rule_problem_errors(spec):
    """Every documented 4xx or 5xx response uses application/problem+json."""
    bad = []
    for path, method, op in operations(spec):
        for code, response in op.get("responses", {}).items():
            if code[0] in "45" and list(response.get("content", {})) != ["application/problem+json"]:
                bad.append(f"{method.upper()} {path} {code}")
    return bad


def rule_secured_operations_document_401(spec):
    """If the API has a security scheme, every secured operation documents 401."""
    bad = []
    for path, method, op in operations(spec):
        if op.get("security") and "401" not in op.get("responses", {}):
            bad.append(f"{method.upper()} {path}")
    return bad


def rule_lists_are_paginated(spec):
    """Every GET on a collection (no {id} at the end) takes limit and cursor."""
    bad = []
    for path, method, op in operations(spec):
        if method == "get" and path not in UNVERSIONED and not path.endswith("}"):
            names = {p.get("name") for p in op.get("parameters", [])}
            if not {"limit", "cursor"} <= names:
                bad.append(f"GET {path}")
    return bad


def rule_creates_are_retry_safe(spec):
    """Every POST that answers 201 Created accepts an Idempotency-Key header."""
    bad = []
    for path, method, op in operations(spec):
        if method == "post" and "201" in op.get("responses", {}):
            names = {p.get("name") for p in op.get("parameters", []) if p.get("in") == "header"}
            if "Idempotency-Key" not in names:
                bad.append(f"POST {path}")
    return bad


RULES = [rule_versioned_paths, rule_resource_names, rule_operation_ids, rule_problem_errors,
         rule_secured_operations_document_401, rule_lists_are_paginated, rule_creates_are_retry_safe]


def main() -> None:
    if len(sys.argv) != 2:
        sys.exit("Usage: python unit06/lint_api.py unit06/openapi.json")
    try:
        with open(sys.argv[1], encoding="utf-8") as file:
            spec = json.load(file)
    except FileNotFoundError:
        sys.exit(f"{sys.argv[1]} not found. Export it first: python unit06/designed_api.py --export-openapi {sys.argv[1]}")
    print(f"Checking {spec['info']['title']} {spec['info']['version']} (OpenAPI {spec['openapi']}), "
          f"{len(list(operations(spec)))} operations")
    failed = 0
    for rule in RULES:
        problems = rule(spec)
        label = rule.__doc__.strip().split("\n")[0]
        print(f"{'OK  ' if not problems else 'FAIL'}  {label}")
        for item in problems:
            print(f"        {item}")
        failed += bool(problems)
    print("All rules pass." if not failed else f"{failed} rule(s) failed.")
    sys.exit(1 if failed else 0)


if __name__ == "__main__":
    main()
  1. Run it:

    python unit06/lint_api.py unit06/openapi.json

What success looks like:

Checking Orchestrate blocked-order explainer 1.1.0 (OpenAPI 3.1.0), 8 operations
OK    Every business path starts with a major version, such as /v1/.
OK    Path segments are lowercase nouns with hyphens; no verbs such as /getOrders.
OK    Every operation has a unique operationId and a summary, so generated clients get good names.
OK    Every documented 4xx or 5xx response uses application/problem+json.
OK    If the API has a security scheme, every secured operation documents 401.
OK    Every GET on a collection (no {id} at the end) takes limit and cursor.
OK    Every POST that answers 201 Created accepts an Idempotency-Key header.
All rules pass.
  1. See a rule fail on purpose. In designed_api.py, find the get_blocked_order decorator and delete , 422: PROBLEM_RESPONSES[422] from its responses=. Save, export again (Step 4), and lint again. You should see:
FAIL  Every documented 4xx or 5xx response uses application/problem+json.
        GET /v1/blocked-orders/{sales_order} 422

FastAPI added its own 422 response in its default format, and the linter caught it. Put the text back, export and lint until all rules pass.

Step 6: Start the server and call it

  1. Start the API:

    python unit06/designed_api.py

What success looks like:

Open http://127.0.0.1:8000/docs in your browser. Press Ctrl+C to stop.
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
  1. Open http://127.0.0.1:8000/docs. Click Authorize, paste your key from .env (without quotes) into the APIKeyHeader box, click Authorize, then Close. Try GET /v1/blocked-orders with Try it out and Execute, then copy next_cursor into the cursor box and execute again.

  2. Open a second terminal (Terminal > New Terminal), turn on .venv there (Step 1), and load your key:

    • Windows (PowerShell):

      $key = (Select-String -Path .env -Pattern '^ORCHESTRATE_API_KEY="(.*)"').Matches[0].Groups[1].Value
    • macOS / Linux:

      key=$(grep '^ORCHESTRATE_API_KEY=' .env | cut -d'"' -f2)
  3. Watch a stream arrive word by word. -N tells curl not to wait for the whole answer. On Windows, type curl.exe (not curl), which recent Windows versions include; PowerShell's own curl name means something else.

    • Windows (PowerShell):

      PowerShell versions treat quotes inside arguments differently, so put the body in a small file first and pass it with @:

      Set-Content -Path body.json -Value '{"sales_order": "4712"}'
      curl.exe -N -X POST http://127.0.0.1:8000/v1/explanations/stream -H "Content-Type: application/json" -H "X-API-Key: $key" -d "@body.json"
    • macOS / Linux:

      curl -N -X POST http://127.0.0.1:8000/v1/explanations/stream -H "Content-Type: application/json" -H "X-API-Key: $key" -d '{"sales_order": "4712"}'

What success looks like (events appear one by one):

event: delta
data: {"text": "Order "}
id: 0

event: delta
data: {"text": "4712 "}
id: 1
...
event: done
data: {"id": "exp_...", "sales_order": "4712", ...}
  1. Start a job and look at the headers (-i prints them):

    • Windows (PowerShell):

      Set-Content -Path body.json -Value '{"sales_orders": ["4711", "4712"]}'
      curl.exe -i -X POST http://127.0.0.1:8000/v1/explanation-jobs -H "Content-Type: application/json" -H "X-API-Key: $key" -d "@body.json"
    • macOS / Linux:

      curl -i -X POST http://127.0.0.1:8000/v1/explanation-jobs -H "Content-Type: application/json" -H "X-API-Key: $key" -d '{"sales_orders": ["4711", "4712"]}'

What success looks like (shortened):

HTTP/1.1 202 Accepted
location: /v1/explanation-jobs/job_62379b22b817
retry-after: 1
...
{"id":"job_62379b22b817","status":"queued","total":2,"done":0,"results":[],"errors":[]}
  1. Poll the job with the location from your output (replace the ID):

    curl -H "X-API-Key: $key" http://127.0.0.1:8000/v1/explanation-jobs/job_62379b22b817

    On Windows use curl.exe. After a second, status is succeeded with two results.

  2. Press Ctrl+C in the first terminal to stop the server.

  3. Save your work with Git:

    git add unit06/designed_api.py unit06/lint_api.py unit06/openapi.json
    git commit -m "Unit 6: REST design for the blocked-orders API, with OpenAPI contract and linter"

What each part of the scripts does:

Part What it does
BlockedOrder, ExplanationRequest, Explanation, Job The request and response shapes; they become schemas in the OpenAPI file
Problem, ApiProblem, problem() The RFC 9457 error shape, raised anywhere and turned into application/problem+json
handle_validation, handle_http Replace FastAPI's default error bodies so every error has the same shape
request_id middleware Adds X-Request-ID to every response; the same ID appears in errors
list_blocked_orders Cursor pagination: limit, an opaque cursor, and next_cursor (null on the last page)
create_explanation 201 Created with Location; an Idempotency-Key returns the stored result, 422 on a different body, 409 while still running
prepare_explanation and stream_explanation Checks run before streaming; then delta events and one final done event with the full object
create_job, run_job, get_job 202 Accepted with Location and Retry-After; a background thread works through the orders; per-item errors
openapi_with_problems Files every error response under application/problem+json in the exported spec
demo() Calls every endpoint in memory and prints the results; no server needed
lint_api.py rules Seven checks on the exported spec: versions, names, operation IDs, error format, 401, pagination, idempotency

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on PATH, or .venv is off Turn on .venv (Step 1); see Set up your computer if Python itself is missing
ModuleNotFoundError: No module named 'fastapi' (or httpx, dotenv) The library isn't in this .venv Check for (.venv) in the prompt, then pip install -r requirements.txt
ModuleNotFoundError: No module named 'fastapi.sse' Your FastAPI is older than 0.135.0 pip install --upgrade "fastapi[standard]"
ORCHESTRATE_API_KEY is missing or shorter than 20 characters No key in .env, or you ran the script from another folder Do Step 2 of Building an AI API; run from the course folder
401 with "title": "Missing or wrong API key" The header is missing or the key differs from .env Load the key again (Step 6.3); restart the server after editing .env
unit06/openapi.json not found You haven't exported the spec yet Run Step 4 first
The stream arrives all at once curl is buffering, or a proxy holds the response Add -N; on a company network, run against 127.0.0.1 only
422 with body: ... JSON decode error on Windows PowerShell changed the quotes in the body Use the body.json file and -d "@body.json" exactly as shown
address already in use Port 8000 is taken, often by an earlier API Stop it with Ctrl+C, or add --port 8080 and use that port
curl fails with a proxy or certificate error A company proxy intercepts local traffic Use the /docs page in the browser instead, or ask IT to exclude 127.0.0.1 from the proxy

The SAP way

Your laptop API follows the same patterns SAP's own APIs use. As of October 2026, here is where each rule shows up on the SAP side.

Contracts on the Business Accelerator Hub. The SAP Business Accelerator Hub (api.sap.com) publishes SAP's API specifications. Log in, open a service, click API Specification, and download EDMX for OData services or JSON or YAML for OpenAPI services. The SAP Cloud SDK generates typed clients from these files. When you build an AI extension, publish your own OpenAPI file the same way: in a known place, versioned, downloadable.

OData's answers to the same questions. Many SAP business APIs use OData. The OData 4.01 protocol covers most of this topic in its own vocabulary:

Design question REST pattern in this topic OData equivalent
Pages limit + opaque cursor $top and $skip; server-driven paging with a next link and opaque $skiptoken; the odata.maxpagesize preference
Long-running work 202 + Location + job resource Prefer: respond-async, 202 Accepted, Location to a status monitor
Lost updates (not needed here: explanations are never edited) If-Match with an ETag; a stale ETag gets 412 Precondition Failed
Errors RFC 9457 problem details OData's own JSON error format

Two practical consequences. First, if your AI service calls SAP OData APIs, use their paging and ETags rather than pulling everything at once. Second, if your service offers OData (for example from CAP, see CAP and side-by-side extensions), follow OData's conventions, not the REST ones above. Don't mix the two styles in one API.

CSRF tokens on writes. The SAP Cloud SDK expects SAP systems to ask for a CSRF token on requests that change data: by default it fetches a token before non-GET requests (with a HEAD request carrying X-CSRF-Token: fetch) and sends it along, for HTTP, OData and OpenAPI clients alike. You can switch this off only for systems that don't require a token. If you call SAP without the SDK, you must handle the token yourself. A read-only AI service never needs it; one that writes back (with human approval, covered in Unit 9) does.

Streaming in the SAP Cloud SDK for AI. The SDK's stream() method on the orchestration client returns chunks based on the server-sent events standard. It supports cancelling with an AbortController (JavaScript), and gives the finish reason and token usage after the stream ends. Your API can relay those chunks as its own delta events, then send the validated done object, as the script does with the sample model.

Long-running resources in SAP AI Core. A model deployment in SAP AI Core is a resource with a lifecycle, not a single call. Through the SDK you create it from a configuration, read its status (such as RUNNING, STOPPED or UNKNOWN), change it by setting a targetStatus such as STOPPED, and delete it only when it is STOPPED or UNKNOWN. That is the job-resource pattern from this topic: ask for a change, then check back.

Gateways. SAP API Management, part of SAP Integration Suite, adds keys, OAuth, rate limits and a developer portal in front of APIs; Building an AI API covered it. A gateway enforces policy on a contract. It doesn't design the contract for you.

Build vs. SAP

Situation Build your own (this topic) Use SAP's
A new AI endpoint for one app or workflow FastAPI with these rules and an exported OpenAPI file CAP if the service is mostly business data with some AI
Your service exposes business entities that Fiori apps will read Not ideal; you would rebuild OData features CAP with OData, so SAP UIs and tools work with it
Reading or writing S/4HANA data Raw HTTP plus your own CSRF and ETag handling SAP Cloud SDK clients generated from Business Accelerator Hub specs
Streaming model answers Your own SSE endpoint relaying chunks SAP Cloud SDK for AI stream() as the source
Keys, quotas and partner access at scale In-process limits, fine for one instance SAP API Management in front of the service
Batch explanations over hundreds of orders A job resource as built here, with a real queue in production An SAP-provided batch feature if your scenario has one; check current docs

Production concerns

Security and authorizations. Every endpoint, including job status and stored explanations, must check that the caller may see that object. A job ID or explanation ID is not a secret; guessable or leaked IDs must still be refused for the wrong user (broken object-level authorization, covered in Building an AI API). On BTP, replace the API key with XSUAA tokens (set up in Deploying AI apps on SAP BTP), and scope idempotency keys per caller, so one caller can't replay another's result by guessing a key.

Data in stored results. Idempotency stores and job results hold model output about customers. Give them an expiry, keep them out of logs, and store them where your data-protection rules allow. The script keeps them in memory, which disappears on restart; production needs a shared store such as a database or cache, or two instances behind a load balancer will disagree.

Jobs that survive restarts. A thread in the web process dies with the process. Production jobs belong in a queue with workers, a persisted status, and a retry policy per item. Cap the batch size (the script allows 1 to 50) and per-caller concurrent jobs.

Streams and proxies. Proxies and gateways sometimes buffer responses, which turns a stream into one late blob. FastAPI's SSE response sets headers that ask proxies not to buffer, and sends keep-alive pings. Test streaming through your real gateway, not just locally.

Cost. Idempotency is also cost control: a retry storm without it multiplies model spend. Log the token usage per request ID and per job.

Evaluation and contract tests. The contract is testable without a model: shapes, status codes, headers. Run the linter and contract tests in CI on every pull request. The next topic, Testing AI applications (Unit 6), builds those tests.

Clean core. The AI service runs side by side on BTP and talks to S/4HANA only through released APIs. Its own versioned contract means S/4HANA upgrades and model swaps don't ripple into callers.

Pitfalls

  • Verbs in paths (/explainOrder, /getJobStatus). They multiply and hide the resource model. Use nouns plus methods.
  • 200 with an error inside. {"status": "error"} with HTTP 200 defeats every client library, gateway and monitor. Use the status code.
  • Parsing detail. If callers branch on error text, every wording fix becomes a breaking change. Give them type.
  • Offset pages over changing data. Blocked orders get released while someone pages through them, so rows silently disappear. Use a cursor.
  • Idempotency keys reused per order. Generating the key from the order number blocks a legitimate second request tomorrow. One random key per action.
  • Erroring mid-stream. After the first event, the client already has a 200. Validate first, then stream; if the model fails mid-stream, send an error event and document it.
  • Treating a model swap as a new API version. Model and prompt versions go in response fields. /v2 is for breaking changes to the shape.
  • A spec nobody checks. An exported OpenAPI file that isn't linted and diffed in pull requests drifts from the code, and then nobody trusts it.

Exercise: add a paginated list of explanations

Add GET /v1/explanations, which lists the explanations created so far, newest first, with the same pagination as blocked orders. Then prove the contract still follows every rule.

  1. Open unit06/designed_api.py. Below the get_explanation function, add a new endpoint with @app.get("/v1/explanations", ...). Copy the decorator style from list_blocked_orders: dependencies=secured, a tags entry, operation_id="listExplanations", a summary, and responses with 401 and 422 from PROBLEM_RESPONSES.
  2. Give it limit and cursor query parameters, exactly like list_blocked_orders.
  3. Return {"items": [...], "next_cursor": ...}. Sort explanations by created_at, newest first, and use the explanation id in the cursor instead of the order number. Add a small ExplanationPage model (copy BlockedOrderPage, with items: list[Explanation]) and use it as response_model.
  4. Add a line to demo() that creates two explanations and then lists them with limit=1, printing both pages.
  5. Run python unit06/designed_api.py --demo and check the new lines.
  6. Export the spec (Step 4) and run python unit06/lint_api.py unit06/openapi.json.
  7. Commit the three files with Git.

Done when the linter prints 9 operations and All rules pass., the demo shows two pages of one explanation each, and openapi.json in Git contains listExplanations. The next topic in Unit 6 writes automated tests against this contract.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the redesigned API make an explanation a resource with an ID, instead of returning text and forgetting it?

    Answer: B. A stored explanation has an address, so it can be fetched again, linked from a ticket and audited. The idempotency store also returns the same resource on a retry. Model cost and streaming don't depend on it.
  2. 2A caller retries POST /v1/explanations with the same Idempotency-Key but a different sales_order. What does the script return, and why?

    Answer: C. The script stores a fingerprint of the body with each key. A different fingerprint means the key was misused, which the IETF draft answers with 422. 409 is for a retry that arrives while the first request with that key is still running.
  3. 3In the streaming endpoint, why does the order lookup run in a FastAPI dependency before the first event?

    Answer: D. Once streaming starts, the HTTP status has been sent. Checking the key, input and order before the stream lets a missing order return a normal 404 problem response instead of a broken stream.
  4. 4Which response tells a caller that a batch was accepted but isn't finished, and where to check?

    Answer: B. RFC 9110 defines 202 as accepted for processing but not completed. The Location header points at the job the caller polls, and Retry-After says how long to wait first.
  5. 5Your team swaps the model behind the explainer for a newer one. The response fields stay the same. What should change?

    Answer: C. A model swap changes wording, not the contract's shape, so it isn't a breaking change. Report the model in a response field so callers and auditors can see which one answered; reserve /v2 for breaking changes to fields or behavior.
  6. 6Why does the script use an opaque cursor rather than ?offset= for blocked orders?

    Answer: B. Blocked orders change between calls, so a numeric offset can skip or repeat rows. A cursor continues after the last item seen, and being opaque, it can change how it works without a new version. OData offers both $skip and server-driven paging.
  7. 7Your AI service will write a note back to an S/4HANA sales order through an OData API, without the SAP Cloud SDK. What must your code handle that a read-only call doesn't?

    Answer: D. The Cloud SDK fetches a CSRF token for non-GET requests by default because SAP systems may require one, so without the SDK you handle it yourself. OData uses If-Match with the ETag to prevent lost updates, answering 412 when it is stale. Writes also need human approval, covered in Unit 9.
  8. 8What does lint_api.py catch when you delete the 422 entry from get_blocked_order?

    Answer: C. Without your own 422 entry, FastAPI documents its default validation error under application/json. The linter's error-format rule flags it, which is the point of checking the exported contract, not just the code.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in