Orchestrate

Testing AI applications: unit and integration tests

Prove an AI service keeps its promises: pytest, fakes and mocks for SAP and the model, contract tests, sandbox integration tests and CI.

Updated Oct 3, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

A test is a small program that checks another program does what it promises. You write it once, and a computer runs it on every change, in seconds, for free.

AI applications need tests more than ordinary software, not less. They have two parts that behave very differently. The plumbing (who may call, what they may send, what happens when the model fails, what gets logged) should behave exactly the same way every time. The model gives a slightly different answer every time, and sometimes a wrong one.

So you test them differently. The plumbing gets exact tests: this input must give that output. The model's answers get rule checks: whatever the wording, the answer must name the order, pick an allowed next step and never claim something only a person may decide. Whether the answers are actually good is a separate job, called evaluation, covered in Unit 8.

The trick that makes this cheap is the test double: a stand-in for the slow, costly or unreliable parts, such as the model and the SAP system. Most tests use the stand-ins. A few tests talk to the real thing, to prove the stand-ins are still honest.

Why it matters to the business

Take the course's running example from earlier in this unit: an AI service that explains blocked sales orders to sales reps and suggests who should act. It sits inside order-to-cash. When it breaks, reps lose time and orders sit longer.

Tests protect that in four ways:

  • Changes stop being scary. Teams change prompts, models and libraries often. Without tests, every change is a manual re-check that nobody has time for, so changes slow down or go out unchecked. With tests, the check takes seconds.
  • Failures are caught before users see them. A test that makes the model "time out" on purpose proves the rep still gets a safe answer instead of an error screen.
  • Promises are enforced. "We never log customer names" and "the AI never says an order is released" become tests that fail the moment someone breaks them. That is evidence you can show an auditor or a works council.
  • Cost stays near zero. The bulk of the tests use a fake model and a fake SAP system. They make no paid model calls and need no SAP access, so they can run on every change.

The cost of skipping tests shows up later, in a production incident, a rollback, or a project that nobody dares to touch.

How SAP does it

SAP's development tools assume you test, and give you the pieces. As of October 2026:

  • CAP has a test helper. The SAP Cloud Application Programming Model (CAP) offers cds.test for Node.js. It starts your CAP service inside the test run, on a free port, with an in-memory database. Tests send requests to it and check the answers, and they can act as mock users to test authorizations. It works with common JavaScript test runners.
  • CAP can stand in for SAP back ends. When a CAP project imports an SAP API, CAP can mock it locally with sample data, as shown in CAP and side-by-side extensions for AI. That is a test double provided by the framework.
  • SAP runs pipelines for you. SAP Continuous Integration and Delivery is a BTP service that SAP describes as pipeline-as-a-service, with templates for SAP-specific projects. A change in your Git repository triggers a pipeline that runs the automated tests and reports back.
  • Quality of AI answers is a separate step. Measuring how good the answers are, across many examples, is evaluation. Unit 8 covers it, including SAP's tooling for it; this topic covers the tests that come first.

The ideas in this topic are not SAP-specific. A Python service on BTP, a CAP app and a Java app are all tested the same way: many fast tests with stand-ins, a few slow tests against the real thing.

The layers of testing at a glance

Each kind of test answers a different question. You need all of them, in very different amounts.

Test Question it answers Uses real model? Uses real SAP? Speed and cost When it runs
Unit test Does this one rule work? ("credit block means credit review") No No Milliseconds, free Every change
API test Does the service keep its promises? (keys, errors, fallback, logs) No, a fake No, a fake Milliseconds, free Every change
Contract test Would a change break the apps that call us, or have our fakes drifted from the real systems? No Partly Fast to daily Every change, plus daily
Integration test Does it really work against the SAP sandbox or test system? No Yes Seconds, needs access Daily or before release
Model output check Do real model answers follow the rules, whatever the wording? Yes No Seconds, small charge Before release, after model or prompt changes
Evaluation (Unit 8) Are the answers actually good and useful? Yes Often Minutes, a charge per run Before release, on a schedule

A healthy project has many tests in the top rows and few in the bottom rows. If most of your checking happens by people clicking through the app, the pyramid is upside down.

Questions to ask

  1. Which tests run automatically on every change, and how long do they take?
  2. What happens in the tests when the model is down, slow or returns nonsense? Show me the test.
  3. Which promises to users and auditors are enforced by tests (no customer data in logs, no "released" claims, authorization checks)?
  4. How do tests run without a paid model call or SAP access? Who keeps the fake SAP data realistic?
  5. When SAP or the model provider changes something, which test notices first, and when?
  6. Do the tests that need keys get them from a secret store, never from the code?
  7. What does a failing test block: the merge, the release, nothing?
  8. Where does testing stop and evaluation start, and who owns each?

Common misconceptions

  • "You can't test AI because the answers change." You can't test the exact wording. You can test everything around it, and you can check every answer against rules.
  • "Tests need the real SAP system to mean anything." Most tests are better without it: faster, free and repeatable. A few integration tests confirm the stand-ins are honest.
  • "Passing tests means the AI is good." Tests prove the service behaves as designed. Evaluation (Unit 8) measures answer quality.
  • "Mocks are cheating." A mock is a controlled experiment. It lets you create failures, like a timeout, that you could never trigger on demand in a real system.
  • "Testing doubles the project cost." Writing tests costs time up front. Not having them costs more, in manual re-checks, incidents and slow changes.
  • "We'll add tests once it works." By then the code is often hard to test. Designing for tests from the start is cheaper.

Key terms

  • Test: a small program that checks another program and reports pass or fail.
  • Unit test: a test of one small piece, such as one rule, with nothing external involved.
  • Integration test: a test that uses a real external system, such as the SAP sandbox.
  • Test double: any stand-in for a real part during tests. Fakes, stubs and mocks are kinds of test doubles.
  • Contract test: a test that checks a promise between two systems still holds: that callers still get the fields they rely on, or that a stand-in still behaves like the real system.
  • Regression: something that used to work and broke after a change.
  • CI (continuous integration): a service that runs the tests automatically on every change, such as GitHub Actions or SAP Continuous Integration and Delivery.
  • Evaluation: measuring how good AI answers are across a set of examples (Unit 8).

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Your team says the AI explainer can't be tested because the model's wording changes every time. What is the best answer?

    Answer: B. The plumbing around the model (keys, errors, fallback, logging) behaves the same every time and can be tested exactly. Model answers can be checked against rules, such as naming the order and never claiming it was released. Answer quality is then measured by evaluation in Unit 8.
  2. 2Why do most tests of an AI service use a fake model and fake SAP data?

    Answer: C. Stand-ins remove model costs, SAP access and network delays, so the tests can run on every change. They also let you simulate a timeout or a broken answer at will, which you can't do with a real system. A few integration tests then confirm the stand-ins are still realistic.
  3. 3Which promise is the best candidate to enforce with an automated test?

    Answer: B. "No customer data or keys in logs" is a precise rule the code must always follow, so a test can check it on every change. Helpfulness is a quality question for evaluation, and model choice and business case are decisions, not behaviour a test can check.
  4. 4A vendor demo shows all tests passing. What should you conclude about the AI's answer quality?

    Answer: D. Tests prove behaviour: the right errors, the fallback, the rules. They don't measure whether answers are accurate and useful across many real cases. That is evaluation, a separate step with its own data set and metrics (Unit 8).
  5. 5Where should most of a project's automated tests sit?

    Answer: A. The healthy shape is many fast, free tests and few slow ones. Tests against real systems and the paid model are valuable but slow and costly, so they run less often. Relying mainly on manual testing turns the pyramid upside down.
  6. 6How do you find out quickly that SAP changed a field your stand-in data relies on?

    Answer: C. A contract test checks that the stand-in still matches the real system. Running it daily against the sandbox spots a change soon after it happens, while the fast tests keep using the stand-in. Calling the real system in every test would make all tests slow and fragile.
  7. 7In SAP's CAP framework, what does cds.test give your developers?

    Answer: B. cds.test starts the CAP service in-process on a free port with an in-memory database, so tests can call it and check the answers. Deployment pipelines come from a CI service such as SAP Continuous Integration and Delivery, and answer quality is evaluation.
Deep layer · 40 min read

Mental model: test what you control exactly, fence in what you don't

An AI service has two kinds of parts. Parts you control: input checks, the fallback rule, the cache, the rate limit, what gets logged, how SAP records are filtered. Parts you don't: the model, the network and the SAP system. Your parts should behave exactly the same on every run, so you test them exactly. The other parts are slow, cost money and change without asking you, so you replace them with test doubles in almost every test.

That only works if the code has a seam: a place where a stand-in can be plugged in without editing the code. The AI API from Building an AI API already has one. create_app(model, ...) accepts any object with a .name and an .answer() method. In production that is the real model; in tests it is a fake. In this topic you add a second seam: a function that reads SAP data takes the HTTP session as an argument, so a test can hand it a mock session.

Then two kinds of tests keep the doubles honest. Contract tests check that the doubles still look like the real thing. Integration tests call the real thing, a few times, on a schedule. And for the model's free text, you don't compare wording; you check rules that every acceptable answer must follow.

How it works

pytest in five ideas

pytest is the standard Python test runner. You met it briefly in Python projects done right. This topic uses five of its features:

Idea What it looks like Why you need it
Discovery Files named test_*.py, functions named test_... pytest finds and runs them; no list to maintain
Plain assert assert body["source"] == "model" On failure, pytest shows both sides of the comparison
Fixtures A function marked @pytest.fixture; a test asks for it by naming it as an argument Shared setup, fresh for every test; shared fixtures live in conftest.py and need no import
Parametrize @pytest.mark.parametrize("number", ["4711", "4712"]) One test function, many cases, each reported on its own
Markers @pytest.mark.integration, then pytest -m "not integration" Choose which tests run where; register them in pytest.ini so a typo is an error, not a silent no-op

pytest's documentation recommends registering custom markers. With --strict-markers, an unknown marker stops the run, which protects you from @pytest.mark.integraton silently doing nothing.

Test doubles: fake, stub, mock

"Mock" is often used for all of them, but the differences help you pick the right one:

Kind What it does In this topic
Fake A small working version with shortcuts FakeModel returns scripted replies and records calls
Stub Gives canned answers, nothing more A fake SAP page with two made-up orders
Mock Records how it was called so the test can check it Mock(spec=requests.Session); assert_called_once_with(...)

Python's built-in unittest.mock provides mocks. Three features do most of the work: return_value (what the mock returns), side_effect (raise an exception, or return a different value on each call) and create_autospec (a mock that refuses calls the real object would refuse). That last one matters. A plain Mock() accepts any method name, including typos, so a test can pass while the real code would crash.

Testing the API in memory

FastAPI's TestClient sends requests to your app directly in memory, without starting a server. It is built on the HTTPX library and is used like requests: client.post("/v1/explain", json=..., headers=...). The FastAPI testing guide notes that the test functions are plain def, not async def, and the calls need no await.

Non-deterministic output: rules, not wording

A model can say "Order 4711 is on hold" or "Sales order 4711 cannot ship yet". Both are fine. A test that compares the text with one expected sentence fails half the time, and people learn to ignore it. Instead, write a function that lists every broken rule:

  • the answer parses and matches the schema (your API already rejects anything else),
  • next_step is one of the allowed values and fits the block reason,
  • the text names the order and stays within the word limit,
  • the text never contains claims only a person may make, such as "released".

Hard rules (schema, allowed values, forbidden claims) are strict. Soft rules (length) get a little slack when you check a real model. Then test the checker itself with a known-bad answer, because a checker that never fails proves nothing.

Contract tests in two directions

flowchart LR
  C[Sales app<br/>the caller] -->|relies on fields| A[Your AI API]
  A -->|relies on fields| S[SAP sandbox<br/>API_SALES_ORDER_SRV]
  T1[test_contract.py] -. checks the API still<br/>publishes those fields .-> A
  T2[test_sandbox_integration.py] -. checks SAP still<br/>sends those fields .-> S
  D[Fake SAP page<br/>conftest.py] -. same field list .-> T2

Toward your callers: the sales app reads sales_order, explanation, next_step, source and request_id. A test reads the OpenAPI contract that FastAPI publishes at /openapi.json and fails if any of those fields stops being required, or if someone adds a new next_step value the sales app can't handle.

Toward SAP: your fake SAP page is only useful while it looks like SAP. Martin Fowler describes contract tests as checks that calls against a test double give the same results as calls to the real service. He suggests running them on the rhythm of the external service, often daily, and treating a failure as a signal to realign the double, not necessarily as a broken build. In this topic one list, FIELDS, drives both the fake SAP page check and the sandbox check, so they can't drift apart.

Where the tests run

flowchart TB
  P[You push a branch] --> F[GitHub Actions: fast tests<br/>unit, API, contract<br/>no key, no model]
  N[Daily schedule or Run workflow] --> I[GitHub Actions: sandbox tests<br/>needs SAP_API_KEY secret]
  L[You, before a release] --> M[Model output checks<br/>RUN_LLM_TESTS=1, small charge]

Fast tests run on every push. Sandbox tests run once a day and on demand, using a secret. Real-model checks run when you choose, because they cost money.

Build it yourself: a test suite for the blocked-orders API

Before you start: complete Set up your computer for this course and Set up for Unit 6. You also need unit06/ai_api.py from Building an AI API, Step 3; this topic tests it without changing it.

You will write 55 automated checks for the blocked-orders AI API and a small SAP reader, plus 5 optional ones that use the real SAP sandbox and the real model. They run in under a second on your computer, then on GitHub on every push.

flowchart LR
  subgraph tests[unit06/tests]
    R[test_rules.py] --> AI[ai_api.py]
    A[test_api.py] -->|TestClient + FakeModel| AI
    C[test_contract.py] --> AI
    M[test_model_output.py] --> AI
    S[test_sap_orders.py] -->|mock session| SO[sap_orders.py]
    I[test_sandbox_integration.py] -->|real call, optional| SAP[(SAP sandbox)]
  end
  SO --> SAP

What you need

  • Your course folder with .venv, unit06/ai_api.py, and the Unit 6 setup done. Free.
  • About 60 to 90 minutes.
  • Optional: your free SAP Business Accelerator Hub key in .env as SAP_API_KEY (from the computer setup topic), for the sandbox tests.
  • Optional: the Unit 5 AICORE_ settings in .env, for the real-model checks. Small per-request charge.
  • Optional: your course repository on GitHub from Git and GitHub for AI engineers, for Step 11. How many GitHub Actions minutes you get for free depends on your GitHub plan and whether the repository is public or private.

Step 1: Open your course folder and install pytest

  1. Open VS Code, choose File > Open Folder and open orchestrate-course.

  2. Open a terminal with Terminal > New Terminal.

  3. Turn on the virtual environment.

    Windows (PowerShell):

    .venv\Scripts\Activate.ps1

    macOS/Linux:

    source .venv/bin/activate

    The prompt now starts with (.venv).

  4. Open requirements.txt, add this line at the end and save:

    pytest
  5. Install it:

    pip install -r requirements.txt
  6. Check it:

    pytest --version

What success looks like:

pytest 9.1.1

Any version 8 or newer works.

Step 2: Tell pytest where the tests are

  1. In the unit06 folder, create a folder named tests (right-click unit06, New Folder).
  2. In unit06 (not in tests), create pytest.ini with this content:
[pytest]
# Settings for the Unit 6 tests. pytest finds this file when you run: pytest unit06
testpaths = tests
pythonpath = .
addopts = --strict-markers -ra
# Newer Starlette versions warn about the library TestClient uses. Harmless for this course.
filterwarnings =
    ignore:Using .httpx. with .starlette.testclient. is deprecated
markers =
    integration: calls the real SAP sandbox; needs SAP_API_KEY in .env (leave out with -m "not integration")
    llm: calls a real model through SAP's orchestration service; costs a little per run (needs RUN_LLM_TESTS=1)

pythonpath = . lets the tests import ai_api and sap_orders from unit06. The two markers label the tests that need outside access, and --strict-markers turns a misspelled marker into an error. -ra prints a short reason for every skipped test at the end.

  1. In unit06, create requirements-test.txt. GitHub uses it in Step 11, because the course's own requirements.txt also holds large libraries such as torch that these tests don't need:
fastapi[standard]
requests
python-dotenv
pytest

Step 3: Add a small SAP reader with a seam

The AI API uses made-up orders. Real projects read them from SAP, so here is a small reader for the sandbox's Sales Order API, the same service you used in Calling your first SAP API. The important design choice: fetch_orders receives the HTTP session as an argument instead of creating it. That is the seam a test uses to pass in a mock.

  1. In unit06, create sap_orders.py:
"""Unit 6: read blocked sales orders from SAP's sandbox, written so it is easy to test.

The network call (fetch_orders) and the business rule (blocked_orders) are separate functions.
Tests give fetch_orders a fake session instead of a real one, so they need no key and no internet.

How to run (from your course folder, with .venv turned on; needs SAP_API_KEY in .env):
    python unit06/sap_orders.py
"""
from decimal import Decimal

import requests

SANDBOX = "https://sandbox.api.sap.com/s4hanacloud/sap/opu/odata/sap/API_SALES_ORDER_SRV"
FIELDS = ["SalesOrder", "SoldToParty", "TotalNetAmount", "TransactionCurrency",
          "DeliveryBlockReason", "HeaderBillingBlockReason"]   # the only fields this code relies on


class SapApiError(Exception):
    """SAP could not be reached, refused the call, or answered in an unexpected shape."""


def fetch_orders(session, api_key: str, top: int = 50, timeout: float = 30) -> list:
    """Read up to `top` sales order headers. `session` is a requests.Session, or a fake one in tests."""
    try:
        response = session.get(f"{SANDBOX}/A_SalesOrder",
                               params={"$top": str(top), "$select": ",".join(FIELDS)},
                               headers={"APIKey": api_key, "Accept": "application/json"},
                               timeout=timeout)
    except requests.RequestException as error:   # no network, proxy, DNS, timeout
        raise SapApiError(f"could not reach SAP ({type(error).__name__})") from error
    if response.status_code != 200:
        raise SapApiError(f"SAP answered HTTP {response.status_code}")
    try:
        return response.json()["d"]["results"]   # OData V2 puts the list under d.results
    except (ValueError, KeyError, TypeError) as error:
        raise SapApiError("SAP answered, but not with the expected OData JSON") from error


def blocked_orders(records: list) -> list:
    """Keep orders with a delivery or billing block (any value) and tidy the fields."""
    result = []
    for record in records:
        delivery = record.get("DeliveryBlockReason") or None   # SAP sends "" when no block is set
        billing = record.get("HeaderBillingBlockReason") or None
        if not (delivery or billing):
            continue
        result.append({"sales_order": record["SalesOrder"],
                       "sold_to_party": record.get("SoldToParty"),
                       "net_value": Decimal(record.get("TotalNetAmount") or "0"),
                       "currency": record.get("TransactionCurrency"),
                       "delivery_block_reason": delivery,
                       "billing_block_reason": billing})
    return result


def main() -> None:
    import os
    import sys

    from dotenv import load_dotenv
    load_dotenv()
    key = os.environ.get("SAP_API_KEY")
    if not key:
        sys.exit("SAP_API_KEY is not set in .env. The tests still run without it: pytest unit06")
    try:
        orders = blocked_orders(fetch_orders(requests.Session(), key, top=50))
    except SapApiError as error:
        sys.exit(f"Stopped: {error}")
    print(f"{len(orders)} blocked order(s) among the first 50 in the sandbox")
    for order in orders[:5]:
        print(f"  {order['sales_order']}: delivery block {order['delivery_block_reason']}, "
              f"billing block {order['billing_block_reason']}")


if __name__ == "__main__":
    main()
  1. If you have SAP_API_KEY in .env, try it (otherwise skip; the tests don't need it):

    python unit06/sap_orders.py

What success looks like (the count and numbers depend on the sandbox data that day):

2 blocked order(s) among the first 50 in the sandbox
  ...: delivery block ..., billing block None

A result of 0 blocked order(s) is valid too: it means none of the first 50 orders has a block today.

Step 4: Shared fixtures in conftest.py

conftest.py holds the setup that several test files share. pytest loads it by itself; tests ask for a fixture by naming it as an argument.

  1. In unit06/tests, create conftest.py:
"""Shared test helpers for Unit 6. pytest loads this file by itself; tests ask for a fixture by
naming it as an argument."""
import json

import pytest
from fastapi.testclient import TestClient

from ai_api import ModelUnavailable, create_app

KEY = "test-key-0123456789abcdef"   # a made-up key, used only inside the tests


class FakeModel:
    """A stand-in for the model. It returns the replies you script, in order, and records every call.
    A reply is either (text, tokens) or an exception to raise."""
    name = "fake-model"

    def __init__(self):
        self.replies = []
        self.calls = []

    def answer(self, sales_order: str, order: dict, max_words: int):
        self.calls.append(sales_order)
        reply = self.replies.pop(0) if self.replies else good_reply()
        if isinstance(reply, Exception):
            raise reply
        return reply


def good_reply(next_step: str = "credit_review", text: str = "Order is over the credit limit.") -> tuple:
    return json.dumps({"explanation": text, "next_step": next_step}), 42


@pytest.fixture
def headers():
    return {"X-API-Key": KEY}


@pytest.fixture
def fake_model():
    return FakeModel()


@pytest.fixture
def make_client():
    """Build a TestClient around any model. Each test gets a fresh app, so caches and rate limits
    never leak from one test into the next."""
    def build(model, **settings):
        settings.setdefault("rate_limit_per_minute", 100)
        return TestClient(create_app(model, KEY, **settings))
    return build


@pytest.fixture
def client(make_client, fake_model):
    return make_client(fake_model)


@pytest.fixture
def sandbox_page():
    """Two records in the shape SAP's sandbox returns (OData V2: the list sits under d.results).
    Field names follow API_SALES_ORDER_SRV; the values are made up."""
    return {"d": {"results": [
        {"SalesOrder": "9000001", "SoldToParty": "10100001", "TotalNetAmount": "17.55",
         "TransactionCurrency": "EUR", "DeliveryBlockReason": "01", "HeaderBillingBlockReason": ""},
        {"SalesOrder": "9000002", "SoldToParty": "10100002", "TotalNetAmount": "980.00",
         "TransactionCurrency": "USD", "DeliveryBlockReason": "", "HeaderBillingBlockReason": ""},
    ]}}

The field names in sandbox_page are the ones SAP's API returns, and the shape matches a recorded record in SAP's own sample project: the list sits under d, amounts are strings, and an unset block reason is an empty string "".

Step 5: Unit tests for the rules

Start with the pure rules: no HTTP, no model. These are the fastest and most precise tests you have.

  1. In unit06/tests, create test_rules.py:
"""Unit tests: the deterministic parts of ai_api.py. No HTTP, no model, no network."""
import pytest
from pydantic import ValidationError

from ai_api import BLOCKED_ORDERS, ExplainRequest, ModelAnswer, fallback_answer


@pytest.mark.parametrize("reason, expected_step", [
    ("Credit limit exceeded", "credit_review"),
    ("Missing export documents", "complete_documents"),
    ("Incomplete delivery address", "fix_master_data"),
    ("A reason nobody planned for", "contact_customer"),   # unknown reasons get a safe default
])
def test_fallback_picks_the_next_step_from_the_reason(reason, expected_step):
    order = {"customer": "Test Co", "reason": reason}
    assert fallback_answer("1", order)["next_step"] == expected_step


@pytest.mark.parametrize("number", sorted(BLOCKED_ORDERS))
def test_fallback_never_claims_the_order_is_released(number):
    text = fallback_answer(number, BLOCKED_ORDERS[number])["explanation"].lower()
    assert number in text
    assert "released" not in text and "will ship" not in text


def test_model_answer_accepts_a_valid_reply():
    answer = ModelAnswer.model_validate_json('{"explanation": "Over the limit.", "next_step": "credit_review"}')
    assert answer.next_step == "credit_review"


@pytest.mark.parametrize("raw", [
    '{"explanation": "I released it.", "next_step": "release_order"}',   # action outside the list
    '{"explanation": "", "next_step": "credit_review"}',                 # empty text
    '{"next_step": "credit_review"}',                                    # field missing
    'Sure! Here is the JSON you asked for.',                             # not JSON at all
])
def test_model_answer_rejects_bad_replies(raw):
    with pytest.raises(ValidationError):
        ModelAnswer.model_validate_json(raw)


@pytest.mark.parametrize("number", ["47-11", "", "12345678901", "4711; DROP TABLE"])
def test_request_rejects_malformed_order_numbers(number):
    with pytest.raises(ValidationError):
        ExplainRequest(sales_order=number)
  1. Run just this file:

    pytest unit06/tests/test_rules.py -v

What success looks like (shortened):

unit06/tests/test_rules.py::test_fallback_picks_the_next_step_from_the_reason[Credit limit exceeded-credit_review] PASSED
...
unit06/tests/test_rules.py::test_request_rejects_malformed_order_numbers[4711; DROP TABLE] PASSED
============================== 17 passed in 0.12s ==============================

Five test functions became 17 tests, because parametrize runs each case separately. -v lists every one by name.

Step 6: API tests with a fake model

Now the whole API, through TestClient, with FakeModel behind it. Each test sets up the replies it needs, so you can create a timeout or a rule-breaking answer on demand.

  1. In unit06/tests, create test_api.py:
"""API tests: every promise of POST /v1/explain, with a fake model behind it. In memory, no server."""
import json
import logging
from unittest.mock import ANY, create_autospec

from ai_api import ModelUnavailable, SampleModel


def reply(next_step="credit_review", text="The order is over the customer's credit limit."):
    return json.dumps({"explanation": text, "next_step": next_step}), 42


def explain(client, headers, number="4711", **extra):
    return client.post("/v1/explain", json={"sales_order": number, **extra}, headers=headers)


# ---------- who may call, and what they may send ----------

def test_health_needs_no_key(client):
    assert client.get("/health").status_code == 200


def test_missing_key_is_401(client):
    response = explain(client, headers={})
    assert response.status_code == 401
    assert response.headers["WWW-Authenticate"] == "APIKey"


def test_wrong_key_is_401(client):
    assert explain(client, headers={"X-API-Key": "wrong-key-0123456789"}).status_code == 401


def test_bad_input_is_422_and_never_reaches_the_model(client, headers, fake_model):
    assert explain(client, headers, number="47-11").status_code == 422
    assert explain(client, headers, max_words=500).status_code == 422
    assert fake_model.calls == []   # rejected before any model cost


def test_unknown_order_is_404(client, headers, fake_model):
    assert explain(client, headers, number="9999").status_code == 404
    assert fake_model.calls == []


# ---------- the three ways an answer can come back ----------

def test_good_model_reply_is_returned(client, headers, fake_model):
    fake_model.replies = [reply()]
    body = explain(client, headers).json()
    assert body["source"] == "model"
    assert body["next_step"] == "credit_review"
    assert body["tokens"] == 42


def test_model_timeout_gives_fallback_not_error(client, headers, fake_model):
    fake_model.replies = [ModelUnavailable("simulated timeout")]
    response = explain(client, headers)
    assert response.status_code == 200
    assert response.json()["fallback_reason"] == "model unavailable"
    assert response.json()["next_step"] == "credit_review"   # 4711 is a credit block


def test_reply_outside_the_schema_gives_fallback(client, headers, fake_model):
    fake_model.replies = [reply(next_step="release_order", text="Done, I released it.")]
    body = explain(client, headers).json()
    assert body["source"] == "fallback"
    assert body["fallback_reason"] == "model answer rejected"
    assert "released" not in body["explanation"]


def test_cache_calls_the_model_once(make_client, headers):
    model = create_autospec(SampleModel, instance=True)   # a mock that only allows SampleModel's methods
    model.name = "mock-model"
    model.answer.return_value = reply()
    client = make_client(model)
    first = explain(client, headers).json()
    second = explain(client, headers).json()
    assert (first["source"], second["source"]) == ("model", "cache")
    model.answer.assert_called_once_with("4711", ANY, 60)   # 60 is the default max_words


def test_fallback_answers_are_not_cached(client, headers, fake_model):
    fake_model.replies = [ModelUnavailable("down"), reply()]
    assert explain(client, headers).json()["source"] == "fallback"
    assert explain(client, headers).json()["source"] == "model"   # tried the model again


# ---------- operations: rate limit, request IDs, logs ----------

def test_rate_limit_returns_429_with_retry_after(make_client, fake_model, headers):
    client = make_client(fake_model, rate_limit_per_minute=2)
    codes = [explain(client, headers).status_code for _ in range(3)]
    assert codes == [200, 200, 429]
    last = explain(client, headers)
    assert int(last.headers["Retry-After"]) > 0


def test_request_id_is_kept_when_valid_and_replaced_when_not(client, headers):
    kept = client.get("/health", headers={"X-Request-ID": "order-check-0001"})
    assert kept.headers["X-Request-ID"] == "order-check-0001"
    replaced = client.get("/health", headers={"X-Request-ID": "<script>"})
    assert replaced.headers["X-Request-ID"] != "<script>"


def test_logs_hold_no_customer_data_or_keys(client, headers, fake_model, caplog):
    caplog.set_level(logging.INFO, logger="ai_api")
    fake_model.replies = [ModelUnavailable("down")]
    explain(client, headers)
    logged = caplog.text
    assert "/v1/explain" in logged                 # the request was logged...
    assert "Made-up Retail GmbH" not in logged     # ...but not the customer
    assert headers["X-API-Key"] not in logged      # ...and not the key
  1. Run it:

    pytest unit06/tests/test_api.py

What success looks like:

unit06/tests/test_api.py .............                                   [100%]
============================== 13 passed in 0.15s ==============================

Notice three techniques. fake_model.calls == [] proves bad input is rejected before any model cost. create_autospec(SampleModel, instance=True) builds a mock that only accepts SampleModel's real methods, and assert_called_once_with proves the cache saved the second model call. caplog is a built-in pytest fixture that captures log lines, so a test can prove the customer name and the key never reach the log.

Step 7: Mock SAP, and contract tests

  1. In unit06/tests, create test_sap_orders.py:
"""Tests for sap_orders.py with a mocked SAP: no key, no internet, and failures on demand."""
from decimal import Decimal
from unittest.mock import Mock

import pytest
import requests

from sap_orders import SANDBOX, SapApiError, blocked_orders, fetch_orders


def fake_session(status=200, payload=None, error=None):
    """A mock requests.Session whose get() returns a canned response, or raises `error`."""
    session = Mock(spec=requests.Session)
    if error:
        session.get.side_effect = error
    else:
        response = Mock(status_code=status)
        response.json.return_value = payload
        session.get.return_value = response
    return session


def test_fetch_sends_key_header_select_and_timeout(sandbox_page):
    session = fake_session(payload=sandbox_page)
    records = fetch_orders(session, "my-key", top=2)
    assert len(records) == 2
    url = session.get.call_args.args[0]
    kwargs = session.get.call_args.kwargs
    assert url == f"{SANDBOX}/A_SalesOrder"
    assert kwargs["headers"]["APIKey"] == "my-key"
    assert kwargs["params"]["$top"] == "2"
    assert "DeliveryBlockReason" in kwargs["params"]["$select"]
    assert kwargs["timeout"] > 0   # a call without a timeout can hang forever


@pytest.mark.parametrize("status", [401, 403, 429, 500])
def test_error_status_raises_sap_api_error(status):
    with pytest.raises(SapApiError, match=str(status)):
        fetch_orders(fake_session(status=status), "my-key")


@pytest.mark.parametrize("error", [requests.Timeout("slow"), requests.ConnectionError("proxy")])
def test_network_problems_raise_sap_api_error(error):
    with pytest.raises(SapApiError, match="could not reach"):
        fetch_orders(fake_session(error=error), "my-key")


def test_unexpected_shape_raises_sap_api_error():
    session = fake_session(payload={"value": []})   # an OData V4 shape, not V2
    with pytest.raises(SapApiError, match="expected OData"):
        fetch_orders(session, "my-key")


def test_blocked_orders_keeps_only_blocked_ones(sandbox_page):
    orders = blocked_orders(sandbox_page["d"]["results"])
    assert [o["sales_order"] for o in orders] == ["9000001"]
    assert orders[0]["net_value"] == Decimal("17.55")   # money as Decimal, never float
    assert orders[0]["billing_block_reason"] is None   # "" from SAP becomes None


def test_blocked_orders_handles_an_empty_page():
    assert blocked_orders([]) == []

Mock(spec=requests.Session) only allows methods a real session has. side_effect makes get() raise a timeout or a connection error, which you can't reliably trigger with a real network. call_args lets the test read exactly what was sent: the URL, the APIKey header, the $top and $select options and the timeout.

  1. In unit06/tests, create test_contract.py:
"""Contract tests: the promises other people build on. If one fails, a caller or a test double
would break, so fix the code or update the contract on purpose, never by accident."""
from sap_orders import FIELDS

# What the sales app reads from every answer. Removing or renaming one breaks the sales app.
FIELDS_CALLERS_USE = {"sales_order", "explanation", "next_step", "source", "request_id"}
NEXT_STEPS_CALLERS_HANDLE = {"credit_review", "complete_documents", "fix_master_data", "contact_customer"}


def response_schema(client):
    """The response schema from the contract FastAPI publishes at /openapi.json."""
    return client.get("/openapi.json").json()["components"]["schemas"]["ExplainResponse"]


def test_explain_endpoint_is_still_published(client):
    assert "post" in client.get("/openapi.json").json()["paths"]["/v1/explain"]


def test_response_still_has_every_field_callers_use(client):
    schema = response_schema(client)
    assert FIELDS_CALLERS_USE <= set(schema["required"])


def test_no_new_next_step_without_telling_callers(client):
    schema = response_schema(client)
    assert set(schema["properties"]["next_step"]["enum"]) == NEXT_STEPS_CALLERS_HANDLE


def test_sap_double_has_every_field_the_code_relies_on(sandbox_page):
    """The fake SAP page in conftest.py must look like SAP. The integration test checks the
    same list against the real sandbox, so the double can't drift away unnoticed."""
    for record in sandbox_page["d"]["results"]:
        assert set(FIELDS) <= set(record)
  1. Run both:

    pytest unit06/tests/test_sap_orders.py unit06/tests/test_contract.py

What success looks like:

unit06/tests/test_sap_orders.py ..........                               [ 71%]
unit06/tests/test_contract.py ....                                       [100%]
============================== 14 passed in 0.14s ==============================

Step 8: Checks for answers that change every run

  1. In unit06/tests, create test_model_output.py:
"""Testing output that changes every run: check the structure and the rules, never the wording."""
import json
import os
import random

import pytest

from ai_api import BLOCKED_ORDERS, NEXT_STEP_BY_REASON, PROMPT_VERSION

FORBIDDEN = ["released", "will ship", "has shipped", "approved"]   # claims only a person may make
OPENINGS = ["Order {n} is on hold.", "Sales order {n} cannot ship yet.", "{n} is blocked right now."]
REASONS = ["The customer is over the credit limit.", "Credit exposure is too high for this customer.",
           "Credit management stopped it."]


class VaryingModel:
    """A fake model that words its answer differently on every call, like a real one."""
    name = "varying-model"

    def answer(self, sales_order, order, max_words):
        text = f"{random.choice(OPENINGS).format(n=sales_order)} {random.choice(REASONS)}"
        return json.dumps({"explanation": text, "next_step": "credit_review"}), 30


def rule_problems(body: dict, number: str, max_words: int, slack: float = 1.0) -> list:
    """Return every broken rule. An empty list means the answer passes."""
    problems = []
    words = len(body["explanation"].split())
    if words > max_words * slack:
        problems.append(f"{words} words, limit {max_words}")
    if number not in body["explanation"]:
        problems.append("does not name the order")
    for phrase in FORBIDDEN:
        if phrase in body["explanation"].lower():
            problems.append(f"claims '{phrase}'")
    reason = BLOCKED_ORDERS[number]["reason"]
    if body["source"] != "fallback" and body["next_step"] != NEXT_STEP_BY_REASON.get(reason):
        problems.append(f"next_step {body['next_step']} does not fit '{reason}'")
    if body["prompt_version"] != PROMPT_VERSION:
        problems.append("prompt_version missing or wrong")
    return problems


@pytest.mark.parametrize("attempt", range(10))   # ten runs, ten different wordings
def test_varying_answers_all_follow_the_rules(make_client, headers, attempt):
    client = make_client(VaryingModel(), cache_seconds=0)
    body = client.post("/v1/explain", json={"sales_order": "4711"}, headers=headers).json()
    assert rule_problems(body, "4711", 60) == []


def test_rule_checker_catches_a_bad_answer():
    """Test the test: a checker that never fails proves nothing."""
    bad = {"explanation": "Good news, I approved and released it.", "next_step": "contact_customer",
           "source": "model", "prompt_version": PROMPT_VERSION}
    problems = rule_problems(bad, "4711", 60)
    assert "does not name the order" in problems
    assert "claims 'released'" in problems
    assert any(p.startswith("next_step") for p in problems)


@pytest.mark.llm
@pytest.mark.parametrize("number", ["4711", "4712"])
def test_real_model_follows_the_rules(make_client, headers, number):
    """Calls the real model through SAP's orchestration service. Opt in with RUN_LLM_TESTS=1."""
    if os.environ.get("RUN_LLM_TESTS") != "1":
        pytest.skip("set RUN_LLM_TESTS=1 to call the real model (small per-request charge)")
    from dotenv import load_dotenv
    load_dotenv()
    if not os.environ.get("AICORE_CLIENT_ID"):
        pytest.skip("AICORE_ settings missing in .env (see 'Set up for Unit 5')")
    from ai_api import OrchestrationModel
    client = make_client(OrchestrationModel("gpt-4o-mini"), cache_seconds=0)
    body = client.post("/v1/explain", json={"sales_order": number}, headers=headers).json()
    assert body["source"] == "model", body.get("fallback_reason")
    assert rule_problems(body, number, 60, slack=1.2) == []   # 20% slack on length: a soft rule

VaryingModel words its answer differently on every call, like a real model. The same rule checker that judges it will later judge the real model in Step 12. cache_seconds=0 turns the cache off so every call reaches the model.

Step 9: Integration tests against the SAP sandbox

  1. In unit06/tests, create test_sandbox_integration.py:
"""Integration tests against SAP's real sandbox. They need SAP_API_KEY in .env and internet access,
so they are marked `integration` and skip themselves when the key is missing."""
import os

import pytest
import requests
from dotenv import load_dotenv

from sap_orders import FIELDS, SapApiError, blocked_orders, fetch_orders

load_dotenv()   # finds .env in your course folder
KEY = os.environ.get("SAP_API_KEY")
pytestmark = [pytest.mark.integration,
              pytest.mark.skipif(not KEY, reason="SAP_API_KEY not set; the sandbox tests need it")]


def test_sandbox_returns_records_with_every_field_we_rely_on():
    records = fetch_orders(requests.Session(), KEY, top=5)
    assert records, "the sandbox returned no orders"
    for record in records:
        missing = set(FIELDS) - set(record)
        assert not missing, f"SAP no longer sends {missing}: update sap_orders.py and the test double"


def test_blocked_filter_runs_on_real_data():
    orders = blocked_orders(fetch_orders(requests.Session(), KEY, top=20))
    assert all(o["delivery_block_reason"] or o["billing_block_reason"] for o in orders)


def test_wrong_key_is_refused():
    with pytest.raises(SapApiError, match="HTTP 4"):   # refused by SAP, not a network error
        fetch_orders(requests.Session(), "not-a-real-key", top=1)

pytestmark applies both markers to every test in the file: integration, and a skip when the key is missing.

  1. Run the whole suite:

    pytest unit06/tests

What success looks like without SAP_API_KEY and without RUN_LLM_TESTS:

============================= test session starts ==============================
platform linux -- Python 3.13.16, pytest-9.1.1, pluggy-1.6.0
rootdir: /home/you/orchestrate-course/unit06
configfile: pytest.ini
collected 60 items

unit06/tests/test_api.py .............                                   [ 21%]
unit06/tests/test_contract.py ....                                       [ 28%]
unit06/tests/test_model_output.py ...........ss                          [ 50%]
unit06/tests/test_rules.py .................                             [ 78%]
unit06/tests/test_sandbox_integration.py sss                             [ 83%]
unit06/tests/test_sap_orders.py ..........                               [100%]

=========================== short test summary info ============================
SKIPPED [2] unit06/tests/test_model_output.py:66: set RUN_LLM_TESTS=1 to call the real model (small per-request charge)
SKIPPED [1] unit06/tests/test_sandbox_integration.py:17: SAP_API_KEY not set; the sandbox tests need it
SKIPPED [1] unit06/tests/test_sandbox_integration.py:25: SAP_API_KEY not set; the sandbox tests need it
SKIPPED [1] unit06/tests/test_sandbox_integration.py:30: SAP_API_KEY not set; the sandbox tests need it
======================== 55 passed, 5 skipped in 0.29s =========================

Each . is a passed test, each s a skipped one. Skips are a valid result here: those tests need a key or a paid model and tell you why. The first lines show your own platform, Python version and folder, and the line numbers in the skip reasons may differ slightly from yours. With SAP_API_KEY in .env and a network that reaches sandbox.api.sap.com, the three sss become ... and the summary says 58 passed, 2 skipped.

  1. Run only the tests that need nothing, as GitHub will in Step 11:

    pytest unit06/tests -m "not integration and not llm"

What success looks like:

55 passed, 5 deselected in 0.31s

"Deselected" means the marker filter left them out on purpose.

Step 10: Make a test fail on purpose

A test you have never seen fail might not test anything. Break the code and watch.

  1. Open unit06/ai_api.py and find fallback_answer. Change the sentence "The responsible team must review it before it can ship." to "It will ship once released." and save.

  2. Run:

    pytest unit06/tests -q

What you should see (shortened):

>       assert "released" not in text and "will ship" not in text
E       AssertionError: assert ('released' not in 'order 4714 ...ce released.'
E         'released' is contained here:
E           ship once released.)
...
FAILED unit06/tests/test_api.py::test_reply_outside_the_schema_gives_fallback
FAILED unit06/tests/test_rules.py::test_fallback_never_claims_the_order_is_released[4711]
...
5 failed, 50 passed, 5 skipped in 0.34s

pytest shows the failing line, the values involved and the exact word that broke the rule.

  1. Change the sentence back, save, and run pytest unit06/tests -q again. It ends with 55 passed, 5 skipped.

Step 11: Run the tests on GitHub for every push

  1. In your course folder (not in unit06), create the folders .github and, inside it, workflows. Watch the leading dot in .github.
  2. In .github/workflows, create unit06-tests.yml:
# Runs the Unit 6 tests on GitHub's computers.
# Fast tests: on every push and pull request. Sandbox tests: once a day and on demand.
name: unit06-tests

on:
  push:
  pull_request:
  schedule:
    - cron: "17 6 * * *"   # every day at 06:17 UTC
  workflow_dispatch:       # adds a "Run workflow" button on the Actions tab

jobs:
  fast-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.x"
      - name: Install the libraries the tests need
        run: |
          python -m pip install --upgrade pip
          pip install -r unit06/requirements-test.txt
      - name: Unit, API and contract tests (no key, no model)
        run: pytest unit06/tests -m "not integration and not llm"

  sandbox-tests:
    if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
    runs-on: ubuntu-latest
    env:
      SAP_API_KEY: ${{ secrets.SAP_API_KEY }}
    steps:
      - uses: actions/checkout@v6
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.x"
      - name: Install the libraries the tests need
        run: pip install -r unit06/requirements-test.txt
      - name: Integration tests against the SAP sandbox
        if: env.SAP_API_KEY != ''
        run: pytest unit06/tests -m integration

The fast job runs on every push and pull request. The sandbox job runs only on the daily schedule or when you click Run workflow. GitHub doesn't allow secrets directly in an if: condition, so the workflow copies the secret into a job variable first and checks that.

  1. Optional: add your sandbox key as a secret. On GitHub, open your repository, click Settings, then in the left menu Secrets and variables > Actions. On the Secrets tab click New repository secret. Name: SAP_API_KEY. Secret: paste your key. Click Add secret.
  1. Save your work on a branch and push it:

    git checkout -b unit06-tests
    git add requirements.txt unit06/pytest.ini unit06/requirements-test.txt unit06/sap_orders.py unit06/tests .github/workflows/unit06-tests.yml
    git commit -m "Add Unit 6 test suite and CI workflow"
    git push -u origin unit06-tests
  2. On GitHub, click the Actions tab. You'll see a run named unit06-tests. Click it, then fast-tests.

What success looks like: a green check next to fast-tests, and in the step Unit, API and contract tests (no key, no model) the line 55 passed, 5 deselected. sandbox-tests shows as skipped on a push; that is expected. To run it now, choose unit06-tests in the left list, click Run workflow, then the green Run workflow button.

  1. Open a pull request from unit06-tests to main. The check appears on the pull request page; merge once it is green.

Step 12 (optional): Check the real model against the rules

This calls the real model twice through SAP's orchestration service, with a small charge per request.

  1. Make sure python check_unit05.py passes.

  2. Turn on the opt-in and run only the llm tests.

    Windows (PowerShell):

    $env:RUN_LLM_TESTS = "1"
    pytest unit06/tests -m llm
    Remove-Item Env:RUN_LLM_TESTS

    macOS/Linux:

    RUN_LLM_TESTS=1 pytest unit06/tests -m llm

What success looks like:

unit06/tests/test_model_output.py ..                                     [100%]
====================== 2 passed, 58 deselected in 4.81s ======================

Your time will differ; it depends on the model. A failure lists the broken rules, for example ['claims \'released\'']; that is a finding about the prompt or model, and exactly what this check is for. If gpt-4o-mini isn't offered in your account, change the model name in the test to one that check_unit05.py lists.

What the code does

Part What it does
pytest.ini Says where tests are, adds unit06 to the import path, registers the integration and llm markers, and makes unknown markers an error
sap_orders.fetch_orders Calls the sandbox with the APIKey header, $top, $select and a timeout; turns every failure into one SapApiError
sap_orders.blocked_orders The business rule: keep orders with a delivery or billing block; amounts as Decimal
conftest.FakeModel A fake model that returns scripted replies (or raises) and records calls
make_client fixture Builds a fresh app for each test, so caches and rate limits don't leak between tests
sandbox_page fixture A stub SAP response in the real OData V2 shape
test_rules.py Parametrized unit tests for the fallback, the answer schema and the input checks
test_api.py The API's promises through TestClient: keys, 422/404, fallback, cache, 429, request IDs, clean logs
test_sap_orders.py Mock(spec=requests.Session) with return_value and side_effect to test success, errors and network failures
test_contract.py Fails if the published OpenAPI contract loses a field callers use, or if the fake SAP page lacks a field the code needs
rule_problems Lists every broken rule in a model answer; tested itself with a known-bad answer
test_sandbox_integration.py Real calls to the sandbox, marked integration, skipped without a key
unit06-tests.yml GitHub Actions: fast tests on every push; sandbox tests daily and on demand

If something goes wrong

What you see What it means What to do
python or pytest is not recognized / command not found Python isn't installed, or .venv isn't active Check for (.venv) in the prompt; activate it (Step 1); see the computer setup topic
ModuleNotFoundError: No module named 'ai_api' pytest didn't find pytest.ini, or ai_api.py isn't in unit06 Check unit06/pytest.ini exists and ai_api.py sits next to it, not in tests
ModuleNotFoundError: No module named 'fastapi' (or requests, dotenv) A library is missing in this Python pip install -r requirements.txt with .venv active
mainloop: caught unexpected SystemExit! You ran pytest unit06, so pytest also loaded the older test_ai_api.py script, which exits when imported Run pytest unit06/tests
'integraton' not found in markers configuration option A misspelled marker; --strict-markers caught it Fix the spelling to integration or llm
fixture 'make_client' not found conftest.py is missing or not in unit06/tests Move it into unit06/tests
Sandbox tests fail with SAP answered HTTP 401 The key is missing, wrong or incomplete Copy it again from the SAP Business Accelerator Hub with Show API Key and update .env
Sandbox tests fail with could not reach SAP (ProxyError) or ConnectionError A network, proxy or firewall blocks sandbox.api.sap.com Try another network or ask IT; the fast tests don't need it
test_wrong_key_is_refused fails with could not reach Same network problem: the call never reached SAP As above
A DeprecationWarning about httpx and starlette.testclient Newer Starlette versions warn about the library TestClient uses Harmless here; pytest.ini hides that one warning
GitHub: fast-tests red with No such file or directory: unit06/requirements-test.txt The file wasn't committed git add unit06/requirements-test.txt, commit and push
GitHub: sandbox-tests green but no tests ran The SAP_API_KEY secret isn't set, so the step is skipped on purpose Add the secret (Step 11, item 3)

The SAP way

The ideas are the same on SAP's stack; the tools differ by runtime. As of October 2026:

CAP services: cds.test. For CAP on Node.js, cds.test starts your service in the test process on a free port, with an in-memory SQLite database deployed for the run. It gives you GET, POST and the other HTTP methods bound to that server, and a preconfigured Chai expect. It works with node --test, Jest, Mocha and Vitest, and the CAP documentation advises avoiding runner-specific features so tests stay portable. data.reset() redeploys the database between tests. For authorizations, tests act as mock users configured in the project, which is how you test that a user without a role gets refused. This is a sketch of a test for the CAP project from CAP and side-by-side extensions for AI:

// Sketch: run in your CAP project folder with Node.js. Service path, entity and user are examples;
// use the names from your own project and the mock users defined in its package.json.
const cds = require('@sap/cds')
const { GET, expect, defaults } = cds.test(__dirname + '/..')

defaults.auth = { username: 'alice' }   // a mock user from your project's configuration

describe('blocked orders service', () => {
  it('lists blocked orders', async () => {
    const { status, data } = await GET `/odata/v4/blocked/BlockedOrders`
    expect(status).to.equal(200)
    expect(data.value).to.be.an('array')
  })
})

Remote SAP services: mocks from CAP. When a CAP project imports an SAP API such as API_SALES_ORDER_SRV, cds watch mocks it locally with sample data, as shown in the CAP topic. That plays the role of sandbox_page here. The same rule applies: keep a few tests against the real sandbox or a test system to confirm the mock still matches.

Pipelines: SAP Continuous Integration and Delivery. SAP describes this BTP service as pipeline-as-a-service with templates for SAP-specific use cases. Webhooks from your Git repository trigger it, so every change gets automated tests and quality feedback. It is an alternative to the GitHub Actions workflow in Step 11 when a customer wants pipelines inside their BTP account. It is a separate BTP service; check your entitlements.

Quality of AI answers. The checks in this topic prove that answers follow rules. They don't measure how good the answers are across many cases. That is evaluation; Unit 8 covers it, including SAP's tooling for it, with an evaluation harness.

Licensing notes. pytest, unittest.mock, FastAPI's TestClient and cds.test are free. GitHub Actions minutes depend on your GitHub plan. SAP Continuous Integration and Delivery and SAP AI Core usage for real-model checks are paid BTP services.

Build vs. SAP

Situation Build it yourself (pytest, GitHub Actions) SAP tooling
Python AI service, like this unit's API pytest with TestClient and fakes; the natural choice No SAP-specific test tool needed
CAP service (Node.js) Possible with plain HTTP calls, but more setup cds.test with mock users and in-memory database
Stand-in for an SAP API Stubs and unittest.mock, as here CAP mocks imported services with sample data
Proving the stand-in is honest Scheduled tests against the sandbox Same idea, against the sandbox or a customer test system
Pipelines in a personal or partner repository GitHub Actions Optional
Pipelines a customer wants inside BTP Possible, but outside their landscape SAP Continuous Integration and Delivery
Measuring answer quality Your own harness (Unit 8) SAP's evaluation tooling (Unit 8)

Most projects mix both: a CAP app tested with cds.test, a Python AI service tested with pytest, and one pipeline that runs both.

Production concerns

  • Secrets in tests. Fast tests need no secrets. Tests that do, read them from .env locally and from the CI secret store in pipelines. Never print a key in a failure message.
  • Test data protection. Fake data is made up. Never copy production orders, customer names or personal data into fixtures. If you record real responses from a test system, scrub them first.
  • Authorizations. Test the refusals, not only the happy path: missing key, wrong key, and in CAP a mock user without the role. A missing refusal test is how an open endpoint ships.
  • Integration targets. Point integration tests at the sandbox or a dedicated test system with read-only users, never at production. Tests that write data need their own clean-up and a system where that is allowed.
  • Flaky tests. A test that fails at random teaches people to ignore red. Keep randomness and real networks out of the fast tests; give integration tests clear skip and failure messages.
  • Cost control. Real-model tests are opt-in, few and small. Run them before releases and after prompt or model changes, not on every push.
  • What blocks a release. Fast tests and contract tests on your own API should block merges. A failing sandbox contract test is a signal to realign the fakes and the code, and should be investigated the same day.
  • Clean core. Tests run outside S/4HANA and read through released APIs, like the service itself. Nothing here modifies the SAP core.

Pitfalls

  • Asserting exact model wording. The test fails at random and gets ignored. Check rules.
  • Mocks without a spec. A plain Mock() accepts typos and wrong arguments, so a test passes while production crashes. Use spec= or create_autospec.
  • Shared state between tests. One app instance for all tests lets the cache or rate limit from one test change the next. Build a fresh app per test, as make_client does.
  • Mocking your own code. Mock the edges (model, network, SAP), not the logic you want to test.
  • Fakes that drift from reality. Without a contract or integration test, the fake SAP page keeps passing after SAP changes. Use one field list for both.
  • Only testing the happy path. The value is in the timeout, the bad answer, the wrong key and the empty page.
  • Tests that need a key to run at all. Contributors and CI from forks can't run them. Make the fast suite key-free.
  • Never seeing a test fail. Break the code once (Step 10) to prove the test notices.

Exercise: test the batch endpoint and a new rule

This builds on the exercise in Building an AI API, which adds POST /v1/explain-batch to ai_api.py. If you haven't done it, do its steps 1 and 2 first.

  1. Open unit06/tests/test_api.py.
  2. Add a test that sets fake_model.replies = [reply(), reply(next_step="complete_documents")], posts {"sales_orders": ["4711", "4712"]} to /v1/explain-batch, and asserts status 200 and that both items in answers have source equal to "model".
  3. Add a second test where the first reply is ModelUnavailable("down"). Assert that the first item in answers is a fallback and the second comes from the model.
  4. Add a test that sends six order numbers and asserts status 422 (the batch allows at most five).
  5. In test_model_output.py, add "guaranteed" to FORBIDDEN. Extend test_rule_checker_catches_a_bad_answer with an answer containing "guaranteed" and assert the checker reports it.
  6. Run pytest unit06/tests -m "not integration and not llm".
  7. Commit on a branch, push, and open a pull request. Wait for fast-tests to go green on GitHub, then merge.

Done when: pytest unit06/tests -m "not integration and not llm" reports at least 58 passed and 0 failed on your computer, and the pull request on GitHub shows a green fast-tests check. Unit 8 reuses rule_problems as the first, rule-based metric of its evaluation harness.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does create_app(model, ...) in ai_api.py make the API easy to test?

    Answer: C. Passing the model in as an argument is a seam: production passes the real model, tests pass FakeModel or a mock. Without it, a test would have to patch the code or call the paid model. The cache and .env don't give tests that control.
  2. 2A teammate writes model = Mock() and the test passes, but production crashes because the code calls model.anwser(...). What would have caught it?

    Answer: B. A plain Mock() accepts any attribute, including typos. create_autospec only allows methods the real class has and checks their arguments, so the misspelled call fails in the test. -v, markers and return_value don't restrict what the mock accepts.
  3. 3How does test_bad_input_is_422_and_never_reaches_the_model prove input checks save money?

    Answer: D. FakeModel records every call. An empty calls list after the 422 answers shows validation stopped the request before any model call, which is where the cost would be. Time, tokens in the error and the rate limit don't prove the model was skipped.
  4. 4The model sometimes writes "Order 4711 is on hold" and sometimes "Sales order 4711 cannot ship yet". How should a test check its answers?

    Answer: C. Both wordings are acceptable, so exact comparison fails at random and gets ignored. Rule checks accept any wording that follows the rules and catch the answers that break them. The checker is itself tested with a known-bad answer so it can't silently pass everything.
  5. 5Why is the same FIELDS list used by sap_orders.py, the fake SAP page check and the sandbox integration test?

    Answer: B. The code only depends on those fields. The contract test checks the fake page has them, and the integration test checks the real sandbox still sends them. If SAP changed, the daily sandbox run fails and tells you to update the code and the fake.
  6. 6What would you do if the daily sandbox-tests job fails with SAP no longer sends {'HeaderBillingBlockReason'} while fast-tests stay green?

    Answer: D. The fast tests use the fake, so they stay green even though the real API changed; that is exactly the gap the contract test covers. Treat it as a signal to realign the code and the fake with the real service. Deleting the test or editing only the fake hides the problem.
  7. 7Why does the workflow copy secrets.SAP_API_KEY into a job variable before the if: check?

    Answer: B. GitHub's documentation says secrets can't be referenced directly in if: conditionals and suggests setting them as environment variables first. The step then checks env.SAP_API_KEY and skips itself when the secret is missing.
  8. 8Which tests should block a merge on every push for this project?

    Answer: A. Fast, key-free tests run in seconds on every push, including from forks that get no secrets. Sandbox tests depend on an external service and run daily; real-model checks cost money and run before releases. Making them block every push would make merges slow, costly and flaky.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in