Orchestrate

Deploying AI apps on SAP BTP

Move an AI service from your laptop to Cloud Foundry on SAP BTP, with keys kept out of code, health checks, logs and updates that don't take it offline.

Updated Oct 2, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An AI service that runs on one laptop helps one person. To help a whole order desk it has to run somewhere managed: always on, reachable over a secure address, and looked after when it crashes. Moving it there is deployment.

On SAP BTP, the usual place is the Cloud Foundry environment. A developer hands over the code with one command. The platform installs what the code needs, starts it in a container, gives it a web address and keeps checking that it answers. If it stops answering, the platform restarts it.

Deployment changes more than the address. Passwords and keys can no longer sit in a file on someone's laptop. Several copies of the app may run at once. Logs have to go somewhere people can read them. Updates have to happen without switching the service off. This topic covers each of those, and the learner deploys the course's AI API to a free BTP trial.

Why it matters to the business

A pilot that only runs on a developer's machine is a demo. Value starts when clerks rely on it at 8 a.m., in every country, without calling the developer. That step is often where AI projects slip, because nobody planned who runs the service.

Take the course's running example: blocked sales orders in order-to-cash. Earlier topics in this unit built an AI API that explains why an order is on hold, and a CAP extension that clerks use. Once the order desk depends on them:

  • Downtime becomes visible. If the explanation service is down, clerks fall back to phone calls and email. The platform's health checks and automatic restarts reduce that, and running two copies avoids a single point of failure.
  • Keys become a security issue. The service holds a key that callers must present, and later the credentials for the model. On a shared platform those must be kept out of code and out of settings that many people can read.
  • Changes need a safe path. Prompts and models change often. A rolling update replaces the running copies one at a time, so clerks keep getting answers while the new version starts.
  • Cost becomes a line item. Each copy reserves memory in the BTP account, and model calls are billed separately. Someone must own both.

Getting this right turns a successful pilot into a service the business can depend on. Getting it wrong usually shows up as a leaked key, a service that silently stopped, or an update that broke every caller at once.

How SAP does it

As of October 2026, SAP BTP offers three runtime environments, as SAP's learning material describes them:

  • Cloud Foundry environment. Supports many languages and runtimes, including Java, Node.js and Python. You push code; buildpacks turn it into a running app. This topic uses it, because the free trial includes it and it asks the least of a beginner.
  • Kyma runtime. A fully managed Kubernetes runtime. Teams that already package apps as containers, or need more control over how they run, often choose it.
  • ABAP environment. For ABAP-based extensions and cloud apps, for example extensions to SAP S/4HANA Cloud.

Around the runtime, SAP provides the managed pieces a real app needs. CAP's deployment guide adds SAP HANA Cloud for data and XSUAA for user sign-in, and packages the whole app as one multitarget application (MTA) that a single command deploys. SAP AI Core credentials reach the app through a service binding rather than a file. SAP's learning material recommends central logging, such as SAP Cloud Logging, for monitoring.

The learner's free BTP trial is enough to practise. Set up for Unit 6 noted its limits: it is not for production or team use, it ends, and its apps stop automatically every day.

A go-live checklist: laptop, trial, production

Use this table to see how far a pilot is from being something the business can depend on.

Concern On a laptop On the BTP trial (this topic) In production
Who can reach it Only that computer Anyone with the address and the key Named users and systems, through sign-in
Where keys live A .env file A bound service, not in code A bound service, rotated on a schedule
How callers prove who they are An API key An API key Company sign-in (XSUAA or Identity Authentication)
Copies running One One, then two in the exercise At least two
If it crashes Someone notices The platform restarts it Restart plus alerts to an on-call owner
Logs The terminal cf logs Central logging with retention rules
Updates Stop and start Rolling update Rolling update from a pipeline, with tests
Model Sample model Sample model SAP AI Core through a service binding
Lifetime While the laptop is on Stops daily; trial expires Owned, budgeted, monitored

Questions to ask

  • Which BTP environment will this run in, Cloud Foundry or Kyma, and who in our company owns that subaccount?
  • Where are the keys and passwords kept, and who can read them? Are any of them in code, manifests or plain environment variables?
  • How do users and other systems sign in? Is it an API key, or company sign-in with roles?
  • How many copies run, and what happens to callers if one of them stops?
  • How do we update it without downtime, and how do we roll back a bad version?
  • Where do logs go, how long are they kept, and do they contain customer data?
  • Who gets called when it breaks at night, and how will they know?
  • Which BTP entitlements does this need (runtime memory, SAP HANA Cloud, SAP AI Core), and are they in our contract?

Common misconceptions

  • "It runs on my laptop, so deployment is a formality." Deployment changes where keys live, how many copies run, and how updates happen. Each of those can break a working app.
  • "The trial is fine for the pilot." SAP rules out productive and team use of the trial, and its apps stop daily. A pilot that users depend on needs a proper account.
  • "Putting the key in an environment variable is safe." Cloud Foundry's own documentation warns against it: such values can show up in command output and platform logs. Use a bound service instead.
  • "More copies means the same behaviour, only faster." Anything an app keeps in memory, such as a cache or a rate-limit counter, is kept per copy. Design for that before scaling out.
  • "An update is a restart." A plain restart takes the service down. A rolling update keeps old copies serving until new ones are healthy.

Key terms

  • Deployment: moving an app to the environment where it will run for its users.
  • Cloud Foundry: an open-source platform for running apps, offered as an environment in SAP BTP. You push code; it runs it.
  • Kyma runtime: SAP's managed Kubernetes runtime on BTP, for container-based apps.
  • Buildpack: the platform tool that recognizes your language and installs what the app needs.
  • Manifest: a small file that tells Cloud Foundry how to run the app: name, memory, start command, health check.
  • Route: the web address the platform gives the app.
  • Health check: a regular test the platform runs to see if the app still answers. A failed check leads to a restart.
  • Service binding: how the platform hands an app the credentials of a service, such as a database or SAP AI Core, without putting them in code.
  • Rolling update: replacing running copies one at a time, so the service stays available.
  • MTA (multitarget application): SAP's way to package an app with several parts, such as a service, a database and sign-in, and deploy them together.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A sponsor says the AI pilot "works on the developer's laptop, so it's ready for the order desk". What is the main gap?

    Answer: B. Deployment is where keys move out of local files, copies are restarted and scaled by the platform, and updates must keep the service up. None of that exists on a laptop. Retraining or rewriting in ABAP isn't what deployment needs.
  2. 2Your team wants to run the order-desk pilot on a developer's BTP trial account. What is the problem?

    Answer: C. The trial has a Cloud Foundry org and space, which is enough to learn deployment. SAP doesn't allow productive or team use, and the apps stop every day for cleanup, so a pilot people depend on needs a proper account.
  3. 3A developer proposes storing the model credentials as a plain Cloud Foundry environment variable. What should you ask for instead?

    Answer: D. Cloud Foundry's documentation warns that plain environment variables can show up in command output and platform logs. Credentials belong in a service instance bound to the app, never in code or the manifest.
  4. 4The order desk needs the explanation service to keep answering during updates. What should the team use?

    Answer: B. A rolling update starts new copies and removes old ones only once the new ones are healthy, so callers keep getting answers. A restart takes the service down, and a second trial account doesn't help production.
  5. 5Which question best tests whether a partner is ready to run your AI service in production?

    Answer: D. Production readiness is about operations: ownership, alerts, rollback, keys and sign-in. Benchmark scores and code size say nothing about whether the service will stay up for the order desk.
  6. 6When is the Kyma runtime a more natural choice than Cloud Foundry on SAP BTP?

    Answer: C. Kyma is SAP's fully managed Kubernetes runtime, a natural home for container-based apps. ABAP extensions go to the ABAP environment, and sign-in needs exist in every runtime.
Deep layer · 40 min read

Mental model: build once, configure per environment, let the platform run it

Deployment hands three jobs to the platform that you did on your laptop by hand: starting the app, configuring it, and watching it.

  • Starting. On your laptop you typed python unit06/ai_api.py. On Cloud Foundry you upload the code once; a buildpack installs the libraries and produces a runnable image, and the platform starts it in a container. You tell it how in a manifest.
  • Configuring. On your laptop the key sat in .env. In the cloud the same code reads its settings from the platform: the port in PORT, credentials from bound services in VCAP_SERVICES. The code stays identical across laptop, test and production; only the configuration changes.
  • Watching. On your laptop you noticed when the app crashed. The platform runs a health check and restarts the app when it fails. It can run several copies and replace them one at a time during updates.

Once you see deployment this way, most production problems sort themselves into one of the three: it didn't start, it was configured wrong, or nobody noticed it failing.

How it works

From cf push to a running app

sequenceDiagram
  participant D as Your laptop
  participant CF as Cloud Foundry
  participant B as Python buildpack
  participant C as Container
  D->>CF: cf push (code + manifest)
  CF->>B: stage: requirements.txt, runtime.txt
  B-->>CF: droplet (app + Python + libraries)
  CF->>C: start "python cf_app.py", PORT, VCAP_SERVICES
  CF->>C: health check GET /health
  C-->>CF: 200 OK
  CF-->>D: running, route assigned
  1. Upload. cf push reads manifest.yml and uploads the folder, minus anything listed in .cfignore.
  2. Staging. The Python buildpack recognizes the app by its requirements.txt. It installs the Python version named in runtime.txt and the listed libraries. The result is a droplet: your code plus everything it needs.
  3. Start. The platform starts the droplet in a container with the manifest's command. It sets PORT, which the app must listen on, and passes service credentials in VCAP_SERVICES.
  4. Health check. The manifest attribute health-check-type can be port (the default), process or http. With http, the platform calls health-check-http-endpoint (for example /health) and expects a success code. If the check fails, the instance is restarted.
  5. Route. The app gets a web address. random-route: true gives the app a random host name, so it doesn't clash with other apps on the same domain.

Configuration: environment variables vs. bound services

Cloud Foundry gives an app two kinds of settings:

Kind Set with Seen in the app as Use it for
User-provided environment variable env: in the manifest, or cf set-env Its own variable, e.g. APP_VERSION Non-secret settings: version labels, feature switches, timeouts
Bound service instance cf create-service or cf create-user-provided-service, then services: in the manifest An entry in VCAP_SERVICES, with credentials Secrets and connections: keys, database, SAP AI Core

Cloud Foundry's documentation is explicit: don't use user-provided environment variables for credentials, because they may appear in command output and platform logs. A user-provided service is the simple alternative when there is no managed service for a secret, such as your own API key. Changes to either take effect only after the app restarts or restages.

Instances, memory and state

The manifest's memory reserves memory per instance; instances sets how many copies run. Two instances of a 256 MB app reserve 512 MB of the account's quota.

Each instance is a separate process with its own memory. Your AI API keeps its cache and rate-limit counters in memory, so with two instances each has its own cache and each allows its own 10 calls per minute. SAP's learning material also tells you to avoid writing files to the container's file system, because it is short-lived. Anything that must survive a restart or be shared between instances belongs in a backing service, such as a database or a cache service.

Updates without downtime

A plain cf push of a new version stops the old one first. cf push --strategy rolling starts new instances, waits until they are healthy, then removes old ones, until all are replaced. Two consequences from Cloud Foundry's documentation:

  • During the update, old and new versions answer at the same route. Changes to your API contract must be backward compatible, or callers get mixed answers.
  • The update needs spare quota for the extra instances, and it doesn't migrate databases for you. cf cancel-deployment stops a rolling update, without a zero-downtime guarantee.

Logs

Whatever the app writes to the terminal (standard output and standard error) becomes its log. cf logs APP --recent shows recent lines; cf logs APP streams them live. Your AI API already writes one line per request with a request ID and no customer data. That is exactly what you want in a platform log.

Multi-part apps: MTAs

One Python app pushes cleanly with cf push. A CAP app with a database, sign-in and a user interface has several parts that must be created and connected in the right order. SAP's answer is the multitarget application: an mta.yaml file that lists the modules and the services they need, built into one archive and deployed with cf deploy. You'll see this in "The SAP way" below.

Build it yourself: deploy the AI API to your BTP trial

You will deploy the AI API from Building an AI API to Cloud Foundry in your free BTP trial. You'll add a small start file that reads its settings from the platform, a manifest, and a key kept in a bound service. You'll first run it on your computer exactly the way Cloud Foundry will, then push it, test it over the internet, and read its logs.

flowchart LR
  P[predeploy_check.py] --> F[unit06/cf-ai-api<br/>cf_app.py + ai_api.py<br/>manifest.yml]
  F -->|Step 4: local, port 8080| L[Your computer]
  F -->|Step 6: cf push| CF[Cloud Foundry<br/>BTP trial, dev space]
  K[orchestrate-ai-api-key<br/>user-provided service] -->|VCAP_SERVICES| CF
  S[smoke_test.py] --> L
  S -->|HTTPS + X-API-Key| CF

Before you start: complete Set up your computer for this course and Set up for Unit 6, including Step 5 (logging in with cf). You also need unit06/ai_api.py and the ORCHESTRATE_API_KEY line in .env from Building an AI API. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder with the Unit 6 setup done (python check_unit06.py ends with All set), and unit06/ai_api.py from the AI API topic.
  • Your SAP BTP trial with its Cloud Foundry dev space. Steps 5 to 8 need it. Without a trial, Steps 1 to 4 still work.
  • About 60 to 90 minutes.
  • Cost: free. The app uses the sample model and 256 MB of the trial's memory. No model calls are made.

Step 1: Open your course folder and check the tools

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. Turn on the virtual environment if the prompt doesn't start with (.venv):

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that the AI API is there and that cf is logged in (the same on every system):

    python -c "import fastapi, dotenv; print('ready')"
    cf target

    What success looks like: ready, then lines with your API endpoint, org and space: dev. If cf target says you are not logged in, run cf login as in Step 5 of the Unit 6 setup. You can do Steps 2 to 4 without logging in.

Run every command in this topic from the course folder unless a step says otherwise.

Step 2: Create the deployment folder and the start file

The deployment gets its own folder, so you upload only what the app needs, not your whole course.

  1. Create the folder and copy the AI API into it:

    • Windows (PowerShell):

      New-Item -ItemType Directory -Force unit06\cf-ai-api
      Copy-Item unit06\ai_api.py unit06\cf-ai-api\ai_api.py
    • macOS / Linux:

      mkdir -p unit06/cf-ai-api
      cp unit06/ai_api.py unit06/cf-ai-api/ai_api.py
  2. In VS Code's file list, right-click unit06/cf-ai-api, choose New File, name it cf_app.py, paste this and save:

"""Start the Unit 6 AI API the way Cloud Foundry expects, on SAP BTP or on your computer.

Cloud Foundry tells the app which port to use in the PORT variable, and hands it the API key
through a bound service (VCAP_SERVICES). On your computer the same file reads .env instead.

How to run on your computer (from your course folder, with .venv turned on):
    python unit06/cf-ai-api/cf_app.py                 # listens on http://127.0.0.1:8080
On Cloud Foundry, manifest.yml runs:  python cf_app.py
"""
import json
import logging
import os
import sys

import uvicorn
from fastapi import FastAPI

from ai_api import PROMPT_VERSION, SampleModel, create_app

SERVICE_NAME = "orchestrate-ai-api-key"   # the user-provided service that holds the key on BTP
ON_CLOUD_FOUNDRY = "VCAP_APPLICATION" in os.environ


def find_api_key() -> tuple[str, str]:
    """Return the API key and where it came from. Never print the key itself."""
    services = json.loads(os.environ.get("VCAP_SERVICES", "{}"))
    for instances in services.values():
        for instance in instances:
            if SERVICE_NAME in (instance.get("name"), instance.get("instance_name")):
                key = instance.get("credentials", {}).get("api_key", "")
                if key:
                    return key, f"service '{SERVICE_NAME}'"
    if not ON_CLOUD_FOUNDRY:   # on your computer: read .env if python-dotenv is installed
        try:
            from dotenv import load_dotenv
            load_dotenv()
        except ImportError:
            pass
    return os.environ.get("ORCHESTRATE_API_KEY", ""), "ORCHESTRATE_API_KEY"


def build_app() -> FastAPI:
    api_key, key_source = find_api_key()
    if len(api_key) < 20:
        sys.exit(f"No API key of 20+ characters found (looked in {key_source}). "
                 f"On BTP: create and bind the service '{SERVICE_NAME}'. Locally: check .env.")
    model = SampleModel()
    app = create_app(model, api_key)

    @app.get("/about")
    def about() -> dict:
        """Which version and which instance answered. No key needed, no customer data."""
        return {"app_version": os.environ.get("APP_VERSION", "dev"),
                "instance": os.environ.get("CF_INSTANCE_INDEX", "local"),
                "model": model.name, "prompt_version": PROMPT_VERSION}

    logging.getLogger("ai_api").info("key from %s, model %s, version %s", key_source, model.name,
                                     os.environ.get("APP_VERSION", "dev"))
    return app


def main() -> None:
    # Cloud Foundry logs whatever the app writes to the terminal (stdout and stderr).
    logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s %(message)s", stream=sys.stdout)
    app = build_app()
    port = int(os.environ.get("PORT", "8080"))
    host = "0.0.0.0" if ON_CLOUD_FOUNDRY else "127.0.0.1"   # in the cloud, listen on the container's network
    print(f"Starting on {host}:{port} (Cloud Foundry: {ON_CLOUD_FOUNDRY})", flush=True)
    uvicorn.run(app, host=host, port=port, access_log=False)


if __name__ == "__main__":
    main()

Your ai_api.py stays unchanged. cf_app.py only decides where the key and port come from, adds an /about endpoint that says which version answered, and starts the server.

Step 3: Tell Cloud Foundry how to run it

You'll create four small files in unit06/cf-ai-api. For each: right-click the folder, choose New File, type the name exactly (including the leading dot in .cfignore), paste the content and save.

  1. requirements.txt: the libraries the cloud must install, with exact versions so the cloud runs what you tested.

    fastapi==0.142.2
    uvicorn==0.54.0

    These are the versions we tested on 2 October 2026. To use the versions in your own .venv instead, run pip show fastapi uvicorn and copy the two Version: numbers.

  2. runtime.txt: the Python version for the cloud. The buildpack reads this file; 3.13.x means "the newest 3.13 it has". SAP's own Python tutorial uses the same line.

    python-3.13.x

    Your laptop may run Python 3.14. That's fine: the app uses nothing that differs between the two.

  3. manifest.yml: how to run the app. Indentation matters in YAML; use spaces, not tabs.

    ---
    applications:
    - name: orchestrate-ai-api
      path: .
      memory: 256M
      instances: 1
      random-route: true
      buildpacks:
      - python_buildpack
      command: python cf_app.py
      health-check-type: http
      health-check-http-endpoint: /health
      env:
        APP_VERSION: "1.0.0"
      services:
      - orchestrate-ai-api-key
  4. .cfignore: files that must never be uploaded.

    .env
    .venv/
    __pycache__/
    *.pyc
  5. Save this check as unit06/predeploy_check.py (in unit06, not in cf-ai-api). It reads the files and tells you what is missing. It uses built-in Python only.

"""Unit 6: check the cf-ai-api folder before you push it to Cloud Foundry.

It only reads files. It needs no account, no network and no extra libraries.
How to run (from your course folder):
    python unit06/predeploy_check.py
"""
import filecmp
import re
import sys
from pathlib import Path

HERE = Path(__file__).resolve().parent          # the unit06 folder
APP = HERE / "cf-ai-api"
problems = 0


def report(ok: bool, label: str, fix: str = "", warn_only: bool = False) -> None:
    global problems
    if ok:
        print(f"  OK       {label}")
    elif warn_only:
        print(f"  WARN     {label} -> {fix}")
    else:
        problems += 1
        print(f"  MISSING  {label} -> {fix}")


def read(name: str) -> str:
    path = APP / name
    return path.read_text(encoding="utf-8") if path.exists() else ""


print(f"Checking {APP}\n")
report(APP.is_dir(), "folder unit06/cf-ai-api", "create it (Step 2)")
for name in ["ai_api.py", "cf_app.py", "requirements.txt", "runtime.txt", "manifest.yml", ".cfignore"]:
    report((APP / name).exists(), name, f"create {name} (Step 2 or 3)")

manifest = read("manifest.yml")
for needed, why in [("name: orchestrate-ai-api", "the app name"),
                    ("health-check-type: http", "an HTTP health check"),
                    ("health-check-http-endpoint: /health", "the /health endpoint"),
                    ("command: python cf_app.py", "the start command"),
                    ("- orchestrate-ai-api-key", "the key service binding")]:
    report(needed in manifest, f"manifest.yml has {why}", f"add the line '{needed}' (Step 3)")
secret_like = re.findall(r"(?i)(api_key|password|secret|token)\s*:", manifest)
report(not secret_like, "manifest.yml holds no secrets", "remove keys and passwords; use the service (Step 5)")

runtime = read("runtime.txt").strip()
report(bool(re.fullmatch(r"python-3\.\d+\.(x|\d+)", runtime)), f"runtime.txt pins a Python version ({runtime or 'empty'})",
       "write one line such as python-3.13.x")

lines = [line.strip() for line in read("requirements.txt").splitlines() if line.strip() and not line.startswith("#")]
unpinned = [line for line in lines if "==" not in line]
report(bool(lines) and not unpinned, "requirements.txt pins every library",
       "use name==version for: " + ", ".join(unpinned or ["fastapi, uvicorn"]))
report(not any(line.lower().startswith("python-dotenv") for line in lines), "requirements.txt leaves out python-dotenv",
       "the cloud reads settings from the platform, not .env", warn_only=True)

cfignore = read(".cfignore")
report(".env" in cfignore.split(), ".cfignore keeps .env out of the upload", "add a line .env")
report(not (APP / ".env").exists(), "no .env file inside cf-ai-api", "delete it; keys go in the service (Step 5)")

app_code = read("cf_app.py")
report("PORT" in app_code and "0.0.0.0" in app_code, "cf_app.py listens on PORT on all interfaces", "copy cf_app.py again (Step 2)")

original = HERE / "ai_api.py"
same = original.exists() and (APP / "ai_api.py").exists() and filecmp.cmp(original, APP / "ai_api.py", shallow=False)
report(same, "ai_api.py matches unit06/ai_api.py", "copy it again so you deploy what you tested (Step 2)", warn_only=True)

print()
if problems:
    sys.exit(f"Fix {problems} item(s) above, then run this check again.")
print("Ready to push.")
  1. Run it (the same on every system):

    python unit06/predeploy_check.py

What success looks like (shortened):

  OK       folder unit06/cf-ai-api
  OK       manifest.yml has an HTTP health check
  OK       manifest.yml holds no secrets
  OK       runtime.txt pins a Python version (python-3.13.x)
  OK       requirements.txt pins every library
  OK       .cfignore keeps .env out of the upload
  OK       ai_api.py matches unit06/ai_api.py

Ready to push.

Any MISSING line names the step to go back to. A WARN line won't stop you, but read it.

Step 4: Run it the cloud way on your computer (no account needed)

Before you push, run the app exactly as Cloud Foundry will: through cf_app.py, on port 8080. Then test it with a smoke test you'll reuse against the cloud.

  1. Save this as unit06/smoke_test.py:
"""Unit 6: smoke-test the AI API after a deployment, on BTP or on your computer.

It calls a few endpoints and checks the answers the API promises. Built-in Python plus python-dotenv.
How to run (from your course folder, with .venv turned on):
    python unit06/smoke_test.py http://127.0.0.1:8080                    # the local copy, key from ORCHESTRATE_API_KEY
    python unit06/smoke_test.py https://YOUR-ROUTE --key-env ORCHESTRATE_CLOUD_API_KEY
    python unit06/smoke_test.py https://YOUR-ROUTE --watch 120           # ask /about every second for 120 seconds
"""
import argparse
import json
import os
import sys
import time
import urllib.error
import urllib.request
from urllib.parse import urlparse

from dotenv import load_dotenv


def call(base: str, path: str, method: str = "GET", body: dict | None = None, key: str | None = None):
    """Send one request. Return (status, parsed JSON or text, headers)."""
    data = json.dumps(body).encode() if body is not None else None
    request = urllib.request.Request(base + path, data=data, method=method)
    request.add_header("Content-Type", "application/json")
    if key:
        request.add_header("X-API-Key", key)
    try:
        with urllib.request.urlopen(request, timeout=20) as response:
            status, raw, headers = response.status, response.read().decode(), response.headers
    except urllib.error.HTTPError as error:
        status, raw, headers = error.code, error.read().decode(), error.headers
    try:
        return status, json.loads(raw), headers
    except json.JSONDecodeError:
        return status, raw, headers


def check(label: str, ok: bool, detail=None) -> bool:
    print(f"  {'PASS' if ok else 'FAIL'}  {label}" + ("" if ok else f"   (got: {str(detail)[:120]})"))
    return ok


def smoke(base: str, key: str) -> bool:
    results = []
    status, body, _ = call(base, "/health")
    results.append(check("GET /health answers 200 without a key", status == 200 and body.get("status") == "ok", body))
    status, body, _ = call(base, "/about")
    results.append(check("GET /about names the version", status == 200 and "app_version" in body, body))
    if status == 200:
        print(f"        version {body['app_version']}, instance {body['instance']}, model {body['model']}")
    status, body, _ = call(base, "/v1/explain", "POST", {"sales_order": "4711"})
    results.append(check("POST /v1/explain without a key -> 401", status == 401, status))
    status, body, headers = call(base, "/v1/explain", "POST", {"sales_order": "4711"}, key)
    results.append(check("POST /v1/explain 4711 with the key -> 200", status == 200 and body.get("source") in ("model", "cache"), body))
    results.append(check("the answer carries a request ID", bool(headers.get("X-Request-ID")), dict(headers)))
    status, body, _ = call(base, "/v1/explain", "POST", {"sales_order": "4713"}, key)
    results.append(check("a model failure still gives a labeled fallback", status == 200 and body.get("source") == "fallback", body))
    if urlparse(base).hostname not in ("127.0.0.1", "localhost"):
        results.append(check("the route uses HTTPS", base.startswith("https://"), base))
    print(f"\n{sum(results)} of {len(results)} passed")
    return all(results)


def watch(base: str, seconds: int) -> None:
    """Ask /about once a second and print what answered, to watch a rolling deployment."""
    failures, seen, end = 0, set(), time.monotonic() + seconds
    while time.monotonic() < end:
        try:
            status, body, _ = call(base, "/about")
            line = f"{body['app_version']} (instance {body['instance']})" if status == 200 else f"HTTP {status}"
        except Exception as error:   # connection refused, timeout, DNS
            status, line = 0, f"no answer: {type(error).__name__}"
        if status != 200:
            failures += 1
        else:
            seen.add(body["app_version"])
        print(time.strftime("%H:%M:%S"), line, flush=True)
        time.sleep(1)
    print(f"\nVersions seen: {', '.join(sorted(seen)) or 'none'}. Failed requests: {failures}.")


def main() -> None:
    parser = argparse.ArgumentParser(description="Smoke-test the Unit 6 AI API.")
    parser.add_argument("url", help="the app's address, e.g. https://orchestrate-ai-api-....hana.ondemand.com")
    parser.add_argument("--key-env", default="ORCHESTRATE_API_KEY", help="name of the .env line that holds the key")
    parser.add_argument("--watch", type=int, metavar="SECONDS", help="only poll /about for this many seconds")
    args = parser.parse_args()
    base = args.url.rstrip("/")
    if not base.startswith(("http://", "https://")):
        base = "https://" + base   # cf prints routes without https://
    if args.watch:
        watch(base, args.watch)
        return
    load_dotenv()
    key = os.environ.get(args.key_env, "")
    if not key:
        sys.exit(f"{args.key_env} is not set in .env. See Step 5.")
    try:
        ok = smoke(base, key)
    except urllib.error.URLError as error:
        sys.exit(f"Could not reach {base}: {error.reason}. Is the app running? Check the address.")
    sys.exit(0 if ok else 1)


if __name__ == "__main__":
    main()
  1. Start the app the cloud way (the same on every system):

    python unit06/cf-ai-api/cf_app.py

    What success looks like:

    INFO ai_api key from ORCHESTRATE_API_KEY, model sample-model, version dev
    Starting on 127.0.0.1:8080 (Cloud Foundry: False)
    INFO:     Uvicorn running on http://127.0.0.1:8080 (Press CTRL+C to quit)
  2. Open a second terminal (Terminal > New Terminal), turn on .venv as in Step 1, and run the smoke test:

    python unit06/smoke_test.py http://127.0.0.1:8080

What success looks like:

  PASS  GET /health answers 200 without a key
  PASS  GET /about names the version
        version dev, instance local, model sample-model
  PASS  POST /v1/explain without a key -> 401
  PASS  POST /v1/explain 4711 with the key -> 200
  PASS  the answer carries a request ID
  PASS  a model failure still gives a labeled fallback

6 of 6 passed

Order 4713 is the sample model's planned failure, so a labeled fallback is the right answer there. Stop the app in the first terminal with Ctrl+C.

Step 5: Create a cloud key and keep it in a service

The cloud copy gets its own key. If one leaks, you replace only that one.

  1. Print a new random key (the same on every system):

    python -c "import secrets; print(secrets.token_urlsafe(32))"
  2. Open .env in the course folder. Add this line at the end, with the new key between the quotes, and save:

    ORCHESTRATE_CLOUD_API_KEY="paste-the-new-key-here"
  3. Create the user-provided service. The -p "api_key" form makes cf ask for the value, so the key doesn't land in your terminal history (the same on every system):

    cf create-user-provided-service orchestrate-ai-api-key -p "api_key"

    When cf asks for api_key, paste the new key and press Enter. You should see a line ending in OK.

  4. Check that it exists:

    cf services

    You should see orchestrate-ai-api-key in the list, with user-provided as its offering.

Step 6: Push the app

  1. Run the check again; it must end with Ready to push.:

    python unit06/predeploy_check.py
  2. Go into the deployment folder and push:

    • Windows (PowerShell):

      cd unit06\cf-ai-api
      cf push
    • macOS / Linux:

      cd unit06/cf-ai-api
      cf push

    cf push finds manifest.yml in the folder. Staging takes a few minutes the first time, while the buildpack installs Python and the libraries. You'll see many lines scroll by.

What success looks like: the output ends with a summary in this shape (your route, dates and numbers differ):

name:              orchestrate-ai-api
requested state:   started
routes:            orchestrate-ai-api-RANDOM-PART.cfapps.YOUR-REGION.hana.ondemand.com
...
     state     since                  cpu    memory
#0   running   2026-10-02T15:40:12Z   0.0%   ...
  1. Copy the address after routes:. That is your app's route.

  2. Go back to the course folder:

    cd ../..

Step 7: Test it over the internet and read its logs

  1. Run the smoke test against the route, with the cloud key. Replace YOUR-ROUTE with the address you copied (the same on every system):

    python unit06/smoke_test.py https://YOUR-ROUTE --key-env ORCHESTRATE_CLOUD_API_KEY

    What success looks like: the same six PASS lines as in Step 4, plus PASS the route uses HTTPS, ending in 7 of 7 passed. The /about line now says version 1.0.0, instance 0.

  2. Read the app's recent logs:

    cf logs orchestrate-ai-api --recent

    Among the platform's lines you should find your app's own, tagged APP/PROC/WEB/0:

    ... [APP/PROC/WEB/0] OUT INFO ai_api key from service 'orchestrate-ai-api-key', model sample-model, version 1.0.0
    ... [APP/PROC/WEB/0] OUT INFO ai_api 0c1f... POST /v1/explain 401 1ms
    ... [APP/PROC/WEB/0] OUT WARNING ai_api 9d3a... model unavailable: sample model: simulated timeout

    The key itself never appears, and neither does any order data. Each line has a request ID you can match with what the caller saw.

  3. Look at the app's state:

    cf app orchestrate-ai-api

    You see one instance, running, and its memory use against the 256 MB you reserved.

Step 8: Stop it when you're done

The trial stops apps daily anyway, but stop yours when you finish so it uses no quota:

cf stop orchestrate-ai-api

To start it again later: cf start orchestrate-ai-api. You'll need it running for the exercise.

What the code does

Part What it does
cf_app.py, find_api_key Looks for the key in VCAP_SERVICES under the service name; on your computer it falls back to .env
ON_CLOUD_FOUNDRY Cloud Foundry always sets VCAP_APPLICATION, so its presence tells the code where it runs
host = "0.0.0.0" In the container the app must accept connections from the platform's router, not only from itself
PORT The port Cloud Foundry chose; 8080 when it isn't set
/about Shows the version and instance number, so you can see which copy answered
requirements.txt, runtime.txt Tell the buildpack exactly what to install
manifest.yml Name, memory, start command, HTTP health check on /health, a non-secret version label, and the key service
.cfignore Keeps .env and local caches out of the upload
predeploy_check.py Catches missing files, secrets in the manifest and unpinned libraries before you push
smoke_test.py Checks health, the 401 without a key, a good answer, a request ID, the fallback and HTTPS

If something goes wrong

What you see What it means What to do
python: command not found or 'python' is not recognized Python isn't on PATH, or .venv is off Turn on .venv (Step 1); on macOS/Linux try python3
ModuleNotFoundError: No module named 'dotenv' or 'fastapi' A library is missing in .venv Check for (.venv) in the prompt, then pip install -r requirements.txt in the course folder
No API key of 20+ characters found locally ORCHESTRATE_API_KEY is missing in .env Add it as in the AI API topic, Step 2
cf target: Not logged in The cf session expired Run cf login again (Unit 6 setup, Step 5)
cf push: Service instance orchestrate-ai-api-key not found Step 5 wasn't done, or the name differs Run cf services; create the service with the exact name
Staging fails with Unsupported Python version or similar The platform's buildpack doesn't have that Python Run cf buildpacks to see what is installed; change runtime.txt, e.g. python-3.12.x
Staging fails during pip install A library name or version is wrong, or blocked Check requirements.txt against pip show fastapi uvicorn in .venv
Instances starting... then crashed The app exited at start Run cf logs orchestrate-ai-api --recent; No API key means the service isn't bound or has no api_key field
Health check failed, app keeps restarting The app doesn't answer /health on PORT Check cf_app.py is the Step 2 version; predeploy_check.py tests this
insufficient resources or a memory quota error The trial's memory is used by other apps Stop apps you don't need: cf apps, then cf stop NAME
Smoke test: FAIL ... with the key and 401 The key in .env differs from the one in the service Recreate it: cf update-user-provided-service orchestrate-ai-api-key -p "api_key", then cf restart orchestrate-ai-api
Smoke test: Could not reach or timed out Wrong address, app stopped, or a network or proxy blocks it Check cf app orchestrate-ai-api; start it; try another network or ask IT about the proxy
CERTIFICATE_VERIFY_FAILED A company proxy inspects HTTPS traffic Ask IT for the proxy's certificate setup for Python

The SAP way

You pushed one Python app with cf push. As of October 2026, here is how the same ideas map to SAP's managed pieces on BTP.

Concern In this topic SAP's way on BTP
Runtime Cloud Foundry, Python buildpack Cloud Foundry (Java, Node.js, Python), or the Kyma runtime for containers
Packaging One app, cf push An MTA (mta.yaml) for apps with several parts, deployed with cf deploy
Callers sign in API key in a user-provided service XSUAA with roles; an application router in front for browser users
Model credentials None (sample model) A bound SAP AI Core instance; the SDKs read VCAP_SERVICES
Database None; cache in memory SAP HANA Cloud, bound to the app
S/4HANA connection None The Destination service, as in CAP and side-by-side extensions
Logs cf logs Central logging such as SAP Cloud Logging

Deploying the CAP extension. CAP's deployment guide describes the path for an app like order-assist from the previous topic. It needs an SAP HANA Cloud instance in the subaccount and the hdi-shared entitlement; on a trial, the database must be started every day. The tools are the MTA Build Tool (npm i -g mbt; Windows also needs GNU Make) and the multiapps plugin for cf. This outline shows the commands the guide gives:

# Outline: run in a CAP project with an SAP HANA Cloud instance running in your subaccount.
cf add-plugin-repo CF-Community https://plugins.cloudfoundry.org
cf install-plugin -f multiapps
npm i -g mbt
cds add hana        # production database configuration
cds add xsuaa       # sign-in and roles from your @requires / @restrict annotations
cds add mta         # writes mta.yaml with the modules and services
npm install --package-lock-only   # freeze dependency versions
cds up              # builds with mbt and deploys with cf deploy

cds up runs mbt build -t gen --mtar mta.tar and then cf deploy gen/mta.tar -f. The generated mta.yaml lists a service module, a database deployer module, and the XSUAA and SAP HANA services they bind to. CAP's guide adds commands for user interfaces, such as cds add approuter for a custom application router.

Replacing the API key with sign-in. SAP's Python tutorial secures a Python app with an XSUAA instance created from an xs-security.json file. The app reads the bound XSUAA credentials and checks each request's token with the sap-xssec library (create_security_context, then check_scope). This sketch shows the shape:

# Sketch: needs an XSUAA instance bound to the app, and the sap-xssec and cfenv libraries.
from cfenv import AppEnv
from sap import xssec

uaa = AppEnv().get_service(name="my-xsuaa").credentials

def caller_may_explain(authorization_header: str) -> bool:
    token = authorization_header.removeprefix("Bearer ")
    context = xssec.create_security_context(token, uaa)
    return context.check_scope("$XSAPPNAME.Explain")   # a scope you define in xs-security.json

Calling the real model from the cloud. SAP Cloud SDK for AI for Python looks for credentials in the AICORE_ environment variables, then in ~/.aicore/config.json, then in VCAP_SERVICES. On Cloud Foundry, you bind an SAP AI Core instance to the app instead of copying AICORE_ values. Your ai_api.py checks for the AICORE_ variables at startup before using --llm; for a bound instance, that check must change to accept VCAP_SERVICES. This needs SAP AI Core with the extended plan, set up in Set up for Unit 5.

Licensing notes. Cloud Foundry runtime memory, SAP HANA Cloud, XSUAA, SAP AI Core and SAP Cloud Logging are entitlements in a BTP contract. The trial covers learning only. Check what your company's contract includes before you size a production deployment.

Build vs. SAP

Situation Do it yourself Use SAP's managed piece
Learning, a demo for your team cf push on the trial, as here Not needed yet
One small stateless service cf push with a manifest Same, in a proper subaccount
An app with a database, sign-in and a UI Hand-ordered cf commands break easily An MTA, with cds up for CAP apps
Callers are people in your company API keys can't tell people apart XSUAA or the Identity Authentication service, with roles
Callers are other systems An API key, rotated Tokens from XSUAA; an API gateway for many consumers, as in Building an AI API
Model access Copied keys in variables A bound SAP AI Core instance
The team already runs Kubernetes Containers on Cloud Foundry are possible but uncommon The Kyma runtime
Repeatable releases Manual cf push A pipeline, for example SAP Continuous Integration and Delivery

Production concerns

  • Credentials. Keys and passwords live in bound services, never in code, manifests or cf set-env. People with developer rights in the space can still read them with cf env, so keep that group small and rotate keys on a schedule.
  • Sign-in and SAP authorizations. An API key says "some caller", not "which clerk". Production callers sign in through XSUAA or the Identity Authentication service, and the app checks roles. Data from S/4HANA must respect the user's own authorizations, not a technical user's.
  • At least two instances. One instance means every platform maintenance or crash is downtime. Run two or more, and size the quota for the extra instances a rolling update creates.
  • State per instance. The AI API's cache and rate limit live in memory, so they multiply with instances and vanish on restart. For a shared limit or cache, use a backing service.
  • Health checks that mean something. /health here says "the process answers". It doesn't call the model, which keeps it cheap. Add a separate, rarer check that the model and S/4HANA connections work, and alert on it.
  • Logs and data protection. Log request IDs, status and timing, not prompts or customer data. Decide retention with your data-protection team before go-live.
  • Evaluation after each release. A smoke test proves the service answers. It doesn't prove answers are good. The evaluation harness in Unit 8 runs after each model or prompt change.
  • Cost. Each instance reserves memory; model calls are billed per use. Watch both, and set alerts on model spend.
  • Clean core. The app runs on BTP and talks to S/4HANA through released APIs and destinations. Deployment changes nothing in S/4HANA.

Pitfalls

  • Listening on 127.0.0.1 in the container. The platform can't reach it, the health check fails, and the app restarts forever. Listen on 0.0.0.0 and PORT.
  • Uploading .env. Without .cfignore, cf push uploads the whole folder, secrets included. Deploy from a dedicated folder with a .cfignore.
  • Putting the key in env: in the manifest. It ends up in Git and in platform output. Use a service.
  • Unpinned libraries. A new library release can break staging or behaviour on the next push. Pin versions and upgrade on purpose.
  • Breaking the API during a rolling update. Old and new versions answer at the same time. Add fields; don't rename or remove them in the same release.
  • Expecting the trial to stay up. Trial apps stop daily. Restart them with cf start before a demo.
  • Forgetting that cf restart and cf restage matter. A changed service or variable takes effect only after one of them.
  • Writing files in the container. The file system is short-lived; files disappear on restart and aren't shared between instances.

Exercise: update without downtime, then write the runbook

You will run two instances, release version 1.1.0 with a rolling update while a second terminal watches the service, and write down how your team deploys and rolls back. The runbook goes into the solution design document in the next topic.

  1. Start the app if it is stopped:

    cf start orchestrate-ai-api
  2. Scale it to two instances:

    cf scale orchestrate-ai-api -i 2

    Run cf app orchestrate-ai-api until it shows #0 and #1, both running.

  3. Open unit06/cf-ai-api/manifest.yml. Change instances: 1 to instances: 2 and APP_VERSION: "1.0.0" to APP_VERSION: "1.1.0". Save.

  4. In a second terminal (with .venv on, in the course folder), start watching. Replace YOUR-ROUTE with your route:

    python unit06/smoke_test.py https://YOUR-ROUTE --watch 240

    You see one line per second, such as 15:52:01 1.0.0 (instance 1).

  5. In the first terminal, push with the rolling strategy:

    • Windows (PowerShell):

      cd unit06\cf-ai-api
      cf push --strategy rolling
      cd ..\..
    • macOS / Linux:

      cd unit06/cf-ai-api
      cf push --strategy rolling
      cd ../..
  6. Watch the second terminal. For a while, lines show both 1.0.0 and 1.1.0, then only 1.1.0. When the watch ends, it prints a summary such as:

    Versions seen: 1.0.0, 1.1.0. Failed requests: 0.

    If you see a few failed requests, note how many. A rolling update aims for none; the runbook should say what your team will accept.

  7. Run the full smoke test once more:

    python unit06/smoke_test.py https://YOUR-ROUTE --key-env ORCHESTRATE_CLOUD_API_KEY

    It should end with 7 of 7 passed and show version 1.1.0.

  8. In unit06, create deployment_runbook.md with four short sections:

    • Deploy: the exact commands, from predeploy_check.py to the smoke test.
    • Roll back: how you would return to 1.0.0 (set the version back and push again with --strategy rolling; cf cancel-deployment orchestrate-ai-api during a running update).
    • Secrets: where the key lives, who can read it, and how to rotate it (cf update-user-provided-service, then cf restart).
    • Checks and limits: the health check, the smoke test, the failed requests you saw in step 6, and what changes per instance (cache and rate limit).
  9. Scale back down and stop the app, so it doesn't use trial quota:

    cf scale orchestrate-ai-api -i 1
    cf stop orchestrate-ai-api
  10. Commit your work from the course folder:

    git add unit06/cf-ai-api unit06/predeploy_check.py unit06/smoke_test.py unit06/deployment_runbook.md
    git commit -m "Deploy the AI API to Cloud Foundry with a rolling update and a runbook"

    git status should not list .env.

No trial? Do steps 3 and 8 only, and in step 8 write the commands you would run. Run python unit06/predeploy_check.py to confirm the folder is ready.

Done when: the watch output shows both versions with the failed-request count recorded, the smoke test passes against version 1.1.0 (or, without a trial, predeploy_check.py prints Ready to push.), and deployment_runbook.md is committed with all four sections filled in.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Your app deploys, then the platform keeps restarting it because the health check fails. The code runs fine on your laptop. What is the most likely cause?

    Answer: C. The platform sends the health check to the port it chose, over the container's network. An app bound to 127.0.0.1 or its own port never answers, so the check fails and the instance restarts. That's why cf_app.py uses 0.0.0.0 and PORT.
  2. 2Why does cf_app.py read the key from VCAP_SERVICES rather than from a variable set with cf set-env?

    Answer: A. Cloud Foundry's documentation advises against user-provided environment variables for credentials and points to service instances instead. Bound credentials are still readable with cf env by space developers, so they are not hidden from your colleagues, only kept out of code and output.
  3. 3You scale the AI API from one to two instances. What happens to its rate limit of 10 calls per minute per key?

    Answer: D. Each instance is its own process with its own memory, so counters and caches aren't shared. A shared limit needs a backing service. The same applies to the cache, which is why the runbook asks you to record it.
  4. 4During a rolling update from 1.0.0 to 1.1.0, what must be true of your API changes?

    Answer: B. Cloud Foundry serves old and new versions at the same route during a rolling update. Changes must be backward compatible, such as adding a field rather than renaming one. Rolling updates don't run database migrations.
  5. 5What does .cfignore protect against in this topic's setup?

    Answer: C. cf push uploads the folder it runs in. .cfignore lists what to leave out, such as .env with your keys. Deploying from a dedicated folder adds a second layer of protection.
  6. 6A colleague wants to deploy the order-assist CAP extension to BTP with its database and sign-in. Which approach fits SAP's documented path?

    Answer: D. CAP's deployment guide adds production configuration with cds add hana, cds add xsuaa and cds add mta, then cds up builds an MTA archive and deploys it with cf deploy. Pushing parts by hand is fragile, and cds watch is for development.
  7. 7The deployed service will move from the sample model to SAP AI Core. How should the credentials reach the app?

    Answer: C. SAP Cloud SDK for AI for Python reads credentials from VCAP_SERVICES when the AICORE_ variables aren't set, so a service binding keeps them out of code and manifests. The startup check in ai_api.py must be changed to allow that.
  8. 8Why does the HTTP health check call /health rather than /v1/explain?

    Answer: B. The platform calls the health check often. /health answers without a key and without calling the model, so it costs nothing and fails only when the process is in trouble. Checks on the model and S/4HANA connections belong in a separate, rarer check with alerts.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in