Move an AI service from your laptop to Cloud Foundry on SAP BTP, with keys kept out of code, health checks, logs and updates that don't take it offline.
An AI service that runs on one laptop helps one person. To help a whole order desk it has to run somewhere managed: always on, reachable over a secure address, and looked after when it crashes. Moving it there is deployment.
On SAP BTP, the usual place is the Cloud Foundry environment. A developer hands over the code with one command. The platform installs what the code needs, starts it in a container, gives it a web address and keeps checking that it answers. If it stops answering, the platform restarts it.
Deployment changes more than the address. Passwords and keys can no longer sit in a file on someone's laptop. Several copies of the app may run at once. Logs have to go somewhere people can read them. Updates have to happen without switching the service off. This topic covers each of those, and the learner deploys the course's AI API to a free BTP trial.
A pilot that only runs on a developer's machine is a demo. Value starts when clerks rely on it at 8 a.m., in every country, without calling the developer. That step is often where AI projects slip, because nobody planned who runs the service.
Take the course's running example: blocked sales orders in order-to-cash. Earlier topics in this unit built an AI API that explains why an order is on hold, and a CAP extension that clerks use. Once the order desk depends on them:
Downtime becomes visible. If the explanation service is down, clerks fall back to phone calls and email. The platform's health checks and automatic restarts reduce that, and running two copies avoids a single point of failure.
Keys become a security issue. The service holds a key that callers must present, and later the credentials for the model. On a shared platform those must be kept out of code and out of settings that many people can read.
Changes need a safe path. Prompts and models change often. A rolling update replaces the running copies one at a time, so clerks keep getting answers while the new version starts.
Cost becomes a line item. Each copy reserves memory in the BTP account, and model calls are billed separately. Someone must own both.
Getting this right turns a successful pilot into a service the business can depend on. Getting it wrong usually shows up as a leaked key, a service that silently stopped, or an update that broke every caller at once.
As of October 2026, SAP BTP offers three runtime environments, as SAP's learning material describes them:
Cloud Foundry environment. Supports many languages and runtimes, including Java, Node.js and Python. You push code; buildpacks turn it into a running app. This topic uses it, because the free trial includes it and it asks the least of a beginner.
Kyma runtime. A fully managed Kubernetes runtime. Teams that already package apps as containers, or need more control over how they run, often choose it.
ABAP environment. For ABAP-based extensions and cloud apps, for example extensions to SAP S/4HANA Cloud.
Around the runtime, SAP provides the managed pieces a real app needs. CAP's deployment guide adds SAP HANA Cloud for data and XSUAA for user sign-in, and packages the whole app as one multitarget application (MTA) that a single command deploys. SAP AI Core credentials reach the app through a service binding rather than a file. SAP's learning material recommends central logging, such as SAP Cloud Logging, for monitoring.
The learner's free BTP trial is enough to practise. Set up for Unit 6 noted its limits: it is not for production or team use, it ends, and its apps stop automatically every day.
"It runs on my laptop, so deployment is a formality." Deployment changes where keys live, how many copies run, and how updates happen. Each of those can break a working app.
"The trial is fine for the pilot." SAP rules out productive and team use of the trial, and its apps stop daily. A pilot that users depend on needs a proper account.
"Putting the key in an environment variable is safe." Cloud Foundry's own documentation warns against it: such values can show up in command output and platform logs. Use a bound service instead.
"More copies means the same behaviour, only faster." Anything an app keeps in memory, such as a cache or a rate-limit counter, is kept per copy. Design for that before scaling out.
"An update is a restart." A plain restart takes the service down. A rolling update keeps old copies serving until new ones are healthy.
Pick one answer for each question. The explanation appears after you choose.
1A sponsor says the AI pilot "works on the developer's laptop, so it's ready for the order desk". What is the main gap?
Answer: B. Deployment is where keys move out of local files, copies are restarted and scaled by the platform, and updates must keep the service up. None of that exists on a laptop. Retraining or rewriting in ABAP isn't what deployment needs.
2Your team wants to run the order-desk pilot on a developer's BTP trial account. What is the problem?
Answer: C. The trial has a Cloud Foundry org and space, which is enough to learn deployment. SAP doesn't allow productive or team use, and the apps stop every day for cleanup, so a pilot people depend on needs a proper account.
3A developer proposes storing the model credentials as a plain Cloud Foundry environment variable. What should you ask for instead?
Answer: D. Cloud Foundry's documentation warns that plain environment variables can show up in command output and platform logs. Credentials belong in a service instance bound to the app, never in code or the manifest.
4The order desk needs the explanation service to keep answering during updates. What should the team use?
Answer: B. A rolling update starts new copies and removes old ones only once the new ones are healthy, so callers keep getting answers. A restart takes the service down, and a second trial account doesn't help production.
5Which question best tests whether a partner is ready to run your AI service in production?
Answer: D. Production readiness is about operations: ownership, alerts, rollback, keys and sign-in. Benchmark scores and code size say nothing about whether the service will stay up for the order desk.
6When is the Kyma runtime a more natural choice than Cloud Foundry on SAP BTP?
Answer: C. Kyma is SAP's fully managed Kubernetes runtime, a natural home for container-based apps. ABAP extensions go to the ABAP environment, and sign-in needs exist in every runtime.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: build once, configure per environment, let the platform run it
Deployment hands three jobs to the platform that you did on your laptop by hand: starting the app, configuring it, and watching it.
Starting. On your laptop you typed python unit06/ai_api.py. On Cloud Foundry you upload the code once; a buildpack installs the libraries and produces a runnable image, and the platform starts it in a container. You tell it how in a manifest.
Configuring. On your laptop the key sat in .env. In the cloud the same code reads its settings from the platform: the port in PORT, credentials from bound services in VCAP_SERVICES. The code stays identical across laptop, test and production; only the configuration changes.
Watching. On your laptop you noticed when the app crashed. The platform runs a health check and restarts the app when it fails. It can run several copies and replace them one at a time during updates.
Once you see deployment this way, most production problems sort themselves into one of the three: it didn't start, it was configured wrong, or nobody noticed it failing.
sequenceDiagram
participant D as Your laptop
participant CF as Cloud Foundry
participant B as Python buildpack
participant C as Container
D->>CF: cf push (code + manifest)
CF->>B: stage: requirements.txt, runtime.txt
B-->>CF: droplet (app + Python + libraries)
CF->>C: start "python cf_app.py", PORT, VCAP_SERVICES
CF->>C: health check GET /health
C-->>CF: 200 OK
CF-->>D: running, route assigned
Upload.cf push reads manifest.yml and uploads the folder, minus anything listed in .cfignore.
Staging. The Python buildpack recognizes the app by its requirements.txt. It installs the Python version named in runtime.txt and the listed libraries. The result is a droplet: your code plus everything it needs.
Start. The platform starts the droplet in a container with the manifest's command. It sets PORT, which the app must listen on, and passes service credentials in VCAP_SERVICES.
Health check. The manifest attribute health-check-type can be port (the default), process or http. With http, the platform calls health-check-http-endpoint (for example /health) and expects a success code. If the check fails, the instance is restarted.
Route. The app gets a web address. random-route: true gives the app a random host name, so it doesn't clash with other apps on the same domain.
#Configuration: environment variables vs. bound services
Cloud Foundry gives an app two kinds of settings:
Kind
Set with
Seen in the app as
Use it for
User-provided environment variable
env: in the manifest, or cf set-env
Its own variable, e.g. APP_VERSION
Non-secret settings: version labels, feature switches, timeouts
Bound service instance
cf create-service or cf create-user-provided-service, then services: in the manifest
An entry in VCAP_SERVICES, with credentials
Secrets and connections: keys, database, SAP AI Core
Cloud Foundry's documentation is explicit: don't use user-provided environment variables for credentials, because they may appear in command output and platform logs. A user-provided service is the simple alternative when there is no managed service for a secret, such as your own API key. Changes to either take effect only after the app restarts or restages.
The manifest's memory reserves memory per instance; instances sets how many copies run. Two instances of a 256 MB app reserve 512 MB of the account's quota.
Each instance is a separate process with its own memory. Your AI API keeps its cache and rate-limit counters in memory, so with two instances each has its own cache and each allows its own 10 calls per minute. SAP's learning material also tells you to avoid writing files to the container's file system, because it is short-lived. Anything that must survive a restart or be shared between instances belongs in a backing service, such as a database or a cache service.
A plain cf push of a new version stops the old one first. cf push --strategy rolling starts new instances, waits until they are healthy, then removes old ones, until all are replaced. Two consequences from Cloud Foundry's documentation:
During the update, old and new versions answer at the same route. Changes to your API contract must be backward compatible, or callers get mixed answers.
The update needs spare quota for the extra instances, and it doesn't migrate databases for you. cf cancel-deployment stops a rolling update, without a zero-downtime guarantee.
Whatever the app writes to the terminal (standard output and standard error) becomes its log. cf logs APP --recent shows recent lines; cf logs APP streams them live. Your AI API already writes one line per request with a request ID and no customer data. That is exactly what you want in a platform log.
One Python app pushes cleanly with cf push. A CAP app with a database, sign-in and a user interface has several parts that must be created and connected in the right order. SAP's answer is the multitarget application: an mta.yaml file that lists the modules and the services they need, built into one archive and deployed with cf deploy. You'll see this in "The SAP way" below.
#Build it yourself: deploy the AI API to your BTP trial
You will deploy the AI API from Building an AI API to Cloud Foundry in your free BTP trial. You'll add a small start file that reads its settings from the platform, a manifest, and a key kept in a bound service. You'll first run it on your computer exactly the way Cloud Foundry will, then push it, test it over the internet, and read its logs.
flowchart LR
P[predeploy_check.py] --> F[unit06/cf-ai-api<br/>cf_app.py + ai_api.py<br/>manifest.yml]
F -->|Step 4: local, port 8080| L[Your computer]
F -->|Step 6: cf push| CF[Cloud Foundry<br/>BTP trial, dev space]
K[orchestrate-ai-api-key<br/>user-provided service] -->|VCAP_SERVICES| CF
S[smoke_test.py] --> L
S -->|HTTPS + X-API-Key| CF
What success looks like:ready, then lines with your API endpoint, org and space: dev. If cf target says you are not logged in, run cf login as in Step 5 of the Unit 6 setup. You can do Steps 2 to 4 without logging in.
Run every command in this topic from the course folder unless a step says otherwise.
#Step 2: Create the deployment folder and the start file
The deployment gets its own folder, so you upload only what the app needs, not your whole course.
In VS Code's file list, right-click unit06/cf-ai-api, choose New File, name it cf_app.py, paste this and save:
"""Start the Unit 6 AI API the way Cloud Foundry expects, on SAP BTP or on your computer.
Cloud Foundry tells the app which port to use in the PORT variable, and hands it the API key
through a bound service (VCAP_SERVICES). On your computer the same file reads .env instead.
How to run on your computer (from your course folder, with .venv turned on):
python unit06/cf-ai-api/cf_app.py # listens on http://127.0.0.1:8080
On Cloud Foundry, manifest.yml runs: python cf_app.py
"""
import json
import logging
import os
import sys
import uvicorn
from fastapi import FastAPI
from ai_api import PROMPT_VERSION, SampleModel, create_app
SERVICE_NAME = "orchestrate-ai-api-key" # the user-provided service that holds the key on BTP
ON_CLOUD_FOUNDRY = "VCAP_APPLICATION" in os.environ
def find_api_key() -> tuple[str, str]:
"""Return the API key and where it came from. Never print the key itself."""
services = json.loads(os.environ.get("VCAP_SERVICES", "{}"))
for instances in services.values():
for instance in instances:
if SERVICE_NAME in (instance.get("name"), instance.get("instance_name")):
key = instance.get("credentials", {}).get("api_key", "")
if key:
return key, f"service '{SERVICE_NAME}'"
if not ON_CLOUD_FOUNDRY: # on your computer: read .env if python-dotenv is installed
try:
from dotenv import load_dotenv
load_dotenv()
except ImportError:
pass
return os.environ.get("ORCHESTRATE_API_KEY", ""), "ORCHESTRATE_API_KEY"
def build_app() -> FastAPI:
api_key, key_source = find_api_key()
if len(api_key) < 20:
sys.exit(f"No API key of 20+ characters found (looked in {key_source}). "
f"On BTP: create and bind the service '{SERVICE_NAME}'. Locally: check .env.")
model = SampleModel()
app = create_app(model, api_key)
@app.get("/about")
def about() -> dict:
"""Which version and which instance answered. No key needed, no customer data."""
return {"app_version": os.environ.get("APP_VERSION", "dev"),
"instance": os.environ.get("CF_INSTANCE_INDEX", "local"),
"model": model.name, "prompt_version": PROMPT_VERSION}
logging.getLogger("ai_api").info("key from %s, model %s, version %s", key_source, model.name,
os.environ.get("APP_VERSION", "dev"))
return app
def main() -> None:
# Cloud Foundry logs whatever the app writes to the terminal (stdout and stderr).
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s %(message)s", stream=sys.stdout)
app = build_app()
port = int(os.environ.get("PORT", "8080"))
host = "0.0.0.0" if ON_CLOUD_FOUNDRY else "127.0.0.1" # in the cloud, listen on the container's network
print(f"Starting on {host}:{port} (Cloud Foundry: {ON_CLOUD_FOUNDRY})", flush=True)
uvicorn.run(app, host=host, port=port, access_log=False)
if __name__ == "__main__":
main()
Your ai_api.py stays unchanged. cf_app.py only decides where the key and port come from, adds an /about endpoint that says which version answered, and starts the server.
You'll create four small files in unit06/cf-ai-api. For each: right-click the folder, choose New File, type the name exactly (including the leading dot in .cfignore), paste the content and save.
requirements.txt: the libraries the cloud must install, with exact versions so the cloud runs what you tested.
fastapi==0.142.2
uvicorn==0.54.0
These are the versions we tested on 2 October 2026. To use the versions in your own .venv instead, run pip show fastapi uvicorn and copy the two Version: numbers.
runtime.txt: the Python version for the cloud. The buildpack reads this file; 3.13.x means "the newest 3.13 it has". SAP's own Python tutorial uses the same line.
python-3.13.x
Your laptop may run Python 3.14. That's fine: the app uses nothing that differs between the two.
manifest.yml: how to run the app. Indentation matters in YAML; use spaces, not tabs.
Save this check as unit06/predeploy_check.py (in unit06, not in cf-ai-api). It reads the files and tells you what is missing. It uses built-in Python only.
"""Unit 6: check the cf-ai-api folder before you push it to Cloud Foundry.
It only reads files. It needs no account, no network and no extra libraries.
How to run (from your course folder):
python unit06/predeploy_check.py
"""
import filecmp
import re
import sys
from pathlib import Path
HERE = Path(__file__).resolve().parent # the unit06 folder
APP = HERE / "cf-ai-api"
problems = 0
def report(ok: bool, label: str, fix: str = "", warn_only: bool = False) -> None:
global problems
if ok:
print(f" OK {label}")
elif warn_only:
print(f" WARN {label} -> {fix}")
else:
problems += 1
print(f" MISSING {label} -> {fix}")
def read(name: str) -> str:
path = APP / name
return path.read_text(encoding="utf-8") if path.exists() else ""
print(f"Checking {APP}\n")
report(APP.is_dir(), "folder unit06/cf-ai-api", "create it (Step 2)")
for name in ["ai_api.py", "cf_app.py", "requirements.txt", "runtime.txt", "manifest.yml", ".cfignore"]:
report((APP / name).exists(), name, f"create {name} (Step 2 or 3)")
manifest = read("manifest.yml")
for needed, why in [("name: orchestrate-ai-api", "the app name"),
("health-check-type: http", "an HTTP health check"),
("health-check-http-endpoint: /health", "the /health endpoint"),
("command: python cf_app.py", "the start command"),
("- orchestrate-ai-api-key", "the key service binding")]:
report(needed in manifest, f"manifest.yml has {why}", f"add the line '{needed}' (Step 3)")
secret_like = re.findall(r"(?i)(api_key|password|secret|token)\s*:", manifest)
report(not secret_like, "manifest.yml holds no secrets", "remove keys and passwords; use the service (Step 5)")
runtime = read("runtime.txt").strip()
report(bool(re.fullmatch(r"python-3\.\d+\.(x|\d+)", runtime)), f"runtime.txt pins a Python version ({runtime or 'empty'})",
"write one line such as python-3.13.x")
lines = [line.strip() for line in read("requirements.txt").splitlines() if line.strip() and not line.startswith("#")]
unpinned = [line for line in lines if "==" not in line]
report(bool(lines) and not unpinned, "requirements.txt pins every library",
"use name==version for: " + ", ".join(unpinned or ["fastapi, uvicorn"]))
report(not any(line.lower().startswith("python-dotenv") for line in lines), "requirements.txt leaves out python-dotenv",
"the cloud reads settings from the platform, not .env", warn_only=True)
cfignore = read(".cfignore")
report(".env" in cfignore.split(), ".cfignore keeps .env out of the upload", "add a line .env")
report(not (APP / ".env").exists(), "no .env file inside cf-ai-api", "delete it; keys go in the service (Step 5)")
app_code = read("cf_app.py")
report("PORT" in app_code and "0.0.0.0" in app_code, "cf_app.py listens on PORT on all interfaces", "copy cf_app.py again (Step 2)")
original = HERE / "ai_api.py"
same = original.exists() and (APP / "ai_api.py").exists() and filecmp.cmp(original, APP / "ai_api.py", shallow=False)
report(same, "ai_api.py matches unit06/ai_api.py", "copy it again so you deploy what you tested (Step 2)", warn_only=True)
print()
if problems:
sys.exit(f"Fix {problems} item(s) above, then run this check again.")
print("Ready to push.")
Run it (the same on every system):
python unit06/predeploy_check.py
What success looks like (shortened):
OK folder unit06/cf-ai-api
OK manifest.yml has an HTTP health check
OK manifest.yml holds no secrets
OK runtime.txt pins a Python version (python-3.13.x)
OK requirements.txt pins every library
OK .cfignore keeps .env out of the upload
OK ai_api.py matches unit06/ai_api.py
Ready to push.
Any MISSING line names the step to go back to. A WARN line won't stop you, but read it.
#Step 4: Run it the cloud way on your computer (no account needed)
Before you push, run the app exactly as Cloud Foundry will: through cf_app.py, on port 8080. Then test it with a smoke test you'll reuse against the cloud.
Save this as unit06/smoke_test.py:
"""Unit 6: smoke-test the AI API after a deployment, on BTP or on your computer.
It calls a few endpoints and checks the answers the API promises. Built-in Python plus python-dotenv.
How to run (from your course folder, with .venv turned on):
python unit06/smoke_test.py http://127.0.0.1:8080 # the local copy, key from ORCHESTRATE_API_KEY
python unit06/smoke_test.py https://YOUR-ROUTE --key-env ORCHESTRATE_CLOUD_API_KEY
python unit06/smoke_test.py https://YOUR-ROUTE --watch 120 # ask /about every second for 120 seconds
"""
import argparse
import json
import os
import sys
import time
import urllib.error
import urllib.request
from urllib.parse import urlparse
from dotenv import load_dotenv
def call(base: str, path: str, method: str = "GET", body: dict | None = None, key: str | None = None):
"""Send one request. Return (status, parsed JSON or text, headers)."""
data = json.dumps(body).encode() if body is not None else None
request = urllib.request.Request(base + path, data=data, method=method)
request.add_header("Content-Type", "application/json")
if key:
request.add_header("X-API-Key", key)
try:
with urllib.request.urlopen(request, timeout=20) as response:
status, raw, headers = response.status, response.read().decode(), response.headers
except urllib.error.HTTPError as error:
status, raw, headers = error.code, error.read().decode(), error.headers
try:
return status, json.loads(raw), headers
except json.JSONDecodeError:
return status, raw, headers
def check(label: str, ok: bool, detail=None) -> bool:
print(f" {'PASS' if ok else 'FAIL'} {label}" + ("" if ok else f" (got: {str(detail)[:120]})"))
return ok
def smoke(base: str, key: str) -> bool:
results = []
status, body, _ = call(base, "/health")
results.append(check("GET /health answers 200 without a key", status == 200 and body.get("status") == "ok", body))
status, body, _ = call(base, "/about")
results.append(check("GET /about names the version", status == 200 and "app_version" in body, body))
if status == 200:
print(f" version {body['app_version']}, instance {body['instance']}, model {body['model']}")
status, body, _ = call(base, "/v1/explain", "POST", {"sales_order": "4711"})
results.append(check("POST /v1/explain without a key -> 401", status == 401, status))
status, body, headers = call(base, "/v1/explain", "POST", {"sales_order": "4711"}, key)
results.append(check("POST /v1/explain 4711 with the key -> 200", status == 200 and body.get("source") in ("model", "cache"), body))
results.append(check("the answer carries a request ID", bool(headers.get("X-Request-ID")), dict(headers)))
status, body, _ = call(base, "/v1/explain", "POST", {"sales_order": "4713"}, key)
results.append(check("a model failure still gives a labeled fallback", status == 200 and body.get("source") == "fallback", body))
if urlparse(base).hostname not in ("127.0.0.1", "localhost"):
results.append(check("the route uses HTTPS", base.startswith("https://"), base))
print(f"\n{sum(results)} of {len(results)} passed")
return all(results)
def watch(base: str, seconds: int) -> None:
"""Ask /about once a second and print what answered, to watch a rolling deployment."""
failures, seen, end = 0, set(), time.monotonic() + seconds
while time.monotonic() < end:
try:
status, body, _ = call(base, "/about")
line = f"{body['app_version']} (instance {body['instance']})" if status == 200 else f"HTTP {status}"
except Exception as error: # connection refused, timeout, DNS
status, line = 0, f"no answer: {type(error).__name__}"
if status != 200:
failures += 1
else:
seen.add(body["app_version"])
print(time.strftime("%H:%M:%S"), line, flush=True)
time.sleep(1)
print(f"\nVersions seen: {', '.join(sorted(seen)) or 'none'}. Failed requests: {failures}.")
def main() -> None:
parser = argparse.ArgumentParser(description="Smoke-test the Unit 6 AI API.")
parser.add_argument("url", help="the app's address, e.g. https://orchestrate-ai-api-....hana.ondemand.com")
parser.add_argument("--key-env", default="ORCHESTRATE_API_KEY", help="name of the .env line that holds the key")
parser.add_argument("--watch", type=int, metavar="SECONDS", help="only poll /about for this many seconds")
args = parser.parse_args()
base = args.url.rstrip("/")
if not base.startswith(("http://", "https://")):
base = "https://" + base # cf prints routes without https://
if args.watch:
watch(base, args.watch)
return
load_dotenv()
key = os.environ.get(args.key_env, "")
if not key:
sys.exit(f"{args.key_env} is not set in .env. See Step 5.")
try:
ok = smoke(base, key)
except urllib.error.URLError as error:
sys.exit(f"Could not reach {base}: {error.reason}. Is the app running? Check the address.")
sys.exit(0 if ok else 1)
if __name__ == "__main__":
main()
Start the app the cloud way (the same on every system):
python unit06/cf-ai-api/cf_app.py
What success looks like:
INFO ai_api key from ORCHESTRATE_API_KEY, model sample-model, version dev
Starting on 127.0.0.1:8080 (Cloud Foundry: False)
INFO: Uvicorn running on http://127.0.0.1:8080 (Press CTRL+C to quit)
Open a second terminal (Terminal > New Terminal), turn on .venv as in Step 1, and run the smoke test:
python unit06/smoke_test.py http://127.0.0.1:8080
What success looks like:
PASS GET /health answers 200 without a key
PASS GET /about names the version
version dev, instance local, model sample-model
PASS POST /v1/explain without a key -> 401
PASS POST /v1/explain 4711 with the key -> 200
PASS the answer carries a request ID
PASS a model failure still gives a labeled fallback
6 of 6 passed
Order 4713 is the sample model's planned failure, so a labeled fallback is the right answer there. Stop the app in the first terminal with Ctrl+C.
#Step 5: Create a cloud key and keep it in a service
The cloud copy gets its own key. If one leaks, you replace only that one.
Print a new random key (the same on every system):
Create the user-provided service. The -p "api_key" form makes cf ask for the value, so the key doesn't land in your terminal history (the same on every system):
Run the check again; it must end with Ready to push.:
python unit06/predeploy_check.py
Go into the deployment folder and push:
Windows (PowerShell):
cd unit06\cf-ai-api
cf push
macOS / Linux:
cd unit06/cf-ai-api
cf push
cf push finds manifest.yml in the folder. Staging takes a few minutes the first time, while the buildpack installs Python and the libraries. You'll see many lines scroll by.
What success looks like: the output ends with a summary in this shape (your route, dates and numbers differ):
name: orchestrate-ai-api
requested state: started
routes: orchestrate-ai-api-RANDOM-PART.cfapps.YOUR-REGION.hana.ondemand.com
...
state since cpu memory
#0 running 2026-10-02T15:40:12Z 0.0% ...
Copy the address after routes:. That is your app's route.
Go back to the course folder:
cd ../..
#Step 7: Test it over the internet and read its logs
Run the smoke test against the route, with the cloud key. Replace YOUR-ROUTE with the address you copied (the same on every system):
What success looks like: the same six PASS lines as in Step 4, plus PASS the route uses HTTPS, ending in 7 of 7 passed. The /about line now says version 1.0.0, instance 0.
Read the app's recent logs:
cf logs orchestrate-ai-api --recent
Among the platform's lines you should find your app's own, tagged APP/PROC/WEB/0:
... [APP/PROC/WEB/0] OUT INFO ai_api key from service 'orchestrate-ai-api-key', model sample-model, version 1.0.0
... [APP/PROC/WEB/0] OUT INFO ai_api 0c1f... POST /v1/explain 401 1ms
... [APP/PROC/WEB/0] OUT WARNING ai_api 9d3a... model unavailable: sample model: simulated timeout
The key itself never appears, and neither does any order data. Each line has a request ID you can match with what the caller saw.
Look at the app's state:
cf app orchestrate-ai-api
You see one instance, running, and its memory use against the 256 MB you reserved.
Deploying the CAP extension. CAP's deployment guide describes the path for an app like order-assist from the previous topic. It needs an SAP HANA Cloud instance in the subaccount and the hdi-shared entitlement; on a trial, the database must be started every day. The tools are the MTA Build Tool (npm i -g mbt; Windows also needs GNU Make) and the multiapps plugin for cf. This outline shows the commands the guide gives:
# Outline: run in a CAP project with an SAP HANA Cloud instance running in your subaccount.
cf add-plugin-repo CF-Community https://plugins.cloudfoundry.org
cf install-plugin -f multiapps
npm i -g mbt
cds add hana # production database configuration
cds add xsuaa # sign-in and roles from your @requires / @restrict annotations
cds add mta # writes mta.yaml with the modules and services
npm install --package-lock-only # freeze dependency versions
cds up # builds with mbt and deploys with cf deploy
cds up runs mbt build -t gen --mtar mta.tar and then cf deploy gen/mta.tar -f. The generated mta.yaml lists a service module, a database deployer module, and the XSUAA and SAP HANA services they bind to. CAP's guide adds commands for user interfaces, such as cds add approuter for a custom application router.
Replacing the API key with sign-in. SAP's Python tutorial secures a Python app with an XSUAA instance created from an xs-security.json file. The app reads the bound XSUAA credentials and checks each request's token with the sap-xssec library (create_security_context, then check_scope). This sketch shows the shape:
# Sketch: needs an XSUAA instance bound to the app, and the sap-xssec and cfenv libraries.
from cfenv import AppEnv
from sap import xssec
uaa = AppEnv().get_service(name="my-xsuaa").credentials
def caller_may_explain(authorization_header: str) -> bool:
token = authorization_header.removeprefix("Bearer ")
context = xssec.create_security_context(token, uaa)
return context.check_scope("$XSAPPNAME.Explain") # a scope you define in xs-security.json
Calling the real model from the cloud. SAP Cloud SDK for AI for Python looks for credentials in the AICORE_ environment variables, then in ~/.aicore/config.json, then in VCAP_SERVICES. On Cloud Foundry, you bind an SAP AI Core instance to the app instead of copying AICORE_ values. Your ai_api.py checks for the AICORE_ variables at startup before using --llm; for a bound instance, that check must change to accept VCAP_SERVICES. This needs SAP AI Core with the extended plan, set up in Set up for Unit 5.
Licensing notes. Cloud Foundry runtime memory, SAP HANA Cloud, XSUAA, SAP AI Core and SAP Cloud Logging are entitlements in a BTP contract. The trial covers learning only. Check what your company's contract includes before you size a production deployment.
Credentials. Keys and passwords live in bound services, never in code, manifests or cf set-env. People with developer rights in the space can still read them with cf env, so keep that group small and rotate keys on a schedule.
Sign-in and SAP authorizations. An API key says "some caller", not "which clerk". Production callers sign in through XSUAA or the Identity Authentication service, and the app checks roles. Data from S/4HANA must respect the user's own authorizations, not a technical user's.
At least two instances. One instance means every platform maintenance or crash is downtime. Run two or more, and size the quota for the extra instances a rolling update creates.
State per instance. The AI API's cache and rate limit live in memory, so they multiply with instances and vanish on restart. For a shared limit or cache, use a backing service.
Health checks that mean something./health here says "the process answers". It doesn't call the model, which keeps it cheap. Add a separate, rarer check that the model and S/4HANA connections work, and alert on it.
Logs and data protection. Log request IDs, status and timing, not prompts or customer data. Decide retention with your data-protection team before go-live.
Evaluation after each release. A smoke test proves the service answers. It doesn't prove answers are good. The evaluation harness in Unit 8 runs after each model or prompt change.
Cost. Each instance reserves memory; model calls are billed per use. Watch both, and set alerts on model spend.
Clean core. The app runs on BTP and talks to S/4HANA through released APIs and destinations. Deployment changes nothing in S/4HANA.
Listening on 127.0.0.1 in the container. The platform can't reach it, the health check fails, and the app restarts forever. Listen on 0.0.0.0 and PORT.
Uploading .env. Without .cfignore, cf push uploads the whole folder, secrets included. Deploy from a dedicated folder with a .cfignore.
Putting the key in env: in the manifest. It ends up in Git and in platform output. Use a service.
Unpinned libraries. A new library release can break staging or behaviour on the next push. Pin versions and upgrade on purpose.
Breaking the API during a rolling update. Old and new versions answer at the same time. Add fields; don't rename or remove them in the same release.
Expecting the trial to stay up. Trial apps stop daily. Restart them with cf start before a demo.
Forgetting that cf restart and cf restage matter. A changed service or variable takes effect only after one of them.
Writing files in the container. The file system is short-lived; files disappear on restart and aren't shared between instances.
#Exercise: update without downtime, then write the runbook
You will run two instances, release version 1.1.0 with a rolling update while a second terminal watches the service, and write down how your team deploys and rolls back. The runbook goes into the solution design document in the next topic.
Start the app if it is stopped:
cf start orchestrate-ai-api
Scale it to two instances:
cf scale orchestrate-ai-api -i 2
Run cf app orchestrate-ai-api until it shows #0 and #1, both running.
Open unit06/cf-ai-api/manifest.yml. Change instances: 1 to instances: 2 and APP_VERSION: "1.0.0" to APP_VERSION: "1.1.0". Save.
In a second terminal (with .venv on, in the course folder), start watching. Replace YOUR-ROUTE with your route:
It should end with 7 of 7 passed and show version 1.1.0.
In unit06, create deployment_runbook.md with four short sections:
Deploy: the exact commands, from predeploy_check.py to the smoke test.
Roll back: how you would return to 1.0.0 (set the version back and push again with --strategy rolling; cf cancel-deployment orchestrate-ai-api during a running update).
Secrets: where the key lives, who can read it, and how to rotate it (cf update-user-provided-service, then cf restart).
Checks and limits: the health check, the smoke test, the failed requests you saw in step 6, and what changes per instance (cache and rate limit).
Scale back down and stop the app, so it doesn't use trial quota:
git add unit06/cf-ai-api unit06/predeploy_check.py unit06/smoke_test.py unit06/deployment_runbook.md
git commit -m "Deploy the AI API to Cloud Foundry with a rolling update and a runbook"
git status should not list .env.
No trial? Do steps 3 and 8 only, and in step 8 write the commands you would run. Run python unit06/predeploy_check.py to confirm the folder is ready.
Done when: the watch output shows both versions with the failed-request count recorded, the smoke test passes against version 1.1.0 (or, without a trial, predeploy_check.py prints Ready to push.), and deployment_runbook.md is committed with all four sections filled in.
Pick one answer for each question. The explanation appears after you choose.
1Your app deploys, then the platform keeps restarting it because the health check fails. The code runs fine on your laptop. What is the most likely cause?
Answer: C. The platform sends the health check to the port it chose, over the container's network. An app bound to 127.0.0.1 or its own port never answers, so the check fails and the instance restarts. That's why cf_app.py uses 0.0.0.0 and PORT.
2Why does cf_app.py read the key from VCAP_SERVICES rather than from a variable set with cf set-env?
Answer: A. Cloud Foundry's documentation advises against user-provided environment variables for credentials and points to service instances instead. Bound credentials are still readable with cf env by space developers, so they are not hidden from your colleagues, only kept out of code and output.
3You scale the AI API from one to two instances. What happens to its rate limit of 10 calls per minute per key?
Answer: D. Each instance is its own process with its own memory, so counters and caches aren't shared. A shared limit needs a backing service. The same applies to the cache, which is why the runbook asks you to record it.
4During a rolling update from 1.0.0 to 1.1.0, what must be true of your API changes?
Answer: B. Cloud Foundry serves old and new versions at the same route during a rolling update. Changes must be backward compatible, such as adding a field rather than renaming one. Rolling updates don't run database migrations.
5What does .cfignore protect against in this topic's setup?
Answer: C. cf push uploads the folder it runs in. .cfignore lists what to leave out, such as .env with your keys. Deploying from a dedicated folder adds a second layer of protection.
6A colleague wants to deploy the order-assist CAP extension to BTP with its database and sign-in. Which approach fits SAP's documented path?
Answer: D. CAP's deployment guide adds production configuration with cds add hana, cds add xsuaa and cds add mta, then cds up builds an MTA archive and deploys it with cf deploy. Pushing parts by hand is fragile, and cds watch is for development.
7The deployed service will move from the sample model to SAP AI Core. How should the credentials reach the app?
Answer: C. SAP Cloud SDK for AI for Python reads credentials from VCAP_SERVICES when the AICORE_ variables aren't set, so a service binding keeps them out of code and manifests. The startup check in ai_api.py must be changed to allow that.
8Why does the HTTP health check call /health rather than /v1/explain?
Answer: B. The platform calls the health check often. /health answers without a key and without calling the model, so it costs nothing and fails only when the process is in trouble. Checks on the model and S/4HANA connections belong in a separate, rarer check with alerts.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Python buildpack (Cloud Foundry documentation)— detection by requirements.txt; runtime.txt with python-3.x.x or 3.x.x patterns; start command from Procfile, cf push -c or the manifest command; the app should listen on PORT
Configuring app deployments: rolling (Cloud Foundry documentation)— cf push --strategy rolling replaces instances after new ones are healthy; old and new versions are served at the same route during the deployment; no database migrations; needs spare quota; cf cancel-deployment
Deploy to Cloud Foundry (CAP documentation, capire)— SAP HANA Cloud instance and hdi-shared entitlement; trial database must be started daily; mbt and the multiapps cf CLI plugin; cds add hana, xsuaa, mta and UI options such as approuter; cds up runs mbt build and cf deploy; freeze dependencies
Describing Runtime Environments (learning.sap.com)— Cloud Foundry environment for many runtimes and languages; Kyma runtime is a fully managed Kubernetes runtime; ABAP environment for ABAP-based extensions