An AI feature on SAP BTP can pass every test and still stop working six weeks after go-live. Nothing in your code changed. The ground under it did.
Four things change under a running AI app on SAP AI Core:
Model versions retire. Providers withdraw versions, and SAP publishes the dates.
Limits fill up. Each model has a requests-per-minute limit. Month-end traffic finds it.
Configurations change. Someone switches the model or a setting, and the switch fails.
Keys age. The credentials your app signs in with must be rotated, and that is your job.
Running AI in production means having a routine that notices each of these before users do. It also means knowing who owns each one. This topic gives leaders that map and gives builders a script that checks all four.
Take the blocked-sales-order assistant from earlier units. It reads a blocked order, checks credit exposure and drafts a release recommendation for the credit team. It went live in August, and the team now clears blocks faster. Then these things happen:
Day 40. The model version the team pinned reaches its retirement date. SAP's documentation says a deployment pinned to a version stops working on that date. The assistant returns errors at 08:00 on a Monday.
Day 58. Month-end. Order volume triples, the model's per-minute limit is reached, and a third of requests are rejected with "too many requests".
Day 63. A developer switches the embedding model and mistypes its name. The deployment fails, and search stops while people look for the cause.
Each of these costs hours of blocked orders, delayed shipments and lost trust in the tool. None needs new technology to prevent. Each needs a date on a calendar, a number on a dashboard or a rule about who may change what.
The cost of the routine is small: a scheduled check, an owner per risk and a few alerts. The cost of skipping it is an outage in the business process the AI was meant to speed up.
As of October 2026, AI apps on SAP BTP usually call models through SAP AI Core and its generative AI hub, often through the orchestration service. SAP gives you these operating controls:
Need
What SAP provides
A production contract
The extended service plan of SAP AI Core includes the generative AI hub, on an enterprise account, with SAP support and an SLA. The standard plan covers production without generative AI.
Separation
Resource groups: separate spaces inside one SAP AI Core tenant, each with its own deployments and configurations.
Model dates
A model catalog API and SAP Note 3437766 list models, versions, rate limits and deprecation dates.
Safe changes
Updating a deployment keeps its address, and SAP AI Core records the last configuration that ran, so you can roll back.
Limits
Per-model requests-per-minute limits, an API to read them, and a support route to raise them.
Approved models only
Model restriction: an allow or deny list of models on an orchestration deployment.
Audit and cost
Security events in audit logs, and a consumption report per model and resource group in the SAP BTP cockpit.
Keys and encryption
Credentials you rotate yourself; since September 2026, customer-managed root keys through SAP Data Custodian Key Management Service.
Two dates matter right now. SAP is decommissioning the first orchestration endpoint on 31 October 2026; apps must use version 2 (/v2/completion). And every model version you pin has its own retirement date to track.
Use this table to assign owners before go-live. "Day two" means everything after launch day.
What changes
What users see
Early signal
Owner
Control
A pinned model version retires
Errors from the assistant
Retirement date in the model catalog
AI platform team
Calendar the date; upgrade and re-test 90 days ahead
Traffic outgrows the rate limit
Slow or rejected requests at peak
Share of limit used at peak
App team
Keep peak well under the limit (this course uses 70%); request increases early
A configuration change fails
Feature stops after a release
Deployment status not "running"
App team
Change through a pipeline; roll back to the last running configuration
Credentials age or leak
Sign-in failures, or misuse
Key age, unusual usage
Security / basis
Rotation schedule; keys in BTP services, not files
Spending drifts
Budget overrun
Consumption report per resource group
Product owner
Monthly review against the cost model
A platform or regional outage
Everything stops
SAP Trust Center status
AI platform team
Decide the recovery target; multi-region only if the process needs it
The last row is a business decision. SAP BTP services run across several availability zones in a region by default. Surviving the loss of a whole region needs a second subaccount in another region, and that doubles parts of the setup. Ask what an hour without the assistant costs before buying it.
"It passed testing, so it will keep working." The model versions, limits and keys under the app change on their own schedule. Production needs checks that run after go-live.
"Using the latest model version removes the risk." It removes the retirement outage, but the model can change behavior under you. Someone still has to re-run the evaluation when it changes.
"A 'too many requests' error means we must buy more." SAP notes that it can also mean temporary back-end load. Retries, caching and fallbacks come first; a quota increase comes after the numbers show the need.
"SAP handles key rotation for us." SAP's documentation says rotating SAP AI Core credentials is the customer's responsibility.
"Multi-region is standard for production." It is a design choice with real cost. Many processes are fine with a single region across availability zones and a tested recovery plan.
Pick one answer for each question. The explanation appears after you choose.
1Your assistant passed all tests in August and nobody has changed its code. Why can it still fail in October?
Answer: C. The topic's point is that the ground under the app moves: versions retire, traffic grows into limits, people change configurations and keys age. Unchanged code does not mean an unchanged system.
2The team pinned a model version for stability. What must someone own as a result?
Answer: D. SAP documents that a deployment pinned to a version stops working on that version's deprecation date. Pinning buys stability, but someone must track the date and upgrade in time.
3At month-end a third of requests fail with "too many requests". What is the sensible first response?
Answer: A. SAP notes this error can mean a rate limit or temporary back-end load, and recommends backoff, caching, fallbacks and workload isolation. A quota increase is right once the numbers show the peak really exceeds the limit.
4Why give production its own resource group in SAP AI Core?
Answer: B. Resource groups separate deployments and configurations, and limits can be set per group. Production then doesn't compete with experiments for the same allocation.
5Who is responsible for rotating the credentials an app uses to reach SAP AI Core?
Answer: C. SAP's documentation states that rotating SAP AI Core credentials and certificates is the customer's responsibility. Put it on the security team's calendar like any other secret.
6What date should every team using SAP's orchestration service check against in October 2026?
Answer: D. SAP announced that the first orchestration endpoint is decommissioned on 31 October 2026. Apps still calling it must move to version 2 before then.
7Your sponsor asks for multi-region deployment "because it's production". What do you ask first?
Answer: B. BTP services already span availability zones within a region. A second region needs a second subaccount and duplicated setup, so the outage cost of the process should decide it.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: your code is still, the platform moves
In Deploying AI apps on SAP BTP you got an app running on Cloud Foundry with health checks and updates that don't take it offline. That part of the system changes only when you push. The AI part behaves differently. It sits on top of a service that changes on four clocks you don't control:
Clock
What moves
How fast
What you can read
Calendar
Model versions are deprecated and retired
Months
retirementDate and deprecated in the model catalog
Traffic
Peak requests approach the per-model limit
Days (month-end)
GET /v2/admin/quota/model against your own peak
People
Configurations and deployments are changed
Any release
Deployment status and latestRunningConfigurationId
Security
Credentials age; policies change
Your rotation period
Your own key inventory
Operating AI on BTP is reading those four clocks on a schedule, comparing them with a written plan, and acting before the user notices. Everything in this topic is one of those three verbs: read, compare, act. The build section turns them into a script you can run every morning.
An SAP AI Core service instance gives you a tenant. Inside it, SAP AI Core's documentation (October 2026) describes two levels:
Tenant level: scenarios, executables and Docker registry secrets. These are shared by every resource group.
Resource-group level: configurations, deployments, executions and artifacts. These belong to exactly one group and can't be shared.
Every tenant gets a default resource group that can't be deleted. You can create more, up to 50 per tenant unless you raise the quota. Each API call names its group in the AI-Resource-Group header.
flowchart TB
subgraph SA[BTP subaccount: production]
subgraph T[SAP AI Core tenant, extended plan]
RG1[Resource group<br/>orders-prod]
RG2[Resource group<br/>default]
RG1 --> D1[Orchestration<br/>deployment]
RG1 --> D2[Embedding model<br/>deployment]
end
APP[Assistant app<br/>Cloud Foundry] -->|AI-Resource-Group: orders-prod| D1
end
Q[Rate limits per model<br/>tenant or group] -.-> RG1
A common layout is one BTP subaccount per environment (development, test, production), each with its own SAP AI Core instance. Inside production, give each app its own resource group. That is what SAP's rate-limit guidance calls isolating workloads: one app's burst then can't use up another's allocation.
A deployment is a running server with a URL. For generative AI you usually have one orchestration deployment (scenario orchestration) and sometimes direct model deployments (scenario foundation-models). SAP's documentation says an orchestration deployment usually already exists in your default group.
SAP AI Launchpad's documentation lists six statuses: Pending, Running, Stopping, Stopped, Dead and Unknown. Only Running serves requests.
flowchart LR
P[PENDING] --> R[RUNNING]
P --> X[DEAD]
R -->|stop| S1[STOPPING] --> S2[STOPPED]
R -->|patch with a bad configuration| X
U[UNKNOWN] -.->|delete| G[deleted]
S2 -->|delete| G
X -->|delete, or patch back within 7 days| G
Three facts from SAP's documentation shape your runbook:
A stopped deployment can't be restarted. To serve again, create a new deployment from the same configuration. The new one gets a new URL.
Only Stopped, Dead or Unknown deployments can be deleted. A Running one must be stopped first.
Deleting frees quota. Each tenant has a default quota for the number of deployments and replicas. Old stopped deployments still count until you delete them.
SAP documents what happens next. The deployment keeps its URL, so your app needs no change. Requests keep working during the transition. If the new configuration is wrong, the deployment can end in DEAD or stay in PENDING. Then requests stop working.
The safety net is a field on the deployment: latestRunningConfigurationId. It holds the last configuration that reached Running. To roll back, PATCH again with that ID. You have seven days: SAP deletes a dead deployment seven days after it reached DEAD.
sequenceDiagram
participant Dev as Release pipeline
participant AIC as SAP AI Core
Dev->>AIC: POST /v2/lm/configurations (new model)
AIC-->>Dev: configurationId B
Dev->>AIC: PATCH deployment, configurationId B
AIC-->>Dev: Deployment modification scheduled
Dev->>AIC: GET deployment (poll)
AIC-->>Dev: status DEAD, latestRunningConfigurationId A
Dev->>AIC: PATCH deployment, configurationId A
AIC-->>Dev: status RUNNING
The lesson: a change isn't finished when the PATCH returns. It is finished when the status reads RUNNING and a smoke test passes. Your pipeline should poll, test and roll back on its own.
Every model version in the generative AI hub has a lifecycle. SAP's Model Lifecycle page gives two upgrade options:
Auto upgrade. Set modelVersion to latest. When SAP AI Core supports a newer version, existing deployments use it automatically. If you don't set a version, latest is the default.
Manual upgrade. Pin a specific modelVersion. It stays fixed until you patch the deployment with a configuration that names a new version. SAP states that a deployment with a pinned version stops working on that version's deprecation date.
Neither is free. Auto upgrade avoids the retirement outage but changes the model under you without a release; answers can shift. Pinning keeps behavior fixed until you choose, but puts a hard date on your calendar.
A practical rule: pin in production, let a test environment run on latest, and run your evaluation harness there whenever the latest version changes. When the test results are good, move production to the new pinned version.
You read the dates from the model catalog:
GET {AI_API_URL}/v2/lm/scenarios/foundation-models/models
AI-Resource-Group: orders-prod
Each model has a versions list. Per version, SAP documents name, isLatest, deprecated, retirementDate, contextLength, cost and more. SAP Note 3437766 holds the same information in human-readable form, including token conversion rates and rate limits.
SAP AI Core limits requests per minute (RPM) per model per tenant. As SAP describes it (October 2026):
All resource groups share the tenant limit by default. You can configure a limit per resource group.
All versions of one model share the same allocation.
Limits reset every minute. Over the limit, requests get HTTP 429.
A 429 can also mean the back-end or hyperscaler is temporarily saturated. Check the Retry-After header.
GET /v2/admin/quota/model returns your limits. Each entry has unit (requestPerMinute), limit, and descriptors with modelName, modelVersion and providerName. Add the AI-Resource-Group header to see a group-level limit.
Your app's side was covered earlier in this unit: retries with backoff and jitter, caching and fallbacks across models. The operations side is planning. Know your peak per model from your traces, compare it with the limit, and request an increase well before month-end. This course uses a 70% headroom line: if the planned peak is above 70% of the limit, act now.
Deployment quota is a separate number: how many deployments and replicas a tenant may have. Raising either limit is a support ticket on component CA-ML-AIC.
Your app signs in with OAuth client credentials from a service key or a service binding, as in Secrets, API authentication and OAuth. Two documented facts matter for operations:
SAP states that you are responsible for rotating SAP AI Core credentials and certificates, according to your own policy.
A service key can use an x.509 certificate instead of a client secret. The certificate has a validity period you set when creating the key, which forces rotation.
Rotation without downtime follows one pattern: create the new key, deploy the app with it, confirm traffic, then delete the old key. In a Cloud Foundry app, bind the service instead of copying a key, and rotate by rebinding.
Three records help you investigate after the fact:
Audit logs. SAP AI Core writes security events to the BTP audit log, such as creating or deleting deployments, executions, resource groups and secrets. Failed sign-ins appear with messages such as an expired or invalid token.
Consumption. Since April 2025, the SAP BTP cockpit shows generative AI consumption per model, for the orchestration service and per resource group. One group per app gives you cost per app for free.
Inference observability. SAP AI Core can store the payloads of model calls, but only when your request headers ask for it. You can store full payloads in an S3 object store or metadata only (model, version, tokens, latency). SAP warns that payloads are stored without masking, so recording personal data needs consent.
These complement the traces your app writes itself. Your traces tell you what the user experienced. SAP's records tell you what the platform saw.
#Build it yourself: a production preflight for SAP AI Core
You will build ai_core_preflight.py, a script that reads the four clocks and compares them with a short plan file. It prints one line per check, marked OK, INFO, WARN or FAIL, and saves a report. It only reads: it never stops, patches or deletes anything. When it finds a problem it tells you the fix, and a person decides.
You run it first on made-up data that contains every kind of problem. Then, if you have the SAP AI Core trial from Unit 5, you run it against your own tenant.
Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. The sample path needs only Python. The optional live run (Step 6) also needs the SAP AI Core trial and the AICORE_ lines in .env from Set up for Unit 5.
flowchart LR
PL[ops_plan.json<br/>your models, peaks, thresholds] --> C[ai_core_preflight.py<br/>compare]
subgraph AIC[SAP AI Core, or --sample]
D[deployments]
CF[configurations]
M[model catalog]
Q[rate limits]
end
D --> C
CF --> C
M --> C
Q --> C
C --> O[OK / INFO / WARN / FAIL<br/>on screen]
C --> R[preflight_report.json]
In VS Code, right-click the unit10 folder, choose New File, name it ai_core_preflight.py, paste the code below and save.
"""Production preflight for an AI app on SAP AI Core.
Reads deployments, configurations, the model catalog and rate limits,
compares them with your operations plan (ops_plan.json) and prints
OK / INFO / WARN / FAIL lines. It only reads: it never changes anything.
python unit10/ai_core_preflight.py --sample made-up data, no account
python unit10/ai_core_preflight.py --live your SAP AI Core (.env)
python unit10/ai_core_preflight.py --live --list-models
"""
import argparse
import json
import os
import sys
from datetime import date, timedelta
from pathlib import Path
HERE = Path(__file__).resolve().parent
ORDER = {"OK": 0, "INFO": 1, "WARN": 2, "FAIL": 3}
SAMPLE_PLAN = {
"app": "blocked-orders-assistant",
"environment": "production",
"access": "orchestration",
"warn_days": 90,
"fail_days": 30,
"headroom": 0.7,
"models": [
{"name": "sample-chat-large", "version": "2025-06-01", "peak_rpm": 50},
{"name": "sample-chat-small", "version": "latest", "peak_rpm": 40},
],
}
def sample_data(today):
"""SAP-shaped responses with made-up names, versions, dates and limits."""
soon = (today + timedelta(days=60)).isoformat()
very_soon = (today + timedelta(days=20)).isoformat()
models = {"count": 3, "resources": [
{"model": "sample-chat-large", "executableId": "sample-provider",
"versions": [
{"name": "2025-06-01", "isLatest": False, "deprecated": False,
"retirementDate": soon},
{"name": "2026-05-01", "isLatest": True, "deprecated": False,
"retirementDate": ""}]},
{"model": "sample-chat-small", "executableId": "sample-provider",
"versions": [
{"name": "2026-02-01", "isLatest": True, "deprecated": False,
"retirementDate": ""}]},
{"model": "sample-embed", "executableId": "sample-provider",
"versions": [
{"name": "1", "isLatest": True, "deprecated": True,
"retirementDate": very_soon}]},
]}
configurations = {"count": 3, "resources": [
{"id": "cfg-orch-1", "name": "orchestration-config",
"scenarioId": "orchestration", "executableId": "orchestration",
"parameterBindings": []},
{"id": "cfg-embed-1", "name": "embed-config",
"scenarioId": "foundation-models", "executableId": "sample-provider",
"parameterBindings": [{"key": "modelName", "value": "sample-embed"},
{"key": "modelVersion", "value": "1"}]},
{"id": "cfg-embed-2", "name": "embed-config-typo",
"scenarioId": "foundation-models", "executableId": "sample-provider",
"parameterBindings": [{"key": "modelName", "value": "sample-embedd"},
{"key": "modelVersion", "value": "1"}]},
]}
deployments = {"count": 4, "resources": [
{"id": "d0rch000000001", "scenarioId": "orchestration",
"configurationId": "cfg-orch-1", "status": "RUNNING",
"latestRunningConfigurationId": "cfg-orch-1"},
{"id": "d3mb000000002", "scenarioId": "foundation-models",
"configurationId": "cfg-embed-1", "status": "RUNNING",
"latestRunningConfigurationId": "cfg-embed-1"},
{"id": "d3mb000000003", "scenarioId": "foundation-models",
"configurationId": "cfg-embed-2", "status": "DEAD",
"lastOperation": "UPDATE",
"latestRunningConfigurationId": "cfg-embed-1"},
{"id": "d0ld000000004", "scenarioId": "foundation-models",
"configurationId": "cfg-embed-1", "status": "STOPPED",
"latestRunningConfigurationId": "cfg-embed-1"},
]}
quota = [{"resourceType": "model", "quotaDetails": [
{"unit": "requestPerMinute", "limit": 60,
"descriptors": {"modelName": "sample-chat-large", "modelVersion": "*"}},
{"unit": "requestPerMinute", "limit": 300,
"descriptors": {"modelName": "sample-chat-small", "modelVersion": "*"}},
]}]
return {"resource_group": "default", "deployments": deployments,
"configurations": configurations, "models": models, "quota": quota}
def live_data():
"""Read the same four lists from SAP AI Core with the .env credentials."""
try:
import requests
from dotenv import load_dotenv
except ImportError as err:
sys.exit(f"Missing library: {err.name}. Run: pip install -r requirements.txt")
load_dotenv()
names = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL",
"AICORE_BASE_URL", "AICORE_RESOURCE_GROUP"]
missing = [n for n in names if not os.getenv(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing))
try:
token = requests.post(
os.environ["AICORE_AUTH_URL"],
data={"grant_type": "client_credentials"},
auth=(os.environ["AICORE_CLIENT_ID"], os.environ["AICORE_CLIENT_SECRET"]),
timeout=30)
except requests.RequestException as err:
sys.exit(f"Could not reach the sign-in URL: {err}")
if token.status_code != 200:
sys.exit(f"Sign-in failed with HTTP {token.status_code}. Check the AICORE_ lines in .env.")
group = os.environ["AICORE_RESOURCE_GROUP"]
base = os.environ["AICORE_BASE_URL"].rstrip("/")
headers = {"Authorization": "Bearer " + token.json()["access_token"],
"AI-Resource-Group": group}
def get(path, optional=False):
try:
resp = requests.get(base + path, headers=headers, timeout=30)
except requests.RequestException as err:
sys.exit(f"Could not reach SAP AI Core: {err}")
if resp.status_code == 200:
return resp.json()
if optional:
print(f" (GET {path} returned HTTP {resp.status_code}; skipping that check)")
return None
sys.exit(f"GET {path} failed with HTTP {resp.status_code}: {resp.text[:200]}")
return {"resource_group": group,
"deployments": get("/lm/deployments"),
"configurations": get("/lm/configurations"),
"models": get("/lm/scenarios/foundation-models/models"),
"quota": get("/admin/quota/model", optional=True)}
def read_date(text):
try:
return date.fromisoformat(str(text)[:10])
except ValueError:
return None
def check_version(model, version_name, today, plan, where):
"""Return (level, message) for one model version against the catalog."""
versions = model.get("versions", [])
if version_name == "latest":
found = next((v for v in versions if v.get("isLatest")), None)
label = f"{model['model']} latest"
else:
found = next((v for v in versions if v.get("name") == version_name), None)
label = f"{model['model']} {version_name}"
if found is None:
return "FAIL", f"{where}: {label} is not in the model catalog"
retire = read_date(found.get("retirementDate", ""))
if retire:
days = (retire - today).days
if days <= plan["fail_days"]:
return "FAIL", f"{where}: {label} retires on {retire} ({days} days)"
if days <= plan["warn_days"]:
return "WARN", f"{where}: {label} retires on {retire} ({days} days); plan the upgrade"
elif found.get("retirementDate"):
return "WARN", f"{where}: {label} has an unreadable retirementDate '{found['retirementDate']}'"
if found.get("deprecated"):
return "WARN", f"{where}: {label} is deprecated; check SAP Note 3437766 for the date"
if version_name == "latest":
return "INFO", (f"{where}: {label} auto-upgrades (now {found.get('name')}); "
"re-run your evaluation when it changes")
if not found.get("isLatest"):
return "OK", f"{where}: {label} is pinned and supported; a newer version exists"
return "OK", f"{where}: {label} is pinned and supported"
def run_checks(plan, data, today):
results = []
add = lambda level, text: results.append({"level": level, "check": text})
catalog = {m["model"]: m for m in data["models"].get("resources", [])}
configs = {c["id"]: c for c in data["configurations"].get("resources", [])}
# 1. Resource group
if plan["environment"] == "production" and data["resource_group"] == "default":
add("WARN", "Resource group 'default' in production: give the app its own group "
"to isolate it and its rate limits")
else:
add("OK", f"Resource group '{data['resource_group']}'")
# 2. Deployments
deployments = data["deployments"].get("resources", [])
running_orch = [d for d in deployments
if d.get("scenarioId") == "orchestration" and d.get("status") == "RUNNING"]
if plan.get("access") == "orchestration":
add("OK" if running_orch else "FAIL",
f"Orchestration deployments running: {len(running_orch)}")
for d in deployments:
status, did = d.get("status", "UNKNOWN"), d.get("id")
if status == "RUNNING":
add("OK", f"Deployment {did} is RUNNING")
elif status == "DEAD":
last = d.get("latestRunningConfigurationId")
msg = f"Deployment {did} is DEAD"
if last and last != d.get("configurationId"):
msg += (f"; roll back with PATCH /v2/lm/deployments/{did} "
f'body {{"configurationId": "{last}"}} (within 7 days)')
add("FAIL", msg)
elif status in ("PENDING", "UNKNOWN"):
add("WARN", f"Deployment {did} is {status}; not serving yet")
else:
add("INFO", f"Deployment {did} is {status}; it can't restart, "
"delete it to free deployment quota")
# 3. Model version behind a direct (foundation-models) deployment
cfg = configs.get(d.get("configurationId"), {})
bind = {b["key"]: b["value"] for b in cfg.get("parameterBindings", [])}
if status == "RUNNING" and "modelName" in bind:
model = catalog.get(bind["modelName"])
if model is None:
add("FAIL", f"Deployment {did}: {bind['modelName']} is not in the model catalog")
else:
add(*check_version(model, bind.get("modelVersion", "latest"),
today, plan, f"Deployment {did}"))
# 4. Models the app sends through orchestration, from the plan
for m in plan["models"]:
model = catalog.get(m["name"])
if model is None:
add("FAIL", f"Plan: {m['name']} is not in this tenant's model catalog")
continue
add(*check_version(model, m.get("version", "latest"), today, plan, "Plan"))
# 5. Rate limits against the planned peak
if data["quota"] is None:
add("INFO", "Rate limits not checked (quota endpoint not available)")
else:
limits = {}
for block in data["quota"]:
for q in block.get("quotaDetails", []):
if q.get("unit") == "requestPerMinute":
limits[q["descriptors"].get("modelName")] = q["limit"]
for m in plan["models"]:
limit = limits.get(m["name"])
if limit is None:
add("WARN", f"Rate limit for {m['name']} not listed; ask before go-live")
continue
share = m["peak_rpm"] / limit
text = f"Rate limit {m['name']}: peak {m['peak_rpm']} of {limit} requests/min ({share:.0%})"
if share > 1:
add("FAIL", text + "; requests will get HTTP 429")
elif share > plan["headroom"]:
add("WARN", text + f"; above the {plan['headroom']:.0%} headroom")
else:
add("OK", text)
return results
def main():
parser = argparse.ArgumentParser(description="Production preflight for SAP AI Core")
mode = parser.add_mutually_exclusive_group(required=True)
mode.add_argument("--sample", action="store_true", help="use made-up data, no account")
mode.add_argument("--live", action="store_true", help="read your SAP AI Core (.env)")
parser.add_argument("--plan", default=str(HERE / "ops_plan.json"))
parser.add_argument("--list-models", action="store_true",
help="print the model catalog and stop")
parser.add_argument("--today", help="pretend today is YYYY-MM-DD")
args = parser.parse_args()
today = date.fromisoformat(args.today) if args.today else date.today()
plan_path = Path(args.plan)
if not plan_path.exists():
plan_path.write_text(json.dumps(SAMPLE_PLAN, indent=2) + "\n", encoding="utf-8")
print(f"Created {plan_path} with the sample plan. Edit it for your app.")
try:
plan = json.loads(plan_path.read_text(encoding="utf-8"))
except json.JSONDecodeError as err:
sys.exit(f"{plan_path.name} is not valid JSON (line {err.lineno}): {err.msg}")
data = sample_data(today) if args.sample else live_data()
if args.list_models:
for m in data["models"].get("resources", []):
for v in m.get("versions", []):
flags = ("latest " if v.get("isLatest") else "") + \
("deprecated " if v.get("deprecated") else "")
print(f"{m['model']:<32} {v.get('name', ''):<14} "
f"retires {v.get('retirementDate') or '-':<12} {flags}")
return 0
results = run_checks(plan, data, today)
print(f"Preflight for {plan['app']} ({plan['environment']}), "
f"{'sample data' if args.sample else 'live'}, {today}")
for r in sorted(results, key=lambda r: -ORDER[r["level"]]):
print(f"{r['level']:<5} {r['check']}")
counts = {lvl: sum(r["level"] == lvl for r in results) for lvl in ORDER}
print("Summary: " + ", ".join(f"{n} {lvl}" for lvl, n in counts.items()))
report = HERE / "preflight_report.json"
report.write_text(json.dumps({"date": today.isoformat(), "app": plan["app"],
"mode": "sample" if args.sample else "live",
"summary": counts, "results": results}, indent=2) + "\n",
encoding="utf-8")
print(f"Saved {report.name}")
return 2 if counts["FAIL"] else 0
if __name__ == "__main__":
sys.exit(main())
Run the script with --sample. The command is the same on every system:
python unit10/ai_core_preflight.py --sample
On the first run it creates unit10/ops_plan.json from the sample plan. You should see output like this (your dates differ; the day counts are the same, because the sample dates are set relative to today):
Created .../unit10/ops_plan.json with the sample plan. Edit it for your app.
Preflight for blocked-orders-assistant (production), sample data, 2026-10-07
FAIL Deployment d3mb000000002: sample-embed 1 retires on 2026-10-27 (20 days)
FAIL Deployment d3mb000000003 is DEAD; roll back with PATCH /v2/lm/deployments/d3mb000000003 body {"configurationId": "cfg-embed-1"} (within 7 days)
WARN Resource group 'default' in production: give the app its own group to isolate it and its rate limits
WARN Plan: sample-chat-large 2025-06-01 retires on 2026-12-06 (60 days); plan the upgrade
WARN Rate limit sample-chat-large: peak 50 of 60 requests/min (83%); above the 70% headroom
INFO Deployment d0ld000000004 is STOPPED; it can't restart, delete it to free deployment quota
INFO Plan: sample-chat-small latest auto-upgrades (now 2026-02-01); re-run your evaluation when it changes
OK Orchestration deployments running: 1
OK Deployment d0rch000000001 is RUNNING
OK Deployment d3mb000000002 is RUNNING
OK Rate limit sample-chat-small: peak 40 of 300 requests/min (13%)
Summary: 4 OK, 2 INFO, 3 WARN, 2 FAIL
Saved preflight_report.json
Check the exit code. Because there is at least one FAIL, it is 2:
Windows (PowerShell):
$LASTEXITCODE
macOS / Linux:
echo $?
A pipeline that runs this script before a release would stop here.
Each line maps to one of the four clocks. Decide what you would do, then compare with this table.
Line
Clock
What it means
Action (by a person)
sample-embed 1 retires ... (20 days)
Calendar
A running deployment is pinned to a version that retires within the 30-day fail line. On that date it stops.
Create a configuration with a supported version, evaluate it, patch the deployment.
d3mb000000003 is DEAD; roll back...
People
Someone patched in a bad configuration (here, a misspelled model name). latestRunningConfigurationId still points to the good one.
Run the printed PATCH within 7 days, then fix the release process.
Resource group 'default' in production
Traffic
Production shares the default group, and with it any group-level limit, with everything else.
Create an orders-prod group and move the app.
sample-chat-large ... (60 days)
Calendar
The model your app requests through orchestration retires inside the 90-day warning window.
Schedule the upgrade and the evaluation run now.
peak 50 of 60 requests/min (83%)
Traffic
Month-end peak would use 83% of the limit. One busy hour and requests get 429.
Add caching or a fallback model, or request a higher limit.
d0ld000000004 is STOPPED
People
A stopped deployment can't serve again but still counts toward deployment quota.
Delete it.
sample-chat-small latest auto-upgrades
Calendar
No retirement outage, but the model can change without a release.
Re-run the evaluation when the version changes.
Open unit10/preflight_report.json in VS Code. It holds the same results in JSON, with the date and a summary. A scheduled job can keep one report per day, so you can show when a problem first appeared.
The plan file is where your team writes down what the app uses. Two of the warnings are fixed by decisions, not by code.
Open unit10/ops_plan.json.
Upgrade the pinned model: change "version": "2025-06-01" to "version": "2026-05-01". The sample catalog has that version.
Lower the month-end peak, as if you had added a cache: change the first "peak_rpm": 50 to "peak_rpm": 40. Save.
Run again:
python unit10/ai_core_preflight.py --sample
You should see the two plan warnings turn to OK:
OK Plan: sample-chat-large 2026-05-01 is pinned and supported
OK Rate limit sample-chat-large: peak 40 of 60 requests/min (67%)
Summary: 6 OK, 2 INFO, 1 WARN, 2 FAIL
The two FAIL lines remain. They describe deployments, and only an action on SAP AI Core fixes them. That is deliberate: the script reports, a person acts.
To see what a misspelled model in your plan looks like, change sample-chat-small to sample-chat-mini, run again, and look for Plan: sample-chat-mini is not in this tenant's model catalog. Change it back afterwards.
#Step 6 (optional): Run it against your SAP AI Core trial
This step reads your own tenant. It needs the trial and the AICORE_ lines in .env from Set up for Unit 5. It makes four GET requests and no model calls, so it has no token cost.
List the models your tenant offers. This shows real names, versions and retirement dates:
You should see one line per model version, like this (names and dates depend on SAP's catalog on the day you run it):
<model name> <version> retires - latest
Open unit10/ops_plan.json. Replace the two sample model names and versions with the models your Unit 5 and Unit 9 scripts use, copied from that list. Set peak_rpm to a number you can defend; if you have traces from Observability for AI systems, use their busiest minute. Set "environment" to "test", because a trial is not production. Save.
Run the live check:
python unit10/ai_core_preflight.py --live
What success looks like: a Preflight for ... live header, an OK for a running orchestration deployment, and one line per plan model. If the trial doesn't allow the rate-limit endpoint, you see a line like (GET /admin/quota/model returned HTTP 403; skipping that check) and an INFO line instead of rate-limit results. That is a valid result: ask your administrator for the limits instead.
An empty result means something too. If you see Orchestration deployments running: 0 with FAIL, your resource group has no running orchestration deployment. Check AICORE_RESOURCE_GROUP in .env; the course uses default.
The default plan: app name, environment, whether it uses orchestration, warning and fail days, headroom, and each model with its version and peak RPM.
sample_data
Made-up responses in the same shape as SAP's documented lists: deployments, configurations, the model catalog and rate limits. Dates are relative to today.
live_data
Gets a token with client credentials, then reads the same four lists with the AI-Resource-Group header. The rate-limit list is optional; any error there becomes a skip, not a crash.
read_date
Reads the first ten characters of retirementDate as a date, or returns None.
check_version
Finds the pinned version (or the one marked isLatest) in the catalog and grades it by retirement date and deprecated flag.
run_checks
Runs the checks in order: resource group, deployment statuses, models behind running direct deployments, plan models, rate limits against peak.
Rollback message
For a DEAD deployment whose latestRunningConfigurationId differs from its current configuration, prints the exact PATCH body that rolls it back. It does not send it.
main
Reads the options, creates the plan on the first run, prints results worst first, saves the report and returns exit code 2 if anything failed.
Your script reads the same APIs SAP's tools use. This section names the SAP pieces a production landscape adds around them. Capabilities are as documented in October 2026.
Extended plan. Production with the generative AI hub. Enterprise account, SAP support with an SLA. Billing is per resource plus model usage: tokens are converted to capacity units at per-model rates listed in SAP Note 3437766.
Standard plan. Production without generative AI. You can upgrade from standard to extended; you can't downgrade.
Trial. SAP's service plan page points to a 30-day trial for exploring the generative AI hub. It is for learning, not for production traffic.
Model availability differs by region. Check SAP Discovery Center for the regions of SAP AI Core and SAP Note 3437766 for which models run where, before you choose the region of your production subaccount.
Model restriction. On an orchestration deployment, set the parameter bindings modelFilterList (model names and, optionally, versions) and modelFilterListType (allow or deny). With an allow list, the deployment refuses any model your architecture board hasn't approved, whatever the app requests. Configuration sketch, following SAP's documented shape:
Credential rotation is the customer's job. Use x.509 service keys if your policy prefers certificates with a set validity.
Encryption. SAP AI Core encrypts data with tenant-specific keys. Since 7 September 2026 it integrates with SAP Data Custodian Key Management Service, so you can supply customer-managed root keys. SAP warns that disabling a root key makes data encrypted with it inaccessible and can disrupt connected systems: treat that key as production infrastructure.
Audit logging. Deployment, execution, resource-group and secret changes are written to the SAP BTP audit log. Route it to your security monitoring.
SAP's reference architecture DevOps with SAP BTP puts three services in the path to production. SAP Continuous Integration and Delivery builds and tests. SAP Cloud Transport Management moves approved changes between subaccounts. SAP Cloud ALM is where you operate. For the AI part, treat SAP AI Core configurations as code: keep the JSON in Git, create configurations from the pipeline, patch deployments from the pipeline, and run a preflight like yours before and after.
SAP Alert Notification service for SAP BTP subscribes to platform events and has a producer API for your own alerts, delivered by email, Slack or to other tools. A scheduled preflight can post its FAIL lines there.
SAP Cloud ALM gives a central operations view across SAP products, including SAP BTP.
The consumption report in the SAP BTP cockpit shows generative AI usage per model, orchestration and resource group.
Platform status is on the SAP Trust Center. Incidents go to SAP support on component CA-ML-AIC, with the region and subaccount.
SAP's Multi-Region HA/DR Resiliency for SAP BTP reference architecture explains the baseline: BTP services run across several availability zones in one region by default. To survive a regional outage you need two subaccounts in different regions, the app and its services in both, a DNS-based load balancer in front, and the SAP Custom Domain service for one stable URL. For AI that means two SAP AI Core instances, with their own deployments and rate limits, kept identical by your pipeline. SAP describes this as a reference design, not a turnkey product.
31 October 2026: SAP decommissions the first orchestration endpoint. Move to /v2/completion with version 2 payloads (SAP Note 3634540). This course has used version 2 since Unit 5.
Every pinned model version: its retirementDate, read from the catalog. Your preflight tracks these.
Security and SAP authorizations. The SAP AI Core credentials act for the whole tenant. Give the preflight its own read-only path where your setup allows it, and never give it the keys that create deployments. If the assistant also calls S/4HANA, the BTP destination and the user's own SAP authorizations decide what it may read, as in Grounding on SAP data without breaking authorizations. Unit 11 covers agent permissions in depth.
Evaluation. Every model version change is a release, even an automatic one. Keep the evaluation set from Unit 8 next to the app, run it on the new version in test, and keep the score in the change record.
Cost. A second region, a second resource group and inference observability all add cost. Price them per business transaction, as in Token economics. The AI business case, next in this unit, needs these operating costs, not only the build cost.
Operations. Write a runbook: one page per FAIL type, with the owner, the command and the check that confirms the fix. Schedule the preflight daily and before every release. Keep a week of reports.
Clean core. Nothing in this topic changes S/4HANA. The AI app, its configurations and its monitoring live on BTP, side by side. Keep it that way: SAP data reaches the model through released APIs, and any action that changes SAP data goes through a person, as in the agent units.
Treating PATCH as done. The response says "scheduled", not "running". Poll the status and smoke-test before you call a change finished.
Missing the 7-day window. A dead deployment is deleted seven days after it died, and with it the easy rollback. Alert on DEAD the same day.
Checking only deployments. Orchestration chooses the model per request. Models named in your app's templates must be in your plan, or their retirement surprises you.
One resource group for everything. Production then shares limits and clutter with every experiment, and the consumption report can't tell apps apart.
Stopping instead of deleting. Stopped deployments can't restart but still count against quota.
Rotating keys by deleting first. Create the new key, deploy, confirm, then delete the old one.
Logging payloads without masking. Inference observability stores what you send it. Decide what may be stored before you turn it on.
Copying limits into slides. Rate limits and model lists change. Read them from the API on the day you decide.
Turn the preflight into the operations section of the blocked-orders assistant. The AI business case topic, next in this unit, reuses your runbook's owners and operating costs.
Make a fresh copy of the sample plan, so your own ops_plan.json stays as it is. Run:
Open unit10/sample_plan.json and add a line after "headroom": 0.7,:
"max_stopped": 0,
Save.
Open unit10/ai_core_preflight.py. In run_checks, just before the line # 4. Models the app sends through orchestration, from the plan, add this check (indented like the lines around it):
stopped = [d for d in deployments if d.get("status") == "STOPPED"]
limit = plan.get("max_stopped", 0)
add("WARN" if len(stopped) > limit else "OK",
f"Stopped deployments: {len(stopped)} (plan allows {limit})")
Save.
Run python unit10/ai_core_preflight.py --sample --plan unit10/sample_plan.json. You should see WARN Stopped deployments: 1 (plan allows 0) and Summary: 4 OK, 2 INFO, 4 WARN, 2 FAIL.
Create unit10/runbook.md with one section per FAIL or WARN line the sample produced (there are now six). For each, write:
what the line means in one sentence;
who owns it (a role, not a name);
the exact action, including the API call or cockpit step from this topic;
how you confirm the fix (which preflight line turns OK).
Add a section Schedule to runbook.md: when the preflight runs (daily time, and before each release), where its FAIL lines go (for example the Alert Notification service), and who reads them.
Add a section Operating cost listing what production adds: the resource group, any second region, inference observability, and the hours per month for the runbook.
Commit:
git add unit10/ai_core_preflight.py unit10/sample_plan.json unit10/runbook.md
git commit -m "Unit 10: preflight check and runbook for the blocked-orders assistant"
Done when the sample run prints the new Stopped deployments line, runbook.md has an owner, action and confirmation for all six FAIL and WARN lines, and it names the schedule and the operating costs.
Pick one answer for each question. The explanation appears after you choose.
1Which of these changes under a running AI app on SAP AI Core without anyone touching the app's code?
Answer: B. Model versions follow the provider's and SAP's calendar, not your release plan. The Cloud Foundry settings, runtime and health check change only when you push.
2You patch a deployment with a new configuration and get "Deployment modification scheduled". What is true?
Answer: C. SAP documents that the update keeps the URL and that the deployment can still end in DEAD or stay PENDING. A change is finished only when the status is RUNNING and a smoke test passes.
3A deployment shows DEAD after an update, and latestRunningConfigurationId differs from configurationId. What do you do?
Answer: A. That field holds the last configuration that reached RUNNING, and patching with it is the documented rollback. After seven days SAP deletes the dead deployment instead of rolling it back.
4Why does the preflight need a plan file listing models, instead of only reading deployments?
Answer: D. An orchestration deployment serves whatever model each request names, so the models live in your app's templates. The plan file brings them into the check; direct foundation-model deployments show theirs in the configuration.
5In check_version, what happens when a plan model uses version latest?
Answer: B. The function finds the version flagged isLatest, still checks its retirement and deprecation, and otherwise returns INFO with a reminder to re-run the evaluation. The script never changes anything in SAP AI Core.
6Peak traffic is 50 requests per minute against a limit of 60. What does SAP's guidance suggest before asking for more?
Answer: C. SAP's rate-limit page lists backoff with jitter, spreading requests, caching, model fallbacks and isolating workloads, then a quota request. Removing retries or sharing the default group makes the problem worse.
7The security team asks how SAP AI Core credentials are rotated. What is the accurate answer?
Answer: C. SAP states that rotating credentials and certificates is the customer's responsibility. Creating the new key before deleting the old one avoids downtime; x.509 keys add a validity period but don't change who rotates.
8Your team wants to store every prompt and response with inference observability for later analysis. What must you settle first?
Answer: B. SAP warns that inference observability stores request, response and feedback without masking, and consent for personal data is the customer's responsibility. Settle that first, or store metadata only.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Model Lifecycle (SAP AI Core documentation, SAP-docs on GitHub)— model versions have deprecation dates; a deployment with a pinned version stops working on that date; auto upgrade with modelVersion latest; manual upgrade by patching with a new configuration; latest is the default
Choose a Model (SAP AI Core documentation, SAP-docs on GitHub)— GET /v2/lm/scenarios/foundation-models/models; versions with isLatest, deprecated, retirementDate; SAP Note 3437766 lists models, token conversion rates, rate limits and deprecation dates
Update a Deployment (SAP AI Core documentation, SAP-docs on GitHub)— PATCH configurationId keeps the URL; requests keep working during the transition; a bad configuration can end DEAD or stuck PENDING; latestRunningConfigurationId for rollback; DEAD deployments can be patched for 7 days, then are deleted
Service Plans (SAP AI Core documentation, SAP-docs on GitHub)— standard plan for production without generative AI, extended plan with generative AI hub; enterprise account, SLA; no downgrade from extended; deployment and replica quotas; 50 resource groups per tenant; ticket on CA-ML-AIC
Manage Resource Groups (SAP AI Core documentation, SAP-docs on GitHub)— default resource group created at onboarding and cannot be deleted; configurations, deployments, executions and artifacts belong to one group; scenarios, executables and Docker registry secrets are shared in the tenant
Rate Limit Management (SAP AI Core documentation, SAP-docs on GitHub)— requests-per-minute limits per model per tenant, shared by resource groups unless configured per group; all versions share a limit; 429 can also mean back-end saturation; Retry-After; GET /v2/admin/quota/model
What's New for SAP AI Core (SAP-docs on GitHub, read 7 October 2026)— first orchestration endpoint decommissioned on 31 October 2026, move to /v2/completion; data encryption with SAP Data Custodian KMS customer-managed root keys (2026-09-07); consumption report per model, orchestration and resource group in the BTP cockpit (2025-04-14); model restriction (2024-09-16)
Multi-Region HA/DR Resiliency for SAP BTP (SAP Architecture Center)— BTP services are multi-AZ by default; multi-region needs two subaccounts in different regions, a DNS load balancer and a custom domain; CI/CD keeps both in sync; reference design, not a turnkey solution