Language models are billed by the token, a small piece of text. A token is often a short word or part of a longer one. Every call is billed twice: once for the tokens you send (the input) and once for the tokens the model writes (the output). Output tokens usually cost more per token.
That makes a token bill easy to read and hard to predict. The useful question is not "what does a token cost?" but "what does one business transaction cost?" For example: what does it cost to explain one blocked sales order to a clerk?
Three numbers multiply to give the monthly bill:
Tokens per call: how long the prompt and the answer are.
Calls per transaction: one call for a simple design, several for an agent that works step by step.
Transactions per month: how many orders, invoices or questions go through.
Three habits keep the bill under control. Set a length budget for each part of the prompt. Cache the part of the prompt that never changes. Ask for answers no longer than they need to be.
Take the running example from this unit: an assistant that explains blocked sales orders to order-to-cash clerks. Each explanation needs the order data from SAP, a few passages from the team's notes, the instructions, and the question.
A cost estimate that stops at "a fraction of a cent per call" misses three things.
Designs multiply calls. An agent that first asks for the order, then for the notes, then answers makes three calls. Each call sends the whole conversation again. In this topic's cost model, the agent design costs about twice as much per order as a design where the app gathers the data and calls the model once.
Volume multiplies everything. A pilot with 200 orders hides what 20,000 orders a month cost. Month-end close can triple a normal week.
The model is only part of the bill. In SAP's own worked example for a retrieval assistant, the model calls are about 55% of the total. The rest is document storage, retrieval, content filtering and data masking.
Cost also shapes decisions that look technical. Retrieving ten passages instead of three makes every answer more expensive. So does a long set of instructions. A cheaper model may need more calls to reach the same answer. Someone who owns the budget should see these trade-offs before they are built in.
As of October 2026, SAP prices AI in two ways, depending on who builds the feature.
AI you build on SAP BTP. Apps and agents that call models through the generative AI hub in SAP AI Core are metered in tokens. SAP converts input and output tokens into "GenAI tokens" at a rate per model, then into BTP capacity units. Your company pays for capacity units under its BTP contract. The generative AI hub is part of SAP AI Core's extended plan. Grounding, content filtering and data masking have their own meters.
AI that SAP builds into its applications. Joule and other SAP-delivered AI features are not billed per token to you. Some come with your cloud subscription; premium capabilities and agents consume AI Units, SAP's prepaid currency for those features. Joule and Joule Studio explains how AI Units are bought and used.
SAP's BTP cockpit shows usage and costs, and can send alerts when spending passes a budget. SAP's documentation is explicit that a budget alert does not stop the spending. Limits that actually stop spending have to be built into your app.
SAP also supports prompt caching in its orchestration service, so a long fixed part of a prompt can be reused across calls. How cached tokens count toward capacity units is in SAP's per-model rate note; ask your account team.
A cost model for one AI feature fits on one page. The figures below come from this topic's hands-on model with made-up rates. Use them for the shape, not the price.
Line
Blocked-orders assistant, simple design
Same, agent design
Model calls per order
1
3
Tokens sent per order (input)
about 1,300
about 3,700
Tokens written per order (output)
about 220
about 300
Cost per order, made-up rates
0.0048
0.0104
With caching of the fixed instructions
0.0034
0.0052
Read the table from the top. The agent design sends almost three times as many tokens for the same answer, because every step sends the conversation again. Caching narrows the gap because the repeated part is billed at a fraction of the normal rate. Neither number means anything until it is multiplied by a real monthly volume and set against the value of the work it saves.
A finished cost model adds three more lines: monthly volume (with peaks), non-model meters such as retrieval and filtering, and the value side, such as minutes saved per order.
"Tokens are words." A token is often shorter than a word. Numbers, codes and many non-English languages use more tokens for the same meaning.
"Input and output cost the same." Output usually costs more per token, so long answers are expensive.
"A cheaper model is always cheaper." If it needs more calls or more retries to reach the same answer, the cost per transaction can rise.
"Caching is free money." Writing to a cache costs extra on some models. If the cache expires before the next call, you pay the premium and get nothing back.
"The budget alert will protect us." On SAP BTP, a budget sends alerts; it does not stop usage.
"The model is the whole bill." Retrieval, filtering, masking and storage are metered too. In SAP's worked example they are close to half.
Pick one answer for each question. The explanation appears after you choose.
1A vendor quotes a price per million tokens for the blocked-orders assistant. What number should the finance lead ask for instead?
Answer: B. The business pays per transaction, and one transaction can involve several calls with different token counts. A per-token price is an input to the estimate, not the estimate.
2The team switches from a single model call to an agent that asks for the order, then the notes, then answers. What happens to cost per order?
Answer: C. Each step of an agent loop sends the instructions and everything gathered so far again. In this topic's model, the agent design costs about twice as much per order as the single-call design.
3Why should a cost estimate for an AI assistant on SAP BTP include more than the model calls?
Answer: A. SAP's own worked example adds grounding storage, retrieval, content filtering and data masking to the model calls, and they come to close to half the total. Leaving them out understates the bill.
4The team sets a budget with alerts in the SAP BTP cockpit. What does that protect against?
Answer: C. SAP's documentation says the system does not suspend costs or usage when a budget is exceeded. Hard limits such as per-user caps or a switch-off must be built into the application.
5When does prompt caching fail to save money?
Answer: D. Writing to the cache can cost more than a normal call on some models. If the cache expires before it is read, you pay the extra and get no discount back.
6A product owner asks for longer, more detailed answers "since they only cost a little more". What should you point out?
Answer: B. Output tokens are usually priced higher than input tokens, and caching applies to the prompt, not the answer. Length should be a decision about what the clerk needs.
7Your business users want Joule to explain blocked orders, instead of a custom app. How is that usage paid for?
Answer: D. SAP-delivered AI features are not billed to you per token. Some capabilities come with the cloud subscription, and premium capabilities and agents consume AI Units, as the Joule topic explains.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
The token bill is a sum over calls, and the unit that matters is the business transaction.
cost per transaction = sum over its calls of
(uncached input x input rate)
+ (cache writes x input rate x write multiplier)
+ (cache reads x input rate x read multiplier)
+ (output x output rate)
cost per month = cost per transaction x transactions per month
Everything in this topic is one of those terms. A longer prompt grows the input term. An agent loop adds calls, and each call carries the history of the ones before it. Caching moves tokens from the first term to the cheaper third. A shorter answer shrinks the last term. On SAP AI Core, the same sums are converted into capacity units at per-model rates, and other modules add their own meters beside them.
Two rules follow:
Measure per transaction, not per token. Anthropic's cost cookbook makes the same point: a model with a higher price per token can be cheaper per task if it finishes in fewer turns.
Count what you send, not what you mean. The model is billed for every token in the request, including tool definitions, examples and history you stopped thinking about long ago.
Every response tells you what was counted. In SAP's orchestration service (Python SDK sap-ai-sdk-gen 7.4.1, read for this topic), the usage object has these fields:
Field
Meaning
Billed as
prompt_tokens
Tokens sent to the model
Input
completion_tokens
Tokens the model wrote
Output
prompt_tokens_details.cached_tokens
Input tokens read from the prompt cache
Input, at the cache read rate where the provider discounts it
prompt_tokens_details.cache_creation_tokens
Input tokens written to the prompt cache
Input, at a premium on models that charge for writes
completion_tokens_details.reasoning_tokens
Tokens a reasoning model "thought" before answering
Output; the SDK notes such tokens are counted in completion tokens for billing
Tokenizers split text into pieces from a fixed vocabulary. The tiktoken README says a token is about 4 bytes of text on average. That rule works for planning English prompts. It breaks for:
Other languages. German umlauts take two bytes in UTF-8, and many scripts take three. A prompt translated for German or Japanese users can cost noticeably more.
Codes and numbers. Order numbers, material numbers and JSON punctuation split into many small tokens.
Different models. Each model family has its own tokenizer. The same prompt has a different count on each.
Use the 4-byte rule to plan, then confirm with the usage numbers a real call returns. This topic's script prints both.
SAP's metering page says output tokens "tend to be slightly more expensive" than input tokens. In SAP's fictitious example the output rate is almost three times the input rate. In the price snapshot in Anthropic's cookbook, output costs five times input for every model listed. Either way, a long answer costs more than a long prompt of the same length.
The ratio of input to output depends on the job. SAP's "typical consumption patterns" table, from the same page:
Use case
Input tokens per request
Output tokens per request
RAG chat
3,500
300
Basic chat
500
100
Summarization
5,000
300
Classification
3,800
10
Generation
500
3,500
Retrieval and classification are input-heavy, so prompt length budgets and caching matter most. Generation is output-heavy, so answer length matters most.
Most prompts start with a long part that never changes: instructions, examples and tool definitions. Prompt caching lets the provider keep that processed prefix for a few minutes and reuse it.
Provider behaviour (from the sources opened for this topic)
Anthropic Claude
OpenAI
How it turns on
cache_control on a block, or automatic at the top level
Automatic for prompts over 1,024 tokens
Minimum prefix
1,024 tokens for Sonnet, 4,096 for Opus and Haiku 4.5
1,024 tokens
Lifetime
5 minutes, refreshed on each hit; 1 hour optional
Not stated in the cookbook
Price of a write
1.25 times input (5 minutes), 2 times (1 hour)
No write premium stated
Price of a read
0.1 times input on the cookbook's models, lower on some newer ones
"Discounted"; no figure in the cookbook
Three things decide whether caching pays:
The prefix must be byte-identical. A timestamp, a user name or a request ID in the system prompt makes every prefix different. Put everything that varies after the fixed part.
The prefix must be long enough. Below the minimum, nothing is cached.
Calls must come before the cache expires. With a write premium, a cache that is never read costs more than no cache at all.
A pipeline design reads SAP and the notes in code, then makes one call. An agent design lets the model ask for each piece. Each step resends everything:
flowchart LR
subgraph Pipeline
P1[Call 1<br/>prefix + order + notes + question] --> A1[Answer]
end
subgraph Agent
C1[Call 1<br/>prefix + question] --> C2[Call 2<br/>prefix + question + order]
C2 --> C3[Call 3<br/>prefix + question + order + notes]
C3 --> A2[Answer]
end
With a fixed prefix of about 1,060 tokens, the agent sends that prefix three times. That is why caching helps agents most: calls two and three come seconds after call one, so they read the prefix from the cache. Anthropic's cookbook adds that intermediate results, such as a large tool output, stay in the context on every later turn unless you remove them.
A prompt length budget gives each part of the prompt a ceiling, the same way a latency budget gives each step a time. For the blocked-orders assistant:
Part
Budget (tokens)
Why it can grow
System instructions and examples
900
Each new rule or example is added "just in case"
Tool definitions
450
Every new tool adds its description and schema to every call
Order data
250
Orders with many items; fields nobody uses
Retrieved notes
200
Raising the number of passages retrieved
Question
60
Pasted emails or documents
Answer (max_tokens)
250
Requests for "more detail"
A budget turns a slow drift into a visible failure. When the notes part goes over, the team decides: fewer passages, shorter passages, or a bigger budget and a higher cost per order.
Anthropic's cookbook treats max_tokens as a backstop for runaway answers, not as the way to shorten normal ones. A model cut off at max_tokens stops mid-sentence. To shorten typical answers, ask for a specific shape and length in the prompt.
SAP's metering page walks through the conversion with fictitious values. Per request:
GenAI tokens = 3,500 / 1,000 x 0.00112 + 300 / 1,000 x 0.00320 = 0.00488
capacity units = 0.00488 x 1.90385 = 0.00929
For 25,000 requests a month, SAP's example totals 232.25 for the model calls. Grounding storage, grounding retrieval, content filtering and data masking add 186.3 CU, for 418.55 CU in all. The model calls are about 55% of the total. The real conversion rates per model are in SAP Note 3437766, and inference observability storage has its own conversion in SAP Note 3720903.
#Build it yourself: a cost model for the blocked-orders assistant
You will build a cost model that counts the tokens in each part of the assistant's prompt, checks them against a length budget, and prices one explained order and one month. It compares the pipeline design with the agent design, with and without prompt caching. You will then switch to SAP's capacity-unit arithmetic, price the token counts you logged in the observability topic, and optionally read the real token usage of one call in SAP AI Core.
flowchart LR
D[Made-up order,<br/>notes, instructions] --> C[Count tokens<br/>per part]
C --> B{Length budget}
B --> P[Price per order<br/>pipeline vs agent]
P --> M[Price per month]
T[traces.jsonl] --> P
Your course folder with .venv from the earlier setup topics.
About 40 to 60 minutes.
Cost: free. No model is called and no account is needed, except in the optional Step 9, which uses your SAP AI Core trial (no charge during the trial; a small per-request charge after it).
#Step 1: Open your course folder and turn on the virtual environment
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
If the prompt doesn't start with (.venv), turn it on:
Windows (PowerShell):
.venv\Scripts\Activate.ps1
macOS / Linux:
source .venv/bin/activate
Check your Python version (the same on every system):
The script holds the assistant's instructions, tool definitions, one made-up blocked order and ten short notes. It counts the tokens in each part with the 4-byte rule, then prices the designs.
In VS Code, right-click unit10, choose New File, name it token_costs.py, paste the code below and save.
"""Unit 10: what does one blocked order cost in tokens, and what does a month cost?
The blocked-orders assistant explains why a sales order is blocked and what the clerk should
check next. This script builds its prompt from made-up, SAP-shaped data, counts the tokens in
each part, checks them against a prompt length budget, and prices one business transaction
("explain one blocked order") and one month, for two designs:
pipeline your code reads the order and the notes, then makes one model call
agent the model asks for the order, then for the notes, then answers: three calls,
and every call sends the whole conversation so far
With no options it uses EXAMPLE rates (made up, see RATES below), calls no model and costs nothing.
python unit10/token_costs.py both designs, with and without prompt caching
python unit10/token_costs.py --notes 10 retrieve 10 note chunks instead of 3
python unit10/token_costs.py --answer-tokens 120 ask for shorter answers
python unit10/token_costs.py --orders 20000 a busier month
python unit10/token_costs.py --sap SAP AI Core arithmetic: GenAI tokens and capacity units
python unit10/token_costs.py --traces unit10/traces.jsonl price the token counts you logged
python unit10/token_costs.py --llm --model NAME one real call in SAP AI Core; print its usage
"""
import argparse
import json
import math
import os
import sys
from pathlib import Path
# ---------------------------------------------------------------------------------------------
# Rates. EXAMPLE VALUES ONLY: replace them with your provider's price list or your SAP contract.
# ---------------------------------------------------------------------------------------------
RATES = {
# Provider-style list prices, in US dollars per million tokens (made up for this course).
"input_per_million": 2.00,
"output_per_million": 10.00,
# Prompt caching multipliers on the input price. These are the ones Anthropic documents for
# its 5-minute cache; other providers use other numbers. Check yours.
"cache_write_multiplier": 1.25,
"cache_read_multiplier": 0.10,
"min_cacheable_tokens": 1024,
# SAP AI Core style: GenAI tokens per 1,000 model tokens, then capacity units per GenAI token.
# These are the fictitious values from SAP's own worked example; real ones are per model,
# in SAP Note 3437766.
"sap_genai_per_1k_input": 0.00112,
"sap_genai_per_1k_output": 0.00320,
"sap_cu_per_genai_token": 1.90385,
}
# The prompt length budget: the most tokens each part of the prompt may use.
BUDGET = {"system": 900, "tools": 450, "order": 250, "notes": 200, "question": 60}
ANSWER_CAP = 250 # max_tokens: the answer is cut off here
# ---------------------------------------------------------------------------------------------
# Made-up, SAP-shaped data. Block reason and status codes are configured per SAP system;
# the meanings below belong to this made-up company only.
# ---------------------------------------------------------------------------------------------
SYSTEM_PROMPT = """You help order-to-cash clerks understand blocked sales orders in SAP S/4HANA.
Answer in plain English, in at most five short sentences, then give one next step.
Use only the order data and the notes you are given. If they do not explain the block, say so.
Never promise a release date. Never release, change or cancel an order yourself.
This company's guide to delivery block reasons (configured in its own SAP system):
- 01 Credit limit: the customer's open items exceed the credit limit. Credit management decides.
- 02 Missing export papers: shipping waits for the customs team to attach the documents.
- 03 Pricing check: a manual price or discount is above the sales rep's limit.
- 04 Customer request: the customer asked us to hold the delivery. Check the order notes.
- 05 Quality hold: one or more items wait for a quality inspection result.
- 06 Dangerous goods: the safety data sheet is missing or out of date for this country.
- 07 Incomplete order: required fields such as the incoterms or the ship-to address are empty.
- 08 Advance payment: the customer must pay before delivery. Check incoming payments.
This company's credit check statuses: A means not checked or released, B means a check failed,
C means partially released. Explain a status in words; do not show the letter alone.
Style rules: name the order number in the first sentence. Use the customer's name, not the ID.
Quote amounts with their currency. If two reasons apply, explain the delivery block first.
If the notes contradict the order data, trust the order data and mention the conflict.
Example 1. Question: Why is order 9000077 blocked? Data: DeliveryBlockReason 02, credit status A.
Answer: Order 9000077 for Contoso Freight is waiting for export papers. Shipping cannot start
until the customs team attaches them. The credit check is fine. Next step: ask the customs team
whether the papers for this order have been requested.
Example 2. Question: What is wrong with order 9000081? Data: no delivery block, credit status C.
Answer: Order 9000081 for Fabrikam Retail has no delivery block. Its credit check is partially
released, so some items may ship while others wait. Next step: check in the credit management
app which items are still held, and tell the customer which ones will ship first.
Example 3. Question: Can order 9000090 ship today? Data: DeliveryBlockReason 05, credit status A.
Answer: Order 9000090 for Litware Medical is on a quality hold. At least one item is waiting for
an inspection result, so it cannot ship until quality releases it. Credit is not the issue.
Next step: ask the quality team when the inspection lot for this order will be decided.
Answer format: first the reason in plain words, then what it means for the customer, then
"Next step:" and one action the clerk can take today. No headings, no lists, no codes alone.
"""
TOOLS = [
{"name": "get_sales_order",
"description": "Read one sales order from SAP: header, block reasons, credit status and items. "
"Use it before explaining any order. Read-only.",
"parameters": {"type": "object", "properties": {
"sales_order": {"type": "string", "description": "The sales order number, for example 9000123"}},
"required": ["sales_order"]}},
{"name": "search_notes",
"description": "Search the order-to-cash team's notes for how to handle a block reason. "
"Returns the most relevant short passages with their source.",
"parameters": {"type": "object", "properties": {
"query": {"type": "string", "description": "What to look for, in plain words"},
"top_k": {"type": "integer", "description": "How many passages to return, 1 to 10"}},
"required": ["query"]}},
{"name": "get_delivery_status",
"description": "Read the outbound deliveries for a sales order and their goods issue status. "
"Use it only when the clerk asks about shipping. Read-only.",
"parameters": {"type": "object", "properties": {
"sales_order": {"type": "string", "description": "The sales order number"}},
"required": ["sales_order"]}},
{"name": "get_customer_credit",
"description": "Read the customer's credit limit, credit exposure and open items. Read-only.",
"parameters": {"type": "object", "properties": {
"customer": {"type": "string", "description": "The customer number (sold-to party)"}},
"required": ["customer"]}},
]
ORDER = {
"SalesOrder": "9000123", "SoldToParty": "10100042", "CustomerName": "Northwind Tools GmbH",
"SalesOrderDate": "2026-10-05", "TotalNetAmount": "48200.00", "TransactionCurrency": "EUR",
"DeliveryBlockReason": "01", "TotalCreditCheckStatus": "B",
"Items": [
{"SalesOrderItem": "10", "Material": "TG-11", "RequestedQuantity": "40", "NetAmount": "28800.00"},
{"SalesOrderItem": "20", "Material": "TG-12", "RequestedQuantity": "20", "NetAmount": "19400.00"},
],
}
NOTES = [
"Credit blocks (reason 01): check the customer's exposure in the credit management app first. "
"If a payment is on its way, ask credit management for a one-time release.",
"Credit check status B after a large new order usually means the order itself pushed exposure over "
"the limit. Splitting the order is not a fix; talk to credit management.",
"Key accounts: for customers in the key account list, the account manager can request a temporary "
"limit increase. Typical turnaround is one business day.",
"Advance payment customers (reason 08) are a different process. Do not ask credit management.",
"If the customer disputes an open invoice, the disputed amount still counts toward exposure "
"until the dispute case is closed.",
"Month-end: credit releases are slower in the last two working days of the month. Tell the "
"customer before promising a date.",
"When a blocked order has items from two plants, the block applies to the whole order, not per item.",
"Customers on payment terms longer than 60 days have a lower credit limit by policy.",
"If the credit limit was changed today, the order may need a new credit check before the block clears.",
"Escalation: if a credit block is older than five working days, inform the sales team lead.",
]
QUESTION = "Why is sales order 9000123 blocked and what should I check next?"
def count_tokens(text: str) -> int:
"""Estimate tokens: about 4 bytes of text per token (the tiktoken README's rule of thumb).
Real counts differ by model and language; --llm shows a real count."""
return max(1, math.ceil(len(text.encode("utf-8")) / 4))
def prompt_parts(n_notes: int) -> dict:
"""The text of each part of the pipeline design's prompt."""
notes = "\n".join(f"[note {i + 1}] {t}" for i, t in enumerate(NOTES[:n_notes]))
return {"system": SYSTEM_PROMPT, "tools": json.dumps(TOOLS), "order": json.dumps(ORDER),
"notes": notes, "question": QUESTION}
def calls_for(design: str, parts: dict, answer_tokens: int) -> list:
"""Return the model calls one transaction makes, as token counts.
Each call: prefix (system + tools, the same every time), rest (everything else sent), out."""
t = {k: count_tokens(v) for k, v in parts.items()}
prefix = t["system"] + t["tools"]
if design == "pipeline":
return [{"prefix": prefix, "rest": t["order"] + t["notes"] + t["question"], "out": answer_tokens}]
# Agent: call 1 asks for the order, call 2 asks for the notes, call 3 answers.
# Every call re-sends what came before: that is why agent loops cost more input tokens.
tool_call = 40 # tokens the model writes to request a tool
history = t["question"]
calls = [{"prefix": prefix, "rest": history, "out": tool_call}]
history += tool_call + t["order"]
calls.append({"prefix": prefix, "rest": history, "out": tool_call})
history += tool_call + t["notes"]
calls.append({"prefix": prefix, "rest": history, "out": answer_tokens})
return calls
def price(calls: list, cache: bool, hit_rate: float, sap: bool) -> dict:
"""Price one transaction. With cache=True the prefix is read from the cache on a hit and
written to it on a miss. In an agent loop, calls after the first are seconds apart, so they
hit the cache for the prefix (and in real systems often for the history as well)."""
r = RATES
total = {"input": 0.0, "cache_write": 0.0, "cache_read": 0.0, "output": 0.0}
for i, c in enumerate(calls):
total["output"] += c["out"]
if not cache or sap or c["prefix"] < r["min_cacheable_tokens"]:
total["input"] += c["prefix"] + c["rest"]
continue
hit = 1.0 if i > 0 else hit_rate
total["cache_read"] += c["prefix"] * hit
total["cache_write"] += c["prefix"] * (1 - hit)
total["input"] += c["rest"]
if sap:
genai = (total["input"] / 1000 * r["sap_genai_per_1k_input"]
+ total["output"] / 1000 * r["sap_genai_per_1k_output"])
total["cost"] = genai * r["sap_cu_per_genai_token"]
else:
pin, pout = r["input_per_million"] / 1e6, r["output_per_million"] / 1e6
total["cost"] = (total["input"] * pin + total["cache_write"] * pin * r["cache_write_multiplier"]
+ total["cache_read"] * pin * r["cache_read_multiplier"] + total["output"] * pout)
return total
def show_budget(parts: dict, answer_tokens: int) -> None:
print("Prompt length budget, pipeline design (estimated tokens)\n")
print(f"{'part':<10}{'tokens':>8}{'budget':>8} check")
for name, text in parts.items():
n = count_tokens(text)
print(f"{name:<10}{n:>8}{BUDGET[name]:>8} {'OK' if n <= BUDGET[name] else 'OVER'}")
total_in = sum(count_tokens(t) for t in parts.values())
print(f"{'input':<10}{total_in:>8}{sum(BUDGET.values()):>8} {'OK' if total_in <= sum(BUDGET.values()) else 'OVER'}")
over = answer_tokens > ANSWER_CAP
print(f"{'answer':<10}{answer_tokens:>8}{ANSWER_CAP:>8} {'OVER: raise max_tokens or ask for less' if over else 'OK'}")
prefix = count_tokens(parts["system"]) + count_tokens(parts["tools"])
if prefix < RATES["min_cacheable_tokens"]:
print(f"\nThe fixed prefix is {prefix} tokens, under the {RATES['min_cacheable_tokens']}-token minimum many "
"providers need before they cache. Caching will not help this prompt.")
def model_costs(args) -> None:
parts = prompt_parts(args.notes)
show_budget(parts, args.answer_tokens)
unit = "CU" if args.sap else "USD"
print(f"\nCost per transaction and per month ({args.orders:,} blocked orders, "
f"{'SAP example conversion rates' if args.sap else 'EXAMPLE rates'})\n")
print(f"{'design':<10}{'cache':<7}{'calls':>6}{'input':>8}{'cached':>8}{'output':>8}"
f"{'per order':>12}{'per month':>12}")
for design in ("pipeline", "agent"):
calls = calls_for(design, parts, args.answer_tokens)
for cache in (False, True):
if cache and args.sap:
continue
p = price(calls, cache, args.hit_rate, args.sap)
cached = p["cache_read"] + p["cache_write"]
per_order = f"{p['cost']:.5f} {unit}"
per_month = f"{p['cost'] * args.orders:,.2f} {unit}"
print(f"{design:<10}{'on' if cache else 'off':<7}{len(calls):>6}{p['input']:>8.0f}{cached:>8.0f}"
f"{p['output']:>8.0f}{per_order:>12}{per_month:>12}")
print("\ninput = tokens billed at the full input rate; cached = prefix tokens written to or read from "
"the cache;\noutput = tokens the model writes. Rates are examples: replace RATES with your own.")
if args.sap:
print("SAP mode shows capacity units (CU). What a CU costs depends on your SAP BTP contract.\n"
"Cached tokens are priced as normal input here: SAP's metering page does not say how they convert.")
def price_traces(path: Path, orders: int) -> None:
"""Price the token counts logged by observe_assistant.py, grouped by prompt version."""
if not path.exists():
sys.exit(f"No file {path}. Run 'python unit10/observe_assistant.py' from the observability "
"topic first, or leave out --traces.")
spans = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
version = {s["trace_id"]: s["attributes"].get("app.prompt.version", "?")
for s in spans if s["parent_id"] is None}
groups = {}
for s in spans:
a = s["attributes"]
if a.get("gen_ai.operation.name") == "chat":
g = groups.setdefault(version.get(s["trace_id"], "?"), {"calls": 0, "in": 0, "out": 0})
g["calls"] += 1
g["in"] += a.get("gen_ai.usage.input_tokens", 0)
g["out"] += a.get("gen_ai.usage.output_tokens", 0)
if not groups:
sys.exit(f"{path} has no chat spans with token counts.")
pin, pout = RATES["input_per_million"] / 1e6, RATES["output_per_million"] / 1e6
print(f"Logged model calls in {path.name}, priced at EXAMPLE rates\n")
print(f"{'prompt':<8}{'calls':>7}{'avg in':>8}{'avg out':>9}{'per call':>12}{'per month':>12}")
for v, g in sorted(groups.items()):
cost = (g["in"] * pin + g["out"] * pout) / g["calls"]
print(f"{v:<8}{g['calls']:>7}{g['in'] / g['calls']:>8.0f}{g['out'] / g['calls']:>9.0f}"
f"{cost:>8.5f} USD{cost * orders:>8,.2f} USD")
print(f"\nper month = per call x {orders:,} calls. Change it with --orders.")
def real_call(args) -> None:
"""One real call through SAP's orchestration service; print the usage SAP reports."""
from dotenv import load_dotenv
load_dotenv()
needed = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
"AICORE_RESOURCE_GROUP"]
missing = [n for n in needed if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', or leave out --llm.")
if not args.model:
sys.exit("Name a model with --model, for example one from 'python unit05/choose_model.py catalog'.")
from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
OrchestrationService, PromptTemplatingModuleConfig,
SystemMessage, Template, UserMessage)
parts = prompt_parts(args.notes)
user = f"Order data: {parts['order']}\n\nNotes:\n{parts['notes']}\n\n{{{{?question}}}}"
template = Template(template=[SystemMessage(content=SYSTEM_PROMPT), UserMessage(content=user)])
config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=args.model, params={"max_tokens": args.answer_tokens},
timeout=60, max_retries=0))))
try:
service = OrchestrationService(config=config)
response = service.run(placeholder_values={"question": QUESTION}, timeout=60)
except Exception as error:
sys.exit(f"The call failed ({type(error).__name__}: {error}). "
"Run python check_unit05.py to find the cause, or leave out --llm.")
usage = response.final_result.usage
estimate = sum(count_tokens(t) for k, t in parts.items() if k != "tools")
print(f"Model: {response.final_result.model}\n")
print(f"prompt_tokens {usage.prompt_tokens:>6} (this script estimated {estimate})")
print(f"completion_tokens {usage.completion_tokens:>6}")
print(f"total_tokens {usage.total_tokens:>6}")
details = usage.prompt_tokens_details
if details:
print(f"cached_tokens {details.cached_tokens or 0:>6}")
print(f"cache_creation {details.cache_creation_tokens or 0:>6}")
if details.cached_tokens or details.cache_creation_tokens:
print("Cache tokens appeared. Check how your model counts them (inside prompt_tokens or not) "
"and how your contract prices them before you add them up.")
reasoning = usage.completion_tokens_details and usage.completion_tokens_details.reasoning_tokens
if reasoning:
print(f"reasoning_tokens {reasoning:>6} (billed as output, usually inside completion_tokens)")
one = price([{"prefix": 0, "rest": usage.prompt_tokens, "out": usage.completion_tokens}], False, 0, False)
print(f"\nAt the EXAMPLE rates this call costs {one['cost']:.5f} USD.")
print("\n" + response.final_result.choices[0].message.content)
def main() -> None:
parser = argparse.ArgumentParser(description="Token cost model for the blocked-orders assistant.")
parser.add_argument("--orders", type=int, default=6000, help="blocked orders explained per month")
parser.add_argument("--notes", type=int, default=3, help="note chunks retrieved per order (1 to 10)")
parser.add_argument("--answer-tokens", type=int, default=220, help="tokens in a typical answer")
parser.add_argument("--hit-rate", type=float, default=0.8,
help="share of first calls that find the prefix in the cache (0 to 1)")
parser.add_argument("--sap", action="store_true", help="convert with SAP AI Core's example rates")
parser.add_argument("--traces", type=Path, help="price the token counts in a traces.jsonl file")
parser.add_argument("--llm", action="store_true", help="make one real call in SAP AI Core")
parser.add_argument("--model", help="model name for --llm")
args = parser.parse_args()
if not 1 <= args.notes <= len(NOTES):
sys.exit(f"--notes must be between 1 and {len(NOTES)}.")
if not 0 <= args.hit_rate <= 1:
sys.exit("--hit-rate must be between 0 and 1.")
if args.traces:
price_traces(args.traces, args.orders)
elif args.llm:
real_call(args)
else:
model_costs(args)
if __name__ == "__main__":
main()
Check that the file is in the right place. In the terminal, list the folder:
Windows (PowerShell):
dir unit10\token_costs.py
macOS / Linux:
ls unit10/token_costs.py
You should see the file name. An error means the file is in another folder or has another name.
Prompt length budget, pipeline design (estimated tokens)
part tokens budget check
system 712 900 OK
tools 349 450 OK
order 114 250 OK
notes 129 200 OK
question 16 60 OK
input 1320 1860 OK
answer 220 250 OK
Cost per transaction and per month (6,000 blocked orders, EXAMPLE rates)
design cache calls input cached output per order per month
pipeline off 1 1320 0 220 0.00484 USD 29.04 USD
pipeline on 1 259 1061 220 0.00342 USD 20.51 USD
agent off 3 3708 0 300 0.01042 USD 62.50 USD
agent on 3 525 3183 300 0.00517 USD 31.05 USD
input = tokens billed at the full input rate; cached = prefix tokens written to or read from the cache;
output = tokens the model writes. Rates are examples: replace RATES with your own.
Read it from the top:
The budget table shows each part of the prompt. All are within budget. The answer is 220 tokens against a cap of 250.
In the cost table, the pipeline design without caching sends 1,320 input tokens and writes 220. The agent design makes 3 calls and sends 3,708 input tokens for the same answer.
With caching, the 1,061-token fixed prefix moves to the cached column. The agent design's cost per order drops by about half, because calls two and three read the prefix from the cache.
The script assumes that 80% of first calls find the prefix in the cache (--hit-rate 0.8). Try the worst case, where the cache has always expired:
Run:
python unit10/token_costs.py --hit-rate 0
Compare the pipeline rows:
design cache calls input cached output per order per month
pipeline off 1 1320 0 220 0.00484 USD 29.04 USD
pipeline on 1 259 1061 220 0.00537 USD 32.22 USD
agent off 3 3708 0 300 0.01042 USD 62.50 USD
agent on 3 525 3183 300 0.00713 USD 42.76 USD
With no hits, the pipeline design pays the 1.25 times write premium on every call and gets nothing back: caching costs more than no caching. The agent still gains, because its later calls come seconds after the first.
What this means for you: caching pays when the same prefix is used again within its lifetime. A clerk working through a queue does that. An assistant used a few times a day does not.
design cache calls input cached output per order per month
pipeline off 1 1320 0 120 0.00384 USD 76.80 USD
pipeline on 1 259 1061 120 0.00242 USD 48.37 USD
agent off 3 3708 0 200 0.00942 USD 188.32 USD
agent on 3 525 3183 200 0.00417 USD 83.49 USD
Cutting 100 output tokens saved 0.001 USD per order in the pipeline design, the same as cutting 500 input tokens would at these rates. Output is the expensive side.
Try --answer-tokens 300. The budget table reports OVER: raise max_tokens or ask for less: the answer would be cut off at 250 tokens.
The cost table now shows capacity units, using the fictitious conversion rates from SAP's metering page:
design cache calls input cached output per order per month
pipeline off 1 1320 0 220 0.00415 CU 24.93 CU
agent off 3 3708 0 300 0.00973 CU 58.41 CU
There are no caching rows. SAP's metering page doesn't say how cached tokens convert, so the script refuses to guess. Look up your model's rates in SAP Note 3437766 (you need an SAP for Me login) and put them in RATES.
Check the arithmetic for the pipeline row by hand: 1,320 / 1,000 x 0.00112 + 220 / 1,000 x 0.00320 = 0.0021824 GenAI tokens. Times 1.90385 gives 0.00415 CU.
The observability topic's script writes unit10/traces.jsonl, with the input and output tokens of every model call. Its --incident option switches half the requests to prompt version v2, which sends the full order history.
If you don't have the file, or want the incident in it, run the observability script again:
Logged model calls in traces.jsonl, priced at EXAMPLE rates
prompt calls avg in avg out per call per month
v1 98 819 94 0.00257 USD 15.44 USD
v2 99 1638 83 0.00410 USD 24.61 USD
per month = per call x 6,000 calls. Change it with --orders.
Prompt v2 doubled the input tokens and raised the cost per call by about 60%. In the observability topic, the same release was the one where answers got worse. A change that costs more and helps less is exactly what token logging is there to catch.
If you see No file unit10/traces.jsonl, run step 1 first.
#Step 9 (optional): Read the usage of one real call in SAP AI Core
This step calls a real model, so it needs your SAP AI Core trial and the AICORE_ lines in .env from Set up for Unit 5. During the trial there is no charge; after it, each call has a small per-request charge.
Check that Unit 5 still works:
python check_unit05.py
Pick a model name your account offers. check_unit05.py prints some, or run python unit05/choose_model.py catalog from the model topic.
Make one call, putting your model's name in place of MODEL_NAME:
You should see the usage SAP reports, the script's own estimate, and the answer. Your numbers depend on the model; this shape is what matters:
Model: MODEL_NAME
prompt_tokens 1012 (this script estimated 971)
completion_tokens 180
total_tokens 1192
cached_tokens 0
cache_creation 0
At the EXAMPLE rates this call costs 0.00382 USD.
Order 9000123 for Northwind Tools GmbH is blocked ...
Compare prompt_tokens with the estimate. The gap is the 4-byte rule being off for this model's tokenizer, plus any formatting added around the messages. A reasoning model also shows reasoning_tokens, billed as output.
The sample numbers above are illustrative. This step was not run against a live SAP AI Core account when this topic was written.
All prices in one place: provider-style rates, cache multipliers, minimum cacheable prefix, and SAP's example conversion rates
BUDGET, ANSWER_CAP
The prompt length budget per part and the answer cap (max_tokens)
SYSTEM_PROMPT, TOOLS, ORDER, NOTES
Made-up, SAP-shaped content; the system prompt and tools are the fixed prefix
count_tokens
Estimates tokens as UTF-8 bytes divided by 4, the tiktoken README's rule of thumb
calls_for
Lists the model calls in one transaction. The agent design adds the order and the notes to the history call by call
price
Applies the cost formula: uncached input, cache writes, cache reads and output. In SAP mode, converts to GenAI tokens and capacity units, without caching
show_budget
Prints each part against its budget and warns when the prefix is too short to cache
price_traces
Reads traces.jsonl, finds each request's prompt version on its root span, and averages tokens per model call
real_call
With --llm, makes one call through SAP's orchestration service and prints final_result.usage
Tokens to capacity units. Generative AI hub use is metered in tokens, converted to GenAI tokens at per-model rates, then to capacity units. The rates are in SAP Note 3437766. The generative AI hub is only in SAP AI Core's extended plan.
No compute charges for foundation models. When you use foundation models and grounding, SAP waives SAP AI Core's compute, storage and baseline charges, and charges accrue on tokens instead. Custom models you train and serve yourself are billed on compute, storage and an hourly baseline.
Other meters. Grounding storage is metered in gigabyte days and retrieval in text blocks; content filtering and data masking have their own metrics. Stored inference records for observability are metered in data volume.
Your contract decides the money. Under SAP BTPEA and CPEA, you prepay cloud credits with an annual commitment; overages are billed in arrears at list price, and volume discounts are available. Pay-As-You-Go has no commitment and is billed monthly at non-discountable rates.
SAP's prompt caching page describes two modes, both in Orchestration V2:
Explicit caching with cache_control breakpoints on content blocks, for Anthropic Claude and Amazon Nova models. Claude supports breakpoints on system, message and tool blocks, up to four per request. The default lifetime is five minutes; selected Anthropic models accept "ttl": "1h".
Implicit caching for OpenAI and Gemini models, on by default.
A cached system message in a request to the orchestration service's /v2/completion endpoint looks like this (a sketch; it needs an SAP AI Core account with the extended plan, set up in Unit 5):
Consumption report. In the SAP BTP cockpit, open the global account, choose Usage, filter by AI Core and export. Model consumption appears under Applications, and resource-group consumption under instance.
Costs and Usage page. On the consumption-based model, the global account's Costs and Usage page shows costs by service and subaccount, monthly trends, and estimates between balance statements. The monthly balance statement is the binding figure.
Budgets. The Budgets tab sets budget limits and email alerts to global account administrators. SAP's documentation states that the system does not suspend costs or service usage when a budget is exceeded, and that alerts can be delayed.
Joule and other AI features inside SAP applications are paid through your subscription and, for premium capabilities and agents, AI Units. The metering above doesn't apply to them. Joule and Joule Studio and Joule agents and agent orchestration cover what consumes AI Units.
Log tokens per request, with context. Record input, output, cached and reasoning tokens with the feature, prompt version, model and tenant on every call, as the observability topic does with gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Without them, a monthly bill can't be traced to a cause.
Enforce limits in code. Set max_tokens on every call, a step limit on every agent loop, and a daily token budget per user or per feature that refuses or queues work when spent. SAP BTP budgets only send alerts.
Count retries and fallbacks. A retry, a fallback model or a hedged request is a second paid call. The latency topic's retries belong in this cost model.
Watch prompts in releases. A new example in the system prompt is a cost change on every call. Review prompt diffs for token impact, as Step 8 showed with prompt v2.
Keep the cache prefix stable. Don't put dates, user names or order numbers in the system prompt. Anthropic's cookbook names dynamic content in the prefix as the most common way to break the cache.
Security and authorizations. Token counts are safe to log; prompt and answer text may hold customer data and must follow your logging rules. Never trade the user's SAP authorizations for a cheaper shared technical user. Response caching, where another user could receive a stored answer, is a different risk, covered in the next topic of this unit.
Peaks and rate limits. Month-end close raises volume and request rates at the same time. Plan cost at peak, and check that the rate limits on your SAP AI Core tenant fit it.
Clean core. All of this lives in your side-by-side app on BTP. S/4HANA needs no change to measure or reduce token cost.
Produce the cost section of a business case for the blocked-orders assistant. The AI business case topic later in this unit uses it.
Open unit10/token_costs.py and change RATES["input_per_million"] and RATES["output_per_million"] to the list prices of a model you might use, from the provider's pricing page. Write the date and the page's address in a comment above RATES.
Choose a realistic month. Ask someone in order-to-cash, or assume 6,000 blocked orders and 15,000 at month-end close.
Run the model for both months, for the design you would build:
Run once more with the answer length you think clerks need, for example --answer-tokens 150.
Create unit10/cost-model.md with a table: month (normal and peak), design, caching on or off, cost per order, cost per month. Add one line for the rates you used and their date.
Below the table, write three sentences: which design you would build and why, what single change saves the most, and which non-model costs (retrieval, filtering, masking) are still missing.
Commit cost-model.md.
Done whencost-model.md shows cost per order and per month for a normal and a peak month, names the rates and their date, and recommends one design with the change that saves the most.
Pick one answer for each question. The explanation appears after you choose.
1In the script's default run, the agent design sends 3,708 input tokens for the same answer that the pipeline design gets from 1,320. What causes most of the difference?
Answer: B. The agent makes three calls, and each one carries the roughly 1,060-token prefix plus everything gathered so far. The answer length and the notes are the same in both designs; the tool definitions are in the prefix of both.
2With --hit-rate 0, the pipeline design with caching costs more than without it. Why?
Answer: C. At a hit rate of zero, each call writes the prefix at 1.25 times the input rate and nothing is ever read at the discounted rate. The agent design still gains because its second and third calls hit the cache.
3Your team adds today's date to the first line of the system prompt so the model knows it. What happens to prompt caching?
Answer: D. Caching reuses a prefix only when it is identical. A changing value at the top makes the prefix different, so calls write a new cache instead of reading the old one. Put changing values after the fixed part.
4A response from an Anthropic model through SAP's orchestration service shows prompt_tokens 3 and cached_tokens 4,657. How should your cost code treat these numbers?
Answer: B. In SAP's example the cached count is far larger than the prompt count, so it can't be included in it. Providers differ on this, so check a real response for your model and price cached tokens at the rate your contract gives them.
5Why does the script refuse to show caching rows with --sap?
Answer: C. SAP AI Core does support prompt caching, but the metering page opened for this topic doesn't give a conversion for cached tokens. Rather than guess, the script prices them as normal input and points you to SAP Note 3437766.
6Retrieving ten notes instead of three pushes the notes part over its budget. Which change brings it back without losing the answer's grounding?
Answer: D. The notes vary per order and come after the fixed prefix, so caching can't help them. Lowering max_tokens affects the answer, not the prompt. Better retrieval, such as reranking, keeps the useful passages and drops the rest.
7The BTP budget for AI Core is exceeded on the 20th of the month. What happens to the assistant's calls?
Answer: A. SAP's documentation states that the system does not suspend costs or usage when a budget is exceeded, and that alerts can be delayed. If you need a hard stop, build per-user or per-feature caps into the app.
8A product owner wants a shorter answer and suggests lowering max_tokens from 250 to 120. What do you do?
Answer: C. The model doesn't plan around max_tokens; when it hits the cap, the answer is cut off mid-sentence. Specify the length and shape you want in the prompt, and keep max_tokens above the longest legitimate answer to stop runaways.
9In SAP's worked example of a RAG assistant, why does the total come to 418.55 CU when the model calls are 232.25?
Answer: B. SAP's example adds grounding storage and retrieval, content filtering and data masking, which together come to 186.3 CU. Those modules are metered separately from tokens and belong in every cost model.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Metering and Pricing for Generative AI (SAP AI Core documentation, SAP-docs on GitHub)— tokens converted to GenAI tokens, then capacity units (CUs) by conversion factors; rates per model in SAP Note 3437766; output tokens tend to be slightly more expensive; worked example with fictitious values (0.00112 and 0.00320 GenAI tokens per 1,000 input and output tokens, 1.90385 CU per GenAI token, 25,000 RAG requests of 3,500 in and 300 out = 232.25 for the model plus 186.3 CU for grounding, filtering and masking = 418.55 CU); typical consumption patterns per use case; generative AI hub only in the extended plan
Prompt Caching (SAP AI Core documentation, SAP-docs on GitHub)— explicit cache_control breakpoints in Orchestration V2 for Anthropic Claude and Amazon Nova models; implicit caching for OpenAI and Gemini models on by default; five-minute default TTL, one hour for selected Anthropic models; usage reports cached_tokens and cache_creation_tokens in prompt_tokens_details
SAP Cloud SDK for AI, generative AI package (sap-ai-sdk-gen 7.4.1 on PyPI)— installed and read in this run: TokenUsage with prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details (cached_tokens, cache_creation_tokens, per-TTL breakdown) and completion_tokens_details (reasoning_tokens); CacheControl with ttl 5m or 1h
Cost optimization on the Claude API (Anthropic, claude-cookbooks on GitHub, September 2026)— measure cost per task, not per token; cache writes at 1.25x (5 minutes) or 2x (1 hour) the input rate and reads at 0.1x on the models used, lower on some newer models; dynamic content in the prefix breaks the cache; Batch API bills every token at 50% off within a 24-hour window; intermediate tokens accumulate in agent loops; max_tokens is a backstop, not a tuning knob
Prompt Caching 101 (OpenAI Cookbook on GitHub)— caching activates automatically for prompts over 1,024 tokens; static content first, variable content last; cached_tokens in usage.prompt_tokens_details; caches scoped to the organization