Orchestrate

Token economics

Price an AI feature per business transaction, not per token, and see how input, output and cached tokens, agent loops and SAP's capacity units add up.

Updated Oct 7, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

Language models are billed by the token, a small piece of text. A token is often a short word or part of a longer one. Every call is billed twice: once for the tokens you send (the input) and once for the tokens the model writes (the output). Output tokens usually cost more per token.

That makes a token bill easy to read and hard to predict. The useful question is not "what does a token cost?" but "what does one business transaction cost?" For example: what does it cost to explain one blocked sales order to a clerk?

Three numbers multiply to give the monthly bill:

  • Tokens per call: how long the prompt and the answer are.
  • Calls per transaction: one call for a simple design, several for an agent that works step by step.
  • Transactions per month: how many orders, invoices or questions go through.

Three habits keep the bill under control. Set a length budget for each part of the prompt. Cache the part of the prompt that never changes. Ask for answers no longer than they need to be.

Why it matters to the business

Take the running example from this unit: an assistant that explains blocked sales orders to order-to-cash clerks. Each explanation needs the order data from SAP, a few passages from the team's notes, the instructions, and the question.

A cost estimate that stops at "a fraction of a cent per call" misses three things.

  • Designs multiply calls. An agent that first asks for the order, then for the notes, then answers makes three calls. Each call sends the whole conversation again. In this topic's cost model, the agent design costs about twice as much per order as a design where the app gathers the data and calls the model once.
  • Volume multiplies everything. A pilot with 200 orders hides what 20,000 orders a month cost. Month-end close can triple a normal week.
  • The model is only part of the bill. In SAP's own worked example for a retrieval assistant, the model calls are about 55% of the total. The rest is document storage, retrieval, content filtering and data masking.

Cost also shapes decisions that look technical. Retrieving ten passages instead of three makes every answer more expensive. So does a long set of instructions. A cheaper model may need more calls to reach the same answer. Someone who owns the budget should see these trade-offs before they are built in.

How SAP does it

As of October 2026, SAP prices AI in two ways, depending on who builds the feature.

AI you build on SAP BTP. Apps and agents that call models through the generative AI hub in SAP AI Core are metered in tokens. SAP converts input and output tokens into "GenAI tokens" at a rate per model, then into BTP capacity units. Your company pays for capacity units under its BTP contract. The generative AI hub is part of SAP AI Core's extended plan. Grounding, content filtering and data masking have their own meters.

AI that SAP builds into its applications. Joule and other SAP-delivered AI features are not billed per token to you. Some come with your cloud subscription; premium capabilities and agents consume AI Units, SAP's prepaid currency for those features. Joule and Joule Studio explains how AI Units are bought and used.

SAP's BTP cockpit shows usage and costs, and can send alerts when spending passes a budget. SAP's documentation is explicit that a budget alert does not stop the spending. Limits that actually stop spending have to be built into your app.

SAP also supports prompt caching in its orchestration service, so a long fixed part of a prompt can be reused across calls. How cached tokens count toward capacity units is in SAP's per-model rate note; ask your account team.

A cost model in five lines

A cost model for one AI feature fits on one page. The figures below come from this topic's hands-on model with made-up rates. Use them for the shape, not the price.

Line Blocked-orders assistant, simple design Same, agent design
Model calls per order 1 3
Tokens sent per order (input) about 1,300 about 3,700
Tokens written per order (output) about 220 about 300
Cost per order, made-up rates 0.0048 0.0104
With caching of the fixed instructions 0.0034 0.0052

Read the table from the top. The agent design sends almost three times as many tokens for the same answer, because every step sends the conversation again. Caching narrows the gap because the repeated part is billed at a fraction of the normal rate. Neither number means anything until it is multiplied by a real monthly volume and set against the value of the work it saves.

A finished cost model adds three more lines: monthly volume (with peaks), non-model meters such as retrieval and filtering, and the value side, such as minutes saved per order.

Questions to ask

  • What is the cost per business transaction, such as one explained order or one matched invoice, not just per token?
  • How many model calls does one transaction make, and does the conversation grow with each call?
  • What is the monthly volume, and what happens at month-end or quarter-end peaks?
  • Which part of the prompt is the same on every call, and is it cached?
  • Is there a length budget for each part of the prompt, and who decided it?
  • What is the longest answer we allow, and does the clerk need all of it?
  • What else is metered besides the model: retrieval, filtering, masking, logging?
  • What stops runaway spending: a per-user limit, a daily cap, a switch-off? Budget alerts alone don't.
  • For SAP-delivered features, do they use AI Units, and how many will our users consume?

Common misconceptions

  • "Tokens are words." A token is often shorter than a word. Numbers, codes and many non-English languages use more tokens for the same meaning.
  • "Input and output cost the same." Output usually costs more per token, so long answers are expensive.
  • "A cheaper model is always cheaper." If it needs more calls or more retries to reach the same answer, the cost per transaction can rise.
  • "Caching is free money." Writing to a cache costs extra on some models. If the cache expires before the next call, you pay the premium and get nothing back.
  • "The budget alert will protect us." On SAP BTP, a budget sends alerts; it does not stop usage.
  • "The model is the whole bill." Retrieval, filtering, masking and storage are metered too. In SAP's worked example they are close to half.

Key terms

  • Token: the piece of text a model reads or writes; about four bytes of English text on average.
  • Input tokens: what you send: instructions, data, retrieved passages, the question.
  • Output tokens: what the model writes, including any hidden reasoning it does first.
  • Prompt caching: reusing a fixed part of the prompt across calls, billed at a lower rate when reused.
  • Cost per transaction: the cost of one complete unit of business work, across all its model calls.
  • Prompt length budget: the most tokens each part of the prompt may use.
  • Capacity units: SAP BTP's billing unit; generative AI hub tokens are converted into them.
  • AI Units: SAP's prepaid currency for premium AI features in its own applications, such as Joule.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A vendor quotes a price per million tokens for the blocked-orders assistant. What number should the finance lead ask for instead?

    Answer: B. The business pays per transaction, and one transaction can involve several calls with different token counts. A per-token price is an input to the estimate, not the estimate.
  2. 2The team switches from a single model call to an agent that asks for the order, then the notes, then answers. What happens to cost per order?

    Answer: C. Each step of an agent loop sends the instructions and everything gathered so far again. In this topic's model, the agent design costs about twice as much per order as the single-call design.
  3. 3Why should a cost estimate for an AI assistant on SAP BTP include more than the model calls?

    Answer: A. SAP's own worked example adds grounding storage, retrieval, content filtering and data masking to the model calls, and they come to close to half the total. Leaving them out understates the bill.
  4. 4The team sets a budget with alerts in the SAP BTP cockpit. What does that protect against?

    Answer: C. SAP's documentation says the system does not suspend costs or usage when a budget is exceeded. Hard limits such as per-user caps or a switch-off must be built into the application.
  5. 5When does prompt caching fail to save money?

    Answer: D. Writing to the cache can cost more than a normal call on some models. If the cache expires before it is read, you pay the extra and get no discount back.
  6. 6A product owner asks for longer, more detailed answers "since they only cost a little more". What should you point out?

    Answer: B. Output tokens are usually priced higher than input tokens, and caching applies to the prompt, not the answer. Length should be a decision about what the clerk needs.
  7. 7Your business users want Joule to explain blocked orders, instead of a custom app. How is that usage paid for?

    Answer: D. SAP-delivered AI features are not billed to you per token. Some capabilities come with the cloud subscription, and premium capabilities and agents consume AI Units, as the Joule topic explains.
Deep layer · 40 min read

Mental model

The token bill is a sum over calls, and the unit that matters is the business transaction.

cost per transaction = sum over its calls of
    (uncached input x input rate)
  + (cache writes   x input rate x write multiplier)
  + (cache reads    x input rate x read multiplier)
  + (output         x output rate)

cost per month = cost per transaction x transactions per month

Everything in this topic is one of those terms. A longer prompt grows the input term. An agent loop adds calls, and each call carries the history of the ones before it. Caching moves tokens from the first term to the cheaper third. A shorter answer shrinks the last term. On SAP AI Core, the same sums are converted into capacity units at per-model rates, and other modules add their own meters beside them.

Two rules follow:

  1. Measure per transaction, not per token. Anthropic's cost cookbook makes the same point: a model with a higher price per token can be cheaper per task if it finishes in fewer turns.
  2. Count what you send, not what you mean. The model is billed for every token in the request, including tool definitions, examples and history you stopped thinking about long ago.

How it works

What gets counted

Every response tells you what was counted. In SAP's orchestration service (Python SDK sap-ai-sdk-gen 7.4.1, read for this topic), the usage object has these fields:

Field Meaning Billed as
prompt_tokens Tokens sent to the model Input
completion_tokens Tokens the model wrote Output
prompt_tokens_details.cached_tokens Input tokens read from the prompt cache Input, at the cache read rate where the provider discounts it
prompt_tokens_details.cache_creation_tokens Input tokens written to the prompt cache Input, at a premium on models that charge for writes
completion_tokens_details.reasoning_tokens Tokens a reasoning model "thought" before answering Output; the SDK notes such tokens are counted in completion tokens for billing

Tokens are not characters

Tokenizers split text into pieces from a fixed vocabulary. The tiktoken README says a token is about 4 bytes of text on average. That rule works for planning English prompts. It breaks for:

  • Other languages. German umlauts take two bytes in UTF-8, and many scripts take three. A prompt translated for German or Japanese users can cost noticeably more.
  • Codes and numbers. Order numbers, material numbers and JSON punctuation split into many small tokens.
  • Different models. Each model family has its own tokenizer. The same prompt has a different count on each.

Use the 4-byte rule to plan, then confirm with the usage numbers a real call returns. This topic's script prints both.

Input, output and why output dominates per token

SAP's metering page says output tokens "tend to be slightly more expensive" than input tokens. In SAP's fictitious example the output rate is almost three times the input rate. In the price snapshot in Anthropic's cookbook, output costs five times input for every model listed. Either way, a long answer costs more than a long prompt of the same length.

The ratio of input to output depends on the job. SAP's "typical consumption patterns" table, from the same page:

Use case Input tokens per request Output tokens per request
RAG chat 3,500 300
Basic chat 500 100
Summarization 5,000 300
Classification 3,800 10
Generation 500 3,500

Retrieval and classification are input-heavy, so prompt length budgets and caching matter most. Generation is output-heavy, so answer length matters most.

Prompt caching

Most prompts start with a long part that never changes: instructions, examples and tool definitions. Prompt caching lets the provider keep that processed prefix for a few minutes and reuse it.

Provider behaviour (from the sources opened for this topic) Anthropic Claude OpenAI
How it turns on cache_control on a block, or automatic at the top level Automatic for prompts over 1,024 tokens
Minimum prefix 1,024 tokens for Sonnet, 4,096 for Opus and Haiku 4.5 1,024 tokens
Lifetime 5 minutes, refreshed on each hit; 1 hour optional Not stated in the cookbook
Price of a write 1.25 times input (5 minutes), 2 times (1 hour) No write premium stated
Price of a read 0.1 times input on the cookbook's models, lower on some newer ones "Discounted"; no figure in the cookbook

Three things decide whether caching pays:

  1. The prefix must be byte-identical. A timestamp, a user name or a request ID in the system prompt makes every prefix different. Put everything that varies after the fixed part.
  2. The prefix must be long enough. Below the minimum, nothing is cached.
  3. Calls must come before the cache expires. With a write premium, a cache that is never read costs more than no cache at all.

Agent loops compound input

A pipeline design reads SAP and the notes in code, then makes one call. An agent design lets the model ask for each piece. Each step resends everything:

flowchart LR
  subgraph Pipeline
    P1[Call 1<br/>prefix + order + notes + question] --> A1[Answer]
  end
  subgraph Agent
    C1[Call 1<br/>prefix + question] --> C2[Call 2<br/>prefix + question + order]
    C2 --> C3[Call 3<br/>prefix + question + order + notes]
    C3 --> A2[Answer]
  end

With a fixed prefix of about 1,060 tokens, the agent sends that prefix three times. That is why caching helps agents most: calls two and three come seconds after call one, so they read the prefix from the cache. Anthropic's cookbook adds that intermediate results, such as a large tool output, stay in the context on every later turn unless you remove them.

Prompt length budgets

A prompt length budget gives each part of the prompt a ceiling, the same way a latency budget gives each step a time. For the blocked-orders assistant:

Part Budget (tokens) Why it can grow
System instructions and examples 900 Each new rule or example is added "just in case"
Tool definitions 450 Every new tool adds its description and schema to every call
Order data 250 Orders with many items; fields nobody uses
Retrieved notes 200 Raising the number of passages retrieved
Question 60 Pasted emails or documents
Answer (max_tokens) 250 Requests for "more detail"

A budget turns a slow drift into a visible failure. When the notes part goes over, the team decides: fewer passages, shorter passages, or a bigger budget and a higher cost per order.

Anthropic's cookbook treats max_tokens as a backstop for runaway answers, not as the way to shorten normal ones. A model cut off at max_tokens stops mid-sentence. To shorten typical answers, ask for a specific shape and length in the prompt.

From tokens to capacity units on SAP AI Core

SAP's metering page walks through the conversion with fictitious values. Per request:

GenAI tokens   = 3,500 / 1,000 x 0.00112 + 300 / 1,000 x 0.00320 = 0.00488
capacity units = 0.00488 x 1.90385 = 0.00929

For 25,000 requests a month, SAP's example totals 232.25 for the model calls. Grounding storage, grounding retrieval, content filtering and data masking add 186.3 CU, for 418.55 CU in all. The model calls are about 55% of the total. The real conversion rates per model are in SAP Note 3437766, and inference observability storage has its own conversion in SAP Note 3720903.

Build it yourself: a cost model for the blocked-orders assistant

You will build a cost model that counts the tokens in each part of the assistant's prompt, checks them against a length budget, and prices one explained order and one month. It compares the pipeline design with the agent design, with and without prompt caching. You will then switch to SAP's capacity-unit arithmetic, price the token counts you logged in the observability topic, and optionally read the real token usage of one call in SAP AI Core.

Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. The main script uses only built-in Python. Step 8 uses the traces.jsonl file from Observability for AI systems. The optional Step 9 also needs the SAP AI Core trial and sap-ai-sdk-gen from Set up for Unit 5.

flowchart LR
  D[Made-up order,<br/>notes, instructions] --> C[Count tokens<br/>per part]
  C --> B{Length budget}
  B --> P[Price per order<br/>pipeline vs agent]
  P --> M[Price per month]
  T[traces.jsonl] --> P

What you need

  • Your course folder with .venv from the earlier setup topics.
  • About 40 to 60 minutes.
  • Cost: free. No model is called and no account is needed, except in the optional Step 9, which uses your SAP AI Core trial (no charge during the trial; a small per-request charge after it).

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check your Python version (the same on every system):

    python --version
    Python 3.14.4

    Any version from 3.10 up works.

Step 2: Create the script

The script holds the assistant's instructions, tool definitions, one made-up blocked order and ten short notes. It counts the tokens in each part with the 4-byte rule, then prices the designs.

  1. In VS Code, right-click unit10, choose New File, name it token_costs.py, paste the code below and save.
"""Unit 10: what does one blocked order cost in tokens, and what does a month cost?

The blocked-orders assistant explains why a sales order is blocked and what the clerk should
check next. This script builds its prompt from made-up, SAP-shaped data, counts the tokens in
each part, checks them against a prompt length budget, and prices one business transaction
("explain one blocked order") and one month, for two designs:

  pipeline   your code reads the order and the notes, then makes one model call
  agent      the model asks for the order, then for the notes, then answers: three calls,
             and every call sends the whole conversation so far

With no options it uses EXAMPLE rates (made up, see RATES below), calls no model and costs nothing.

  python unit10/token_costs.py                     both designs, with and without prompt caching
  python unit10/token_costs.py --notes 10          retrieve 10 note chunks instead of 3
  python unit10/token_costs.py --answer-tokens 120 ask for shorter answers
  python unit10/token_costs.py --orders 20000      a busier month
  python unit10/token_costs.py --sap               SAP AI Core arithmetic: GenAI tokens and capacity units
  python unit10/token_costs.py --traces unit10/traces.jsonl   price the token counts you logged
  python unit10/token_costs.py --llm --model NAME  one real call in SAP AI Core; print its usage
"""
import argparse
import json
import math
import os
import sys
from pathlib import Path

# ---------------------------------------------------------------------------------------------
# Rates. EXAMPLE VALUES ONLY: replace them with your provider's price list or your SAP contract.
# ---------------------------------------------------------------------------------------------
RATES = {
    # Provider-style list prices, in US dollars per million tokens (made up for this course).
    "input_per_million": 2.00,
    "output_per_million": 10.00,
    # Prompt caching multipliers on the input price. These are the ones Anthropic documents for
    # its 5-minute cache; other providers use other numbers. Check yours.
    "cache_write_multiplier": 1.25,
    "cache_read_multiplier": 0.10,
    "min_cacheable_tokens": 1024,
    # SAP AI Core style: GenAI tokens per 1,000 model tokens, then capacity units per GenAI token.
    # These are the fictitious values from SAP's own worked example; real ones are per model,
    # in SAP Note 3437766.
    "sap_genai_per_1k_input": 0.00112,
    "sap_genai_per_1k_output": 0.00320,
    "sap_cu_per_genai_token": 1.90385,
}

# The prompt length budget: the most tokens each part of the prompt may use.
BUDGET = {"system": 900, "tools": 450, "order": 250, "notes": 200, "question": 60}
ANSWER_CAP = 250  # max_tokens: the answer is cut off here

# ---------------------------------------------------------------------------------------------
# Made-up, SAP-shaped data. Block reason and status codes are configured per SAP system;
# the meanings below belong to this made-up company only.
# ---------------------------------------------------------------------------------------------
SYSTEM_PROMPT = """You help order-to-cash clerks understand blocked sales orders in SAP S/4HANA.
Answer in plain English, in at most five short sentences, then give one next step.
Use only the order data and the notes you are given. If they do not explain the block, say so.
Never promise a release date. Never release, change or cancel an order yourself.

This company's guide to delivery block reasons (configured in its own SAP system):
- 01 Credit limit: the customer's open items exceed the credit limit. Credit management decides.
- 02 Missing export papers: shipping waits for the customs team to attach the documents.
- 03 Pricing check: a manual price or discount is above the sales rep's limit.
- 04 Customer request: the customer asked us to hold the delivery. Check the order notes.
- 05 Quality hold: one or more items wait for a quality inspection result.
- 06 Dangerous goods: the safety data sheet is missing or out of date for this country.
- 07 Incomplete order: required fields such as the incoterms or the ship-to address are empty.
- 08 Advance payment: the customer must pay before delivery. Check incoming payments.

This company's credit check statuses: A means not checked or released, B means a check failed,
C means partially released. Explain a status in words; do not show the letter alone.

Style rules: name the order number in the first sentence. Use the customer's name, not the ID.
Quote amounts with their currency. If two reasons apply, explain the delivery block first.
If the notes contradict the order data, trust the order data and mention the conflict.

Example 1. Question: Why is order 9000077 blocked? Data: DeliveryBlockReason 02, credit status A.
Answer: Order 9000077 for Contoso Freight is waiting for export papers. Shipping cannot start
until the customs team attaches them. The credit check is fine. Next step: ask the customs team
whether the papers for this order have been requested.

Example 2. Question: What is wrong with order 9000081? Data: no delivery block, credit status C.
Answer: Order 9000081 for Fabrikam Retail has no delivery block. Its credit check is partially
released, so some items may ship while others wait. Next step: check in the credit management
app which items are still held, and tell the customer which ones will ship first.

Example 3. Question: Can order 9000090 ship today? Data: DeliveryBlockReason 05, credit status A.
Answer: Order 9000090 for Litware Medical is on a quality hold. At least one item is waiting for
an inspection result, so it cannot ship until quality releases it. Credit is not the issue.
Next step: ask the quality team when the inspection lot for this order will be decided.

Answer format: first the reason in plain words, then what it means for the customer, then
"Next step:" and one action the clerk can take today. No headings, no lists, no codes alone.
"""

TOOLS = [
    {"name": "get_sales_order",
     "description": "Read one sales order from SAP: header, block reasons, credit status and items. "
                    "Use it before explaining any order. Read-only.",
     "parameters": {"type": "object", "properties": {
         "sales_order": {"type": "string", "description": "The sales order number, for example 9000123"}},
         "required": ["sales_order"]}},
    {"name": "search_notes",
     "description": "Search the order-to-cash team's notes for how to handle a block reason. "
                    "Returns the most relevant short passages with their source.",
     "parameters": {"type": "object", "properties": {
         "query": {"type": "string", "description": "What to look for, in plain words"},
         "top_k": {"type": "integer", "description": "How many passages to return, 1 to 10"}},
         "required": ["query"]}},
    {"name": "get_delivery_status",
     "description": "Read the outbound deliveries for a sales order and their goods issue status. "
                    "Use it only when the clerk asks about shipping. Read-only.",
     "parameters": {"type": "object", "properties": {
         "sales_order": {"type": "string", "description": "The sales order number"}},
         "required": ["sales_order"]}},
    {"name": "get_customer_credit",
     "description": "Read the customer's credit limit, credit exposure and open items. Read-only.",
     "parameters": {"type": "object", "properties": {
         "customer": {"type": "string", "description": "The customer number (sold-to party)"}},
         "required": ["customer"]}},
]

ORDER = {
    "SalesOrder": "9000123", "SoldToParty": "10100042", "CustomerName": "Northwind Tools GmbH",
    "SalesOrderDate": "2026-10-05", "TotalNetAmount": "48200.00", "TransactionCurrency": "EUR",
    "DeliveryBlockReason": "01", "TotalCreditCheckStatus": "B",
    "Items": [
        {"SalesOrderItem": "10", "Material": "TG-11", "RequestedQuantity": "40", "NetAmount": "28800.00"},
        {"SalesOrderItem": "20", "Material": "TG-12", "RequestedQuantity": "20", "NetAmount": "19400.00"},
    ],
}

NOTES = [
    "Credit blocks (reason 01): check the customer's exposure in the credit management app first. "
    "If a payment is on its way, ask credit management for a one-time release.",
    "Credit check status B after a large new order usually means the order itself pushed exposure over "
    "the limit. Splitting the order is not a fix; talk to credit management.",
    "Key accounts: for customers in the key account list, the account manager can request a temporary "
    "limit increase. Typical turnaround is one business day.",
    "Advance payment customers (reason 08) are a different process. Do not ask credit management.",
    "If the customer disputes an open invoice, the disputed amount still counts toward exposure "
    "until the dispute case is closed.",
    "Month-end: credit releases are slower in the last two working days of the month. Tell the "
    "customer before promising a date.",
    "When a blocked order has items from two plants, the block applies to the whole order, not per item.",
    "Customers on payment terms longer than 60 days have a lower credit limit by policy.",
    "If the credit limit was changed today, the order may need a new credit check before the block clears.",
    "Escalation: if a credit block is older than five working days, inform the sales team lead.",
]

QUESTION = "Why is sales order 9000123 blocked and what should I check next?"


def count_tokens(text: str) -> int:
    """Estimate tokens: about 4 bytes of text per token (the tiktoken README's rule of thumb).
    Real counts differ by model and language; --llm shows a real count."""
    return max(1, math.ceil(len(text.encode("utf-8")) / 4))


def prompt_parts(n_notes: int) -> dict:
    """The text of each part of the pipeline design's prompt."""
    notes = "\n".join(f"[note {i + 1}] {t}" for i, t in enumerate(NOTES[:n_notes]))
    return {"system": SYSTEM_PROMPT, "tools": json.dumps(TOOLS), "order": json.dumps(ORDER),
            "notes": notes, "question": QUESTION}


def calls_for(design: str, parts: dict, answer_tokens: int) -> list:
    """Return the model calls one transaction makes, as token counts.
    Each call: prefix (system + tools, the same every time), rest (everything else sent), out."""
    t = {k: count_tokens(v) for k, v in parts.items()}
    prefix = t["system"] + t["tools"]
    if design == "pipeline":
        return [{"prefix": prefix, "rest": t["order"] + t["notes"] + t["question"], "out": answer_tokens}]
    # Agent: call 1 asks for the order, call 2 asks for the notes, call 3 answers.
    # Every call re-sends what came before: that is why agent loops cost more input tokens.
    tool_call = 40  # tokens the model writes to request a tool
    history = t["question"]
    calls = [{"prefix": prefix, "rest": history, "out": tool_call}]
    history += tool_call + t["order"]
    calls.append({"prefix": prefix, "rest": history, "out": tool_call})
    history += tool_call + t["notes"]
    calls.append({"prefix": prefix, "rest": history, "out": answer_tokens})
    return calls


def price(calls: list, cache: bool, hit_rate: float, sap: bool) -> dict:
    """Price one transaction. With cache=True the prefix is read from the cache on a hit and
    written to it on a miss. In an agent loop, calls after the first are seconds apart, so they
    hit the cache for the prefix (and in real systems often for the history as well)."""
    r = RATES
    total = {"input": 0.0, "cache_write": 0.0, "cache_read": 0.0, "output": 0.0}
    for i, c in enumerate(calls):
        total["output"] += c["out"]
        if not cache or sap or c["prefix"] < r["min_cacheable_tokens"]:
            total["input"] += c["prefix"] + c["rest"]
            continue
        hit = 1.0 if i > 0 else hit_rate
        total["cache_read"] += c["prefix"] * hit
        total["cache_write"] += c["prefix"] * (1 - hit)
        total["input"] += c["rest"]
    if sap:
        genai = (total["input"] / 1000 * r["sap_genai_per_1k_input"]
                 + total["output"] / 1000 * r["sap_genai_per_1k_output"])
        total["cost"] = genai * r["sap_cu_per_genai_token"]
    else:
        pin, pout = r["input_per_million"] / 1e6, r["output_per_million"] / 1e6
        total["cost"] = (total["input"] * pin + total["cache_write"] * pin * r["cache_write_multiplier"]
                         + total["cache_read"] * pin * r["cache_read_multiplier"] + total["output"] * pout)
    return total


def show_budget(parts: dict, answer_tokens: int) -> None:
    print("Prompt length budget, pipeline design (estimated tokens)\n")
    print(f"{'part':<10}{'tokens':>8}{'budget':>8}  check")
    for name, text in parts.items():
        n = count_tokens(text)
        print(f"{name:<10}{n:>8}{BUDGET[name]:>8}  {'OK' if n <= BUDGET[name] else 'OVER'}")
    total_in = sum(count_tokens(t) for t in parts.values())
    print(f"{'input':<10}{total_in:>8}{sum(BUDGET.values()):>8}  {'OK' if total_in <= sum(BUDGET.values()) else 'OVER'}")
    over = answer_tokens > ANSWER_CAP
    print(f"{'answer':<10}{answer_tokens:>8}{ANSWER_CAP:>8}  {'OVER: raise max_tokens or ask for less' if over else 'OK'}")
    prefix = count_tokens(parts["system"]) + count_tokens(parts["tools"])
    if prefix < RATES["min_cacheable_tokens"]:
        print(f"\nThe fixed prefix is {prefix} tokens, under the {RATES['min_cacheable_tokens']}-token minimum many "
              "providers need before they cache. Caching will not help this prompt.")


def model_costs(args) -> None:
    parts = prompt_parts(args.notes)
    show_budget(parts, args.answer_tokens)
    unit = "CU" if args.sap else "USD"
    print(f"\nCost per transaction and per month ({args.orders:,} blocked orders, "
          f"{'SAP example conversion rates' if args.sap else 'EXAMPLE rates'})\n")
    print(f"{'design':<10}{'cache':<7}{'calls':>6}{'input':>8}{'cached':>8}{'output':>8}"
          f"{'per order':>12}{'per month':>12}")
    for design in ("pipeline", "agent"):
        calls = calls_for(design, parts, args.answer_tokens)
        for cache in (False, True):
            if cache and args.sap:
                continue
            p = price(calls, cache, args.hit_rate, args.sap)
            cached = p["cache_read"] + p["cache_write"]
            per_order = f"{p['cost']:.5f} {unit}"
            per_month = f"{p['cost'] * args.orders:,.2f} {unit}"
            print(f"{design:<10}{'on' if cache else 'off':<7}{len(calls):>6}{p['input']:>8.0f}{cached:>8.0f}"
                  f"{p['output']:>8.0f}{per_order:>12}{per_month:>12}")
    print("\ninput = tokens billed at the full input rate; cached = prefix tokens written to or read from "
          "the cache;\noutput = tokens the model writes. Rates are examples: replace RATES with your own.")
    if args.sap:
        print("SAP mode shows capacity units (CU). What a CU costs depends on your SAP BTP contract.\n"
              "Cached tokens are priced as normal input here: SAP's metering page does not say how they convert.")


def price_traces(path: Path, orders: int) -> None:
    """Price the token counts logged by observe_assistant.py, grouped by prompt version."""
    if not path.exists():
        sys.exit(f"No file {path}. Run 'python unit10/observe_assistant.py' from the observability "
                 "topic first, or leave out --traces.")
    spans = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
    version = {s["trace_id"]: s["attributes"].get("app.prompt.version", "?")
               for s in spans if s["parent_id"] is None}
    groups = {}
    for s in spans:
        a = s["attributes"]
        if a.get("gen_ai.operation.name") == "chat":
            g = groups.setdefault(version.get(s["trace_id"], "?"), {"calls": 0, "in": 0, "out": 0})
            g["calls"] += 1
            g["in"] += a.get("gen_ai.usage.input_tokens", 0)
            g["out"] += a.get("gen_ai.usage.output_tokens", 0)
    if not groups:
        sys.exit(f"{path} has no chat spans with token counts.")
    pin, pout = RATES["input_per_million"] / 1e6, RATES["output_per_million"] / 1e6
    print(f"Logged model calls in {path.name}, priced at EXAMPLE rates\n")
    print(f"{'prompt':<8}{'calls':>7}{'avg in':>8}{'avg out':>9}{'per call':>12}{'per month':>12}")
    for v, g in sorted(groups.items()):
        cost = (g["in"] * pin + g["out"] * pout) / g["calls"]
        print(f"{v:<8}{g['calls']:>7}{g['in'] / g['calls']:>8.0f}{g['out'] / g['calls']:>9.0f}"
              f"{cost:>8.5f} USD{cost * orders:>8,.2f} USD")
    print(f"\nper month = per call x {orders:,} calls. Change it with --orders.")


def real_call(args) -> None:
    """One real call through SAP's orchestration service; print the usage SAP reports."""
    from dotenv import load_dotenv
    load_dotenv()
    needed = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
              "AICORE_RESOURCE_GROUP"]
    missing = [n for n in needed if not os.environ.get(n)]
    if missing:
        sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', or leave out --llm.")
    if not args.model:
        sys.exit("Name a model with --model, for example one from 'python unit05/choose_model.py catalog'.")
    from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
                                             OrchestrationService, PromptTemplatingModuleConfig,
                                             SystemMessage, Template, UserMessage)
    parts = prompt_parts(args.notes)
    user = f"Order data: {parts['order']}\n\nNotes:\n{parts['notes']}\n\n{{{{?question}}}}"
    template = Template(template=[SystemMessage(content=SYSTEM_PROMPT), UserMessage(content=user)])
    config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
        prompt=template, model=LLMModelDetails(name=args.model, params={"max_tokens": args.answer_tokens},
                                               timeout=60, max_retries=0))))
    try:
        service = OrchestrationService(config=config)
        response = service.run(placeholder_values={"question": QUESTION}, timeout=60)
    except Exception as error:
        sys.exit(f"The call failed ({type(error).__name__}: {error}). "
                 "Run python check_unit05.py to find the cause, or leave out --llm.")
    usage = response.final_result.usage
    estimate = sum(count_tokens(t) for k, t in parts.items() if k != "tools")
    print(f"Model: {response.final_result.model}\n")
    print(f"prompt_tokens       {usage.prompt_tokens:>6}   (this script estimated {estimate})")
    print(f"completion_tokens   {usage.completion_tokens:>6}")
    print(f"total_tokens        {usage.total_tokens:>6}")
    details = usage.prompt_tokens_details
    if details:
        print(f"cached_tokens       {details.cached_tokens or 0:>6}")
        print(f"cache_creation      {details.cache_creation_tokens or 0:>6}")
        if details.cached_tokens or details.cache_creation_tokens:
            print("Cache tokens appeared. Check how your model counts them (inside prompt_tokens or not) "
                  "and how your contract prices them before you add them up.")
    reasoning = usage.completion_tokens_details and usage.completion_tokens_details.reasoning_tokens
    if reasoning:
        print(f"reasoning_tokens    {reasoning:>6}   (billed as output, usually inside completion_tokens)")
    one = price([{"prefix": 0, "rest": usage.prompt_tokens, "out": usage.completion_tokens}], False, 0, False)
    print(f"\nAt the EXAMPLE rates this call costs {one['cost']:.5f} USD.")
    print("\n" + response.final_result.choices[0].message.content)


def main() -> None:
    parser = argparse.ArgumentParser(description="Token cost model for the blocked-orders assistant.")
    parser.add_argument("--orders", type=int, default=6000, help="blocked orders explained per month")
    parser.add_argument("--notes", type=int, default=3, help="note chunks retrieved per order (1 to 10)")
    parser.add_argument("--answer-tokens", type=int, default=220, help="tokens in a typical answer")
    parser.add_argument("--hit-rate", type=float, default=0.8,
                        help="share of first calls that find the prefix in the cache (0 to 1)")
    parser.add_argument("--sap", action="store_true", help="convert with SAP AI Core's example rates")
    parser.add_argument("--traces", type=Path, help="price the token counts in a traces.jsonl file")
    parser.add_argument("--llm", action="store_true", help="make one real call in SAP AI Core")
    parser.add_argument("--model", help="model name for --llm")
    args = parser.parse_args()
    if not 1 <= args.notes <= len(NOTES):
        sys.exit(f"--notes must be between 1 and {len(NOTES)}.")
    if not 0 <= args.hit_rate <= 1:
        sys.exit("--hit-rate must be between 0 and 1.")
    if args.traces:
        price_traces(args.traces, args.orders)
    elif args.llm:
        real_call(args)
    else:
        model_costs(args)


if __name__ == "__main__":
    main()
  1. Check that the file is in the right place. In the terminal, list the folder:

    • Windows (PowerShell):

      dir unit10\token_costs.py
    • macOS / Linux:

      ls unit10/token_costs.py

    You should see the file name. An error means the file is in another folder or has another name.

Step 3: Run the cost model

  1. Run the script with no options:

    python unit10/token_costs.py
  2. You should see this:

    Prompt length budget, pipeline design (estimated tokens)
    
    part        tokens  budget  check
    system         712     900  OK
    tools          349     450  OK
    order          114     250  OK
    notes          129     200  OK
    question        16      60  OK
    input         1320    1860  OK
    answer         220     250  OK
    
    Cost per transaction and per month (6,000 blocked orders, EXAMPLE rates)
    
    design    cache   calls   input  cached  output   per order   per month
    pipeline  off         1    1320       0     220 0.00484 USD   29.04 USD
    pipeline  on          1     259    1061     220 0.00342 USD   20.51 USD
    agent     off         3    3708       0     300 0.01042 USD   62.50 USD
    agent     on          3     525    3183     300 0.00517 USD   31.05 USD
    
    input = tokens billed at the full input rate; cached = prefix tokens written to or read from the cache;
    output = tokens the model writes. Rates are examples: replace RATES with your own.
  3. Read it from the top:

    • The budget table shows each part of the prompt. All are within budget. The answer is 220 tokens against a cap of 250.
    • In the cost table, the pipeline design without caching sends 1,320 input tokens and writes 220. The agent design makes 3 calls and sends 3,708 input tokens for the same answer.
    • With caching, the 1,061-token fixed prefix moves to the cached column. The agent design's cost per order drops by about half, because calls two and three read the prefix from the cache.

Step 4: Break the length budget

Retrieval often grows quietly: someone raises the number of passages "to be safe".

  1. Retrieve all ten notes instead of three:

    python unit10/token_costs.py --notes 10
  2. The notes line now fails its budget:

    notes          325     200  OVER
  3. Look at the cost table. The pipeline design without caching rises from 0.00484 to 0.00523 USD per order:

    design    cache   calls   input  cached  output   per order   per month
    pipeline  off         1    1516       0     220 0.00523 USD   31.39 USD
    pipeline  on          1     455    1061     220 0.00381 USD   22.86 USD
    agent     off         3    3904       0     300 0.01081 USD   64.85 USD
    agent     on          3     721    3183     300 0.00557 USD   33.40 USD

    The notes come after the fixed prefix, so caching doesn't help with them. Only retrieving fewer or shorter passages does.

Step 5: See when caching costs more

The script assumes that 80% of first calls find the prefix in the cache (--hit-rate 0.8). Try the worst case, where the cache has always expired:

  1. Run:

    python unit10/token_costs.py --hit-rate 0
  2. Compare the pipeline rows:

    design    cache   calls   input  cached  output   per order   per month
    pipeline  off         1    1320       0     220 0.00484 USD   29.04 USD
    pipeline  on          1     259    1061     220 0.00537 USD   32.22 USD
    agent     off         3    3708       0     300 0.01042 USD   62.50 USD
    agent     on          3     525    3183     300 0.00713 USD   42.76 USD

    With no hits, the pipeline design pays the 1.25 times write premium on every call and gets nothing back: caching costs more than no caching. The agent still gains, because its later calls come seconds after the first.

  3. What this means for you: caching pays when the same prefix is used again within its lifetime. A clerk working through a queue does that. An assistant used a few times a day does not.

Step 6: Shorter answers and a busier month

  1. Ask for 120-token answers and a month with 20,000 blocked orders:

    python unit10/token_costs.py --answer-tokens 120 --orders 20000
  2. Compare with Step 3:

    design    cache   calls   input  cached  output   per order   per month
    pipeline  off         1    1320       0     120 0.00384 USD   76.80 USD
    pipeline  on          1     259    1061     120 0.00242 USD   48.37 USD
    agent     off         3    3708       0     200 0.00942 USD  188.32 USD
    agent     on          3     525    3183     200 0.00417 USD   83.49 USD

    Cutting 100 output tokens saved 0.001 USD per order in the pipeline design, the same as cutting 500 input tokens would at these rates. Output is the expensive side.

  3. Try --answer-tokens 300. The budget table reports OVER: raise max_tokens or ask for less: the answer would be cut off at 250 tokens.

Step 7: Switch to SAP's capacity-unit arithmetic

  1. Run:

    python unit10/token_costs.py --sap
  2. The cost table now shows capacity units, using the fictitious conversion rates from SAP's metering page:

    design    cache   calls   input  cached  output   per order   per month
    pipeline  off         1    1320       0     220  0.00415 CU    24.93 CU
    agent     off         3    3708       0     300  0.00973 CU    58.41 CU
  3. There are no caching rows. SAP's metering page doesn't say how cached tokens convert, so the script refuses to guess. Look up your model's rates in SAP Note 3437766 (you need an SAP for Me login) and put them in RATES.

  4. Check the arithmetic for the pipeline row by hand: 1,320 / 1,000 x 0.00112 + 220 / 1,000 x 0.00320 = 0.0021824 GenAI tokens. Times 1.90385 gives 0.00415 CU.

Step 8: Price the tokens you logged

The observability topic's script writes unit10/traces.jsonl, with the input and output tokens of every model call. Its --incident option switches half the requests to prompt version v2, which sends the full order history.

  1. If you don't have the file, or want the incident in it, run the observability script again:

    python unit10/observe_assistant.py --incident
  2. Price the logged calls:

    python unit10/token_costs.py --traces unit10/traces.jsonl
  3. You should see this:

    Logged model calls in traces.jsonl, priced at EXAMPLE rates
    
    prompt    calls  avg in  avg out    per call   per month
    v1           98     819       94 0.00257 USD   15.44 USD
    v2           99    1638       83 0.00410 USD   24.61 USD
    
    per month = per call x 6,000 calls. Change it with --orders.

    Prompt v2 doubled the input tokens and raised the cost per call by about 60%. In the observability topic, the same release was the one where answers got worse. A change that costs more and helps less is exactly what token logging is there to catch.

  4. If you see No file unit10/traces.jsonl, run step 1 first.

Step 9 (optional): Read the usage of one real call in SAP AI Core

This step calls a real model, so it needs your SAP AI Core trial and the AICORE_ lines in .env from Set up for Unit 5. During the trial there is no charge; after it, each call has a small per-request charge.

  1. Check that Unit 5 still works:

    python check_unit05.py
  2. Pick a model name your account offers. check_unit05.py prints some, or run python unit05/choose_model.py catalog from the model topic.

  3. Make one call, putting your model's name in place of MODEL_NAME:

    python unit10/token_costs.py --llm --model MODEL_NAME
  4. You should see the usage SAP reports, the script's own estimate, and the answer. Your numbers depend on the model; this shape is what matters:

    Model: MODEL_NAME
    
    prompt_tokens         1012   (this script estimated 971)
    completion_tokens      180
    total_tokens          1192
    cached_tokens            0
    cache_creation           0
    
    At the EXAMPLE rates this call costs 0.00382 USD.
    
    Order 9000123 for Northwind Tools GmbH is blocked ...
  5. Compare prompt_tokens with the estimate. The gap is the 4-byte rule being off for this model's tokenizer, plus any formatting added around the messages. A reasoning model also shows reasoning_tokens, billed as output.

The sample numbers above are illustrative. This step was not run against a live SAP AI Core account when this topic was written.

Step 10: Save your work

  1. traces.jsonl is output, not code. If you haven't already, keep it out of Git by adding this line to .gitignore in the course folder:

    unit10/*.jsonl
  2. Save:

    git add .gitignore unit10/token_costs.py
    git commit -m "Unit 10: token cost model for the blocked-orders assistant"

How the code works

Part What it does
RATES All prices in one place: provider-style rates, cache multipliers, minimum cacheable prefix, and SAP's example conversion rates
BUDGET, ANSWER_CAP The prompt length budget per part and the answer cap (max_tokens)
SYSTEM_PROMPT, TOOLS, ORDER, NOTES Made-up, SAP-shaped content; the system prompt and tools are the fixed prefix
count_tokens Estimates tokens as UTF-8 bytes divided by 4, the tiktoken README's rule of thumb
calls_for Lists the model calls in one transaction. The agent design adds the order and the notes to the history call by call
price Applies the cost formula: uncached input, cache writes, cache reads and output. In SAP mode, converts to GenAI tokens and capacity units, without caching
show_budget Prints each part against its budget and warns when the prefix is too short to cache
price_traces Reads traces.jsonl, finds each request's prompt version on its root span, and averages tokens per model call
real_call With --llm, makes one call through SAP's orchestration service and prints final_result.usage

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
can't open file ... token_costs.py You are not in the course folder, or the file is named differently Run cd to orchestrate-course; check the file is in unit10
SyntaxError near the top of the file Part of the code was not pasted Select all in the file, delete, and paste the whole block again
Your numbers differ from the ones shown You changed RATES, BUDGET or the sample text Expected after edits. Paste the code again to get the published numbers
--notes must be between 1 and 10. You asked for more notes than the sample has Use a number from 1 to 10
No file unit10/traces.jsonl The observability script hasn't run in this folder Run python unit10/observe_assistant.py --incident first
Missing in .env: AICORE_... Step 9 can't find your SAP AI Core details Run unit05/key_to_env.py from the Unit 5 setup, or skip Step 9
ModuleNotFoundError: No module named 'gen_ai_hub' or 'dotenv' The Unit 5 libraries aren't in this Python Check for (.venv) in the prompt, then pip install -r requirements.txt
The call failed (AIAPIAuthenticatorException ...) Wrong or expired key, or the trial has ended Run python check_unit05.py, which names the cause
The call failed with a connection or proxy error The network blocks SAP AI Core Try another network, or ask IT to allow SAP AI Core's API address

The SAP way

As of October 2026, from SAP's documentation:

Metering on SAP AI Core

  • Tokens to capacity units. Generative AI hub use is metered in tokens, converted to GenAI tokens at per-model rates, then to capacity units. The rates are in SAP Note 3437766. The generative AI hub is only in SAP AI Core's extended plan.
  • No compute charges for foundation models. When you use foundation models and grounding, SAP waives SAP AI Core's compute, storage and baseline charges, and charges accrue on tokens instead. Custom models you train and serve yourself are billed on compute, storage and an hourly baseline.
  • Other meters. Grounding storage is metered in gigabyte days and retrieval in text blocks; content filtering and data masking have their own metrics. Stored inference records for observability are metered in data volume.
  • Your contract decides the money. Under SAP BTPEA and CPEA, you prepay cloud credits with an annual commitment; overages are billed in arrears at list price, and volume discounts are available. Pay-As-You-Go has no commitment and is billed monthly at non-discountable rates.

Prompt caching in the orchestration service

SAP's prompt caching page describes two modes, both in Orchestration V2:

  • Explicit caching with cache_control breakpoints on content blocks, for Anthropic Claude and Amazon Nova models. Claude supports breakpoints on system, message and tool blocks, up to four per request. The default lifetime is five minutes; selected Anthropic models accept "ttl": "1h".
  • Implicit caching for OpenAI and Gemini models, on by default.

A cached system message in a request to the orchestration service's /v2/completion endpoint looks like this (a sketch; it needs an SAP AI Core account with the extended plan, set up in Unit 5):

{
  "role": "system",
  "content": [
    {
      "type": "text",
      "text": "You help order-to-cash clerks understand blocked sales orders ...",
      "cache_control": { "type": "ephemeral" }
    }
  ]
}

The Python SDK has a matching CacheControl class with ttl set to "5m" or "1h". Responses report cache use in prompt_tokens_details.

Watching and limiting spend

  • Consumption report. In the SAP BTP cockpit, open the global account, choose Usage, filter by AI Core and export. Model consumption appears under Applications, and resource-group consumption under instance.
  • Costs and Usage page. On the consumption-based model, the global account's Costs and Usage page shows costs by service and subaccount, monthly trends, and estimates between balance statements. The monthly balance statement is the binding figure.
  • Budgets. The Budgets tab sets budget limits and email alerts to global account administrators. SAP's documentation states that the system does not suspend costs or service usage when a budget is exceeded, and that alerts can be delayed.

SAP-delivered AI

Joule and other AI features inside SAP applications are paid through your subscription and, for premium capabilities and agents, AI Units. The metering above doesn't apply to them. Joule and Joule Studio and Joule agents and agent orchestration cover what consumes AI Units.

Build vs. SAP

Situation Use Why
Early estimate before any code A spreadsheet or this topic's script with your contract rates Shows cost per transaction and the effect of design choices
Custom app on BTP using models in the generative AI hub Token usage from each response, logged per request; BTP Usage export for the bill Your own logs explain why; the BTP report says how much
Long fixed instructions or many tools, steady traffic Prompt caching in the orchestration service Repeated prefix billed at the cache rate where the model supports it
Overnight jobs, such as classifying all open orders A provider's batch interface, where your model and contract offer one Anthropic's cookbook describes a 50% discount for work that can wait up to 24 hours
Joule skills and agents in SAP applications SAP's subscription and AI Units Not metered in tokens to you; plan in AI Units
Hard spending limits Your own caps in the app BTP budgets alert but don't stop usage

Production concerns

  • Log tokens per request, with context. Record input, output, cached and reasoning tokens with the feature, prompt version, model and tenant on every call, as the observability topic does with gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Without them, a monthly bill can't be traced to a cause.
  • Enforce limits in code. Set max_tokens on every call, a step limit on every agent loop, and a daily token budget per user or per feature that refuses or queues work when spent. SAP BTP budgets only send alerts.
  • Count retries and fallbacks. A retry, a fallback model or a hedged request is a second paid call. The latency topic's retries belong in this cost model.
  • Watch prompts in releases. A new example in the system prompt is a cost change on every call. Review prompt diffs for token impact, as Step 8 showed with prompt v2.
  • Keep the cache prefix stable. Don't put dates, user names or order numbers in the system prompt. Anthropic's cookbook names dynamic content in the prefix as the most common way to break the cache.
  • Security and authorizations. Token counts are safe to log; prompt and answer text may hold customer data and must follow your logging rules. Never trade the user's SAP authorizations for a cheaper shared technical user. Response caching, where another user could receive a stored answer, is a different risk, covered in the next topic of this unit.
  • Peaks and rate limits. Month-end close raises volume and request rates at the same time. Plan cost at peak, and check that the rate limits on your SAP AI Core tenant fit it.
  • Clean core. All of this lives in your side-by-side app on BTP. S/4HANA needs no change to measure or reduce token cost.

Pitfalls

  • Pricing per token, not per transaction. The per-token rate tells you little until it is multiplied by calls per transaction and volume.
  • Forgetting the invisible tokens. Tool definitions, examples, history and hidden reasoning are all billed.
  • Adding cached tokens twice. Some providers count cached tokens inside the prompt count, others outside it. Check before summing.
  • Caching a prefix that changes. A timestamp at the top of the system prompt makes every call a cache write.
  • Caching rarely used prompts. With a write premium and no reads, caching costs more.
  • Using max_tokens to shorten answers. It cuts answers off. Ask for the length you want instead.
  • Using made-up or old rates in a business case. Rates differ by model and contract, and change. Use your current contract rates.
  • Leaving out the other meters. Retrieval, filtering and masking are close to half of SAP's own example.

Exercise

Produce the cost section of a business case for the blocked-orders assistant. The AI business case topic later in this unit uses it.

  1. Open unit10/token_costs.py and change RATES["input_per_million"] and RATES["output_per_million"] to the list prices of a model you might use, from the provider's pricing page. Write the date and the page's address in a comment above RATES.

  2. Choose a realistic month. Ask someone in order-to-cash, or assume 6,000 blocked orders and 15,000 at month-end close.

  3. Run the model for both months, for the design you would build:

    python unit10/token_costs.py --orders 6000
    python unit10/token_costs.py --orders 15000
  4. Run once more with the answer length you think clerks need, for example --answer-tokens 150.

  5. Create unit10/cost-model.md with a table: month (normal and peak), design, caching on or off, cost per order, cost per month. Add one line for the rates you used and their date.

  6. Below the table, write three sentences: which design you would build and why, what single change saves the most, and which non-model costs (retrieval, filtering, masking) are still missing.

  7. Commit cost-model.md.

Done when cost-model.md shows cost per order and per month for a normal and a peak month, names the rates and their date, and recommends one design with the change that saves the most.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the script's default run, the agent design sends 3,708 input tokens for the same answer that the pipeline design gets from 1,320. What causes most of the difference?

    Answer: B. The agent makes three calls, and each one carries the roughly 1,060-token prefix plus everything gathered so far. The answer length and the notes are the same in both designs; the tool definitions are in the prefix of both.
  2. 2With --hit-rate 0, the pipeline design with caching costs more than without it. Why?

    Answer: C. At a hit rate of zero, each call writes the prefix at 1.25 times the input rate and nothing is ever read at the discounted rate. The agent design still gains because its second and third calls hit the cache.
  3. 3Your team adds today's date to the first line of the system prompt so the model knows it. What happens to prompt caching?

    Answer: D. Caching reuses a prefix only when it is identical. A changing value at the top makes the prefix different, so calls write a new cache instead of reading the old one. Put changing values after the fixed part.
  4. 4A response from an Anthropic model through SAP's orchestration service shows prompt_tokens 3 and cached_tokens 4,657. How should your cost code treat these numbers?

    Answer: B. In SAP's example the cached count is far larger than the prompt count, so it can't be included in it. Providers differ on this, so check a real response for your model and price cached tokens at the rate your contract gives them.
  5. 5Why does the script refuse to show caching rows with --sap?

    Answer: C. SAP AI Core does support prompt caching, but the metering page opened for this topic doesn't give a conversion for cached tokens. Rather than guess, the script prices them as normal input and points you to SAP Note 3437766.
  6. 6Retrieving ten notes instead of three pushes the notes part over its budget. Which change brings it back without losing the answer's grounding?

    Answer: D. The notes vary per order and come after the fixed prefix, so caching can't help them. Lowering max_tokens affects the answer, not the prompt. Better retrieval, such as reranking, keeps the useful passages and drops the rest.
  7. 7The BTP budget for AI Core is exceeded on the 20th of the month. What happens to the assistant's calls?

    Answer: A. SAP's documentation states that the system does not suspend costs or usage when a budget is exceeded, and that alerts can be delayed. If you need a hard stop, build per-user or per-feature caps into the app.
  8. 8A product owner wants a shorter answer and suggests lowering max_tokens from 250 to 120. What do you do?

    Answer: C. The model doesn't plan around max_tokens; when it hits the cap, the answer is cut off mid-sentence. Specify the length and shape you want in the prompt, and keep max_tokens above the longest legitimate answer to stop runaways.
  9. 9In SAP's worked example of a RAG assistant, why does the total come to 418.55 CU when the model calls are 232.25?

    Answer: B. SAP's example adds grounding storage and retrieval, content filtering and data masking, which together come to 186.3 CU. Those modules are metered separately from tokens and belong in every cost model.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in