Orchestrate

Caching for LLM applications

Use prompt, response, semantic and embedding caches to cut cost and wait time, without serving a stale answer or another user's data.

Updated Oct 7, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

A cache keeps the result of earlier work so you don't pay for it twice. AI applications use four kinds, and they are easy to confuse.

  • Prompt cache. The model provider keeps the processed start of your prompt, such as long instructions, for a few minutes. The model still writes a new answer every time. It is cheaper and faster, and the answer is still fresh.
  • Response cache. Your application stores whole answers. When the same question comes again, it returns the stored answer and skips the model.
  • Semantic cache. A response cache that also matches questions worded differently but meaning the same thing.
  • Embedding cache. Stores the number lists (embeddings) your search uses, so the same text isn't converted twice.

The first saves money with little risk. The other three can return an answer that was right for someone else, or right an hour ago. Every response cache needs two rules: who may reuse an answer, and when an answer stops being true.

Why it matters to the business

Take the running example of this unit: an assistant that explains blocked sales orders to order-to-cash clerks.

Clerks ask the same things again and again. "What does delivery block 01 mean?" comes up every day. Two clerks often look at the same urgent order in the same hour. Month-end close multiplies both.

A cache turns repeated questions into instant answers. In this topic's hands-on replay of one made-up morning, an exact-match cache answers 30% of questions without calling the model. Adding a carefully limited semantic cache raises that to 39%. Each hit costs nothing in model tokens and comes back in a fraction of a second instead of seconds.

The same replay shows three ways it goes wrong:

  • Stale answers. Credit management releases an order at 9:20. The cache still says "blocked" at 9:25, because it stored the answer at 9:01. A clerk tells the customer the wrong thing.
  • Leaked answers. A clerk in another sales organization asks about the same order. A cache shared by everyone returns the answer written for a colleague who may see it, including the customer's name and order value.
  • Wrong matches. A semantic cache decides that "Why is order 9000124 blocked?" means the same as "Why is order 9000123 blocked?" The wording is almost identical; the orders are not. In the replay, a semantic cache without safeguards reaches a 70% hit rate, and 11 of 23 answers are wrong.

A higher hit rate is worth nothing if the answers are wrong. The business question is not "how much can we cache?" but "which answers are safe to reuse, for whom, and for how long?"

How SAP does it

As of October 2026, from the SAP sources opened for this topic:

  • Prompt caching is available in the orchestration service of the generative AI hub in SAP AI Core. SAP's Python SDK added a CacheControl setting in September 2026, so developers can mark the fixed part of a prompt for caching. Responses report how many tokens came from the cache. Token economics covers which model families SAP documents for this and how it affects the bill.
  • Response and semantic caches are not a managed feature in the SAP AI Core sources opened for this topic. Your team builds them in its own application on SAP BTP. SAP HANA Cloud, with its vector engine, can store cached questions as vectors and find similar ones. A community member has published an open-source package that does this; it is not an SAP product.
  • Knowing when an answer is stale is where SAP helps most. SAP S/4HANA can publish business events, such as "sales order changed", through SAP Event Mesh. Your app can listen for them and drop stored answers about that order.

SAP-delivered AI, such as Joule, manages its own internals. These choices apply to AI you build.

Which cache for which job

Cache What is reused Saves Main risk Safe when
Prompt cache The processed start of the prompt Input cost and time to first answer Low; costs extra if the cache expires unused The start of the prompt is identical across calls
Response cache Whole answers to identical questions The whole model call Stale or leaked answers Answers are shared only between users with the same access, and dropped when data changes
Semantic cache Answers to similar questions The whole model call, more often Wrong matches, plus the response cache risks Limited to general questions, or exact IDs must match too
Embedding cache Vectors for text already seen Embedding calls Mixing vectors from two models The model name is part of the key

A practical order: switch on prompt caching first, add an exact response cache for general questions second, and consider a semantic cache last, with an evaluation set to prove it.

Questions to ask

  • Which of the four caches does this design use, and what does each one store?
  • For a response cache: what is in the key? Does it include the user's access, the prompt version and the model?
  • Can two users with different SAP authorizations ever receive the same stored answer?
  • How does a stored answer find out that the order, invoice or document changed? What is the maximum age if that signal is lost?
  • What does the hit rate look like on our real questions, and how many cached answers were checked against fresh ones?
  • Where are cached answers stored, for how long, and does that follow our data retention and privacy rules?
  • When we call models through SAP's generative AI hub, how is the provider's prompt cache isolated between SAP customers? (The sources opened for this topic don't say; ask SAP.)
  • Who can clear the cache in an incident, and how fast?

Common misconceptions

  • "Prompt caching means the model reuses old answers." It reuses the processed prompt, not the answer. The model writes a new answer each time.
  • "A higher hit rate is always better." Only if the hits are correct. A loose semantic cache scores high and answers wrongly.
  • "Caching is a performance detail for engineers." Who may reuse which answer is an access-control decision. It belongs in the security review.
  • "A time limit is enough." Within the limit, the answer can be wrong the minute the data changes. Use change signals where you can, and the time limit as a backstop.
  • "Similar questions have similar answers." Two questions about different orders look almost the same to an embedding model, but need different answers.
  • "The provider's cache is private by design." The two providers checked for this topic isolate caches per organization today. Researchers found shared caches across customers at several providers in 2024, so ask.

Key terms

  • Cache hit / miss: the stored result was found and used / not found, so the work was done.
  • Hit rate: the share of requests answered from the cache.
  • Cache key: everything used to decide whether a stored answer matches a new request.
  • Prompt caching: the provider reuses the processed start of a prompt across calls.
  • Response cache: stores whole answers for identical requests.
  • Semantic cache: returns a stored answer when a new question is similar enough in meaning.
  • Embedding cache: stores the vectors computed for text.
  • Invalidation: removing stored answers that are no longer true.
  • Time to live (TTL): the longest time a stored item may be reused.
  • Business event: a message SAP S/4HANA sends when something happens, such as a sales order changing.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1The team says it "turned on prompt caching" for the blocked-orders assistant. What does that change for clerks?

    Answer: C. Prompt caching reuses the processed start of the prompt, such as the fixed instructions. The model still reads the current order and writes a new answer, so freshness and access are unchanged.
  2. 2A vendor proposes a response cache with a 70% hit rate. What should you ask before agreeing?

    Answer: C. In this topic's replay, the 70% hit rate came from a semantic cache that answered 11 of 23 questions wrongly. A hit rate means something only next to a measured error rate.
  3. 3Credit management releases a blocked order at 9:20. At 9:25 a clerk still sees "blocked". What is the most likely cause?

    Answer: B. A response cache keeps the answer until it is invalidated or expires. Without a change signal, such as an S/4HANA business event, the stored "blocked" answer outlives the release.
  4. 4Clerks in two sales organizations use the same assistant. What must the response cache guarantee?

    Answer: D. A shared cache can hand one clerk an answer written for a colleague who is allowed to see that order. Reuse must be limited to users whose authorizations would produce the same answer.
  5. 5Why should a semantic cache not treat "Why is order 9000124 blocked?" as a match for the same question about order 9000123?

    Answer: B. Embedding models judge meaning mostly from words and give the exact digits little weight. The two questions look almost the same, so identifiers such as order numbers must match exactly before a stored answer is reused.
  6. 6Where does SAP give the most direct help with keeping cached answers fresh?

    Answer: C. S/4HANA can publish events such as SalesOrder.Changed through SAP Event Mesh, and your app can drop stored answers about that order. A managed response or semantic cache was not found in the SAP AI Core sources for this topic.
  7. 7Which order of adoption fits most teams?

    Answer: A. Prompt caching saves money without changing who sees what. Exact caching of general questions is the next safest. A semantic cache has the most ways to go wrong, so it comes last, with an evaluation set.
Deep layer · 40 min read

Mental model

A cache entry is a promise: "this stored answer is the one you would compute now, for this user."

Everything in this topic follows from keeping that promise. It breaks in exactly two ways:

  1. The key is too small. Something that changes the answer is missing from the key, so a request matches an entry it shouldn't. Leave out the user's authorization and one clerk gets another's answer. Leave out the order number, as a semantic match does in effect, and a clerk gets an answer about another order. Leave out the prompt version and a fixed prompt keeps serving answers from the broken one.
  2. The entry outlives its truth. The data behind the answer changed, and nothing removed the entry.

Prompt caching is the exception that proves the rule. The provider keeps only the processed input, keyed on the exact bytes of the prefix. The answer is computed fresh every time, so neither failure applies to what the user sees. That is why it is the safe first step, and the other three caches need design.

How it works

Four caches in one request

flowchart LR
  Q[Clerk's question] --> R{Response cache<br/>exact key}
  R -- hit --> A[Answer]
  R -- miss --> E[Embedding cache] --> S{Semantic cache<br/>scope + IDs + similarity}
  S -- hit --> A
  S -- miss --> D[Read SAP with<br/>the user's identity]
  D --> M[Model call<br/>prompt cache at provider]
  M --> W[Store answer] --> A
  V[SAP business event] -. drop entries .-> R
  V -. drop entries .-> S

The cheapest check runs first. An exact hit costs a dictionary lookup. A semantic lookup costs an embedding, which the embedding cache often saves too. Only a miss on both reads SAP and calls the model, and the model call itself uses the provider's prompt cache.

Prompt caching: reuse the processed prefix

A model turns the prompt into internal state before it writes a word. Providers can keep that state for a prefix and reuse it when a later request starts with exactly the same content. From the Anthropic and OpenAI documentation opened for this topic (October 2026):

Anthropic Claude OpenAI
How it turns on cache_control on a block (up to 4 breakpoints), or one top-level setting Automatic on prompts above the minimum
What must match The exact prefix, built in the order tools, system, messages The exact prefix
Minimum length 512 to 4,096 tokens, depending on the model 1,024 tokens
Lifetime 5 minutes, refreshed free on every hit; 1 hour optional Model-dependent; retention options documented per model
Price Writes 1.25x input (5 min) or 2x (1 hour); reads 0.1x on most models Reads at a fraction of input; write pricing depends on the model
Isolation Per organization; per workspace on the Claude API Not shared across organizations
Usage fields cache_read_input_tokens, cache_creation_input_tokens, input_tokens Cached tokens in the input token details

Three rules decide whether it works:

  1. Fixed content first, changing content last. Anthropic's documentation shows that a change to the tool definitions invalidates everything after them. A timestamp or user name at the top of the system prompt makes every request a new prefix and every call a cache write.
  2. Reach the minimum. Below it, nothing is cached and no error is raised. Check the usage fields: both cache counters at zero means no caching.
  3. Reuse within the lifetime. A five-minute cache helps a clerk working through a queue, not a report run twice a day.

Anthropic's documentation also warns that input_tokens counts only the tokens after the last breakpoint, so total input is the sum of three fields. Token economics shows how the counters differ between providers and how to price them.

Response caching: the key is the design

A response cache maps a key to a stored answer. Everything that can change the answer belongs in the key:

Part of the key Why Example
Normalized question Trivial differences shouldn't cause misses lower case, single spaces, no final ?
User's authorization scope Users who may see different data get different answers the sales organizations the clerk may display
Prompt version and model A new prompt or model must not reuse old answers prompt-v3, the model name
Generation settings Temperature and output format change the answer temperature=0, JSON or text
Data version, where cheap Ties the answer to the data it was built from an order's ETag or last-changed time

Two design choices matter beyond the key:

  • Authorization scope, not user ID. Keying on the user ID is safe but shares nothing. Keying on the part of the authorization that changes the answer (here, the sales organizations) lets clerks with the same access share answers. Get this wrong and you have built a way around SAP authorizations: on a cache hit, your app never reads SAP, so SAP never checks anything. Grounding on SAP data with authorizations explains why reads must run as the user.
  • What not to store. Don't cache refusals ("you are not authorized"): if access is granted later, the refusal hides the order. Don't cache answers to questions that contain personal data unless your retention rules allow it. Don't cache errors.

Semantic caching: similar is not the same

A semantic cache embeds the question, searches stored questions by vector similarity, and returns the best stored answer above a threshold. Embeddings and semantic similarity explains the similarity score. Libraries differ in how they state the threshold: RedisVL's SemanticCache uses a cosine distance with a default of 0.1, where lower is stricter. SAP HANA Cloud's COSINE_SIMILARITY returns values up to 1, where higher is more similar.

The problem is that embeddings blur exactly what business questions turn on. "Why is order 9000123 blocked?" and "Why is order 9000124 blocked?" differ in one digit. An embedding model sees two nearly identical sentences. So does a block code: "What does delivery block 01 mean?" against "block 05".

Two filters make a semantic cache safe enough to try:

  1. Scope filter. Only search entries written for the same authorization scope. RedisVL supports this with filterable_fields and a filter_expression on lookup.
  2. Identifier filter. Extract the identifiers (order numbers, codes, dates, amounts) and require them to match exactly. The embedding then only decides whether the wording means the same thing.

Then measure. A threshold is a trade-off between hit rate and wrong answers, and only an evaluation set of real question pairs shows where it sits. The community author of a HANA-based semantic cache warns against it as a day-one feature for the same reason.

Embedding caching: the safe one

An embedding depends only on the text and the model. Key it on both, for example a hash of the model name and the text, and it never goes stale. RedisVL's EmbeddingsCache keys on exact content plus model name. The one trap is changing the embedding model: vectors from two models are not comparable, so the model name must be in the key or the cache must be emptied.

Invalidation: three signals

Signal How it works Good for Weakness
Time to live Each entry expires after a fixed time Everything, as a backstop Wrong until it expires; short TTLs lower the hit rate
Change events S/4HANA publishes an event; your app drops entries about that object Order-specific answers Events can be delayed or lost; not every change has an event
Version check Read the object's ETag or change time; include it in the key High-stakes answers Costs one SAP read per request, though far cheaper than a model call

SAP's learning material for event-based replication lists two sales order events, sap.s4.beh.salesorder.v1.SalesOrder.Created.v1 and sap.s4.beh.salesorder.v1.SalesOrder.Changed.v1, delivered through SAP Event Mesh, and notes that deletions are not covered in that scenario. That gap is why a TTL backstop stays even when you use events.

For S/4HANA Cloud OData APIs, an SAP blog explains that a GET returns an ETag that changes whenever the document changes. That makes the ETag a natural data version for a cache key. ETag behaviour depends on the API, so check the one you call.

A cache is also an attack surface

  • Timing. A cache hit is faster than a miss, so response time can reveal that someone asked a question before. Stanford researchers audited 17 model API providers in late 2024 and detected prompt caching at 8, with caches shared globally across users at 7. After disclosure, at least five changed their systems or documentation. Your own response cache has the same property: scope it so a hit can only reveal what the user could see anyway.
  • Poisoning. If an attacker gets a bad answer stored, for example through a prompt injection in a document, every later hit serves it. Cache only answers that passed your output checks, and keep a way to purge entries.
  • Retention. Cached answers are stored data. Customer names and amounts in them fall under the same retention and deletion rules as any other copy.

Build it yourself: four caches and three failures

You will replay one made-up morning of 23 questions from three order-to-cash clerks against the blocked-orders assistant. The script answers each question from the "model" (a stand-in that costs tokens and time) or from a cache, and checks every answer against a fresh one. You will see the hit rate go up, then make the cache fail in three ways: a stale answer, a leaked answer and a wrong semantic match. Last, you will see why prompt layout decides whether the provider's prompt cache works.

Before you start: complete Set up your computer for this course and Set up for Unit 10. They create your orchestrate-course folder with .venv and the unit10 folder. The script uses only built-in Python: nothing to install.

flowchart LR
  T[23 questions<br/>+ 1 SAP change] --> C{Cache?}
  C -- hit --> A[Stored answer]
  C -- miss --> M[Stand-in model] --> A
  A --> K[Compare with<br/>fresh answer]
  K --> R[Hit rate, wrong<br/>answers, cost]

What you need

  • Your course folder with .venv from the earlier setup topics.
  • About 40 to 50 minutes.
  • Cost: free. No account, no key, no network.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check your Python version (the same command on every system):

    python --version
    Python 3.14.4

    Any version from 3.10 up works.

Step 2: Create the script

  1. In VS Code, right-click the unit10 folder, choose New File, name it llm_cache.py, paste the code below and save.
"""Unit 10: caching for the blocked-orders assistant, and the ways a cache can go wrong.

The script replays one made-up morning of questions from three order-to-cash clerks. Each
question either goes to the "model" (a stand-in function that costs tokens and time) or is
answered from a cache. After every request the script also computes the answer the clerk
SHOULD have received from fresh data, so it can count wrong answers, not just hits.

Everything is made up and runs offline with built-in Python. No model is called; no account.

  python unit10/llm_cache.py                       no cache: the baseline
  python unit10/llm_cache.py --cache exact         exact-match response cache
  python unit10/llm_cache.py --cache semantic      exact cache first, then a semantic cache
  python unit10/llm_cache.py --cache semantic --no-id-check       let similar wording match other IDs
  python unit10/llm_cache.py --cache exact --no-invalidate        ignore SAP change events
  python unit10/llm_cache.py --cache exact --shared               one cache for all users, no check
  python unit10/llm_cache.py --cache exact --ttl 10               reuse an answer for 10 minutes at most
  python unit10/llm_cache.py --cache semantic --log               print one line per request
  python unit10/llm_cache.py --prompt-layout       which prompt layouts can use the provider's prompt cache
"""
import argparse
import hashlib
import math
import re
import sys

# ---------------------------------------------------------------------------------------------
# EXAMPLE rates and timings, made up for this course. Replace them with your own.
# ---------------------------------------------------------------------------------------------
INPUT_PER_MILLION = 2.00          # USD per million input tokens
OUTPUT_PER_MILLION = 10.00        # USD per million output tokens
EMBED_PER_MILLION = 0.02          # USD per million tokens sent to the embedding model
TOKENS_IN, TOKENS_OUT = 1320, 220  # one model call, as in the token economics topic
MODEL_SECONDS = 4.0               # time for one model answer
CACHE_SECONDS = 0.05              # time for a cache lookup

# ---------------------------------------------------------------------------------------------
# Made-up, SAP-shaped data. Block reason codes are configured per SAP system; these meanings
# belong to this made-up company only.
# ---------------------------------------------------------------------------------------------
BLOCK_GUIDE = {
    "01": "credit limit: the customer's open items exceed the credit limit, and credit management decides",
    "02": "missing export papers: shipping waits for the customs team to attach the documents",
    "05": "quality hold: an item waits for a quality inspection result",
}

ORDERS = {  # the "SAP system": sales order -> current data. 'version' plays the role of an ETag.
    "9000123": {"SalesOrganization": "1010", "CustomerName": "Northwind Tools GmbH",
                "TotalNetAmount": "48200.00 EUR", "DeliveryBlockReason": "01", "version": 1},
    "9000124": {"SalesOrganization": "1010", "CustomerName": "Contoso Freight",
                "TotalNetAmount": "12900.00 EUR", "DeliveryBlockReason": "05", "version": 1},
    "9000130": {"SalesOrganization": "1710", "CustomerName": "Fabrikam Retail",
                "TotalNetAmount": "7300.00 USD", "DeliveryBlockReason": "02", "version": 1},
}

USERS = {  # which sales organizations each clerk may see (their SAP authorization, simplified)
    "anna": {"1010"},
    "ben": {"1010"},
    "chen": {"1710"},
}

# One morning: (minute, user, question). An entry with user "EVENT" is a change in SAP.
TRAFFIC = [
    (1, "anna", "Why is order 9000123 blocked?"),
    (2, "ben", "What does delivery block 01 mean?"),
    (3, "anna", "What does delivery block 01 mean?"),
    (5, "ben", "why is order 9000123 blocked"),
    (6, "chen", "What does delivery block 02 mean?"),
    (8, "anna", "What is block reason 01?"),
    (9, "ben", "Explain delivery block code 01"),
    (11, "anna", "Why is order 9000124 blocked?"),
    (12, "chen", "Why is order 9000130 blocked?"),
    (14, "ben", "Why is order 9000124 blocked?"),
    (15, "chen", "Why is order 9000123 blocked?"),
    (16, "anna", "What does delivery block 05 mean?"),
    (18, "ben", "What is block reason 05?"),
    (20, "EVENT", "sap.s4.beh.salesorder.v1.SalesOrder.Changed.v1 9000123 DeliveryBlockReason="),
    (21, "anna", "Why is order 9000123 blocked?"),
    (22, "ben", "What does delivery block 02 mean?"),
    (24, "chen", "What does delivery block 01 mean?"),
    (25, "ben", "Why is order 9000123 blocked?"),
    (27, "chen", "Why is order 9000130 blocked?"),
    (29, "anna", "Explain delivery block code 05"),
    (31, "ben", "What is block reason 02?"),
    (33, "anna", "Why is order 9000124 blocked?"),
    (35, "chen", "What does delivery block 05 mean?"),
    (38, "ben", "Why is order 9000124 blocked?"),
]

ORDER_RE = re.compile(r"\b(9\d{6})\b")
ID_RE = re.compile(r"\b\d+\b")  # any number in a question: an order number or a block code


# ---------------------------------------------------------------------------------------------
# The assistant without a cache
# ---------------------------------------------------------------------------------------------
def fresh_answer(user: str, question: str) -> str:
    """What the assistant answers from current data. Order questions check the user's
    authorization first, the way a real app reads SAP with the user's own identity."""
    match = ORDER_RE.search(question)
    if match:
        number = match.group(1)
        order = ORDERS.get(number)
        if order is None:
            return f"I can't find order {number}."
        if order["SalesOrganization"] not in USERS[user]:
            return f"You are not authorized to see order {number}."
        reason = order["DeliveryBlockReason"]
        if not reason:
            return f"Order {number} for {order['CustomerName']} has no delivery block any more."
        return (f"Order {number} for {order['CustomerName']} ({order['TotalNetAmount']}) is blocked: "
                f"{BLOCK_GUIDE[reason]}.")
    code = re.search(r"\b(0\d)\b", question)
    if code and code.group(1) in BLOCK_GUIDE:
        return f"Delivery block {code.group(1)} means {BLOCK_GUIDE[code.group(1)]}."
    return "I can only explain blocked orders and delivery block reasons."


# ---------------------------------------------------------------------------------------------
# A toy embedding model. Real systems call an embedding model (see the embeddings topic); this
# stand-in counts words after a few clean-ups, so that similar questions get similar vectors.
# Like real embedding models, it gives numbers little weight: "order 9000123" and
# "order 9000124" come out almost identical.
# ---------------------------------------------------------------------------------------------
STOP = {"what", "does", "is", "a", "the", "why", "please", "me", "of"}
SYNONYMS = {"code": "reason", "explain": "mean", "means": "mean", "delivery": "block"}
EMBED_MODEL = "toy-bag-of-words-v1"


def embed(text: str) -> dict:
    words = re.findall(r"[a-z0-9]+", text.lower())
    words = [SYNONYMS.get(w, w) for w in words if w not in STOP]
    vector = {}
    for w in words:
        vector[w] = vector.get(w, 0) + (0.3 if w.isdigit() else 1.0)
    norm = math.sqrt(sum(v * v for v in vector.values())) or 1.0
    return {w: v / norm for w, v in vector.items()}


def cosine(a: dict, b: dict) -> float:
    return sum(v * b.get(w, 0.0) for w, v in a.items())


class EmbeddingCache:
    """Exact-match cache for embeddings: same model + same text = same vector, so it never
    goes stale. The model name is part of the key: vectors from two models don't mix."""

    def __init__(self):
        self.store, self.hits, self.misses, self.tokens = {}, 0, 0, 0

    def get(self, text: str) -> dict:
        key = hashlib.sha256(f"{EMBED_MODEL}\n{text}".encode("utf-8")).hexdigest()
        if key in self.store:
            self.hits += 1
        else:
            self.misses += 1
            self.tokens += max(1, len(text.encode("utf-8")) // 4)
            self.store[key] = embed(text)
        return self.store[key]


# ---------------------------------------------------------------------------------------------
# The response cache
# ---------------------------------------------------------------------------------------------
def normalize(question: str) -> str:
    """Make trivially different questions identical: case, spaces, final punctuation."""
    return re.sub(r"\s+", " ", question.lower()).strip(" ?.!")


def identifiers(question: str) -> tuple:
    """Order numbers and codes in the question. Embeddings blur them, so they are matched exactly."""
    return tuple(sorted(ID_RE.findall(question)))


def scope_of(user: str) -> str:
    """The part of the user's authorization that changes the answer. Two users with the same
    scope may share answers; users with different scopes never do."""
    return ",".join(sorted(USERS[user]))


class ResponseCache:
    def __init__(self, args):
        self.args = args
        self.exact = {}      # key -> entry
        self.semantic = []   # entries with a vector
        self.embeddings = EmbeddingCache()

    def key(self, user: str, question: str) -> str:
        scope = "shared" if self.args.shared else scope_of(user)
        # A real key also holds the prompt version and the model name: a new prompt or model
        # must not reuse answers written by the old one.
        return f"prompt-v3|model-x|{scope}|{normalize(question)}"

    def lookup(self, user: str, question: str, minute: int):
        """Return (answer, kind) on a hit, (None, None) on a miss."""
        entry = self.exact.get(self.key(user, question))
        if entry and self.fresh(entry, minute):
            return entry["answer"], "exact"
        if self.args.cache == "semantic":
            vector = self.embeddings.get(normalize(question))
            scope = "shared" if self.args.shared else scope_of(user)
            ids = identifiers(question)
            best, best_score = None, 0.0
            for e in self.semantic:
                if e["scope"] != scope or not self.fresh(e, minute):
                    continue  # the scope filter: only answers written for the same scope
                if not self.args.no_id_check and e["ids"] != ids:
                    continue  # the identifier filter: order numbers and codes must match exactly
                score = cosine(vector, e["vector"])
                if score > best_score:
                    best, best_score = e, score
            if best and best_score >= self.args.threshold:
                return best["answer"], f"semantic {best_score:.2f}"
        return None, None

    def store(self, user: str, question: str, answer: str, minute: int):
        if answer.startswith("You are not authorized"):
            return  # never cache a refusal: it would hide the order from an authorized user later
        order = ORDER_RE.search(question)
        entry = {"answer": answer, "minute": minute, "order": order.group(1) if order else None,
                 "scope": "shared" if self.args.shared else scope_of(user), "ids": identifiers(question)}
        self.exact[self.key(user, question)] = entry
        if self.args.cache == "semantic":
            self.semantic.append(dict(entry, vector=self.embeddings.get(normalize(question))))

    def fresh(self, entry: dict, minute: int) -> bool:
        return self.args.ttl == 0 or minute - entry["minute"] < self.args.ttl

    def invalidate(self, order_number: str) -> int:
        """Called on a SalesOrder.Changed event: drop every answer about that order."""
        before = len(self.exact) + len(self.semantic)
        self.exact = {k: e for k, e in self.exact.items() if e["order"] != order_number}
        self.semantic = [e for e in self.semantic if e["order"] != order_number]
        return before - len(self.exact) - len(self.semantic)


# ---------------------------------------------------------------------------------------------
# Replay the morning
# ---------------------------------------------------------------------------------------------
def apply_event(text: str) -> str:
    """Change the made-up SAP data the way the event describes, and return the order number."""
    _, number, change = text.split(" ")
    field, value = change.split("=")
    ORDERS[number][field] = value
    ORDERS[number]["version"] += 1
    return number


def classify(user: str, question: str, answer: str, truth: str, changed: set) -> str:
    """Compare the answer the clerk got with the fresh one."""
    if answer == truth:
        return "ok"
    shown = ORDER_RE.search(answer)
    if shown and ORDERS[shown.group(1)]["SalesOrganization"] not in USERS[user]:
        return "LEAKED"   # the clerk saw an order outside their authorization
    asked = ORDER_RE.search(question)
    if asked and shown and asked.group(1) == shown.group(1) and asked.group(1) in changed:
        return "STALE"    # right order, but data from before a change in SAP
    return "WRONG"        # an answer about a different order or block code


def replay(args) -> None:
    cache = ResponseCache(args) if args.cache != "none" else None
    stats = {"requests": 0, "model": 0, "exact": 0, "semantic": 0, "stale": 0, "leaked": 0, "wrong": 0}
    seconds = 0.0
    changed = set()  # orders changed in SAP since the morning started
    for minute, user, text in TRAFFIC:
        if user == "EVENT":
            number = apply_event(text)
            changed.add(number)
            dropped = cache.invalidate(number) if cache and not args.no_invalidate else 0
            if args.log:
                print(f"{minute:>3}  EVENT {text.split(' ')[0].split('.')[-2]} for {number}: "
                      f"{'ignored' if args.no_invalidate else f'{dropped} cache entries dropped'}")
            continue
        stats["requests"] += 1
        truth = fresh_answer(user, text)
        answer, kind = cache.lookup(user, text, minute) if cache else (None, None)
        if answer is None:
            stats["model"] += 1
            seconds += MODEL_SECONDS
            answer, kind = truth, "model"
            if cache:
                cache.store(user, text, answer, minute)
        else:
            seconds += CACHE_SECONDS
            stats[kind.split()[0]] += 1
        verdict = classify(user, text, answer, truth, changed)
        if verdict != "ok":
            stats[verdict.lower()] += 1
        if args.log:
            print(f"{minute:>3}  {user:<5} {kind:<14} {verdict:<7} {text}")
    report(args, stats, seconds, cache)


def report(args, s: dict, seconds: float, cache) -> None:
    if args.log:
        print()
    mode = args.cache
    flags = [f for f, on in (("no invalidation", args.no_invalidate), ("shared across users", args.shared),
                             ("no identifier check", args.no_id_check and args.cache == "semantic")) if on]
    print(f"Cache: {mode}{' | ' + ', '.join(flags) if flags else ''}"
          f"{f' | threshold {args.threshold}' if args.cache == 'semantic' else ''}"
          f"{f' | ttl {args.ttl} min' if args.ttl else ''}\n")
    hits = s["exact"] + s["semantic"]
    print(f"requests            {s['requests']:>5}")
    print(f"model calls         {s['model']:>5}")
    print(f"cache hits          {hits:>5}   (exact {s['exact']}, semantic {s['semantic']})")
    print(f"hit rate            {hits / s['requests']:>5.0%}")
    bad = s["stale"] + s["leaked"] + s["wrong"]
    print(f"wrong answers       {bad:>5}   (stale {s['stale']}, leaked {s['leaked']}, wrong order or code {s['wrong']})")
    model_cost = s["model"] * (TOKENS_IN * INPUT_PER_MILLION + TOKENS_OUT * OUTPUT_PER_MILLION) / 1e6
    embed_cost = (cache.embeddings.tokens if cache else 0) * EMBED_PER_MILLION / 1e6
    print(f"\nmodel cost          {model_cost:.5f} USD")
    if cache and args.cache == "semantic":
        e = cache.embeddings
        print(f"embedding calls     {e.misses:>5}   (embedding cache hits {e.hits}; cost {embed_cost:.6f} USD)")
    print(f"average wait        {seconds / s['requests']:>5.2f} s per request")
    print("\nRates and timings are EXAMPLES. 'wrong answers' compares each answer with a fresh one.")


# ---------------------------------------------------------------------------------------------
# Prompt layout: what a provider's prompt cache can reuse
# ---------------------------------------------------------------------------------------------
SYSTEM = ("You help order-to-cash clerks understand blocked sales orders in SAP S/4HANA. "
          "Use only the data you are given. Never change an order.\n" + "\n".join(
              f"Block {k}: {v}." for k, v in BLOCK_GUIDE.items()))


def layout(kind: str, user: str, minute: int, question: str) -> list:
    """Return the request as a list of blocks, in the order they are sent."""
    stamp = f"Current time: 2026-10-07 08:{minute:02d}. User: {user}."
    if kind == "bad":
        return [stamp + "\n" + SYSTEM, question]          # changing text FIRST
    return [SYSTEM, stamp + "\n" + question]              # fixed text first, changing text last


def prompt_layout() -> None:
    calls = [("anna", 1, "Why is order 9000123 blocked?"), ("ben", 2, "Why is order 9000124 blocked?"),
             ("anna", 3, "What does delivery block 01 mean?")]
    print("A provider's prompt cache reuses the longest identical START of a request.\n")
    for kind in ("bad", "good"):
        print(f"Layout '{kind}': " + ("time and user name at the top of the system prompt" if kind == "bad"
                                     else "fixed instructions first, time, user and question last"))
        first = None
        for user, minute, q in calls:
            blocks = layout(kind, user, minute, q)
            digest = hashlib.sha256(blocks[0].encode("utf-8")).hexdigest()[:12]
            if first is None:
                same = "first call: cache write"
            else:
                same = "same prefix: cache read" if digest == first else "prefix changed: cache write"
            first = first or digest
            print(f"  call by {user:<5} first block {digest}  {same}")
        print()
    print("Only the 'good' layout lets calls 2 and 3 read the cached prefix. Real providers also need\n"
          "a minimum prefix length (from hundreds to thousands of tokens, by model) before they cache.")


def main() -> None:
    parser = argparse.ArgumentParser(description="Response, semantic and embedding caching, offline.")
    parser.add_argument("--cache", choices=["none", "exact", "semantic"], default="none")
    parser.add_argument("--threshold", type=float, default=0.9, help="semantic similarity needed for a hit (0 to 1)")
    parser.add_argument("--ttl", type=int, default=0, help="minutes an answer may be reused; 0 = no limit")
    parser.add_argument("--no-id-check", action="store_true", help="semantic hits need not match order numbers or codes")
    parser.add_argument("--no-invalidate", action="store_true", help="ignore SalesOrder.Changed events")
    parser.add_argument("--shared", action="store_true", help="one cache for every user, no authorization scope")
    parser.add_argument("--log", action="store_true", help="print one line per request")
    parser.add_argument("--prompt-layout", action="store_true", help="show which prompt layouts can use prompt caching")
    args = parser.parse_args()
    if not 0 < args.threshold <= 1:
        sys.exit("--threshold must be above 0 and at most 1.")
    if args.ttl < 0:
        sys.exit("--ttl must be 0 or more minutes.")
    if args.prompt_layout:
        prompt_layout()
    else:
        replay(args)


if __name__ == "__main__":
    main()
  1. Check that the file is in the right place:

    • Windows (PowerShell):

      dir unit10\llm_cache.py
    • macOS / Linux:

      ls unit10/llm_cache.py

    You should see the file name. An error means the file is in another folder or has another name.

Step 3: Run the baseline without a cache

  1. Run:

    python unit10/llm_cache.py
  2. You should see this:

    Cache: none
    
    requests               23
    model calls            23
    cache hits              0   (exact 0, semantic 0)
    hit rate               0%
    wrong answers           0   (stale 0, leaked 0, wrong order or code 0)
    
    model cost          0.11132 USD
    average wait         4.00 s per request
    
    Rates and timings are EXAMPLES. 'wrong answers' compares each answer with a fresh one.
  3. Every question goes to the model: 23 calls, four seconds each, no wrong answers. This is the line every cache is measured against.

Step 4: Add an exact response cache

  1. Run:

    python unit10/llm_cache.py --cache exact
  2. You should see:

    Cache: exact
    
    requests               23
    model calls            16
    cache hits              7   (exact 7, semantic 0)
    hit rate              30%
    wrong answers           0   (stale 0, leaked 0, wrong order or code 0)
    
    model cost          0.07744 USD
    average wait         2.80 s per request
  3. Seven questions were exact repeats once lower case, spaces and the question mark were normalized, so seven model calls were saved. "why is order 9000123 blocked" without a question mark matched "Why is order 9000123 blocked?".

  4. There are still no wrong answers. The key holds the clerk's authorization scope, and the order change at minute 20 dropped the stored answers about order 9000123.

Step 5: Add a semantic cache

  1. Run with one line per request:

    python unit10/llm_cache.py --cache semantic --log
  2. Look for the two semantic lines in the log:

      9  ben   semantic 0.91  ok      Explain delivery block code 01
     29  anna  semantic 0.91  ok      Explain delivery block code 05

    "Explain delivery block code 01" matched "What is block reason 01?" with a similarity of 0.91, above the 0.9 threshold. The block codes matched exactly, so the answer was right.

  3. The summary at the end:

    Cache: semantic | threshold 0.9
    
    requests               23
    model calls            14
    cache hits              9   (exact 7, semantic 2)
    hit rate              39%
    wrong answers           0   (stale 0, leaked 0, wrong order or code 0)
    
    model cost          0.06776 USD
    embedding calls        11   (embedding cache hits 18; cost 0.000001 USD)
    average wait         2.45 s per request
  4. The embedding cache line matters too. Each question is embedded when it is looked up and again when it is stored; the embedding cache served 18 of those 29 requests, so only 11 reached the embedding model.

Step 6: Break it: a wrong semantic match

  1. Turn off the identifier filter, so similar wording is enough even when the order number or code differs:

    python unit10/llm_cache.py --cache semantic --no-id-check --log
  2. Look at the log:

     11  anna  semantic 0.96  WRONG   Why is order 9000124 blocked?
     16  anna  semantic 0.98  WRONG   What does delivery block 05 mean?

    Anna asked about order 9000124 and got the stored answer about 9000123: similarity 0.96. She asked about block 05 and got the answer for block 01: similarity 0.98. The wrong matches scored higher than the right ones in Step 5.

  3. The summary:

    cache hits             16   (exact 4, semantic 12)
    hit rate              70%
    wrong answers          11   (stale 0, leaked 0, wrong order or code 11)
  4. What this means for you: the best hit rate came with the most wrong answers. Raising the threshold doesn't fix it here, because the wrong pairs are the most similar ones. Matching identifiers exactly does.

Step 7: Break it: a stale answer

  1. Ignore the change event from SAP:

    python unit10/llm_cache.py --cache exact --no-invalidate --log
  2. Look at the lines after the event:

     20  EVENT Changed for 9000123: ignored
     21  anna  exact          STALE   Why is order 9000123 blocked?
     25  ben   exact          STALE   Why is order 9000123 blocked?

    At minute 20 the delivery block on order 9000123 was removed. Anna and Ben were still told it is blocked, from the answer stored at minute 1.

  3. Now add a 10-minute time to live as a backstop, still ignoring the event:

    python unit10/llm_cache.py --cache exact --no-invalidate --ttl 10
    cache hits              5   (exact 5, semantic 0)
    hit rate              22%
    wrong answers           0   (stale 0, leaked 0, wrong order or code 0)

    The stale answers are gone, but the hit rate fell from 30% to 22%, because good answers expire too. In this replay the old answer was already more than 10 minutes old when it would have been wrong; a change closer to the stored time would still slip through. Events keep the hit rate and fix staleness; a TTL catches what events miss.

Step 8: Break it: a leaked answer

  1. Use one cache for everyone, without the authorization scope in the key:

    python unit10/llm_cache.py --cache exact --shared --log
  2. Find this line:

     15  chen  exact          LEAKED  Why is order 9000123 blocked?

    Chen works for sales organization 1710 and may not see order 9000123. Without a cache, the assistant reads SAP as Chen and refuses. With a shared cache, the stored answer written for Anna is returned before SAP is ever asked, including the customer's name and the order value.

  3. The summary shows the trap: a higher hit rate (48%) and one leak.

    cache hits             11   (exact 11, semantic 0)
    hit rate              48%
    wrong answers           1   (stale 0, leaked 1, wrong order or code 0)

Step 9: See which prompt layouts can use the provider's prompt cache

  1. Run:

    python unit10/llm_cache.py --prompt-layout
  2. You should see:

    A provider's prompt cache reuses the longest identical START of a request.
    
    Layout 'bad': time and user name at the top of the system prompt
      call by anna  first block bbc1f5b6e41a  first call: cache write
      call by ben   first block b4d9e1f174f0  prefix changed: cache write
      call by anna  first block 11c16f510281  prefix changed: cache write
    
    Layout 'good': fixed instructions first, time, user and question last
      call by anna  first block 60c0b9876255  first call: cache write
      call by ben   first block 60c0b9876255  same prefix: cache read
      call by anna  first block 60c0b9876255  same prefix: cache read
    
    Only the 'good' layout lets calls 2 and 3 read the cached prefix. Real providers also need
    a minimum prefix length (from hundreds to thousands of tokens, by model) before they cache.
  3. The 12-character codes are fingerprints (hashes) of the first block. One changed character changes the fingerprint. In the bad layout, a timestamp at the top makes every call a new prefix, so every call pays to write the cache and none reads it.

Step 10: Save your work

  1. Save the script with Git:

    git add unit10/llm_cache.py
    git commit -m "Unit 10: response, semantic and embedding caching with failure cases"

How the code works

Part What it does
Rates and timings Made-up prices and seconds per model call and per cache lookup, reused from the token economics topic
ORDERS, USERS, TRAFFIC The "SAP system", each clerk's sales organizations, and the morning's questions with one SalesOrder.Changed event
fresh_answer The assistant without a cache: checks the clerk's authorization, then answers from current data. Used as the truth for every request
embed, cosine A toy embedding: counts words after removing filler words and merging synonyms, with numbers weighted low as in real models
EmbeddingCache Stores vectors under a hash of model name plus text, and counts hits and misses
normalize, identifiers, scope_of The parts of the key: tidied question, exact numbers in it, and the clerk's authorization scope
ResponseCache.lookup Exact lookup first; then, for --cache semantic, the most similar stored question in the same scope with the same identifiers, above the threshold
ResponseCache.store Stores answers, but never refusals
ResponseCache.invalidate Drops every entry about an order when its change event arrives
classify Compares the answer given with the fresh answer: LEAKED if it shows an order outside the clerk's scope, STALE if it is the right order before a change, otherwise WRONG
prompt_layout Fingerprints the first block of two prompt layouts to show which one keeps a stable prefix

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't installed, or the terminal can't find it Windows: repeat Unit 1, Step 1, then open a new terminal. macOS/Linux: use python3 until .venv is active
can't open file ... llm_cache.py You are not in the course folder, or the file has another name Run cd to orchestrate-course; check the file is in unit10
SyntaxError or IndentationError Part of the code was not pasted, or the indentation changed Select all in the file, delete, and paste the whole block again
ModuleNotFoundError A line was changed to import a library this script doesn't use Paste the code again; the script needs only built-in Python
error: unrecognized arguments An option is misspelled Check the spelling; python unit10/llm_cache.py --help lists all options
--threshold must be above 0 and at most 1. The threshold is outside the valid range Use a value such as 0.85 or 0.95
Your numbers differ from the ones shown You changed the traffic, the rates or the threshold Expected after edits. Paste the code again to get the published numbers
A proxy or network error Not possible here: the script makes no network calls If you see one, you are running a different file

The SAP way

As of October 2026, from the SAP sources opened for this topic.

Prompt caching in the orchestration service

SAP's Python SDK, sap-ai-sdk-gen, gained a CacheControl class in its Orchestration V2 module through a pull request merged on 18 September 2026. In version 7.4.1, read for this topic, CacheControl takes type="ephemeral" and an optional ttl of "5m" (the default) or "1h". It attaches to a text or image part of a message, or to a tool definition. The usage object reports prompt_tokens_details.cached_tokens (read from the cache) and cache_creation_tokens (written to it).

A cached system message looks like this. It is a sketch: it needs an SAP AI Core account with the generative AI hub, set up in Set up for Unit 5, and a model that supports explicit caching.

# Sketch: needs SAP AI Core with the generative AI hub (Unit 5) and a model that supports caching.
from gen_ai_hub.orchestration_v2 import CacheControl, SystemMessage, TextPart, Template, UserMessage

SYSTEM_PROMPT = "You help order-to-cash clerks understand blocked sales orders ..."  # long and fixed

template = Template(template=[
    SystemMessage(content=[TextPart(text=SYSTEM_PROMPT,
                                    cache_control=CacheControl(type="ephemeral", ttl="5m"))]),
    UserMessage(content="Order data: {{?order}}\n\nQuestion: {{?question}}"),  # changes every call
])

Token economics lists which model families SAP documents for explicit and implicit caching, and how cached tokens appear in usage.

Response and semantic caches on SAP BTP

No managed response or semantic cache was found in the SAP AI Core sources opened for this topic, so you build one in your app on SAP BTP:

  • Store. SAP HANA Cloud can hold the cache table. Its vector engine has a REAL_VECTOR column type and a COSINE_SIMILARITY function, so a semantic lookup is a SELECT TOP 1 ... ORDER BY COSINE_SIMILARITY(...) DESC with WHERE clauses for scope and identifiers. SAP HANA Cloud vector engine covers the setup.
  • Community package. A community member published langchain-hana-cache (v0.1.0), a semantic cache for LangChain on HANA Cloud that stores the prompt hash, text, vector, model identifier and response in one table, with a similarity threshold and a TTL. It is not an SAP product; review it like any third-party code, and check that you can add a scope filter.
  • Embeddings. The same table design works for an embedding cache: key on the model name and a hash of the text.

Invalidation from S/4HANA

  • Business events. S/4HANA can publish sap.s4.beh.salesorder.v1.SalesOrder.Created.v1 and sap.s4.beh.salesorder.v1.SalesOrder.Changed.v1 through SAP Event Mesh, as SAP's learning material describes for sales order replication. Your app subscribes and drops entries for that order. Deletion is not covered in that scenario, so keep a TTL.
  • ETags. S/4HANA Cloud OData APIs return an ETag that changes with the document. Reading it and putting it in the key ties each answer to one version of the order, at the cost of a small SAP read per request.

Setting up Event Mesh is beyond this topic; here the events and ETags matter as the freshness signal for your cache.

SAP-delivered AI

Joule and other AI features inside SAP applications manage their own caching. Nothing in this topic needs configuring for them.

Build vs. SAP

Need Use Why
Long fixed instructions or many tools, steady traffic Prompt caching in the orchestration service (CacheControl) Cheaper, faster input; answers stay fresh
Repeated general questions ("what does block 01 mean?") Your own exact response cache, keyed on scope, prompt version and model Skips the model entirely; low risk when general
Repeated order-specific questions Exact cache plus invalidation by S/4HANA events, with a TTL backstop Order data changes; events keep answers current
Paraphrased questions at volume A semantic cache in HANA Cloud, with scope and identifier filters, after an evaluation Higher hit rate, but only with proof that matches are right
Re-embedding the same documents or questions An embedding cache keyed on model and text Never stale; saves embedding calls
Joule skills and agents Nothing to build SAP manages them

Production concerns

  • Authorizations. A hit returns data without reading SAP, so SAP's checks never run. Put the authorization scope in the key, check access before lookup for sensitive objects, and have the security review sign off the key design.
  • Evaluation. Track hit rate and wrong-hit rate together. Sample cached answers daily, recompute them fresh, and compare, as classify does. Re-run your evaluation set whenever the threshold, embedding model or prompt changes. Building an evaluation harness has the tooling.
  • Observability. Log on every request whether it was an exact hit, semantic hit (with score) or miss, plus prompt-cache read and write tokens. Observability for AI systems shows where these attributes go in a trace.
  • Cost. Count embedding calls, cache storage and the extra SAP reads for version checks. A prompt cache that expires before reuse costs more than none on providers that charge for writes.
  • Data protection. Cached answers with names and amounts are personal and business data at rest. Set retention, encrypt, and include the cache in deletion requests.
  • Operations. Have a purge button per prompt version, per object and for everything. Bump the prompt version in the key on every release, so old answers are never served by a new prompt.
  • Poisoning. Store only answers that passed your output checks; never cache answers built from content that failed a prompt-injection check.
  • Clean core. The cache lives in your side-by-side app on BTP. S/4HANA only publishes the events it already offers.

Pitfalls

  • Leaving the authorization scope out of the key. One cache for everyone turns a cache hit into a way around SAP authorizations.
  • Trusting similarity for identifiers. Order numbers, codes and amounts must match exactly; embeddings blur them.
  • Judging by hit rate alone. The highest hit rate in this topic came with the most wrong answers.
  • TTL as the only freshness rule. Answers are wrong from the moment data changes until expiry.
  • Caching refusals or errors. A stored "not authorized" outlives a granted authorization.
  • Dynamic content at the top of the prompt. A timestamp or user name first means no prompt-cache reads.
  • Forgetting the prompt version. A fixed prompt keeps serving answers written by the broken one.
  • Changing the embedding model without clearing the cache. Vectors from two models can't be compared.

Exercise

Write the caching policy for the blocked-orders assistant and test it on new traffic. The AI business case and the BTP production topics later in this unit use it.

  1. Open unit10/llm_cache.py and add six questions at the end of TRAFFIC: two repeats of earlier general questions in new words, two about order 9000130 from Chen, and two from Anna about an order she may not see.
  2. Add a second event, for example (40, "EVENT", "sap.s4.beh.salesorder.v1.SalesOrder.Changed.v1 9000124 DeliveryBlockReason="), and put at least one question about 9000124 after it.
  3. Run the script with --cache exact, --cache semantic and --cache semantic --no-id-check. Note the hit rate and wrong answers for each.
  4. Run --cache semantic with --threshold 0.85 and --threshold 0.95. Note what changes.
  5. Create unit10/cache-policy.md with:
    • a table: run, hit rate, wrong answers, model cost;
    • which caches you would switch on, and the threshold you chose;
    • the cache key, part by part, with one line on why each part is there;
    • the invalidation rules: which events, which TTL, and what happens on deletion;
    • two lines for the security review: who can see a cached answer, and how to purge it.
  6. Commit cache-policy.md and your changed script.

Done when cache-policy.md shows the results of at least five runs, names a cache key with the authorization scope in it, and gives an invalidation rule with both an event and a TTL, and your chosen setting shows zero wrong answers on your extended traffic.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why is prompt caching safe to switch on where a response cache needs a design review?

    Answer: B. A provider's prompt cache keeps the processed prefix, not the answer. The model still reads the current data and writes a new answer, so the stale and leaked failures of a response cache don't apply to what the user sees.
  2. 2In Step 9, why does the "bad" layout never read the prompt cache?

    Answer: C. Prompt caches match the exact start of the request. Changing content at the top gives every call a new prefix, so each call writes the cache and none reads it. Put fixed content first and changing content last.
  3. 3In Step 8, Chen received the answer about order 9000123. Why didn't SAP stop it?

    Answer: D. Authorization is checked when the app reads SAP as the user. A shared cache returns a stored answer without that read, so the check never runs. The authorization scope must be in the key.
  4. 4In Step 6, the wrong semantic matches scored 0.96 and 0.98, higher than the right ones. What fixes this?

    Answer: B. The wrong pairs differ only in an identifier, which embeddings weigh lightly, so they look most alike. A stricter threshold would block the right matches first. Matching identifiers exactly lets the embedding judge only the wording.
  5. 5Your app listens to SalesOrder.Changed events to drop cached answers. Why keep a TTL as well?

    Answer: A. SAP's learning material lists created and changed events for sales orders, not deletions, and any event can arrive late or not at all. A TTL bounds how long a missed change can stay wrong, at the cost of some hit rate.
  6. 6The team switches the embedding model used by the semantic cache. What must happen to the cache?

    Answer: C. Vectors from two embedding models can't be compared, so old entries would match badly. Keying on model name plus text, as RedisVL's EmbeddingsCache does, or emptying the cache, avoids mixing them.
  7. 7Which response from SAP's Python SDK shows that the provider read the prefix from its cache?

    Answer: D. In sap-ai-sdk-gen 7.4.1, cached_tokens counts tokens read from the prompt cache and cache_creation_tokens counts tokens written to it. A write means the cache was filled, not reused.
  8. 8A product owner wants the semantic cache because it reached a 70% hit rate in testing. What do you do first?

    Answer: C. In this topic's replay, the 70% hit rate came with 11 wrong answers out of 23. A cache is only worth its hit rate if the hits are correct, and only an evaluation set of real questions shows that.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in