How retrieval-augmented generation works as eight stages from documents to cited answer, when it beats fine-tuning, and a minimal RAG you build and run.
A large language model answers from what it learned in training. It has never read your company's credit policy, your tolerance rules for supplier invoices or last week's change to the payment run. Ask it anyway and it will often produce a fluent, confident and wrong answer.
Retrieval-augmented generation (RAG) fixes this with a lookup before the answer. When someone asks a question, the system first retrieves the few passages from your own documents that best match it. It augments the prompt by pasting those passages in. Then the model generates an answer from them, and points to the passages it used.
Think of an open-book exam. The model is the student; RAG is the book and the rule "answer only from the book, and cite the page." The quality of the answer depends heavily on whether the right page was found.
The idea comes from a 2020 research paper by Lewis and colleagues. It paired a language model with a searchable index of Wikipedia passages. The authors noted that showing where answers come from, and updating what a model knows, were open problems for models without such an index.
Most questions employees ask an assistant are about your rules, not general knowledge. RAG is how an assistant answers them.
Value. A sales rep asks "who can release this credit-blocked order and what do I tell the customer?". The assistant finds the order-to-cash guide and answers in two sentences, with a link. No search through a shared drive, no ticket to the credit team.
Fresh knowledge without retraining. When the invoice tolerance changes, you update one document and re-index it. The next answer uses the new rule. Changing what a fine-tuned model "knows" means training it again.
Traceability. Each answer cites its sources. An auditor, a manager or the user can check the passage behind a claim. That is the basis for trust and for fixing wrong answers.
Risk. RAG reduces made-up answers; it doesn't remove them. If retrieval finds the wrong passage, the model can answer the wrong question convincingly. If the document store holds files a user shouldn't see, RAG can leak them. Both risks are design choices, covered below.
Cost. Each answer costs one search plus one model call with a longer prompt. Indexing costs a one-time embedding pass per document, repeated only when documents change.
SAP's Architecture Center describes RAG in three steps: encode the question, retrieve similar chunks with the SAP HANA Cloud vector engine, and generate the response with a model. It offers two routes: use SAP's managed RAG services, or build your own on SAP BTP. SAP recommends the managed services where they fit.
As of October 2026, the main SAP offerings are:
Offering
What it does
Who it's for
Grounding in the generative AI hub (SAP AI Core, orchestration service)
A grounding module that ingests documents, chunks and embeds them into the SAP HANA vector engine, and adds the best chunks to the prompt. Added to orchestration in November 2024, per SAP's release notes.
Teams building their own AI apps on SAP BTP
Document grounding in Joule
Lets Joule answer from business documents held in SAP and third-party repositories. SAP lists it as an AI feature with a flat fee.
Companies that use Joule and want it to know their policies
SAP HANA Cloud vector engine
Stores embeddings in database tables and searches them with SQL
Custom RAG next to business data (Unit 7 has a topic on it)
The grounding module reads documents from Microsoft SharePoint, Amazon S3 and SFTP, or takes text you send it directly. SAP has since added SAP Document Management (2025) and ServiceNow (2026) as repository types. It can also search help.sap.com for SAP product questions.
Every RAG system, bought or built, has the same eight stages. Two run ahead of time, when documents arrive. The rest run for every question.
Stage
When
Plain words
What goes wrong if it's weak
1. Ingest
Ahead of time
Collect the documents and their owners
Outdated or duplicate files give outdated answers
2. Chunk
Ahead of time
Cut each document into passages of a few paragraphs
Passages too big or too small hide the answer
3. Embed
Ahead of time and per question
Turn each passage into numbers that capture meaning
Model misses your terms (e.g. "credit hold" vs. "credit block")
4. Index
Ahead of time
Store passages and numbers in a vector store
Slow search, no way to filter by company code
5. Retrieve
Per question
Find the passages closest in meaning to the question
Wrong passage, and everything after it is wrong
6. Augment
Per question
Paste the passages into the prompt, numbered
Too many passages bury the right one
7. Generate
Per question
The model answers using only those passages
Model adds outside "knowledge"
8. Cite
Per question
Show which passage supports each claim
Users can't check, and nobody can debug
The rest of Unit 7 deepens one stage at a time: chunking, vector databases, hybrid search and filters, re-ranking, SAP HANA Cloud, and grounding with authorizations.
Live numbers: "how many orders are blocked right now?"
A tool call to an SAP API
The answer is in a system, not a document
A fixed style, format or narrow skill
Fine-tuning (or a better prompt first)
It's about behavior, not facts
A short, stable rule set
Put it in the prompt
Nothing to retrieve
SAP Learning notes that RAG is often more economical than fine-tuning because you update a knowledge base instead of retraining. Many real assistants combine RAG for documents with tool calls for live SAP data.
"RAG trains the model on our documents." It doesn't. The model is unchanged. Your documents are searched at question time and pasted into the prompt.
"With RAG, the assistant can't make things up." It can. A wrong passage or a model that ignores the instruction still produces wrong answers. Citations make errors visible; they don't prevent them.
"Put all our documents in and it will work." Old, duplicate and contradictory documents give contradictory answers. Curating the source documents is part of the project.
"RAG is a product you buy." It's an architecture. SAP's grounding module and Joule's document grounding package it, but the eight stages and their choices remain.
"A bigger model fixes poor answers." If the right passage isn't retrieved, no model can answer from it. Most RAG quality problems are retrieval problems.
Pick one answer for each question. The explanation appears after you choose.
1What does RAG change about the language model itself?
Answer: B. The model stays as it is. RAG searches your documents for each question and pastes the best passages into the prompt, so updating knowledge means updating documents, not retraining.
2The invoice tolerance policy changed yesterday. What has to happen for a RAG assistant to answer with the new rule?
Answer: C. In RAG, knowledge lives in the indexed documents. Re-indexing the changed document is enough; the next question retrieves the new rule.
3A user asks "how many sales orders are credit-blocked right now?" Which approach fits best?
Answer: D. The answer is a live number in a system, not text in a document. RAG fits policies and guides; live data needs a call to the system that holds it.
4A pilot gives confident wrong answers. Where should the team look first?
Answer: A. If the right passage isn't retrieved, no model can answer from it. Most RAG quality problems are retrieval problems, so measure retrieval before changing models.
5Why should every RAG answer show its sources?
Answer: B. Citations make errors visible and checkable. They don't prevent errors, but they let users verify and let the team find which passage caused a bad answer.
6What is the main security risk to raise before indexing a shared drive?
Answer: D. RAG pastes whatever it retrieves into the prompt. If the store holds files a user can't open today, the assistant can leak them unless access rules are enforced at retrieval.
7Which SAP offering lets a team building its own app on SAP BTP add document retrieval to model calls?
Answer: C. The orchestration service's grounding module ingests, chunks, embeds and retrieves documents for your own app. Joule's document grounding serves Joule users rather than custom apps.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
RAG is a search engine whose results are read by a model instead of a person.
Everything else follows from that. A search engine is judged by whether the right result is in the top few. So is RAG. A person skims results and ignores the irrelevant ones; a model tends to use whatever you hand it. So RAG needs a stricter filter than a search page, and an instruction that says "only these, and cite them."
The 2020 paper by Lewis and colleagues framed it as two kinds of memory. Parametric memory is what the model learned in its weights. Non-parametric memory is an index you can read, change and audit. RAG answers from the second, using the first only for language.
That gives you the one rule for debugging: check retrieval before generation. If the right chunk is in the top few, the problem is in the prompt or model. If it isn't, nothing downstream can fix it.
The pipeline has an indexing path that runs when documents change and a query path that runs for every question. They share one thing: the embedding model. Questions and chunks must be embedded by the same model, or their numbers aren't comparable.
flowchart LR
subgraph Indexing["Indexing (when documents change)"]
D[Documents] --> I[1 Ingest] --> C[2 Chunk] --> E1[3 Embed] --> X[(4 Index)]
end
subgraph Query["Query (every question)"]
Q[Question] --> E2[3 Embed] --> R[5 Retrieve]
X --> R
R --> A[6 Augment] --> G[7 Generate] --> CI[8 Cite]
end
Read documents from where they live and keep metadata with them: file name, owner, process, company code, language, last change date. Metadata is what later lets you filter, cite and enforce access. Ingesting also means deciding what not to load: drafts, superseded versions and duplicates.
A model's prompt has room for a few passages, not whole manuals. You cut documents into chunks. Too large, and one chunk mixes topics and dilutes the match. Too small, and a rule loses its context ("the tolerance is 2 percent" of what?).
A good starting point is structure-aware chunking: cut at headings, and prefix each chunk with its document title and section name. The chunk then still makes sense on its own. Chunk sizes, overlap, tables and PDFs are the next Unit 7 topic.
An embedding model turns each chunk into a vector, as in Embeddings and semantic similarity. The question is embedded the same way at query time. Change the embedding model and you must re-embed every chunk.
A vector store saves chunks, vectors and metadata, and finds the nearest vectors quickly. You installed Chroma and set up SAP HANA Cloud in Set up for Unit 7. How the index finds neighbors fast is covered later in Unit 7.
Embed the question, ask for the top k nearest chunks, optionally with a metadata filter (only procure-to-pay guides). Two settings matter from day one:
k: how many chunks. Too few misses the answer; too many adds noise and cost.
A minimum similarity: below it, treat the chunk as irrelevant. If nothing passes, the honest answer is "I can't find this", and you don't need to call the model at all.
Build the prompt: a system instruction, then the chunks numbered[1], [2], [3], then the question. Numbering is what makes citations possible. This is context engineering from Prompt and context engineering, with the context chosen by search.
The model answers. The system instruction does three jobs: use only the sources, cite them by number, and say a fixed phrase when they don't contain the answer. A fixed phrase is easy to detect in code and in tests.
Show the user which source each number points to, and check the citations in code. An answer with no citations, or one citing [7] when you sent three sources, is a warning sign you can catch automatically.
"Always answer in this format", narrow classification
Prompt first, then fine-tuning
Classify a ticket into one of eight queues
Small, stable rules
Put them in the prompt
Five escalation rules
#Build it yourself: a minimal RAG over SAP process guides
Before you start: complete Set up your computer for this course and Set up for Unit 7. They create your orchestrate-course folder and .venv, and install Chroma and Sentence Transformers. For the optional real model call, you also need the AICORE_ lines in .env and sap-ai-sdk-gen from Set up for Unit 5. This walkthrough doesn't repeat those steps.
You will build one Python script that runs all eight stages over eight short, made-up company guides about blocked sales orders, blocked invoices and MRP exceptions. Each stage is one function, so later Unit 7 topics can replace one stage at a time. You will also measure retrieval with six test questions, the seed of your evaluation set for Unit 8.
flowchart LR
F[rag_docs/<br/>8 guides] --> S[rag_minimal.py<br/>chunk, embed]
S --> V[(rag_store/<br/>Chroma)]
Q[Your question] --> S
V --> P[Numbered prompt]
P --> M{--llm?}
M -->|no| O[Show prompt]
M -->|sample| A[Made-up answer]
M -->|MODEL| L[SAP orchestration] --> A2[Answer + citations]
Your course folder with the Unit 7 setup done. Cost: free.
Optional: the AICORE_ lines in .env (Unit 5) for a real model call. Cost: a small per-request charge on your SAP AI Core account, unless your trial or plan covers it.
About 60 minutes.
#Step 1: Open your course folder and turn on the virtual environment
Open a terminal (on Windows, PowerShell; on macOS, Terminal).
Go to your course folder and turn on .venv.
Windows (PowerShell):
cd $HOME\orchestrate-course
.\.venv\Scripts\Activate.ps1
macOS/Linux:
cd ~/orchestrate-course
source .venv/bin/activate
Check that the libraries are there:
pip show chromadb sentence-transformers
You should see a Name: and Version: block for each. If one is missing, run pip install -r requirements.txt. No new libraries are needed for this topic.
In VS Code, right-click the unit07 folder, choose New File, name it rag_minimal.py, paste the code below and save.
"""A minimal RAG pipeline over made-up SAP process guides: ingest, chunk, embed, index,
retrieve, augment, generate, cite. Each stage is one function, so later Unit 7 topics can
swap one stage at a time.
Run from your course folder (orchestrate-course), with .venv turned on:
python unit07/rag_minimal.py "Why is the invoice blocked for payment?" # retrieve + show prompt
python unit07/rag_minimal.py "Why is the invoice blocked for payment?" --llm sample # made-up answer, no account
python unit07/rag_minimal.py "Why is the invoice blocked for payment?" --llm MODEL # real model via SAP orchestration
python unit07/rag_minimal.py "What does reschedule in mean?" --process plan-to-produce
python unit07/rag_minimal.py --eval # did retrieval find the right guide?
Add --offline to any command to use toy embeddings instead of the Unit 3 model.
"""
import argparse
import hashlib
import math
import os
import re
import shutil
import sys
from pathlib import Path
HERE = Path(__file__).resolve().parent
DOCS = HERE / "rag_docs" # your documents: .md or .txt files, one per guide
STORE = HERE / "rag_store" # the Chroma index; rebuildable, keep it out of Git
MODEL = "sentence-transformers/all-MiniLM-L6-v2" # the Unit 3 model; 384 numbers per text
MAX_CHARS = 700 # longest chunk we keep before splitting by paragraph
# Made-up company guides, shaped like the course's running examples. Written to rag_docs/ on first run.
SAMPLE_DOCS = {
"o2c-credit-blocks.md": """process: order-to-cash
# Guide: sales orders blocked by credit checks
## When an order is blocked
An order is held for delivery when the customer's open receivables plus the new order exceed the credit limit agreed with credit management. The order stays saved; only the delivery is stopped.
## Who releases it
Only the credit management team may release a credit block. Sales may not release it, even for a key account. A release is recorded with the name of the person and the reason.
## What sales should tell the customer
Tell the customer the order is confirmed but waiting for a credit review. Do not promise a delivery date until the block is released. Typical review time is one working day.
""",
"p2p-invoice-blocks.md": """process: procure-to-pay
# Guide: supplier invoices blocked for payment
## Quantity differences
An invoice is blocked for payment when the invoiced quantity is higher than the quantity posted in the goods receipt. Accounts payable waits for the missing goods receipt or a credit memo from the supplier.
## Price differences
An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance. In this company the tolerance is 2 percent or 50 euros per line, whichever is lower.
## Three-way match
Before payment, the purchase order, the goods receipt and the invoice are compared. Only invoices that pass this three-way match are paid in the next payment run.
""",
"p2p-payment-runs.md": """process: procure-to-pay
# Guide: payment runs
## Schedule
The payment run takes place every Tuesday and Friday. Invoices released by 12:00 the day before are included.
## Urgent payments
An urgent payment outside the run needs approval from the head of accounts payable and a second signature from treasury.
""",
"p2p-new-suppliers.md": """process: procure-to-pay
# Guide: setting up a new supplier
## Required documents
A new supplier needs a signed supplier code of conduct, bank details on company letterhead and a tax registration certificate.
## Duplicate check
Before creating a supplier, purchasing searches for existing suppliers with the same tax number or bank account to avoid duplicates.
""",
"o2c-pricing-blocks.md": """process: order-to-cash
# Guide: orders blocked for pricing
## Missing prices
An order with a line that has no valid price is blocked for billing. Sales must maintain the price or apply an approved manual price before the order can be billed.
## Manual discounts
A manual discount above 10 percent needs approval from the sales manager. Until it is approved, the order is blocked for delivery.
""",
"p2p-goods-receipt.md": """process: procure-to-pay
# Guide: posting goods receipts
## Partial deliveries
Post a goods receipt only for the quantity that physically arrived. The open quantity stays on the purchase order until the rest is delivered.
## Damaged goods
Damaged goods are posted to blocked stock and the supplier is informed the same day. They are never posted as unrestricted stock.
""",
"ptp-mrp-exceptions.md": """process: plan-to-produce
# Guide: MRP exception messages
## Reschedule in
A reschedule-in exception means the planning run suggests moving a planned order or purchase order earlier, because demand appears before the receipt.
## Reschedule out
A reschedule-out exception means the receipt arrives earlier than needed. Planners may move it later to reduce stock.
## How planners work the list
Planners review exceptions daily, starting with materials that block production in the next five working days.
""",
"ptp-safety-stock.md": """process: plan-to-produce
# Guide: safety stock
## Why we keep it
Safety stock covers unexpected demand and late supplier deliveries. The planning run does not plan to use it for normal demand.
## Who changes it
Safety stock levels are reviewed every quarter by the materials planner and approved by the supply chain manager.
""",
}
# Questions with the guide that should answer them. A first, tiny evaluation set (Unit 8 grows it).
EVAL = [
("Who is allowed to release an order that failed the credit check?", "o2c-credit-blocks.md"),
("The invoice quantity is more than what we received. What happens?", "p2p-invoice-blocks.md"),
("What does it mean when MRP says to move an order earlier?", "ptp-mrp-exceptions.md"),
("How much can the invoice price differ from the PO price?", "p2p-invoice-blocks.md"),
("What should I tell a customer whose order is on credit hold?", "o2c-credit-blocks.md"),
("On which days are suppliers paid?", "p2p-payment-runs.md"),
]
SYSTEM = ("You answer questions from employees using only the numbered sources provided. "
"Cite every statement with the source number in square brackets, like [1]. "
"If the sources do not contain the answer, say: I can't find this in the guides. "
"Do not use outside knowledge.")
# ---------- 1. ingest ----------
def ingest() -> list:
"""Read every .md and .txt file in rag_docs/. Writes the sample guides on the first run."""
if not DOCS.exists():
DOCS.mkdir(parents=True)
for name, text in SAMPLE_DOCS.items():
(DOCS / name).write_text(text, encoding="utf-8")
print(f"Wrote {len(SAMPLE_DOCS)} sample guides to {DOCS.name}/")
docs = []
for path in sorted(DOCS.glob("*")):
if path.suffix.lower() in (".md", ".txt"):
docs.append((path.name, path.read_text(encoding="utf-8")))
if not docs:
sys.exit(f"No .md or .txt files in {DOCS}. Delete the empty folder to get the samples back.")
return docs
# ---------- 2. chunk ----------
def chunk(name: str, text: str) -> list:
"""Split one guide into chunks at '## ' headings. Each chunk starts with the guide title and
section name, so it still makes sense on its own. Long sections are split by paragraph."""
process = "unknown"
match = re.match(r"process:\s*(\S+)", text)
if match:
process = match.group(1)
text = text[match.end():]
title_match = re.search(r"^# (.+)$", text, flags=re.M)
title = title_match.group(1).strip() if title_match else name
chunks = []
for part in re.split(r"^## ", text, flags=re.M)[1:]:
heading, _, body = part.partition("\n")
pieces, current = [], ""
for paragraph in [p.strip() for p in body.split("\n\n") if p.strip()]:
if current and len(current) + len(paragraph) > MAX_CHARS:
pieces.append(current)
current = ""
current = (current + "\n\n" + paragraph).strip()
if current:
pieces.append(current)
for piece in pieces:
chunks.append({"id": f"{name}#{len(chunks) + 1}",
"text": f"{title} > {heading.strip()}\n{piece}",
"meta": {"source": name, "section": heading.strip(), "process": process}})
return chunks
# ---------- 3. embed ----------
STOP = set("a an and are as at be by can do does for from has how i in is it it's may of on or our "
"should so that the their this to was we what when where which who why will with".split())
def toy_embed(text: str, size: int = 256) -> list:
"""A tiny stand-in for a real model: hash each word into one of 256 slots, skipping common
words such as 'the' and 'what'. Matches words, not meaning."""
vector = [0.0] * size
for word in re.findall(r"[a-z]+", text.lower()):
if word in STOP:
continue
slot = int(hashlib.md5(word.encode()).hexdigest(), 16) % size
vector[slot] += 1.0
length = math.sqrt(sum(v * v for v in vector)) or 1.0
return [v / length for v in vector]
def get_embedder(offline: bool):
"""Return (name, function that turns a list of texts into a list of vectors)."""
if offline:
return "toy-hash-256", lambda texts: [toy_embed(t) for t in texts]
try:
from sentence_transformers import SentenceTransformer
except ImportError:
sys.exit("sentence-transformers is not installed. See Set up for Unit 3, or add --offline.")
print(f"Loading {MODEL} (from disk if you ran Unit 3; otherwise it downloads once)...")
try:
model = SentenceTransformer(MODEL)
except Exception as error: # usually a blocked download on a company network
sys.exit(f"Could not load the model ({type(error).__name__}). Check your network, or add --offline.")
return "minilm-384", lambda texts: model.encode(texts, normalize_embeddings=True).tolist()
# ---------- 4. index ----------
def index(chunks: list, embedder: tuple, rebuild: bool):
"""Store chunks and their embeddings in Chroma. Rebuilds when the documents changed."""
try:
import chromadb
from chromadb.config import Settings
except ImportError:
sys.exit("chromadb is not installed. Run: pip install -r requirements.txt")
name, embed = embedder
fingerprint = hashlib.sha256("".join(c["id"] + c["text"] for c in chunks).encode()).hexdigest()[:12]
marker = STORE / f"fingerprint-{name}.txt"
if rebuild or (marker.exists() and marker.read_text() != fingerprint):
shutil.rmtree(STORE, ignore_errors=True)
print("Documents changed (or --reindex): rebuilding the index")
client = chromadb.PersistentClient(path=str(STORE), settings=Settings(anonymized_telemetry=False))
collection = client.get_or_create_collection(
name=f"guides_{name}", embedding_function=None, configuration={"hnsw": {"space": "cosine"}})
if collection.count() == 0:
collection.add(ids=[c["id"] for c in chunks], documents=[c["text"] for c in chunks],
metadatas=[c["meta"] for c in chunks], embeddings=embed([c["text"] for c in chunks]))
marker.write_text(fingerprint)
print(f"Indexed {collection.count()} chunks in {STORE.name}/ (collection {collection.name})")
return collection
# ---------- 5. retrieve ----------
def retrieve(collection, embed, question: str, top: int, min_sim: float, process: str = "") -> list:
"""Return up to `top` chunks closest in meaning to the question, above a minimum similarity."""
where = {"process": process} if process else None
found = collection.query(query_embeddings=embed([question]), n_results=top, where=where)
hits = []
for cid, text, meta, dist in zip(found["ids"][0], found["documents"][0],
found["metadatas"][0], found["distances"][0]):
similarity = 1 - dist # Chroma's cosine distance is 1 minus cosine similarity
if similarity >= min_sim:
hits.append({"id": cid, "text": text, "source": meta["source"], "similarity": similarity})
return hits
# ---------- 6. augment ----------
def augment(hits: list) -> str:
"""Number the retrieved chunks so the model can cite them as [1], [2], ..."""
return "\n\n".join(f"[{n}] (from {h['id']})\n{h['text']}" for n, h in enumerate(hits, 1))
# ---------- 7. generate ----------
def generate(model: str, question: str, context: str) -> str:
"""Call a model through SAP's orchestration service, or return a made-up answer with --llm sample."""
if model == "sample":
first = context.split("\n")[2] if context.count("\n") >= 2 else context
return "[sample answer, no model called] " + first.split(". ")[0].rstrip(".") + ". [1]"
from dotenv import load_dotenv
load_dotenv()
missing = [n for n in ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL",
"AICORE_BASE_URL", "AICORE_RESOURCE_GROUP"] if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5'. Or use --llm sample.")
from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
OrchestrationService, PromptTemplatingModuleConfig,
SystemMessage, Template, UserMessage)
template = Template(template=[SystemMessage(content=SYSTEM),
UserMessage(content="Sources:\n{{?context}}\n\nQuestion: {{?question}}")])
config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=model, timeout=60, max_retries=1))))
service = None
try: # creating the service fetches a token, so it can fail too
service = OrchestrationService(config=config)
result = service.run(placeholder_values={"context": context, "question": question})
except Exception as error:
sys.exit(f"The model call failed: {type(error).__name__}: {str(error)[:300]}")
finally:
if service is not None:
service.close_http_connection()
return result.final_result.choices[0].message.content or ""
# ---------- 8. cite ----------
def check_citations(answer: str, hits: list) -> list:
"""Return warnings: no citations, or citations to sources that were never provided."""
cited = {int(n) for n in re.findall(r"\[(\d+)\]", answer)}
warnings = []
if not cited and "can't find" not in answer.lower():
warnings.append("The answer cites no source. Treat it as unverified.")
for n in sorted(cited):
if not 1 <= n <= len(hits):
warnings.append(f"The answer cites [{n}], but only {len(hits)} sources were provided.")
return warnings
# ---------- running it ----------
def run_eval(collection, embed, top: int) -> None:
"""For each EVAL question: is the expected guide among the top chunks? (hit rate at k)"""
found = 0
for question, expected in EVAL:
hits = retrieve(collection, embed, question, top, min_sim=-1.0)
sources = [h["source"] for h in hits]
ok = expected in sources
found += ok
print(f"{'ok ' if ok else 'MISS '} {question}\n expected {expected}; got {', '.join(sources)}")
print(f"\nHit rate at {top}: {found}/{len(EVAL)} questions found the right guide in the top {top}.")
def main() -> None:
parser = argparse.ArgumentParser(description="A minimal RAG pipeline over made-up SAP process guides.")
parser.add_argument("question", nargs="?", default="Who can release an order blocked by the credit check?")
parser.add_argument("--process", choices=["order-to-cash", "procure-to-pay", "plan-to-produce"],
help="only search guides for this process (a metadata filter)")
parser.add_argument("--top", type=int, default=3, help="how many chunks to retrieve (default 3)")
parser.add_argument("--min-sim", type=float, default=0.2,
help="ignore chunks below this similarity (default 0.2)")
parser.add_argument("--llm", metavar="MODEL", help="generate an answer: 'sample' (no account) or a model name")
parser.add_argument("--offline", action="store_true", help="toy embeddings; no model download")
parser.add_argument("--reindex", action="store_true", help="delete and rebuild the index")
parser.add_argument("--eval", action="store_true", help="check retrieval against the built-in questions")
args = parser.parse_args()
docs = ingest()
chunks = [c for name, text in docs for c in chunk(name, text)]
print(f"{len(docs)} guides -> {len(chunks)} chunks")
embedder = get_embedder(args.offline)
collection = index(chunks, embedder, args.reindex)
embed = embedder[1]
if args.eval:
run_eval(collection, embed, args.top)
return
hits = retrieve(collection, embed, args.question, args.top, args.min_sim, args.process or "")
print(f'\nQuestion: "{args.question}"' + (f" (only {args.process})" if args.process else ""))
if not hits:
print(f"No chunk reached similarity {args.min_sim}. Answer: I can't find this in the guides.")
print("(No model was called. Saying 'not found' is the correct result for an out-of-scope question.)")
return
print("\nRetrieved:")
for n, h in enumerate(hits, 1):
print(f" [{n}] {h['id']:<28} similarity {h['similarity']:.2f}")
context = augment(hits)
if not args.llm:
print("\nPrompt the model would receive (system instruction, then this):\n")
print(f"Sources:\n{context}\n\nQuestion: {args.question}")
print("\nNo model called. Add --llm sample (no account) or --llm MODEL to generate an answer.")
return
answer = generate(args.llm, args.question, context)
print(f"\nAnswer:\n{answer}\n\nSources:")
for n, h in enumerate(hits, 1):
print(f" [{n}] {h['id']} ({h['source']})")
for warning in check_citations(answer, hits):
print(f"WARNING: {warning}")
if __name__ == "__main__":
main()
You don't need to understand every line yet. The table after Step 8 explains each part.
#Step 3: Index the guides and ask your first question
Run the script with the real embedding model from Unit 3:
python unit07/rag_minimal.py
If the model can't be downloaded on your network, run the same command with --offline. It uses toy embeddings that match words, not meaning:
python unit07/rag_minimal.py --offline
On the first run you should see something like this (the output below is from --offline; with the real model the similarity numbers differ):
Wrote 8 sample guides to rag_docs/
8 guides -> 19 chunks
Indexed 19 chunks in rag_store/ (collection guides_toy-hash-256)
Question: "Who can release an order blocked by the credit check?"
Retrieved:
[1] o2c-credit-blocks.md#1 similarity 0.61
[2] o2c-credit-blocks.md#2 similarity 0.56
[3] o2c-credit-blocks.md#3 similarity 0.41
Prompt the model would receive (system instruction, then this):
Sources:
[1] (from o2c-credit-blocks.md#1)
Guide: sales orders blocked by credit checks > When an order is blocked
An order is held for delivery when the customer's open receivables plus ...
[2] (from o2c-credit-blocks.md#2)
Guide: sales orders blocked by credit checks > Who releases it
Only the credit management team may release a credit block. ...
Question: Who can release an order blocked by the credit check?
No model called. Add --llm sample (no account) or --llm MODEL to generate an answer.
Stages 1 to 6 have run. Open the unit07/rag_docs folder: the eight guides are plain Markdown files. Open one and compare it with its chunks: each ## section became one chunk, prefixed with the guide's title. The index is saved in unit07/rag_store, so the next run skips the embedding pass.
#Step 4: Generate an answer, with and without an account
Without an account, use the made-up answer. It copies the first sentence of source 1 and cites it, so you can see the citation check work:
python unit07/rag_minimal.py "Why is the invoice blocked for payment?" --llm sample
Answer:
[sample answer, no model called] An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance. [1]
Sources:
[1] p2p-invoice-blocks.md#2 (p2p-invoice-blocks.md)
[2] p2p-invoice-blocks.md#3 (p2p-invoice-blocks.md)
[3] p2p-invoice-blocks.md#1 (p2p-invoice-blocks.md)
With your SAP AI Core key from Unit 5, pass a model name from your catalog. Choosing and calling LLMs shows how to list the names your account offers:
python unit07/rag_minimal.py "Why is the invoice blocked for payment?" --llm MODEL_NAME
Replace MODEL_NAME with a name from your catalog. A good answer names both the quantity and the price reason and cites [1], [2] or [3] for each. If the script prints WARNING: The answer cites no source, the model ignored the instruction; that is worth recording for Unit 8.
python unit07/rag_minimal.py "What is the parental leave policy?"
Question: "What is the parental leave policy?"
No chunk reached similarity 0.2. Answer: I can't find this in the guides.
(No model was called. Saying 'not found' is the correct result for an out-of-scope question.)
This empty result is correct. The minimum similarity stopped the pipeline before the model, so there was nothing for it to make up.
Now switch the threshold off and watch what goes wrong:
python unit07/rag_minimal.py "What is the parental leave policy?" --min-sim -1 --llm sample
The script now retrieves the three "least bad" chunks, about pricing or safety stock, and the sample answer cites one of them. It looks grounded and is nonsense. A real model given the system instruction will usually say it can't find the answer, but not always. The threshold is your first defense.
Ask a planning question, but only search plan-to-produce guides:
python unit07/rag_minimal.py "What does reschedule in mean?" --process plan-to-produce
Question: "What does reschedule in mean?" (only plan-to-produce)
Retrieved:
[1] ptp-mrp-exceptions.md#2 similarity 0.28
[2] ptp-mrp-exceptions.md#1 similarity 0.27
The filter is a where condition on the chunk metadata, applied inside the search. Only two chunks passed the threshold, so only two are sent. In production, the same mechanism carries company codes and authorizations; a later Unit 7 topic covers that.
ok Who is allowed to release an order that failed the credit check?
expected o2c-credit-blocks.md; got o2c-credit-blocks.md, o2c-credit-blocks.md, o2c-credit-blocks.md
...
MISS On which days are suppliers paid?
expected p2p-payment-runs.md; got p2p-new-suppliers.md, p2p-new-suppliers.md, ptp-mrp-exceptions.md
Hit rate at 3: 5/6 questions found the right guide in the top 3.
This is from --offline. The toy embeddings miss "suppliers paid" because the guide says "payment run", a different word. A model that understands meaning should close that gap; run without --offline to compare. The hit rate (how often the right document is in the top k) is the first retrieval metric. You'll build on it in Unit 8.
In unit07/rag_docs, create a new file o2c-returns.md with this text and save:
process: order-to-cash
# Guide: customer returns
## Approval
Returns above 1,000 euros need approval from the sales manager before the credit memo is created.
Ask about it. The script notices the documents changed and rebuilds the index:
You should see Documents changed (or --reindex): rebuilding the index and o2c-returns.md#1 as source 1. Then try the same question in other words, such as "Who approves a large customer return?". With --offline, the toy embeddings miss it because the words differ; the real model should find it.
Keep the rebuildable index out of Git. Open .gitignore, add this line at the end and save:
SAP packages the same pipeline as the grounding module of the orchestration service in the generative AI hub (SAP AI Core). As of October 2026, SAP documents it like this:
RAG stage
SAP AI Core grounding
Added (per SAP release notes)
Ingest, chunk, embed, index
Pipeline API: reads documents from a source, segments them into chunks, generates embeddings and stores them in a vector database. Sources include Microsoft SharePoint, Amazon S3 and SFTP; SAP Document Management (September 2025) and ServiceNow (March 2026) were added later as repository types
November 2024
Index your own chunks
Vector API: manages collections and documents in the vector database, so you can send chunks you prepared yourself
December 2024
Retrieve
Retrieval API: searches data repositories and returns relevant chunks
December 2024
Augment and generate
Grounding module in orchestration: retrieves chunks for an input placeholder and puts them in an output placeholder in your prompt template
November 2024
The SDK reference says grounding uses the SAP HANA vector engine underneath. A filter can point to vector repositories (your documents) or to the help.sap.com type, which searches SAP's help portal. SAP Architecture Center recommends these managed services where they fit, and building custom RAG on SAP BTP when you need more control.
The orchestration v2 classes ship in SAP's Python SDK, sap-ai-sdk-gen (version 7.4.1 on PyPI as of September 2026). The grounding configuration has the same shape as your script: an input placeholder (the question), an output placeholder (the retrieved context), and a limit like your --top.
Sketch (unit07/grounding_sketch.py):
"""SKETCH: the same RAG flow, with SAP's grounding module doing ingest-to-retrieve.
Needs an SAP AI Core service key with orchestration and a grounding data repository
(set up with the Pipeline API or Vector API). Builds the configuration; runs it only with --run.
"""
import argparse
from gen_ai_hub.orchestration_v2 import (DataRepositoryType, DocumentGroundingConfig,
DocumentGroundingFilter, DocumentGroundingPlaceholders,
GroundingModuleConfig, GroundingSearchConfig, GroundingType,
LLMModelDetails, ModuleConfig, OrchestrationConfig,
PromptTemplatingModuleConfig, SystemMessage, Template,
UserMessage)
def make_config(model: str, repository_id: str) -> OrchestrationConfig:
grounding = GroundingModuleConfig(
type=GroundingType.DOCUMENT_GROUNDING_SERVICE,
config=DocumentGroundingConfig(
placeholders=DocumentGroundingPlaceholders(input=["question"], output="context"),
filters=[DocumentGroundingFilter(
id="process-guides",
data_repository_type=DataRepositoryType.VECTOR, # your own documents
data_repositories=[repository_id], # from the Retrieval API
search_config=GroundingSearchConfig(max_chunk_count=3))])) # like --top 3
template = Template(template=[
SystemMessage(content="Answer only from the sources. If they don't contain the answer, say so."),
UserMessage(content="Sources:\n{{?context}}\n\nQuestion: {{?question}}")])
return OrchestrationConfig(modules=ModuleConfig(
grounding=grounding,
prompt_templating=PromptTemplatingModuleConfig(prompt=template, model=LLMModelDetails(name=model))))
def main() -> None:
parser = argparse.ArgumentParser(description="Sketch: SAP orchestration with document grounding.")
parser.add_argument("question", nargs="?", default="Who can release an order blocked by the credit check?")
parser.add_argument("--model", default="MODEL_NAME")
parser.add_argument("--repository", default="REPOSITORY_ID")
parser.add_argument("--run", action="store_true", help="really call SAP AI Core (needs .env and a repository)")
args = parser.parse_args()
config = make_config(args.model, args.repository)
print("Configuration built. Grounding filter:", config.modules.grounding.config.filters[0].id)
if args.run:
from dotenv import load_dotenv
from gen_ai_hub.orchestration_v2 import OrchestrationService
load_dotenv()
service = OrchestrationService(config=config)
result = service.run(placeholder_values={"question": args.question})
print(result.final_result.choices[0].message.content)
if __name__ == "__main__":
main()
The JSON this sends puts grounding in a module with "type": "document_grounding_service", "placeholders": {"input": ["question"], "output": "context"} and a filter with "data_repository_type": "vector". Compare it with your script: the Pipeline API replaces ingest, chunk, embed and index; the grounding module replaces retrieve and augment; {{?context}} is filled by SAP instead of by augment().
Joule. For assistants that end users reach through Joule, SAP offers document grounding as a Joule capability that draws on documents in SAP and third-party repositories. You configure sources; you don't write the pipeline.
Licensing notes. Chroma and Sentence Transformers are free. SAP AI Core usage (model calls, grounding) is billed against your SAP BTP account; check the current metrics in your contract. SAP lists Joule document grounding as an AI feature with a flat fee.
Documents in SharePoint, S3, SFTP, SAP Document Management
You write connectors
Built-in pipelines
Configured sources
Custom chunking, re-ranking, hybrid search
Full control
Within the module's settings
Not exposed
Answers inside SAP apps for end users
You build the UI
You build the UI
Built in
Operations, scaling, monitoring
Yours
SAP runs the pipeline and store
SAP runs it
Mixed with live SAP data via tools
Yours to wire
Combine with orchestration and your app
Joule's other capabilities
Cost
Hosting plus model calls
BTP consumption
Flat fee per SAP's listing
A common path: prototype with your own script to learn what your documents need, then move ingestion and retrieval to the grounding module once the choices are clear.
Authorizations. RAG can leak anything in the index. Index only what the audience may read, or store access labels (company code, sales organization, role) as metadata and filter on them at retrieval, using the user's identity, not a choice in the request. A later Unit 7 topic covers grounding with SAP authorizations.
Prompt injection through documents. A retrieved chunk is untrusted input. A document saying "ignore your instructions" ends up in the prompt. Keep instructions in the system message, and treat source text as data.
Freshness. Record each document's version and index date. Re-index on change, and remove chunks of deleted documents; a stale chunk gives a confident old answer.
Evaluation. Track retrieval (hit rate at k) separately from answer quality (correct, grounded, cited). Your EVAL list is the start; Unit 8 builds a harness.
Cost. Each extra chunk in k is paid on every call. Measure whether k=5 beats k=3 before you pay for it.
Logging. Log the question, retrieved chunk IDs, similarities, model and answer, minus personal data, as in Errors, logging and debugging. Without chunk IDs you can't debug a bad answer.
Clean core. RAG reads documents and, if needed, SAP data through released APIs. It doesn't modify SAP. Keep it side by side on SAP BTP, as in CAP and side-by-side extensions.
Skipping the threshold. Top k always returns k chunks, even for nonsense questions. Without a minimum similarity, the model gets irrelevant text and is invited to use it.
Changing the embedding model without re-indexing. Old and new vectors aren't comparable. Rebuild the index whenever the model changes.
Chunks without context. "The tolerance is 2 percent" alone is ambiguous. Prefix chunks with title and section.
Debugging the model first. Print what was retrieved before you change prompts or models.
No "not found" path. Users trust an assistant more when it says "I can't find this" than when it guesses.
Trusting citations blindly. A model can cite [1] for a claim that [1] doesn't support. Citation checks catch missing or out-of-range numbers, not wrong ones; that needs evaluation.
Dumping every document in. Drafts and old versions produce contradictions. Curate the sources.
Extend your RAG with a second test set and a "not found" check. Its output, unit07/rag_eval.txt, is the seed of your evaluation set in Unit 8.
Open unit07/rag_minimal.py and find the EVAL list.
Add four questions about your new o2c-returns.md guide and the existing guides, each with the expected file name. Save.
Run python unit07/rag_minimal.py --eval (add --offline if needed) and note the hit rate.
Run three questions the guides don't cover (for example about travel expenses, holidays and IT passwords). Note whether each prints I can't find this in the guides.
If a question you expected to work was rejected, run it with --min-sim -1, note its similarity, and decide whether to change the default threshold.
Then add one line at the end of the file, by hand, with your three out-of-scope results and the threshold you chose.
Commit unit07/rag_minimal.py, unit07/rag_docs and unit07/rag_eval.txt.
Done when--eval covers at least ten questions, all three out-of-scope questions return "I can't find this in the guides", and unit07/rag_eval.txt is in your repository with the hit rate and chosen threshold.
Pick one answer for each question. The explanation appears after you choose.
1Your assistant answers a credit-block question wrongly. What do you check first?
Answer: B. If the right chunk isn't retrieved, nothing downstream can fix the answer. Check retrieval first, then the prompt and the model.
2Why must questions and chunks use the same embedding model?
Answer: C. Each model maps text into its own space. Comparing a question vector from one model with chunk vectors from another gives meaningless similarities, so changing models means re-indexing.
3In rag_minimal.py, why does chunk() prefix each chunk with the guide title and section?
Answer: A. A sentence like "the tolerance is 2 percent" is ambiguous alone. The title and section restore its context for both the embedding and the model.
4An out-of-scope question still gets a confident, cited answer. What change in the script prevents this?
Answer: D. Top k always returns k chunks, even for unrelated questions. The minimum similarity stops the pipeline with "I can't find this" before the model sees irrelevant text.
5What does check_citations() catch, and what does it miss?
Answer: B. The check only looks at the numbers. A model can cite [1] for something [1] doesn't say; catching that needs evaluation of groundedness, covered in Unit 8.
6In SAP's grounding module, what fills the output placeholder in your prompt template?
Answer: C. You name an input placeholder for the question and an output placeholder for the context. The grounding module retrieves chunks and puts them into the output placeholder, the job augment() does in your script.
7The index holds guides for several company codes, and a user may see only one. Where should that rule be enforced?
Answer: B. Anything retrieved reaches the prompt and can leak. Filtering at retrieval, based on who the user is rather than what they type, keeps forbidden chunks out of the prompt entirely.
8Which part of your script does SAP's Pipeline API replace?
Answer: C. SAP documents the Pipeline API as segmenting documents into chunks, generating embeddings and storing them in a vector database. Retrieval and augmentation are done by the grounding module.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Retrieval Augmented Generation (RAG) (SAP Architecture Center)— question encoding, document retrieval with similarity search in SAP HANA Cloud vector engine, response generation; managed SAP RAG services vs. custom-built on SAP BTP; SAP recommends standardized services where possible
Exploring RAG and Grounding Use Cases (SAP Learning)— grounding as anchoring outputs in reliable, context-specific data; RAG as one way to ground; often more economical than fine-tuning; IT support, HR policy and warranty examples
SAP AI Core (product guide PDF, help.sap.com)— grounding module added to orchestration 2024-11-04; Pipeline API chunks and embeds into a vector database; Vector API manages collections and documents; Retrieval API returns relevant chunks; SAP Document Management (2025-09-01) and ServiceNow (2026-03-01) added as repositories
Document Grounding (generative AI hub SDK reference)— grounding uses the SAP HANA vector engine; VECTOR repositories from SharePoint, S3, SFTP or chunks via the Vector API; URL type supports only help.sap.com; input_params, output_param, filters, max_chunk_count
sap-ai-sdk-gen (PyPI)— SAP Cloud SDK for AI (Python); version 7.4.1 of 23 September 2026; orchestration service access; Apache 2.0
Document Grounding (sap.com)— a Joule capability that draws on business documents in SAP and third-party repositories; listed as an AI feature with a flat fee
Query and Get (Chroma documentation)— query with query_embeddings, n_results (default 10), where metadata filter; returns documents, metadatas and distances by default