Check every claim in a RAG answer against the passages it was given, verify citations, count honest "I can't find it" answers, and report hallucination risk with numbers.
A hallucination is a confident statement that isn't backed by the facts or by the sources the AI was given. It reads exactly like a correct sentence. That is what makes it dangerous.
For a RAG assistant, the most useful test is groundedness: can every claim in the answer be traced to a passage the search actually returned? You can measure it. Split each answer into its claims, check each claim against the passages, and count:
How many claims are backed by the passages (often called faithfulness).
How many answers contain at least one unbacked claim (the hallucination rate a business user feels).
Whether the citations point to the right passage, so a user who clicks "source" sees the proof.
How often the assistant says "I can't find that" when the answer really isn't there, and how often it gives up when the answer is there.
Groundedness is not the same as truth. A grounded answer can still be wrong if the source document is wrong or outdated. But it turns "the AI sometimes makes things up" into a number you can track, set a target for and report.
Take the running procure-to-pay example. An accounts payable clerk asks why the invoice for purchase order 4500017311 is blocked. The case note says the supplier delivered 90 of 100 pieces and was asked for a credit memo. Three bad answers, all fluent:
"The supplier delivered 80 of 100 pieces." The clerk disputes the wrong quantity with the supplier.
"The supplier will send the credit memo within 5 working days." Nothing says so. The clerk promises a date to the business.
"Most companies set safety stock to two weeks of demand." It may even be true somewhere, but it isn't in your notes and isn't your policy.
Each costs something different: rework, a broken promise, or a decision made on general knowledge instead of company rules. Measuring groundedness pays off in three ways:
A risk number for go-live decisions. "1 in 10 answers had an unsupported claim on 200 test questions" is something a process owner can accept, reject or put controls around.
The right fix. If the right passage was retrieved but the model added things, fix the prompt or model. If the passage was missing, fix the search (Evaluating retrieval).
Safe "I don't know". An assistant that admits a gap costs a few minutes. One that invents an answer can cost a payment, a customer or an audit finding.
As of the SAP AI Core service guide dated 4 September 2026:
Grounding in the generative AI hub retrieves passages from your document repositories and inserts them into the prompt. In SAP's Python SDK, the passages that were used are returned in the response under the grounding module result. That matters: you can store exactly what the model saw and check the answer against it.
Evaluations in SAP AI Core compares prompt and model configurations. SAP lists system-defined metrics such as ROUGE, BLEU and COMET, and custom LLM-as-a-judge metrics with your own rating criteria. A groundedness check can be written as such a custom judge metric.
SAP Learning material on evaluating LLM use cases names groundedness and fact-checking metrics as crucial for RAG: they verify that a response is supported by the retrieved documents.
In the SAP documentation we could open in this run, we did not find a built-in metric named groundedness or faithfulness. Plan to define the check yourself, either as a custom judge metric in SAP AI Core or in your own harness, and calibrate it against people as LLM evaluation fundamentals showed.
"Delivered 80 of 100 pieces" when the note says 90
Checking each claim against the retrieved passages
Adds what the source doesn't say
"Credit memo within 5 working days"
Same check: the claim has no support
Brings in outside knowledge
"Most companies set it to two weeks"
Same check; also an instruction to use only the sources
Wrong citation
A correct sentence pointing to the wrong note
Checking the claim against the cited passage only
Answers what it can't know
Names the "cheapest supplier" from a note about expediting
Test questions with no answer in the documents
Gives up too easily
"I can't find PRC-112" when the note is right there
Test questions that do have an answer
Researchers group the first three as unfaithful to the context or to the facts. For a business owner the useful split is simpler: wrong, unbacked and unhelpfully silent. Each needs its own number.
A good one-page report for a steering committee has five lines:
Hallucination rate per answer, with a range: "5 of 10 answers, somewhere between 24% and 76% given so few tests". A small test set gives a wide range; say so.
Severity split: contradictions of the source are usually worse than harmless extra detail.
Abstention: share of unanswerable questions where it said "I can't find it", and share of answerable ones where it gave up.
How the judge was checked: "the automatic check caught 5 of 6 hallucinated claims a person found, and raised 1 false alarm in 14".
The control: what happens when a claim can't be backed (hide it, flag it, route to a person).
"RAG stops hallucinations." It gives the model the right facts. It doesn't stop the model from changing a number or adding a promise. Measure it.
"Grounded means true." It means backed by the retrieved text. If the document is outdated, a grounded answer is outdated too.
"An answer with citations is verified." Citations can point to the wrong passage. In one research benchmark, about half of the answers from strong models were not fully supported by the passages they cited.
"The AI judge is the ground truth." A judge is a measuring instrument. Check how many real hallucinations it catches before trusting its numbers.
"Fewer 'I don't know' answers is always better." A model that always guesses scores well on accuracy-only tests and badly in production. Score confident errors as worse than honest abstentions.
"One percentage is enough." "5% hallucination" on 20 questions and on 2,000 questions mean very different things. Ask for the range.
Pick one answer for each question. The explanation appears after you choose.
1A RAG assistant's answer is fully grounded. What does that tell you?
Answer: B. Groundedness checks the answer against the retrieved passages, nothing more. If a source document is outdated or wrong, a grounded answer repeats the error, so document quality still needs its own owner.
2The assistant says the credit memo will arrive "within 5 working days", but no note mentions a date. What kind of problem is this?
Answer: C. The claim isn't contradicted by the note, it simply has no support in it. These invented details are what a claim-by-claim groundedness check catches, and they often sound the most helpful.
3Which test questions show whether the assistant knows when to say "I can't find it"?
Answer: D. Abstention can only be measured on questions the documents can't answer. Pair them with answerable questions to also catch an assistant that gives up too easily.
4A vendor reports "4% hallucination rate". What should you ask first?
Answer: A. A rate on 25 answers has a wide range, and a rate from an unchecked AI judge may miss many hallucinations. The number of tests, the labelers and the judge's catch rate decide how much the 4% is worth.
5Why should a scoring scheme for an SAP assistant punish confident errors more than "I don't know"?
Answer: C. If a wrong answer and an abstention both score zero, guessing always looks better. In a process like accounts payable, a wrong answer can cause a wrong payment or dispute, so the score should reflect that cost.
6SAP's grounding returns the passages used in its response. Why does that matter for hallucination control?
Answer: B. A groundedness check needs the exact passages the model received. Logging them with each answer lets you score test runs and trace a production complaint to the evidence.
7The automatic judge flags 5 of the 6 hallucinated claims a person found. How should you report this?
Answer: D. The hallucination rate is measured by the judge, so its blind spots are part of the risk. Reporting the catch rate and false alarms tells leaders how far to trust the headline number.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Treat groundedness as claim-level entailment against the evidence the model was given, and treat abstention as a separate decision with a price on wrong answers.
An answer is a bag of claims. Each claim is checked on its own: does the retrieved text support it, contradict it, or say nothing?
Citations are a second, stricter check: does the cited passage support the claim, not just some passage?
Some questions should not be answered. Whether the system declines them, and whether it declines answerable ones, is measured on purpose-built test cases.
The checker is a judge, and judges are measured against people before their numbers are believed. That is the lesson of LLM evaluation fundamentals, applied to one specific grading task.
Retrieval evaluation asked "did the right passage reach the model?" (Evaluating retrieval). This topic asks the next question: "given those passages, did the model stay within them, and did it know when to stop?"
Huang and colleagues' survey (ACM Transactions on Information Systems) splits hallucinations into two families:
Factuality hallucination: the output conflicts with real-world facts (factual contradiction) or states things that can't be verified (factual fabrication).
Faithfulness hallucination: the output departs from the user's instructions (instruction inconsistency), from the provided context (context inconsistency), or from itself (logical inconsistency).
In an enterprise RAG system, you can test faithfulness to the context directly, because you hold the context. Factuality against "the world" is mostly out of scope: your process documents are the source of truth, and a claim from general knowledge is a defect even when it is true. So the lab labels each claim with one of three verdicts:
A paragraph can hold one wrong number among five correct facts. Scoring the whole answer as one unit hides that. Two well-known methods split first:
Ragas faithfulness extracts the claims in a response, checks whether each can be inferred from the retrieved context, and divides supported claims by all claims. It needs no reference answer, only the context. Its documented example scores 0.5 for an answer with one correct and one wrong claim.
FActScore (Min et al., EMNLP 2023) breaks long text into atomic facts and reports the share supported by a knowledge source. On biographies, ChatGPT scored 58%.
Both use a model to do the splitting. The lab splits by sentence, which is predictable and good enough when you control the answer format. If your assistant writes long sentences with several facts, ask the generator to write one fact per sentence, or use a model to split.
Numbers and codes must match; most content words must appear
Free, fast, explainable
Misses negation and paraphrase
Small NLI model
A classifier scores whether the passage entails the claim
Cheap, runs locally
Needs a download; English-only in the open HHEM version
LLM judge
A model reads passage and claim and gives a verdict
Handles paraphrase and logic
Costs per call; has its own errors
An example of the middle option: Vectara's HHEM-2.1-Open takes (premise, hypothesis) pairs and returns a score from 0 to 1, with 0.5 as the threshold between consistent and hallucinated. Its model card says it is based on FLAN-T5-base, about 0.1B parameters, Apache 2.0, English-only, and runs on a CPU. Ragas offers a faithfulness variant that uses it for the verification step.
Whichever you choose, measure it against people first: how many hallucinated claims does it catch, and how many good claims does it flag? The lab's agree command does exactly that.
The ALCE benchmark (Gao et al., EMNLP 2023) defines two citation metrics, judged with an NLI model:
Citation recall, per statement: 1 if it has at least one citation and the cited passages together entail it.
Citation precision, per citation: a citation counts only if its statement passes recall and the citation isn't irrelevant. A citation is irrelevant when it neither supports the claim alone nor is needed alongside the others.
The ALCE authors report that on their ELI5 dataset, around half of the answers from their ChatGPT and GPT-4 baselines were not fully supported by the cited passages. Treat citations as claims that need checking, not as proof.
The lab reports citation recall over the claims that are grounded. That separates two failures: "the sentence is made up" (groundedness) and "the sentence is right but points to the wrong note" (citation).
#Step 4: measure abstention, and price wrong answers
Kalai, Nachum, Vempala and Zhang ("Why Language Models Hallucinate", OpenAI, September 2025) argue that hallucinations persist partly because evaluations grade right or wrong, 1 or 0. Under that scoring, "I don't know" always earns 0, while a guess sometimes earns 1, so guessing wins. Their proposal is to state a confidence target in the instructions: answer only if you are more than t confident, because a mistake costs t / (1 − t) points and a correct answer earns 1.
Applied to your own evaluation:
Add test questions the documents can't answer, and count how often the assistant declines.
Add answerable questions and count false abstentions, so you don't reward a system that refuses everything.
Score outcomes with a penalty for wrong answers. With t = 0.75, the penalty is 0.75 / 0.25 = 3. The lab's --penalty 3 uses that.
Anthropic's guidance on reducing hallucinations points the same way from the prompt side: give the model explicit permission to say "I don't know", restrict it to the provided documents, and have it cite quotes for each claim and retract claims it can't support. The same page notes these steps reduce hallucinations but don't eliminate them, which is why you still measure.
A hallucination rate from 10 answers is a rough estimate. The lab reports a Wilson 95% interval, a standard way to put a range around a rate measured on few cases. With 5 hallucinated answers out of 10, it spans about 24% to 76%. With 50 out of 100, the same 50% rate spans roughly 40% to 60%. The width is the honest message to the business: more test cases, narrower range.
#Build it yourself: score groundedness, citations and abstention
Before you start: complete Set up your computer for this course and Set up for Unit 8. They create your orchestrate-course folder with its .venv and unit08 folder. The script itself uses only Python's built-in modules. The optional model judge uses anthropic and python-dotenv from Unit 1.
You will build groundedness_lab.py. It holds 12 test cases: a question, the note IDs the search returned, the assistant's answer with citations like [D07], and a person's verdict on every claim. The notes are a subset of the made-up SAP help notes from Evaluating retrieval. The answers are deliberately seeded with each kind of hallucination from the table above. The script splits the answers into claims, judges them, checks citations, counts abstentions, compares its judge with the person, and writes a report.
flowchart LR
I[init<br/>cases file] --> CH[check<br/>claim verdicts]
CH --> AG[agree<br/>judge vs person]
AG --> SC[score<br/>rates and report]
SC --> E[edit cases<br/>and rerun]
E --> CH
In VS Code, right-click the unit08 folder, choose New File, name it groundedness_lab.py, paste the code below and save.
"""Unit 8: measure hallucinations and groundedness in RAG answers.
Each test case holds a question, the notes the search retrieved, the assistant's answer with
citations like [D07], and a person's verdict on every claim in the answer. The script splits
answers into claims, has a judge check each claim against the retrieved notes, checks the
citations, counts abstentions ("I can't find ..."), and compares the judge with the person.
How to run (from your course folder, with .venv turned on):
python unit08/groundedness_lab.py init # write the test cases to a file
python unit08/groundedness_lab.py check # claim-by-claim verdicts
python unit08/groundedness_lab.py agree # does the judge agree with the person?
python unit08/groundedness_lab.py score # groundedness, citations, abstention
python unit08/groundedness_lab.py score --penalty 3 # score that punishes confident errors
python unit08/groundedness_lab.py score --report # also save unit08/groundedness_report.md
Add --judge llm to check, agree or score to use a model as judge (needs ANTHROPIC_API_KEY in .env).
"""
import argparse
import json
import math
import os
import re
import sys
from datetime import date
from pathlib import Path
HERE = Path(__file__).resolve().parent
CASES_FILE = HERE / "grounded_cases.jsonl"
REPORT_FILE = HERE / "groundedness_report.md"
VERDICTS = ["supported", "contradicted", "unsupported"]
# ---------------------------------------------------------------------------
# 1. The notes the search can return: made-up SAP help notes (the same IDs as evaluating-retrieval).
# ---------------------------------------------------------------------------
NOTES = {
"D01": "A sales order is blocked for delivery when the customer's credit exposure is above the "
"credit limit. Exposure includes open orders, deliveries and unpaid invoices. The credit "
"analyst reviews the open items. If the overrun is below 2 percent the analyst releases "
"the order; above that the credit manager approves the release.",
"D02": "Sales can put a delivery block on an order, for example while export papers are missing "
"or because the customer asked to wait. Once the reason is cleared, remove the delivery "
"block in the order. The delivery is created in the next delivery run.",
"D03": "A billing block stops the invoice for a delivered order, usually because a price or "
"discount has to be checked first. After the pricing review, release the billing block.",
"D04": "Error PRC-112 means a mandatory price condition is missing on the order item. Maintain "
"the condition record for the customer and material, then redetermine prices on the order.",
"D05": "An order is incomplete when mandatory data is missing, such as the payment terms or the "
"ship-to party. Incomplete orders cannot be delivered or billed. Open the incompletion "
"log on the order to see which fields are missing and fill them.",
"D06": "Each supplier invoice is checked against the purchase order and the goods receipt. A "
"quantity or price variance above the tolerance blocks the invoice for payment until "
"someone resolves it.",
"D07": "Case: quantity variance on purchase order 4500017311. The supplier delivered 90 of 100 "
"pieces but invoiced all 100. The invoice was blocked for payment. Accounts payable asked "
"the supplier for a credit memo for the 10 missing pieces.",
"D08": "The vendor billed a higher price than the one agreed on the purchase order. If the price "
"variance is above tolerance, the buyer contacts the vendor and asks for a corrected "
"invoice or a credit memo.",
"D09": "When the variance is resolved, the accounts payable clerk releases the blocked invoice. "
"Released invoices are picked up by the next payment run if they are due.",
"D10": "The due date of an invoice is calculated from the baseline date and the payment terms of "
"the supplier. The payment run pays every invoice that is due and not blocked.",
"D12": "The planning run flags exception messages: missing parts, late receipts and orders to "
"reschedule. Planners review the exceptions every morning, starting with critical components.",
"D13": "The planning run showed a shortage of component RM-4711 for production in week 38. The "
"planner moved the production order forward and expedited the supplier.",
"D15": "Safety stock is a buffer quantity that planning keeps for uncertain demand. When available "
"stock falls below safety stock, the planning run creates a planned order.",
"D17": "A change to a customer's bank details needs a second person to approve it. This control "
"reduces the risk of fraud in refunds and credit memos.",
}
# ---------------------------------------------------------------------------
# 2. The test cases. "labels" is a person's verdict on each claim (each sentence that is not an
# "I can't find" sentence), in order. "answerable" says whether the notes hold the answer.
# ---------------------------------------------------------------------------
BUILT_IN_CASES = [
{"id": "c01", "question": "Why is the invoice for purchase order 4500017311 blocked?",
"retrieved": ["D07", "D06"], "answerable": True,
"answer": "The supplier delivered 90 of 100 pieces but invoiced all 100 [D07]. A quantity variance "
"above the tolerance blocks the invoice for payment [D06]. Accounts payable asked the "
"supplier for a credit memo for the 10 missing pieces [D07].",
"labels": ["supported", "supported", "supported"]},
{"id": "c02", "question": "Why is the invoice for purchase order 4500017311 blocked?",
"retrieved": ["D07", "D06"], "answerable": True,
"answer": "The supplier delivered 80 of 100 pieces [D07]. The invoice stays blocked until the "
"plant manager signs a waiver [D06].",
"labels": ["contradicted", "unsupported"]},
{"id": "c03", "question": "Customer Orion's order is 1.5 percent over the credit limit. Who can release it?",
"retrieved": ["D01"], "answerable": True,
"answer": "If the overrun is below 2 percent, the credit analyst releases the order [D01]. "
"At 1.5 percent, the credit analyst can release Orion's order [D01].",
"labels": ["supported", "supported"]},
{"id": "c04", "question": "What is the credit limit of customer 10023?",
"retrieved": ["D01", "D17"], "answerable": False,
"answer": "I can't find the credit limit of customer 10023 in the retrieved notes.",
"labels": []},
{"id": "c05", "question": "The vendor charged more than we agreed. What happens next?",
"retrieved": ["D08", "D06"], "answerable": True,
"answer": "If the price variance is above tolerance, the buyer contacts the vendor [D08]. "
"The buyer requests an amended bill or a credit note from the vendor [D08].",
"labels": ["supported", "supported"]},
{"id": "c06", "question": "When will the credit memo for purchase order 4500017311 arrive?",
"retrieved": ["D07"], "answerable": False,
"answer": "The supplier will send the credit memo for the 10 missing pieces within 5 working days [D07].",
"labels": ["unsupported"]},
{"id": "c07", "question": "Why can we not send the goods to the client?",
"retrieved": ["D02", "D01", "D05"], "answerable": True,
"answer": "Sales may have put a delivery block on the order, for example while export papers are "
"missing [D02]. The order may be blocked for delivery because credit exposure is above "
"the credit limit [D01]. Incomplete orders cannot be delivered [D05].",
"labels": ["supported", "supported", "supported"]},
{"id": "c08", "question": "Can an incomplete order be billed?",
"retrieved": ["D05"], "answerable": True,
"answer": "Yes, incomplete orders can be billed once they are delivered [D05].",
"labels": ["contradicted"]},
{"id": "c09", "question": "When does the supplier get their money?",
"retrieved": ["D10", "D09"], "answerable": True,
"answer": "The payment run pays every invoice that is due and not blocked [D10]. The due date is "
"calculated from the baseline date and the payment terms of the supplier [D09].",
"labels": ["supported", "supported"]},
{"id": "c10", "question": "What is safety stock?",
"retrieved": ["D15"], "answerable": True,
"answer": "Safety stock is a buffer quantity that planning keeps for uncertain demand [D15]. "
"Most companies set it to two weeks of average demand.",
"labels": ["supported", "unsupported"]},
{"id": "c11", "question": "Which supplier is cheapest for component RM-4711?",
"retrieved": ["D13", "D12"], "answerable": False,
"answer": "The planner expedited the supplier of component RM-4711 [D13]. That supplier offers "
"the lowest price for RM-4711 [D13].",
"labels": ["supported", "unsupported"]},
{"id": "c12", "question": "What does error PRC-112 mean?",
"retrieved": ["D04", "D03"], "answerable": True,
"answer": "I can't find information about PRC-112 in the retrieved notes.",
"labels": []},
]
# ---------------------------------------------------------------------------
# 3. Splitting an answer into claims and citations.
# ---------------------------------------------------------------------------
ABSTAIN = re.compile(r"\b(i can't find|i cannot find|i don't know|not in the retrieved|no information)\b", re.I)
CITE = re.compile(r"\[(D\d+)\]")
STOP = set("a an the and or but of to in on for from by with at as is are was were be been it its this "
"that these those there their they them we our you your if then than so may can will "
"has have had do does not no yes once when while which who what why how all every each "
"most also into out up".split())
def split_claims(answer: str) -> tuple[list[dict], bool]:
"""Return the claims (sentence text plus cited note IDs) and whether the answer abstains."""
claims, abstained = [], False
for sentence in re.split(r"(?<=[.!?])\s+", answer.strip()):
if not sentence:
continue
if ABSTAIN.search(sentence):
abstained = True
continue
text = re.sub(r"\s+([.!?,])", r"\1", CITE.sub("", sentence)).strip()
claims.append({"text": text, "cites": CITE.findall(sentence)})
return claims, abstained
def numbers(text: str) -> set[str]:
"""Numbers and codes that contain a digit: 90, 1.5, 4500017311, PRC-112, RM-4711."""
return {n.rstrip(".,").lower() for n in re.findall(r"[A-Za-z]*-?\d[\d.,-]*", text)}
def words(text: str) -> set[str]:
out = set()
for w in re.findall(r"[a-z][a-z'-]+", text.lower()):
w = w.replace("'s", "")
if w in STOP or len(w) < 3:
continue
out.add(w[:-1] if w.endswith("s") and not w.endswith("ss") else w) # plural to singular
return out
# ---------------------------------------------------------------------------
# 4. Judges: does the evidence support the claim?
# ---------------------------------------------------------------------------
def rules_judge(claim: str, evidence: str) -> str:
"""A cheap, transparent judge: numbers must match, and most content words must appear."""
missing = numbers(claim) - numbers(evidence)
if missing:
return "contradicted" if numbers(evidence) else "unsupported"
claim_words = words(claim)
if not claim_words:
return "supported"
coverage = len(claim_words & words(evidence)) / len(claim_words)
return "supported" if coverage >= 0.75 else "unsupported"
JUDGE_PROMPT = (
"You check whether a claim is supported by evidence. Use only the evidence, not your own "
"knowledge. 'supported': the evidence states or clearly implies the claim. 'contradicted': the "
"evidence says something that conflicts with the claim. 'unsupported': the evidence neither "
"supports nor contradicts it. Think briefly, then end with one line exactly like: VERDICT: supported")
def call_llm(system: str, user: str) -> str:
"""One request to a model through Anthropic's SDK, as in Unit 1. Replace only this to change provider."""
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
message = client.messages.create(
model=os.environ.get("LLM_MODEL", "claude-opus-5-5"), max_tokens=400,
system=system, messages=[{"role": "user", "content": user}])
return "".join(block.text for block in message.content if block.type == "text")
def llm_judge(claim: str, evidence: str) -> str:
reply = call_llm(JUDGE_PROMPT, f"Evidence:\n{evidence}\n\nClaim:\n{claim}")
found = re.findall(r"VERDICT:\s*(supported|contradicted|unsupported)", reply, re.I)
return found[-1].lower() if found else "unsupported" # no verdict line: treat as not supported
JUDGES = {"rules": rules_judge, "llm": llm_judge}
_cache: dict = {}
def judge(name: str, claim: str, evidence: str) -> str:
key = (name, claim, evidence)
if key not in _cache:
_cache[key] = JUDGES[name](claim, evidence)
return _cache[key]
def evidence_for(case: dict, ids: list[str]) -> str:
"""The question counts as evidence too: a claim may repeat what the user said."""
notes = " ".join(NOTES[i] for i in ids if i in NOTES)
return f"Question: {case['question']}\n{notes}"
# ---------------------------------------------------------------------------
# 5. Checking one case: groundedness against all retrieved notes, citations against cited notes.
# ---------------------------------------------------------------------------
def check_case(case: dict, judge_name: str) -> dict:
claims, abstained = split_claims(case["answer"])
context = evidence_for(case, case["retrieved"])
for c in claims:
c["verdict"] = judge(judge_name, c["text"], context)
cited_ok = bool(c["cites"]) and judge(judge_name, c["text"], evidence_for(case, c["cites"])) == "supported"
c["citation_ok"] = cited_ok # citation recall: cited notes alone support the claim
c["cite_relevant"] = []
for cid in c["cites"]: # citation precision: does this citation pull its weight?
alone = judge(judge_name, c["text"], evidence_for(case, [cid])) == "supported"
others = [o for o in c["cites"] if o != cid]
without = bool(others) and judge(judge_name, c["text"], evidence_for(case, others)) == "supported"
c["cite_relevant"].append(cited_ok and (alone or not without))
grounded = all(c["verdict"] == "supported" for c in claims)
if abstained and not claims:
outcome = "abstained"
elif grounded and case["answerable"]:
outcome = "correct"
else:
outcome = "wrong" # an unsupported claim, or a confident answer the notes can't back
return {"case": case, "claims": claims, "abstained": abstained, "outcome": outcome}
# ---------------------------------------------------------------------------
# 6. Loading and validating the cases file.
# ---------------------------------------------------------------------------
def load_cases() -> list[dict]:
if not CASES_FILE.exists():
sys.exit(f"Missing {CASES_FILE.relative_to(HERE.parent)}. Run: python unit08/groundedness_lab.py init")
cases = []
for n, line in enumerate(CASES_FILE.read_text(encoding="utf-8").splitlines(), start=1):
if not line.strip():
continue
try:
case = json.loads(line)
except json.JSONDecodeError:
sys.exit(f"{CASES_FILE.name}: line {n} is not valid JSON")
claims, _ = split_claims(case.get("answer", ""))
if len(claims) != len(case.get("labels", [])):
sys.exit(f"{case.get('id', 'line ' + str(n))}: the answer has {len(claims)} claims "
f"but {len(case.get('labels', []))} labels")
unknown = [i for i in case.get("retrieved", []) if i not in NOTES]
if unknown or any(v not in VERDICTS for v in case["labels"]):
sys.exit(f"{case['id']}: unknown note ID {unknown} or a label not in {VERDICTS}")
cases.append(case)
return cases
def wilson(hits: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% interval for a rate measured on n cases (Wilson score interval)."""
if n == 0:
return 0.0, 0.0
p = hits / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return max(0.0, centre - half), min(1.0, centre + half)
# ---------------------------------------------------------------------------
# 7. Commands.
# ---------------------------------------------------------------------------
def cmd_init(force: bool) -> None:
if CASES_FILE.exists() and not force:
sys.exit(f"{CASES_FILE.relative_to(HERE.parent)} already exists. Add --force to overwrite your edits.")
with CASES_FILE.open("w", encoding="utf-8") as f:
for case in BUILT_IN_CASES:
f.write(json.dumps(case) + "\n")
claims = sum(len(c["labels"]) for c in BUILT_IN_CASES)
print(f"Wrote {len(BUILT_IN_CASES)} test cases to {CASES_FILE.relative_to(HERE.parent)}")
print(f" claims labeled by a person: {claims}")
print(f" answerable: {sum(c['answerable'] for c in BUILT_IN_CASES)}, "
f"not answerable from the notes: {sum(not c['answerable'] for c in BUILT_IN_CASES)}")
def cmd_check(judge_name: str) -> None:
for r in (check_case(c, judge_name) for c in load_cases()):
case = r["case"]
print(f"{case['id']} {case['question']}")
if r["abstained"]:
print(" (abstains: says it can't find the answer)")
for c, label in zip(r["claims"], case["labels"]):
flag = "" if c["verdict"] == label else f" <- person said {label}"
cites = ",".join(c["cites"]) or "no citation"
print(f" {c['verdict']:<13} cite {'ok ' if c['citation_ok'] else 'BAD'} [{cites}] "
f"{c['text'][:60]}{flag}")
print(f" outcome: {r['outcome']}\n")
def cmd_agree(judge_name: str) -> None:
pairs = []
for r in (check_case(c, judge_name) for c in load_cases()):
pairs += [(label, c["verdict"]) for c, label in zip(r["claims"], r["case"]["labels"])]
print(f"Judge '{judge_name}' against the person's labels, {len(pairs)} claims\n")
print("person \\ judge " + "".join(f"{v:>14}" for v in VERDICTS))
for p in VERDICTS:
print(f"{p:<15} " + "".join(f"{sum(1 for a, b in pairs if a == p and b == v):>14}" for v in VERDICTS))
agree = sum(a == b for a, b in pairs)
bad = [(a, b) for a, b in pairs if a != "supported"]
good = [(a, b) for a, b in pairs if a == "supported"]
caught = sum(b != "supported" for _, b in bad)
false_alarms = sum(b != "supported" for _, b in good)
print(f"\nExact agreement: {agree}/{len(pairs)} = {agree / len(pairs):.0%}")
print(f"Hallucinated claims caught: {caught}/{len(bad)} (a missed one reaches the user)")
print(f"False alarms on good claims: {false_alarms}/{len(good)} (each one costs a reviewer's time)")
def cmd_score(judge_name: str, penalty: float, report: bool) -> None:
results = [check_case(c, judge_name) for c in load_cases()]
claims = [c for r in results for c in r["claims"]]
answered = [r for r in results if r["claims"]]
hallucinated = [r for r in answered if any(c["verdict"] != "supported" for c in r["claims"])]
contradicted = sum(c["verdict"] == "contradicted" for c in claims)
unsupported = sum(c["verdict"] == "unsupported" for c in claims)
cites = [ok for c in claims for ok in c["cite_relevant"]]
supported_claims = [c for c in claims if c["verdict"] == "supported"]
cite_recall = sum(c["citation_ok"] for c in supported_claims) / max(1, len(supported_claims))
unanswerable = [r for r in results if not r["case"]["answerable"]]
answerable = [r for r in results if r["case"]["answerable"]]
right_abstain = sum(r["outcome"] == "abstained" for r in unanswerable)
false_abstain = sum(r["outcome"] == "abstained" for r in answerable)
outcomes = {o: sum(r["outcome"] == o for r in results) for o in ("correct", "abstained", "wrong")}
lo, hi = wilson(len(hallucinated), len(answered))
by_person = sum(any(v != "supported" for v in r["case"]["labels"]) for r in answered)
rows = [
("Claims supported by the retrieved notes (faithfulness)",
f"{len(supported_claims)}/{len(claims)} = {len(supported_claims) / max(1, len(claims)):.0%}"),
(" of the rest: contradicted / unsupported", f"{contradicted} / {unsupported}"),
("Answers with at least one unsupported claim (hallucination rate)",
f"{len(hallucinated)}/{len(answered)} = {len(hallucinated) / max(1, len(answered)):.0%}, "
f"95% interval {lo:.0%} to {hi:.0%}"),
(" the same, by the person's labels", f"{by_person}/{len(answered)}"),
("Supported claims whose citations back them (citation recall)", f"{cite_recall:.0%}"),
("Citations that pull their weight (citation precision)",
f"{sum(cites)}/{len(cites)} = {sum(cites) / max(1, len(cites)):.0%}"),
("Unanswerable questions where it said it can't find it", f"{right_abstain}/{len(unanswerable)}"),
("Answerable questions where it wrongly gave up", f"{false_abstain}/{len(answerable)}"),
]
width = max(len(name) for name, _ in rows)
print(f"{len(results)} cases, {len(claims)} claims, judge '{judge_name}'\n")
for name, value in rows:
print(f"{name:<{width}} {value}")
score = outcomes["correct"] - penalty * outcomes["wrong"]
print(f"\nOutcomes: correct {outcomes['correct']}, abstained {outcomes['abstained']}, wrong {outcomes['wrong']}")
print(f"Accuracy-only score (wrong costs 0): {outcomes['correct']}/{len(results)}")
print(f"Risk-weighted score (correct +1, abstain 0, wrong -{penalty:g}): {score:+g}")
if report:
lines = [f"# Groundedness report ({date.today().isoformat()})", "",
f"{len(results)} test cases, {len(claims)} claims, judge `{judge_name}`.", "",
"| Measure | Result |", "| --- | --- |"]
lines += [f"| {name.strip()} | {value} |" for name, value in rows]
lines += ["", f"Outcomes: correct {outcomes['correct']}, abstained {outcomes['abstained']}, "
f"wrong {outcomes['wrong']}. Risk-weighted score with penalty {penalty:g}: {score:+g}.", ""]
REPORT_FILE.write_text("\n".join(lines), encoding="utf-8")
print(f"\nSaved the report to {REPORT_FILE.relative_to(HERE.parent)}")
def main() -> None:
parser = argparse.ArgumentParser(description="Measure hallucinations and groundedness in RAG answers.")
parser.add_argument("command", choices=["init", "check", "agree", "score"])
parser.add_argument("--judge", choices=sorted(JUDGES), default="rules", help="default: rules")
parser.add_argument("--penalty", type=float, default=1.0, help="points lost per wrong answer (default 1)")
parser.add_argument("--report", action="store_true", help="save unit08/groundedness_report.md")
parser.add_argument("--force", action="store_true", help="init: overwrite the cases file")
args = parser.parse_args()
if args.judge == "llm":
try:
from dotenv import load_dotenv
import anthropic # noqa: F401 (only checks that it is installed)
except ImportError:
sys.exit("Missing library. Run: pip install -r requirements.txt")
load_dotenv() # ANTHROPIC_API_KEY comes from .env, as in Unit 1
if not os.getenv("ANTHROPIC_API_KEY"):
sys.exit("ANTHROPIC_API_KEY is not set in .env (see Set up your computer). Or use --judge rules.")
if args.command == "init":
cmd_init(args.force)
elif args.command == "check":
cmd_check(args.judge)
elif args.command == "agree":
cmd_agree(args.judge)
else:
cmd_score(args.judge, args.penalty, args.report)
if __name__ == "__main__":
main()
Read the parts of the script:
Part
What it does
NOTES
The made-up SAP help notes the search can return, with the IDs used in Evaluating retrieval
BUILT_IN_CASES
12 test cases: question, retrieved note IDs, the answer, whether the notes can answer it, and a person's label per claim
split_claims
Splits an answer into sentences, pulls out [Dxx] citations, and spots abstention sentences such as "I can't find"
numbers, words
Find numbers and codes (90, 4500017311, RM-4711) and content words, ignoring small words like "the"
rules_judge
contradicted if a number in the claim differs from the evidence; supported if at least 75% of content words appear; else unsupported
call_llm, llm_judge
One Anthropic API call, as in Unit 1; reads the last VERDICT: line
judge
Remembers verdicts so the same check is never sent twice
evidence_for
Joins the question and the chosen notes into one evidence text
check_case
Judges every claim against all retrieved notes, then against its cited notes, then each citation alone; decides the outcome: correct, abstained or wrong
load_cases
Reads grounded_cases.jsonl and stops with a clear message if a line is broken or labels don't match claims
wilson
The 95% interval around a rate
cmd_check, cmd_agree, cmd_score
The three reports: per claim, judge against person, and the summary
Wrote 12 test cases to unit08/grounded_cases.jsonl
claims labeled by a person: 20
answerable: 9, not answerable from the notes: 3
Open unit08/grounded_cases.jsonl. Each line is one case:
{"id": "c06", "question": "When will the credit memo for purchase order 4500017311 arrive?", "retrieved": ["D07"], "answerable": false, "answer": "The supplier will send the credit memo for the 10 missing pieces within 5 working days [D07].", "labels": ["unsupported"]}
labels has one verdict per claim, in order. An "I can't find" sentence is not a claim, so c04 and c12 have empty labels.
Running init again stops with already exists. That protects your edits. Add --force only to get the original cases back.
c01 Why is the invoice for purchase order 4500017311 blocked?
supported cite ok [D07] The supplier delivered 90 of 100 pieces but invoiced all 100
supported cite ok [D06] A quantity variance above the tolerance blocks the invoice f
supported cite ok [D07] Accounts payable asked the supplier for a credit memo for th
outcome: correct
c02 Why is the invoice for purchase order 4500017311 blocked?
contradicted cite BAD [D07] The supplier delivered 80 of 100 pieces.
unsupported cite BAD [D06] The invoice stays blocked until the plant manager signs a wa
outcome: wrong
...
c08 Can an incomplete order be billed?
supported cite ok [D05] Yes, incomplete orders can be billed once they are delivered <- person said contradicted
outcome: correct
c09 When does the supplier get their money?
supported cite ok [D10] The payment run pays every invoice that is due and not block
supported cite BAD [D09] The due date is calculated from the baseline date and the pa
outcome: correct
...
c12 What does error PRC-112 mean?
(abstains: says it can't find the answer)
outcome: abstained
Read four cases closely:
c02 shows both kinds of error. "80 of 100" conflicts with the note's 90, so it is contradicted. "Plant manager signs a waiver" appears nowhere, so it is unsupported.
c08 is a miss. The note says incomplete orders cannot be billed; the answer says they can. Almost every word matches, so the word-overlap judge says supported. Negation is a classic blind spot of simple judges.
c09 is grounded but mis-cited. The second sentence is true to the retrieved notes, but it cites D09, which says nothing about baseline dates. A user who clicks the citation finds no proof.
c12 abstains, but the answer was there (D04 explains PRC-112). That is a false abstention: safe, but useless.
A line ending in <- person said ... is a disagreement between the judge and the person.
Judge 'rules' against the person's labels, 20 claims
person \ judge supported contradicted unsupported
supported 13 0 1
contradicted 1 1 0
unsupported 0 1 3
Exact agreement: 17/20 = 85%
Hallucinated claims caught: 5/6 (a missed one reaches the user)
False alarms on good claims: 1/14 (each one costs a reviewer's time)
Read it:
The miss is c08, the negation. In production that answer would reach the clerk unflagged.
The false alarm is c05: "requests an amended bill or a credit note" means the same as the note's "asks for a corrected invoice or a credit memo", but the words differ. The rules judge can't see paraphrase.
One disagreement doesn't matter much: c06 is unsupported to the person and contradicted to the judge, because the claim's "5" is not among the note's numbers. Both flag it.
For a groundedness judge, caught is the number to watch: a missed hallucination is the expensive error. False alarms cost review time.
12 cases, 20 claims, judge 'rules'
Claims supported by the retrieved notes (faithfulness) 14/20 = 70%
of the rest: contradicted / unsupported 2 / 4
Answers with at least one unsupported claim (hallucination rate) 5/10 = 50%, 95% interval 24% to 76%
the same, by the person's labels 5/10
Supported claims whose citations back them (citation recall) 93%
Citations that pull their weight (citation precision) 13/19 = 68%
Unanswerable questions where it said it can't find it 1/3
Answerable questions where it wrongly gave up 1/9
Outcomes: correct 5, abstained 2, wrong 5
Accuracy-only score (wrong costs 0): 5/12
Risk-weighted score (correct +1, abstain 0, wrong -1): +0
Read it:
50% looks the same for judge and person, but it isn't the same five answers. The judge flags c05 (false alarm) and misses c08. Equal totals can hide offsetting errors; that is why Step 5 comes first.
The interval is wide: 24% to 76% on 10 answered cases. This test set is built to contain every failure type, so its rate says nothing about a real assistant. Real numbers need real questions, and more of them.
Citation precision is low because citations attached to unsupported claims count as wrong, as in ALCE.
Now price confident errors. A penalty of 3 matches "answer only if you are more than 75% confident":
The last line becomes Risk-weighted score (correct +1, abstain 0, wrong -3): -10.
Imagine you improve the prompt so the assistant declines c06 instead of inventing a date. Open unit08/grounded_cases.jsonl, find the c06 line, and change two fields:
"answer": "I can't find when the credit memo will arrive in the retrieved notes.", "labels": []
The accuracy-only score didn't move: 5 of 12 either way. The risk-weighted score improved from -10 to -7, and "Unanswerable questions where it said it can't find it" rose to 2/3. Only a score that prices wrong answers rewards the safer behaviour. This is the argument of "Why Language Models Hallucinate", on your own data.
The model name comes from LLM_MODEL in .env, as in Unit 1, and defaults to claude-opus-5-5. The command sends about 35 short requests.
Compare its caught and false-alarm counts with the rules judge. Look at c08 (negation) and c05 (paraphrase): those are where a model judge should do better. If a verdict changes between two runs, you have seen judge non-determinism.
We could not run Step 7 in our test environment without a key, so we have no sample output for it. A model judge that disagrees with the person is not automatically wrong: read the case, and fix the label if the person was.
On SAP BTP, the answer and its evidence usually come from orchestration with grounding in the generative AI hub. Measuring groundedness there takes three pieces.
In SAP's Python SDK documentation for document grounding, the grounding module takes the user's question as an input parameter and writes the retrieved passages into an output parameter. The prompt template inserts them with a placeholder such as {{ ?grounding_response }}. A search setting max_chunk_count limits how many chunks are retrieved. After the call, the retrieved context is available in the response under module_results.grounding.
# Sketch: needs SAP AI Core (extended plan) with grounding configured as in Unit 7.
# `response` is the orchestration result; `answer_text` is the model's reply from your Unit 5 code.
evidence = response.module_results.grounding.data["grounding_result"] # the passages the model saw
record = {"question": question, "evidence": evidence, "answer": answer_text}
# Store `record` with the same access rules as the documents, then judge claims against `evidence`.
Two practical points. First, judge against what the model actually received, not against a fresh search: indexes change. Second, put a stable document ID in each chunk's metadata, as Evaluating retrieval recommended, so citations can be checked against documents.
Ask for one fact per sentence, a document ID after each sentence, and an explicit "I can't find that in the documents" when the context doesn't answer the question. These are the same steps Anthropic's guidance lists to reduce hallucinations, and they make the output easy to split and check. Test the abstention wording against your split_claims pattern, or your counts will be wrong.
SAP AI Core Evaluations, as of the 4 September 2026 guide, compares orchestration configurations with system-defined metrics or custom LLM-as-a-judge metrics with rating criteria. A groundedness metric can be written as rating criteria, for example:
Rate whether every sentence of the answer is supported by the provided context.
3 = every sentence is supported. 2 = one sentence adds detail not in the context.
1 = a sentence contradicts the context, or most sentences are unsupported.
Use only the context, not general knowledge.
We did not find a system-defined groundedness or faithfulness metric in the SAP documentation we could open; check the current metric list in your tenant before writing your own. A whole-answer rating is coarser than this lab's claim-by-claim check, so calibrate it the same way: run it on your labeled cases and count caught hallucinations and false alarms.
Returned in the orchestration response under the grounding module result
Claim-level groundedness
Sentence or model-based claim splitting, rules, NLI or LLM judge
Custom LLM-as-a-judge metric; claim splitting is up to your criteria
Citation checks
Per-claim and per-citation checks as in ALCE
Not found as a built-in metric; build it
Abstention and false abstention
Purpose-built unanswerable cases, penalty scoring
Write the cases into your evaluation dataset; scoring is yours
Judge calibration
agree against human labels
Run the custom metric on labeled cases and compare
Cost
Rules and HHEM free on a laptop; LLM judge per call
Extended plan of SAP AI Core, plus model calls for judge metrics
The usual split on SAP projects: production answers through orchestration with grounding, groundedness scoring in your own harness, optionally mirrored as a custom judge metric in SAP AI Core Evaluations so the platform team sees the same number. The next topic, building an evaluation harness, combines retrieval scores, groundedness and judge checks into one run.
Authorizations. Evidence and answers contain business data. A groundedness log is as sensitive as the documents behind it. Apply the same access rules, and judge only with a model the data is allowed to go to (Grounding on SAP data without breaking authorizations).
Runtime guard, not only offline test. You can run a cheap judge on every live answer and hide or flag unsupported sentences before the user sees them. Measure its latency and its catch rate first; a guard that misses negations gives false comfort.
Severity by process. A wrong quantity on a disputed invoice is worse than a vague sentence in a how-to answer. Weight test cases or set separate targets by process and claim type (numbers, dates, amounts, IDs).
Pin the judge. If the judge model changes, its verdicts change. Pin the version and rerun agree after any change.
Grow the set from incidents. Every user complaint about a made-up answer becomes a labeled case. That is how the interval narrows where it matters.
Cost. A claim-level LLM judge makes several calls per answer. Use rules or a small NLI model for every answer, and an LLM judge on a sample or on flagged answers.
Clean core. Groundedness checks read answers and evidence. They never need to write to S/4HANA. Any automatic correction stays in the assistant layer.
Scoring whole answers. One wrong number in five correct sentences disappears. Split into claims.
Testing only answerable questions. You will never see the assistant invent an answer it doesn't have. Add unanswerable cases.
Rewarding refusal. A system that always says "I can't find it" has zero hallucinations. Always report false abstentions next to the hallucination rate.
Trusting citations. A cited sentence can point to the wrong passage. Check claims against cited passages separately.
Judging against a fresh search. Re-retrieving at test time checks against evidence the model never saw. Log and use the original passages.
Matching totals, mismatched cases. A judge can reach the same rate as people with different errors. Compare case by case.
Reporting a rate without its range. "50%" from 10 answers and from 1,000 answers are different claims.
Treating outside knowledge as fine because it's true. In an enterprise assistant, an unsupported claim is a defect even when it happens to be correct, because nobody can trace it to company rules.
Make the judge better on its known blind spot, and add cases from your own retrieval work, ready for the evaluation harness later in this unit.
Before you start: complete the Build it yourself steps above, so that unit08/groundedness_lab.py and unit08/grounded_cases.jsonl exist.
Open unit08/groundedness_lab.py and find the line def rules_judge(claim: str, evidence: str) -> str:.
Directly above that line, paste this helper. It finds the evidence sentence closest to the claim and checks whether exactly one of the two contains a negation word:
NEGATIONS = {"not", "cannot", "never", "no"}
def flips_negation(claim: str, evidence: str) -> bool:
"""True if the claim and its closest evidence sentence disagree on not / cannot / never / no."""
def negated(text: str) -> bool:
return bool(NEGATIONS & set(re.findall(r"[a-z]+", text.lower())))
sentences = re.split(r"(?<=[.!?])\s+", evidence)
closest = max(sentences, key=lambda s: len(words(claim) & words(s)))
return negated(claim) != negated(closest)
Inside rules_judge, find the line that starts with coverage = . Directly below it, at the same indentation, add:
if coverage >= 0.75 and flips_negation(claim, evidence):
return "contradicted"
Save the file.
Run:
python unit08/groundedness_lab.py agree
What success looks like (last lines):
Exact agreement: 18/20 = 90%
Hallucinated claims caught: 6/6 (a missed one reaches the user)
False alarms on good claims: 1/14 (each one costs a reviewer's time)
c08 is now caught, and no new false alarm appeared. The c05 paraphrase is still flagged: rules can't fix that one. Note the limit too: this check looks at one sentence and four negation words, so a double negative or "unless" would fool it.
Add four new cases to unit08/grounded_cases.jsonl, IDs c13 to c16, using the note IDs in NOTES: one fully grounded answer, one with a wrong number, one that should abstain and does, and one that gives up although the note has the answer. Give each claim a label.
Run the score with the penalty and save the report:
Open unit08/groundedness_report.md and add three sentences at the end: the hallucination rate with its range, how many hallucinated claims your judge caught, and one control you would put in the product for unsupported claims.
Done whenagree reports 6 of 6 hallucinated claims caught on the original cases, your file has 16 cases that load without errors, and unit08/groundedness_report.md holds the score table with your three sentences.
Pick one answer for each question. The explanation appears after you choose.
1Why does the lab check each claim separately instead of rating the whole answer?
Answer: B. Faithfulness is the share of supported claims, so a single wrong number shows up as a failed claim. A whole-answer rating can average it away, which is how c02's "80 of 100" would slip through.
2In c09, the second sentence is true to the retrieved notes but cites D09, which doesn't mention baseline dates. How does the lab classify it?
Answer: C. Groundedness is checked against all retrieved notes, citations against the cited note only. Keeping them apart tells you whether to fix what the model says or how it points to sources.
3The rules judge marks c08 ("incomplete orders can be billed") as supported, while the note says they cannot. What does this show?
Answer: A. Almost every word matches, so coverage is high, but the meaning is reversed. Only comparing the judge with human labels reveals this blind spot before it reaches users.
4Judge and person both report a 50% hallucination rate on the ten answered cases. Why is that not proof the judge is reliable?
Answer: D. The judge misses c08 and wrongly flags c05, so the totals match by accident. Compare verdicts case by case, as the agree command does, before trusting a judge's headline number.
5You change the prompt so the assistant declines c06 instead of inventing a date. Which number shows the improvement?
Answer: B. Under accuracy-only scoring an abstention and a wrong answer both earn nothing, so the change looks neutral. A penalty for wrong answers, as proposed in "Why Language Models Hallucinate", makes the safer behaviour visible.
6What does a penalty of 3 for wrong answers correspond to?
Answer: C. With a confidence target t, a mistake costs t / (1 - t) points. For t = 0.75 that is 3, so guessing only pays when the system is more than 75% likely to be right.
7On SAP AI Core, what should a groundedness check compare each answer against?
Answer: D. The answer can only be faithful to what the model received. Re-running the search may return different passages if the index changed, so log the grounding result with each answer.
8Your assistant scores zero hallucinations on 200 test questions, but users complain it rarely helps. What do you check first?
Answer: D. A system that declines everything never hallucinates. Reporting false abstentions next to the hallucination rate shows whether safety came at the cost of usefulness.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Faithfulness (Ragas documentation)— claims supported by the retrieved context divided by all claims in the response; 0 to 1; no reference answer needed; FaithfulnessWithHHEM variant uses Vectara's open classifier
Enabling Large Language Models to Generate Text with Citations (Gao et al., EMNLP 2023)— ALCE; citation recall (statement has a citation and the cited passages entail it) and citation precision (a citation is irrelevant if it neither supports alone nor is needed); NLI judge; about half of ELI5 answers from ChatGPT and GPT-4 baselines not fully supported
Reduce hallucinations (Claude Platform Docs)— allow "I don't know", direct quotes, verify with citations and retract unsupported claims, best-of-N consistency, restrict to provided documents; these reduce but don't eliminate hallucinations
Document Grounding (SAP Cloud SDK for AI, Python)— GroundingModule with DocumentGroundingFilter and GroundingFilterSearch(max_chunk_count); output_param inserted with {{ ?grounding_response }}; retrieved context in response.module_results.grounding
SAP AI Core service guide (PDF, 4 September 2026)— Evaluations with system-defined metrics such as ROUGE, BLEU, COMET and tool-calling metrics, or custom LLM-as-a-judge metrics with rating criteria; grounding in the generative AI hub
Evaluating and Testing Your LLM Use Case (SAP Learning)— groundedness and fact-checking metrics called crucial for RAG, verifying the response is supported by retrieved source documents; LLM-as-a-judge and rule-based checks