SAP Knowledge Graph and Business Data Cloud: giving AI a map of the business and data it can trust
What SAP Knowledge Graph and SAP Business Data Cloud add to grounding, how they fit together, and how to build a small business knowledge graph yourself.
Earlier topics in this unit grounded AI answers on documents and records. That works when the answer sits in one place. Many business questions don't work that way.
"Why is order 5004 blocked?" needs the order, its customer, the customer's credit limit, their unpaid invoices and their other open orders. Those facts live in different tables, sometimes in different systems. A search that returns "the three most similar texts" often brings back the order and misses the rest.
SAP offers two building blocks for this:
SAP Knowledge Graph is a map of the business. It records what things are (a sales order, a customer, a field called SoldToParty) and how they connect (an order is sold to a customer). AI uses the map to understand a question and find the right data.
SAP Business Data Cloud (BDC) is SAP's managed data platform. It delivers data products: ready-made, described and governed sets of business data, such as sales orders from S/4HANA, that other tools and AI can read without building their own extracts.
Put simply: Business Data Cloud supplies trustworthy data, and the Knowledge Graph supplies its meaning and connections.
Better answers to questions that cross objects. Credit, delivery and payment questions in order-to-cash almost always span several business objects. A map of relationships lets the assistant collect all the needed facts, not just the closest text.
Fewer wrong guesses about what data means. S/4HANA is huge. SAP Learning describes its metadata as about 452,000 ABAP tables, 80,000 CDS views and 7.3 million fields. A language model alone can't know which field holds a delivery block. A graph that links business words to technical fields reduces those guesses.
Less extraction work. Many AI projects start by building custom extracts from SAP. Data products in BDC are meant to replace much of that with packaged, documented datasets. That moves effort from plumbing to the use case.
The cost is real. BDC is a paid subscription. The full Knowledge Graph capability in SAP HANA Cloud is not available on the free trial or free tier. Graph modelling also takes skilled people. Leaders should expect a data workstream next to the AI workstream.
A concrete example. A credit analyst asks, "Which blocked orders belong to customers with overdue invoices?" With a graph over sales orders, customers and receivables, the system follows the links and returns the exact list, each with the overdue amount. A plain text search returns whichever records share the most words with the question.
As of October 2026, SAP describes these offerings. Names and scope change often, so confirm details with SAP before you plan.
SAP Knowledge Graph. SAP calls it the semantic backbone of its AI architecture. It links business language to SAP metadata: APIs, business semantics, data product descriptions and customer extensions. SAP describes it grounding Joule and Joule agents, and helping agents pick the right API and filters for a request such as "show me overdue orders".
SAP HANA Cloud knowledge graph engine. A database feature that stores graphs in the open RDF standard and queries them with SPARQL. Your own team can use it to build a graph over your data. SAP's tutorial states it is not available for trial or free tier users.
SAP Business Data Cloud. A fully managed SaaS offering that combines SAP Datasphere, SAP Analytics Cloud, SAP Business Warehouse, SAP Databricks and intelligent applications. Its data products are self-describing, read-only datasets. SAP-managed data products are filled from S/4HANA by SAP's own pipelines.
BDC Connect. Shares data products with partner platforms without copying them. SAP named Databricks, Snowflake, Google BigQuery and Microsoft Fabric in May 2026, with Amazon Athena planned for the second half of 2026.
The link between them. In May 2026 SAP announced Joule agents in BDC that generate business insights using SAP Knowledge Graph, and SAP HANA Cloud as BDC's database for graph, vector and other workloads.
You want Joule to answer well on standard SAP data
Joule with SAP Knowledge Graph as SAP delivers it
SAP builds and maintains this map; you configure, not model
AI or analytics needs clean S/4HANA data outside S/4HANA
SAP Business Data Cloud data products
Packaged, described datasets replace custom extracts
Your data science team works in Databricks or Snowflake
BDC Connect
Shares data products without another copy to govern
Your own app must answer multi-hop questions on your data
A knowledge graph you build (for example in SAP HANA Cloud)
You control the model of your business and its rules
Questions are about documents and policies, not records
Document retrieval from earlier Unit 7 topics
A graph adds little when the answer is in one passage
A pilot with a small budget
A small graph on sample data, as in the deep layer
Proves the value before any subscription decision
Most programs end up with more than one row. A common pattern is BDC for the data, a graph for meaning and links, and document retrieval for policy text.
"A knowledge graph replaces our data platform." It describes and connects data. It still needs reliable data underneath, which is the job of BDC or another platform.
"BDC is just a new name for SAP Datasphere." Datasphere is one component. BDC also includes SAP Analytics Cloud, SAP Business Warehouse, SAP Databricks, SAP-managed data products and intelligent applications.
"If it's in BDC, SAP authorizations come with it." A copy or share of data has its own access rules. Ask how the S/4HANA rules are reflected, as covered in the previous topic on grounding with authorizations.
"Graphs make the model stop hallucinating." They reduce one cause: missing or misread facts. The model can still misstate the facts it was given, so answers still need evaluation.
"We can try the graph engine on the free tier." SAP states the HANA Cloud knowledge graph feature is not available on trial or free tier. The BDC trial is a separate, shared 30-day tenant.
Pick one answer for each question. The explanation appears after you choose.
1Why do questions like "Why is order 5004 blocked?" often defeat plain text search?
Answer: B. The answer needs the order, its customer, the credit limit, open invoices and other orders. Text search returns the most similar records, which often means the order alone. A graph follows the links between them.
2Which description fits SAP Knowledge Graph best?
Answer: C. SAP describes it as a semantic backbone that links business words to metadata such as APIs, fields and data products. It helps AI understand questions and find the right data. It doesn't store or replace the business data itself.
3What does SAP Business Data Cloud add for an AI program?
Answer: D. BDC delivers data products: described, read-only datasets such as S/4HANA sales orders, filled by SAP pipelines. That cuts the custom extraction work many AI projects start with. Access rules still need design.
4Your data science team works in Snowflake and needs S/4HANA sales data. What should you ask about first?
Answer: A. BDC Connect shares data products with partner platforms such as Snowflake without copying them. That avoids another extract to build and govern. Manual exports create exactly the copies BDC is meant to avoid.
5A team plans to prototype the HANA Cloud knowledge graph engine on the free tier. What is the risk?
Answer: B. SAP's tutorial says the knowledge graph feature is not available for trial or free tier users. A prototype needs a paid HANA Cloud instance, or an open-source library on sample data as in the deep layer.
6Data is now shared through BDC. What should a leader still ask about access?
Answer: C. A copy or share of data has its own access rules. Someone must show how the S/4HANA rules are carried over and tested. Instructions to the model don't enforce access.
7Which statement about graphs and hallucinations is accurate?
Answer: D. A graph supplies the right linked facts, which removes one common cause of wrong answers. The model can still misstate what it was given. That is why the evaluation work in Unit 8 still matters.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Retrieval finds things that look like the question. A graph finds things that are connected to the answer.
Every earlier Unit 7 topic ranked pieces of text or records by similarity, by keywords, or both. That ranking is the right tool when one passage holds the answer. It breaks down when the answer is a chain: order to customer, customer to open items, customer to other orders. The links in that chain are often missing from the question. "Why is order 5004 blocked?" never mentions invoice 9001, yet invoice 9001 is part of the reason.
A knowledge graph stores those links explicitly, so you can follow them instead of guessing them. It has three layers, and SAP's offerings map onto them:
Layer
Holds
Example
SAP offering
Metadata
What data exists, where, and what each field means
DeliveryBlockReason in the sales order data product is "delivery block reason", also called "on hold"
SAP Knowledge Graph (metadata, APIs, data product descriptions)
Ontology
Business concepts and allowed relations
A sales order is sold to a customer
SAP Knowledge Graph; or your own ontology
Records
The facts themselves, linked
Order 5004 is sold to customer 10100001
Data products in SAP Business Data Cloud; a graph in SAP HANA Cloud
Keep one sentence in mind for the rest of the topic: Business Data Cloud is where trustworthy data comes from; the knowledge graph is how an AI system knows what that data means and how it connects.
The W3C standard for this is RDF (Resource Description Framework). Every thing and every relation gets a URI, a globally unique name that looks like a web address. A short prefix such as orch: stands for the long part. Values such as amounts and dates are literals with a type, for example "2026-08-15"^^xsd:date.
Because the predicate is itself a named thing, you can describe it too: orch:soldTo has the label "is sold to", a domain orch:SalesOrder and a range orch:Customer. That description is the ontology. It is what lets a program, or a model, know which links make sense.
SPARQL is the W3C query language for RDF. A query is a pattern of triples with variables (written ?name), and the engine returns every way the pattern matches:
Read it as: "find orders with a block reason, their customer, and any open item owed by that customer that was due before 1 October". The shared variable ?cis the join. You don't name join keys, because the link is already a fact in the graph. SPARQL also has OPTIONAL (keep the row even if part of the pattern is missing), aggregates such as SUM and GROUP_CONCAT, and property paths such as orch:inSalesOrg/orch:id, which follow two links in one step.
A graph-backed assistant does four things. The deep difference from vector RAG is in steps 1 and 2.
flowchart LR
Q[Question + user] --> L[1 Link words<br/>to graph terms]
L --> P[2 Pick a query<br/>pattern]
P --> A[3 Add the user's<br/>access filter]
A --> G[(Knowledge graph)]
G --> F[4 Facts as numbered<br/>context]
F --> M[Model writes answer<br/>with citations]
Link words in the question to graph terms. "Order 5004" maps to a sales order; "overdue" maps to the due date field, through the metadata layer's labels and synonyms. This is the job SAP describes for SAP Knowledge Graph when it finds the right filter for "show me overdue orders".
Pick a query pattern. Safer systems use a small set of tested, parameterized queries (templates). More flexible systems let a model write SPARQL or SQL, then validate it before running. The template approach is easier to secure and evaluate.
Add the access filter. The user's authorizations become part of the query, for example "only sales organizations this user may display". The previous topic explained why this has to happen before ranking or generation, never after.
Turn rows into facts the model can cite, then generate, exactly as in RAG fundamentals.
Graph retrieval and vector retrieval combine well. SAP Learning describes SAP Knowledge Graph as combining GraphRAG with vector-based retrieval over data in SAP HANA and SAP Datasphere. A common shape: vector search finds the right entity or document; the graph expands it to its neighbours; the model answers from both.
A graph is only as good as the data under it. Building it from hand-made extracts recreates every problem of point-to-point integration. SAP Business Data Cloud addresses this with data products:
What they are. SAP Learning describes them as predefined, self-describing datasets combining business data and metadata, built for large-scale, read-only analytical use. SAP's Architecture Center adds clear ownership, schema, authorization rules and lifecycle management.
How SAP-managed ones are filled. BDC's foundation services replicate data from S/4HANA, using ABAP CDS views to harmonize and join source tables, with initial and delta loads. The data lands as files in the SAP HANA Cloud data lake object store.
How they are shared. Through the open Delta Sharing protocol, which gives other tools access without copying. Customers install data products into SAP Datasphere or SAP Databricks; BDC Connect extends sharing to partner platforms.
How they describe themselves. With two open specifications. Open Resource Discovery (ORD) describes the data product, its output ports (Delta Sharing APIs), input ports and entity types. CSN Interop describes each entity and its elements (fields), with annotations for semantics.
That self-description is exactly the metadata layer of a knowledge graph. When SAP says its Knowledge Graph connects data product metadata with API metadata and business semantics, this is the material it works from.
flowchart LR
S4[S/4HANA<br/>CDS views] --> FS[BDC foundation<br/>services]
FS --> DP[Data products<br/>files + ORD/CSN metadata]
DP -->|Delta Sharing| DS[Datasphere /<br/>SAP Databricks / partners]
DP -->|metadata| KG[Knowledge graph<br/>meaning + links]
DS --> KG
KG --> AI[Joule agents /<br/>your AI app]
#Build it yourself: a business knowledge graph for grounding
You will build business_graph.py. It creates a small knowledge graph with all three layers: metadata for three made-up data products, an ontology for order-to-cash, and SAP-shaped records for customers, sales orders and receivables. You will search the metadata like an analyst, answer multi-hop questions with SPARQL for users with different authorizations, and measure how many needed facts a graph finds compared with keyword retrieval. The comparison report feeds Unit 8, where you evaluate retrieval properly.
flowchart LR
B[build] --> T[(business_graph.ttl)]
T --> F[find: search metadata]
T --> A[ask: multi-hop SPARQL<br/>+ access filter]
T --> C[compare: graph vs. keyword]
T --> S[sparql: your own query]
C --> R[graph_compare_report.txt<br/>for Unit 8]
The data, all made up:
Record
Details
Customer 10100001, Made-up Retail GmbH
Sales org 1010, credit limit 40,000 EUR; open items 9001 (18,000, overdue) and 9002 (7,000, not yet due); orders 5001 and 5004, both blocked
Customer 10100002, Sample Foods AG
Sales org 1010, credit limit 20,000 EUR; open item 9004; orders 5002 (not blocked) and 5005 (blocked, pricing)
Customer 17100001, Example Distribution Inc.
Sales org 1710, credit limit 80,000 USD; open item 9003 (52,000, overdue); orders 5003 (blocked) and 5006
The users are the same as in the previous topic: anna may display sales org 1010, ben may display 1010 and 1710, and carl has no roles yet.
The block reason codes Z1 and Z2 are lab codes. Real delivery block reasons are configured in each SAP system, so never hard-code a code's meaning from one system into another.
Your course folder with the Unit 7 setup done. Cost: free.
One new library, rdflib (free, open source). python-dotenv (from the course setup) and gen_ai_hub (from Unit 5) are needed only for --llm MODEL.
About 45 minutes. No account and no network after installing rdflib, unless you use --llm MODEL (small per-request charge on your SAP AI Core account).
In VS Code, right-click the unit07 folder, choose New File, name it business_graph.py, paste the code below and save.
"""business_graph.py: a tiny business knowledge graph for grounding AI answers.
Orchestrate course, Unit 7. Builds an RDF graph with three layers (data product
metadata, business ontology, sample records), searches the metadata, answers
multi-hop questions with SPARQL, and compares that with a keyword retriever.
Commands:
python unit07/business_graph.py build
python unit07/business_graph.py find "credit hold"
python unit07/business_graph.py ask "Why is order 5004 blocked?" --user anna
python unit07/business_graph.py compare
python unit07/business_graph.py sparql "SELECT ?o WHERE { ?o a orch:SalesOrder }"
Add --llm sample (no account) or --llm MODEL (SAP orchestration, Unit 5 setup) to "ask".
All data is made up. Field names marked "SAP API field" follow the S/4HANA sales order API.
"""
import argparse
import os
import re
import sys
from pathlib import Path
try:
from rdflib import RDF, RDFS, Graph, Literal, Namespace
from rdflib.namespace import XSD
except ImportError:
sys.exit("rdflib is not installed. Run: pip install rdflib (then add it to requirements.txt)")
HERE = Path(__file__).resolve().parent
TTL_FILE = HERE / "business_graph.ttl"
REPORT_FILE = HERE / "graph_compare_report.txt"
TODAY = "2026-10-01" # fixed "today" so the overdue logic gives the same result every run
ORCH = Namespace("https://orchestrate.example/ontology#") # classes and relations
META = Namespace("https://orchestrate.example/meta#") # data products, entities, fields
DATA = Namespace("https://orchestrate.example/data/") # the records themselves
PREFIXES = {"orch": ORCH, "meta": META, "data": DATA, "rdfs": RDFS, "xsd": XSD}
# ---------------------------------------------------------------------------
# 1. Sample data (made up, SAP-shaped)
# ---------------------------------------------------------------------------
CUSTOMERS = [ # Customer, name, sales org, credit limit, currency
("10100001", "Made-up Retail GmbH", "1010", 40000, "EUR"),
("10100002", "Sample Foods AG", "1010", 20000, "EUR"),
("17100001", "Example Distribution Inc.", "1710", 80000, "USD"),
]
ORDERS = [ # SalesOrder, SoldToParty, SalesOrganization, TotalNetAmount, currency, DeliveryBlockReason
("5001", "10100001", "1010", 12500, "EUR", "Z1"),
("5002", "10100002", "1010", 4200, "EUR", ""),
("5003", "17100001", "1710", 38000, "USD", "Z1"),
("5004", "10100001", "1010", 9800, "EUR", "Z1"),
("5005", "10100002", "1010", 3100, "EUR", "Z2"),
("5006", "17100001", "1710", 15000, "USD", ""),
]
BLOCK_CODES = { # lab codes: real block reason codes are configured per system
"Z1": "Credit limit exceeded (lab code)",
"Z2": "Pricing incomplete (lab code)",
}
OPEN_ITEMS = [ # item, customer, amount, currency, due date
("9001", "10100001", 18000, "EUR", "2026-08-15"),
("9002", "10100001", 7000, "EUR", "2026-10-20"),
("9003", "17100001", 52000, "USD", "2026-09-01"),
("9004", "10100002", 1200, "EUR", "2026-10-10"),
]
USERS = {"anna": {"1010"}, "ben": {"1010", "1710"}, "carl": set()} # sales orgs each user may display
# Metadata layer: data products -> entities -> fields, with labels and synonyms.
DATA_PRODUCTS = {
"SalesOrderLab": ("Sales orders (lab data product)", "SalesOrder", [
("SalesOrder", "Sales order number", "SAP API field", ["order", "order number"]),
("SoldToParty", "Sold-to party", "SAP API field", ["customer", "buyer"]),
("SalesOrganization", "Sales organization", "SAP API field", ["sales org"]),
("TotalNetAmount", "Net value of the order", "SAP API field", ["order value", "amount"]),
("TransactionCurrency", "Document currency", "SAP API field", ["currency"]),
("DeliveryBlockReason", "Delivery block reason", "SAP API field",
["blocked", "block", "credit hold", "on hold"]),
]),
"CustomerLab": ("Customers (lab data product)", "Customer", [
("Customer", "Customer number", "lab field", ["customer id"]),
("CustomerName", "Customer name", "lab field", ["name"]),
("CreditLimit", "Credit limit", "lab field", ["credit line", "credit hold", "limit"]),
]),
"OpenItemLab": ("Receivables open items (lab data product)", "OpenItem", [
("OpenItem", "Open item number", "lab field", ["invoice", "receivable"]),
("DueDate", "Net due date", "lab field", ["overdue", "late payment", "due"]),
("Amount", "Open amount", "lab field", ["outstanding", "exposure"]),
]),
}
# ---------------------------------------------------------------------------
# 2. Build the graph
# ---------------------------------------------------------------------------
def new_graph() -> Graph:
g = Graph()
for prefix, ns in PREFIXES.items():
g.bind(prefix, ns)
return g
def build_graph() -> tuple[Graph, dict]:
g = new_graph()
counts = {}
# Layer 1: metadata (what data exists, where, and what the fields mean)
before = len(g)
for dp_id, (dp_label, entity, fields) in DATA_PRODUCTS.items():
dp, ent = META[dp_id], META[entity]
g.add((dp, RDF.type, META.DataProduct))
g.add((dp, RDFS.label, Literal(dp_label)))
g.add((dp, META.providesEntity, ent))
g.add((ent, RDF.type, META.Entity))
g.add((ent, META.describesClass, ORCH[entity]))
for name, label, origin, synonyms in fields:
f = META[f"{entity}.{name}"]
g.add((f, RDF.type, META.Field))
g.add((f, META.fieldOf, ent))
g.add((f, META.technicalName, Literal(name)))
g.add((f, RDFS.label, Literal(label)))
g.add((f, META.origin, Literal(origin)))
for s in synonyms:
g.add((f, META.synonym, Literal(s)))
counts["metadata"] = len(g) - before
# Layer 2: ontology (the business meaning: classes and how they relate)
before = len(g)
for cls, label in [("SalesOrder", "Sales order"), ("Customer", "Customer"),
("OpenItem", "Receivables open item"), ("SalesOrg", "Sales organization")]:
g.add((ORCH[cls], RDF.type, RDFS.Class))
g.add((ORCH[cls], RDFS.label, Literal(label)))
for prop, dom, rng, label in [
("soldTo", "SalesOrder", "Customer", "is sold to"),
("inSalesOrg", "SalesOrder", "SalesOrg", "belongs to sales organization"),
("customerSalesOrg", "Customer", "SalesOrg", "is served by sales organization"),
("owedBy", "OpenItem", "Customer", "is owed by"),
]:
g.add((ORCH[prop], RDF.type, RDF.Property))
g.add((ORCH[prop], RDFS.domain, ORCH[dom]))
g.add((ORCH[prop], RDFS.range, ORCH[rng]))
g.add((ORCH[prop], RDFS.label, Literal(label)))
counts["ontology"] = len(g) - before
# Layer 3: instance data (the records, linked by the relations above)
before = len(g)
for cid, name, org, limit, cur in CUSTOMERS:
c = DATA[f"customer/{cid}"]
g.add((c, RDF.type, ORCH.Customer))
g.add((c, ORCH.id, Literal(cid)))
g.add((c, RDFS.label, Literal(name)))
g.add((c, ORCH.customerSalesOrg, DATA[f"salesorg/{org}"]))
g.add((c, ORCH.creditLimit, Literal(limit, datatype=XSD.decimal)))
g.add((c, ORCH.currency, Literal(cur)))
for org in sorted({o[2] for o in ORDERS}):
g.add((DATA[f"salesorg/{org}"], RDF.type, ORCH.SalesOrg))
g.add((DATA[f"salesorg/{org}"], ORCH.id, Literal(org)))
for so, cust, org, amount, cur, block in ORDERS:
o = DATA[f"salesorder/{so}"]
g.add((o, RDF.type, ORCH.SalesOrder))
g.add((o, ORCH.id, Literal(so)))
g.add((o, ORCH.soldTo, DATA[f"customer/{cust}"]))
g.add((o, ORCH.inSalesOrg, DATA[f"salesorg/{org}"]))
g.add((o, ORCH.netAmount, Literal(amount, datatype=XSD.decimal)))
g.add((o, ORCH.currency, Literal(cur)))
if block:
g.add((o, ORCH.deliveryBlockReason, Literal(block)))
g.add((o, ORCH.blockText, Literal(BLOCK_CODES[block])))
for item, cust, amount, cur, due in OPEN_ITEMS:
i = DATA[f"openitem/{item}"]
g.add((i, RDF.type, ORCH.OpenItem))
g.add((i, ORCH.id, Literal(item)))
g.add((i, ORCH.owedBy, DATA[f"customer/{cust}"]))
g.add((i, ORCH.amount, Literal(amount, datatype=XSD.decimal)))
g.add((i, ORCH.currency, Literal(cur)))
g.add((i, ORCH.dueDate, Literal(due, datatype=XSD.date)))
counts["records"] = len(g) - before
return g, counts
def load_graph() -> Graph:
if not TTL_FILE.exists():
sys.exit(f"{TTL_FILE.name} not found. Run first: python unit07/business_graph.py build")
g = new_graph()
g.parse(TTL_FILE, format="turtle")
return g
# ---------------------------------------------------------------------------
# 3. Queries (SPARQL, with values bound safely through initBindings)
# ---------------------------------------------------------------------------
Q_FIND_FIELD = """
SELECT ?dpLabel ?entity ?name ?label ?origin WHERE {
?f a meta:Field ; meta:technicalName ?name ; rdfs:label ?label ; meta:origin ?origin ;
meta:fieldOf ?ent .
?dp meta:providesEntity ?ent ; rdfs:label ?dpLabel .
BIND(STRAFTER(STR(?ent), "#") AS ?entity)
OPTIONAL { ?f meta:synonym ?syn }
FILTER(CONTAINS(LCASE(?label), ?term) || CONTAINS(LCASE(STR(?syn)), ?term)
|| CONTAINS(LCASE(?name), ?term))
} GROUP BY ?dpLabel ?entity ?name ?label ?origin ORDER BY ?entity ?name
"""
# One order, its customer, the customer's open items and other open orders: three hops.
Q_ORDER_CONTEXT = """
SELECT ?org ?amount ?cur ?block ?blockText ?cid ?cname ?limit
?item ?itemAmount ?due ?other ?otherAmount ?otherBlock WHERE {
?o a orch:SalesOrder ; orch:id ?orderId ; orch:inSalesOrg/orch:id ?org ;
orch:netAmount ?amount ; orch:currency ?cur ; orch:soldTo ?c .
OPTIONAL { ?o orch:deliveryBlockReason ?block ; orch:blockText ?blockText }
?c orch:id ?cid ; rdfs:label ?cname ; orch:creditLimit ?limit .
OPTIONAL { ?i orch:owedBy ?c ; orch:id ?item ; orch:amount ?itemAmount ; orch:dueDate ?due }
OPTIONAL { ?o2 orch:soldTo ?c ; orch:id ?other ; orch:netAmount ?otherAmount .
FILTER(?o2 != ?o)
OPTIONAL { ?o2 orch:deliveryBlockReason ?otherBlock } }
FILTER(?org IN (%ALLOWED%))
}
"""
# Portfolio question: blocked orders whose customer has at least one overdue open item.
Q_BLOCKED_WITH_OVERDUE = """
SELECT ?orderId ?org ?cname ?cur (SUM(?amt) AS ?overdue) (MIN(?due) AS ?oldest)
(GROUP_CONCAT(?itemId; separator=", ") AS ?items) WHERE {
?o a orch:SalesOrder ; orch:id ?orderId ; orch:deliveryBlockReason ?b ;
orch:inSalesOrg/orch:id ?org ; orch:soldTo ?c .
?c rdfs:label ?cname ; orch:currency ?cur .
?i orch:owedBy ?c ; orch:id ?itemId ; orch:amount ?amt ; orch:dueDate ?due .
FILTER(?due < ?today)
FILTER(?org IN (%ALLOWED%))
} GROUP BY ?orderId ?org ?cname ?cur ORDER BY ?orderId
"""
def num(value) -> str:
"""Format an amount from the graph: 12500.0 -> 12,500."""
return f"{float(value):,.0f}"
def allowed_filter(user: str) -> str:
"""Turn the user's sales orgs into a SPARQL IN list. No orgs means nothing matches."""
if user not in USERS:
sys.exit(f"Unknown user '{user}'. Choose one of: {', '.join(USERS)}")
orgs = sorted(USERS[user])
return ", ".join(f'"{o}"' for o in orgs) if orgs else '"-none-"'
def order_context(g: Graph, order_id: str, user: str) -> list[str]:
q = Q_ORDER_CONTEXT.replace("%ALLOWED%", allowed_filter(user))
rows = list(g.query(q, initNs=PREFIXES, initBindings={"orderId": Literal(order_id)}))
if not rows:
return []
r = rows[0]
facts = [f"Order {order_id}: sales org {r.org}, net value {num(r.amount)} {r.cur}, "
f"sold to {r.cname} ({r.cid})."]
facts.append(f"Order {order_id} delivery block: {r.blockText} [code {r.block}]."
if r.block else f"Order {order_id} has no delivery block.")
facts.append(f"Customer {r.cid} credit limit: {num(r.limit)} {r.cur}.")
items = {(str(x.item), str(x.itemAmount), str(x.due)) for x in rows if x.item is not None}
open_total = 0
for item, amt, due in sorted(items):
status = "overdue" if due < TODAY else "not yet due"
facts.append(f"Open item {item} for customer {r.cid}: {num(amt)} {r.cur}, due {due} ({status}).")
open_total += float(amt)
others = {(str(x.other), str(x.otherAmount), str(x.otherBlock or "")) for x in rows
if x.other is not None}
order_total = float(r.amount)
for other, amt, blk in sorted(others):
facts.append(f"Other open order {other} for customer {r.cid}: {num(amt)} {r.cur}"
+ (f", also blocked [code {blk}]." if blk else ", not blocked."))
order_total += float(amt)
exposure = open_total + order_total
facts.append(f"Computed exposure for customer {r.cid}: open items {num(open_total)} + open orders "
f"{num(order_total)} = {num(exposure)} {r.cur} against a limit of {num(r.limit)} {r.cur}.")
return facts
def blocked_with_overdue(g: Graph, user: str) -> list[str]:
q = Q_BLOCKED_WITH_OVERDUE.replace("%ALLOWED%", allowed_filter(user))
rows = g.query(q, initNs=PREFIXES, initBindings={"today": Literal(TODAY, datatype=XSD.date)})
return [f"Blocked order {r.orderId} (sales org {r.org}, {r.cname}): customer has {num(r.overdue)} "
f"{r.cur} overdue in open item {r.items}, oldest due {r.oldest}." for r in rows]
# ---------------------------------------------------------------------------
# 4. A keyword retriever over flattened records, for comparison
# ---------------------------------------------------------------------------
def flat_records() -> list[tuple[str, str]]:
names = {c[0]: c[1] for c in CUSTOMERS}
recs = []
for cid, name, org, limit, cur in CUSTOMERS:
recs.append((f"customer/{cid}", f"Customer {cid} {name} sales org {org} credit limit {limit} {cur}"))
for so, cust, org, amount, cur, block in ORDERS:
txt = f"Sales order {so} sold to {cust} {names[cust]} sales org {org} net value {amount} {cur}"
recs.append((f"salesorder/{so}", txt + (f" blocked {BLOCK_CODES[block]}" if block else "")))
for item, cust, amount, cur, due in OPEN_ITEMS:
recs.append((f"openitem/{item}", f"Open item {item} customer {cust} amount {amount} {cur} due {due}"))
return recs
def keyword_top_k(question: str, k: int = 3) -> list[str]:
words = set(re.findall(r"[a-z0-9]+", question.lower())) - {"the", "is", "why", "and", "which", "with", "to"}
scored = []
for rid, txt in flat_records():
overlap = len(words & set(re.findall(r"[a-z0-9]+", txt.lower())))
scored.append((overlap, rid))
scored.sort(key=lambda x: (-x[0], x[1]))
return [rid for score, rid in scored[:k] if score > 0]
def graph_ids(facts: list[str]) -> set[str]:
ids = set()
for f in facts:
ids |= {f"salesorder/{m}" for m in re.findall(r"[Oo]rder (\d{4})", f)}
ids |= {f"openitem/{m}" for m in re.findall(r"[Oo]pen item (\d{4})", f)}
ids |= {f"customer/{m}" for m in re.findall(r"\b(\d{8})\b", f)}
return ids
# Each question lists the records a complete answer needs (the "gold" set).
COMPARE_SET = [
("Why is order 5004 blocked?", "5004",
{"salesorder/5004", "customer/10100001", "openitem/9001", "openitem/9002", "salesorder/5001"}),
("Why is order 5003 blocked?", "5003",
{"salesorder/5003", "customer/17100001", "openitem/9003", "salesorder/5006"}),
("Why is order 5005 blocked?", "5005",
{"salesorder/5005", "customer/10100002", "openitem/9004", "salesorder/5002"}),
("Which blocked orders belong to customers with overdue invoices?", None,
{"salesorder/5001", "salesorder/5004", "salesorder/5003", "openitem/9001", "openitem/9003"}),
]
# ---------------------------------------------------------------------------
# 5. Optional model call (same pattern as earlier Unit 7 topics)
# ---------------------------------------------------------------------------
SYSTEM = ("You explain SAP sales order situations to a business user. Use only the numbered facts. "
"Cite facts like [2]. If the facts don't answer the question, say so. Never invent numbers.")
def generate(model: str, question: str, facts: list[str]) -> str:
if model == "sample":
if not facts:
return "[sample answer, no model called] I can't find that in the data you may see."
blocked = next((f for f in facts if "delivery block:" in f), facts[0])
return f"[sample answer, no model called] {blocked.split(' [code')[0]}. See the exposure fact [{len(facts)}]."
from dotenv import load_dotenv
load_dotenv()
missing = [n for n in ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL",
"AICORE_BASE_URL", "AICORE_RESOURCE_GROUP"] if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5'. Or use --llm sample.")
from gen_ai_hub.orchestration_v2 import (LLMModelDetails, ModuleConfig, OrchestrationConfig,
OrchestrationService, PromptTemplatingModuleConfig,
SystemMessage, Template, UserMessage)
context = "\n".join(f"[{n}] {f}" for n, f in enumerate(facts, 1)) or "(no facts)"
template = Template(template=[SystemMessage(content=SYSTEM),
UserMessage(content="Facts:\n{{?context}}\n\nQuestion: {{?question}}")])
config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=model, timeout=60, max_retries=1))))
service = None
try:
service = OrchestrationService(config=config)
result = service.run(placeholder_values={"context": context, "question": question})
except Exception as error:
sys.exit(f"The model call failed: {type(error).__name__}: {str(error)[:300]}")
finally:
if service is not None:
service.close_http_connection()
return result.final_result.choices[0].message.content or ""
# ---------------------------------------------------------------------------
# 6. Commands
# ---------------------------------------------------------------------------
def cmd_build(args) -> None:
g, counts = build_graph()
g.serialize(destination=TTL_FILE, format="turtle")
print(f"Wrote {TTL_FILE.name}: {len(g)} triples")
for layer, n in counts.items():
print(f" {layer:<9}{n:>4} triples")
def cmd_find(args) -> None:
g = load_graph()
rows = list(g.query(Q_FIND_FIELD, initNs=PREFIXES,
initBindings={"term": Literal(args.term.lower())}))
print(f'Fields matching "{args.term}":')
for r in rows:
print(f" {r.dpLabel} > {r.entity}.{r.name} ({r.label}; {r.origin})")
print(f"{len(rows)} field(s) found.")
def cmd_ask(args) -> None:
g = load_graph()
match = re.search(r"\border (\d{4})\b", args.question.lower())
if match:
facts = order_context(g, match.group(1), args.user)
route = f"order context for {match.group(1)} (3 hops: order > customer > open items, other orders)"
elif "overdue" in args.question.lower() and "blocked" in args.question.lower():
facts = blocked_with_overdue(g, args.user)
route = "blocked orders whose customer has overdue items (join + aggregate)"
else:
sys.exit("This lab understands questions about one order ('order 5004') or "
"'blocked orders ... overdue'. Try the commands in the topic.")
print(f"User: {args.user} (sales orgs: {', '.join(sorted(USERS[args.user])) or 'none'})")
print(f"Route: {route}")
print("Facts from the graph:")
for n, f in enumerate(facts, 1):
print(f" [{n}] {f}")
if not facts:
print(" (none: no record exists, or none you are allowed to see)")
if args.llm:
print("\nAnswer:\n" + generate(args.llm, args.question, facts))
def cmd_compare(args) -> None:
g = load_graph()
lines = [f"Graph vs. keyword top-3, user ben (all sales orgs), today = {TODAY}", ""]
totals = {"keyword": 0, "graph": 0}
gold_total = 0
for question, order_id, gold in COMPARE_SET:
kw = set(keyword_top_k(question))
facts = order_context(g, order_id, "ben") if order_id else blocked_with_overdue(g, "ben")
gr = graph_ids(facts)
k_hit, g_hit = len(kw & gold), len(gr & gold)
totals["keyword"] += k_hit
totals["graph"] += g_hit
gold_total += len(gold)
lines.append(question)
lines.append(f" needed {len(gold)} records | keyword found {k_hit} | graph found {g_hit}")
missed = sorted(gold - kw)
if missed:
lines.append(f" keyword missed: {', '.join(missed)}")
lines.append("")
lines.append(f"Coverage: keyword {totals['keyword']}/{gold_total} = {totals['keyword'] / gold_total:.0%}, "
f"graph {totals['graph']}/{gold_total} = {totals['graph'] / gold_total:.0%}")
text = "\n".join(lines)
print(text)
REPORT_FILE.write_text(text + "\n", encoding="utf-8")
print(f"\nSaved {REPORT_FILE.name}")
def cmd_sparql(args) -> None:
if args.hana_sql:
prefixes = "".join(f"PREFIX {p}: <{ns}> " for p, ns in PREFIXES.items())
body = (prefixes + args.query).replace("'", "''")
print("-- Sketch: the same query in SAP HANA Cloud (knowledge graph engine enabled,")
print("-- graph loaded into the default graph). Not runnable on the trial or free tier.")
print(f"SELECT * FROM SPARQL_TABLE('{body}');")
return
g = load_graph()
try:
result = g.query(args.query, initNs=PREFIXES)
except Exception as error:
sys.exit(f"SPARQL error: {error}")
rows = list(result)
for r in rows:
print(" " + " | ".join(str(v) for v in r))
print(f"{len(rows)} row(s).")
def main() -> None:
p = argparse.ArgumentParser(description="A tiny business knowledge graph (Orchestrate Unit 7)")
sub = p.add_subparsers(dest="cmd", required=True)
sub.add_parser("build", help="build the graph and write business_graph.ttl").set_defaults(fn=cmd_build)
f = sub.add_parser("find", help="search the metadata layer for a field")
f.add_argument("term")
f.set_defaults(fn=cmd_find)
a = sub.add_parser("ask", help="answer a question from the graph")
a.add_argument("question")
a.add_argument("--user", default="ben", help="anna, ben or carl")
a.add_argument("--llm", metavar="MODEL", help="'sample' (no account) or a model name")
a.set_defaults(fn=cmd_ask)
sub.add_parser("compare", help="graph vs. keyword retrieval on four questions").set_defaults(fn=cmd_compare)
s = sub.add_parser("sparql", help="run your own SPARQL query")
s.add_argument("query")
s.add_argument("--hana-sql", action="store_true", help="print the HANA Cloud SQL wrapper instead")
s.set_defaults(fn=cmd_sparql)
args = p.parse_args()
args.fn(args)
if __name__ == "__main__":
main()
Fields matching "credit hold":
Customers (lab data product) > Customer.CreditLimit (Credit limit; lab field)
Sales orders (lab data product) > SalesOrder.DeliveryBlockReason (Delivery block reason; SAP API field)
2 field(s) found.
Fields matching "net value":
Sales orders (lab data product) > SalesOrder.TotalNetAmount (Net value of the order; SAP API field)
1 field(s) found.
Nobody types DeliveryBlockReason into a chat. The synonyms in the metadata layer connect the business phrase to the technical field and to the data product that holds it. "SAP API field" marks names that follow the S/4HANA sales order API; "lab field" marks names made up for this lab.
A search for a word nobody added, such as find "incoterms", prints 0 field(s) found. That is a valid empty result: the graph only knows what someone described.
#Step 6: Ask a multi-hop question as different users
Ask about order 5004 as anna, with the no-account sample answer:
python unit07/business_graph.py ask "Why is order 5004 blocked?" --user anna --llm sample
You should see:
User: anna (sales orgs: 1010)
Route: order context for 5004 (3 hops: order > customer > open items, other orders)
Facts from the graph:
[1] Order 5004: sales org 1010, net value 9,800 EUR, sold to Made-up Retail GmbH (10100001).
[2] Order 5004 delivery block: Credit limit exceeded (lab code) [code Z1].
[3] Customer 10100001 credit limit: 40,000 EUR.
[4] Open item 9001 for customer 10100001: 18,000 EUR, due 2026-08-15 (overdue).
[5] Open item 9002 for customer 10100001: 7,000 EUR, due 2026-10-20 (not yet due).
[6] Other open order 5001 for customer 10100001: 12,500 EUR, also blocked [code Z1].
[7] Computed exposure for customer 10100001: open items 25,000 + open orders 22,300 = 47,300 EUR against a limit of 40,000 EUR.
Answer:
[sample answer, no model called] Order 5004 delivery block: Credit limit exceeded (lab code). See the exposure fact [7].
One query followed three links and collected everything a credit analyst would check. The exposure is computed in code, not by the model, so the number is exact.
Now ask about the US order as anna, then as ben:
python unit07/business_graph.py ask "Why is order 5003 blocked?" --user anna
python unit07/business_graph.py ask "Why is order 5003 blocked?" --user ben
For anna you should see:
Facts from the graph:
(none: no record exists, or none you are allowed to see)
For ben you see six facts, ending with an exposure of 105,000 USD against a limit of 80,000 USD. The empty result for anna is correct: order 5003 is in sales org 1710. The message doesn't say which of the two reasons applies, so it doesn't reveal that the order exists.
Ask a portfolio question that has no single starting record:
python unit07/business_graph.py ask "Which blocked orders belong to customers with overdue invoices?" --user ben
You should see orders 5001, 5003 and 5004, each with the overdue amount and open item. Order 5005 is blocked too, but its customer has nothing overdue, so it is correctly left out. Run it again with --user carl and you get the empty result.
Optional, with the Unit 5 keys in .env: replace --llm sample with --llm MODEL_NAME, using a model name available in your generative AI hub. The model receives only the numbered facts and is told to cite them.
Graph vs. keyword top-3, user ben (all sales orgs), today = 2026-10-01
Why is order 5004 blocked?
needed 5 records | keyword found 2 | graph found 5
keyword missed: customer/10100001, openitem/9001, openitem/9002
Why is order 5003 blocked?
needed 4 records | keyword found 1 | graph found 4
keyword missed: customer/17100001, openitem/9003, salesorder/5006
Why is order 5005 blocked?
needed 4 records | keyword found 1 | graph found 4
keyword missed: customer/10100002, openitem/9004, salesorder/5002
Which blocked orders belong to customers with overdue invoices?
needed 5 records | keyword found 3 | graph found 5
keyword missed: openitem/9001, openitem/9003
Coverage: keyword 7/18 = 39%, graph 18/18 = 100%
Saved graph_compare_report.txt
The keyword retriever finds the order but misses the customer and the open items, because the question shares no words with them. It finds the three blocked orders in the last question only because they contain the word "blocked", and it misses the overdue evidence.
Read the 100% with care. The graph's queries were written for exactly these questions, on clean data with perfect links. In real systems, links are missing, records are duplicated and questions arrive in forms no template expects. The fair conclusion is narrower: when the links exist and the question matches a pattern, a graph retrieves complete multi-hop evidence that similarity search misses. Unit 8 shows how to measure this on realistic data.
python unit07/business_graph.py sparql "SELECT ?id ?amt WHERE { ?o a orch:SalesOrder ; orch:id ?id ; orch:netAmount ?amt } ORDER BY DESC(?amt)"
You should see six rows, starting with 5003 | 38000.0, then 6 row(s).
See how the same query would look in SAP HANA Cloud:
python unit07/business_graph.py sparql "SELECT ?id WHERE { ?o a orch:Customer ; orch:id ?id }" --hana-sql
This prints a SELECT * FROM SPARQL_TABLE('...') statement with the prefixes written out. It is a sketch: it needs a HANA Cloud instance with the knowledge graph engine enabled, which the trial and free tier don't offer. Nothing is sent anywhere.
Save your work with Git. Add the generated files to .gitignore first if you prefer to rebuild them; here they are small and made up, so committing them is fine:
As of October 2026, SAP offers the same idea at three levels. Check the status of each feature when you plan, because SAP labels several as planned rather than generally available.
SAP Knowledge Graph is SAP's own graph of its business semantics. SAP Learning describes it covering S/4HANA's metadata model and helping Joule agents disambiguate terms, link concepts to entities and reason over relationships. SAP's Architecture Center lists the layers it connects: API metadata, business semantics, data product metadata and customer-specific extensions. It also describes the graph supporting API discovery across tens of thousands of endpoints.
For a builder, the practical point is this: you don't model SAP's standard semantics yourself. When Joule or a Joule agent answers on standard SAP objects, SAP's graph does the linking. Your modelling effort goes into your own extensions and non-SAP data. In May 2026 SAP announced Joule agents in Business Data Cloud that generate business insights using SAP Knowledge Graph.
#SAP HANA Cloud knowledge graph engine: your own graph, in SAP's database
For your own graphs, SAP HANA Cloud includes a knowledge graph engine that stores RDF and runs SPARQL next to your relational and vector data. SAP's developer tutorial gives the setup:
The instance must be version 2025.2 or later.
In SAP HANA Cloud Central, open Manage Configuration, then Advanced Settings, and select the Triple Store option.
The feature is not available for trial or free tier users. The Unit 7 setup used the trial, so this part stays a sketch for most learners.
Two SQL entry points matter. SPARQL_EXECUTE is a procedure that runs any SPARQL request, including updates. SPARQL_TABLE is a table function that returns SPARQL results as rows you can join with ordinary SQL. The sketch below loads the lab's Turtle file and runs one query. It needs a paid HANA Cloud instance with the triple store enabled.
# sketch: load business_graph.ttl into SAP HANA Cloud and query it (paid instance only)
import os
from dotenv import load_dotenv
from hdbcli import dbapi
from rdflib import Graph
load_dotenv()
conn = dbapi.connect(address=os.environ["HANA_DB_ADDRESS"], port=int(os.environ["HANA_DB_PORT"]),
user=os.environ["HANA_DB_USER"], password=os.environ["HANA_DB_PASSWORD"])
cur = conn.cursor()
# 1. Turn the Turtle file into N-Triples and insert them into a named graph
nt = Graph().parse("unit07/business_graph.ttl", format="turtle").serialize(format="nt")
cur.callproc("SPARQL_EXECUTE", ("INSERT DATA { GRAPH <kg_orchestrate> { " + nt + " } }", "", "?", None))
# 2. Query it as rows with SQL
cur.execute("""SELECT * FROM SPARQL_TABLE('
PREFIX orch: <https://orchestrate.example/ontology#>
SELECT ?id ?amt FROM <kg_orchestrate>
WHERE { ?o a orch:SalesOrder ; orch:id ?id ; orch:netAmount ?amt }')""")
for row in cur.fetchall():
print(row)
conn.close()
SAP's H2 2025 innovation guide also announced that the engine would generate knowledge graphs automatically from SAP HANA Cloud metadata, showing tables, columns and their relationships, with general availability planned for Q1 2026. Check whether your release has it before you plan around it. For LangChain users, the langchain_hana package documents a HanaRdfGraph class over the same engine.
#SAP Business Data Cloud: the records, as data products
Business Data Cloud is where SAP wants the records to come from. As of October 2026:
Components. SAP Datasphere, SAP Analytics Cloud, SAP Business Warehouse, SAP Databricks, foundation services, data products and intelligent applications. In May 2026 SAP added SAP HANA Cloud as BDC's database for transactional, analytical and multi-model work (graph, vector, spatial), plus Reltio for master data and SAP Master Data Governance.
Data product types. SAP's FAQ lists SAP-managed, customer-managed and third-party data products. SAP-managed ones are activated in packages from the BDC cockpit. A data product studio for modelling custom data products was announced with general availability planned for H1 2026.
Consumption. Install a data product into SAP Datasphere or SAP Databricks, or share it through BDC Connect. SAP lists Databricks, Snowflake, Google BigQuery and Microsoft Fabric, with Amazon Athena planned for H2 2026. Under the hood, sharing uses the open Delta Sharing protocol.
AI. SAP announced deeper SAP AI Core integration for batch inference on data products, and Joule agents for data product discovery, modelling, insights and SAP Analytics Cloud stories.
Because Delta Sharing is an open protocol, a Python program with the open-source delta-sharing client can read a shared table when a provider gives it a credential file. This is the shape, not something you can run without a BDC share:
# sketch: read a shared table with the open-source Delta Sharing client (pip install delta-sharing)
import delta_sharing
profile = "config.share" # credential file from the data provider; keep it out of Git
client = delta_sharing.SharingClient(profile)
print(client.list_all_tables()) # discover what the share offers
df = delta_sharing.load_as_pandas(f"{profile}#<share>.<schema>.<table>")
print(df.head())
To try BDC itself, SAP offers a free 30-day basic trial on a shared tenant with preconfigured sample data. Some features are restricted, and content doesn't carry over between trial periods.
Business Data Cloud is a paid SaaS subscription. SAP's launch FAQ described capacity units as the pricing metric, shared across components. Existing Datasphere or SAP Analytics Cloud contracts don't automatically cover BDC. Confirm the current model with SAP.
SAP HANA Cloud knowledge graph engine runs on a paid HANA Cloud instance; the trial and free tier don't include it. Graph data adds to what the instance must hold.
SAP Knowledge Graph has no separate price in the sources used for this topic. Ask SAP which of your Joule, BDC or other licences include it.
Edition prerequisites. SAP's 2025 FAQ listed S/4HANA Cloud public and private editions as sources for SAP-managed data products, and none for on-premise S/4HANA at the time. Check the current list for your edition.
Security and SAP authorizations. A graph collapses many tables into one queryable space, which makes over-sharing easy. Put the user's authorizations into every query, as allowed_filter() does, and derive them from the user's real roles, not from a request parameter. Treat graph exports like any other copy of SAP data. Model-written SPARQL needs validation: limit it to SELECT, add the access filter yourself after generation, and never let a model's query run INSERT or DELETE. Use value binding (initBindings, or bind parameters in SQL) for every user-supplied value.
Evaluation. Score retrieval on facts needed, not only on answer wording. The compare command is a tiny version of that. Unit 8 builds labelled query sets and measures recall on them. Also test the link step: how often does "overdue" or a customer name map to the right graph term?
Data quality. Graphs magnify bad master data. Duplicate customers split exposure across two nodes, and a missing link silently drops a fact. Profile key relationships (every order has exactly one sold-to party) before loading, and alert when the rules break.
Freshness. Data products are replicated with initial and delta loads, so they lag S/4HANA. For a decision like releasing a credit block, confirm the current status live in S/4HANA as the user before acting, as the previous topic recommended.
Cost. Graph data adds to the HANA Cloud instance's size, and size drives cost. BDC consumes capacity units. Keep the graph to the entities your use cases need, and load records by reference where you can.
Operations. Version the ontology like code and review changes. Schedule rebuilds or incremental updates when source data changes, and monitor triple counts per layer, as the build command prints.
Clean core. Read S/4HANA through released APIs, CDS views or BDC data products. Don't add custom tables in S/4HANA just to feed a graph. Keep the graph and its logic on SAP BTP or in BDC, side by side with the core.
Modelling the whole enterprise first. Graph projects stall when they try to model everything. Start with the five to ten entities behind your top questions.
Rebuilding SAP's semantics. If Joule and SAP Knowledge Graph already cover standard objects, spend your effort on your own extensions instead.
Letting the model compute numbers. Exposure, totals and overdue days belong in queries or code. The model explains them.
Trusting code values across systems. Block reasons, statuses and document types are configured per system. Map them per system in the metadata layer.
Leaking existence. "You may not see order 5003" tells the user it exists. Return the same empty result for "missing" and "not allowed".
Mistaking a planned feature for an available one. Several BDC and HANA Cloud features in this topic were announced with planned dates. Check the release you have.
Assuming the trial covers it. The HANA Cloud knowledge graph engine needs a paid instance. Plan the prototype on open-source tools, as here.
Add a fourth data product, deliveries, and a question that needs it.
In business_graph.py, add a list DELIVERIES with two made-up outbound deliveries, for example ("8001", "5002", "2026-09-28") and ("8002", "5006", "2026-09-30"): delivery number, sales order, goods issue date.
Add "DeliveryLab" to DATA_PRODUCTS with entity Delivery and three lab fields with labels and synonyms such as "shipped" and "goods issue".
In build_graph(), add the class orch:Delivery, a relation orch:delivers from delivery to sales order, and the delivery records.
Write a query that lists, for one order, its deliveries and goods issue dates, with the same access filter as the other queries. Add it to ask for questions containing "shipped", checked before the order-number route so it wins.
Add your question and its needed records to COMPARE_SET, then run build and compare again.
Commit the script and the new graph_compare_report.txt. Unit 8 uses this report as the starting point for retrieval evaluation.
Done whenfind "shipped" returns your new field, ask "What has shipped for order 5006?" --user ben lists delivery 8002, the same question as anna returns the empty result, and compare shows your new question with the graph finding every needed record.
Pick one answer for each question. The explanation appears after you choose.
1What is the core difference between vector retrieval and graph retrieval?
Answer: B. Similarity search ranks items that look like the question. A graph stores links such as "order sold to customer" and follows them, so it reaches facts the question never mentions, such as the customer's open items.
2In the lab, why does the order context query use OPTIONAL for open items and other orders?
Answer: C. Without OPTIONAL, any missing part removes the whole row, so an order whose customer has no open items would vanish. OPTIONAL keeps the row and leaves those variables empty.
3How does the lab stop anna from seeing order 5003?
Answer: D. allowed_filter() puts the user's sales organizations into the SPARQL query, so rows from 1710 never come back. Filtering inside the query is enforcement; instructions to the model or filtering after generation are not.
4What do BDC foundation services do for SAP-managed data products?
Answer: B. Foundation services run SAP's pipelines: CDS views harmonize source tables, initial and delta loads replicate them, and the files land in the SAP HANA Cloud data lake object store. Data products are then shared with Delta Sharing.
5Which two specifications describe a data product in BDC, and what does each cover?
Answer: C. Open Resource Discovery describes the data product, its output and input ports and entity types. CSN Interop describes each entity and its elements with semantic annotations. Together they are the metadata a knowledge graph can link.
6Your team wants to prototype the HANA Cloud knowledge graph engine on the BTP trial from Unit 7 setup. What do you tell them?
Answer: B. SAP's tutorial states the knowledge graph feature is not available for trial or free tier users. An open-source library on sample data proves the idea, and the later move needs a paid instance with the Triple Store option.
7A colleague proposes letting the model write SPARQL against the production graph. What is the safest design?
Answer: D. SPARQL also has updates such as INSERT and DELETE, so restrict the query type. The access filter must be added by your code, not trusted to the model, and user-supplied values should be bound, never pasted into query text.
8The compare command shows the graph at 100% coverage. What is the right conclusion?
Answer: B. The queries were written for these questions on clean, fully linked data. Real data has gaps and questions arrive in unexpected forms. The fair claim is narrower, and Unit 8 measures it on realistic data.
9Credit release decisions use a data product refreshed by delta loads. What should the application do before acting?
Answer: C. Data products are replicated and lag S/4HANA. For a decision that changes a business outcome, check the live status as the user, which also applies that user's authorizations.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Foundation Layer, AI-native north star architecture (SAP Architecture Center, updated 2026-05-13)— Knowledge Graph links natural language to structured metadata, API discovery, layers of API metadata, business semantics, data product metadata and customer extensions; filter parameters for "show me overdue orders"; data products with ownership, schema, authorization rules and lifecycle; access by SAC, Datasphere, partner platforms and agents via natural-language-to-SQL
Accelerate the Autonomous Enterprise with SAP Business Data Cloud (SAP News, May 2026)— SAP HANA Cloud as AI database in BDC; Reltio; SAP Master Data Governance; SAP AI Core batch inference on data products; Joule agents for data product discovery and insights using SAP Knowledge Graph; BDC Connect partners Snowflake, Databricks, Google BigQuery, Microsoft Fabric, Amazon Athena (Athena planned H2 2026); zero-copy sharing
SAP Business Data Cloud - FAQs (SAP Community)— written around the February 2025 launch; components; SAP-managed, customer-managed and third-party data products; activation in BDC cockpit; capacity units as pricing metric; supported S/4HANA editions at the time; no free trial at launch
SAP Innovation Guide H2 2025 (sap.com)— HANA Cloud knowledge graph engine generates knowledge graphs from HANA Cloud metadata (planned GA Q1 2026); data product studio (planned H1 2026); bidirectional data product sharing between BDC and HANA Cloud (planned Q1 2026); BDC Connect for Snowflake and Google BigQuery planned H1 2026, Databricks GA