Orchestrate

Data security and PII in AI systems

Find where personal data flows and piles up in an AI assistant, keep it minimal, teach detectors your SAP IDs, and erase one person from every store.

Updated Oct 8, 2026Foundational 9 minDeep 40 min
Foundational layer · 9 min read

The 60-second version

An AI assistant is a new place where personal data goes. A credit controller asks about a blocked order, and the customer's name, email and payment history travel to a model. They also land in a log, a search index, a cache and a test set.

Data security for AI means three things:

  • Send less. Give the model only the fields the task needs, and mask the rest.
  • Know where it piles up. Every copy of personal data needs an owner, a reason and an end date.
  • Be able to delete it. When a person's data must go, it must go from the AI's stores too, not only from SAP.

Most AI privacy failures are not clever attacks. They are copies nobody listed.

Why it matters to the business

Personal data in an SAP landscape is everywhere: business partners, contact persons, employees, bank details, free-text notes. In order-to-cash, a customer who is a sole trader is a person. Their customer number identifies them as surely as their name.

The risk shows up in three ways:

  • Disclosure. OWASP lists sensitive information disclosure as a top risk for AI applications. An answer can show one user another person's data, or an attacker can extract it with prompt injection.
  • Retention. SAP systems already block and delete personal data when its purpose ends. An AI log that keeps every prompt forever quietly undoes that work.
  • Cost of rework. Retrofitting deletion into a vector index or an evaluation set after go-live is slow and expensive. Designing it in costs little.

Example. A blocked-orders assistant goes live in credit management. Six months later, a former customer asks to have their data erased. SAP's blocking and deletion handles the business partner. But the assistant's prompt log, its search index of credit notes and its answer cache still hold her name and history. Nobody can say how many copies exist. That is the gap this topic closes.

How SAP does it

As of October 2026, SAP's material splits the work into several pieces. Check current documentation and contracts before you rely on any of them.

  • Model providers and training. SAP Learning separates inference (a model processes your request) from training (a model is built from data). SAP states that it applies contractual and technical controls so external model providers don't use customer data to train their models. Some services offer optional learning from customer data; SAP says those depend on the service and its settings, and customers decide.
  • Where data is processed. SAP Learning describes three patterns: embedded AI runs inside the SAP application, side-by-side AI runs on SAP BTP, and generative AI requests go through the generative AI hub to SAP-managed infrastructure or approved external providers. Regional options are "subject to availability". For European customers SAP offers an EU AI Cloud.
  • Masking before the model. The orchestration service in the generative AI hub has a data masking module. It can anonymize (replace for good) or pseudonymize (replace, then restore in the answer). The guardrails topic shows how to configure it.
  • Masking as a service. SAP Data Privacy Integration offers anonymization and pseudonymization of free text, files and images on SAP BTP.
  • Blocking and deletion in SAP. SAP Information Lifecycle Management (ILM) blocks personal data when its business purpose ends and deletes it when retention periods expire. SAP Data Retention Manager does deletion orchestration for applications on SAP BTP.
  • Who read what. Read Access Logging records who accessed sensitive data in SAP and when.

One line from SAP Learning matters most: customers stay responsible for storage, governance and deletion in their own environments. Your side-by-side assistant is your environment.

Where personal data lives in an AI assistant

Use this table as a checklist in design reviews. Every row is a store that needs a retention rule and a deletion path.

Store What lands there Typical owner Question to settle
Prompt and answer log Questions, grounding data, answers App team How long do we keep full text, and who can read it?
Vector index Chunks of notes and documents, plus their embeddings App team Can we find and delete all chunks for one person?
Answer cache Reused answers keyed by question App team Does a cached answer outlive the person's data in SAP?
Evaluation set Real questions and expected answers AI team Did we use real customer data, and is that allowed?
Traces and monitoring Prompts in spans, error messages Operations Is content capture switched on in production?
Model provider The request while it is processed Provider, under contract What do the contract and service say about storage?
Fine-tuned model Patterns, sometimes verbatim text, from training data AI team Could the model repeat what it was trained on?

Two rows surprise people:

  • Embeddings are not anonymous. Researchers recovered 92% of short texts exactly from their embeddings, and most names from embedded clinical notes. Treat a vector index like the text it came from.
  • Pseudonymized is still personal. Under the GDPR, data that can be linked back to a person with extra information is still personal data. Masking with a reversible map lowers risk; it doesn't take the data out of scope.

Questions to ask

  1. Which fields does each AI use case send to the model, and why is each one needed?
  2. Which personal data types do we mask before the model? Have we tested the masker on our own data, including SAP IDs and German text?
  3. Where is personal data stored by the assistant: logs, index, cache, eval set, traces? Who owns each store?
  4. What is the retention period for each store? Does it match the SAP retention rules for the same data?
  5. When SAP blocks a business partner, does the assistant stop using that person's data?
  6. Can we erase one person from every AI store, and prove it? How long does it take?
  7. Which model provider and region process our requests? What does the contract say about storage and training?
  8. Is prompt content captured in production traces? Who can read them?
  9. Has our data protection officer reviewed the design before go-live?

Common misconceptions

  • "The provider doesn't train on our data, so we're done." Training is one risk. Logs, caches, indexes and traces you run are the bigger ones, and they're yours.
  • "We mask, so the data is anonymous." Reversible masking is pseudonymization. The data is still personal data, and the map back is itself sensitive.
  • "Embeddings are just numbers." Text can be reconstructed from embeddings. Protect and delete them like the source text.
  • "A customer number isn't personal data." For a sole trader or a contact person, it points straight to a human.
  • "Deleting in SAP deletes everywhere." SAP's blocking and deletion covers SAP's own data. Copies in a side-by-side app need their own deletion path.
  • "Masking tools catch everything." Detectors miss formats they don't know and flag things that aren't personal data. They need testing like any other component.

Key terms

  • Personal data / PII: information about an identified or identifiable person. PII (personally identifiable information) is the common US term.
  • Data minimisation: using only the data a purpose needs.
  • Anonymization: removing the link to a person for good.
  • Pseudonymization: replacing identifiers so the link can be restored only with extra information kept apart.
  • End of purpose: the point where data is no longer needed for its original business purpose.
  • Blocking: restricting access to data whose purpose has ended but which must be kept for legal retention.
  • Retention period: how long data must be kept, for example for audits, before deletion.
  • Data subject: the person the data is about.
  • Read Access Logging: SAP's record of who read sensitive data and when.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1An assistant's provider contract says customer data isn't used for training. What is the main remaining risk?

    Answer: B. Training is one way data spreads, but the stores you build around the model are the larger risk. SAP Learning states that customers remain responsible for storage, governance and deletion in their own environments.
  2. 2A credit assistant needs to explain why an order is blocked. Which design best follows data minimisation?

    Answer: C. Data minimisation means using only what the purpose needs. Region and log cleanup help, but they don't reduce what the model sees in the first place.
  3. 3Your team masks names with numbered placeholders and restores them in the answer. How should you treat the masked data?

    Answer: B. Reversible masking is pseudonymization, and the GDPR treats pseudonymized data as personal data when extra information can restore the link. The mapping itself needs protection.
  4. 4A vendor says its vector index is safe because it stores embeddings, not text. What should you answer?

    Answer: D. Researchers recovered 92% of short inputs exactly from their embeddings, and most names from embedded clinical notes. Treat a vector index like the text it came from.
  5. 5SAP blocks a business partner whose purpose has ended. What must the assistant's design ensure?

    Answer: C. SAP's blocking and deletion covers SAP's own data. Copies in a side-by-side app need their own retention and deletion paths that match SAP's rules.
  6. 6For a sole trader in order-to-cash, which item is personal data?

    Answer: B. A sole trader is a person, so any identifier that leads to them is personal data, including the customer number. Masking tools need to know your SAP ID formats to catch it.
  7. 7Which question best tests whether an AI team is ready for a deletion request?

    Answer: D. Deletion must reach logs, indexes, caches, eval sets and traces, and you need a record that it happened. The other questions matter but don't show readiness for erasure.
Deep layer · 40 min read

Mental model: personal data is a flow with copies

Picture personal data as water moving through your assistant. It enters from SAP and from users. It leaves through the model call and the answer. Along the way, every component that writes to disk is a tank where some of it stays.

Three controls follow from that picture:

  1. Narrow the inflow. Minimise fields, then mask what's left. This is the cheapest control and protects every tank downstream.
  2. Inventory the tanks. For each store, know which data types land there, who reads it and how long it stays.
  3. Drain on demand. For one data subject, find and delete every copy, then prove it without creating a new copy of their data.

Detection, masking and deletion all depend on one thing: knowing what personal data looks like in your data. Generic detectors know names and emails. They don't know that 10100001 is a customer number.

How it works

flowchart LR
  SAP[(SAP record)] --> MIN[Minimise fields]
  U[User question] --> MASK[Detect + mask]
  MIN --> MASK
  MASK --> M[Model]
  M --> A[Answer]
  MASK -.-> LOG[(Prompt log)]
  A -.-> CACHE[(Answer cache)]
  DOCS[Notes, PDFs] --> IDX[(Vector index)]
  IDX --> MASK
  LOG -.-> EVAL[(Eval set)]

Solid lines are the request path. Dotted lines are copies that persist.

Minimisation

Minimisation is an allowlist of fields per task. To explain a credit block, the model needs the credit limit, credit segment and block reason. It doesn't need the birthday, street or bank account. Build the prompt from an explicit list of fields, not from "the whole record". A new field added to the SAP API later then stays out by default.

Detection

A detector such as Presidio combines two methods:

  • Patterns (regular expressions, sometimes with checksums) for structured data such as emails, phone numbers and IBANs.
  • Named-entity recognition (NER), a trained language model that tags names, places and dates in free text.

Each finding has a type, a position and a confidence score. A context word near a pattern match, such as "customer" before an 8-digit number, raises the score. You add your own recognizers for formats the detector doesn't know, such as SAP customer, supplier and personnel numbers.

Number ranges for these IDs are configured per SAP system. The patterns in this topic match its made-up data only. Copy the real ranges from your own system's configuration.

Masking

The masker replaces each finding. Anonymization throws the original away. Pseudonymization keeps a map from placeholder to value so the answer can be restored for an authorized user. The map is personal data, so it lives in memory for one request or in a store with its own protection. The guardrails topic builds this map.

Stores and their risks

Store Why it holds personal data Main control
Prompt log Full prompts and answers for debugging Redact before writing; short retention; restricted access
Vector index Chunks of notes, and embeddings computed from them Subject metadata on every chunk; delete text and vector together
Answer cache Answers reused across users Key by user's authorization scope; expire with the source
Eval set Real questions copied from logs Use made-up or masked data; record its origin
Traces Prompt content in spans Leave content capture off in production unless approved
Fine-tuned model Training text the model may memorise Mask training data; test for regurgitation

The vector index needs the most care. Morris and colleagues showed that Vec2Text, an inversion method, recovered 92% of 32-token inputs exactly from GTR-base embeddings. On clinical notes, it recovered 89% of full names. Their conclusion: treat embeddings as highly sensitive data. Deleting a chunk's text but keeping its vector is not deletion.

Erasure

Erasure is a search problem first. You must find every row about one person. Two things make it hard:

  • Chunking cuts the link. A long credit note is split into chunks. The chunk that says "Ms Becker prefers calls after 4 pm" may not contain the customer number.
  • Free text varies. "Anna Becker", "Ms Becker" and "A. Becker" are the same person to a human and three strings to code.

The fix is to attach the data subject's ID as metadata to every chunk, log row and cache entry when it's written. Deletion then becomes a filter on that field. A wider text search afterwards catches what the metadata missed, for a human to review.

The erasure record proves the deletion happened. It must not become a new copy of the person's data. Store a keyed hash of the ID, the date and the counts per store. A plain hash of an 8-digit number is not enough: anyone can hash all 100 million candidates and find the match.

Logging

Application logs are a store too. A logging filter runs on every log line before it's written and masks personal data. It's a backstop: the safer design is not to log full prompts at all in production, and to log IDs and decisions instead.

Build it yourself: a privacy lab for the blocked-orders assistant

You will build one script, privacy_lab.py, with six commands. It teaches Presidio your SAP ID formats, builds a minimised and masked prompt, creates the made-up stores of an assistant, lists the personal data in each, erases one customer from all of them and scores the detector on labelled examples. A final command shows a log filter that keeps personal data out of log files.

Before you start: complete Set up your computer for this course and Set up for Unit 11. You need Presidio and the spaCy model en_core_web_lg from Step 2 of the Unit 11 setup, and python-dotenv from the course setup.

flowchart LR
  D[detect] --> P[prompt]
  P --> I[init stores]
  I --> V[inventory]
  V --> E[erase]
  E --> S[evaluate]
  S --> L[log-demo]

What you need

  • Your course folder orchestrate-course with its .venv, and Presidio installed from the Unit 11 setup.
  • About 45 minutes.
  • Cost: free. No account, no key, no model call. Everything runs on your computer.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that Presidio and python-dotenv are installed:

    python -c "import presidio_analyzer, dotenv; print('ready')"

    You should see ready. If you see ModuleNotFoundError, go back to Step 2 of Set up for Unit 11.

Run every command in this topic from the course folder, not from inside unit11.

Step 2: Create the lab

  1. In VS Code's file list, right-click unit11, choose New File, and name it privacy_lab.py.
  2. Paste the code below and save.
"""Privacy lab: find, minimise, inventory and erase personal data in an AI assistant.

Run every command from your course folder:
    python unit11/privacy_lab.py detect --sample          # SAP-aware detection on made-up notes
    python unit11/privacy_lab.py detect "Customer 10100001 called"
    python unit11/privacy_lab.py prompt --customer 10100001
    python unit11/privacy_lab.py init                     # create the assistant's made-up stores
    python unit11/privacy_lab.py inventory
    python unit11/privacy_lab.py erase --customer 10100001 --dry-run
    python unit11/privacy_lab.py erase --customer 10100001
    python unit11/privacy_lab.py evaluate
    python unit11/privacy_lab.py log-demo

Everything runs on your computer with Presidio and spaCy from the Unit 11 setup. No text leaves
your machine. All people, numbers and notes are invented.
"""
import argparse
import hashlib
import hmac
import json
import logging
import os
from datetime import date
from pathlib import Path

from dotenv import load_dotenv
from presidio_analyzer import AnalyzerEngine, Pattern, PatternRecognizer
from presidio_anonymizer import AnonymizerEngine

HERE = Path(__file__).parent
STORE_FILE = HERE / "privacy_store.json"
LOG_FILE = HERE / "privacy_lab.log"
REPORT_FILE = HERE / "erasure_log.jsonl"

# ---------- 1. Your SAP system's ID formats ----------
# Number ranges are configured per SAP system. These match this lab's made-up data only.
# For a real project, copy the ranges from your own system's configuration.
SAP_IDS = {
    "SAP_CUSTOMER_ID": (r"\b101\d{5}\b", ["customer", "sold-to", "payer", "debtor", "kunde", "debitor"]),
    "SAP_SUPPLIER_ID": (r"\b173\d{5}\b", ["supplier", "vendor", "creditor", "lieferant", "kreditor"]),
    "SAP_PERSONNEL_ID": (r"\b500\d{5}\b", ["employee", "personnel", "pernr", "mitarbeiter"]),
}
STANDARD = ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "IBAN_CODE"]

# ---------- 2. Made-up SAP master data (stands in for business partners in S/4HANA) ----------
CUSTOMERS = {
    "10100001": {"Name": "Anna Becker", "Type": "sole trader", "Email": "anna.becker@example.com",
                 "Phone": "+49 30 1234 5678", "IBAN": "DE89 3704 0044 0532 0130 00",
                 "CreditLimit": 250000, "CreditSegment": "A", "BlockReason": "credit limit exceeded",
                 "Birthday": "1984-03-12", "Street": "Lindenstrasse 4", "City": "Berlin"},
    "10100002": {"Name": "Nordwind Logistik GmbH", "Type": "company", "Contact": "Marco Rossi",
                 "Email": "m.rossi@nordwind.example", "Phone": "+49 40 555 0101",
                 "IBAN": "DE02 1203 0000 0000 2020 51", "CreditLimit": 900000, "CreditSegment": "B",
                 "BlockReason": "overdue items", "Street": "Hafenweg 9", "City": "Hamburg"},
}


def build_analyzer(sap_aware: bool = True) -> AnalyzerEngine:
    """Presidio's analyzer, optionally taught your SAP ID formats."""
    analyzer = AnalyzerEngine()
    if sap_aware:
        for entity, (regex, context) in SAP_IDS.items():
            analyzer.registry.add_recognizer(PatternRecognizer(
                supported_entity=entity,
                patterns=[Pattern(name=entity.lower(), regex=regex, score=0.5)],
                context=context,  # nearby words like "customer" raise the score
            ))
    return analyzer


def entities(sap_aware: bool) -> list:
    return STANDARD + (list(SAP_IDS) if sap_aware else [])


def find(analyzer, text: str, sap_aware: bool = True) -> list:
    return sorted(analyzer.analyze(text=text, entities=entities(sap_aware), language="en"),
                  key=lambda f: f.start)


# ---------- 3. detect: what does the detector see? ----------
SAMPLE_NOTES = [
    "Customer 10100001 called about blocked order 9000001. Contact: Anna Becker, "
    "anna.becker@example.com, phone +49 30 1234 5678.",
    "Supplier 17300042 asks why invoice 90012345 is parked. Approved by employee 50000123.",
]


def cmd_detect(args) -> None:
    analyzer, anonymizer = build_analyzer(), AnonymizerEngine()
    texts = SAMPLE_NOTES if args.sample or not args.text else [args.text]
    for number, text in enumerate(texts, start=1):
        findings = find(analyzer, text)
        print(f"\nText {number}: {len(findings)} finding(s)")
        for f in findings:
            print(f"  {f.entity_type:<17} {f.score:.2f}  {text[f.start:f.end]}")
        print("  Masked:", anonymizer.anonymize(text=text, analyzer_results=findings).text)


# ---------- 4. prompt: minimise first, then mask ----------
# The fields this task (explain a credit block) needs. Everything else stays in SAP.
NEEDED_FOR_CREDIT_BLOCK = ["Type", "CreditLimit", "CreditSegment", "BlockReason"]


def cmd_prompt(args) -> None:
    record = CUSTOMERS[args.customer]
    kept = {k: v for k, v in record.items() if k in NEEDED_FOR_CREDIT_BLOCK}
    print(f"SAP record has {len(record)} fields; the prompt gets {len(kept)}: {', '.join(kept)}")
    question = f"Customer {args.customer} ({record.get('Name')}) asks why order 9000001 is blocked."
    analyzer, anonymizer = build_analyzer(), AnonymizerEngine()
    masked = anonymizer.anonymize(text=question, analyzer_results=find(analyzer, question)).text
    print("Prompt sent to the model:")
    print(f"  Question: {masked}")
    print(f"  Record:   {json.dumps(kept)}")


# ---------- 5. init: the assistant's stores, where personal data piles up ----------
def fake_vector(text: str) -> list:
    """Stands in for an embedding: numbers computed from the text. Real ones can be inverted."""
    digest = hashlib.sha256(text.encode()).digest()
    return [round(b / 255, 3) for b in digest[:6]]


def cmd_init(args) -> None:
    chunks = [
        ("10100001", "Credit note: Anna Becker, sole trader, asked to raise the limit to 300,000 EUR."),
        ("10100001", "Customer 10100001 paid late twice in 2026; reminder sent to anna.becker@example.com."),
        ("10100002", "Nordwind Logistik GmbH: overdue items on invoice 90012345, contact Marco Rossi."),
        # A long note split in two: this half lost the customer number when it was chunked.
        (None, "Ms Becker prefers calls after 4 pm and asked us not to email her reminders."),
        (None, "Policy: orders above the credit limit are blocked until credit management releases them."),
    ]
    store = {
        "prompt_log": [
            {"id": 1, "user": "CC_MUELLER", "customer": "10100001",
             "prompt": "Why is order 9000001 for Anna Becker blocked?",
             "answer": "Customer 10100001 exceeded the credit limit of 250,000 EUR."},
            {"id": 2, "user": "CC_MUELLER", "customer": "10100002",
             "prompt": "Why is order 9000002 blocked?",
             "answer": "Nordwind Logistik GmbH has overdue items. Call Marco Rossi on +49 40 555 0101."},
            {"id": 3, "user": "CC_SCHMIDT", "customer": None,
             "prompt": "What is our credit block policy?",
             "answer": "Orders above the credit limit are blocked until credit management releases them."},
        ],
        "vector_index": [{"id": i, "customer": c, "text": t, "vector": fake_vector(t)}
                         for i, (c, t) in enumerate(chunks, start=1)],
        "answer_cache": [
            {"key": "why is order 9000001 blocked", "customer": "10100001",
             "answer": "Customer 10100001 exceeded the credit limit of 250,000 EUR."},
            {"key": "what is our credit block policy", "customer": None,
             "answer": "Orders above the credit limit are blocked until credit management releases them."},
        ],
        "eval_set": [
            {"id": "E1", "customer": "10100001", "question": "Why is Anna Becker's order 9000001 blocked?",
             "expected": "credit limit exceeded"},
            {"id": "E2", "customer": "10100002", "question": "Why is order 9000002 blocked?",
             "expected": "overdue items"},
        ],
    }
    STORE_FILE.write_text(json.dumps(store, indent=2), encoding="utf-8")
    print(f"Created {STORE_FILE.name} with " +
          ", ".join(f"{name} ({len(rows)})" for name, rows in store.items()))


def load_store() -> dict:
    if not STORE_FILE.exists():
        raise SystemExit("No store yet. Run: python unit11/privacy_lab.py init")
    return json.loads(STORE_FILE.read_text(encoding="utf-8"))


def texts_of(row: dict) -> str:
    return " ".join(str(v) for k, v in row.items() if isinstance(v, str) and k not in ("id", "user"))


# ---------- 6. inventory: which store holds which kinds of personal data? ----------
def cmd_inventory(args) -> None:
    store, analyzer = load_store(), build_analyzer()
    print(f"{'store':<14} {'rows':>4} {'with PII':>8}  types found")
    for name, rows in store.items():
        types, hits = set(), 0
        for row in rows:
            found = {f.entity_type for f in find(analyzer, texts_of(row))}
            hits += bool(found)
            types |= found
        print(f"{name:<14} {len(rows):>4} {hits:>8}  {', '.join(sorted(types)) or '-'}")
    if any("vector" in row for row in store.get("vector_index", [])):
        print("\nNote: vector_index also holds vectors computed from the same text. "
              "Treat them as personal data too.")


# ---------- 7. erase: one data subject, every store ----------
def identifiers(customer: str) -> list:
    record = CUSTOMERS.get(customer)
    if record is None:
        raise SystemExit(f"Unknown customer {customer}. Known: {', '.join(CUSTOMERS)}")
    values = [customer] + [record[k] for k in ("Name", "Contact", "Email", "Phone") if k in record]
    return [v for v in values if v]


def matches(row: dict, customer: str, ids: list) -> bool:
    text = texts_of(row)
    return row.get("customer") == customer or any(v in text for v in ids)


def residual_hits(store: dict, customer: str) -> list:
    """A wider, case-insensitive search after deletion: identifiers plus surnames."""
    record = CUSTOMERS[customer]
    terms = [v.lower() for v in identifiers(customer)]
    person = record.get("Contact") or record["Name"]  # the human behind the account
    terms.append(person.split()[-1].lower())          # surname, e.g. "becker"
    return [f"{name} row {row.get('id', row.get('key'))}" for name, rows in store.items()
            for row in rows if any(t in texts_of(row).lower() for t in set(terms))]


def cmd_erase(args) -> None:
    store, ids = load_store(), identifiers(args.customer)
    counts = {}
    for name, rows in store.items():
        doomed = [r for r in rows if matches(r, args.customer, ids)]
        counts[name] = len(doomed)
        store[name] = [r for r in rows if r not in doomed]
    residual = residual_hits(store, args.customer)
    verb = "Would delete" if args.dry_run else "Deleted"
    for name, n in counts.items():
        print(f"{verb} {n} row(s) from {name}")
    print(f"Possible leftovers after erasure: {len(residual)}")
    for hit in residual:
        print(f"  check by hand: {hit}")
    if args.dry_run:
        print("Dry run: nothing changed.")
        return
    STORE_FILE.write_text(json.dumps(store, indent=2), encoding="utf-8")
    # The erasure record proves the deletion without keeping the personal data itself.
    # A plain hash of an 8-digit number is easy to reverse by trying every number,
    # so the hash is keyed with a secret from .env (HMAC).
    load_dotenv()
    key = os.getenv("ERASURE_LOG_KEY")
    if not key:
        print("Note: ERASURE_LOG_KEY is not set in .env; using a lab-only key.")
        key = "lab-only-key"
    subject = hmac.new(key.encode(), args.customer.encode(), hashlib.sha256).hexdigest()[:12]
    entry = {"date": date.today().isoformat(), "subject_hash": subject, "deleted": counts,
             "leftovers": len(residual)}
    with REPORT_FILE.open("a", encoding="utf-8") as f:
        f.write(json.dumps(entry) + "\n")
    print(f"Logged to {REPORT_FILE.name}: {json.dumps(entry)}")


# ---------- 8. evaluate: is the detector good enough for this data? ----------
# Each case: text, and the personal data a reviewer marked in it (type, exact text).
GOLD = [
    ("Customer 10100001 called about order 9000001.", [("SAP_CUSTOMER_ID", "10100001")]),
    ("Sold-to 10100003 disputes invoice 90012346.", [("SAP_CUSTOMER_ID", "10100003")]),
    ("Supplier 17300042 sent a new bank account.", [("SAP_SUPPLIER_ID", "17300042")]),
    ("Approved by employee 50000123 on Monday.", [("SAP_PERSONNEL_ID", "50000123")]),
    ("Contact Anna Becker at anna.becker@example.com.",
     [("PERSON", "Anna Becker"), ("EMAIL_ADDRESS", "anna.becker@example.com")]),
    ("Marco Rossi asked for a call back on +49 40 555 0101.",
     [("PERSON", "Marco Rossi"), ("PHONE_NUMBER", "+49 40 555 0101")]),
    ("Pay to IBAN DE89 3704 0044 0532 0130 00 by Friday.", [("IBAN_CODE", "DE89 3704 0044 0532 0130 00")]),
    ("Order 9000002 waits for credit release; invoice 90012345 is parked.", []),
    ("Plant 1010 and sales organization 1710 are affected.", []),
    ("Kunde 10100002 hat offene Posten.", [("SAP_CUSTOMER_ID", "10100002")]),
]


def score(sap_aware: bool, analyzer) -> tuple:
    tp = fp = fn = 0
    misses = []
    for text, gold in GOLD:
        got = {(f.entity_type, text[f.start:f.end]) for f in find(analyzer, text, sap_aware)}
        want = set(gold)
        tp += len(got & want)
        fp += len(got - want)
        fn += len(want - got)
        misses += [f"missed {t} '{v}'" for t, v in sorted(want - got)]
        misses += [f"false alarm {t} '{v}'" for t, v in sorted(got - want)]
    return tp, fp, fn, misses


def cmd_evaluate(args) -> None:
    analyzer = build_analyzer()
    for label, sap_aware in (("Presidio defaults", False), ("SAP-aware", True)):
        tp, fp, fn, problems = score(sap_aware, analyzer)
        recall = tp / (tp + fn) if tp + fn else 1.0
        precision = tp / (tp + fp) if tp + fp else 1.0
        print(f"{label:<18} recall {recall:.0%}  precision {precision:.0%}  "
              f"(found {tp}, missed {fn}, false alarms {fp})")
        for p in problems:
            print(f"    {p}")


# ---------- 9. log-demo: keep personal data out of application logs ----------
class RedactPersonalData(logging.Filter):
    """Masks personal data in every log message before any handler writes it."""

    def __init__(self, analyzer):
        super().__init__()
        self.analyzer, self.anonymizer = analyzer, AnonymizerEngine()

    def filter(self, record: logging.LogRecord) -> bool:
        text = record.getMessage()
        findings = find(self.analyzer, text)
        record.msg = self.anonymizer.anonymize(text=text, analyzer_results=findings).text
        record.args = ()
        return True


def cmd_log_demo(args) -> None:
    log = logging.getLogger("assistant")
    log.setLevel(logging.INFO)
    handler = logging.FileHandler(LOG_FILE, mode="w", encoding="utf-8")
    handler.setFormatter(logging.Formatter("%(levelname)s %(message)s"))
    if not args.no_redact:
        handler.addFilter(RedactPersonalData(build_analyzer()))
    log.addHandler(handler)
    log.info("question from CC_MUELLER: Why is order 9000001 for Anna Becker blocked?")
    log.info("grounding: customer 10100001, email anna.becker@example.com, limit 250000")
    handler.close()
    print(f"{LOG_FILE.name}:")
    print(LOG_FILE.read_text(encoding="utf-8").rstrip())


def main() -> None:
    parser = argparse.ArgumentParser(description="Personal data lab for an SAP-shaped assistant.")
    sub = parser.add_subparsers(dest="command", required=True)
    p = sub.add_parser("detect", help="find and mask personal data in text")
    p.add_argument("text", nargs="?")
    p.add_argument("--sample", action="store_true")
    p.set_defaults(func=cmd_detect)
    p = sub.add_parser("prompt", help="build a minimised, masked prompt for one customer")
    p.add_argument("--customer", default="10100001", choices=list(CUSTOMERS))
    p.set_defaults(func=cmd_prompt)
    sub.add_parser("init", help="create the made-up stores").set_defaults(func=cmd_init)
    sub.add_parser("inventory", help="list personal data per store").set_defaults(func=cmd_inventory)
    p = sub.add_parser("erase", help="delete one customer's data from every store")
    p.add_argument("--customer", required=True)
    p.add_argument("--dry-run", action="store_true")
    p.set_defaults(func=cmd_erase)
    sub.add_parser("evaluate", help="score the detector on labelled text").set_defaults(func=cmd_evaluate)
    p = sub.add_parser("log-demo", help="write two log lines through the redaction filter")
    p.add_argument("--no-redact", action="store_true")
    p.set_defaults(func=cmd_log_demo)
    args = parser.parse_args()
    args.func(args)


if __name__ == "__main__":
    main()

Step 3: Teach Presidio your SAP IDs

  1. Run the detector on the two made-up notes:

    python unit11/privacy_lab.py detect --sample

What success looks like (the first run takes about 10 seconds while spaCy loads):

Text 1: 4 finding(s)
  SAP_CUSTOMER_ID   0.85  10100001
  PERSON            0.85  Anna Becker
  EMAIL_ADDRESS     1.00  anna.becker@example.com
  PHONE_NUMBER      0.75  +49 30 1234 5678
  Masked: Customer <SAP_CUSTOMER_ID> called about blocked order 9000001. Contact: <PERSON>, <EMAIL_ADDRESS>, phone <PHONE_NUMBER>.

Text 2: 2 finding(s)
  SAP_SUPPLIER_ID   0.85  17300042
  SAP_PERSONNEL_ID  0.85  50000123
  Masked: Supplier <SAP_SUPPLIER_ID> asks why invoice 90012345 is parked. Approved by employee <SAP_PERSONNEL_ID>.

In the Unit 11 setup, the customer number stayed visible. Now it's masked. The order number 9000001 and invoice 90012345 stay, because they identify documents, not people.

  1. See what context does. Run:

    python unit11/privacy_lab.py detect "The account 10100001 is new"
    Text 1: 1 finding(s)
      SAP_CUSTOMER_ID   0.50  10100001
      Masked: The account <SAP_CUSTOMER_ID> is new

    The score is 0.50, the pattern's base score. In the sample, "Customer" in front raised it to 0.85. If you later set a minimum score, context words decide which matches survive.

If Presidio prints a long message about publicsuffix.org first, it's harmless; see the troubleshooting table.

Step 4: Minimise, then mask

  1. Build the prompt the assistant would send for customer 10100001:

    python unit11/privacy_lab.py prompt --customer 10100001

What success looks like:

SAP record has 11 fields; the prompt gets 4: Type, CreditLimit, CreditSegment, BlockReason
Prompt sent to the model:
  Question: Customer <SAP_CUSTOMER_ID> (<PERSON>) asks why order 9000001 is blocked.
  Record:   {"Type": "sole trader", "CreditLimit": 250000, "CreditSegment": "A", "BlockReason": "credit limit exceeded"}

Seven fields, including the IBAN, birthday and address, never leave. Masking then catches what the user typed. Minimisation did more work than masking here, and it can't miss.

Step 5: Create the assistant's stores and take inventory

  1. Create the made-up stores:

    python unit11/privacy_lab.py init
    Created privacy_store.json with prompt_log (3), vector_index (5), answer_cache (2), eval_set (2)
  2. List which personal data sits in each store:

    python unit11/privacy_lab.py inventory

What success looks like:

store          rows with PII  types found
prompt_log        3        2  PERSON, PHONE_NUMBER, SAP_CUSTOMER_ID
vector_index      5        4  EMAIL_ADDRESS, PERSON, SAP_CUSTOMER_ID
answer_cache      2        1  SAP_CUSTOMER_ID
eval_set          2        2  PERSON, SAP_CUSTOMER_ID

Note: vector_index also holds vectors computed from the same text. Treat them as personal data too.

Every store holds personal data, including the eval set, which was copied from real-looking questions. This table is the start of the assistant's record of processing: one row per store, with an owner and a retention period added by your team. Open unit11/privacy_store.json in VS Code to see the rows.

Step 6: Erase one customer from every store

  1. Do a dry run first. It shows what would be deleted without changing anything:

    python unit11/privacy_lab.py erase --customer 10100001 --dry-run

What success looks like:

Would delete 1 row(s) from prompt_log
Would delete 2 row(s) from vector_index
Would delete 1 row(s) from answer_cache
Would delete 1 row(s) from eval_set
Possible leftovers after erasure: 1
  check by hand: vector_index row 4
Dry run: nothing changed.
  1. Look at row 4 in unit11/privacy_store.json. It says "Ms Becker prefers calls after 4 pm…" and its customer field is null. When the note was chunked, this half lost the customer number. Exact matching on the ID, full name and email missed it; the wider surname search found it.

  2. Run the real erasure:

    python unit11/privacy_lab.py erase --customer 10100001
    Deleted 1 row(s) from prompt_log
    Deleted 2 row(s) from vector_index
    Deleted 1 row(s) from answer_cache
    Deleted 1 row(s) from eval_set
    Possible leftovers after erasure: 1
      check by hand: vector_index row 4
    Note: ERASURE_LOG_KEY is not set in .env; using a lab-only key.
    Logged to erasure_log.jsonl: {"date": "2026-10-08", "subject_hash": "be6d5e70af43", "deleted": {"prompt_log": 1, "vector_index": 2, "answer_cache": 1, "eval_set": 1}, "leftovers": 1}

    Your date will be today's. The erasure log proves what was deleted and when, without the name, email or customer number.

  3. Give the log its own secret. Open the .env file in your course folder and add a line with a long random text of your own:

    ERASURE_LOG_KEY=replace-this-with-a-long-random-text
  1. Run init and the erasure again. The note about the lab-only key is gone, and subject_hash changes, because the key changed.

Step 7: Score the detector

  1. Run the evaluation on ten labelled sentences:

    python unit11/privacy_lab.py evaluate

What success looks like:

Presidio defaults  recall 50%  precision 83%  (found 5, missed 5, false alarms 1)
    missed SAP_CUSTOMER_ID '10100001'
    missed SAP_CUSTOMER_ID '10100003'
    missed SAP_SUPPLIER_ID '17300042'
    missed SAP_PERSONNEL_ID '50000123'
    missed SAP_CUSTOMER_ID '10100002'
    false alarm PERSON 'offene Posten'
SAP-aware          recall 100%  precision 91%  (found 10, missed 0, false alarms 1)
    false alarm PERSON 'offene Posten'

Read it in two parts. Your recognizers lifted recall from 50% to 100% on this set. The false alarm comes from a German sentence: the English name model tagged "offene Posten" (open items) as a person. Ten sentences prove little; a real evaluation set has hundreds, in every language your users write.

Step 8: Keep personal data out of the log file

  1. Write two log lines through the redaction filter:

    python unit11/privacy_lab.py log-demo
    privacy_lab.log:
    INFO question from CC_MUELLER: Why is order 9000001 for <PERSON> blocked?
    INFO grounding: customer <SAP_CUSTOMER_ID>, email <EMAIL_ADDRESS>, limit 250000
  2. Compare without the filter:

    python unit11/privacy_lab.py log-demo --no-redact

    The same lines now show the name, customer number and email. Without the filter, every debug session copies personal data into a file that's rarely cleaned up.

Step 9: Save your work in Git

  1. The stores and logs hold (made-up) personal data, so keep them out of Git. Open .gitignore in the course folder, add these lines at the end and save:

    unit11/privacy_store.json
    unit11/privacy_lab.log
    unit11/erasure_log.jsonl
  2. Commit:

    git add .gitignore unit11/privacy_lab.py
    git commit -m "Unit 11: privacy lab"

How the code works

Part What it does
SAP_IDS, build_analyzer Your SAP ID formats as Presidio PatternRecognizers, each with a regex, a base score of 0.5 and context words in English and German.
CUSTOMERS Made-up business partner records standing in for S/4HANA master data. One sole trader, one company with a contact person.
cmd_detect Runs the SAP-aware analyzer and masks each finding with its type.
NEEDED_FOR_CREDIT_BLOCK, cmd_prompt Minimisation: an allowlist of fields for one task, then masking of the question.
cmd_init, fake_vector Writes four stores to privacy_store.json. The "vectors" are numbers computed from the text, standing in for embeddings.
cmd_inventory Scans every text field of every row and lists the personal data types per store.
identifiers, matches The values that identify one customer, and the rule for deleting a row: same customer field, or an exact identifier in the text.
residual_hits A wider, case-insensitive search after deletion, including the surname, for a human to review.
cmd_erase Deletes matching rows from every store, reports counts, and appends an HMAC-keyed erasure record. --dry-run changes nothing.
GOLD, score, cmd_evaluate Ten labelled sentences, including two with no personal data, and recall and precision for default and SAP-aware detection.
RedactPersonalData, cmd_log_demo A logging.Filter that masks each log message before the file handler writes it.

If something goes wrong

What you see What it means What to do
python: command not found or 'python' is not recognized Python isn't on your path, or .venv isn't active Turn on .venv (Step 1). On macOS/Linux without .venv, try python3.
ModuleNotFoundError: No module named 'presidio_analyzer' Presidio isn't in this environment Complete Step 2 of Set up for Unit 11.
ModuleNotFoundError: No module named 'dotenv' python-dotenv is missing Add python-dotenv to requirements.txt, then pip install -r requirements.txt.
OSError: [E050] Can't find model 'en_core_web_lg' Presidio's language model isn't downloaded Run python -m spacy download en_core_web_lg.
Long Exception reading Public Suffix List or ProxyError messages, then normal output A library behind Presidio tried to fetch a list online, which your network or proxy blocked, and used its built-in copy Harmless. Your text is never sent.
No store yet. Run: python unit11/privacy_lab.py init inventory or erase ran before init Run init first.
Unknown customer ... The lab only knows 10100001 and 10100002 Use one of those numbers.
Erasure shows 0 deletions You already erased that customer Run init to recreate the stores.
IndentationError or SyntaxError Part of the code was lost when pasting Delete the file contents, copy the whole block again, and save.
Your scores differ slightly A newer spaCy model or Presidio version tags names differently Compare the missed and false alarm lines; that's what to watch when you upgrade.
You never needed an API key Correct: no model is called ERASURE_LOG_KEY is a secret you invent, not a key from a service.

The SAP way

As of October 2026, SAP covers this topic with several separate offerings. Check current documentation and your contracts before you design around any of them.

Model calls through the generative AI hub

SAP Learning describes generative AI requests as routed through the generative AI hub to SAP-managed infrastructure or approved external providers, depending on configuration. It states that SAP applies contractual and technical controls so external providers don't use customer data to train or retrain their models, and that data is processed only to fulfil the request. Whether data is stored depends on the service and its configuration.

For your design, that means:

  • Ask for the service-level facts in writing: which provider and region serve each model you deploy, and what is stored and for how long. The SAP Learning pages are summaries, not contracts.
  • Pick the region deliberately. SAP lists regional deployment options as "subject to availability", and offers the EU AI Cloud for European requirements.
  • Treat optional learning features as a decision. SAP says some services can improve from customer data where the customer enables it. Record who decided, and why.

Masking in the orchestration service

The orchestration service's data masking module anonymizes or pseudonymizes sensitive data before the model call. SAP Learning notes that unmasking in the response works only with pseudonymization. The guardrails topic covers configuration, including custom patterns for company-specific IDs and masking of grounding input. Add your SAP ID formats there, exactly as SAP_IDS does in the lab, and test them on your own data.

SAP Data Privacy Integration

SAP Data Privacy Integration offers anonymization of unstructured data such as free text, files and images. Its pseudonymization returns a metadata report that links each pseudonym to the real value. That report is the re-identification key, so store and delete it like personal data. SAP lists the offering under an enterprise service plan; its page notes that AI units "are not currently required", subject to change.

Blocking and deletion in SAP

SAP's ILM material describes the lifecycle: when a business purpose ends, personal data must be deleted, unless retention rules apply; then it is blocked until deletion. Two ILM functions carry this: retention management for transactional documents and simplified blocking and deletion for master data such as business partners. SAP's example uses a ten-year retention period from the end of business.

On SAP BTP, SAP Data Retention Manager is a reuse service for deletion orchestration. You define business purposes, then residence and retention rules per legal ground. It identifies data subjects that can be blocked or deleted and triggers deletion once the end of purpose is reached.

What this means for a side-by-side assistant:

  • Respect blocking at read time. If the assistant reads SAP through APIs with the user's identity, SAP's own authorization checks decide whether blocked data is visible, as in Grounding on SAP data with authorizations. Test this with a blocked test partner; don't assume it.
  • Mirror deletion in your stores. Your index, cache and logs are outside ILM. Give them an erasure path like the lab's, keyed on the business partner ID, and connect it to the same process that deletes in SAP. Whether your app can register with Data Retention Manager depends on its integration options; check its documentation.
  • Keep retention in step. A prompt log kept for two years about data SAP blocks after six months defeats the purpose of blocking.

Read Access Logging

SAP's Read Access Logging records who accessed sensitive data and when, including the content of UI elements, usually to meet legal requirements. It records the identity that made the access. If your assistant reads SAP with a shared technical user, and RAL is set up for the channel it uses, the log names the technical user, not the person who asked. Check with your SAP security team which channels RAL covers in your system. That's one more reason to propagate the user's identity, and to keep your own audit record of which user asked about which business partner.

Licensing and cost

Masking in orchestration is metered on top of model tokens, as covered in the guardrails topic. Data Privacy Integration, ILM and Data Retention Manager are separate products with their own licensing. Ask your SAP account team for current terms; this topic doesn't quote prices.

Build vs. SAP

Need Build it yourself SAP offering Use
Mask personal data before the model Presidio with custom recognizers Orchestration data masking SAP masking in production on the generative AI hub; Presidio to learn, test and cover side paths
Anonymize free text, files, images as a service Presidio analyzer and anonymizer SAP Data Privacy Integration DPI when you need a managed service on BTP; Presidio when you control the runtime
Field minimisation An allowlist in your prompt builder Not a module; your code Always your code
Block and delete in SAP Not appropriate ILM, simplified blocking and deletion Always SAP
Delete from AI stores Erasure by subject ID across stores, as in the lab Data Retention Manager for BTP apps, where integration fits Your code, triggered by the same process as SAP deletion
Read audit Your own access log in the app Read Access Logging in SAP Both: RAL for SAP, your log for the assistant
Detector quality A labelled evaluation set, as in the lab Not provided as a metric Always your eval set

Production concerns

  • Security and SAP authorizations. Least privilege comes first: the assistant reads only what the user may read, through the user's identity. Masking and minimisation reduce what reaches the model; they don't replace authorization checks. Agent permissions are covered later in Unit 11.
  • Data protection impact assessment. An assistant that processes personal data at scale is a candidate for a formal privacy review. Bring the store table from Step 5 and the data flow diagram to your data protection officer before go-live, not after.
  • Evaluation. Keep the detector evaluation in your evaluation harness. Track recall per data type and language. Re-run it on every Presidio, spaCy or SAP masking change.
  • Retention. Give every store a retention period and an automatic purge. Default prompt logs to IDs and decisions, not full text. Make full-text debug logging a time-boxed switch with an owner.
  • Traces. OpenTelemetry-based tracing tools can record full prompts in spans. Keep content capture off in production unless the trace store meets the same rules as your prompt log. See Observability for AI systems.
  • Eval sets and fine-tuning. Build eval sets from made-up or masked data. OWASP lists training data that includes sensitive data as a disclosure path, so mask before fine-tuning, which comes up again in Unit 12.
  • Caches. A cached answer must not be served to a user who couldn't see the source data, and must expire when the source is blocked or deleted. See Caching for LLM applications.
  • Operations. Run a test erasure every quarter on a test subject, and measure the time it takes. Keep the erasure log; it is your evidence.
  • Clean core. Every control in the lab lives in your side-by-side app. ILM, blocking and Read Access Logging are standard SAP functions you configure, not modify.

Pitfalls

  • Masking without minimising. A detector misses some values. A field that never leaves SAP can't be missed.
  • Trusting generic detectors with SAP data. Out of the box, Presidio missed every SAP ID in the lab's evaluation.
  • Losing the subject ID at ingestion. Chunking split a note and dropped the customer number. Without metadata on every chunk, erasure relies on text search and a human.
  • Deleting text but not vectors. An orphaned embedding can be inverted back toward its text.
  • Plain hashes as pseudonyms. A hash of an 8-digit ID can be reversed by trying every number. Use a keyed hash and protect the key.
  • Logging the masked prompt and the map together. The placeholder map turns the masked log back into personal data.
  • English models on German text. The lab tagged "offene Posten" as a person. Test in every language your users write.
  • Real data in eval sets. Teams copy production questions into tests, which then live in Git forever.

Exercise

Fix the leftover from Step 6 at its source: give every chunk its data subject's ID when it is written. Your findings feed the red-teaming topic at the end of Unit 11.

  1. Open unit11/privacy_lab.py. In cmd_init, find this line:

            (None, "Ms Becker prefers calls after 4 pm and asked us not to email her reminders."),
  2. Replace None with "10100001", so the line starts with ("10100001", "Ms Becker. Keep the rest of the line. Save.

  3. Recreate the stores and erase again:

    python unit11/privacy_lab.py init
    python unit11/privacy_lab.py erase --customer 10100001
  4. Check the output. vector_index now shows Deleted 3 row(s), and the leftovers line reads Possible leftovers after erasure: 0.

  5. Run python unit11/privacy_lab.py inventory and check that answer_cache shows - under types found: the only cached answer with personal data was this customer's.

  6. Open unit11/findings.md and add two rows: F-07, area data security, what happened chunking dropped the customer ID; erasure missed a chunk naming the customer, severity high, status fixed: subject ID on every chunk; and F-08, area data security, what happened default detector missed all SAP IDs and tagged German text as a name, severity medium, status mitigated: SAP recognizers; German eval cases open.

  7. Commit: git add unit11/privacy_lab.py unit11/findings.md, then git commit -m "Unit 11: subject ID on every chunk".

Done when the erasure reports 3 rows deleted from vector_index and 0 leftovers, and findings.md has rows F-07 and F-08.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the lab build the prompt from NEEDED_FOR_CREDIT_BLOCK before masking?

    Answer: B. Minimisation is an allowlist, so the IBAN, birthday and address never reach the prompt. Masking is a detector with misses, so it works best on what minimisation leaves.
  2. 2What does context=["customer", "sold-to", ...] change in the SAP_CUSTOMER_ID recognizer?

    Answer: C. The pattern alone gave 0.50, and "Customer" in front raised it to 0.85. Context changes confidence, so a minimum score would decide which matches survive.
  3. 3The dry run reports one leftover in vector_index row 4. What caused it?

    Answer: D. The note was chunked, and the second half had customer: null and no full name or email. Exact matching missed it; the wider surname search flagged it for review. The exercise fixes it at ingestion.
  4. 4Why does the erasure log store an HMAC of the customer number instead of a plain SHA-256 hash?

    Answer: B. There are only 100 million 8-digit numbers, so anyone can hash them all and find the match. A keyed hash needs the secret from .env, which outsiders don't have. HMAC is not encryption and can't be decrypted.
  5. 5A team deletes chunk text from its vector index but keeps the embeddings "for search quality". What is the problem?

    Answer: C. Vec2Text recovered 92% of short inputs exactly and 89% of full names from embedded clinical notes. Delete text and vector together, keyed on the subject ID.
  6. 6Your evaluation shows "false alarm PERSON 'offene Posten'". What should you do?

    Answer: B. The English name model misread German words. Removing PERSON would drop real names, and blanket score changes hide the problem. More German cases show how big it is before you change the setup.
  7. 7Your assistant reads S/4HANA with a shared technical user, and RAL covers that channel. Whose access does the log show?

    Answer: D. RAL records the identity that made the access, where it is set up for that channel. With a shared technical user, the person is lost, so propagate the user's identity and keep your own audit record in the assistant.
  8. 8SAP Data Privacy Integration pseudonymizes a note and returns a metadata report. How should you treat that report?

    Answer: C. The report is the re-identification key. Storing it with the masked note undoes the masking, so protect, retain and delete it like the original data.
  9. 9A customer's data is blocked in SAP after their purpose ended. What should the assistant's own stores do?

    Answer: B. ILM blocks and deletes SAP's own data, not copies in your app. Your logs, index and cache need a matching retention rule and an erasure path triggered by the same process.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in