Before an AI assistant can answer from your manuals, someone has to prepare them. That means two jobs.
First, turn files into clean text. A PDF looks tidy on screen, but inside it is mostly loose words placed on a page. Page headers, page numbers, broken words and tables come out jumbled.
Second, cut the text into chunks: passages small enough to search precisely and to fit in the model's prompt, but large enough to make sense alone. Each chunk gets labels, called metadata: which document, which page, which section, which process, which version.
These choices decide what the assistant can find. If the answer is split across two chunks, or a table loses its header, the right passage never reaches the model. No prompt or bigger model fixes that.
In RAG fundamentals you saw the rule: check retrieval before generation. Chunking is the biggest lever on retrieval that your own team controls.
Take the running procure-to-pay example. An accounts payable clerk asks, "What is the price tolerance for company code 2000?" The answer sits in a table in the invoice handling manual. Cut that manual badly, and the table rows land in one chunk while the words "price tolerance" sit in another. The search finds the wrong passage, and the assistant answers from it, confidently.
Three business effects follow:
Quality. Badly prepared documents produce wrong or "can't find" answers, and users stop trusting the assistant.
Cost. Every retrieved chunk is paid for in every model call. Chunks that are too big, or that repeat each other, raise the bill without raising quality.
Control. Metadata is what lets you filter by company code or process, show the page a fact came from, and remove old versions. Without it you can't audit an answer.
SAP AI Core grounding (generative AI hub). The Pipeline API reads documents from a connected repository, segments them into chunks and creates embeddings for you. SAP Learning describes the document store as holding files such as PDFs and text files. This is the low-effort path; you accept SAP's chunking.
The Vector API is the path when you want control. Your team prepares the chunks and uploads them. Since February 2026, SAP documents metadata on documents, collections and chunks, and filtering on that metadata at search time.
SAP HANA Cloud includes a text splitter in its Python machine learning client (hana-ml), with settings for chunk size, overlap and splitting method. Useful when the text already lives in the database.
SAP Document AI is a different tool for a different job. It extracts structured fields from documents such as invoices and purchase orders. If you want a field value posted to SAP, that is extraction, not RAG.
Paragraphs first, then lines, sentences, words, until pieces fit
General text without clear headings
Can still separate a table from its heading
Structure-aware
At headings; keeps tables whole; labels each chunk with its section
Manuals, policies, process guides
Needs clean structure, so preparation matters more
Two settings go with every approach:
Chunk size: how long a chunk may be. Small chunks match specific questions; large chunks carry more context but more noise and cost.
Overlap: how much of the previous chunk is repeated at the start of the next. It protects answers that fall on a boundary, at the price of storing and sending duplicates.
There is no universal best setting. Chroma's July 2024 study found differences of up to 9 percent in recall between strategies on the same data. The only reliable answer is to test with real questions on your own documents, which the deep layer shows how to do in an afternoon.
"Just upload the PDFs." PDFs carry no notion of paragraphs, headers or tables. Somebody has to rebuild that structure, or the assistant reads noise.
"Bigger chunks are safer because they hold more." They also hold more irrelevant text, cost more per call, and can exceed what the embedding model reads. Many small embedding models cut off long input.
"More overlap is always better." Overlap duplicates text. Chroma's study found that reducing overlap improved efficiency scores, and a common default of large chunks with heavy overlap had the lowest efficiency scores.
"Chunking is a one-time technical detail." Changing it means re-indexing everything, and it moves answer quality. Treat it as a design decision with a test behind it.
"RAG can read our invoices into SAP." Pulling fields from invoices is extraction, a job for SAP Document AI, not retrieval.
Pick one answer for each question. The explanation appears after you choose.
1Why does chunking matter so much for answer quality?
Answer: B. The model answers from the chunks it is given. If chunking splits or buries the answer, retrieval misses it, and no prompt or larger model can recover it.
2A clerk asks for the price tolerance of company code 2000, and the answer is in a table. What usually goes wrong with careless chunking?
Answer: C. Fixed-size cutting can place "price tolerance" in one chunk and the rows in another. The search then matches the wrong passage, and the answer is built on it.
3Your team proposes large chunks with heavy overlap "to be safe". What is the best response?
Answer: D. Bigger chunks and more overlap raise cost and noise, and Chroma's study found such a default scored lowest on efficiency. Overlap can still help with boundaries, so the answer is a measured test, not a rule.
4Which metadata most helps you audit an answer and retire old content?
Answer: A. Source, page and section let you show where a fact came from. Version and valid-from date let you find and remove outdated chunks when a manual changes.
5You want full control over how your manuals are cut, but you will run on SAP AI Core. Which path fits?
Answer: B. The Pipeline API chunks for you, which is less effort but less control. The Vector API takes chunks you prepared, with your metadata, which SAP documents for filtering since February 2026.
6Finance wants invoice amounts and vendor numbers read from PDFs and posted to SAP. What is this?
Answer: D. Pulling structured fields from invoices is extraction. SAP Document AI is built for that; RAG chunking is for answering questions from text.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
A chunk is the unit of evidence. Retrieval ranks chunks, the prompt carries chunks, citations point to chunks. So each chunk must pass one test: could a colleague answer the question from this chunk alone, and say where it came from?
That test explains every rule in this topic:
Cut at meaning boundaries (sections, paragraphs, table ends), not at character counts, so a chunk holds one complete idea.
Put context into the text (document title and section path), so "the tolerance is 2 percent" still says of what.
Keep labels outside the text (metadata: page, version, process), so code can filter, cite and expire chunks.
Measure with questions whose answers you know, because intuition about chunk size is unreliable.
Document preparation is what makes those rules possible. You can't cut at headings you can't see.
The indexing path from RAG fundamentals had "ingest" and "chunk" as two boxes. In practice there are five steps, and most failures happen in the first two.
flowchart LR
P[PDF, Word, HTML] --> X[1 Extract text]
X --> C[2 Clean and rebuild structure]
C --> S[3 Split into chunks]
S --> M[4 Add context and metadata]
M --> E[5 Embed and index]
Q[Test questions] --> V[Measure hit rate]
E --> V
V -.adjust.-> S
The pypdf documentation is blunt: a PDF has no semantic layer. It holds instructions to draw text at positions. There is no "header", "paragraph" or "table", and in a table, each cell is just text placed at coordinates. Extraction tools rebuild reading order from those positions:
Plain extraction joins text in order. Table cells come out separated by single spaces, so "2000 France 5 percent Controller" loses its columns.
Layout extraction (pypdf's extraction_mode="layout") keeps the spacing as rendered, so columns stay visible as wide gaps.
Scanned PDFs are images. They contain no text to extract; they need OCR first. pypdf recommends using OCR software directly in that case.
Extracted text carries noise that hurts retrieval:
Noise
Example
Fix
Repeated page headers and footers
"ACME Corp - Internal …", "Page 2 of 3" on every page
Drop lines that repeat on every page (compare with digits masked)
Hyphenated line breaks
"pur-" / "chase order"
Re-join a word when a line ends in a hyphen and the next starts lowercase
Hard line breaks inside paragraphs
One sentence spread over three lines
Join lines until a blank line
Tables as loose text
Columns separated only by spaces
Detect column gaps and rebuild rows as Markdown tables
Lost headings
"2.2 Price differences" looks like any other line
Detect numbering patterns or font sizes; mark as headings
The output is clean Markdown: headings, paragraphs and tables you can read and check. This is the step teams skip, and it is the step that decides whether structure-aware chunking is possible.
SAP documentation is usually easier: help.sap.com pages are HTML, which already marks headings and tables. For SAP's own product documentation, SAP's grounding module can also search help.sap.com directly, as RAG fundamentals showed. Your own process guides, often PDFs or Word files in SharePoint, are where preparation work lands.
Fixed size. Every N characters, stepping back by the overlap. Simple and fast, but it cuts mid-word and mid-table.
Recursive. LangChain's RecursiveCharacterTextSplitter tries separators in order, by default paragraph break, line break, space, then single characters, until each piece fits. It then packs pieces up to the size. It respects paragraphs, but it doesn't know a table belongs to the heading above it.
Structure-aware. Cut at headings. Keep each table whole with its header row. Split a long section by paragraph, repeating the section label on each part.
Size is measured in characters or tokens. Embedding models have an input limit: the model card for all-MiniLM-L6-v2, the Unit 3 model, says input longer than 256 word pieces is truncated. Anything past that point is silently ignored when the chunk is embedded, though it still goes into the prompt.
Context in the text. Prefix each chunk with Document title > Section. It changes the embedding, so it helps matching. Anthropic's 2024 "contextual retrieval" write-up takes this further: a model writes a short, chunk-specific context line before embedding. In their tests, failed retrievals at top 20 fell by 35 percent with contextual embeddings alone, and by 49 percent when combined with keyword search.
Metadata beside the text. Source, page, section, process, version, valid-from date, language. It doesn't change the embedding. It lets code filter ("only procure-to-pay"), cite ("manual p. 2, section 3.1") and expire ("drop version 3.1").
You can't judge chunking by looking at it. You need test questions with known answers. For chunking, the most useful check is at chunk level, not document level: is there one retrieved chunk that contains the whole answer? Chroma's technical report goes further and scores at token level: recall (how much of the answer was retrieved) and precision (how much of what was retrieved was answer). Two of its findings are worth remembering:
Recursive splitting at about 200 tokens with no overlap performed well across their data.
A popular default of 800 tokens with 400 overlap had slightly below-average recall and the lowest efficiency scores.
Your documents are not theirs, so treat these as starting points to test, not settings to copy.
You will build chunk_lab.py, one script that takes a PDF manual through every step: extract, clean, chunk three ways, and measure. It writes its own sample PDF, a made-up vendor invoice handling manual with page headers, a hyphenated word and two tables, so you see real extraction problems without hunting for a file. At the end it exports the winning chunks with metadata to chunks.jsonl, which later Unit 7 topics load into a vector store.
flowchart LR
A[Sample PDF<br/>or your PDF] --> B[extract<br/>raw.txt, clean.md]
B --> C{strategy}
C --> F[fixed]
C --> R[recursive]
C --> H[headings]
F --> M[compare<br/>hit rate, cost]
R --> M
H --> M
H --> J[export<br/>chunks.jsonl]
In VS Code, right-click the unit07 folder, choose New File, name it chunk_lab.py, paste the code below and save.
"""Chunk lab: turn a PDF manual into clean text, cut it into chunks three ways, and measure which
way lets retrieval find the answers. Builds on rag_minimal.py from "RAG fundamentals".
Run from your course folder (orchestrate-course), with .venv turned on:
python unit07/chunk_lab.py extract # PDF -> raw text and clean Markdown
python unit07/chunk_lab.py show --strategy headings # print the chunks one strategy makes
python unit07/chunk_lab.py show --strategy fixed --source raw --question "When are vendors paid?"
python unit07/chunk_lab.py compare # every strategy against the test questions
python unit07/chunk_lab.py export --strategy headings # write chunks.jsonl for later topics
Add --offline to use toy embeddings instead of the Unit 3 model. Add --pdf PATH to use your own PDF.
"""
import argparse
import hashlib
import json
import math
import re
import sys
import textwrap
from pathlib import Path
HERE = Path(__file__).resolve().parent
DOCS = HERE / "chunk_docs"
SAMPLE_PDF = DOCS / "p2p-invoice-manual.pdf"
MODEL = "sentence-transformers/all-MiniLM-L6-v2" # the Unit 3 model
# ---------- the made-up sample manual (procure-to-pay), written to a real PDF ----------
HEADER = "ACME Corp - Internal - Vendor invoice handling manual"
SAMPLE_PAGES = [
[("title", "Vendor invoice handling manual"),
("p", "Version 3.2. Valid from 2026-07-01. Owner: accounts payable. This manual replaces version 3.1."),
("h", "1 Purpose and scope"),
("p", "This manual explains why vendor invoices are blocked for payment, who may release them and "
"how fast. It applies to all company codes in Europe. It does not cover travel expenses."),
("h", "2 Why invoices are blocked"),
("p", "The system compares each invoice with the purchase order and the goods receipt. This "
"comparison is called the three-way match. When a line fails the match, the invoice is "
"blocked for payment until someone releases it."),
("h", "2.1 Quantity differences"),
("p", "If the invoiced quantity is higher than the quantity received, the invoice is blocked. "
"The buyer checks with the warehouse whether a second delivery is on its way."),
("h", "2.2 Price differences"),
("lines", ["An invoice is blocked when the invoice price differs from the pur-",
"chase order price by more than the tolerance for the company code:"]),
("table", [["Company code", "Tolerance", "Approver"],
["1000 Germany", "2 percent", "AP team lead"],
["2000 France", "5 percent", "Controller"],
["3000 Spain", "3 percent", "AP team lead"]])],
[("h", "2.3 Missing goods receipt"),
("p", "If no goods receipt has been posted, the invoice waits for the goods receipt. Do not "
"release it by hand. Ask the requester to confirm the delivery in the system first."),
("h", "3 Releasing blocked invoices"),
("h", "3.1 Who may release"),
("p", "The accounts payable clerk may release blocks up to 10,000 euros per invoice. Above "
"10,000 euros, only the head of accounts payable may release the block. Nobody may "
"release an invoice they created themselves."),
("h", "3.2 Release deadlines"),
("p", "Every block must be worked within the deadline below, or it is escalated:"),
("table", [["Block reason", "Release within", "Escalate to"],
["Quantity difference", "5 working days", "Buyer"],
["Price difference", "3 working days", "Controller"],
["Missing goods receipt", "10 working days", "Requester's manager"]]),
("p", "Escalations are sent by e-mail every Monday morning.")],
[("h", "4 Duplicate invoices"),
("p", "An invoice is a suspected duplicate when it has the same vendor, the same reference "
"and the same amount as an invoice from the last 90 days. Suspected duplicates are "
"never paid automatically. The clerk compares both documents and rejects one."),
("h", "5 Payment runs"),
("p", "Released invoices are paid in the payment run every Tuesday and Friday. An invoice "
"released after 12:00 on a payment day waits for the next run. Urgent payments outside "
"the run need approval from the treasury team."),
("h", "6 Archiving"),
("p", "Paid invoices are archived after the payment run. Keep invoices for ten years, or "
"longer where local law requires it.")],
]
# Test questions, each with phrases that must ALL appear in one retrieved chunk to count as found.
EVAL = [
("What is the price tolerance for company code 2000?", ["2000", "5 percent"]),
("Who may release an invoice above 10,000 euros?", ["head of accounts payable"]),
("When are vendors paid?", ["tuesday and friday"]),
("How long must invoices be kept?", ["ten years"]),
("What should I do when the goods receipt is missing?", ["waits for the goods receipt"]),
("How fast must a quantity difference block be released?", ["quantity difference", "5 working days"]),
("How do we recognize a duplicate invoice?", ["same vendor", "same amount"]),
("Why is an invoice blocked when the price is different from the purchase order?",
["invoice price differs from the purchase order price"]),
]
def write_pdf(pages: list, path: Path) -> None:
"""Write a small PDF by hand (no extra library). Each page gets a header and a page-number footer,
like a real manual. Table cells are placed at fixed x positions, as most PDFs do."""
def esc(s):
return s.replace("\\", "\\\\").replace("(", "\\(").replace(")", "\\)")
objects = {1: "<< /Type /Catalog /Pages 2 0 R >>",
3: "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>"}
kids = []
for number, page in enumerate(pages, 1):
items, y = [(50, 805, 8, HEADER)], 770
for kind, value in page:
if kind == "title":
items.append((50, y, 18, value)); y -= 30
elif kind == "h":
y -= 8; items.append((50, y, 13, value)); y -= 22
elif kind in ("p", "lines"):
for line in (textwrap.wrap(value, 92) if kind == "p" else value):
items.append((50, y, 10, line)); y -= 14
y -= 10
elif kind == "table":
for row in value:
for x, cell in zip((50, 230, 370), row):
items.append((x, y, 10, cell))
y -= 15
y -= 10
items.append((260, 30, 8, f"Page {number} of {len(pages)}"))
stream = "\n".join(f"BT /F1 {s} Tf {x} {y} Td ({esc(t)}) Tj ET" for x, y, s, t in items)
page_id, content_id = 2 + 2 * number, 3 + 2 * number
kids.append(f"{page_id} 0 R")
objects[page_id] = ("<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] "
f"/Resources << /Font << /F1 3 0 R >> >> /Contents {content_id} 0 R >>")
objects[content_id] = (f"<< /Length {len(stream.encode('latin-1'))} >>\n"
f"stream\n{stream}\nendstream")
objects[2] = f"<< /Type /Pages /Kids [{' '.join(kids)}] /Count {len(pages)} >>"
data, offsets = b"%PDF-1.4\n", []
for i in range(1, max(objects) + 1):
offsets.append(len(data))
data += f"{i} 0 obj\n{objects[i]}\nendobj\n".encode("latin-1")
xref = len(data)
data += f"xref\n0 {len(offsets) + 1}\n0000000000 65535 f \n".encode()
data += "".join(f"{o:010d} 00000 n \n" for o in offsets).encode()
data += f"trailer\n<< /Size {len(offsets) + 1} /Root 1 0 R >>\nstartxref\n{xref}\n%%EOF\n".encode()
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(data)
# ---------- 1. extract: PDF -> text per page ----------
def extract(pdf: Path, mode: str) -> list:
"""Return one string per page. mode 'plain' is pypdf's default; 'layout' keeps column spacing."""
try:
from pypdf import PdfReader
except ImportError:
sys.exit("pypdf is not installed. Add 'pypdf' to requirements.txt and run: pip install -r requirements.txt")
try:
reader = PdfReader(pdf)
return [page.extract_text(extraction_mode=mode) or "" for page in reader.pages]
except Exception as error:
sys.exit(f"Could not read {pdf.name}: {type(error).__name__}: {error}")
# ---------- 2. clean: remove headers and footers, rebuild paragraphs, headings and tables ----------
def clean(pages: list) -> list:
"""Turn layout-mode page texts into blocks: {'kind': 'h'|'p'|'table', 'text', 'page'}."""
def shape(line): # "Page 2 of 3" and "Page 3 of 3" have the same shape
return re.sub(r"\d+", "#", line.strip())
seen = {}
for text in pages:
for line in {shape(l) for l in text.splitlines() if l.strip()}:
seen[line] = seen.get(line, 0) + 1
repeated = {line for line, count in seen.items() if len(pages) > 1 and count >= len(pages)}
blocks = []
for number, text in enumerate(pages, 1):
paragraph = []
def flush():
if paragraph:
joined = paragraph[0]
for nxt in paragraph[1:]:
if re.search(r"[a-z]-$", joined) and nxt[:1].islower():
joined = joined[:-1] + nxt # re-join a word split across lines
else:
joined += " " + nxt
blocks.append({"kind": "p", "text": joined, "page": number})
paragraph.clear()
for raw in text.splitlines():
line = raw.strip()
if not line or shape(line) in repeated:
flush()
continue
cells = re.split(r"\s{3,}", line)
if len(cells) >= 2: # columns -> a table row
flush()
row = "| " + " | ".join(cells) + " |"
if blocks and blocks[-1]["kind"] == "table":
blocks[-1]["text"] += "\n" + row
else:
blocks.append({"kind": "table", "text": row, "page": number})
elif re.match(r"^\d+(\.\d+)*\s+[A-Z][^.]{0,80}$", line): # "2.2 Price differences"
flush()
blocks.append({"kind": "h", "text": line, "page": number})
else:
paragraph.append(line)
flush()
return blocks
def to_markdown(blocks: list) -> str:
"""Render blocks as Markdown, so you can read what the chunker will see."""
out = []
for b in blocks:
if b["kind"] == "h":
out.append("#" * (b["text"].split()[0].count(".") + 2) + " " + b["text"])
elif b["kind"] == "table":
rows = b["text"].split("\n")
divider = "|" + "---|" * (rows[0].count("|") - 1)
out.append("\n".join([rows[0], divider] + rows[1:]))
else:
out.append(b["text"])
return "\n\n".join(out) + "\n"
# ---------- 3. chunk: three strategies ----------
def chunk_fixed(text: str, size: int, overlap: int) -> list:
"""Cut every `size` characters, stepping back `overlap` characters each time. Ignores words."""
step = max(1, size - overlap)
return [text[i:i + size] for i in range(0, len(text), step) if text[i:i + size].strip()]
def chunk_recursive(text: str, size: int, overlap: int,
separators=("\n\n", "\n", ". ", " ")) -> list:
"""Split at the biggest separator that works (paragraph, line, sentence, word), then pack the
pieces into chunks of at most `size` characters, repeating up to `overlap` characters of the
previous chunk's end. The same idea as LangChain's RecursiveCharacterTextSplitter."""
def pieces(t, seps):
if len(t) <= size or not seps:
return [t]
sep, rest = seps[0], seps[1:]
parts = t.split(sep)
out = []
for i, part in enumerate(parts):
part = part + (sep if i < len(parts) - 1 else "")
out.extend(pieces(part, rest) if len(part) > size else [part])
return out
chunks, current = [], []
for piece in pieces(text, separators):
if current and sum(map(len, current)) + len(piece) > size:
chunks.append("".join(current).strip())
tail = []
for prev in reversed(current): # carry the end of the chunk forward
if sum(map(len, tail)) + len(prev) > overlap:
break
tail.insert(0, prev)
current = tail
current.append(piece)
if current:
chunks.append("".join(current).strip())
return [c for c in chunks if c]
def chunk_headings(blocks: list, title: str, size: int) -> list:
"""Structure-aware: one chunk per section, prefixed with 'title > section'. Tables stay whole
with their header row; long sections are split by paragraph, each part keeping the prefix."""
sections, current = [], {"heading": "", "parts": [], "page": blocks[0]["page"] if blocks else 1}
for b in blocks:
if b["kind"] == "h":
if current["parts"] or current["heading"]:
sections.append(current)
current = {"heading": b["text"], "parts": [], "page": b["page"]}
else:
current["parts"].append(b["text"])
sections.append(current)
chunks = []
for s in sections:
if not s["parts"]:
continue # a heading with no text of its own
heading = s["heading"] or "Introduction"
prefix = f"{title} > {heading}\n"
group = ""
for part in s["parts"]:
if group and len(prefix) + len(group) + len(part) > size:
chunks.append({"text": prefix + group, "section": heading, "page": s["page"]})
group = ""
group = (group + "\n\n" + part).strip()
if group:
chunks.append({"text": prefix + group, "section": heading, "page": s["page"]})
return chunks
# ---------- 4. embed (same choices as rag_minimal.py) ----------
STOP = set("a an and are as at be by can do does for from has how i in is it it's may of on or our "
"should so that the their this to was we what when where which who why will with".split())
def toy_embed(text: str, size: int = 256) -> list:
"""Hash each word into one of 256 slots, skipping common words. Matches words, not meaning."""
vector = [0.0] * size
for word in re.findall(r"[a-z0-9]+", text.lower()):
if word in STOP:
continue
vector[int(hashlib.md5(word.encode()).hexdigest(), 16) % size] += 1.0
length = math.sqrt(sum(v * v for v in vector)) or 1.0
return [v / length for v in vector]
def get_embedder(offline: bool):
"""Return (embed function, token counter or None)."""
if offline:
return (lambda texts: [toy_embed(t) for t in texts]), None
try:
from sentence_transformers import SentenceTransformer
except ImportError:
sys.exit("sentence-transformers is not installed. See Set up for Unit 3, or add --offline.")
print(f"Loading {MODEL} (from disk if you ran Unit 3)...")
try:
model = SentenceTransformer(MODEL)
except Exception as error:
sys.exit(f"Could not load the model ({type(error).__name__}). Check your network, or add --offline.")
limit = model.max_seq_length
def too_long(texts):
return sum(1 for t in texts if len(model.tokenizer(t)["input_ids"]) > limit), limit
return (lambda texts: model.encode(texts, normalize_embeddings=True).tolist()), too_long
def top_k(question_vec: list, chunk_vecs: list, k: int) -> list:
"""Indexes of the k chunks with the highest cosine similarity (vectors are already normalized)."""
scores = [sum(a * b for a, b in zip(question_vec, v)) for v in chunk_vecs]
return sorted(range(len(scores)), key=lambda i: -scores[i])[:k]
# ---------- 5. measure ----------
def norm(text: str) -> str:
return re.sub(r"\s+", " ", text.lower())
def evaluate(texts: list, embed, k: int) -> tuple:
"""Hit rate: share of EVAL questions where one top-k chunk holds all expected phrases.
Also returns the average number of characters the k chunks would add to each prompt."""
vecs = embed(texts)
q_vecs = embed([q for q, _ in EVAL])
hits, sent, misses = 0, 0, []
for (question, phrases), qv in zip(EVAL, q_vecs):
chosen = [texts[i] for i in top_k(qv, vecs, k)]
sent += sum(len(c) for c in chosen)
if any(all(p in norm(c) for p in phrases) for c in chosen):
hits += 1
else:
misses.append(question)
return hits, sent / len(EVAL), misses
# ---------- running it ----------
def load(pdf: Path) -> tuple:
"""Return (raw text, clean blocks, title). Writes the sample PDF on first use."""
if pdf == SAMPLE_PDF and not pdf.exists():
write_pdf(SAMPLE_PAGES, pdf)
print(f"Wrote the sample manual to {pdf.relative_to(HERE.parent)}")
if not pdf.exists():
sys.exit(f"File not found: {pdf}")
raw = "\n".join(extract(pdf, "plain"))
blocks = clean(extract(pdf, "layout"))
if not raw.strip():
sys.exit("No text found. The PDF may be a scan (an image of text); it needs OCR first.")
first = next((b["text"] for b in blocks if b["kind"] == "p"), pdf.stem)
title = first if len(first) < 60 else pdf.stem # the sample's title line is its first paragraph
return raw, blocks, title
def make_chunks(strategy: str, source: str, size: int, overlap: int, raw: str, blocks: list,
title: str) -> list:
"""Return chunk dicts with text, section and page for any strategy."""
if strategy == "headings":
return chunk_headings(blocks, title, size)
text = raw if source == "raw" else to_markdown(blocks)
cut = chunk_fixed if strategy == "fixed" else chunk_recursive
return [{"text": t, "section": "", "page": None} for t in cut(text, size, overlap)]
def metadata(pdf: Path, raw: str, strategy: str, chunk: dict) -> dict:
"""Labels stored with each chunk: for filters, citations and freshness checks."""
version = re.search(r"Version (\d+(?:\.\d+)*)", raw)
valid = re.search(r"Valid from (\d{4}-\d{2}-\d{2})", raw)
return {"source": pdf.name, "page": chunk["page"] or 0, "section": chunk["section"],
"process": "procure-to-pay", "doc_version": version.group(1) if version else "",
"valid_from": valid.group(1) if valid else "", "language": "en", "chunker": strategy}
def main() -> None:
parser = argparse.ArgumentParser(description="Prepare a PDF and compare chunking strategies.")
parser.add_argument("action", choices=["extract", "show", "compare", "export"])
parser.add_argument("--pdf", type=Path, default=SAMPLE_PDF, help="your own PDF (default: the sample)")
parser.add_argument("--strategy", choices=["fixed", "recursive", "headings"], default="headings")
parser.add_argument("--source", choices=["raw", "clean"], default="clean",
help="cut the raw extracted text or the cleaned text (fixed and recursive only)")
parser.add_argument("--size", type=int, default=600, help="maximum characters per chunk (default 600)")
parser.add_argument("--overlap", type=int, default=0, help="characters repeated between chunks (default 0)")
parser.add_argument("--top", type=int, default=3, help="chunks retrieved per question (default 3)")
parser.add_argument("--question", help="with show: also print the top chunks for this question")
parser.add_argument("--offline", action="store_true", help="toy embeddings; no model download")
args = parser.parse_args()
if args.overlap >= args.size:
sys.exit("--overlap must be smaller than --size.")
raw, blocks, title = load(args.pdf)
if args.action == "extract":
stem = args.pdf.with_suffix("")
Path(f"{stem}.raw.txt").write_text(raw, encoding="utf-8")
Path(f"{stem}.clean.md").write_text(to_markdown(blocks), encoding="utf-8")
kinds = {k: sum(b["kind"] == k for b in blocks) for k in ("h", "p", "table")}
print(f"Raw text: {len(raw)} characters -> {Path(f'{stem}.raw.txt').name}")
print(f"Clean Markdown: {kinds['h']} headings, {kinds['p']} paragraphs, {kinds['table']} tables "
f"-> {Path(f'{stem}.clean.md').name}")
print("Open both files side by side and compare them.")
return
if args.action == "show":
chunks = make_chunks(args.strategy, args.source, args.size, args.overlap, raw, blocks, title)
print(f"{args.strategy} on {args.source if args.strategy != 'headings' else 'clean'} text, "
f"size {args.size}, overlap {args.overlap}: {len(chunks)} chunks\n")
if args.question:
embed, _ = get_embedder(args.offline)
texts = [c["text"] for c in chunks]
order = top_k(embed([args.question])[0], embed(texts), args.top)
print(f'Top {args.top} for "{args.question}":\n')
chunks = [chunks[i] for i in order]
for n, c in enumerate(chunks, 1):
where = f"page {c['page']}, " if c["page"] else ""
print(f"--- chunk {n} ({where}{len(c['text'])} chars) ---\n{c['text']}\n")
return
if args.action == "compare":
if args.pdf != SAMPLE_PDF:
print("Note: the test questions are about the sample manual. Write your own EVAL for your PDF.")
embed, too_long = get_embedder(args.offline)
runs = [("fixed", "raw", 300, 0), ("fixed", "raw", 300, 100), ("fixed", "clean", 300, 0),
("recursive", "clean", 300, 50), ("recursive", "clean", 800, 100),
("fixed", "clean", 1500, 0), ("headings", "clean", 600, 0)]
print(f"{'strategy':<10} {'source':<6} {'size':>5} {'overlap':>7} {'chunks':>6} "
f"{'avg chars':>9} {'hit@' + str(args.top):>6} {'chars sent':>10}")
for strategy, source, size, overlap in runs:
chunks = make_chunks(strategy, source, size, overlap, raw, blocks, title)
texts = [c["text"] for c in chunks]
hits, sent, misses = evaluate(texts, embed, args.top)
print(f"{strategy:<10} {source:<6} {size:>5} {overlap:>7} {len(texts):>6} "
f"{sum(map(len, texts)) / len(texts):>9.0f} {f'{hits}/{len(EVAL)}':>6} {sent:>10.0f}")
if too_long:
count, limit = too_long(texts)
if count:
print(f" {count} chunk(s) exceed the model's {limit} tokens; their ends are ignored")
for question in misses:
print(f" MISS {question}")
print("\nhit@k: questions where one of the top chunks holds the answer. "
"chars sent: what those chunks add to every prompt.")
return
if args.action == "export":
chunks = make_chunks(args.strategy, args.source, args.size, args.overlap, raw, blocks, title)
out = HERE / "chunks.jsonl"
with out.open("w", encoding="utf-8") as f:
for n, c in enumerate(chunks, 1):
record = {"id": f"{args.pdf.name}#{n}", "text": c["text"],
"metadata": metadata(args.pdf, raw, args.strategy, c)}
f.write(json.dumps(record, ensure_ascii=False) + "\n")
print(f"Wrote {len(chunks)} chunks to {out.relative_to(HERE.parent)}")
print("First record:\n" + json.dumps(json.loads(out.read_text(encoding='utf-8').splitlines()[0]),
indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
The table after Step 7 explains each part.
#Step 3: Extract the PDF and compare raw with clean
Run:
python unit07/chunk_lab.py extract
Wrote the sample manual to unit07/chunk_docs/p2p-invoice-manual.pdf
Raw text: 2655 characters -> p2p-invoice-manual.raw.txt
Clean Markdown: 11 headings, 13 paragraphs, 2 tables -> p2p-invoice-manual.clean.md
Open both files side by side and compare them.
Open unit07/chunk_docs/p2p-invoice-manual.pdf in your PDF viewer. It looks like a normal three-page manual.
Open p2p-invoice-manual.raw.txt in VS Code. This is pypdf's plain extraction. Find the price table:
chase order price by more than the tolerance for the company code:
Company code Tolerance Approver
1000 Germany 2 percent AP team lead
2000 France 5 percent Controller
3000 Spain 3 percent AP team lead
Page 1 of 3
ACME Corp - Internal - Vendor invoice handling manual
2.3 Missing goods receipt
Four problems in eight lines: the word "purchase" is broken over two lines, the columns are gone, a page footer and the next page's header sit in the middle of the text, and nothing marks 2.3 as a heading.
Open p2p-invoice-manual.clean.md. The same part now reads:
### 2.2 Price differences
An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
| Company code | Tolerance | Approver |
|---|---|---|
| 1000 Germany | 2 percent | AP team lead |
| 2000 France | 5 percent | Controller |
| 3000 Spain | 3 percent | AP team lead |
### 2.3 Missing goods receipt
The clean() function read the layout extraction, dropped lines that repeat on every page, re-joined "pur-chase", turned column gaps into table rows and marked numbered lines as headings.
headings on clean text, size 600, overlap 0: 11 chunks
--- chunk 5 (page 1, 354 chars) ---
Vendor invoice handling manual > 2.2 Price differences
An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
| Company code | Tolerance | Approver |
| 1000 Germany | 2 percent | AP team lead |
| 2000 France | 5 percent | Controller |
| 3000 Spain | 3 percent | AP team lead |
(Only chunk 5 is shown here; you'll see all 11.) The table stays with its sentence and its section label, and the chunk knows its page.
Recursive chunks at 300 characters, with the top 3 for a question. Add --offline if the Unit 3 model can't load on your network:
python unit07/chunk_lab.py show --strategy recursive --size 300 --overlap 50 --question "What is the price tolerance for company code 2000?" --offline
--- chunk 1 (158 chars) ---
### 2.2 Price differences
An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
--- chunk 2 (211 chars) ---
| Company code | Tolerance | Approver |
|---|---|---|
| 1000 Germany | 2 percent | AP team lead |
| 2000 France | 5 percent | Controller |
| 3000 Spain | 3 percent | AP team lead |
### 2.3 Missing goods receipt
The table got separated from its sentence and glued to the next heading. Here both pieces still made the top 3, but a model now has to guess that the table belongs to "price differences".
Fixed-size chunks on the raw text, for a payment question:
python unit07/chunk_lab.py show --strategy fixed --source raw --question "When are vendors paid?" --offline --top 2
Look at where the chunks end. The first one stops in the middle of a time, after 12:, and the next starts with 00 on a payment day. Fixed-size cutting doesn't know what a word is.
strategy source size overlap chunks avg chars hit@3 chars sent
fixed raw 300 0 9 295 6/8 883
MISS How do we recognize a duplicate invoice?
MISS Why is an invoice blocked when the price is different from the purchase order?
fixed raw 300 100 14 279 7/8 822
MISS Why is an invoice blocked when the price is different from the purchase order?
fixed clean 300 0 9 290 5/8 878
MISS Who may release an invoice above 10,000 euros?
MISS What should I do when the goods receipt is missing?
MISS How do we recognize a duplicate invoice?
recursive clean 300 50 13 217 7/8 648
MISS When are vendors paid?
recursive clean 800 100 4 685 8/8 2085
fixed clean 1500 0 2 1306 7/8 2611
MISS Who may release an invoice above 10,000 euros?
headings clean 600 0 11 260 8/8 812
Read the table row by row:
hit@3 counts questions where one of the top 3 chunks contains all the expected phrases from EVAL.
chars sent is what those 3 chunks add to every prompt: a direct proxy for cost.
fixed raw can never answer the price question: in the raw text, the phrase is broken as "pur- chase".
Overlap fixed one boundary miss for fixed-size chunks (6 to 7), at the price of 5 more chunks.
recursive 800 and headings both found 8 of 8, but recursive sends about 2.5 times as many characters per question.
fixed 1500 cuts the manual into two huge chunks and still misses a question.
Now run it with the real embedding model from Unit 3 (drop --offline):
python unit07/chunk_lab.py compare
Your numbers will differ from the ones above, because the model matches meaning, not words. If a line says chunk(s) exceed the model's 256 tokens; their ends are ignored, those chunks were too long for the embedding model. Compare which strategy wins with both embedders.
Try --top 1. With only one chunk allowed, every strategy loses some questions. That is the honest test: does the best chunk hold the answer?
Wrote 11 chunks to unit07/chunks.jsonl
First record:
{
"id": "p2p-invoice-manual.pdf#1",
"text": "Vendor invoice handling manual > Introduction\nVendor invoice handling manual\n\nVersion 3.2. Valid from 2026-07-01. Owner: accounts payable. This manual replaces version 3.1.",
"metadata": {
"source": "p2p-invoice-manual.pdf",
"page": 1,
"section": "Introduction",
"process": "procure-to-pay",
"doc_version": "3.2",
"valid_from": "2026-07-01",
"language": "en",
"chunker": "headings"
}
}
Open unit07/chunks.jsonl. Each line is one chunk. The metadata block is what the next topics use for filtering (process, company code), citing (page, section) and freshness (version, valid-from date). Fixed and recursive chunks get "page": 0, meaning "unknown": a blind cut loses track of where it came from unless you add extra bookkeeping.
If the clean file has no headings, your PDF probably doesn't number its headings. That is the heuristic's limit: everything lands in one "Introduction" section, split by paragraph.
If the script prints No text found, the PDF is a scan. It needs OCR first; that isn't covered here.
Keep generated files out of Git. Open .gitignore, add these lines at the end and save:
As of October 2026, SAP gives you three places where chunking happens.
SAP AI Core grounding, Pipeline API. SAP's release notes describe the Pipeline API (added November 2024) as segmenting data into chunks and generating embeddings. Since February 2025 it also has endpoints to check the processing status of each document. SAP Learning describes the document store as holding files such as PDFs and text files. You point it at a repository; SAP extracts, chunks and embeds. Check the current Pipeline API documentation for the file types and any chunking options your release supports; this course could not confirm a list of configurable chunking settings.
SAP AI Core grounding, Vector API. You send chunks you prepared. The SAP Cloud SDK for AI documents a document as a list of chunks, each with content and key/value metadata, plus key/value metadata on the document itself. Since February 2026, SAP's release notes list metadata on documents, collections and chunks, and metadata filtering in search. That is exactly what chunks.jsonl holds.
Sketch (unit07/vector_api_body_sketch.py, run after Step 6):
"""SKETCH: turn chunks.jsonl into one document-with-chunks body for SAP AI Core's Vector API.
It only builds and prints the JSON; it sends nothing. Sending it needs an SAP AI Core service key,
a resource group and a collection (Unit 7, "SAP HANA Cloud vector engine", and Unit 5 setup).
Check the field names against the Vector API reference for your release before you send it.
"""
import json
from pathlib import Path
HERE = Path(__file__).resolve().parent
records = [json.loads(line) for line in (HERE / "chunks.jsonl").read_text(encoding="utf-8").splitlines()]
first = records[0]["metadata"]
document_keys = ("source", "process", "doc_version", "valid_from", "language") # same for the whole file
body = {"documents": [{
"metadata": [{"key": k, "value": [str(first[k])]} for k in document_keys],
"chunks": [{"content": r["text"],
"metadata": [{"key": "page", "value": [str(r["metadata"]["page"])]},
{"key": "section", "value": [r["metadata"]["section"]]}]}
for r in records]}]}
print(json.dumps(body, indent=2, ensure_ascii=False)[:900] + "\n...")
print(f"\n{len(records)} chunks in one document. Nothing was sent.")
It prints the start of the request body and ends with 11 chunks in one document. Nothing was sent. Compare it with chunks.jsonl: document-level labels (source, process, version) are stored once on the document; page and section go on each chunk.
SAP HANA Cloud text splitter. The hana-ml Python client has a TextSplitter that runs inside SAP HANA Cloud. Its documented settings include chunk_size, overlap, split_type (default recursive), doc_type (default plain) and language (default auto). Input is a table with ID and TEXT columns. Use it when the text already sits in the database, for example service notes or long texts copied from SAP, and you want chunks without moving data out.
SAP Document AI is not a chunker. It extracts structured fields from documents such as invoices and purchase orders. If the goal is "read the PDF and fill SAP fields", use extraction; if it is "answer questions from the PDF", use RAG.
Licensing notes. pypdf, Sentence Transformers and Chroma are free. SAP AI Core grounding usage is billed against your SAP BTP account; check the metrics in your contract. hana-ml runs against an SAP HANA Cloud instance you pay for (or your trial).
A practical path: use this lab to learn what your documents need. If SAP's pipeline handles them well on your test questions, use it. If tables or metadata suffer, prepare chunks yourself and upload them with the Vector API.
Security and authorizations. Chunking can mix content with different audiences in one chunk. Keep access labels (company code, sales organization, confidentiality) in chunk metadata and filter on them from the user's identity. Never place two audiences' text in one chunk. Unit 7's topic on grounding with SAP authorizations builds on this.
Prompt injection. Documents may contain text written to steer the model. Cleaning doesn't remove that; treat chunk text as data, keep instructions in the system message, and log chunk IDs.
Versioning. Store doc_version and valid_from. When a manual changes, delete its old chunks before indexing the new ones; mixed versions produce contradictory answers.
Re-indexing cost. Changing chunk size or strategy means re-chunking and re-embedding everything. Decide with a test, record the settings in metadata (chunker), and change them on purpose.
Evaluation. Keep the chunking test set with the code, and re-run it when documents or settings change. Track hit rate and characters sent together.
Embedding limits. Check chunk length against your embedding model's input limit. The lab counts tokens for the Unit 3 model; do the same for SAP-hosted embedding models.
Personal data. Manuals rarely hold personal data, but exports, tickets and e-mails do. Decide what may be indexed before you chunk it.
Clean core. Preparation happens outside SAP on documents and released APIs. Nothing here modifies SAP; keep it side by side on SAP BTP.
Tune chunking for one more document and record the result. The report, unit07/chunking_report.txt, feeds the vector database and evaluation topics later in Unit 7 and Unit 8.
Open unit07/chunk_lab.py. In SAMPLE_PAGES, add a fourth page with a section 7 Credit memos and a short table (three rows: amount range, approver, deadline). Save.
Delete the old sample so it is rewritten: delete the folder unit07/chunk_docs in VS Code (right-click, Delete).
Add two questions to EVAL whose answers are in your new table, each with the phrases that must appear. Save.
Run python unit07/chunk_lab.py compare --offline and note the hit rate and characters sent for each row.
Add one row to the runs list in main(), for example ("headings", "clean", 300, 0), and run compare again. Does a smaller limit split your table? Check with show --size 300.
Choose a strategy and size. Save the comparison to a file:
Add one line at the end of the file, by hand: the strategy and size you chose, and why, in one sentence.
Run python unit07/chunk_lab.py export --strategy headings with your chosen size (add --size N), then commit unit07/chunk_lab.py, unit07/chunks.jsonl and unit07/chunking_report.txt.
Done whenEVAL has ten questions, unit07/chunking_report.txt shows at least eight comparison rows plus your one-line decision, and unit07/chunks.jsonl holds chunks from four pages with page and section metadata.
Pick one answer for each question. The explanation appears after you choose.
1Why does pypdf's plain extraction lose the columns of the price table?
Answer: C. The pypdf documentation says PDFs have no semantic layer and tables are absolutely positioned text. Layout mode keeps the spacing, which clean() turns back into table rows.
2In clean(), how are page headers and footers such as "Page 2 of 3" removed?
Answer: B. Masking digits makes "Page 2 of 3" and "Page 3 of 3" identical, and lines that appear on every page are treated as headers or footers. Deleting fixed line positions would break on PDFs with different layouts.
3The fixed-raw run can never find the answer to the price question. Why?
Answer: D. Raw extraction keeps "pur-" and "chase" on separate lines, so the expected phrase doesn't exist in any raw chunk. Cleaning re-joins the word before chunking.
4Recursive chunking at 800 and headings at 600 both find 8 of 8. Which do you choose, and why?
Answer: A. Hit rate ties, but recursive 800 sends about 2.5 times the characters into every prompt, which is paid on every call. Track quality and cost together.
5The compare run with the real model warns that some chunks exceed 256 tokens. What happens to them?
Answer: C. The model card says input longer than 256 word pieces is truncated. The cut-off text can't help matching, yet the whole chunk still costs prompt tokens when retrieved.
6What is the difference between the title > section prefix and the metadata block?
Answer: B. Text in the chunk is embedded, so the prefix helps matching. Metadata sits beside the text, where code can filter by process, show page and section, and remove old versions.
7A manual moves from version 3.1 to 3.2 and users get contradictory answers. What do you do?
Answer: D. Old and new chunks both match, so the model sees conflicting rules. Version metadata lets you remove the old chunks cleanly before indexing the new ones.
8You run on SAP AI Core and need your own company-code labels on every chunk. Which route fits?
Answer: C. The Vector API takes your chunks with key/value metadata on documents and chunks, and SAP documents metadata filtering since February 2026. The Pipeline API chunks for you, so you control less of what each chunk carries.
9Finance asks you to read invoice amounts from supplier PDFs into SAP. What do you tell them?
Answer: B. Pulling structured fields from invoices is document extraction. SAP Document AI is built for that; RAG chunking is for answering questions from text.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Splitting recursively (LangChain documentation)— default separators paragraph, line, space, character, tried in order until chunks are small enough; chunk_size and chunk_overlap; package langchain-text-splitters
Extract Text from a PDF (pypdf documentation)— PDFs have no semantic layer (no header, footer, table or paragraph information); tables are absolutely positioned text; extraction_mode layout; scanned PDFs need OCR software
Contextual Retrieval (Anthropic engineering, September 2024)— chunks lose context (revenue grew 3 percent example); prepending chunk-specific context before embedding and BM25 reduced failed retrievals by 35 percent (embeddings only) and 49 percent (with BM25), measured as 1 minus recall at 20
SAP AI Core (product guide PDF, help.sap.com)— Pipeline API segments data into chunks and generates embeddings (2024-11-04); Pipeline API document status endpoints (2025-02-03); Vector API metadata for documents, collections and chunks, and metadata filtering (2026-02-16)
Document Grounding (SAP Cloud SDK for AI, Java)— Vector API documents hold chunks with content and key/value metadata; document metadata also key/value; retrieval filters with data repository type vector and maxChunkCount
TextSplitter (hana-ml documentation, help.sap.com)— SAP HANA Cloud text splitter in the Python machine learning client; chunk_size, overlap, split_type (default recursive), doc_type (default plain), language (default auto); input ID and TEXT columns