Orchestrate

Chunking and document preparation: from PDF to retrievable passages

How to turn PDFs and manuals into clean, well-labelled chunks, choose size, overlap and structure, and measure which choice retrieves best.

Updated Oct 3, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

Before an AI assistant can answer from your manuals, someone has to prepare them. That means two jobs.

First, turn files into clean text. A PDF looks tidy on screen, but inside it is mostly loose words placed on a page. Page headers, page numbers, broken words and tables come out jumbled.

Second, cut the text into chunks: passages small enough to search precisely and to fit in the model's prompt, but large enough to make sense alone. Each chunk gets labels, called metadata: which document, which page, which section, which process, which version.

These choices decide what the assistant can find. If the answer is split across two chunks, or a table loses its header, the right passage never reaches the model. No prompt or bigger model fixes that.

Why it matters to the business

In RAG fundamentals you saw the rule: check retrieval before generation. Chunking is the biggest lever on retrieval that your own team controls.

Take the running procure-to-pay example. An accounts payable clerk asks, "What is the price tolerance for company code 2000?" The answer sits in a table in the invoice handling manual. Cut that manual badly, and the table rows land in one chunk while the words "price tolerance" sit in another. The search finds the wrong passage, and the assistant answers from it, confidently.

Three business effects follow:

  • Quality. Badly prepared documents produce wrong or "can't find" answers, and users stop trusting the assistant.
  • Cost. Every retrieved chunk is paid for in every model call. Chunks that are too big, or that repeat each other, raise the bill without raising quality.
  • Control. Metadata is what lets you filter by company code or process, show the page a fact came from, and remove old versions. Without it you can't audit an answer.

How SAP does it

As of October 2026:

  • SAP AI Core grounding (generative AI hub). The Pipeline API reads documents from a connected repository, segments them into chunks and creates embeddings for you. SAP Learning describes the document store as holding files such as PDFs and text files. This is the low-effort path; you accept SAP's chunking.
  • The Vector API is the path when you want control. Your team prepares the chunks and uploads them. Since February 2026, SAP documents metadata on documents, collections and chunks, and filtering on that metadata at search time.
  • SAP HANA Cloud includes a text splitter in its Python machine learning client (hana-ml), with settings for chunk size, overlap and splitting method. Useful when the text already lives in the database.
  • SAP Document AI is a different tool for a different job. It extracts structured fields from documents such as invoices and purchase orders. If you want a field value posted to SAP, that is extraction, not RAG.

Three ways to cut a document

Approach How it cuts Good for Watch out for
Fixed size Every N characters, regardless of content Quick tests on plain prose Cuts words, sentences and tables in half
Recursive Paragraphs first, then lines, sentences, words, until pieces fit General text without clear headings Can still separate a table from its heading
Structure-aware At headings; keeps tables whole; labels each chunk with its section Manuals, policies, process guides Needs clean structure, so preparation matters more

Two settings go with every approach:

  • Chunk size: how long a chunk may be. Small chunks match specific questions; large chunks carry more context but more noise and cost.
  • Overlap: how much of the previous chunk is repeated at the start of the next. It protects answers that fall on a boundary, at the price of storing and sending duplicates.

There is no universal best setting. Chroma's July 2024 study found differences of up to 9 percent in recall between strategies on the same data. The only reliable answer is to test with real questions on your own documents, which the deep layer shows how to do in an afternoon.

Questions to ask

  • Which document types will we index (PDF, Word, wiki pages, scans), and who owns each source?
  • Are any of the PDFs scans? Who runs OCR (text recognition) on them, and who checks its quality?
  • How do we chunk, and why? Was the choice tested against real user questions, with a hit rate we can see?
  • What metadata does each chunk carry: document, version, valid-from date, page, section, process, company code, language?
  • How are old versions removed when a manual is updated?
  • How are tables handled? Can the team show a table chunk that still has its header row?
  • If we use SAP's grounding pipeline, do we accept its default chunking, or prepare chunks ourselves through the Vector API?

Common misconceptions

  • "Just upload the PDFs." PDFs carry no notion of paragraphs, headers or tables. Somebody has to rebuild that structure, or the assistant reads noise.
  • "Bigger chunks are safer because they hold more." They also hold more irrelevant text, cost more per call, and can exceed what the embedding model reads. Many small embedding models cut off long input.
  • "More overlap is always better." Overlap duplicates text. Chroma's study found that reducing overlap improved efficiency scores, and a common default of large chunks with heavy overlap had the lowest efficiency scores.
  • "Chunking is a one-time technical detail." Changing it means re-indexing everything, and it moves answer quality. Treat it as a design decision with a test behind it.
  • "RAG can read our invoices into SAP." Pulling fields from invoices is extraction, a job for SAP Document AI, not retrieval.

Key terms

  • Document preparation: turning files into clean, structured text before chunking.
  • Chunk: a passage cut from a document, stored and searched as one unit.
  • Chunk size: the maximum length of a chunk, in characters or tokens.
  • Overlap: text repeated between neighbouring chunks.
  • Structure-aware chunking: cutting at headings and keeping tables and lists whole.
  • Metadata: labels stored with a chunk, such as source, page, section, version and process.
  • OCR (optical character recognition): turning an image of text, such as a scan, into text.
  • Hit rate at k: the share of test questions whose answer is in one of the top k retrieved chunks.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does chunking matter so much for answer quality?

    Answer: B. The model answers from the chunks it is given. If chunking splits or buries the answer, retrieval misses it, and no prompt or larger model can recover it.
  2. 2A clerk asks for the price tolerance of company code 2000, and the answer is in a table. What usually goes wrong with careless chunking?

    Answer: C. Fixed-size cutting can place "price tolerance" in one chunk and the rows in another. The search then matches the wrong passage, and the answer is built on it.
  3. 3Your team proposes large chunks with heavy overlap "to be safe". What is the best response?

    Answer: D. Bigger chunks and more overlap raise cost and noise, and Chroma's study found such a default scored lowest on efficiency. Overlap can still help with boundaries, so the answer is a measured test, not a rule.
  4. 4Which metadata most helps you audit an answer and retire old content?

    Answer: A. Source, page and section let you show where a fact came from. Version and valid-from date let you find and remove outdated chunks when a manual changes.
  5. 5You want full control over how your manuals are cut, but you will run on SAP AI Core. Which path fits?

    Answer: B. The Pipeline API chunks for you, which is less effort but less control. The Vector API takes chunks you prepared, with your metadata, which SAP documents for filtering since February 2026.
  6. 6Finance wants invoice amounts and vendor numbers read from PDFs and posted to SAP. What is this?

    Answer: D. Pulling structured fields from invoices is extraction. SAP Document AI is built for that; RAG chunking is for answering questions from text.
Deep layer · 40 min read

Mental model

A chunk is the unit of evidence. Retrieval ranks chunks, the prompt carries chunks, citations point to chunks. So each chunk must pass one test: could a colleague answer the question from this chunk alone, and say where it came from?

That test explains every rule in this topic:

  • Cut at meaning boundaries (sections, paragraphs, table ends), not at character counts, so a chunk holds one complete idea.
  • Put context into the text (document title and section path), so "the tolerance is 2 percent" still says of what.
  • Keep labels outside the text (metadata: page, version, process), so code can filter, cite and expire chunks.
  • Measure with questions whose answers you know, because intuition about chunk size is unreliable.

Document preparation is what makes those rules possible. You can't cut at headings you can't see.

How it works

The indexing path from RAG fundamentals had "ingest" and "chunk" as two boxes. In practice there are five steps, and most failures happen in the first two.

flowchart LR
  P[PDF, Word, HTML] --> X[1 Extract text]
  X --> C[2 Clean and rebuild structure]
  C --> S[3 Split into chunks]
  S --> M[4 Add context and metadata]
  M --> E[5 Embed and index]
  Q[Test questions] --> V[Measure hit rate]
  E --> V
  V -.adjust.-> S

1. Extract text

The pypdf documentation is blunt: a PDF has no semantic layer. It holds instructions to draw text at positions. There is no "header", "paragraph" or "table", and in a table, each cell is just text placed at coordinates. Extraction tools rebuild reading order from those positions:

  • Plain extraction joins text in order. Table cells come out separated by single spaces, so "2000 France 5 percent Controller" loses its columns.
  • Layout extraction (pypdf's extraction_mode="layout") keeps the spacing as rendered, so columns stay visible as wide gaps.
  • Scanned PDFs are images. They contain no text to extract; they need OCR first. pypdf recommends using OCR software directly in that case.

2. Clean and rebuild structure

Extracted text carries noise that hurts retrieval:

Noise Example Fix
Repeated page headers and footers "ACME Corp - Internal …", "Page 2 of 3" on every page Drop lines that repeat on every page (compare with digits masked)
Hyphenated line breaks "pur-" / "chase order" Re-join a word when a line ends in a hyphen and the next starts lowercase
Hard line breaks inside paragraphs One sentence spread over three lines Join lines until a blank line
Tables as loose text Columns separated only by spaces Detect column gaps and rebuild rows as Markdown tables
Lost headings "2.2 Price differences" looks like any other line Detect numbering patterns or font sizes; mark as headings

The output is clean Markdown: headings, paragraphs and tables you can read and check. This is the step teams skip, and it is the step that decides whether structure-aware chunking is possible.

SAP documentation is usually easier: help.sap.com pages are HTML, which already marks headings and tables. For SAP's own product documentation, SAP's grounding module can also search help.sap.com directly, as RAG fundamentals showed. Your own process guides, often PDFs or Word files in SharePoint, are where preparation work lands.

3. Split into chunks

Three families, from blunt to careful:

  • Fixed size. Every N characters, stepping back by the overlap. Simple and fast, but it cuts mid-word and mid-table.
  • Recursive. LangChain's RecursiveCharacterTextSplitter tries separators in order, by default paragraph break, line break, space, then single characters, until each piece fits. It then packs pieces up to the size. It respects paragraphs, but it doesn't know a table belongs to the heading above it.
  • Structure-aware. Cut at headings. Keep each table whole with its header row. Split a long section by paragraph, repeating the section label on each part.

Size is measured in characters or tokens. Embedding models have an input limit: the model card for all-MiniLM-L6-v2, the Unit 3 model, says input longer than 256 word pieces is truncated. Anything past that point is silently ignored when the chunk is embedded, though it still goes into the prompt.

4. Add context and metadata

Two different things, often confused:

  • Context in the text. Prefix each chunk with Document title > Section. It changes the embedding, so it helps matching. Anthropic's 2024 "contextual retrieval" write-up takes this further: a model writes a short, chunk-specific context line before embedding. In their tests, failed retrievals at top 20 fell by 35 percent with contextual embeddings alone, and by 49 percent when combined with keyword search.
  • Metadata beside the text. Source, page, section, process, version, valid-from date, language. It doesn't change the embedding. It lets code filter ("only procure-to-pay"), cite ("manual p. 2, section 3.1") and expire ("drop version 3.1").

5. Measure

You can't judge chunking by looking at it. You need test questions with known answers. For chunking, the most useful check is at chunk level, not document level: is there one retrieved chunk that contains the whole answer? Chroma's technical report goes further and scores at token level: recall (how much of the answer was retrieved) and precision (how much of what was retrieved was answer). Two of its findings are worth remembering:

  • Recursive splitting at about 200 tokens with no overlap performed well across their data.
  • A popular default of 800 tokens with 400 overlap had slightly below-average recall and the lowest efficiency scores.

Your documents are not theirs, so treat these as starting points to test, not settings to copy.

Build it yourself: a chunk lab for a PDF manual

Before you start: complete Set up your computer for this course and Set up for Unit 7. This topic builds on RAG fundamentals; you don't need its code to run this one, but you'll reuse the ideas.

You will build chunk_lab.py, one script that takes a PDF manual through every step: extract, clean, chunk three ways, and measure. It writes its own sample PDF, a made-up vendor invoice handling manual with page headers, a hyphenated word and two tables, so you see real extraction problems without hunting for a file. At the end it exports the winning chunks with metadata to chunks.jsonl, which later Unit 7 topics load into a vector store.

flowchart LR
  A[Sample PDF<br/>or your PDF] --> B[extract<br/>raw.txt, clean.md]
  B --> C{strategy}
  C --> F[fixed]
  C --> R[recursive]
  C --> H[headings]
  F --> M[compare<br/>hit rate, cost]
  R --> M
  H --> M
  H --> J[export<br/>chunks.jsonl]

What you need

  • Your course folder with the Unit 7 setup done. Cost: free.
  • One new free library, pypdf.
  • About 60 minutes. No account and no model calls; every step runs on your computer.

Step 1: Open your course folder and add pypdf

  1. Open a terminal (on Windows, PowerShell; on macOS, Terminal) and turn on your environment.

    Windows (PowerShell):

    cd $HOME\orchestrate-course
    .\.venv\Scripts\Activate.ps1

    macOS/Linux:

    cd ~/orchestrate-course
    source .venv/bin/activate
  2. Open requirements.txt in VS Code, add this line at the end and save:

    pypdf
  3. Install it (the same on every system):

    pip install -r requirements.txt
  4. Check it:

    pip show pypdf

You should see Name: pypdf and a Version: line (6.19.0 was current when this topic was written).

Step 2: Create the script

  1. In VS Code, right-click the unit07 folder, choose New File, name it chunk_lab.py, paste the code below and save.
"""Chunk lab: turn a PDF manual into clean text, cut it into chunks three ways, and measure which
way lets retrieval find the answers. Builds on rag_minimal.py from "RAG fundamentals".

Run from your course folder (orchestrate-course), with .venv turned on:
    python unit07/chunk_lab.py extract                              # PDF -> raw text and clean Markdown
    python unit07/chunk_lab.py show --strategy headings             # print the chunks one strategy makes
    python unit07/chunk_lab.py show --strategy fixed --source raw --question "When are vendors paid?"
    python unit07/chunk_lab.py compare                              # every strategy against the test questions
    python unit07/chunk_lab.py export --strategy headings           # write chunks.jsonl for later topics
Add --offline to use toy embeddings instead of the Unit 3 model. Add --pdf PATH to use your own PDF.
"""
import argparse
import hashlib
import json
import math
import re
import sys
import textwrap
from pathlib import Path

HERE = Path(__file__).resolve().parent
DOCS = HERE / "chunk_docs"
SAMPLE_PDF = DOCS / "p2p-invoice-manual.pdf"
MODEL = "sentence-transformers/all-MiniLM-L6-v2"   # the Unit 3 model

# ---------- the made-up sample manual (procure-to-pay), written to a real PDF ----------

HEADER = "ACME Corp - Internal - Vendor invoice handling manual"
SAMPLE_PAGES = [
    [("title", "Vendor invoice handling manual"),
     ("p", "Version 3.2. Valid from 2026-07-01. Owner: accounts payable. This manual replaces version 3.1."),
     ("h", "1 Purpose and scope"),
     ("p", "This manual explains why vendor invoices are blocked for payment, who may release them and "
           "how fast. It applies to all company codes in Europe. It does not cover travel expenses."),
     ("h", "2 Why invoices are blocked"),
     ("p", "The system compares each invoice with the purchase order and the goods receipt. This "
           "comparison is called the three-way match. When a line fails the match, the invoice is "
           "blocked for payment until someone releases it."),
     ("h", "2.1 Quantity differences"),
     ("p", "If the invoiced quantity is higher than the quantity received, the invoice is blocked. "
           "The buyer checks with the warehouse whether a second delivery is on its way."),
     ("h", "2.2 Price differences"),
     ("lines", ["An invoice is blocked when the invoice price differs from the pur-",
                "chase order price by more than the tolerance for the company code:"]),
     ("table", [["Company code", "Tolerance", "Approver"],
                ["1000 Germany", "2 percent", "AP team lead"],
                ["2000 France", "5 percent", "Controller"],
                ["3000 Spain", "3 percent", "AP team lead"]])],
    [("h", "2.3 Missing goods receipt"),
     ("p", "If no goods receipt has been posted, the invoice waits for the goods receipt. Do not "
           "release it by hand. Ask the requester to confirm the delivery in the system first."),
     ("h", "3 Releasing blocked invoices"),
     ("h", "3.1 Who may release"),
     ("p", "The accounts payable clerk may release blocks up to 10,000 euros per invoice. Above "
           "10,000 euros, only the head of accounts payable may release the block. Nobody may "
           "release an invoice they created themselves."),
     ("h", "3.2 Release deadlines"),
     ("p", "Every block must be worked within the deadline below, or it is escalated:"),
     ("table", [["Block reason", "Release within", "Escalate to"],
                ["Quantity difference", "5 working days", "Buyer"],
                ["Price difference", "3 working days", "Controller"],
                ["Missing goods receipt", "10 working days", "Requester's manager"]]),
     ("p", "Escalations are sent by e-mail every Monday morning.")],
    [("h", "4 Duplicate invoices"),
     ("p", "An invoice is a suspected duplicate when it has the same vendor, the same reference "
           "and the same amount as an invoice from the last 90 days. Suspected duplicates are "
           "never paid automatically. The clerk compares both documents and rejects one."),
     ("h", "5 Payment runs"),
     ("p", "Released invoices are paid in the payment run every Tuesday and Friday. An invoice "
           "released after 12:00 on a payment day waits for the next run. Urgent payments outside "
           "the run need approval from the treasury team."),
     ("h", "6 Archiving"),
     ("p", "Paid invoices are archived after the payment run. Keep invoices for ten years, or "
           "longer where local law requires it.")],
]

# Test questions, each with phrases that must ALL appear in one retrieved chunk to count as found.
EVAL = [
    ("What is the price tolerance for company code 2000?", ["2000", "5 percent"]),
    ("Who may release an invoice above 10,000 euros?", ["head of accounts payable"]),
    ("When are vendors paid?", ["tuesday and friday"]),
    ("How long must invoices be kept?", ["ten years"]),
    ("What should I do when the goods receipt is missing?", ["waits for the goods receipt"]),
    ("How fast must a quantity difference block be released?", ["quantity difference", "5 working days"]),
    ("How do we recognize a duplicate invoice?", ["same vendor", "same amount"]),
    ("Why is an invoice blocked when the price is different from the purchase order?",
     ["invoice price differs from the purchase order price"]),
]


def write_pdf(pages: list, path: Path) -> None:
    """Write a small PDF by hand (no extra library). Each page gets a header and a page-number footer,
    like a real manual. Table cells are placed at fixed x positions, as most PDFs do."""
    def esc(s):
        return s.replace("\\", "\\\\").replace("(", "\\(").replace(")", "\\)")
    objects = {1: "<< /Type /Catalog /Pages 2 0 R >>",
               3: "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>"}
    kids = []
    for number, page in enumerate(pages, 1):
        items, y = [(50, 805, 8, HEADER)], 770
        for kind, value in page:
            if kind == "title":
                items.append((50, y, 18, value)); y -= 30
            elif kind == "h":
                y -= 8; items.append((50, y, 13, value)); y -= 22
            elif kind in ("p", "lines"):
                for line in (textwrap.wrap(value, 92) if kind == "p" else value):
                    items.append((50, y, 10, line)); y -= 14
                y -= 10
            elif kind == "table":
                for row in value:
                    for x, cell in zip((50, 230, 370), row):
                        items.append((x, y, 10, cell))
                    y -= 15
                y -= 10
        items.append((260, 30, 8, f"Page {number} of {len(pages)}"))
        stream = "\n".join(f"BT /F1 {s} Tf {x} {y} Td ({esc(t)}) Tj ET" for x, y, s, t in items)
        page_id, content_id = 2 + 2 * number, 3 + 2 * number
        kids.append(f"{page_id} 0 R")
        objects[page_id] = ("<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] "
                            f"/Resources << /Font << /F1 3 0 R >> >> /Contents {content_id} 0 R >>")
        objects[content_id] = (f"<< /Length {len(stream.encode('latin-1'))} >>\n"
                               f"stream\n{stream}\nendstream")
    objects[2] = f"<< /Type /Pages /Kids [{' '.join(kids)}] /Count {len(pages)} >>"
    data, offsets = b"%PDF-1.4\n", []
    for i in range(1, max(objects) + 1):
        offsets.append(len(data))
        data += f"{i} 0 obj\n{objects[i]}\nendobj\n".encode("latin-1")
    xref = len(data)
    data += f"xref\n0 {len(offsets) + 1}\n0000000000 65535 f \n".encode()
    data += "".join(f"{o:010d} 00000 n \n" for o in offsets).encode()
    data += f"trailer\n<< /Size {len(offsets) + 1} /Root 1 0 R >>\nstartxref\n{xref}\n%%EOF\n".encode()
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_bytes(data)


# ---------- 1. extract: PDF -> text per page ----------

def extract(pdf: Path, mode: str) -> list:
    """Return one string per page. mode 'plain' is pypdf's default; 'layout' keeps column spacing."""
    try:
        from pypdf import PdfReader
    except ImportError:
        sys.exit("pypdf is not installed. Add 'pypdf' to requirements.txt and run: pip install -r requirements.txt")
    try:
        reader = PdfReader(pdf)
        return [page.extract_text(extraction_mode=mode) or "" for page in reader.pages]
    except Exception as error:
        sys.exit(f"Could not read {pdf.name}: {type(error).__name__}: {error}")


# ---------- 2. clean: remove headers and footers, rebuild paragraphs, headings and tables ----------

def clean(pages: list) -> list:
    """Turn layout-mode page texts into blocks: {'kind': 'h'|'p'|'table', 'text', 'page'}."""
    def shape(line):          # "Page 2 of 3" and "Page 3 of 3" have the same shape
        return re.sub(r"\d+", "#", line.strip())
    seen = {}
    for text in pages:
        for line in {shape(l) for l in text.splitlines() if l.strip()}:
            seen[line] = seen.get(line, 0) + 1
    repeated = {line for line, count in seen.items() if len(pages) > 1 and count >= len(pages)}

    blocks = []
    for number, text in enumerate(pages, 1):
        paragraph = []

        def flush():
            if paragraph:
                joined = paragraph[0]
                for nxt in paragraph[1:]:
                    if re.search(r"[a-z]-$", joined) and nxt[:1].islower():
                        joined = joined[:-1] + nxt          # re-join a word split across lines
                    else:
                        joined += " " + nxt
                blocks.append({"kind": "p", "text": joined, "page": number})
                paragraph.clear()

        for raw in text.splitlines():
            line = raw.strip()
            if not line or shape(line) in repeated:
                flush()
                continue
            cells = re.split(r"\s{3,}", line)
            if len(cells) >= 2:                              # columns -> a table row
                flush()
                row = "| " + " | ".join(cells) + " |"
                if blocks and blocks[-1]["kind"] == "table":
                    blocks[-1]["text"] += "\n" + row
                else:
                    blocks.append({"kind": "table", "text": row, "page": number})
            elif re.match(r"^\d+(\.\d+)*\s+[A-Z][^.]{0,80}$", line):   # "2.2 Price differences"
                flush()
                blocks.append({"kind": "h", "text": line, "page": number})
            else:
                paragraph.append(line)
        flush()
    return blocks


def to_markdown(blocks: list) -> str:
    """Render blocks as Markdown, so you can read what the chunker will see."""
    out = []
    for b in blocks:
        if b["kind"] == "h":
            out.append("#" * (b["text"].split()[0].count(".") + 2) + " " + b["text"])
        elif b["kind"] == "table":
            rows = b["text"].split("\n")
            divider = "|" + "---|" * (rows[0].count("|") - 1)
            out.append("\n".join([rows[0], divider] + rows[1:]))
        else:
            out.append(b["text"])
    return "\n\n".join(out) + "\n"


# ---------- 3. chunk: three strategies ----------

def chunk_fixed(text: str, size: int, overlap: int) -> list:
    """Cut every `size` characters, stepping back `overlap` characters each time. Ignores words."""
    step = max(1, size - overlap)
    return [text[i:i + size] for i in range(0, len(text), step) if text[i:i + size].strip()]


def chunk_recursive(text: str, size: int, overlap: int,
                    separators=("\n\n", "\n", ". ", " ")) -> list:
    """Split at the biggest separator that works (paragraph, line, sentence, word), then pack the
    pieces into chunks of at most `size` characters, repeating up to `overlap` characters of the
    previous chunk's end. The same idea as LangChain's RecursiveCharacterTextSplitter."""
    def pieces(t, seps):
        if len(t) <= size or not seps:
            return [t]
        sep, rest = seps[0], seps[1:]
        parts = t.split(sep)
        out = []
        for i, part in enumerate(parts):
            part = part + (sep if i < len(parts) - 1 else "")
            out.extend(pieces(part, rest) if len(part) > size else [part])
        return out

    chunks, current = [], []
    for piece in pieces(text, separators):
        if current and sum(map(len, current)) + len(piece) > size:
            chunks.append("".join(current).strip())
            tail = []
            for prev in reversed(current):              # carry the end of the chunk forward
                if sum(map(len, tail)) + len(prev) > overlap:
                    break
                tail.insert(0, prev)
            current = tail
        current.append(piece)
    if current:
        chunks.append("".join(current).strip())
    return [c for c in chunks if c]


def chunk_headings(blocks: list, title: str, size: int) -> list:
    """Structure-aware: one chunk per section, prefixed with 'title > section'. Tables stay whole
    with their header row; long sections are split by paragraph, each part keeping the prefix."""
    sections, current = [], {"heading": "", "parts": [], "page": blocks[0]["page"] if blocks else 1}
    for b in blocks:
        if b["kind"] == "h":
            if current["parts"] or current["heading"]:
                sections.append(current)
            current = {"heading": b["text"], "parts": [], "page": b["page"]}
        else:
            current["parts"].append(b["text"])
    sections.append(current)

    chunks = []
    for s in sections:
        if not s["parts"]:
            continue                                  # a heading with no text of its own
        heading = s["heading"] or "Introduction"
        prefix = f"{title} > {heading}\n"
        group = ""
        for part in s["parts"]:
            if group and len(prefix) + len(group) + len(part) > size:
                chunks.append({"text": prefix + group, "section": heading, "page": s["page"]})
                group = ""
            group = (group + "\n\n" + part).strip()
        if group:
            chunks.append({"text": prefix + group, "section": heading, "page": s["page"]})
    return chunks


# ---------- 4. embed (same choices as rag_minimal.py) ----------

STOP = set("a an and are as at be by can do does for from has how i in is it it's may of on or our "
           "should so that the their this to was we what when where which who why will with".split())


def toy_embed(text: str, size: int = 256) -> list:
    """Hash each word into one of 256 slots, skipping common words. Matches words, not meaning."""
    vector = [0.0] * size
    for word in re.findall(r"[a-z0-9]+", text.lower()):
        if word in STOP:
            continue
        vector[int(hashlib.md5(word.encode()).hexdigest(), 16) % size] += 1.0
    length = math.sqrt(sum(v * v for v in vector)) or 1.0
    return [v / length for v in vector]


def get_embedder(offline: bool):
    """Return (embed function, token counter or None)."""
    if offline:
        return (lambda texts: [toy_embed(t) for t in texts]), None
    try:
        from sentence_transformers import SentenceTransformer
    except ImportError:
        sys.exit("sentence-transformers is not installed. See Set up for Unit 3, or add --offline.")
    print(f"Loading {MODEL} (from disk if you ran Unit 3)...")
    try:
        model = SentenceTransformer(MODEL)
    except Exception as error:
        sys.exit(f"Could not load the model ({type(error).__name__}). Check your network, or add --offline.")
    limit = model.max_seq_length

    def too_long(texts):
        return sum(1 for t in texts if len(model.tokenizer(t)["input_ids"]) > limit), limit
    return (lambda texts: model.encode(texts, normalize_embeddings=True).tolist()), too_long


def top_k(question_vec: list, chunk_vecs: list, k: int) -> list:
    """Indexes of the k chunks with the highest cosine similarity (vectors are already normalized)."""
    scores = [sum(a * b for a, b in zip(question_vec, v)) for v in chunk_vecs]
    return sorted(range(len(scores)), key=lambda i: -scores[i])[:k]


# ---------- 5. measure ----------

def norm(text: str) -> str:
    return re.sub(r"\s+", " ", text.lower())


def evaluate(texts: list, embed, k: int) -> tuple:
    """Hit rate: share of EVAL questions where one top-k chunk holds all expected phrases.
    Also returns the average number of characters the k chunks would add to each prompt."""
    vecs = embed(texts)
    q_vecs = embed([q for q, _ in EVAL])
    hits, sent, misses = 0, 0, []
    for (question, phrases), qv in zip(EVAL, q_vecs):
        chosen = [texts[i] for i in top_k(qv, vecs, k)]
        sent += sum(len(c) for c in chosen)
        if any(all(p in norm(c) for p in phrases) for c in chosen):
            hits += 1
        else:
            misses.append(question)
    return hits, sent / len(EVAL), misses


# ---------- running it ----------

def load(pdf: Path) -> tuple:
    """Return (raw text, clean blocks, title). Writes the sample PDF on first use."""
    if pdf == SAMPLE_PDF and not pdf.exists():
        write_pdf(SAMPLE_PAGES, pdf)
        print(f"Wrote the sample manual to {pdf.relative_to(HERE.parent)}")
    if not pdf.exists():
        sys.exit(f"File not found: {pdf}")
    raw = "\n".join(extract(pdf, "plain"))
    blocks = clean(extract(pdf, "layout"))
    if not raw.strip():
        sys.exit("No text found. The PDF may be a scan (an image of text); it needs OCR first.")
    first = next((b["text"] for b in blocks if b["kind"] == "p"), pdf.stem)
    title = first if len(first) < 60 else pdf.stem       # the sample's title line is its first paragraph
    return raw, blocks, title


def make_chunks(strategy: str, source: str, size: int, overlap: int, raw: str, blocks: list,
                title: str) -> list:
    """Return chunk dicts with text, section and page for any strategy."""
    if strategy == "headings":
        return chunk_headings(blocks, title, size)
    text = raw if source == "raw" else to_markdown(blocks)
    cut = chunk_fixed if strategy == "fixed" else chunk_recursive
    return [{"text": t, "section": "", "page": None} for t in cut(text, size, overlap)]


def metadata(pdf: Path, raw: str, strategy: str, chunk: dict) -> dict:
    """Labels stored with each chunk: for filters, citations and freshness checks."""
    version = re.search(r"Version (\d+(?:\.\d+)*)", raw)
    valid = re.search(r"Valid from (\d{4}-\d{2}-\d{2})", raw)
    return {"source": pdf.name, "page": chunk["page"] or 0, "section": chunk["section"],
            "process": "procure-to-pay", "doc_version": version.group(1) if version else "",
            "valid_from": valid.group(1) if valid else "", "language": "en", "chunker": strategy}


def main() -> None:
    parser = argparse.ArgumentParser(description="Prepare a PDF and compare chunking strategies.")
    parser.add_argument("action", choices=["extract", "show", "compare", "export"])
    parser.add_argument("--pdf", type=Path, default=SAMPLE_PDF, help="your own PDF (default: the sample)")
    parser.add_argument("--strategy", choices=["fixed", "recursive", "headings"], default="headings")
    parser.add_argument("--source", choices=["raw", "clean"], default="clean",
                        help="cut the raw extracted text or the cleaned text (fixed and recursive only)")
    parser.add_argument("--size", type=int, default=600, help="maximum characters per chunk (default 600)")
    parser.add_argument("--overlap", type=int, default=0, help="characters repeated between chunks (default 0)")
    parser.add_argument("--top", type=int, default=3, help="chunks retrieved per question (default 3)")
    parser.add_argument("--question", help="with show: also print the top chunks for this question")
    parser.add_argument("--offline", action="store_true", help="toy embeddings; no model download")
    args = parser.parse_args()
    if args.overlap >= args.size:
        sys.exit("--overlap must be smaller than --size.")

    raw, blocks, title = load(args.pdf)

    if args.action == "extract":
        stem = args.pdf.with_suffix("")
        Path(f"{stem}.raw.txt").write_text(raw, encoding="utf-8")
        Path(f"{stem}.clean.md").write_text(to_markdown(blocks), encoding="utf-8")
        kinds = {k: sum(b["kind"] == k for b in blocks) for k in ("h", "p", "table")}
        print(f"Raw text: {len(raw)} characters -> {Path(f'{stem}.raw.txt').name}")
        print(f"Clean Markdown: {kinds['h']} headings, {kinds['p']} paragraphs, {kinds['table']} tables "
              f"-> {Path(f'{stem}.clean.md').name}")
        print("Open both files side by side and compare them.")
        return

    if args.action == "show":
        chunks = make_chunks(args.strategy, args.source, args.size, args.overlap, raw, blocks, title)
        print(f"{args.strategy} on {args.source if args.strategy != 'headings' else 'clean'} text, "
              f"size {args.size}, overlap {args.overlap}: {len(chunks)} chunks\n")
        if args.question:
            embed, _ = get_embedder(args.offline)
            texts = [c["text"] for c in chunks]
            order = top_k(embed([args.question])[0], embed(texts), args.top)
            print(f'Top {args.top} for "{args.question}":\n')
            chunks = [chunks[i] for i in order]
        for n, c in enumerate(chunks, 1):
            where = f"page {c['page']}, " if c["page"] else ""
            print(f"--- chunk {n} ({where}{len(c['text'])} chars) ---\n{c['text']}\n")
        return

    if args.action == "compare":
        if args.pdf != SAMPLE_PDF:
            print("Note: the test questions are about the sample manual. Write your own EVAL for your PDF.")
        embed, too_long = get_embedder(args.offline)
        runs = [("fixed", "raw", 300, 0), ("fixed", "raw", 300, 100), ("fixed", "clean", 300, 0),
                ("recursive", "clean", 300, 50), ("recursive", "clean", 800, 100),
                ("fixed", "clean", 1500, 0), ("headings", "clean", 600, 0)]
        print(f"{'strategy':<10} {'source':<6} {'size':>5} {'overlap':>7} {'chunks':>6} "
              f"{'avg chars':>9} {'hit@' + str(args.top):>6} {'chars sent':>10}")
        for strategy, source, size, overlap in runs:
            chunks = make_chunks(strategy, source, size, overlap, raw, blocks, title)
            texts = [c["text"] for c in chunks]
            hits, sent, misses = evaluate(texts, embed, args.top)
            print(f"{strategy:<10} {source:<6} {size:>5} {overlap:>7} {len(texts):>6} "
                  f"{sum(map(len, texts)) / len(texts):>9.0f} {f'{hits}/{len(EVAL)}':>6} {sent:>10.0f}")
            if too_long:
                count, limit = too_long(texts)
                if count:
                    print(f"           {count} chunk(s) exceed the model's {limit} tokens; their ends are ignored")
            for question in misses:
                print(f"           MISS {question}")
        print("\nhit@k: questions where one of the top chunks holds the answer. "
              "chars sent: what those chunks add to every prompt.")
        return

    if args.action == "export":
        chunks = make_chunks(args.strategy, args.source, args.size, args.overlap, raw, blocks, title)
        out = HERE / "chunks.jsonl"
        with out.open("w", encoding="utf-8") as f:
            for n, c in enumerate(chunks, 1):
                record = {"id": f"{args.pdf.name}#{n}", "text": c["text"],
                          "metadata": metadata(args.pdf, raw, args.strategy, c)}
                f.write(json.dumps(record, ensure_ascii=False) + "\n")
        print(f"Wrote {len(chunks)} chunks to {out.relative_to(HERE.parent)}")
        print("First record:\n" + json.dumps(json.loads(out.read_text(encoding='utf-8').splitlines()[0]),
                                              indent=2, ensure_ascii=False))


if __name__ == "__main__":
    main()

The table after Step 7 explains each part.

Step 3: Extract the PDF and compare raw with clean

  1. Run:

    python unit07/chunk_lab.py extract
    Wrote the sample manual to unit07/chunk_docs/p2p-invoice-manual.pdf
    Raw text: 2655 characters -> p2p-invoice-manual.raw.txt
    Clean Markdown: 11 headings, 13 paragraphs, 2 tables -> p2p-invoice-manual.clean.md
    Open both files side by side and compare them.
  2. Open unit07/chunk_docs/p2p-invoice-manual.pdf in your PDF viewer. It looks like a normal three-page manual.

  3. Open p2p-invoice-manual.raw.txt in VS Code. This is pypdf's plain extraction. Find the price table:

    chase order price by more than the tolerance for the company code:
    Company code Tolerance Approver
    1000 Germany 2 percent AP team lead
    2000 France 5 percent Controller
    3000 Spain 3 percent AP team lead
    Page 1 of 3
    ACME Corp - Internal - Vendor invoice handling manual
    2.3 Missing goods receipt

    Four problems in eight lines: the word "purchase" is broken over two lines, the columns are gone, a page footer and the next page's header sit in the middle of the text, and nothing marks 2.3 as a heading.

  4. Open p2p-invoice-manual.clean.md. The same part now reads:

    ### 2.2 Price differences
    
    An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
    
    | Company code | Tolerance | Approver |
    |---|---|---|
    | 1000 Germany | 2 percent | AP team lead |
    | 2000 France | 5 percent | Controller |
    | 3000 Spain | 3 percent | AP team lead |
    
    ### 2.3 Missing goods receipt

    The clean() function read the layout extraction, dropped lines that repeat on every page, re-joined "pur-chase", turned column gaps into table rows and marked numbered lines as headings.

Step 4: Look at the chunks each strategy makes

  1. Structure-aware chunks (the default):

    python unit07/chunk_lab.py show
    headings on clean text, size 600, overlap 0: 11 chunks
    
    --- chunk 5 (page 1, 354 chars) ---
    Vendor invoice handling manual > 2.2 Price differences
    An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
    
    | Company code | Tolerance | Approver |
    | 1000 Germany | 2 percent | AP team lead |
    | 2000 France | 5 percent | Controller |
    | 3000 Spain | 3 percent | AP team lead |

    (Only chunk 5 is shown here; you'll see all 11.) The table stays with its sentence and its section label, and the chunk knows its page.

  2. Recursive chunks at 300 characters, with the top 3 for a question. Add --offline if the Unit 3 model can't load on your network:

    python unit07/chunk_lab.py show --strategy recursive --size 300 --overlap 50 --question "What is the price tolerance for company code 2000?" --offline
    --- chunk 1 (158 chars) ---
    ### 2.2 Price differences
    
    An invoice is blocked when the invoice price differs from the purchase order price by more than the tolerance for the company code:
    
    --- chunk 2 (211 chars) ---
    | Company code | Tolerance | Approver |
    |---|---|---|
    | 1000 Germany | 2 percent | AP team lead |
    | 2000 France | 5 percent | Controller |
    | 3000 Spain | 3 percent | AP team lead |
    
    ### 2.3 Missing goods receipt

    The table got separated from its sentence and glued to the next heading. Here both pieces still made the top 3, but a model now has to guess that the table belongs to "price differences".

  3. Fixed-size chunks on the raw text, for a payment question:

    python unit07/chunk_lab.py show --strategy fixed --source raw --question "When are vendors paid?" --offline --top 2

    Look at where the chunks end. The first one stops in the middle of a time, after 12:, and the next starts with 00 on a payment day. Fixed-size cutting doesn't know what a word is.

Step 5: Measure every strategy

  1. Run the comparison:

    python unit07/chunk_lab.py compare --offline
    strategy   source  size overlap chunks avg chars  hit@3 chars sent
    fixed      raw      300       0      9       295    6/8        883
               MISS How do we recognize a duplicate invoice?
               MISS Why is an invoice blocked when the price is different from the purchase order?
    fixed      raw      300     100     14       279    7/8        822
               MISS Why is an invoice blocked when the price is different from the purchase order?
    fixed      clean    300       0      9       290    5/8        878
               MISS Who may release an invoice above 10,000 euros?
               MISS What should I do when the goods receipt is missing?
               MISS How do we recognize a duplicate invoice?
    recursive  clean    300      50     13       217    7/8        648
               MISS When are vendors paid?
    recursive  clean    800     100      4       685    8/8       2085
    fixed      clean   1500       0      2      1306    7/8       2611
               MISS Who may release an invoice above 10,000 euros?
    headings   clean    600       0     11       260    8/8        812
  2. Read the table row by row:

    • hit@3 counts questions where one of the top 3 chunks contains all the expected phrases from EVAL.
    • chars sent is what those 3 chunks add to every prompt: a direct proxy for cost.
    • fixed raw can never answer the price question: in the raw text, the phrase is broken as "pur- chase".
    • Overlap fixed one boundary miss for fixed-size chunks (6 to 7), at the price of 5 more chunks.
    • recursive 800 and headings both found 8 of 8, but recursive sends about 2.5 times as many characters per question.
    • fixed 1500 cuts the manual into two huge chunks and still misses a question.
  3. Now run it with the real embedding model from Unit 3 (drop --offline):

    python unit07/chunk_lab.py compare

    Your numbers will differ from the ones above, because the model matches meaning, not words. If a line says chunk(s) exceed the model's 256 tokens; their ends are ignored, those chunks were too long for the embedding model. Compare which strategy wins with both embedders.

  4. Try --top 1. With only one chunk allowed, every strategy loses some questions. That is the honest test: does the best chunk hold the answer?

Step 6: Export the chunks with metadata

  1. Export the structure-aware chunks:

    python unit07/chunk_lab.py export --strategy headings
    Wrote 11 chunks to unit07/chunks.jsonl
    First record:
    {
      "id": "p2p-invoice-manual.pdf#1",
      "text": "Vendor invoice handling manual > Introduction\nVendor invoice handling manual\n\nVersion 3.2. Valid from 2026-07-01. Owner: accounts payable. This manual replaces version 3.1.",
      "metadata": {
        "source": "p2p-invoice-manual.pdf",
        "page": 1,
        "section": "Introduction",
        "process": "procure-to-pay",
        "doc_version": "3.2",
        "valid_from": "2026-07-01",
        "language": "en",
        "chunker": "headings"
      }
    }
  2. Open unit07/chunks.jsonl. Each line is one chunk. The metadata block is what the next topics use for filtering (process, company code), citing (page, section) and freshness (version, valid-from date). Fixed and recursive chunks get "page": 0, meaning "unknown": a blind cut loses track of where it came from unless you add extra bookkeeping.

Step 7: Try your own PDF and save your work

  1. Copy a PDF you are allowed to use (a public product manual, not a confidential document) into unit07/chunk_docs/, for example my-guide.pdf.

  2. Extract it and look at the clean Markdown:

    python unit07/chunk_lab.py extract --pdf unit07/chunk_docs/my-guide.pdf

    If the clean file has no headings, your PDF probably doesn't number its headings. That is the heuristic's limit: everything lands in one "Introduction" section, split by paragraph.

  3. If the script prints No text found, the PDF is a scan. It needs OCR first; that isn't covered here.

  4. Keep generated files out of Git. Open .gitignore, add these lines at the end and save:

    unit07/chunk_docs/*.raw.txt
    unit07/chunk_docs/*.clean.md
  5. Save your work:

    git add requirements.txt .gitignore unit07/chunk_lab.py unit07/chunk_docs/p2p-invoice-manual.pdf unit07/chunks.jsonl
    git commit -m "Unit 7: chunk lab with PDF preparation and chunking comparison"

Don't commit your own PDFs unless you may publish them.

What the code does

Part What it does
SAMPLE_PAGES, write_pdf A made-up three-page manual, written as a real PDF by hand, with a header and page number on every page
EVAL Eight questions, each with phrases that must all appear in one retrieved chunk
extract pypdf text per page, in plain or layout mode
clean Drops repeated headers and footers, re-joins hyphenated words, joins paragraph lines, turns column gaps into table rows, marks numbered lines as headings
to_markdown Writes the blocks as Markdown so you can check them
chunk_fixed Cuts every size characters, stepping back by overlap
chunk_recursive Splits by paragraph, line, sentence, word until pieces fit, then packs them with overlap
chunk_headings One chunk per section with a title > section prefix; tables stay whole; long sections split by paragraph
toy_embed, get_embedder Word-hash vectors for --offline, or the Unit 3 model plus a token counter for the length warning
top_k, evaluate Cosine similarity in memory; hit rate at k and characters sent
metadata, export Labels each chunk (source, page, section, process, version, valid-from date) and writes chunks.jsonl

If something goes wrong

What you see What it means What to do
python is not recognized / command not found Python isn't on your PATH, or the terminal was open before you installed it Close and reopen the terminal; see Set up your computer for this course
pypdf is not installed The library is missing in the Python you're using Check for (.venv) in the prompt, then pip install -r requirements.txt
pip install fails with a proxy or SSL error The company network blocks the package index Ask IT to allow pypi.org, or try on another network
sentence-transformers is not installed Unit 3 setup wasn't done in this environment Install it as in Set up for Unit 3, or add --offline
Could not load the model (...) The model download is blocked by a proxy or no network Add --offline, or ask IT to allow the model download
File not found: ... The --pdf path is wrong Run from the course folder; use a path like unit07/chunk_docs/my-guide.pdf
No text found The PDF is a scan (an image of text) Use a PDF with real text, or run OCR on it first
--overlap must be smaller than --size Overlap would repeat the whole chunk Lower --overlap
The clean Markdown has no tables or headings Your PDF's layout doesn't match the simple rules in clean() Read the raw and clean files; adjust the heading pattern, or try a layout-aware parsing library

The SAP way

As of October 2026, SAP gives you three places where chunking happens.

SAP AI Core grounding, Pipeline API. SAP's release notes describe the Pipeline API (added November 2024) as segmenting data into chunks and generating embeddings. Since February 2025 it also has endpoints to check the processing status of each document. SAP Learning describes the document store as holding files such as PDFs and text files. You point it at a repository; SAP extracts, chunks and embeds. Check the current Pipeline API documentation for the file types and any chunking options your release supports; this course could not confirm a list of configurable chunking settings.

SAP AI Core grounding, Vector API. You send chunks you prepared. The SAP Cloud SDK for AI documents a document as a list of chunks, each with content and key/value metadata, plus key/value metadata on the document itself. Since February 2026, SAP's release notes list metadata on documents, collections and chunks, and metadata filtering in search. That is exactly what chunks.jsonl holds.

Sketch (unit07/vector_api_body_sketch.py, run after Step 6):

"""SKETCH: turn chunks.jsonl into one document-with-chunks body for SAP AI Core's Vector API.
It only builds and prints the JSON; it sends nothing. Sending it needs an SAP AI Core service key,
a resource group and a collection (Unit 7, "SAP HANA Cloud vector engine", and Unit 5 setup).
Check the field names against the Vector API reference for your release before you send it.
"""
import json
from pathlib import Path

HERE = Path(__file__).resolve().parent
records = [json.loads(line) for line in (HERE / "chunks.jsonl").read_text(encoding="utf-8").splitlines()]
first = records[0]["metadata"]
document_keys = ("source", "process", "doc_version", "valid_from", "language")   # same for the whole file
body = {"documents": [{
    "metadata": [{"key": k, "value": [str(first[k])]} for k in document_keys],
    "chunks": [{"content": r["text"],
                "metadata": [{"key": "page", "value": [str(r["metadata"]["page"])]},
                             {"key": "section", "value": [r["metadata"]["section"]]}]}
               for r in records]}]}
print(json.dumps(body, indent=2, ensure_ascii=False)[:900] + "\n...")
print(f"\n{len(records)} chunks in one document. Nothing was sent.")

It prints the start of the request body and ends with 11 chunks in one document. Nothing was sent. Compare it with chunks.jsonl: document-level labels (source, process, version) are stored once on the document; page and section go on each chunk.

SAP HANA Cloud text splitter. The hana-ml Python client has a TextSplitter that runs inside SAP HANA Cloud. Its documented settings include chunk_size, overlap, split_type (default recursive), doc_type (default plain) and language (default auto). Input is a table with ID and TEXT columns. Use it when the text already sits in the database, for example service notes or long texts copied from SAP, and you want chunks without moving data out.

SAP Document AI is not a chunker. It extracts structured fields from documents such as invoices and purchase orders. If the goal is "read the PDF and fill SAP fields", use extraction; if it is "answer questions from the PDF", use RAG.

Licensing notes. pypdf, Sentence Transformers and Chroma are free. SAP AI Core grounding usage is billed against your SAP BTP account; check the metrics in your contract. hana-ml runs against an SAP HANA Cloud instance you pay for (or your trial).

Build vs. SAP

Situation Your own preparation (this topic) Pipeline API Vector API with your chunks HANA TextSplitter
Learning, offline tests Best Needs SAP AI Core Needs SAP AI Core Needs SAP HANA Cloud
Messy PDFs with tables Full control of cleaning SAP's extraction Your cleaning, SAP's store Text must already be clean
Custom metadata (company code, version) Yours Check what the pipeline sets Key/value per document and chunk Columns in your own tables
Documents in SharePoint, S3, SFTP You write connectors Built-in You write connectors You load the table
Operations and re-indexing Yours SAP runs the pipeline Shared: you prepare, SAP stores Database jobs
Effort Highest Lowest Medium Low if data is in HANA

A practical path: use this lab to learn what your documents need. If SAP's pipeline handles them well on your test questions, use it. If tables or metadata suffer, prepare chunks yourself and upload them with the Vector API.

Production concerns

  • Security and authorizations. Chunking can mix content with different audiences in one chunk. Keep access labels (company code, sales organization, confidentiality) in chunk metadata and filter on them from the user's identity. Never place two audiences' text in one chunk. Unit 7's topic on grounding with SAP authorizations builds on this.
  • Prompt injection. Documents may contain text written to steer the model. Cleaning doesn't remove that; treat chunk text as data, keep instructions in the system message, and log chunk IDs.
  • Versioning. Store doc_version and valid_from. When a manual changes, delete its old chunks before indexing the new ones; mixed versions produce contradictory answers.
  • Re-indexing cost. Changing chunk size or strategy means re-chunking and re-embedding everything. Decide with a test, record the settings in metadata (chunker), and change them on purpose.
  • Evaluation. Keep the chunking test set with the code, and re-run it when documents or settings change. Track hit rate and characters sent together.
  • Embedding limits. Check chunk length against your embedding model's input limit. The lab counts tokens for the Unit 3 model; do the same for SAP-hosted embedding models.
  • Personal data. Manuals rarely hold personal data, but exports, tickets and e-mails do. Decide what may be indexed before you chunk it.
  • Clean core. Preparation happens outside SAP on documents and released APIs. Nothing here modifies SAP; keep it side by side on SAP BTP.

Pitfalls

  • Chunking raw PDF text. Headers, page numbers and broken words land in chunks and in answers. Clean first, and read the clean file.
  • Orphaned tables. A table cut away from its heading answers nothing. Keep tables whole, with their header row and a section label.
  • Chunks past the embedding limit. The end of a long chunk is never embedded, so questions about it don't match.
  • Copying someone else's settings. 200, 512 or 1,000 tokens are starting points. Test on your documents.
  • Overlap as a reflex. Overlap costs storage and prompt tokens and returns near-duplicate chunks. Use it when boundaries cut answers.
  • No page or section in metadata. Without them you can't cite precisely or debug a wrong answer.
  • Scans treated as text. A scan yields no text, and an empty document is easy to miss in a big batch. Count characters per file and flag zeros.
  • Judging chunks by eye. Readable chunks can still retrieve badly. Measure.

Exercise

Tune chunking for one more document and record the result. The report, unit07/chunking_report.txt, feeds the vector database and evaluation topics later in Unit 7 and Unit 8.

  1. Open unit07/chunk_lab.py. In SAMPLE_PAGES, add a fourth page with a section 7 Credit memos and a short table (three rows: amount range, approver, deadline). Save.

  2. Delete the old sample so it is rewritten: delete the folder unit07/chunk_docs in VS Code (right-click, Delete).

  3. Add two questions to EVAL whose answers are in your new table, each with the phrases that must appear. Save.

  4. Run python unit07/chunk_lab.py compare --offline and note the hit rate and characters sent for each row.

  5. Add one row to the runs list in main(), for example ("headings", "clean", 300, 0), and run compare again. Does a smaller limit split your table? Check with show --size 300.

  6. Choose a strategy and size. Save the comparison to a file:

    Windows (PowerShell):

    python unit07/chunk_lab.py compare --offline | Out-File -Encoding utf8 unit07/chunking_report.txt

    macOS/Linux:

    python unit07/chunk_lab.py compare --offline > unit07/chunking_report.txt
  7. Add one line at the end of the file, by hand: the strategy and size you chose, and why, in one sentence.

  8. Run python unit07/chunk_lab.py export --strategy headings with your chosen size (add --size N), then commit unit07/chunk_lab.py, unit07/chunks.jsonl and unit07/chunking_report.txt.

Done when EVAL has ten questions, unit07/chunking_report.txt shows at least eight comparison rows plus your one-line decision, and unit07/chunks.jsonl holds chunks from four pages with page and section metadata.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does pypdf's plain extraction lose the columns of the price table?

    Answer: C. The pypdf documentation says PDFs have no semantic layer and tables are absolutely positioned text. Layout mode keeps the spacing, which clean() turns back into table rows.
  2. 2In clean(), how are page headers and footers such as "Page 2 of 3" removed?

    Answer: B. Masking digits makes "Page 2 of 3" and "Page 3 of 3" identical, and lines that appear on every page are treated as headers or footers. Deleting fixed line positions would break on PDFs with different layouts.
  3. 3The fixed-raw run can never find the answer to the price question. Why?

    Answer: D. Raw extraction keeps "pur-" and "chase" on separate lines, so the expected phrase doesn't exist in any raw chunk. Cleaning re-joins the word before chunking.
  4. 4Recursive chunking at 800 and headings at 600 both find 8 of 8. Which do you choose, and why?

    Answer: A. Hit rate ties, but recursive 800 sends about 2.5 times the characters into every prompt, which is paid on every call. Track quality and cost together.
  5. 5The compare run with the real model warns that some chunks exceed 256 tokens. What happens to them?

    Answer: C. The model card says input longer than 256 word pieces is truncated. The cut-off text can't help matching, yet the whole chunk still costs prompt tokens when retrieved.
  6. 6What is the difference between the title > section prefix and the metadata block?

    Answer: B. Text in the chunk is embedded, so the prefix helps matching. Metadata sits beside the text, where code can filter by process, show page and section, and remove old versions.
  7. 7A manual moves from version 3.1 to 3.2 and users get contradictory answers. What do you do?

    Answer: D. Old and new chunks both match, so the model sees conflicting rules. Version metadata lets you remove the old chunks cleanly before indexing the new ones.
  8. 8You run on SAP AI Core and need your own company-code labels on every chunk. Which route fits?

    Answer: C. The Vector API takes your chunks with key/value metadata on documents and chunks, and SAP documents metadata filtering since February 2026. The Pipeline API chunks for you, so you control less of what each chunk carries.
  9. 9Finance asks you to read invoice amounts from supplier PDFs into SAP. What do you tell them?

    Answer: B. Pulling structured fields from invoices is document extraction. SAP Document AI is built for that; RAG chunking is for answering questions from text.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in