Orchestrate

Document AI, built and bought: from supplier invoice PDF to checked SAP data

How document AI turns invoices and orders into checked fields, what SAP Document AI offers, and how to build, measure and route an extractor yourself.

Updated Oct 9, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

A large share of business still arrives as documents: supplier invoices, customer purchase orders, delivery notes, payment advices. Someone reads each one and types the numbers into SAP.

Document AI does that reading. It takes a PDF, a scan or a photo and returns fields: invoice number, date, supplier, purchase order number, amounts, line items. It then checks those fields and decides whether a person must look.

The hard part is not reading. Modern models read invoices well. The hard part is knowing when the reading is wrong, before a wrong amount is posted and paid. Good document AI is a pipeline: read, fill a fixed form, check, route the doubtful ones to a person, then post.

You can buy this. SAP sells SAP Document AI as a service, and S/4HANA Cloud uses it inside some standard apps. You can also build it with open tools and a language model. This topic helps you choose, and tells you what to measure either way.

Why it matters to the business

Take procure-to-pay. Accounts payable receives supplier invoices as PDFs by email. A clerk keys in each one. Then SAP compares the invoice with the purchase order and the goods receipt: the three-way match. If the price or quantity differs, the invoice is blocked and someone investigates.

Document AI attacks the first step, the keying. The value is real:

  • Time. SAP's product page claims "up to 70%" less document processing time, for a modelled company with stated assumptions. Treat that as SAP's estimate, not your baseline.
  • Speed to pay. Faster entry means fewer missed early-payment discounts and fewer late-payment complaints.
  • Scale. Volume can grow without hiring in step.

The risk is just as real. A model that misreads a total with confidence creates a silent error. The invoice passes, is paid, and nobody notices until the supplier statement disagrees. A missed field that stops for review costs a minute. A wrong field that sails through costs money and trust.

So the business question is not "how accurate is it?" It is: what share can go straight through, and how many errors slip through unseen? Ask for both numbers, measured on your own documents.

How SAP does it

As of October 2026, SAP's offering is called SAP Document AI. Its API keeps an older name, Document Information Extraction. There are three ways to meet it.

1. Inside SAP applications. SAP's Q1 2026 release highlights list two features as generally available in S/4HANA Cloud Public Edition:

  • Sales order creation from unstructured data. A sales rep uploads a PDF or image purchase order; SAP Document AI extracts the data and the system proposes a sales order for review.
  • Payment advice processing with SAP Document AI. It extracts amounts, references and currencies from payment advices. SAP claims 70% less processing time and 83% less template maintenance.

Here you don't build anything. You switch on and configure a standard process.

2. As a service on SAP BTP. You send documents through its API or a web interface and get fields back. Per SAP's product page, it covers invoices, purchase orders and delivery notes out of the box, and reads more than 35 file formats. You define your own schema (the list of fields you want) for other documents. Extraction blends pretrained models with large language models, and reviewers confirm or correct results in a web interface. Results can be pushed into S/4HANA.

3. Embedded edition. SAP also lists an embedded edition, which reads documents from channels such as email inboxes and processes them into a target system such as S/4HANA. It is activated with AI Units.

On plans, SAP's pages list Free, Base, Premium and Embedded, and the product FAQ mentions Premium Plus. "Instant learning" from user corrections needs Premium or Premium Plus. Prices are not public on the pages we opened; confirm them with SAP before you build a business case.

Built or bought: a decision guide

Your situation Lean towards Why
Standard S/4HANA Cloud process already covers the document, such as PDF purchase orders to sales orders The embedded SAP feature No build, SAP maintains it, review steps are part of the app
Common business documents, many suppliers, many languages, posting into SAP SAP Document AI service OCR, schemas, review interface and SAP integration come ready
An unusual document type, such as a quality certificate or a customs form SAP Document AI with a custom schema, or a build Test both on 50 of your real documents before deciding
Data may not leave a site, or you need full control of the model A build with a model you host You own the model, the checks and the operations
A few hundred documents a year Neither, yet A clerk may still be cheaper than any project

Whichever you pick, the checks after extraction are yours to design. A supplier's PDF never knows your purchase orders.

Questions to ask

Ask your team:

  • How many documents of each type arrive per month, in which formats and languages?
  • What does one manual entry cost today, and how often is it wrong? That is the baseline.
  • Which fields are critical, where an error costs money (amount, bank details, purchase order)? Which are merely helpful?
  • Who reviews the doubtful ones, and how fast must they do it?

Ask a vendor or partner, SAP included:

  • What straight-through rate and what silent error rate did you measure, on whose documents?
  • How is confidence calculated per field, and can we set thresholds per field?
  • Where are documents stored, for how long, and in which region? SAP's page says seven days by default, configurable.
  • What happens when a supplier changes its layout? Who notices, and who fixes it?
  • How does the extracted data reach SAP, and which SAP checks still run on it?

Common misconceptions

  • "Document AI posts invoices for us." It proposes fields. SAP's own checks, such as the three-way match, still run, and a person still handles the exceptions.
  • "99% accuracy means 99% of invoices are right." Accuracy is usually per field. If each of 10 fields is 99% right and errors are independent, a whole invoice is right only about 90% of the time.
  • "High confidence means correct." Confidence is the system's estimate. You must measure how often high-confidence fields are still wrong.
  • "Templates are obsolete." A template is cheap and exact for one stable layout. It breaks quietly when the layout moves.
  • "A three-way match exception means the AI failed." Often the AI read correctly, and the supplier really billed a different price. That exception is the process working.
  • "Generative AI can't be audited." Keep the document, the extracted fields, the checks and the reviewer's decision. That trail is auditable.

Key terms

  • Document AI (intelligent document processing): software that turns documents into structured data and checks it.
  • OCR (optical character recognition): turning an image of text into text a computer can read.
  • Schema: the fixed list of fields to extract, each with a name, type and description.
  • Template: extraction rules tied to one exact layout.
  • Confidence score: the system's own estimate, per field, that a value is right.
  • Straight-through processing: documents that pass every check and need no person.
  • Silent error: a wrong value that passed every check.
  • Enrichment: matching extracted values to SAP master data, such as finding the supplier number.
  • Three-way match: comparing invoice, purchase order and goods receipt before payment.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A vendor says its document AI is "98% accurate". What should a finance leader ask next?

    Answer: C. Per-field accuracy hides the two numbers that decide value and risk. The straight-through rate drives savings, and silent errors drive losses. Both must be measured on your own documents.
  2. 2Your company runs S/4HANA Cloud Public Edition and wants customer PDF purchase orders turned into sales orders. What is the sensible first option?

    Answer: A. As of SAP's Q1 2026 highlights, sales order creation from PDF or image purchase orders is generally available in S/4HANA Cloud Public Edition and uses SAP Document AI. A standard feature needs no build and keeps the review step inside the app.
  3. 3Why is a silent error worse than a field the system flags as missing?

    Answer: D. A flagged field costs a reviewer a minute. A wrong value that passes every check reaches SAP and may be paid, and nobody looks until something else disagrees.
  4. 4An invoice is extracted correctly, but SAP blocks it because the billed price is above the purchase order price. What does this tell you?

    Answer: B. Extraction and the three-way match do different jobs. When the reading is right and the price differs, the exception is genuine and belongs with a person.
  5. 5When does building your own document AI make more sense than buying SAP Document AI?

    Answer: C. SAP Document AI ships OCR, schemas, review and SAP integration for common documents. A build earns its cost when you must host the model yourself or control it fully. At very low volume, neither may pay off.
  6. 6Which question best protects you when a vendor proposes document AI?

    Answer: D. Layouts change without warning. You need to know who notices and how extraction keeps working, because that is where quality quietly drops after go-live.
Deep layer · 40 min read

Mental model: read, fill, check, route

Think of document AI as a careful new clerk with a form.

  1. Read. Get text from the document. If it's a scan, that needs OCR first.
  2. Fill. Copy values into a fixed form: the schema. The form never changes shape, whatever the supplier's layout.
  3. Check. Test every value against things you know: the document itself, arithmetic, your SAP master data and open purchase orders.
  4. Route. If every check passes, send it on. If any fails, a person looks.

Most builders over-invest in step 2 and under-invest in step 3. A modern model fills the form well most of the time. The checks decide whether "most of the time" is safe. Your goal is not zero extraction errors. It is zero silent errors: wrong values that pass every check.

How it works

The pipeline

flowchart LR
  A[Intake<br/>email, upload, API] --> B[Digitize<br/>text layer or OCR]
  B --> C[Classify<br/>which document type]
  C --> D[Extract<br/>fill the schema]
  D --> E[Check and enrich<br/>rules, master data]
  E -->|all pass| F[Post to SAP<br/>via API]
  E -->|any fail| G[Human review]
  G --> F
  G -.corrections.-> D
  • Intake collects documents from inboxes, uploads or other systems.
  • Digitize gets text. A PDF made by software usually has a text layer you can read directly. A scan or phone photo is an image, and needs OCR. The chunking topic showed that pypdf reads text layers but can't read images.
  • Classify decides the document type, so the right schema is used.
  • Extract fills the schema.
  • Check and enrich validates values and adds SAP keys, such as the supplier number.
  • Review shows a person the document beside the doubtful fields.
  • Post sends approved data to SAP through an API. SAP's own checks run again there.
  • Corrections from review can improve later extraction.

Three ways to fill the schema

Approach How it works Strong when Weak when
Template Rules for one layout: "the number after Invoice No.:" One supplier, stable layout, high volume Any layout change; every new supplier needs new rules
Pretrained extraction model A model trained on many invoices finds standard fields Common fields on common document types Unusual documents or fields it wasn't trained for
Language model with a schema A general model reads the text and fills a JSON schema from field descriptions New layouts, new fields, many languages Cost per page, speed, and confident mistakes

SAP Document AI offers all three, as its tutorial shows (see The SAP way). In your own build, the third approach uses structured outputs, from structured outputs and function calling. Ollama's format parameter takes a JSON schema and constrains the answer to it. That guarantees the shape of the answer, not the truth of the values.

Confidence: two different things

There are two ways to decide "auto or review":

  • Model confidence. The extraction system scores each field. SAP's architecture guide describes scores from 0 to 100% per field and a threshold, "typically 90% for critical fields". It is useful, but it is the model grading itself.
  • Independent checks. Tests the model can't talk its way past. Is the invoice number printed in the document? Do net plus tax equal gross? Do the lines add up to net? Does the purchase order exist, for this supplier?

Use both when you can. In your build today, you will use only independent checks, because a small local model gives no reliable per-field score. You'll see how far checks alone get you.

Enrichment and the three-way match

Extraction gives you what the supplier printed. SAP needs its own keys: the supplier number, the purchase order item, the company code. Enrichment maps one to the other using master data.

Then the three-way match compares the invoice with the purchase order price and the goods receipt quantity. SAP's enrichment patterns guide lists exactly this as an enrichment example. Keep two kinds of "stop" apart:

  • Extraction stop: the reading may be wrong. A person checks the document.
  • Business stop: the reading is right, and the supplier billed something different from what was ordered or received. A buyer or clerk resolves it with the supplier.

Mixing them up wrecks your metrics. A genuine price difference is not an AI error.

Measuring it

Measure on a test set: documents with the right answers written down by a person. Track:

Metric Question it answers
Field accuracy, per field Which fields does it struggle with?
Documents with every field right How often is the whole invoice right?
Straight-through rate How much work does it remove?
Silent errors How much risk does it add?
Seconds and cost per document Can we afford it at our volume?

LLM evaluation fundamentals and building an evaluation harness cover the general method. Here it is applied to documents.

Build it yourself: an invoice extractor with checks

You will build a small accounts payable pipeline. A script makes eight made-up supplier invoices as PDFs, from four suppliers with four different layouts. A second script reads them, fills a fixed schema two ways (hand-written rules and a local model), checks every invoice, runs a three-way match against made-up purchase orders, and scores everything against the right answers.

Before you start: complete Set up your computer for this course and Set up for Unit 12: local models. They install Python, Ollama and the ollama library, and pull qwen3:0.6b. This walkthrough doesn't repeat those steps. It also uses pypdf, which chunking and document preparation added in Unit 7; Step 2 adds it if you skipped that topic.

flowchart LR
  S3[Step 3<br/>make 8 invoices] --> S4[Step 4<br/>see the text]
  S4 --> S5[Step 5<br/>rules]
  S5 --> S6[Step 6<br/>local model]
  S6 --> S8[Step 8<br/>read results]
  S8 --> S9[Step 9<br/>save in Git]

What you need

  • Your course folder orchestrate-course with its .venv, Ollama running and qwen3:0.6b pulled, from the Unit 12 setup.
  • Optional: about 1.4 GB of free disk for qwen3:1.7b, the next size up (Step 7).
  • 30 to 45 minutes.
  • Cost: free. No account, no key, no internet after the model download. Everything runs on your computer.

Step 1: Open your course folder

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal with Terminal > New Terminal. Turn on the virtual environment if the prompt doesn't start with (.venv):

    Windows (PowerShell):

    .venv\Scripts\Activate.ps1

    macOS or Linux:

    source .venv/bin/activate
  3. Check that the unit12 folder exists. If it doesn't, create it:

    mkdir unit12

Run every command in this topic from the course folder, not from inside unit12.

Step 2: Check your libraries

This topic needs no new library. Check that the two it uses are installed:

pip show pypdf ollama

You should see two blocks, one starting Name: pypdf and one starting Name: ollama. If you see WARNING: Package(s) not found: pypdf:

  1. Open requirements.txt, add this line at the end, and save:

    pypdf
  2. Install it:

    pip install -r requirements.txt

If ollama is missing, go back to Set up for Unit 12, which adds it.

Step 3: Make the invoices

This script writes eight one-page PDF invoices. Each of four suppliers labels things differently:

Layout Supplier What makes it different
A Brightline Components Ltd Clean English labels and ISO dates
B Kessler Antriebstechnik GmbH German labels, dates like 14.09.2026, amounts like 1.457,75
C Pacific Valve Co. US style: Sep 15, 2026, dollar signs, "Customer PO"
D Atlas Facility Services A service invoice in sentences; one has no purchase order at all

It also writes truth.json, the right answer for every field, and open_pos.json, made-up purchase orders with goods receipts. Two invoices are planted to fail the three-way match: one bills a higher price than ordered, one bills more than was received.

  1. In VS Code's file list, right-click unit12, choose New File, name it make_invoices.py, paste the code below and save.
"""Make eight made-up supplier invoices as PDFs, plus the right answers and the open purchase orders.

Run it from your course folder:
    python unit12/make_invoices.py

It writes:
    unit12/invoices/INV-*.pdf      eight invoices from four suppliers, each with its own layout
    unit12/invoices/truth.json     the correct field values for every invoice (your test set)
    unit12/open_pos.json           purchase orders with goods receipts, for a three-way match

Everything is made up. No library beyond Python itself, no account, no internet.
"""
import json
from pathlib import Path

FOLDER = Path("unit12") / "invoices"

# The right answers. Field names are this course's own, not SAP field names.
TRUTH = [
    {"file": "INV-01.pdf", "layout": "A", "invoice_number": "BL-24117", "invoice_date": "2026-09-02",
     "supplier_name": "Brightline Components Ltd", "po_number": "4500000101", "currency": "EUR",
     "lines": [{"description": "Hex bolt M8 x 40, zinc", "quantity": 500, "unit_price": 0.42},
               {"description": "Washer M8, steel", "quantity": 500, "unit_price": 0.06}],
     "net_amount": 240.00, "tax_amount": 45.60, "gross_amount": 285.60},
    {"file": "INV-02.pdf", "layout": "A", "invoice_number": "BL-24130", "invoice_date": "2026-09-09",
     "supplier_name": "Brightline Components Ltd", "po_number": "4500000102", "currency": "EUR",
     "lines": [{"description": "Bearing 6204-2RS", "quantity": 40, "unit_price": 3.85}],
     "net_amount": 154.00, "tax_amount": 29.26, "gross_amount": 183.26},
    {"file": "INV-03.pdf", "layout": "B", "invoice_number": "RE-2026-0815", "invoice_date": "2026-09-14",
     "supplier_name": "Kessler Antriebstechnik GmbH", "po_number": "4500000103", "currency": "EUR",
     "lines": [{"description": "Getriebemotor GM-40", "quantity": 2, "unit_price": 612.50}],
     "net_amount": 1225.00, "tax_amount": 232.75, "gross_amount": 1457.75},
    {"file": "INV-04.pdf", "layout": "B", "invoice_number": "RE-2026-0822", "invoice_date": "2026-09-18",
     "supplier_name": "Kessler Antriebstechnik GmbH", "po_number": "4500000104", "currency": "EUR",
     "lines": [{"description": "Kupplung K-25", "quantity": 10, "unit_price": 48.90}],
     "net_amount": 489.00, "tax_amount": 92.91, "gross_amount": 581.91},
    {"file": "INV-05.pdf", "layout": "C", "invoice_number": "PV-88213", "invoice_date": "2026-09-15",
     "supplier_name": "Pacific Valve Co.", "po_number": "4500000105", "currency": "USD",
     "lines": [{"description": "Ball valve 2in, stainless", "quantity": 12, "unit_price": 185.00},
               {"description": "Gasket kit 2in", "quantity": 12, "unit_price": 19.50}],
     "net_amount": 2454.00, "tax_amount": 196.32, "gross_amount": 2650.32},
    {"file": "INV-06.pdf", "layout": "C", "invoice_number": "PV-88240", "invoice_date": "2026-09-21",
     "supplier_name": "Pacific Valve Co.", "po_number": "4500000106", "currency": "USD",
     "lines": [{"description": "Check valve 1in", "quantity": 30, "unit_price": 64.00}],
     "net_amount": 1920.00, "tax_amount": 153.60, "gross_amount": 2073.60},
    {"file": "INV-07.pdf", "layout": "D", "invoice_number": "AFS/0931", "invoice_date": "2026-09-30",
     "supplier_name": "Atlas Facility Services", "po_number": None, "currency": "EUR",
     "lines": [{"description": "Warehouse cleaning, September", "quantity": 1, "unit_price": 1850.00}],
     "net_amount": 1850.00, "tax_amount": 351.50, "gross_amount": 2201.50},
    {"file": "INV-08.pdf", "layout": "D", "invoice_number": "AFS/0932", "invoice_date": "2026-09-30",
     "supplier_name": "Atlas Facility Services", "po_number": "4500000108", "currency": "EUR",
     "lines": [{"description": "Forklift inspection", "quantity": 3, "unit_price": 140.00}],
     "net_amount": 420.00, "tax_amount": 79.80, "gross_amount": 499.80},
]

# Made-up purchase orders and what the warehouse received (goods receipts).
# Two invoices are planted to fail the three-way match: INV-04 bills a higher price,
# INV-06 bills more than was received.
OPEN_POS = {
    "4500000101": {"supplier": "Brightline Components Ltd", "lines": [
        {"ordered": 500, "received": 500, "unit_price": 0.42}, {"ordered": 500, "received": 500, "unit_price": 0.06}]},
    "4500000102": {"supplier": "Brightline Components Ltd", "lines": [
        {"ordered": 40, "received": 40, "unit_price": 3.85}]},
    "4500000103": {"supplier": "Kessler Antriebstechnik GmbH", "lines": [
        {"ordered": 2, "received": 2, "unit_price": 612.50}]},
    "4500000104": {"supplier": "Kessler Antriebstechnik GmbH", "lines": [
        {"ordered": 10, "received": 10, "unit_price": 45.00}]},
    "4500000105": {"supplier": "Pacific Valve Co.", "lines": [
        {"ordered": 12, "received": 12, "unit_price": 185.00}, {"ordered": 12, "received": 12, "unit_price": 19.50}]},
    "4500000106": {"supplier": "Pacific Valve Co.", "lines": [
        {"ordered": 30, "received": 24, "unit_price": 64.00}]},
    "4500000108": {"supplier": "Atlas Facility Services", "lines": [
        {"ordered": 3, "received": 3, "unit_price": 140.00}]},
}


def de_amount(x: float) -> str:
    """1457.75 -> '1.457,75' (German style)."""
    return f"{x:,.2f}".replace(",", "#").replace(".", ",").replace("#", ".")


def us_date(iso: str) -> str:
    months = "Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec".split()
    y, m, d = iso.split("-")
    return f"{months[int(m) - 1]} {int(d)}, {y}"


def page_items(t: dict) -> list:
    """Return (x, y, font size, text) for one invoice. Each layout labels things differently."""
    L, items, y = t["layout"], [], 790
    def put(x, text, size=10):
        items.append((x, y, size, text))
    if L == "A":
        put(50, t["supplier_name"], 16); y -= 18
        put(50, "Unit 4, Riverside Park, Leeds"); y -= 40
        put(50, "INVOICE", 14); y -= 24
        put(50, f"Invoice No.: {t['invoice_number']}"); put(330, f"Invoice date: {t['invoice_date']}"); y -= 16
        put(50, f"PO number: {t['po_number']}"); put(330, "Currency: EUR"); y -= 16
        put(50, "Bill to: Orchestrate Demo Plant 1010"); y -= 36
        put(50, "Item"); put(300, "Qty"); put(370, "Unit price"); put(470, "Amount"); y -= 16
        for ln in t["lines"]:
            put(50, ln["description"]); put(300, str(ln["quantity"]))
            put(370, f"{ln['unit_price']:.2f}"); put(470, f"{ln['quantity'] * ln['unit_price']:.2f}"); y -= 16
        y -= 16
        put(370, "Net amount:"); put(470, f"{t['net_amount']:.2f}"); y -= 16
        put(370, "VAT 19%:"); put(470, f"{t['tax_amount']:.2f}"); y -= 16
        put(370, "Total due:"); put(470, f"{t['gross_amount']:.2f}")
    elif L == "B":
        put(330, t["supplier_name"], 12); y -= 16
        put(330, "Industriestr. 12, 70565 Stuttgart"); y -= 50
        put(50, "Rechnung / Invoice", 14); y -= 24
        put(50, f"Rechnung Nr. {t['invoice_number']}"); y -= 16
        d = t["invoice_date"].split("-")
        put(50, f"Datum: {d[2]}.{d[1]}.{d[0]}"); y -= 16
        put(50, f"Ihre Bestellung / Your order: {t['po_number']}"); y -= 36
        put(50, "Pos."); put(90, "Bezeichnung"); put(290, "Menge"); put(350, "Einzelpreis"); put(460, "Gesamt"); y -= 16
        for i, ln in enumerate(t["lines"], 1):
            put(50, f"{i * 10}"); put(90, ln["description"]); put(290, f"{ln['quantity']} St")
            put(350, de_amount(ln["unit_price"])); put(460, de_amount(ln["quantity"] * ln["unit_price"])); y -= 16
        y -= 16
        put(350, "Netto"); put(460, de_amount(t["net_amount"])); y -= 16
        put(350, "MwSt 19 %"); put(460, de_amount(t["tax_amount"])); y -= 16
        put(350, "Gesamtbetrag EUR"); put(460, de_amount(t["gross_amount"])); y -= 30
        put(50, "Zahlbar innerhalb 30 Tagen ohne Abzug.", 8)
    elif L == "C":
        put(50, t["supplier_name"].upper(), 16); y -= 18
        put(50, "1200 Harbor Blvd, Oakland CA"); put(400, f"INVOICE # {t['invoice_number']}", 12); y -= 16
        put(400, f"Date: {us_date(t['invoice_date'])}"); y -= 16
        put(400, f"Customer PO: {t['po_number']}"); y -= 40
        put(50, "Description"); put(280, "Quantity"); put(360, "Price"); put(460, "Ext. price"); y -= 16
        for ln in t["lines"]:
            put(50, ln["description"]); put(280, str(ln["quantity"]))
            put(360, f"${ln['unit_price']:,.2f}"); put(460, f"${ln['quantity'] * ln['unit_price']:,.2f}"); y -= 16
        y -= 16
        put(360, "Subtotal"); put(460, f"${t['net_amount']:,.2f}"); y -= 16
        put(360, "Sales tax 8%"); put(460, f"${t['tax_amount']:,.2f}"); y -= 16
        put(360, "Amount due (USD)"); put(460, f"${t['gross_amount']:,.2f}")
    else:  # D: a service invoice, with or without a purchase order
        put(50, "Atlas Facility Services", 14); y -= 16
        put(50, "Service invoice for work at plant 1010"); y -= 40
        put(50, f"Bill number {t['invoice_number']}, issued {t['invoice_date']}"); y -= 16
        ref = t["po_number"] or "none given (framework agreement FA-12)"
        put(50, f"Purchase order reference: {ref}"); y -= 30
        for ln in t["lines"]:
            put(50, f"{ln['description']}: {ln['quantity']} x EUR {ln['unit_price']:.2f}"); y -= 16
        y -= 10
        put(50, f"Subtotal EUR {t['net_amount']:.2f}, VAT EUR {t['tax_amount']:.2f}"); y -= 16
        put(50, f"Please pay EUR {t['gross_amount']:.2f} by bank transfer within 14 days.")
    return items


def write_pdf(items: list, path: Path) -> None:
    """Write a one-page PDF by hand: text placed at x, y positions, as most PDFs do."""
    def esc(s):
        return s.replace("\\", "\\\\").replace("(", "\\(").replace(")", "\\)")
    stream = "\n".join(f"BT /F1 {s} Tf {x} {y} Td ({esc(t)}) Tj ET" for x, y, s, t in items)
    objects = [
        "<< /Type /Catalog /Pages 2 0 R >>",
        "<< /Type /Pages /Kids [4 0 R] /Count 1 >>",
        "<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>",
        "<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] /Resources << /Font << /F1 3 0 R >> >> /Contents 5 0 R >>",
        f"<< /Length {len(stream.encode('latin-1'))} >>\nstream\n{stream}\nendstream",
    ]
    data, offsets = b"%PDF-1.4\n", []
    for i, obj in enumerate(objects, 1):
        offsets.append(len(data))
        data += f"{i} 0 obj\n{obj}\nendobj\n".encode("latin-1")
    xref = len(data)
    data += f"xref\n0 {len(objects) + 1}\n0000000000 65535 f \n".encode()
    data += "".join(f"{o:010d} 00000 n \n" for o in offsets).encode()
    data += f"trailer\n<< /Size {len(objects) + 1} /Root 1 0 R >>\nstartxref\n{xref}\n%%EOF\n".encode()
    path.write_bytes(data)


def main() -> None:
    FOLDER.mkdir(parents=True, exist_ok=True)
    for t in TRUTH:
        write_pdf(page_items(t), FOLDER / t["file"])
    (FOLDER / "truth.json").write_text(json.dumps(TRUTH, indent=2), encoding="utf-8")
    (FOLDER.parent / "open_pos.json").write_text(json.dumps(OPEN_POS, indent=2), encoding="utf-8")
    suppliers = len({t["supplier_name"] for t in TRUTH})
    print(f"Wrote {len(TRUTH)} invoices to {FOLDER}/ from {suppliers} suppliers")
    print(f"Wrote the right answers to {FOLDER / 'truth.json'}")
    print(f"Wrote {len(OPEN_POS)} purchase orders to {FOLDER.parent / 'open_pos.json'}")


if __name__ == "__main__":
    main()
  1. Run it:

    python unit12/make_invoices.py

    On macOS or Linux, use python3 if python isn't found.

You should see:

Wrote 8 invoices to unit12/invoices/ from 4 suppliers
Wrote the right answers to unit12/invoices/truth.json
Wrote 7 purchase orders to unit12/open_pos.json

On Windows the paths show backslashes. Open unit12/invoices/INV-03.pdf in your browser or PDF viewer to see what a clerk would see.

Step 4: Create the extractor and look at what the computer sees

  1. Right-click unit12, choose New File, name it extract_invoices.py, paste the code below and save.
"""Extract fields from supplier invoice PDFs two ways, check them, and score them against the right answers.

Run it from your course folder (after: python unit12/make_invoices.py):
    python unit12/extract_invoices.py                        # rules only: no model needed
    python unit12/extract_invoices.py --method llm           # a local model through Ollama
    python unit12/extract_invoices.py --method both          # compare the two
    python unit12/extract_invoices.py --method llm --model qwen3:1.7b
    python unit12/extract_invoices.py --show INV-03.pdf      # see the text and the result for one file

Pipeline for every PDF: read the text (pypdf) -> extract fields (rules or model) -> check them
(our own checks, which decide "auto" or "review") -> three-way match against made-up purchase
orders and goods receipts -> compare with truth.json. Results go to unit12/extraction_results.csv.
"""
import argparse
import csv
import json
import re
import sys
import time
from datetime import date
from pathlib import Path

FOLDER = Path("unit12") / "invoices"
POS_FILE = Path("unit12") / "open_pos.json"
OUT_CSV = Path("unit12") / "extraction_results.csv"
FIELDS = ["invoice_number", "invoice_date", "supplier_name", "po_number", "currency",
          "net_amount", "tax_amount", "gross_amount", "lines"]
AMOUNTS = ["net_amount", "tax_amount", "gross_amount"]

# Each field has a plain description. The model reads them as instructions,
# the same idea as field descriptions in an SAP Document AI schema.
DESCRIPTIONS = {
    "invoice_number": "The supplier's own number for this invoice, exactly as printed.",
    "invoice_date": "The date the invoice was issued, written as YYYY-MM-DD.",
    "supplier_name": "The company that sent the invoice (not the customer it is billed to).",
    "po_number": "Our purchase order number that the invoice refers to, or an empty string if none is printed.",
    "currency": "Three-letter ISO currency code, for example EUR or USD.",
    "net_amount": "Total before tax, as a plain number with a dot for decimals.",
    "tax_amount": "The tax (VAT, MwSt or sales tax) amount, as a plain number.",
    "gross_amount": "The total to pay including tax, as a plain number.",
    "lines": "Every invoice line with its description, quantity and unit price as plain numbers.",
}
LINE_SCHEMA = {"type": "object", "properties": {
    "description": {"type": "string"}, "quantity": {"type": "number"}, "unit_price": {"type": "number"}},
    "required": ["description", "quantity", "unit_price"]}
SCHEMA = {"type": "object", "properties": {
    "invoice_number": {"type": "string", "description": DESCRIPTIONS["invoice_number"]},
    "invoice_date": {"type": "string", "description": DESCRIPTIONS["invoice_date"]},
    "supplier_name": {"type": "string", "description": DESCRIPTIONS["supplier_name"]},
    "po_number": {"type": "string", "description": DESCRIPTIONS["po_number"]},
    "currency": {"type": "string", "description": DESCRIPTIONS["currency"]},
    "net_amount": {"type": "number", "description": DESCRIPTIONS["net_amount"]},
    "tax_amount": {"type": "number", "description": DESCRIPTIONS["tax_amount"]},
    "gross_amount": {"type": "number", "description": DESCRIPTIONS["gross_amount"]},
    "lines": {"type": "array", "items": LINE_SCHEMA, "description": DESCRIPTIONS["lines"]},
}, "required": FIELDS}

SYSTEM = ("You extract fields from one supplier invoice for an accounts payable clerk. "
          "Use only what is printed in the invoice text. If a text field is not printed, use an empty string. "
          "Never guess a purchase order number. Return JSON that matches the schema.\n\nFields:\n"
          + "\n".join(f"- {name}: {text}" for name, text in DESCRIPTIONS.items()))


# ---------- 1. read: PDF -> text ----------

def read_text(pdf: Path) -> str:
    try:
        from pypdf import PdfReader
    except ImportError:
        sys.exit("pypdf is not installed. Run: pip install -r requirements.txt")
    return "\n".join(page.extract_text() or "" for page in PdfReader(pdf).pages)


def to_number(value):
    """Turn 1457.75, '1.457,75', '$2,650.32' or '285.60' into a float; None if it isn't a number."""
    if isinstance(value, (int, float)):
        return float(value)
    if not isinstance(value, str):
        return None
    s = re.sub(r"[^\d,.\-]", "", value)
    if "," in s and (s.rfind(",") > s.rfind(".")):      # German style: 1.457,75
        s = s.replace(".", "").replace(",", ".")
    else:                                               # English style: 2,650.32
        s = s.replace(",", "")
    try:
        return float(s)
    except ValueError:
        return None


# ---------- 2a. extract with rules (a "template" for one layout) ----------

def extract_rules(text: str) -> dict:
    """Regular expressions written for supplier layout A only, like a classic extraction template."""
    def find(pattern):
        m = re.search(pattern, text)
        return m.group(1) if m else None
    lines = [{"description": m.group(1), "quantity": float(m.group(2)), "unit_price": float(m.group(3))}
             for m in re.finditer(r"^(.+?) (\d+) (\d+\.\d{2}) \d+\.\d{2}$", text, re.MULTILINE)]
    return {
        "invoice_number": find(r"Invoice No\.:\s*(\S+)"),
        "invoice_date": find(r"Invoice date:\s*(\d{4}-\d{2}-\d{2})"),
        "supplier_name": text.strip().splitlines()[0] if text.strip() else None,
        "po_number": find(r"PO number:\s*(\d+)"),
        "currency": find(r"Currency:\s*([A-Z]{3})"),
        "net_amount": to_number(find(r"Net amount:\s*([\d.,]+)")),
        "tax_amount": to_number(find(r"VAT \d+%:\s*([\d.,]+)")),
        "gross_amount": to_number(find(r"Total due:\s*([\d.,]+)")),
        "lines": lines,
    }


# ---------- 2b. extract with a local model ----------

def make_client():
    try:
        import ollama
    except ImportError:
        sys.exit("The ollama library isn't installed. Run: pip install -r requirements.txt")
    return ollama.Client()      # talks to Ollama on this computer (http://localhost:11434)


def extract_llm(text: str, client, model: str) -> dict:
    import ollama
    try:
        response = client.chat(
            model=model,
            messages=[{"role": "system", "content": SYSTEM},
                      {"role": "user", "content": f"Invoice text:\n{text}\n\nReturn the fields as JSON."}],
            format=SCHEMA,                                # the answer must match this JSON schema
            think=False,                                  # answer straight away
            options={"temperature": 0, "num_ctx": 4096},
        )
    except ConnectionError:
        sys.exit("Can't reach Ollama. Start the Ollama app (or run: ollama serve), then try again.")
    except ollama.ResponseError as err:
        if err.status_code == 404:
            sys.exit(f"Ollama doesn't have {model}. Run: ollama pull {model}")
        raise
    try:
        data = json.loads(response.message.content)
    except json.JSONDecodeError:
        data = {}
    if not isinstance(data, dict):
        data = {}
    for name in AMOUNTS:                                  # tidy numbers, whatever the model sent
        data[name] = to_number(data.get(name))
    data["lines"] = [{"description": ln.get("description"), "quantity": to_number(ln.get("quantity")),
                      "unit_price": to_number(ln.get("unit_price"))}
                     for ln in (data.get("lines") or []) if isinstance(ln, dict)]
    for name in ("invoice_number", "invoice_date", "supplier_name", "po_number", "currency"):
        value = data.get(name)
        data[name] = str(value).strip() if str(value).strip().lower() not in ("", "none", "null") else None
    return data


# ---------- 3. check: decide "auto" or "review" ----------

def close(a, b) -> bool:
    return a is not None and b is not None and abs(a - b) < 0.005


def checks(doc: dict, text: str, pos: dict) -> list:
    """Our own checks. Any reason returned sends the invoice to a person."""
    reasons = []
    for name in ("invoice_number", "invoice_date", "supplier_name", "currency", "gross_amount"):
        if doc.get(name) in (None, ""):
            reasons.append(f"missing {name}")
    if not doc.get("lines"):
        reasons.append("no lines")
    if doc.get("invoice_number") and doc["invoice_number"] not in text:
        reasons.append("invoice number not found in the document")
    try:
        date.fromisoformat(doc.get("invoice_date") or "")
    except ValueError:
        reasons.append("date is not YYYY-MM-DD")
    printed = {round(n, 2) for n in map(to_number, re.findall(r"\d[\d.,]*", text)) if n is not None}
    for name in ("tax_amount", "gross_amount"):
        if doc.get(name) is not None and round(doc[name], 2) not in printed:
            reasons.append(f"{name} not found in the document")
    if not close((doc.get("net_amount") or 0) + (doc.get("tax_amount") or 0), doc.get("gross_amount")):
        reasons.append("net + tax does not equal gross")
    line_sum = sum((ln["quantity"] or 0) * (ln["unit_price"] or 0) for ln in doc.get("lines") or [])
    if doc.get("lines") and not close(round(line_sum, 2), doc.get("net_amount")):
        reasons.append("lines do not add up to net")
    po = doc.get("po_number")
    if po is None:
        reasons.append("no purchase order: needs an approver")
    elif po not in text:
        reasons.append("PO number not found in the document")
    elif po not in pos:
        reasons.append("PO number unknown")
    elif pos[po]["supplier"].casefold() != (doc.get("supplier_name") or "").casefold():
        reasons.append("supplier differs from the PO")
    return reasons


def three_way(doc: dict, pos: dict) -> str:
    """Compare invoice lines with the PO price and the goods receipt quantity, line by line."""
    po = pos.get(doc.get("po_number") or "")
    if not po:
        return "n/a"
    issues = []
    for i, (inv, ordered) in enumerate(zip(doc["lines"], po["lines"]), 1):
        if not close(inv["unit_price"], ordered["unit_price"]):
            issues.append(f"line {i} price {inv['unit_price']} vs PO {ordered['unit_price']}")
        if (inv["quantity"] or 0) > ordered["received"]:
            issues.append(f"line {i} qty {inv['quantity'] or 0:g} vs received {ordered['received']}")
    if len(doc["lines"]) != len(po["lines"]):
        issues.append("line count differs from PO")
    return "match" if not issues else "exception: " + "; ".join(issues)


# ---------- 4. score against the right answers ----------

def correct(name: str, got, want) -> bool:
    if name in AMOUNTS:
        return close(got, want)
    if name == "lines":
        return len(got or []) == len(want) and all(
            close(g["quantity"], w["quantity"]) and close(g["unit_price"], w["unit_price"])
            for g, w in zip(got, want))
    if got is None or want is None:
        return got is None and want is None
    return str(got).strip().casefold() == str(want).strip().casefold()


def run(method: str, truth: list, pos: dict, model: str, client) -> list:
    rows = []
    print(f"\n=== Method: {method}" + (f" ({model})" if method == "llm" else "") + " ===")
    print(f"{'File':<12}{'Fields right':>13}  {'Route':<7} {'Three-way match / reasons'}")
    for t in truth:
        text = read_text(FOLDER / t["file"])
        started = time.perf_counter()
        doc = extract_rules(text) if method == "rules" else extract_llm(text, client, model)
        seconds = time.perf_counter() - started
        right = {name: correct(name, doc.get(name), t[name]) for name in FIELDS}
        reasons = checks(doc, text, pos)
        route = "review" if reasons else "auto"
        match = three_way(doc, pos) if route == "auto" else "; ".join(reasons[:2]) + (
            f" (+{len(reasons) - 2} more)" if len(reasons) > 2 else "")
        print(f"{t['file']:<12}{sum(right.values()):>9} of 9  {route:<7} {match}")
        rows.append({"method": method, "file": t["file"], "layout": t["layout"], "seconds": round(seconds, 2),
                     "fields_right": sum(right.values()), "route": route,
                     "detail": three_way(doc, pos) if route == "auto" else "; ".join(reasons),
                     **{f"ok_{name}": int(ok) for name, ok in right.items()}})
    summarize(rows)
    return rows


def summarize(rows: list) -> None:
    n = len(rows)
    print("\nField accuracy:  " + "  ".join(
        f"{name.replace('_amount', '').replace('invoice_', 'inv_')} {sum(r[f'ok_{name}'] for r in rows)}/{n}"
        for name in FIELDS))
    all_right = sum(r["fields_right"] == len(FIELDS) for r in rows)
    auto = [r for r in rows if r["route"] == "auto"]
    silent = sum(r["fields_right"] < len(FIELDS) for r in auto)
    print(f"Invoices with every field right: {all_right} of {n}")
    print(f"Sent straight through (auto):    {len(auto)} of {n}")
    print(f"Silent errors (auto but wrong):  {silent}   <- the number that hurts in production")
    print(f"Average seconds per invoice:     {sum(r['seconds'] for r in rows) / n:.2f}")


def show(file: str, method: str, pos: dict, model: str, client) -> None:
    if not (FOLDER / file).exists():
        sys.exit(f"No file {FOLDER / file}. Use a name like INV-03.pdf")
    text = read_text(FOLDER / file)
    print(f"--- Text pypdf read from {file} ---\n{text}\n")
    for m in (["rules", "llm"] if method == "both" else [method]):
        doc = extract_rules(text) if m == "rules" else extract_llm(text, client, model)
        print(f"--- Extracted with {m} ---\n{json.dumps(doc, indent=2)}")
        print(f"Checks: {checks(doc, text, pos) or 'all passed'}\n")


def main() -> None:
    parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument("--method", choices=["rules", "llm", "both"], default="rules")
    parser.add_argument("--model", default="qwen3:0.6b", help="an Ollama model you have pulled")
    parser.add_argument("--show", metavar="FILE", help="print the text and extracted fields for one invoice")
    args = parser.parse_args()

    if not (FOLDER / "truth.json").exists():
        sys.exit("No invoices yet. Run first: python unit12/make_invoices.py")
    truth = json.loads((FOLDER / "truth.json").read_text(encoding="utf-8"))
    pos = json.loads(POS_FILE.read_text(encoding="utf-8"))
    client = make_client() if args.method in ("llm", "both") else None

    if args.show:
        show(args.show, args.method, pos, args.model, client)
        return
    rows = []
    for method in (["rules", "llm"] if args.method == "both" else [args.method]):
        rows += run(method, truth, pos, args.model, client)

    with OUT_CSV.open("w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=list(rows[0]))
        writer.writeheader()
        writer.writerows(rows)
    print(f"\nSaved {len(rows)} rows to {OUT_CSV}")


if __name__ == "__main__":
    main()
  1. Look at one German invoice the way the script sees it:

    python unit12/extract_invoices.py --show INV-03.pdf

You should see the text pypdf read, then what the rules found:

--- Text pypdf read from INV-03.pdf ---
Kessler Antriebstechnik GmbH
Industriestr. 12, 70565 Stuttgart
Rechnung / Invoice
Rechnung Nr. RE-2026-0815
Datum: 14.09.2026
Ihre Bestellung / Your order: 4500000103
Pos. Bezeichnung Menge Einzelpreis Gesamt
10 Getriebemotor GM-40 2 St 612,50 1.225,00
Netto 1.225,00
MwSt 19 % 232,75
Gesamtbetrag EUR 1.457,75
Zahlbar innerhalb 30 Tagen ohne Abzug.

--- Extracted with rules ---
{
  "invoice_number": null,
  "invoice_date": null,
  "supplier_name": "Kessler Antriebstechnik GmbH",
  ...

Notice two things. The text has no table: the columns are just words in a row, because a PDF only places text at positions. And the rules found almost nothing, because they look for English labels from layout A.

Step 5: Run the rules (the template approach)

python unit12/extract_invoices.py

You should see:

=== Method: rules ===
File         Fields right  Route   Three-way match / reasons
INV-01.pdf          9 of 9  auto    match
INV-02.pdf          9 of 9  auto    match
INV-03.pdf          1 of 9  review  missing invoice_number; missing invoice_date (+6 more)
INV-04.pdf          1 of 9  review  missing invoice_number; missing invoice_date (+6 more)
INV-05.pdf          1 of 9  review  missing invoice_number; missing invoice_date (+6 more)
INV-06.pdf          1 of 9  review  missing invoice_number; missing invoice_date (+6 more)
INV-07.pdf          2 of 9  review  missing invoice_number; missing invoice_date (+6 more)
INV-08.pdf          1 of 9  review  missing invoice_number; missing invoice_date (+6 more)

Field accuracy:  inv_number 2/8  inv_date 2/8  supplier_name 8/8  po_number 3/8  currency 2/8  net 2/8  tax 2/8  gross 2/8  lines 2/8
Invoices with every field right: 2 of 8
Sent straight through (auto):    2 of 8
Silent errors (auto but wrong):  0   <- the number that hurts in production
Average seconds per invoice:     0.00

Saved 8 rows to unit12/extraction_results.csv

This is the template story in one table. Perfect on the layout it was written for, useless on the others. But look at the silent errors: zero. The rules fail loudly. Every miss becomes a missing field, and the checks send it to review. INV-07 scores 2 because its right answer for the purchase order is "none", and the rules found none.

Step 6: Run the local model

The model reads each invoice's text and fills the JSON schema. The descriptions in DESCRIPTIONS are its instructions, one per field.

  1. Make sure Ollama is running. On Windows and macOS, start the Ollama app. On Linux, it usually runs as a service already. Check:

    ollama list

    You should see qwen3:0.6b in the list. If not, run ollama pull qwen3:0.6b.

  2. Run both methods side by side:

    python unit12/extract_invoices.py --method both

The rules table comes first, the same as Step 5. Then a second table, === Method: llm (qwen3:0.6b) ===, in the same format. The first invoice is slower while the model loads.

Your numbers depend on the model and your computer, so we don't print ours here. Read your table with these questions:

  • Field accuracy: which fields does the model get wrong? Dates and amounts in German format are good places to look.
  • Straight through: how many invoices passed every check?
  • Silent errors: is it zero? If not, run --show on that file (Step 8) and find which wrong value passed every check.
  • Three-way match: if INV-04 and INV-06 went through, do they show exception:? Those are the planted business differences.

Step 7: Try a bigger model (optional)

A larger model usually reads better and runs slower. The Unit 12 setup lists qwen3:1.7b at about 1.4 GB.

ollama pull qwen3:1.7b
python unit12/extract_invoices.py --method llm --model qwen3:1.7b

Compare field accuracy, silent errors and seconds per invoice with the 0.6b run. Note that each run overwrites extraction_results.csv.

Step 8: Read the results

  1. Look closely at any invoice the model got wrong. For example:

    python unit12/extract_invoices.py --method llm --show INV-07.pdf

    You see the text, the model's JSON and the checks. INV-07 is the trap: it has no purchase order, only "framework agreement FA-12". A model that writes FA-12 into po_number has guessed. The check PO number unknown catches it, because FA-12 is not in open_pos.json.

  2. Open unit12/extraction_results.csv in Excel or VS Code. Each row is one invoice and one method, with a 1 or 0 for each field and the full list of check results in detail.

  3. Ask three questions of your results:

    • Which checks caught the errors? Which errors, if any, passed every check?
    • Which review reasons were extraction stops, and which were business stops (no purchase order: needs an approver)?
    • If INV-04 and INV-06 reached the three-way match, the exception: text is the process working, not the model failing.

Step 9: Save your work in Git

git add unit12/make_invoices.py unit12/extract_invoices.py unit12/invoices unit12/open_pos.json unit12/extraction_results.csv
git commit -m "Unit 12: invoice extraction with checks"

The invoices are made up and small, so they can live in Git as your test set. Never commit real supplier invoices.

How the code works

Part of the script What it does
TRUTH, OPEN_POS in make_invoices.py The right answers and the made-up purchase orders with goods receipts
page_items() Places text for each of the four layouts, with different labels, dates and number formats
write_pdf() Writes a one-page PDF by hand, so no extra library is needed
read_text() Reads the PDF's text layer with pypdf
to_number() Turns 1.457,75, $2,650.32 or 285.60 into one number
extract_rules() Regular expressions for layout A only: the template approach
SCHEMA, DESCRIPTIONS, SYSTEM The fixed form, and the field descriptions the model reads as instructions
extract_llm() Sends the text to Ollama with format=SCHEMA, think=False and temperature 0, then tidies the values
checks() Independent checks: required fields, values printed in the document, arithmetic, known purchase order and matching supplier
three_way() Compares each invoice line with the PO price and the received quantity
correct(), summarize() Scores against truth.json and prints accuracy, straight-through rate and silent errors

If something goes wrong

What you see What it means What to do
python is not recognized, or command not found Python isn't on your path, or the terminal is in the wrong place Reopen the course folder in VS Code; on macOS or Linux try python3; see Set up your computer
pypdf is not installed or The ollama library isn't installed The library is missing in the Python you're using Check for (.venv) in the prompt, then pip install -r requirements.txt (Step 2)
No invoices yet. Run first: python unit12/make_invoices.py Step 3 hasn't run, or you ran from inside unit12 Run the commands from the course folder
Can't reach Ollama Ollama isn't running Start the Ollama app, or run ollama serve in a second terminal
Ollama doesn't have qwen3:0.6b The model isn't downloaded ollama pull qwen3:0.6b
ollama pull fails or hangs at work A proxy or firewall blocks the download Ask IT; the Unit 12 setup covers Ollama and proxies
Every LLM invoice shows missing ... The model returned something that isn't valid JSON Update Ollama to a current version (structured outputs need it), then try --model qwen3:1.7b
No file unit12/invoices/INV-9.pdf A typo in the --show file name Use the names from unit12/invoices, such as INV-03.pdf

The SAP way

As of October 2026, SAP's sources describe SAP Document AI as a service on SAP BTP, an embedded edition, and features inside SAP applications.

What the service does

From SAP's product page and its Architecture Center pages:

  • Intake by API, through a web interface, or automatically from sources such as email.
  • OCR with handwriting and barcode detection, text in more than 100 languages, more than 35 file formats.
  • Preconfigured content for invoices, purchase orders and delivery notes. The SAP tutorial names invoice, payment advice and purchase order as standard document types. Q1 2026 added a business partner document type for the embedded and premium editions.
  • Schemas that you define: header fields and line item fields, versioned.
  • Blended extraction: pretrained models plus large language models. The extraction pipeline calls SAP AI Core and the generative AI hub, with prompts managed in the Prompt Registry, per the reference architecture.
  • Confidence per field, from 0 to 100%. Q1 2026 added custom low, medium and high thresholds per field in schemas.
  • Review in a web interface, side by side with the document, and instant learning from corrections (Premium and Premium Plus, per the product FAQ).
  • Enrichment with master data, and outbound delivery to S/4HANA through SAP BTP's Connectivity and Destination services, or the Cloud Connector for on-premise systems.
  • Storage: documents kept seven days by default, configurable, per the product page, which also mentions an EU-only access option.

Schemas and the three extraction approaches

SAP's tutorial "Use Trial to Extract Information from Standard Documents with Generative AI" maps exactly onto the three approaches above. In Schema Configuration, each field has a setup type:

Setup type in SAP's schema What extracts the field
auto with a default extractor, such as grossAmount SAP's pretrained models
auto without a default extractor A large language model
manual A template

The field description is the prompt for the language model. That is the same job your DESCRIPTIONS dictionary does. A schema moves from DRAFT to ACTIVE when you activate it, and must be deactivated before you change it. Uploaded documents move from PENDING to DONE.

Calling the API

The service is reached through the REST "Document Information Extraction API" (an OData v2 API also exists). An older SAP TechEd exercise shows the pattern: get an OAuth token with the service key, POST a file to /document/jobs under the base path /document-information-extraction/v1, then GET /document/jobs/{id} until the job leaves PENDING. That exercise also names the statuses READY, FAILED, CONFIRMED and DONE.

"""SKETCH: send one invoice PDF to SAP Document AI and print the job result.

Not something to run now. It needs an SAP BTP subaccount with an SAP Document AI
instance and a service key. Option names follow an older SAP TechEd exercise:
check them in the API reference for your instance before use.
"""
import json
import os
import time

import requests
from dotenv import load_dotenv

load_dotenv()                                    # DOX_* values from the service key, kept in .env
BASE = os.environ["DOX_URL"].rstrip("/") + "/document-information-extraction/v1"
TOKEN_URL = os.environ["DOX_UAA_URL"].rstrip("/") + "/oauth/token"

token = requests.post(TOKEN_URL, data={"grant_type": "client_credentials"},
                      auth=(os.environ["DOX_CLIENT_ID"], os.environ["DOX_CLIENT_SECRET"]),
                      timeout=30).json()["access_token"]
headers = {"Authorization": f"Bearer {token}"}

# The schema you activated is selected with a further option: take its exact name
# from the API reference, don't guess it.
options = {"clientId": "default", "documentType": "invoice"}
with open("unit12/invoices/INV-03.pdf", "rb") as pdf:
    job = requests.post(f"{BASE}/document/jobs", headers=headers, timeout=60,
                        files={"file": ("INV-03.pdf", pdf, "application/pdf")},
                        data={"options": json.dumps(options)}).json()

result = {"status": "PENDING"}
while result.get("status") == "PENDING":
    time.sleep(5)
    result = requests.get(f"{BASE}/document/jobs/{job['id']}", headers=headers, timeout=30).json()
print(json.dumps(result, indent=2)[:3000])     # fields with values and confidence, if DONE

To compare SAP Document AI with your build, upload the eight made-up invoices into a trial schema, export the results, and score them with the same truth.json. Same test set, same metrics: that is a fair comparison.

Inside S/4HANA Cloud

SAP's Q1 2026 highlights list two generally available features in S/4HANA Cloud Public Edition that use SAP Document AI: sales order creation from PDF or image purchase orders, and payment advice processing. Where AI creates value in S/4HANA walks through the sales order flow, scope item 4X9: upload, extraction, a proposed sales order request, review of missing fields, simulate, create.

These run inside the standard app. You configure them; you don't call the API yourself.

Licensing notes

The product page lists Free, Base, Premium and Embedded plans, and the FAQ mentions Premium Plus. The embedded edition page says it is activated with AI Units. Commercial details sit in SAP Discovery Center and SAP Store. The AI business case explains AI Units and BTP credits. Get a quote for your volume before comparing costs.

Build vs. SAP

Situation Choose Why
S/4HANA Cloud standard feature covers the document Embedded SAP feature No build, review built into the app, SAP maintains it
Common business documents, scans, many languages, posting to SAP SAP Document AI OCR, schemas, review interface, enrichment and SAP connectivity included
New document type, data may go to SAP's service SAP Document AI with a custom schema and LLM fields Field descriptions instead of code; test on real samples
Data must stay on a device or site Own build with a local model Model and documents never leave your boundary
Need custom checks on SAP data before posting Either, plus your own CAP app or integration flow SAP's enrichment patterns guide routes custom checks to CAP or Integration Suite
One stable layout at very high volume A template, in either Cheap, fast and exact while the layout holds

Production concerns

  • Test set first. Collect 50 to 200 real documents per type, with right answers keyed by a person and approved by data protection. Re-run it for every model, prompt or schema change.
  • Per-field thresholds. Not every field matters equally. Set strict rules for amounts, bank details and the purchase order; looser rules for descriptions. SAP's guide mentions about 90% for critical fields; measure what your threshold actually catches.
  • Independent checks are cheap insurance. Values printed in the document, arithmetic, known purchase order for this supplier, duplicate invoice number. They catch errors that confidence scores miss.
  • Never guess the purchase order. A guessed PO can pass a three-way match against the wrong order. Treat "not printed" as a valid answer.
  • Security and authorizations. Documents carry bank details and personal data. Limit who can see the review queue, log access, and set retention to the minimum. Posting into S/4HANA uses a technical user or the reviewer's identity; give it only the authorizations for the posting it does. See agent permissions and SAP authorizations and data security and PII.
  • Prompt injection. A document is untrusted input. A line of text in an invoice can carry instructions. Keep the model's job to filling a schema, with no tools, and let your checks decide. See prompt injection and tool poisoning.
  • Human approval. A person approves anything that fails a check. Nothing the model produces changes SAP data without passing the checks and SAP's own posting validation.
  • Monitoring. Track straight-through rate and review reasons per supplier per week. A sudden drop for one supplier usually means a new layout.
  • Cost. For a build, count model time per page and reviewer minutes. For SAP Document AI, count plan or AI Unit consumption. Reviewer time is often the biggest line.
  • Clean core. Extraction runs beside SAP and posts through released APIs or standard apps. Nothing here modifies the SAP system.

Pitfalls

  • Scoring per field only. Report documents fully right and silent errors too.
  • Testing on one supplier. Your test set must cover your real mix of layouts, languages and scans.
  • Treating schema-valid JSON as correct. Structured output guarantees the shape, not the values.
  • Counting business exceptions as AI errors. Keep extraction stops and three-way match exceptions apart.
  • Trusting model confidence without measuring it. Check how often "high confidence" values are still wrong on your test set.
  • Forgetting scans. pypdf can't read images. Real inboxes hold scans and phone photos, which need OCR or a vision model; Ollama's structured outputs also work with vision models.
  • No owner for layout changes. Someone must watch the review reasons and update schemas or descriptions.

Exercise

Add a new supplier and see which approach copes. The result goes into your Unit 12 decision log and feeds the buy-vs-build topic later in this unit.

  1. Open unit12/make_invoices.py. Copy the two layout C entries in TRUTH and change them into a new supplier: a new supplier_name, new invoice numbers such as NS-1001 and NS-1002, files INV-09.pdf and INV-10.pdf, new purchase order numbers 4500000109 and 4500000110, and keep "layout": "C". Change the lines, then recalculate net_amount, tax_amount (8%, as layout C prints) and gross_amount so they add up.

  2. Add the two purchase orders to OPEN_POS, with the new supplier name, prices equal to the invoice and received quantities equal to the invoiced ones.

  3. Run python unit12/make_invoices.py, then python unit12/extract_invoices.py --method both. Note the results for the two new files.

  4. Change one field description in DESCRIPTIONS in extract_invoices.py, for example make invoice_date say "The date printed next to Date, Datum or issued, written as YYYY-MM-DD". Run the LLM method again and compare.

  5. Open unit12/model_log.md (create it if it doesn't exist) and add a section ## Document AI decision with a table: method, model, documents fully right, straight-through, silent errors, seconds per invoice. Add rows for rules, qwen3:0.6b and, if you ran it, qwen3:1.7b.

  6. Under the table, write two sentences: which approach would you put in front of the clerks, and which check caught the most errors?

  7. Commit your changes:

    git add unit12/make_invoices.py unit12/extract_invoices.py unit12/invoices unit12/open_pos.json unit12/model_log.md
    git commit -m "Unit 12: document AI decision"

Done when make_invoices.py writes 10 invoices, model_log.md has a ## Document AI decision table with your own measured numbers for at least two methods and a two-sentence choice, and both are committed.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1In the read, fill, check, route model, what is the main job of the checks?

    Answer: C. Structured output already fixes the shape of the answer. The checks test the values against the document, arithmetic and SAP data, and send anything doubtful to a person. The target is zero silent errors, not zero extraction errors.
  2. 2Your rules extractor scores 9 of 9 on layout A and 1 or 2 of 9 on the other layouts, with zero silent errors. Why zero?

    Answer: A. Templates fail loudly. A regex that doesn't match returns nothing, and a missing required field fails the checks. A language model fails differently: it may return a plausible but wrong value.
  3. 3What does Ollama's format=SCHEMA guarantee in extract_llm()?

    Answer: D. Structured outputs constrain the answer to the schema's structure. The values can still be wrong, which is why checks() exists.
  4. 4The model returns tax 196.23 and gross 2650.23 for an invoice printed with 196.32 and 2,650.32. Net plus tax still equals gross. Which check catches it?

    Answer: B. A consistent misread passes the arithmetic. Checking that tax and gross appear in the document's own numbers catches it, because 2650.23 is not printed anywhere.
  5. 5INV-07 says "Purchase order reference: none given (framework agreement FA-12)". The model puts FA-12 in po_number. What happens in the pipeline, and why is it right?

    Answer: C. The value is printed in the text, so the "found in the document" check passes. The known-PO check fails, because FA-12 isn't in the open orders. A guessed purchase order must never pass silently.
  6. 6In SAP Document AI's schema configuration, what does a field with setup type auto and no default extractor use?

    Answer: C. Per SAP's tutorial, auto with a default extractor uses pretrained models, auto without one uses a large language model, and manual uses a template. The field description acts as the prompt.
  7. 7SAP's enrichment guide describes per-field confidence with a threshold around 90% for critical fields. What should you still do before relying on it?

    Answer: D. Confidence is the system grading itself. Only your labelled documents show how many values above the threshold are still wrong, which is the silent error rate for your mix.
  8. 8A supplier's invoice is extracted perfectly but stops with "line 1 price 48.9 vs PO 45.0". What do you do?

    Answer: B. The reading is right; the supplier billed a different price. That is a business exception for a person. The agent never changes SAP data such as a purchase order on its own.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in