Orchestrate

Tool design for agents

Design the tools an agent calls so it picks the right one, reads small useful results, recovers from errors, and can only change data through a person.

Updated Oct 5, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An agent can only do what its tools let it do. A tool is an action your software offers the model, such as "explain why this order is blocked". The model never sees the code behind it. It sees three things: a name, a short description and a list of inputs. From those alone it decides which tool to call and what to pass.

So writing a tool is closer to writing a job description than to writing an API. Anthropic's engineering team calls tools a contract between predictable software and an unpredictable agent. SAP's Joule Studio course says the same about Joule agents: the quality of a tool's description directly decides how accurately the agent behaves.

Most first drafts mirror the system's API: one tool per endpoint, named after the endpoint, returning the whole record. This topic shows why that fails, and what a designed tool looks like instead: fewer tools, each named for a job, with results the model can read and errors that tell it how to recover.

Why it matters to the business

Take the running example: clerks asking an agent about blocked sales orders in order-to-cash. The same SAP data can be offered in two ways.

  • Mirrored: get_order, get_order_details, get_order_status, get_bp, get_credit, query and update_order. Each returns what the API returns: about 30 header fields, many of them codes like "B" or "01".
  • Designed: five tools, such as sales_order_get_block_summary, which returns the customer's name, the order value and the block reason in words, and release_request_create, which drafts a release for a person to approve.

The difference shows up in four places:

  • Right answers. If three tools sound alike, the model guesses between them. A wrong tool means a wrong or incomplete answer, delivered with confidence.
  • Running cost. Every tool description is sent on every model call, and every result is re-sent on every later step. In the lab, one mirrored "list the customer's orders" call returns about 5,700 characters; the designed search returns about 600.
  • Risk. update_order lets the model change any field of any order. A designed set offers no such tool: the model can only draft a request, and a person decides.
  • Reuse. Well-designed tools can be offered to other agents, to Joule, or through the Model Context Protocol (MCP). A mirror of the API needs a clever prompt everywhere it goes.

The decision for a leader: treat the tool set as a product with an owner, a review and a test, not as plumbing a developer wires up on the side.

How SAP does it

As of October 2026, tools show up in two places in SAP's AI stack.

  • SAP's generative AI hub (orchestration service). For custom-built agents, your developers define tools in code. SAP's Python SDK can build a tool from a Python function, using its signature and its documentation text as the description the model reads. Tools can also be written as a JSON Schema, with a strict option that makes the model's inputs follow it. Running the loop around the tools is your own code; the SDK documentation says it has no built-in abstraction for that.
  • Joule Studio. For Joule agents, SAP's learning course lists the tool types an agent can use: Joule skills (governed, reusable functions over SAP and non-SAP sources), document grounding, a calculator, a human-in-the-loop approval step, other Joule agents, and MCP servers. The course gives three principles that match this topic: only offer a validated capability, describe its purpose in plain business language, and write the agent's instructions as a tool usage guide for a new team member.

Either way, someone in your team writes tool names and descriptions, and that writing decides behavior. Joule Studio's product scope and editions are changing quickly; check the current terms before you plan around them (covered later in Unit 9).

Mirrored or designed: a side-by-side

Mirrored API tools Designed tools
How many One per endpoint; grows with the API A few, one per job a clerk does
Names get_order, get_bp, query sales_order_get_block_summary, customer_get_address_gaps
Descriptions "Get an order." What it does, when to use it, what it returns, what it can't do
Inputs id, filter, a free-form fields object sales_order (digits only), a fixed list of choices
Results Whole records, codes, internal IDs The fields the next decision needs, with names and texts
Long lists Everything at once A limit, a total, and a note on how to narrow
Errors 400, 404 "Digits only, for example '4711'. Remove the prefix and call again."
Changes to data update_order changes the order A draft request; a person approves
Safety labels None Marked read-only or not, so the client can ask for confirmation

In the lab, a crude test of "which tool would be picked first" over ten clerk requests scores the mirrored set 4 out of 10 and the designed set 10 out of 10. The test is deliberately simple, so the gap is the lesson, not the exact numbers.

Questions to ask

  • Which jobs, in the clerk's words, does each tool do? Is any job covered by two tools?
  • Who reviews tool names and descriptions, and is there a test that shows the agent picks the right tool?
  • What does each tool return? Could a clerk read it, or is it codes and internal IDs?
  • Which tools can change data? Does a person approve every change, and is that enforced in code rather than only in the prompt?
  • What happens when a tool fails? Does the agent get a message it can act on?
  • How many tools does the agent see at once, and how much of each model call do their descriptions take up?
  • When a tool's description changes, who re-runs the test before it goes live?

Common misconceptions

  • "Give the agent the whole API and it will figure it out." More, overlapping tools make choosing harder and cost more on every call. OpenAI's guide suggests fewer than 20 tools at a time as a soft limit.
  • "Descriptions are documentation; the code is what matters." For the model, the description is the only documentation. Anthropic's documentation calls detailed descriptions by far the most important factor in tool performance.
  • "Return everything, just in case." Large results crowd out the question and are re-sent on every step. Return what the next decision needs.
  • "A label saying read-only makes a tool safe." MCP's safety labels are hints a client can use to ask for confirmation. They don't stop a tool from doing harm; checks in code and authorizations do.
  • "A clever prompt can fix poor tools." Instructions help the model choose, but a tool that returns codes or fails with "404" still leaves it stuck.

Key terms

  • Tool: an action your software offers a model, described by a name, a description and an input schema.
  • Tool definition: the name, description and input schema the model reads. It is sent with every model call.
  • Mirrored tool: a tool that copies one API call one to one, with the API's names and results.
  • Consolidated tool: one tool that does a whole job, even if it makes several API calls inside.
  • Namespacing: starting tool names with the object or system, such as sales_order_ or customer_, so similar tools are easy to tell apart.
  • Strict mode: a setting that makes the model's inputs follow the schema exactly.
  • Tool annotations: MCP's labels for how a tool behaves, such as read-only or destructive. Hints, not enforcement.
  • Tool-selection test: a list of requests with the right tool for each, used to measure whether the model picks correctly.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1What does a model actually use to decide which tool to call?

    Answer: B. The model never sees the code. It chooses from the name, the description and the inputs, so those are what decide behavior. That is why SAP's Joule Studio course says description quality determines how accurately an agent behaves.
  2. 2A developer proposes one agent tool for every endpoint of the sales order API. What is the main concern?

    Answer: C. A mirrored set grows with the API and has overlapping tools, so the model guesses between them, and every definition is sent on every call. Designed sets offer a few tools, each named for one job.
  3. 3An agent's tool returns the whole sales order record, about 30 fields with codes. What is the better design?

    Answer: B. Large results crowd the model's context and are re-sent on every later step, and codes mean nothing to the model. Return what the next decision needs, in words a clerk would understand.
  4. 4The clerks want the agent to release blocked orders. Which tool design fits?

    Answer: D. A draft request lets the agent do the preparation while a person makes the decision. A general update tool lets the model change anything, and a prompt warning is not a control.
  5. 5A vendor says their MCP tools are safe because each is labeled read-only. What should you ask?

    Answer: B. MCP's labels are hints a client can use, for example to ask for confirmation, and clients must treat them as untrusted unless the server is trusted. Real protection comes from code checks and authorizations.
  6. 6A developer improves one tool's description before go-live. What should happen next?

    Answer: D. Descriptions decide which tool the model picks, so a small wording change can improve or break selection. Treat a description change like a code change and re-run the test.
Deep layer · 40 min read

Mental model: an interface for a reader who can't ask questions

A human developer who meets a new API reads the docs, tries a call, asks a colleague. A model gets one chance: the definitions in the request. Treat each part of a tool as a message to that reader:

  • The definition is a prompt. The name, description and parameter descriptions are text the model reads on every call. Anthropic's advice is to describe a tool the way you would to a new hire; OpenAI's version is the "intern test": could someone use it correctly with only what you wrote?
  • The result is context. Whatever a tool returns becomes part of every later model call in the loop. Each field is a cost, and each code is a puzzle.
  • The error is an instruction. When a call fails, the error text is the only guidance the model gets for its next attempt.

Anthropic's engineering post frames tools as a contract between deterministic systems and non-deterministic agents. Your code is the deterministic side, so it carries the guarantees: validation, authorizations, approval. The model is the side that chooses, so the definitions must make the right choice easy.

How it works

What the model sees, and what it doesn't

flowchart LR
  subgraph Sent[Sent to the model]
    N[Name] --- D[Description] --- P[Input schema]
  end
  subgraph Kept[Kept in your code]
    C[Implementation] --- V[Validation] --- A[Annotations<br/>and approval]
  end
  Sent --> M{Model picks<br/>tool + arguments}
  M --> Kept
  Kept --> R[Result or error<br/>back to the model]

The model sees only the left box. Your code keeps the right box. Every design decision below is about making the left box clear and the right box strict.

Decision 1: what one tool should do

Anthropic's post warns against tools that merely wrap existing endpoints. Its example: instead of list_contacts, which returns everyone, offer search_contacts. OpenAI's guide says the same in other words: combine functions that are always called in sequence, and don't make the model fill in arguments your code already knows.

For blocked orders, that means:

  • sales_order_get_block_summary reads the order header and the customer's name and the block text, because a clerk always wants them together.
  • customer_get_credit_exposure adds the order's value to the open items and applies the approval rule in code. The model no longer does arithmetic or interprets a policy from the prompt.

There is one place to stop consolidating: reads and writes. Claude's documentation suggests grouping related actions into one tool with an action parameter. That reduces selection errors, but for SAP data keep reading and changing in separate tools. A write tool needs its own authorization check, its own approval and its own audit line, and a separate tool makes those easy to enforce and easy to leave out of a read-only agent.

Decision 2: names

  • Use the pattern object, verb, detail: sales_order_get_block_summary, customer_get_address_gaps. Anthropic recommends namespacing by service or resource, so tools from different systems don't collide.
  • Stay inside the allowed characters. Claude's documentation gives the pattern ^[a-zA-Z0-9_-]{1,128}$; the lab's lint uses a stricter 64-character limit to keep names short.
  • Avoid near-twins. get_order, get_order_details and get_order_status force the model to guess the difference from one word.

Decision 3: descriptions

Claude's documentation recommends at least three to four sentences per tool and calls detailed descriptions the most important factor in tool performance. A useful order:

  1. What it does, in the clerk's words.
  2. What it returns, so the model knows whether the answer will be there.
  3. When to use it, and when not to, especially if another tool is close.
  4. Limits: read-only, maximum rows, what it can't see.

SAP's Joule Studio course asks for the same: the tool's purpose in plain business language. It also puts sequencing in the agent's instructions, with a worked example: first look up the customer ID, then pass it to the order list tool. In your own code you can often remove such a sequence by consolidating; where you can't, say it in both places.

Anthropic notes that even small refinements to descriptions can produce large improvements, which is why the description deserves a test like code does.

Decision 4: parameters

  • Unambiguous names. Anthropic's example is user_id instead of user. For SAP data, sales_order and customer say more than id.
  • Enums for closed choices. OpenAI's guide recommends enums and object structure to make invalid states impossible. The lab's detail input only accepts "concise" or "detailed".
  • Strict-ready schemas. OpenAI's strict mode needs additionalProperties: false, every field listed in required, and null for "not given". The lab's sales_order_search has "limit": {"type": ["integer", "null"]} for that reason.
  • Formats the schema can't express. A schema can say "string"; it can't say "digits only, no SO- prefix". Put that in the parameter description. Anthropic's advanced tool use post makes the same point with dates: a schema can't say which date format to use, and in Anthropic's tests, adding example inputs improved accuracy on complex parameters.
  • Don't ask for what you know. If your code knows the user's sales organization, set it in code. Every input the model fills is a chance to get it wrong.

Decision 5: results

Claude's documentation asks for "high-signal" results: the fields the model needs to reason about its next step, with stable identifiers. Anthropic's post adds three techniques:

  • Names over codes. Return "block_reason": "Blocked by the credit check." rather than "TotalCreditCheckStatus": "B". In SAP, what a block code means is configured per system, so translate it in code using that system's texts.
  • A detail switch. A response_format input with "concise" and "detailed" lets the model ask for more only when it needs it. The lab calls it detail.
  • Limits that steer. Paginate, filter or truncate with sensible defaults, and when you cut a list short, say so and say how to narrow it. Anthropic notes that Claude Code limits tool responses to 25,000 tokens by default. The lab's search returns 5 of 7 orders with a note: set blocked_only, or raise limit.

Decision 6: errors

The MCP specification separates two kinds of failure. Protocol errors (unknown tool, malformed request) go back as standard errors. Tool execution errors (an API failure, invalid input data, a business rule) go back inside the result with isError: true, so the model can read them. As Agents from first principles showed, an error the model can read lets it retry with corrected input, ask the user, or explain the limitation.

Write each error as an instruction: what was wrong, what is expected, an example, and which tool to try instead.

Mirrored Designed
{"error": "400 Bad Request"} sales_order must be digits only, for example '4711'. You sent "SO-4711". Remove any prefix or spaces and call again.
{"error": "404"} Sales order 4799 does not exist or you may not see it. Check the number with the user; to find a customer's orders, use sales_order_search.

The second designed error also avoids leaking whether an order exists to someone who isn't allowed to see it.

Decision 7: safety labels and writes

MCP defines four annotations. The MCP blog post on them gives their defaults:

Hint Question it answers Default if missing
readOnlyHint Does the tool change its environment? false (it may write)
destructiveHint If it writes, can the change be destructive rather than additive? true
idempotentHint Is calling it again with the same arguments safe? false
openWorldHint Does it reach an open world of external entities? true

Declare them explicitly, because the defaults assume the worst. A client can use them to skip confirmation for trusted read-only tools or to ask before destructive ones. But the specification says clients must treat annotations as untrusted unless they come from a trusted server, and the blog post is plain that hints are not enforcement. Your code still validates inputs and checks authorizations; the MCP specification lists both as things servers must do.

For SAP data, the lab's pattern is: no tool that changes a record. release_request_create drafts a request, is marked not read-only, not destructive and idempotent, says in its description that a person approves, and returns the same draft if called twice. How the approval is enforced and how the loop pauses for it is the subject of the next Unit 9 topics.

Decision 8: how many tools

OpenAI's guide suggests fewer than 20 functions at the start of a turn, as a soft limit to test against. Anthropic's post warns that too many or overlapping tools distract agents. Size matters too: Anthropic's advanced tool use post describes 58 tools from five servers taking about 55K tokens before the conversation starts, and reports that letting the model search for tools on demand improved selection accuracy in its tests. If an agent needs many tools, offer the few a task needs, or give it a way to look tools up.

Descriptions cost tokens. The designed set in the lab is about 850 tokens of definitions against about 300 for the mirrored one. That is a good trade, because a clear description saves wrong calls, but it is a cost on every call.

Measuring a tool set

You can't judge a tool set by reading it. Anthropic's post lists the metrics worth tracking when you run an agent over test tasks: accuracy, runtime, the number of tool calls, token use and tool errors. It also recommends a held-out set of tasks you don't tune against. The simplest place to start is a tool-selection test: requests a clerk would make, each with the right first tool. The lab builds one, and Building an evaluation harness shows how to grow it.

Build it yourself: lint, compare and test two tool sets

You will build unit09/tool_design.py, one script with three commands that compare a mirrored tool set with a designed one over the same made-up SAP data:

  • lint checks every tool definition against the rules in this topic, without a model.
  • respond calls tools from both sets with good and bad input, and shows what the model would read back.
  • select runs a ten-request tool-selection test, either with a crude word-matching stand-in (--sample) or with a real model through SAP's orchestration service (--model).
flowchart LR
  M[mirror set<br/>7 tools] --> L[lint]
  D[designed set<br/>5 tools] --> L
  M --> R[respond]
  D --> R
  M --> S[select<br/>10 requests]
  D --> S
  S --> SC[score per set]

Nothing in the script changes any data. The records are made up, shaped like SAP's sales order API.

Before you start: complete Set up your computer for this course and Set up for Unit 9, which creates the unit09 folder. For the real-model path only, also Set up for Unit 5 for the SAP AI Core keys and sap-ai-sdk-gen. This walkthrough doesn't repeat those steps.

What you need

  • Your course folder with its .venv. No new libraries.
  • About 40 minutes.
  • lint, respond and select --sample are free and need no account.
  • select --model makes 20 model calls (ten requests, two tool sets): a small per-request charge on SAP AI Core with the generative AI hub, using a model in your catalog that supports tool calling.

Step 1: Open your course folder and turn on the virtual environment

  1. Open VS Code, choose File > Open Folder, and open orchestrate-course.

  2. Open a terminal: Terminal > New Terminal.

  3. If the prompt doesn't start with (.venv), turn it on:

    • Windows (PowerShell):

      .venv\Scripts\Activate.ps1
    • macOS / Linux:

      source .venv/bin/activate
  4. Check that the unit09 folder exists (same command on every system):

    python -c "import pathlib; print(pathlib.Path('unit09').is_dir())"

    You should see True. If you see False, create it: right-click in VS Code's file list, choose New Folder and name it unit09.

Step 2: Save the script

  1. In VS Code's file list, right-click unit09, choose New File and name it tool_design.py.
  2. Paste the code below and save.
"""Unit 9: tool design for agents. Two tool sets over the same made-up SAP data, compared three ways.

Set "mirror":   tools that copy the API one to one, the way a first draft often looks.
Set "designed": the same data, shaped for a model to choose, call and read.

Commands (run from your course folder, with .venv turned on):
    python unit09/tool_design.py lint                        # check both sets against design rules (no account)
    python unit09/tool_design.py lint --set designed         # one set only, with every finding listed
    python unit09/tool_design.py respond                     # what each set returns, for good and bad input
    python unit09/tool_design.py select --sample             # tool-choice test with a crude word-matching stand-in
    python unit09/tool_design.py select --model MODEL_NAME   # the same test with a real model (SAP AI Core)

The real path reads the AICORE_ lines in .env (see "Set up for Unit 5").
Nothing in this file changes any data: release_request_create only drafts a request for a person.
"""
import argparse
import json
import os
import re
import sys

# ---------- made-up data, shaped like SAP's sales order API (A_SalesOrder in API_SALES_ORDER_SRV) ----------


def order(number, customer, amount, delivery_block="", credit_status="", requested="2026-10-12"):
    """A made-up order header with field names from A_SalesOrder. Code values are made up: what each
    code means is configured per system, so never assume a code means the same everywhere."""
    return {"__metadata": {"type": "API_SALES_ORDER_SRV.A_SalesOrderType"},
            "SalesOrder": number, "SalesOrderType": "OR", "SalesOrganization": "1010",
            "DistributionChannel": "10", "OrganizationDivision": "00", "SoldToParty": customer,
            "CreationDate": "2026-10-01", "CreatedByUser": "CB9980000046", "PurchaseOrderByCustomer": f"PO-{number}",
            "SalesOrderDate": "2026-10-01", "TotalNetAmount": amount, "TransactionCurrency": "EUR",
            "PricingDate": "2026-10-01", "RequestedDeliveryDate": requested, "ShippingCondition": "01",
            "IncotermsClassification": "EXW", "IncotermsTransferLocation": "Walldorf",
            "CustomerPaymentTerms": "0004", "HeaderBillingBlockReason": "", "DeliveryBlockReason": delivery_block,
            "OverallSDProcessStatus": "A", "TotalCreditCheckStatus": credit_status,
            "OverallTotalDeliveryStatus": "A", "OverallSDDocumentRejectionSts": "A"}


ORDERS = {o["SalesOrder"]: o for o in [
    order("4711", "10023", "1800.00", credit_status="B"),
    order("4723", "10051", "640.00", delivery_block="01"),
    order("4725", "10077", "3900.00", delivery_block="02"),
    order("4730", "10023", "520.00"), order("4731", "10023", "75.00"), order("4732", "10023", "1210.00"),
    order("4733", "10023", "300.00"), order("4734", "10023", "980.00"), order("4735", "10023", "60.00"),
]}
# Made-up texts a clerk would read for each blocked order. In a real system they come from the block
# reasons configured in that system, read in the user's language.
BLOCK_TEXT = {"4711": "Blocked by the credit check.",
              "4723": "Incomplete: delivery address data missing.",
              "4725": "Pricing: the customer disputes the price."}
# A made-up customer master and credit lookup (not an SAP API). An empty string means the value is missing.
CUSTOMERS = {
    "10023": {"name": "Nordhafen Tools GmbH", "street": "Hauptstrasse 5", "postal_code": "69190",
              "city": "Walldorf", "country": "DE", "credit_limit": 50000.00, "open_items": 50700.00},
    "10051": {"name": "Donau Retail AG", "street": "Ringstrasse 12", "postal_code": "", "city": "Vienna",
              "country": "AT", "credit_limit": 30000.00, "open_items": 4100.00},
    "10077": {"name": "Lac Leman Instruments SA", "street": "Rue de Rive 3", "postal_code": "1204",
              "city": "Geneva", "country": "CH", "credit_limit": 80000.00, "open_items": 12000.00},
}
RELEASE_REQUESTS = {}   # drafts for a person to approve; nothing here ever reaches SAP


def is_blocked(o: dict) -> bool:
    """The course's rule from Unit 1: an order counts as blocked if any block field has a value."""
    return any(o[f] for f in ("DeliveryBlockReason", "HeaderBillingBlockReason", "TotalCreditCheckStatus"))


# ---------- set 1: "mirror" tools, one per API call, as a first draft often looks ----------

def obj(props: dict, required: list) -> dict:
    return {"type": "object", "properties": props, "required": required}


MIRROR = [
    {"name": "get_order", "description": "Get an order.",
     "parameters": obj({"id": {"type": "string"}}, ["id"])},
    {"name": "get_order_details", "description": "Get order details.",
     "parameters": obj({"id": {"type": "string"}}, ["id"])},
    {"name": "get_order_status", "description": "Get order status.",
     "parameters": obj({"id": {"type": "string"}}, ["id"])},
    {"name": "get_bp", "description": "Get business partner.",
     "parameters": obj({"id": {"type": "string"}}, ["id"])},
    {"name": "get_credit", "description": "Get credit.",
     "parameters": obj({"id": {"type": "string"}}, ["id"])},
    {"name": "query", "description": "Query an entity set with a filter.",
     "parameters": obj({"entity": {"type": "string"}, "filter": {"type": "string"}}, ["entity"])},
    {"name": "update_order", "description": "Update an order.",
     "parameters": obj({"id": {"type": "string"}, "fields": {"type": "object"}}, ["id", "fields"])},
]


def run_mirror(name: str, args: dict) -> dict:
    """Mirror tools pass the API straight through: whole records, codes, and bare status errors."""
    key = str(args.get("id", ""))
    if name in ("get_order", "get_order_details", "get_order_status"):
        if not key.isdigit():
            return {"error": "400 Bad Request"}
        if key not in ORDERS:
            return {"error": "404"}
        record = ORDERS[key]
        if name == "get_order_status":
            return {k: v for k, v in record.items() if k.endswith(("Status", "Sts", "BlockReason"))}
        if name == "get_order_details":
            item = {"SalesOrderItem": "10", "Material": "TG12", "RequestedQuantity": "1",
                    "NetAmount": record["TotalNetAmount"]}
            return {**record, "to_Item": {"results": [item]}}
        return record
    if name in ("get_bp", "get_credit"):
        if key not in CUSTOMERS:
            return {"error": "404"}
        c = CUSTOMERS[key]
        if name == "get_credit":
            return {"limit": c["credit_limit"], "open": c["open_items"]}
        return {"BusinessPartner": key, **{k: v for k, v in c.items() if k not in ("credit_limit", "open_items")}}
    if name == "query":
        match = re.search(r"SoldToParty eq '(\d+)'", str(args.get("filter", "")))
        rows = [o for o in ORDERS.values() if not match or o["SoldToParty"] == match.group(1)]
        return {"d": {"results": rows}}
    if name == "update_order":
        return {"error": "403"}   # this lab never writes; a real mirror tool would have changed the order
    return {"error": "unknown tool"}


# ---------- set 2: "designed" tools: fewer, named for the job, with results a model can use ----------

READ = {"readOnlyHint": True, "openWorldHint": False}

DESIGNED = [
    {"name": "sales_order_get_block_summary",
     "description": "Explain why one sales order is blocked: credit, pricing, incomplete data or another reason. "
                    "Returns the customer, net value, whether the order is blocked and the block reason as text "
                    "a clerk would read. Use it first whenever the user "
                    "asks about a specific order number. It only reads; it changes nothing.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order", "detail"],
                    "properties": {
                        "sales_order": {"type": "string", "description": "Sales order number, digits only, "
                                                                          "for example '4711'."},
                        "detail": {"type": "string", "enum": ["concise", "detailed"],
                                   "description": "Use 'concise' unless you need dates, the customer's PO number "
                                                  "or the sales organization."}}},
     "annotations": READ},
    {"name": "customer_get_credit_exposure",
     "description": "Check a customer's credit exposure against the credit limit, and who must approve. "
                    "Exposure is open items plus the order's net value, computed for you. Use it when an order "
                    "is blocked by the credit check or the user asks how far over the limit a customer is. "
                    "Pass the order number when there is one so its value is included.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["customer", "sales_order"],
                    "properties": {
                        "customer": {"type": "string", "description": "Customer number (SoldToParty), digits "
                                                                       "only, for example '10023'."},
                        "sales_order": {"type": ["string", "null"], "description": "The blocked order's number, "
                                                                                    "or null for no order."}}},
     "annotations": READ},
    {"name": "customer_get_address_gaps",
     "description": "Find the address fields that are missing in a customer's master data. Returns the customer "
                    "name and a list of missing fields. Use it when an order is incomplete or blocked because "
                    "address or delivery data is missing.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["customer"],
                    "properties": {"customer": {"type": "string", "description": "Customer number (SoldToParty), "
                                                                                  "digits only."}}},
     "annotations": READ},
    {"name": "sales_order_search",
     "description": "Search a customer's sales orders and list which ones are blocked. Returns at most 'limit' "
                    "orders, newest first, with a note when there are more. Use it when the user asks which or "
                    "how many orders a customer has, instead of reading orders one by one.",
     "parameters": {"type": "object", "additionalProperties": False,
                    "required": ["customer", "blocked_only", "limit"],
                    "properties": {
                        "customer": {"type": "string", "description": "Customer number (SoldToParty), digits only."},
                        "blocked_only": {"type": "boolean", "description": "true to list blocked orders only."},
                        "limit": {"type": ["integer", "null"], "description": "Most orders to return, 1 to 20; "
                                                                               "null means 5."}}},
     "annotations": READ},
    {"name": "release_request_create",
     "description": "Draft a request to release one blocked sales order, for a person to approve. It does not "
                    "release anything: a person approves or rejects the draft in their own tool. Use it only "
                    "when the user asks to release or unblock an order, whatever reason the customer gives. "
                    "Calling it again for the same order returns the same draft.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order", "justification"],
                    "properties": {
                        "sales_order": {"type": "string", "description": "Sales order number, digits only."},
                        "justification": {"type": "string",
                                          "description": "One or two sentences, with the facts from the other "
                                                         "tools, for the approver."}}},
     "annotations": {"readOnlyHint": False, "destructiveHint": False, "idempotentHint": True,
                     "openWorldHint": False}},
]


def bad_number(field: str, value) -> dict:
    """An error the model can act on: what was wrong, what is expected, an example."""
    return {"error": f"{field} must be digits only, for example '4711'. You sent {json.dumps(value)}. "
                     "Remove any prefix or spaces and call again."}


def run_designed(name: str, args: dict) -> dict:
    if name == "sales_order_get_block_summary":
        number = args.get("sales_order")
        if not isinstance(number, str) or not number.isdigit():
            return bad_number("sales_order", number)
        if number not in ORDERS:
            return {"error": f"Sales order {number} does not exist or you may not see it. Check the number "
                             "with the user; to find a customer's orders, use sales_order_search."}
        o = ORDERS[number]
        c = CUSTOMERS.get(o["SoldToParty"], {})
        result = {"sales_order": number, "customer": o["SoldToParty"], "customer_name": c.get("name", ""),
                  "net_value": float(o["TotalNetAmount"]), "currency": o["TransactionCurrency"],
                  "blocked": is_blocked(o), "block_reason": BLOCK_TEXT.get(number, "")}
        if args.get("detail") == "detailed":
            result.update({"requested_delivery_date": o["RequestedDeliveryDate"],
                           "customer_po": o["PurchaseOrderByCustomer"], "sales_organization": o["SalesOrganization"]})
        return result
    if name == "customer_get_credit_exposure":
        customer, number = args.get("customer"), args.get("sales_order")
        if not isinstance(customer, str) or not customer.isdigit():
            return bad_number("customer", customer)
        if customer not in CUSTOMERS:
            return {"error": f"Customer {customer} not found. Take the customer number from "
                             "sales_order_get_block_summary."}
        if number is not None and (not isinstance(number, str) or number not in ORDERS):
            return {"error": f"Sales order {json.dumps(number)} not found. Pass null to check the customer only."}
        c = CUSTOMERS[customer]
        order_value = float(ORDERS[number]["TotalNetAmount"]) if number else 0.0
        exposure = c["open_items"] + order_value
        over = round((exposure / c["credit_limit"] - 1) * 100, 1)
        # The approval rule (made up for the course) lives here, in code, not in the prompt.
        approver = ("none needed" if exposure <= c["credit_limit"]
                    else "CREDIT_MANAGER" if exposure <= c["credit_limit"] * 1.05 else "HEAD_OF_FINANCE")
        return {"customer": customer, "customer_name": c["name"], "credit_limit": c["credit_limit"],
                "open_items": c["open_items"], "order_value": order_value, "exposure": exposure,
                "percent_over_limit": max(over, 0.0), "approver": approver, "currency": "EUR"}
    if name == "customer_get_address_gaps":
        customer = args.get("customer")
        if not isinstance(customer, str) or not customer.isdigit():
            return bad_number("customer", customer)
        if customer not in CUSTOMERS:
            return {"error": f"Customer {customer} not found."}
        c = CUSTOMERS[customer]
        missing = [f for f in ("street", "postal_code", "city", "country") if c[f] == ""]
        return {"customer": customer, "customer_name": c["name"], "missing_fields": missing,
                "complete": not missing}
    if name == "sales_order_search":
        customer, limit = args.get("customer"), args.get("limit")
        if not isinstance(customer, str) or not customer.isdigit():
            return bad_number("customer", customer)
        limit = 5 if limit is None else limit
        if not isinstance(limit, int) or not 1 <= limit <= 20:
            return {"error": f"limit must be a whole number from 1 to 20, or null for 5. You sent {limit!r}."}
        rows = [o for o in ORDERS.values() if o["SoldToParty"] == customer
                and (is_blocked(o) or not args.get("blocked_only"))]
        rows.sort(key=lambda o: o["SalesOrder"], reverse=True)
        shown = [{"sales_order": o["SalesOrder"], "net_value": float(o["TotalNetAmount"]),
                  "blocked": is_blocked(o), "block_reason": BLOCK_TEXT.get(o["SalesOrder"], "")}
                 for o in rows[:limit]]
        result = {"customer": customer, "total": len(rows), "showing": len(shown), "orders": shown}
        if len(rows) > limit:
            result["note"] = (f"{len(rows) - limit} more not shown. Set blocked_only to true, or raise limit "
                              "(up to 20), rather than reading orders one by one.")
        return result
    if name == "release_request_create":
        number, why = args.get("sales_order"), str(args.get("justification", "")).strip()
        if not isinstance(number, str) or not number.isdigit():
            return bad_number("sales_order", number)
        if number not in ORDERS or not is_blocked(ORDERS[number]):
            return {"error": f"Sales order {number} is not blocked, so there is nothing to release."}
        if len(why) < 20:
            return {"error": "justification is too short. Give the approver the facts, for example the "
                             "exposure and limit from customer_get_credit_exposure."}
        draft = RELEASE_REQUESTS.setdefault(number, {"request_id": f"RR-{number}", "sales_order": number,
                                                     "status": "waiting_for_approval", "justification": why})
        return {**draft, "message": "Draft created. A person must approve it; nothing has been released."}
    return {"error": f"Unknown tool {name!r}. Available: " + ", ".join(t["name"] for t in DESIGNED)}


SETS = {"mirror": (MIRROR, run_mirror), "designed": (DESIGNED, run_designed)}


def model_view(tool: dict) -> dict:
    """What the model is sent: name, description and parameters. Annotations stay with your code."""
    return {k: tool[k] for k in ("name", "description", "parameters")}


# ---------- lint: design rules you can check without a model ----------

GENERIC_PARAMS = {"id", "key", "data", "value", "fields", "filter", "entity", "user", "input", "query"}
WRITE_WORDS = ("update", "create", "delete", "release", "post", "change", "set", "approve")


def lint_tool(t: dict) -> list:
    """Return (rule, message) pairs for one tool."""
    out, name, desc = [], t["name"], t.get("description", "")
    params = t.get("parameters", {})
    props = params.get("properties", {})
    if not re.fullmatch(r"[a-zA-Z0-9_-]{1,64}", name):
        out.append(("name", "use letters, digits, _ or - only, at most 64 characters"))
    if "_" not in name or len(name.split("_")) < 3:
        out.append(("name", "name the object and the job, e.g. sales_order_get_block_summary"))
    sentences = len([s for s in re.split(r"[.!?]\s", desc + " ") if s.strip()])
    if sentences < 3:
        out.append(("description", f"{sentences} sentence(s); say what it does, when to use it, what it returns"))
    if not re.search(r"\bUse it\b", desc):
        out.append(("description", "no 'Use it when/for ...': the model can't tell when to pick it"))
    for p, spec in props.items():
        if not spec.get("description"):
            out.append(("parameter", f"'{p}' has no description"))
        if p in GENERIC_PARAMS:
            out.append(("parameter", f"'{p}' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'"))
        if spec.get("type") == "object" and "properties" not in spec:
            out.append(("parameter", f"'{p}' is a free-form object; list the allowed fields"))
    if params.get("additionalProperties") is not False:
        out.append(("strict", "set additionalProperties to false"))
    if set(params.get("required", [])) != set(props):
        out.append(("strict", "list every parameter in required; use null for 'not given'"))
    ann = t.get("annotations")
    if ann is None or "readOnlyHint" not in ann:
        out.append(("safety", "declare readOnlyHint; when it is missing, MCP's default says the tool may write"))
    writes = (ann or {}).get("readOnlyHint") is False or any(w in name.lower() for w in WRITE_WORDS)
    if writes and "approve" not in desc.lower():
        out.append(("safety", "a tool that writes must say a person approves, and your code must enforce it"))
    return out


def lint_set(tools: list) -> list:
    """Rules about the set as a whole."""
    out, names = [], [t["name"] for t in tools]
    if len(tools) > 20:
        out.append(("set", f"{len(tools)} tools; keep the set small (OpenAI suggests fewer than 20)"))
    for i, a in enumerate(names):
        for b in names[i + 1:]:
            ta, tb = set(a.split("_")), set(b.split("_"))
            if len(ta & tb) / len(ta | tb) >= 0.5 and len(ta & tb) >= 2:
                out.append(("set", f"'{a}' and '{b}' overlap; merge them or make the difference obvious"))
    size = len(json.dumps([model_view(t) for t in tools]))
    out.append(("info", f"definitions are about {size // 4} tokens, sent with every model call"))
    return out


def cmd_lint(args) -> None:
    for set_name in (["mirror", "designed"] if args.set == "both" else [args.set]):
        tools = SETS[set_name][0]
        per_tool = {t["name"]: lint_tool(t) for t in tools}
        set_findings = lint_set(tools)
        problems = sum(len(v) for v in per_tool.values()) + sum(1 for r, _ in set_findings if r != "info")
        print(f"== {set_name}: {len(tools)} tools, {problems} finding(s)")
        for name, findings in per_tool.items():
            mark = "ok " if not findings else f"{len(findings)}x "
            print(f"  {mark:4}{name}")
            if args.set != "both":
                for rule, msg in findings:
                    print(f"        [{rule}] {msg}")
        for rule, msg in set_findings:
            print(f"  [{rule}] {msg}")
        print()
    if args.set == "both":
        print("Add --set mirror or --set designed to see every finding.")


# ---------- respond: what the model would read back ----------

def show(set_name: str, tool: str, args: dict) -> None:
    result = SETS[set_name][1](tool, args)
    text = json.dumps(result)
    shown = text if len(text) <= 300 else text[:300] + f"... ({len(text) - 300} more characters)"
    print(f"  {tool}({json.dumps(args)})\n    -> {len(text)} characters: {shown}")


def cmd_respond(args) -> None:
    print("1. Read order 4711 (blocked by the credit check)")
    show("mirror", "get_order", {"id": "4711"})
    show("designed", "sales_order_get_block_summary", {"sales_order": "4711", "detail": "concise"})
    print("\n2. A slightly wrong order number")
    show("mirror", "get_order", {"id": "SO-4711"})
    show("designed", "sales_order_get_block_summary", {"sales_order": "SO-4711", "detail": "concise"})
    print("\n3. Is customer 10023 over the credit limit, with order 4711?")
    show("mirror", "get_credit", {"id": "10023"})
    show("designed", "customer_get_credit_exposure", {"customer": "10023", "sales_order": "4711"})
    print("\n4. Which orders does customer 10023 have?")
    show("mirror", "query", {"entity": "A_SalesOrder", "filter": "SoldToParty eq '10023'"})
    show("designed", "sales_order_search", {"customer": "10023", "blocked_only": False, "limit": None})
    print("\n5. Release order 4711, twice")
    show("mirror", "update_order", {"id": "4711", "fields": {"TotalCreditCheckStatus": ""}})
    why = "Exposure 52,500 EUR is 5% over the 50,000 EUR limit; customer pays on time."
    show("designed", "release_request_create", {"sales_order": "4711", "justification": why})
    show("designed", "release_request_create", {"sales_order": "4711", "justification": why})
    print("\nThe mirror set returns codes, whole records and bare status numbers. The designed set returns "
          "what the next decision needs, and errors that say how to recover.")


# ---------- select: does the model pick the right tool? ----------

# Each case: a request, and the tools that would be right in each set (None: the set has no right tool).
CASES = [
    ("Why is order 4711 blocked?",
     {"mirror": ["get_order", "get_order_status"], "designed": ["sales_order_get_block_summary"]}),
    ("How far over its credit limit is customer 10023?",
     {"mirror": ["get_credit"], "designed": ["customer_get_credit_exposure"]}),
    ("Which orders of customer 10051 are blocked?",
     {"mirror": ["query"], "designed": ["sales_order_search"]}),
    ("What address data is missing for customer 10051?",
     {"mirror": ["get_bp"], "designed": ["customer_get_address_gaps"]}),
    ("Please get order 4711 released, the customer always pays.",
     {"mirror": None, "designed": ["release_request_create"]}),
    ("Is order 4725 blocked for pricing or for credit?",
     {"mirror": ["get_order", "get_order_status"], "designed": ["sales_order_get_block_summary"]}),
    ("How many orders does customer 10023 have?",
     {"mirror": ["query"], "designed": ["sales_order_search"]}),
    ("Does customer 10023 need the head of finance to approve?",
     {"mirror": ["get_credit"], "designed": ["customer_get_credit_exposure"]}),
    ("Order 4723 is incomplete. What is missing for its customer?",
     {"mirror": ["get_order", "get_bp"], "designed": ["sales_order_get_block_summary", "customer_get_address_gaps"]}),
    ("Unblock order 4723 now that the postal code is fixed.",
     {"mirror": None, "designed": ["release_request_create"]}),
]

STOP = set("a an the is are of for to and or it its this that what how does do in on with by be i me my "
           "now please get has have any one only when whether there their them you your use".split())


def words(text: str) -> set:
    out = set()
    for w in re.findall(r"[a-z]+", text.lower()):
        if w in STOP:
            continue
        out.add(w[:5])   # compare first five letters, so "orders" matches "order" and "released" "release"
    return out


def sample_pick(question: str, tools: list) -> str:
    """A crude stand-in for a model: pick the tool whose name and description share the most words with
    the request. Ties go to the tool listed first. It knows nothing about meaning; it only shows that
    the words in a definition are what selection has to go on."""
    q = words(question)
    scores = [(len(q & words(t["name"].replace("_", " ") + " " + t["description"])), -i, t["name"])
              for i, t in enumerate(tools)]
    return max(scores)[2]


def real_picker(model: str, tools: list):
    """Ask a real model through SAP's orchestration service which tool it would call first."""
    from dotenv import load_dotenv
    load_dotenv()
    names = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
             "AICORE_RESOURCE_GROUP"]
    missing = [n for n in names if not os.environ.get(n)]
    if missing:
        sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5. "
                 "Or use --sample to try without an account.")
    from gen_ai_hub.orchestration_v2 import (FunctionObject, FunctionTool, LLMModelDetails, ModuleConfig,
                                             OrchestrationConfig, OrchestrationService,
                                             PromptTemplatingModuleConfig, SystemMessage, Template, UserMessage)
    strict = all(t["parameters"].get("additionalProperties") is False for t in tools)
    fts = [FunctionTool(function=FunctionObject(name=t["name"], description=t["description"],
                                                parameters=t["parameters"], strict=strict)) for t in tools]
    template = Template(template=[SystemMessage(content="You help SAP order-to-cash clerks with blocked sales "
                                                        "orders. Call the one tool that best fits the request."),
                                  UserMessage(content="{{?question}}")], tools=fts)
    config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
        prompt=template, model=LLMModelDetails(name=model, params={"temperature": 0}, timeout=60, max_retries=1))))
    service = OrchestrationService(config=config)

    def pick(question: str) -> str:
        message = service.run(placeholder_values={"question": question}).final_result.choices[0].message
        return message.tool_calls[0].function.name if message.tool_calls else "(no tool)"
    return pick, service


def cmd_select(args) -> None:
    print("Which tool is picked first for each request?" + (" [sample: word-matching stand-in]" if args.sample
                                                           else f" [model: {args.model}]") + "\n")
    totals = {}
    for set_name in ("mirror", "designed"):
        tools = SETS[set_name][0]
        service = None
        if args.sample:
            pick = lambda q, tools=tools: sample_pick(q, tools)
        else:
            pick, service = real_picker(args.model, [model_view(t) for t in tools])
        right = 0
        print(f"== {set_name}")
        try:
            for question, expected in CASES:
                ok_names = expected[set_name]
                try:
                    chosen = pick(question)
                except Exception as error:
                    sys.exit(f"The call failed: {type(error).__name__}: {str(error)[:400]}")
                good = ok_names is not None and chosen in ok_names
                right += good
                note = "" if ok_names is not None else "  (this set has no safe tool for it)"
                print(f"  {'right' if good else 'WRONG'}  {chosen:<33} <- {question}{note}")
        finally:
            if service is not None:
                service.close_http_connection()
        totals[set_name] = right
        print()
    for set_name, right in totals.items():
        print(f"{set_name:<9} {right}/{len(CASES)} right")
    if args.sample:
        print("\n[sample] A word-matching stand-in, not a model. Run with --model for a real result.")


def main() -> None:
    parser = argparse.ArgumentParser(description="Tool design for agents: lint, respond, select.")
    sub = parser.add_subparsers(dest="command", required=True)
    p = sub.add_parser("lint", help="check tool definitions against design rules")
    p.add_argument("--set", choices=["both", "mirror", "designed"], default="both")
    sub.add_parser("respond", help="compare what each set returns")
    p = sub.add_parser("select", help="test which tool is picked for ten requests")
    p.add_argument("--sample", action="store_true", help="use the word-matching stand-in (no account)")
    p.add_argument("--model", default="gpt-4o-mini", help="model name from your catalog")
    args = parser.parse_args()
    {"lint": cmd_lint, "respond": cmd_respond, "select": cmd_select}[args.command](args)


if __name__ == "__main__":
    main()

Step 3: Lint both tool sets

A linter is a program that checks code or definitions against rules without running them. Run:

python unit09/tool_design.py lint

You should see:

== mirror: 7 tools, 57 finding(s)
  7x  get_order
  6x  get_order_details
  6x  get_order_status
  7x  get_bp
  7x  get_credit
  10x query
  11x update_order
  [set] 'get_order' and 'get_order_details' overlap; merge them or make the difference obvious
  [set] 'get_order' and 'get_order_status' overlap; merge them or make the difference obvious
  [set] 'get_order_details' and 'get_order_status' overlap; merge them or make the difference obvious
  [info] definitions are about 296 tokens, sent with every model call

== designed: 5 tools, 0 finding(s)
  ok  sales_order_get_block_summary
  ok  customer_get_credit_exposure
  ok  customer_get_address_gaps
  ok  sales_order_search
  ok  release_request_create
  [info] definitions are about 853 tokens, sent with every model call

Add --set mirror or --set designed to see every finding.

Now see the findings for the mirrored set:

python unit09/tool_design.py lint --set mirror

The last tool shows why mirrored write tools are dangerous:

  11x update_order
        [name] name the object and the job, e.g. sales_order_get_block_summary
        [description] 1 sentence(s); say what it does, when to use it, what it returns
        [description] no 'Use it when/for ...': the model can't tell when to pick it
        [parameter] 'id' has no description
        [parameter] 'id' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'
        [parameter] 'fields' has no description
        [parameter] 'fields' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'
        [parameter] 'fields' is a free-form object; list the allowed fields
        [strict] set additionalProperties to false
        [safety] declare readOnlyHint; when it is missing, MCP's default says the tool may write
        [safety] a tool that writes must say a person approves, and your code must enforce it

A clean lint does not prove a tool set works. It catches the easy mistakes so the test in Step 5 can focus on behavior.

Step 4: Compare what each set returns

python unit09/tool_design.py respond

You will see five pairs. The first two look like this (results longer than 300 characters are cut, with a count of what was left out):

1. Read order 4711 (blocked by the credit check)
  get_order({"id": "4711"})
    -> 810 characters: {"__metadata": {"type": "API_SALES_ORDER_SRV.A_SalesOrderType"}, "SalesOrder": "4711", "SalesOrderType": "OR", "SalesOrganization": "1010", "DistributionChannel": "10", "OrganizationDivision": "00", "SoldToParty": "10023", "CreationDate": "2026-10-01", "CreatedByUser": "CB9980000046", "PurchaseOrder... (510 more characters)
  sales_order_get_block_summary({"sales_order": "4711", "detail": "concise"})
    -> 190 characters: {"sales_order": "4711", "customer": "10023", "customer_name": "Nordhafen Tools GmbH", "net_value": 1800.0, "currency": "EUR", "blocked": true, "block_reason": "Blocked by the credit check."}

2. A slightly wrong order number
  get_order({"id": "SO-4711"})
    -> 28 characters: {"error": "400 Bad Request"}
  sales_order_get_block_summary({"sales_order": "SO-4711", "detail": "concise"})
    -> 131 characters: {"error": "sales_order must be digits only, for example '4711'. You sent \"SO-4711\". Remove any prefix or spaces and call again."}

Look for three more things in the rest of the output:

  • Pair 3: the designed credit tool returns "exposure": 52500.0, "percent_over_limit": 5.0 and "approver": "CREDIT_MANAGER". Your code did the arithmetic and applied the rule; the mirrored tool returns two bare numbers.
  • Pair 4: the mirrored query returns about 5,700 characters for one customer; the designed search returns about 600, with "total": 7, "showing": 5 and a note on how to narrow.
  • Pair 5: calling release_request_create twice returns the same "request_id": "RR-4711". That is what idempotent means: a repeated call does no extra harm.

Step 5: Run the tool-selection test with the stand-in

python unit09/tool_design.py select --sample

You should see:

Which tool is picked first for each request? [sample: word-matching stand-in]

== mirror
  right  get_order                         <- Why is order 4711 blocked?
  right  get_credit                        <- How far over its credit limit is customer 10023?
  WRONG  get_order                         <- Which orders of customer 10051 are blocked?
  WRONG  get_order                         <- What address data is missing for customer 10051?
  WRONG  get_order                         <- Please get order 4711 released, the customer always pays.  (this set has no safe tool for it)
  right  get_order                         <- Is order 4725 blocked for pricing or for credit?
  WRONG  get_order                         <- How many orders does customer 10023 have?
  WRONG  get_order                         <- Does customer 10023 need the head of finance to approve?
  right  get_order                         <- Order 4723 is incomplete. What is missing for its customer?
  WRONG  get_order                         <- Unblock order 4723 now that the postal code is fixed.  (this set has no safe tool for it)

== designed
  right  sales_order_get_block_summary     <- Why is order 4711 blocked?
  ...
  right  release_request_create            <- Unblock order 4723 now that the postal code is fixed.

mirror    4/10 right
designed  10/10 right

[sample] A word-matching stand-in, not a model. Run with --model for a real result.

Note the two release requests. The mirrored set has no safe tool for them: its only candidate, update_order, would change the order directly. The test counts both as misses on purpose.

Step 6 (optional): Run the test with a real model

This step needs your SAP AI Core keys in .env, as set up in Unit 5. Replace MODEL_NAME with a model from your catalog that supports tool calling:

python unit09/tool_design.py select --model MODEL_NAME

The output has the same shape as Step 5, with [model: MODEL_NAME] in the header. Compare the two scores, and look at which requests each set got wrong. The mirrored set's two release requests always count as misses, because it has no safe tool for them. Run it twice: a model's choices can vary between runs, which is one reason to keep the test.

Step 7: Save your work in Git

git add unit09/tool_design.py
git commit -m "Unit 9: tool design lint, responses and selection test"

You should see a line like 1 file changed.

What each part of the script does

Part What it does
order(), ORDERS, CUSTOMERS, BLOCK_TEXT Made-up data. Order headers use the field names of SAP's A_SalesOrder; codes and texts are invented
MIRROR, run_mirror Seven tools copied from the API: short names, one-line descriptions, whole records, bare errors
DESIGNED, run_designed Five tools named for jobs, with strict schemas, annotations, small results and errors that say how to recover
bad_number One shared error message for malformed numbers, so every tool says the same thing
model_view Strips annotations before definitions go to the model: they are for your code and the client
lint_tool, lint_set The design rules as checks: names, descriptions, parameters, strict mode, safety, overlap, size
cmd_respond Calls both sets side by side and prints the size of each result
CASES Ten clerk requests with the right tools per set; None means the set has no safe tool
sample_pick The word-matching stand-in for --sample
real_picker Sends one request at a time with the tools through SAP's orchestration service and reads the first tool call

If something goes wrong

What you see What it means What to do
python: command not found or 'python' is not recognized Python isn't on your path, or the virtual environment is off Turn on .venv (Step 1). On macOS/Linux try python3. See Set up your computer
can't open file ... tool_design.py You're not in the course folder, or the file has another name Run pwd (macOS/Linux) or Get-Location (Windows); it must end in orchestrate-course
ModuleNotFoundError: No module named 'dotenv' or 'gen_ai_hub' Only for --model: a library is missing pip install -r requirements.txt from Unit 5's setup; --sample needs neither
Missing in .env: AICORE_... Keys aren't in .env, or .env isn't in the course folder Follow Set up for Unit 5, Step 5. Or use --sample
The call failed: ... 401 or 403 Wrong or expired key, or wrong resource group Recreate the service key and check AICORE_RESOURCE_GROUP
The call failed: ... 404 or a model error The model name isn't deployed in your orchestration setup, or doesn't support tool calling Use a model name from your catalog that supports tools
The call failed: ... ConnectionError or a timeout Network or proxy blocks the call Try another network, or ask IT to allow your SAP AI Core URL. Use --sample meanwhile
Every real answer is (no tool) The model answered in text instead of calling a tool Check that it supports tool calling; try another model

The SAP way

As of October 2026, this is how the design decisions map onto SAP's stack.

Tools in the orchestration service

SAP's Python SDK documentation (Orchestration Service V2) offers three ways to define a tool:

  • The @function_tool() decorator: the function's signature and docstring describe the tool to the model. That makes the docstring your description, so write it with the same care, three or four sentences, when to use it.
  • FunctionTool with a FunctionObject of name, description, parameters and strict, as the lab's real_picker does. FunctionTool.from_function(..., strict=True) builds one from a Python function.
  • A plain JSON Schema dictionary, useful when the tool runs somewhere else.

The model's choice arrives in response.final_result.choices[0].message.tool_calls; your code runs the tool and returns a ToolChatMessage with the tool_call_id. The SDK has no built-in agentic loop, as Agents from first principles showed, so every design decision here, including validation, limits and approval, sits in your code. Orchestration modules such as data masking and content filtering, from SAP Generative AI Hub and the orchestration service, apply to each call of the loop.

Tools in Joule Studio

SAP's Joule Studio course lists six tool types for custom Joule agents: Joule skills, document grounding, a calculator, human in the loop, other Joule agents, and MCP servers. Its three principles line up with this topic:

SAP's principle In this topic
It must be a validated capability Test a tool on its own before an agent sees it; run the selection test after
Define its purpose in plain business language Decision 3: what, returns, when, limits
Instructions as a tool usage guide for a new team member Selection strategy, sequencing, source priority and limits in the agent's instructions

The course's examples of good instructions name tools in the instructions: "use the GetOrderStatus tool" for delivery status, and a discount tool that must not be used for VIP customers, with a hand-over to a human sales manager. That last example is a business rule. Put it in the instructions so the agent can explain it, and also enforce it in the tool's code, as customer_get_credit_exposure does with the approval rule. Building Joule agents is covered later in Unit 9.

From SAP's API to a designed tool

The mirrored records in the lab use the real field list of A_SalesOrder from API_SALES_ORDER_SRV, as in SAP's own sample record: SalesOrder, SoldToParty, TotalNetAmount, DeliveryBlockReason, TotalCreditCheckStatus, OverallSDProcessStatus and more than 20 others. A designed tool calls the same released API, as in Calling your first SAP API, and then:

  • selects the fields the job needs, for example with OData $select;
  • turns codes into texts using your system's configuration, because block reasons and status values are configured per system;
  • adds what the clerk would look up next, such as the customer's name.

Wrapping SAP APIs as tools with the user's identity, authorizations and audit logging is covered later in Unit 9.

MCP servers

When you offer tools through MCP, the same rules hold, and annotations become part of the definition. The server in Set up for Unit 9 already marked its tools with ToolAnnotations(read_only_hint=True) and returned errors with ToolError, which the client receives as an error result. This topic used the 2025-06-18 version of the MCP specification's tools page; the protocol topic later in Unit 9 covers the current version.

Licensing

Tool definitions travel with each orchestration call, so longer definitions make every step of the loop a larger request in the generative AI hub; Set up for Unit 5 covers access and plans. For Joule Studio, access and pricing were changing as of October 2026; check SAP's current terms before planning.

Build vs. SAP

Need Build it yourself SAP
Define a tool JSON Schema in your code, any vendor's tool calling API FunctionTool, @function_tool() or a schema in the orchestration SDK
Descriptions Yours to write and test Yours to write too: docstrings in the SDK, the description form in Joule Studio
Validation and business rules Your tool code Your tool code, or the governed logic inside a Joule skill
Approval before changes A draft-and-approve tool and your own approval step Joule Studio's human-in-the-loop tool for Joule agents
Reuse across agents An MCP server MCP servers as a Joule Studio tool type
Testing tool selection A selection test like the lab's, in your harness Check what Joule Studio's own testing offers; keep your own cases either way

A rule of thumb: whichever runtime runs the agent, the tool set is your design. Neither SAP nor a framework can write a description that knows your clerks' jobs.

Production concerns

  • Security and SAP authorizations. A tool runs with some identity. Prefer the calling user's own SAP authorizations for read tools, so a tool can never return more than the user could see. Validate every argument in code, as the MCP specification requires of servers. Unit 11 covers agent permissions in depth.
  • Errors that don't leak. Helpful is not the same as revealing. "Does not exist or you may not see it" helps the model without telling an unauthorized user that an order exists.
  • Writes. No tool changes an SAP record without a person's approval, enforced in code. Prefer tools that create a request over tools that change the record, make them idempotent, and log who asked and what was proposed.
  • Annotations are not controls. Declare them so clients can ask for confirmation. Don't rely on them from servers you don't trust, and never instead of authorizations.
  • Prompt injection through results. Any text a tool returns, such as an order note, can carry instructions. Keep free text out of results unless the job needs it, and label it as data.
  • Evaluation and change control. A description change is a behavior change. Version tool definitions with the agent's instructions, and re-run the selection test before each release. Keep held-out cases you don't tune against.
  • Cost. Count definition tokens and result sizes per call. Offer only the tools a task needs, cap list results, and keep a detail switch for the rare case that needs more.
  • Operations. Log every tool call with its arguments, result size, duration, errors and tool-definition version. Rising error rates on one tool usually mean its description or its error messages need work.
  • Clean core. Tools call released SAP APIs from BTP. Code-to-text translation and business rules live in the tool layer, not in custom code inside S/4HANA.

Pitfalls

  • Mirroring the API. One tool per endpoint, named after it, is the most common first draft and the hardest for a model to use.
  • Near-twin tools. If two tools differ by one word, merge them or say in both descriptions when to use which.
  • One-line descriptions. "Get credit." gives the model nothing to choose with. Say what, returns, when and limits.
  • Generic parameter names. id, data and filter invite the wrong value. Name what goes in.
  • Raw codes in results. The model will guess what "B" means. Translate codes with your system's configured texts.
  • Unbounded lists. Always set a default limit, return the total, and say how to narrow.
  • Status-code errors. "404" teaches nothing. Say what to do next.
  • A general update tool. Any tool that can change any field of any record is too much power for a model.
  • Rules only in the prompt. Put rules in the instructions so the agent can explain them, and in code so they hold.
  • Untested descriptions. Change a description, re-run the test.

Exercise: add a pricing tool and test the set

You will add a sixth tool for pricing disputes, two test requests for it, and a short design note. The note feeds the next Unit 9 topics, where tools run in multi-step agents and over real SAP APIs.

  1. Open unit09/tool_design.py and find the line RELEASE_REQUESTS = {}.

  2. Just above it, add made-up price data for order 4725:

    PRICES = {"4725": {"sales_order": "4725", "list_price": 4300.00, "discount": 400.00, "net_price": 3900.00,
                       "currency": "EUR"}}
  3. Find {"name": "release_request_create", inside DESIGNED. Just above that line, at the same indentation, add the new tool:

        {"name": "sales_order_get_price_conditions",
         "description": "Read the price conditions of one sales order: list price, discount and net price. "
                        "Use it when an order is blocked for pricing or the customer disputes a price. "
                        "It only reads; it changes nothing.",
         "parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order"],
                        "properties": {"sales_order": {"type": "string",
                                                       "description": "Sales order number, digits only."}}},
         "annotations": READ},
  4. In run_designed, find the last line, return {"error": f"Unknown tool {name!r}. .... Just above it, at the same indentation as the other if name == lines, add:

        if name == "sales_order_get_price_conditions":
            number = args.get("sales_order")
            if not isinstance(number, str) or not number.isdigit():
                return bad_number("sales_order", number)
            return PRICES.get(number) or {"error": f"No price conditions found for sales order {number}."}
  5. In CASES, just before the closing ], add two requests:

        ("What discount did order 4725 get?",
         {"mirror": ["get_order_details"], "designed": ["sales_order_get_price_conditions"]}),
        ("The customer disputes the price on order 4725. What was agreed?",
         {"mirror": ["get_order_details"], "designed": ["sales_order_get_price_conditions"]}),
  6. Run the lint for the designed set:

    python unit09/tool_design.py lint --set designed

    You should see 6 tools, 0 finding(s) and the definitions growing to about 970 tokens. If the new tool has findings, fix its definition until it has none.

  7. Run the selection test:

    python unit09/tool_design.py select --sample

    You should see designed 12/12 right. If you have keys, also run it with --model MODEL_NAME.

  8. Now break it on purpose: change the new tool's description to "Get prices.", run steps 6 and 7 again, and note what changed. Then undo the change.

  9. In unit09, create tool_design_notes.md with these headings:

    • Scores: the mirror and designed scores, with --sample and, if you ran it, with a real model.
    • The broken description: what the lint and the test showed in step 8.
    • Results: for the price tool, which fields you return and which you leave out, and why.
    • Writes: which of the six tools may change anything, and how a person approves it.
    • Open questions: at least two, such as which SAP API the price tool would call and which identity it would use.
  10. Save your work:

    git add unit09/tool_design.py unit09/tool_design_notes.md
    git commit -m "Unit 9: pricing tool and tool design notes"

Done when: lint --set designed shows 6 tools, 0 finding(s), select --sample shows designed 12/12 right, and tool_design_notes.md covers scores, the broken description, results, writes and at least two open questions.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does Anthropic's guidance say to avoid tools that merely wrap API endpoints?

    Answer: B. One tool per endpoint gives the model near-twin choices and whole-record results. Consolidated tools named for a job, like search instead of list-everything, make the right choice obvious and keep results small.
  2. 2Claude's documentation suggests grouping related actions into one tool with an action parameter. Where does this topic stop consolidating for SAP data, and why?

    Answer: D. Fewer tools reduce selection errors, but a write needs its own authorization check, approval and audit line. Separate write tools are easy to enforce and easy to leave out of a read-only agent.
  3. 3In the lab, why does customer_get_credit_exposure compute the exposure and the approver in code?

    Answer: A. OpenAI's guide says to offload work to your code, and a rule that decides who approves must hold every time. The tool returns the exposure, the percentage and the approver, and the model explains them.
  4. 4A tool must accept "no limit given". How do you keep its schema ready for strict mode?

    Answer: B. OpenAI's strict mode needs every field in required and additionalProperties: false. Optional values are expressed as null, as in the lab's "type": ["integer", "null"].
  5. 5Your search tool finds 400 orders for a customer. What should it return?

    Answer: C. Anthropic recommends sensible defaults for pagination and truncation, and a note that steers the agent to narrower searches. The lab's search returns 5 of 7 with a note to set blocked_only or raise limit.
  6. 6An agent keeps calling get_order with "SO-4711" and getting "400 Bad Request". What is the best fix?

    Answer: D. The error text is the only guidance the model gets for its next try. Say what was wrong, what is expected and give an example, and put the format in the parameter description so the first call is right.
  7. 7An MCP server from a third party marks its delete_items tool with readOnlyHint: true. What follows?

    Answer: B. The MCP specification says clients must treat annotations as untrusted unless they come from trusted servers, and hints are not enforcement. Note also the default when the hint is missing is false, meaning the tool may write.
  8. 8The clerks ask the agent to release order 4711. Which tool design does this topic recommend?

    Answer: C. A draft request lets the agent prepare the case while a person decides, and idempotency means a repeated call creates no second request. A prompt warning is not a control, and an update tool still changes SAP data without approval.

Sources

  • Writing effective tools for AI agents, using AI agents (Anthropic Engineering, 11 September 2025) — tools as a contract between deterministic systems and non-deterministic agents; don't merely wrap API endpoints; consolidate (search_contacts instead of list_contacts); namespacing; meaningful names over cryptic identifiers; response_format concise or detailed; pagination, filtering, truncation with defaults; actionable errors; describe tools as to a new hire; unambiguous parameter names such as user_id; evaluation metrics and held-out test sets
  • Define tools (Claude Platform documentation) — descriptions of at least 3 to 4 sentences, the most important factor in tool performance; say what, when, parameters and caveats; group related actions with an action parameter; namespacing; name regex; return high-signal fields and stable identifiers; input_examples; strict mode
  • Introducing advanced tool use on the Claude Developer Platform (Anthropic Engineering) — 58 tools in five servers took about 55K tokens of definitions; tool search improved selection accuracy in Anthropic's tests; examples capture format conventions a schema can't express
  • Function calling (OpenAI API documentation) — describe purpose and each parameter; the intern test; enums and object structure to prevent invalid states; don't make the model fill arguments you already know; combine functions always called in sequence; aim for fewer than 20 functions at the start of a turn (soft); strict mode needs additionalProperties false, all fields required, null for optional
  • Tools (Model Context Protocol specification, version 2025-06-18) — tool fields name, title, description, inputSchema, outputSchema, annotations; annotations untrusted unless from trusted servers; protocol errors vs tool execution errors with isError; servers must validate inputs, apply access controls, rate limit and sanitize outputs; a human in the loop able to deny invocations
  • Tool annotations as risk vocabulary: what hints can and can't do (MCP blog, 16 March 2026) — readOnlyHint (default false), destructiveHint (default true), idempotentHint (default false), openWorldHint (default true); hints are not enforcement; clients treat hints from untrusted servers as untrusted
  • Orchestration Service V2 API (SAP Cloud SDK for AI, Python) — tools from the @function_tool() decorator (signature and docstring describe the tool), FunctionTool with FunctionObject (name, description, parameters, strict), or a JSON Schema dictionary; tool_calls in the response; ToolChatMessage; no built-in abstraction for the agentic loop
  • Leveraging tools and Model Context Protocol (MCP) for custom Joule agents (SAP Learning) — tool types Joule skills, document grounding, calculator, human in the loop, Joule agents and MCP; description quality determines agent accuracy; three principles (validated capability, purpose in plain business language, instructions as a tool usage guide for a new team member); tool selection, sequencing, data source priority and limitations
  • A_SalesOrder('69') test record (SAP-samples/s4hana-ext-geo-report-app on GitHub) — the full field list of a sales order header from API_SALES_ORDER_SRV, used for the lab's mirror records (DeliveryBlockReason, TotalCreditCheckStatus, OverallSDProcessStatus and others)

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in