Design the tools an agent calls so it picks the right one, reads small useful results, recovers from errors, and can only change data through a person.
An agent can only do what its tools let it do. A tool is an action your software offers the model, such as "explain why this order is blocked". The model never sees the code behind it. It sees three things: a name, a short description and a list of inputs. From those alone it decides which tool to call and what to pass.
So writing a tool is closer to writing a job description than to writing an API. Anthropic's engineering team calls tools a contract between predictable software and an unpredictable agent. SAP's Joule Studio course says the same about Joule agents: the quality of a tool's description directly decides how accurately the agent behaves.
Most first drafts mirror the system's API: one tool per endpoint, named after the endpoint, returning the whole record. This topic shows why that fails, and what a designed tool looks like instead: fewer tools, each named for a job, with results the model can read and errors that tell it how to recover.
Take the running example: clerks asking an agent about blocked sales orders in order-to-cash. The same SAP data can be offered in two ways.
Mirrored:get_order, get_order_details, get_order_status, get_bp, get_credit, query and update_order. Each returns what the API returns: about 30 header fields, many of them codes like "B" or "01".
Designed: five tools, such as sales_order_get_block_summary, which returns the customer's name, the order value and the block reason in words, and release_request_create, which drafts a release for a person to approve.
The difference shows up in four places:
Right answers. If three tools sound alike, the model guesses between them. A wrong tool means a wrong or incomplete answer, delivered with confidence.
Running cost. Every tool description is sent on every model call, and every result is re-sent on every later step. In the lab, one mirrored "list the customer's orders" call returns about 5,700 characters; the designed search returns about 600.
Risk.update_order lets the model change any field of any order. A designed set offers no such tool: the model can only draft a request, and a person decides.
Reuse. Well-designed tools can be offered to other agents, to Joule, or through the Model Context Protocol (MCP). A mirror of the API needs a clever prompt everywhere it goes.
The decision for a leader: treat the tool set as a product with an owner, a review and a test, not as plumbing a developer wires up on the side.
As of October 2026, tools show up in two places in SAP's AI stack.
SAP's generative AI hub (orchestration service). For custom-built agents, your developers define tools in code. SAP's Python SDK can build a tool from a Python function, using its signature and its documentation text as the description the model reads. Tools can also be written as a JSON Schema, with a strict option that makes the model's inputs follow it. Running the loop around the tools is your own code; the SDK documentation says it has no built-in abstraction for that.
Joule Studio. For Joule agents, SAP's learning course lists the tool types an agent can use: Joule skills (governed, reusable functions over SAP and non-SAP sources), document grounding, a calculator, a human-in-the-loop approval step, other Joule agents, and MCP servers. The course gives three principles that match this topic: only offer a validated capability, describe its purpose in plain business language, and write the agent's instructions as a tool usage guide for a new team member.
Either way, someone in your team writes tool names and descriptions, and that writing decides behavior. Joule Studio's product scope and editions are changing quickly; check the current terms before you plan around them (covered later in Unit 9).
What it does, when to use it, what it returns, what it can't do
Inputs
id, filter, a free-form fields object
sales_order (digits only), a fixed list of choices
Results
Whole records, codes, internal IDs
The fields the next decision needs, with names and texts
Long lists
Everything at once
A limit, a total, and a note on how to narrow
Errors
400, 404
"Digits only, for example '4711'. Remove the prefix and call again."
Changes to data
update_order changes the order
A draft request; a person approves
Safety labels
None
Marked read-only or not, so the client can ask for confirmation
In the lab, a crude test of "which tool would be picked first" over ten clerk requests scores the mirrored set 4 out of 10 and the designed set 10 out of 10. The test is deliberately simple, so the gap is the lesson, not the exact numbers.
"Give the agent the whole API and it will figure it out." More, overlapping tools make choosing harder and cost more on every call. OpenAI's guide suggests fewer than 20 tools at a time as a soft limit.
"Descriptions are documentation; the code is what matters." For the model, the description is the only documentation. Anthropic's documentation calls detailed descriptions by far the most important factor in tool performance.
"Return everything, just in case." Large results crowd out the question and are re-sent on every step. Return what the next decision needs.
"A label saying read-only makes a tool safe." MCP's safety labels are hints a client can use to ask for confirmation. They don't stop a tool from doing harm; checks in code and authorizations do.
"A clever prompt can fix poor tools." Instructions help the model choose, but a tool that returns codes or fails with "404" still leaves it stuck.
Pick one answer for each question. The explanation appears after you choose.
1What does a model actually use to decide which tool to call?
Answer: B. The model never sees the code. It chooses from the name, the description and the inputs, so those are what decide behavior. That is why SAP's Joule Studio course says description quality determines how accurately an agent behaves.
2A developer proposes one agent tool for every endpoint of the sales order API. What is the main concern?
Answer: C. A mirrored set grows with the API and has overlapping tools, so the model guesses between them, and every definition is sent on every call. Designed sets offer a few tools, each named for one job.
3An agent's tool returns the whole sales order record, about 30 fields with codes. What is the better design?
Answer: B. Large results crowd the model's context and are re-sent on every later step, and codes mean nothing to the model. Return what the next decision needs, in words a clerk would understand.
4The clerks want the agent to release blocked orders. Which tool design fits?
Answer: D. A draft request lets the agent do the preparation while a person makes the decision. A general update tool lets the model change anything, and a prompt warning is not a control.
5A vendor says their MCP tools are safe because each is labeled read-only. What should you ask?
Answer: B. MCP's labels are hints a client can use, for example to ask for confirmation, and clients must treat them as untrusted unless the server is trusted. Real protection comes from code checks and authorizations.
6A developer improves one tool's description before go-live. What should happen next?
Answer: D. Descriptions decide which tool the model picks, so a small wording change can improve or break selection. Treat a description change like a code change and re-run the test.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Deep layer · 40 min read
#Mental model: an interface for a reader who can't ask questions
A human developer who meets a new API reads the docs, tries a call, asks a colleague. A model gets one chance: the definitions in the request. Treat each part of a tool as a message to that reader:
The definition is a prompt. The name, description and parameter descriptions are text the model reads on every call. Anthropic's advice is to describe a tool the way you would to a new hire; OpenAI's version is the "intern test": could someone use it correctly with only what you wrote?
The result is context. Whatever a tool returns becomes part of every later model call in the loop. Each field is a cost, and each code is a puzzle.
The error is an instruction. When a call fails, the error text is the only guidance the model gets for its next attempt.
Anthropic's engineering post frames tools as a contract between deterministic systems and non-deterministic agents. Your code is the deterministic side, so it carries the guarantees: validation, authorizations, approval. The model is the side that chooses, so the definitions must make the right choice easy.
flowchart LR
subgraph Sent[Sent to the model]
N[Name] --- D[Description] --- P[Input schema]
end
subgraph Kept[Kept in your code]
C[Implementation] --- V[Validation] --- A[Annotations<br/>and approval]
end
Sent --> M{Model picks<br/>tool + arguments}
M --> Kept
Kept --> R[Result or error<br/>back to the model]
The model sees only the left box. Your code keeps the right box. Every design decision below is about making the left box clear and the right box strict.
Anthropic's post warns against tools that merely wrap existing endpoints. Its example: instead of list_contacts, which returns everyone, offer search_contacts. OpenAI's guide says the same in other words: combine functions that are always called in sequence, and don't make the model fill in arguments your code already knows.
For blocked orders, that means:
sales_order_get_block_summary reads the order header and the customer's name and the block text, because a clerk always wants them together.
customer_get_credit_exposure adds the order's value to the open items and applies the approval rule in code. The model no longer does arithmetic or interprets a policy from the prompt.
There is one place to stop consolidating: reads and writes. Claude's documentation suggests grouping related actions into one tool with an action parameter. That reduces selection errors, but for SAP data keep reading and changing in separate tools. A write tool needs its own authorization check, its own approval and its own audit line, and a separate tool makes those easy to enforce and easy to leave out of a read-only agent.
Use the pattern object, verb, detail: sales_order_get_block_summary, customer_get_address_gaps. Anthropic recommends namespacing by service or resource, so tools from different systems don't collide.
Stay inside the allowed characters. Claude's documentation gives the pattern ^[a-zA-Z0-9_-]{1,128}$; the lab's lint uses a stricter 64-character limit to keep names short.
Avoid near-twins. get_order, get_order_details and get_order_status force the model to guess the difference from one word.
Claude's documentation recommends at least three to four sentences per tool and calls detailed descriptions the most important factor in tool performance. A useful order:
What it does, in the clerk's words.
What it returns, so the model knows whether the answer will be there.
When to use it, and when not to, especially if another tool is close.
Limits: read-only, maximum rows, what it can't see.
SAP's Joule Studio course asks for the same: the tool's purpose in plain business language. It also puts sequencing in the agent's instructions, with a worked example: first look up the customer ID, then pass it to the order list tool. In your own code you can often remove such a sequence by consolidating; where you can't, say it in both places.
Anthropic notes that even small refinements to descriptions can produce large improvements, which is why the description deserves a test like code does.
Unambiguous names. Anthropic's example is user_id instead of user. For SAP data, sales_order and customer say more than id.
Enums for closed choices. OpenAI's guide recommends enums and object structure to make invalid states impossible. The lab's detail input only accepts "concise" or "detailed".
Strict-ready schemas. OpenAI's strict mode needs additionalProperties: false, every field listed in required, and null for "not given". The lab's sales_order_search has "limit": {"type": ["integer", "null"]} for that reason.
Formats the schema can't express. A schema can say "string"; it can't say "digits only, no SO- prefix". Put that in the parameter description. Anthropic's advanced tool use post makes the same point with dates: a schema can't say which date format to use, and in Anthropic's tests, adding example inputs improved accuracy on complex parameters.
Don't ask for what you know. If your code knows the user's sales organization, set it in code. Every input the model fills is a chance to get it wrong.
Claude's documentation asks for "high-signal" results: the fields the model needs to reason about its next step, with stable identifiers. Anthropic's post adds three techniques:
Names over codes. Return "block_reason": "Blocked by the credit check." rather than "TotalCreditCheckStatus": "B". In SAP, what a block code means is configured per system, so translate it in code using that system's texts.
A detail switch. A response_format input with "concise" and "detailed" lets the model ask for more only when it needs it. The lab calls it detail.
Limits that steer. Paginate, filter or truncate with sensible defaults, and when you cut a list short, say so and say how to narrow it. Anthropic notes that Claude Code limits tool responses to 25,000 tokens by default. The lab's search returns 5 of 7 orders with a note: set blocked_only, or raise limit.
The MCP specification separates two kinds of failure. Protocol errors (unknown tool, malformed request) go back as standard errors. Tool execution errors (an API failure, invalid input data, a business rule) go back inside the result with isError: true, so the model can read them. As Agents from first principles showed, an error the model can read lets it retry with corrected input, ask the user, or explain the limitation.
Write each error as an instruction: what was wrong, what is expected, an example, and which tool to try instead.
Mirrored
Designed
{"error": "400 Bad Request"}
sales_order must be digits only, for example '4711'. You sent "SO-4711". Remove any prefix or spaces and call again.
{"error": "404"}
Sales order 4799 does not exist or you may not see it. Check the number with the user; to find a customer's orders, use sales_order_search.
The second designed error also avoids leaking whether an order exists to someone who isn't allowed to see it.
MCP defines four annotations. The MCP blog post on them gives their defaults:
Hint
Question it answers
Default if missing
readOnlyHint
Does the tool change its environment?
false (it may write)
destructiveHint
If it writes, can the change be destructive rather than additive?
true
idempotentHint
Is calling it again with the same arguments safe?
false
openWorldHint
Does it reach an open world of external entities?
true
Declare them explicitly, because the defaults assume the worst. A client can use them to skip confirmation for trusted read-only tools or to ask before destructive ones. But the specification says clients must treat annotations as untrusted unless they come from a trusted server, and the blog post is plain that hints are not enforcement. Your code still validates inputs and checks authorizations; the MCP specification lists both as things servers must do.
For SAP data, the lab's pattern is: no tool that changes a record. release_request_create drafts a request, is marked not read-only, not destructive and idempotent, says in its description that a person approves, and returns the same draft if called twice. How the approval is enforced and how the loop pauses for it is the subject of the next Unit 9 topics.
OpenAI's guide suggests fewer than 20 functions at the start of a turn, as a soft limit to test against. Anthropic's post warns that too many or overlapping tools distract agents. Size matters too: Anthropic's advanced tool use post describes 58 tools from five servers taking about 55K tokens before the conversation starts, and reports that letting the model search for tools on demand improved selection accuracy in its tests. If an agent needs many tools, offer the few a task needs, or give it a way to look tools up.
Descriptions cost tokens. The designed set in the lab is about 850 tokens of definitions against about 300 for the mirrored one. That is a good trade, because a clear description saves wrong calls, but it is a cost on every call.
You can't judge a tool set by reading it. Anthropic's post lists the metrics worth tracking when you run an agent over test tasks: accuracy, runtime, the number of tool calls, token use and tool errors. It also recommends a held-out set of tasks you don't tune against. The simplest place to start is a tool-selection test: requests a clerk would make, each with the right first tool. The lab builds one, and Building an evaluation harness shows how to grow it.
#Build it yourself: lint, compare and test two tool sets
You will build unit09/tool_design.py, one script with three commands that compare a mirrored tool set with a designed one over the same made-up SAP data:
lint checks every tool definition against the rules in this topic, without a model.
respond calls tools from both sets with good and bad input, and shows what the model would read back.
select runs a ten-request tool-selection test, either with a crude word-matching stand-in (--sample) or with a real model through SAP's orchestration service (--model).
flowchart LR
M[mirror set<br/>7 tools] --> L[lint]
D[designed set<br/>5 tools] --> L
M --> R[respond]
D --> R
M --> S[select<br/>10 requests]
D --> S
S --> SC[score per set]
Nothing in the script changes any data. The records are made up, shaped like SAP's sales order API.
Your course folder with its .venv. No new libraries.
About 40 minutes.
lint, respond and select --sample are free and need no account.
select --model makes 20 model calls (ten requests, two tool sets): a small per-request charge on SAP AI Core with the generative AI hub, using a model in your catalog that supports tool calling.
#Step 1: Open your course folder and turn on the virtual environment
Open VS Code, choose File > Open Folder, and open orchestrate-course.
Open a terminal: Terminal > New Terminal.
If the prompt doesn't start with (.venv), turn it on:
Windows (PowerShell):
.venv\Scripts\Activate.ps1
macOS / Linux:
source .venv/bin/activate
Check that the unit09 folder exists (same command on every system):
In VS Code's file list, right-click unit09, choose New File and name it tool_design.py.
Paste the code below and save.
"""Unit 9: tool design for agents. Two tool sets over the same made-up SAP data, compared three ways.
Set "mirror": tools that copy the API one to one, the way a first draft often looks.
Set "designed": the same data, shaped for a model to choose, call and read.
Commands (run from your course folder, with .venv turned on):
python unit09/tool_design.py lint # check both sets against design rules (no account)
python unit09/tool_design.py lint --set designed # one set only, with every finding listed
python unit09/tool_design.py respond # what each set returns, for good and bad input
python unit09/tool_design.py select --sample # tool-choice test with a crude word-matching stand-in
python unit09/tool_design.py select --model MODEL_NAME # the same test with a real model (SAP AI Core)
The real path reads the AICORE_ lines in .env (see "Set up for Unit 5").
Nothing in this file changes any data: release_request_create only drafts a request for a person.
"""
import argparse
import json
import os
import re
import sys
# ---------- made-up data, shaped like SAP's sales order API (A_SalesOrder in API_SALES_ORDER_SRV) ----------
def order(number, customer, amount, delivery_block="", credit_status="", requested="2026-10-12"):
"""A made-up order header with field names from A_SalesOrder. Code values are made up: what each
code means is configured per system, so never assume a code means the same everywhere."""
return {"__metadata": {"type": "API_SALES_ORDER_SRV.A_SalesOrderType"},
"SalesOrder": number, "SalesOrderType": "OR", "SalesOrganization": "1010",
"DistributionChannel": "10", "OrganizationDivision": "00", "SoldToParty": customer,
"CreationDate": "2026-10-01", "CreatedByUser": "CB9980000046", "PurchaseOrderByCustomer": f"PO-{number}",
"SalesOrderDate": "2026-10-01", "TotalNetAmount": amount, "TransactionCurrency": "EUR",
"PricingDate": "2026-10-01", "RequestedDeliveryDate": requested, "ShippingCondition": "01",
"IncotermsClassification": "EXW", "IncotermsTransferLocation": "Walldorf",
"CustomerPaymentTerms": "0004", "HeaderBillingBlockReason": "", "DeliveryBlockReason": delivery_block,
"OverallSDProcessStatus": "A", "TotalCreditCheckStatus": credit_status,
"OverallTotalDeliveryStatus": "A", "OverallSDDocumentRejectionSts": "A"}
ORDERS = {o["SalesOrder"]: o for o in [
order("4711", "10023", "1800.00", credit_status="B"),
order("4723", "10051", "640.00", delivery_block="01"),
order("4725", "10077", "3900.00", delivery_block="02"),
order("4730", "10023", "520.00"), order("4731", "10023", "75.00"), order("4732", "10023", "1210.00"),
order("4733", "10023", "300.00"), order("4734", "10023", "980.00"), order("4735", "10023", "60.00"),
]}
# Made-up texts a clerk would read for each blocked order. In a real system they come from the block
# reasons configured in that system, read in the user's language.
BLOCK_TEXT = {"4711": "Blocked by the credit check.",
"4723": "Incomplete: delivery address data missing.",
"4725": "Pricing: the customer disputes the price."}
# A made-up customer master and credit lookup (not an SAP API). An empty string means the value is missing.
CUSTOMERS = {
"10023": {"name": "Nordhafen Tools GmbH", "street": "Hauptstrasse 5", "postal_code": "69190",
"city": "Walldorf", "country": "DE", "credit_limit": 50000.00, "open_items": 50700.00},
"10051": {"name": "Donau Retail AG", "street": "Ringstrasse 12", "postal_code": "", "city": "Vienna",
"country": "AT", "credit_limit": 30000.00, "open_items": 4100.00},
"10077": {"name": "Lac Leman Instruments SA", "street": "Rue de Rive 3", "postal_code": "1204",
"city": "Geneva", "country": "CH", "credit_limit": 80000.00, "open_items": 12000.00},
}
RELEASE_REQUESTS = {} # drafts for a person to approve; nothing here ever reaches SAP
def is_blocked(o: dict) -> bool:
"""The course's rule from Unit 1: an order counts as blocked if any block field has a value."""
return any(o[f] for f in ("DeliveryBlockReason", "HeaderBillingBlockReason", "TotalCreditCheckStatus"))
# ---------- set 1: "mirror" tools, one per API call, as a first draft often looks ----------
def obj(props: dict, required: list) -> dict:
return {"type": "object", "properties": props, "required": required}
MIRROR = [
{"name": "get_order", "description": "Get an order.",
"parameters": obj({"id": {"type": "string"}}, ["id"])},
{"name": "get_order_details", "description": "Get order details.",
"parameters": obj({"id": {"type": "string"}}, ["id"])},
{"name": "get_order_status", "description": "Get order status.",
"parameters": obj({"id": {"type": "string"}}, ["id"])},
{"name": "get_bp", "description": "Get business partner.",
"parameters": obj({"id": {"type": "string"}}, ["id"])},
{"name": "get_credit", "description": "Get credit.",
"parameters": obj({"id": {"type": "string"}}, ["id"])},
{"name": "query", "description": "Query an entity set with a filter.",
"parameters": obj({"entity": {"type": "string"}, "filter": {"type": "string"}}, ["entity"])},
{"name": "update_order", "description": "Update an order.",
"parameters": obj({"id": {"type": "string"}, "fields": {"type": "object"}}, ["id", "fields"])},
]
def run_mirror(name: str, args: dict) -> dict:
"""Mirror tools pass the API straight through: whole records, codes, and bare status errors."""
key = str(args.get("id", ""))
if name in ("get_order", "get_order_details", "get_order_status"):
if not key.isdigit():
return {"error": "400 Bad Request"}
if key not in ORDERS:
return {"error": "404"}
record = ORDERS[key]
if name == "get_order_status":
return {k: v for k, v in record.items() if k.endswith(("Status", "Sts", "BlockReason"))}
if name == "get_order_details":
item = {"SalesOrderItem": "10", "Material": "TG12", "RequestedQuantity": "1",
"NetAmount": record["TotalNetAmount"]}
return {**record, "to_Item": {"results": [item]}}
return record
if name in ("get_bp", "get_credit"):
if key not in CUSTOMERS:
return {"error": "404"}
c = CUSTOMERS[key]
if name == "get_credit":
return {"limit": c["credit_limit"], "open": c["open_items"]}
return {"BusinessPartner": key, **{k: v for k, v in c.items() if k not in ("credit_limit", "open_items")}}
if name == "query":
match = re.search(r"SoldToParty eq '(\d+)'", str(args.get("filter", "")))
rows = [o for o in ORDERS.values() if not match or o["SoldToParty"] == match.group(1)]
return {"d": {"results": rows}}
if name == "update_order":
return {"error": "403"} # this lab never writes; a real mirror tool would have changed the order
return {"error": "unknown tool"}
# ---------- set 2: "designed" tools: fewer, named for the job, with results a model can use ----------
READ = {"readOnlyHint": True, "openWorldHint": False}
DESIGNED = [
{"name": "sales_order_get_block_summary",
"description": "Explain why one sales order is blocked: credit, pricing, incomplete data or another reason. "
"Returns the customer, net value, whether the order is blocked and the block reason as text "
"a clerk would read. Use it first whenever the user "
"asks about a specific order number. It only reads; it changes nothing.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order", "detail"],
"properties": {
"sales_order": {"type": "string", "description": "Sales order number, digits only, "
"for example '4711'."},
"detail": {"type": "string", "enum": ["concise", "detailed"],
"description": "Use 'concise' unless you need dates, the customer's PO number "
"or the sales organization."}}},
"annotations": READ},
{"name": "customer_get_credit_exposure",
"description": "Check a customer's credit exposure against the credit limit, and who must approve. "
"Exposure is open items plus the order's net value, computed for you. Use it when an order "
"is blocked by the credit check or the user asks how far over the limit a customer is. "
"Pass the order number when there is one so its value is included.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["customer", "sales_order"],
"properties": {
"customer": {"type": "string", "description": "Customer number (SoldToParty), digits "
"only, for example '10023'."},
"sales_order": {"type": ["string", "null"], "description": "The blocked order's number, "
"or null for no order."}}},
"annotations": READ},
{"name": "customer_get_address_gaps",
"description": "Find the address fields that are missing in a customer's master data. Returns the customer "
"name and a list of missing fields. Use it when an order is incomplete or blocked because "
"address or delivery data is missing.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["customer"],
"properties": {"customer": {"type": "string", "description": "Customer number (SoldToParty), "
"digits only."}}},
"annotations": READ},
{"name": "sales_order_search",
"description": "Search a customer's sales orders and list which ones are blocked. Returns at most 'limit' "
"orders, newest first, with a note when there are more. Use it when the user asks which or "
"how many orders a customer has, instead of reading orders one by one.",
"parameters": {"type": "object", "additionalProperties": False,
"required": ["customer", "blocked_only", "limit"],
"properties": {
"customer": {"type": "string", "description": "Customer number (SoldToParty), digits only."},
"blocked_only": {"type": "boolean", "description": "true to list blocked orders only."},
"limit": {"type": ["integer", "null"], "description": "Most orders to return, 1 to 20; "
"null means 5."}}},
"annotations": READ},
{"name": "release_request_create",
"description": "Draft a request to release one blocked sales order, for a person to approve. It does not "
"release anything: a person approves or rejects the draft in their own tool. Use it only "
"when the user asks to release or unblock an order, whatever reason the customer gives. "
"Calling it again for the same order returns the same draft.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order", "justification"],
"properties": {
"sales_order": {"type": "string", "description": "Sales order number, digits only."},
"justification": {"type": "string",
"description": "One or two sentences, with the facts from the other "
"tools, for the approver."}}},
"annotations": {"readOnlyHint": False, "destructiveHint": False, "idempotentHint": True,
"openWorldHint": False}},
]
def bad_number(field: str, value) -> dict:
"""An error the model can act on: what was wrong, what is expected, an example."""
return {"error": f"{field} must be digits only, for example '4711'. You sent {json.dumps(value)}. "
"Remove any prefix or spaces and call again."}
def run_designed(name: str, args: dict) -> dict:
if name == "sales_order_get_block_summary":
number = args.get("sales_order")
if not isinstance(number, str) or not number.isdigit():
return bad_number("sales_order", number)
if number not in ORDERS:
return {"error": f"Sales order {number} does not exist or you may not see it. Check the number "
"with the user; to find a customer's orders, use sales_order_search."}
o = ORDERS[number]
c = CUSTOMERS.get(o["SoldToParty"], {})
result = {"sales_order": number, "customer": o["SoldToParty"], "customer_name": c.get("name", ""),
"net_value": float(o["TotalNetAmount"]), "currency": o["TransactionCurrency"],
"blocked": is_blocked(o), "block_reason": BLOCK_TEXT.get(number, "")}
if args.get("detail") == "detailed":
result.update({"requested_delivery_date": o["RequestedDeliveryDate"],
"customer_po": o["PurchaseOrderByCustomer"], "sales_organization": o["SalesOrganization"]})
return result
if name == "customer_get_credit_exposure":
customer, number = args.get("customer"), args.get("sales_order")
if not isinstance(customer, str) or not customer.isdigit():
return bad_number("customer", customer)
if customer not in CUSTOMERS:
return {"error": f"Customer {customer} not found. Take the customer number from "
"sales_order_get_block_summary."}
if number is not None and (not isinstance(number, str) or number not in ORDERS):
return {"error": f"Sales order {json.dumps(number)} not found. Pass null to check the customer only."}
c = CUSTOMERS[customer]
order_value = float(ORDERS[number]["TotalNetAmount"]) if number else 0.0
exposure = c["open_items"] + order_value
over = round((exposure / c["credit_limit"] - 1) * 100, 1)
# The approval rule (made up for the course) lives here, in code, not in the prompt.
approver = ("none needed" if exposure <= c["credit_limit"]
else "CREDIT_MANAGER" if exposure <= c["credit_limit"] * 1.05 else "HEAD_OF_FINANCE")
return {"customer": customer, "customer_name": c["name"], "credit_limit": c["credit_limit"],
"open_items": c["open_items"], "order_value": order_value, "exposure": exposure,
"percent_over_limit": max(over, 0.0), "approver": approver, "currency": "EUR"}
if name == "customer_get_address_gaps":
customer = args.get("customer")
if not isinstance(customer, str) or not customer.isdigit():
return bad_number("customer", customer)
if customer not in CUSTOMERS:
return {"error": f"Customer {customer} not found."}
c = CUSTOMERS[customer]
missing = [f for f in ("street", "postal_code", "city", "country") if c[f] == ""]
return {"customer": customer, "customer_name": c["name"], "missing_fields": missing,
"complete": not missing}
if name == "sales_order_search":
customer, limit = args.get("customer"), args.get("limit")
if not isinstance(customer, str) or not customer.isdigit():
return bad_number("customer", customer)
limit = 5 if limit is None else limit
if not isinstance(limit, int) or not 1 <= limit <= 20:
return {"error": f"limit must be a whole number from 1 to 20, or null for 5. You sent {limit!r}."}
rows = [o for o in ORDERS.values() if o["SoldToParty"] == customer
and (is_blocked(o) or not args.get("blocked_only"))]
rows.sort(key=lambda o: o["SalesOrder"], reverse=True)
shown = [{"sales_order": o["SalesOrder"], "net_value": float(o["TotalNetAmount"]),
"blocked": is_blocked(o), "block_reason": BLOCK_TEXT.get(o["SalesOrder"], "")}
for o in rows[:limit]]
result = {"customer": customer, "total": len(rows), "showing": len(shown), "orders": shown}
if len(rows) > limit:
result["note"] = (f"{len(rows) - limit} more not shown. Set blocked_only to true, or raise limit "
"(up to 20), rather than reading orders one by one.")
return result
if name == "release_request_create":
number, why = args.get("sales_order"), str(args.get("justification", "")).strip()
if not isinstance(number, str) or not number.isdigit():
return bad_number("sales_order", number)
if number not in ORDERS or not is_blocked(ORDERS[number]):
return {"error": f"Sales order {number} is not blocked, so there is nothing to release."}
if len(why) < 20:
return {"error": "justification is too short. Give the approver the facts, for example the "
"exposure and limit from customer_get_credit_exposure."}
draft = RELEASE_REQUESTS.setdefault(number, {"request_id": f"RR-{number}", "sales_order": number,
"status": "waiting_for_approval", "justification": why})
return {**draft, "message": "Draft created. A person must approve it; nothing has been released."}
return {"error": f"Unknown tool {name!r}. Available: " + ", ".join(t["name"] for t in DESIGNED)}
SETS = {"mirror": (MIRROR, run_mirror), "designed": (DESIGNED, run_designed)}
def model_view(tool: dict) -> dict:
"""What the model is sent: name, description and parameters. Annotations stay with your code."""
return {k: tool[k] for k in ("name", "description", "parameters")}
# ---------- lint: design rules you can check without a model ----------
GENERIC_PARAMS = {"id", "key", "data", "value", "fields", "filter", "entity", "user", "input", "query"}
WRITE_WORDS = ("update", "create", "delete", "release", "post", "change", "set", "approve")
def lint_tool(t: dict) -> list:
"""Return (rule, message) pairs for one tool."""
out, name, desc = [], t["name"], t.get("description", "")
params = t.get("parameters", {})
props = params.get("properties", {})
if not re.fullmatch(r"[a-zA-Z0-9_-]{1,64}", name):
out.append(("name", "use letters, digits, _ or - only, at most 64 characters"))
if "_" not in name or len(name.split("_")) < 3:
out.append(("name", "name the object and the job, e.g. sales_order_get_block_summary"))
sentences = len([s for s in re.split(r"[.!?]\s", desc + " ") if s.strip()])
if sentences < 3:
out.append(("description", f"{sentences} sentence(s); say what it does, when to use it, what it returns"))
if not re.search(r"\bUse it\b", desc):
out.append(("description", "no 'Use it when/for ...': the model can't tell when to pick it"))
for p, spec in props.items():
if not spec.get("description"):
out.append(("parameter", f"'{p}' has no description"))
if p in GENERIC_PARAMS:
out.append(("parameter", f"'{p}' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'"))
if spec.get("type") == "object" and "properties" not in spec:
out.append(("parameter", f"'{p}' is a free-form object; list the allowed fields"))
if params.get("additionalProperties") is not False:
out.append(("strict", "set additionalProperties to false"))
if set(params.get("required", [])) != set(props):
out.append(("strict", "list every parameter in required; use null for 'not given'"))
ann = t.get("annotations")
if ann is None or "readOnlyHint" not in ann:
out.append(("safety", "declare readOnlyHint; when it is missing, MCP's default says the tool may write"))
writes = (ann or {}).get("readOnlyHint") is False or any(w in name.lower() for w in WRITE_WORDS)
if writes and "approve" not in desc.lower():
out.append(("safety", "a tool that writes must say a person approves, and your code must enforce it"))
return out
def lint_set(tools: list) -> list:
"""Rules about the set as a whole."""
out, names = [], [t["name"] for t in tools]
if len(tools) > 20:
out.append(("set", f"{len(tools)} tools; keep the set small (OpenAI suggests fewer than 20)"))
for i, a in enumerate(names):
for b in names[i + 1:]:
ta, tb = set(a.split("_")), set(b.split("_"))
if len(ta & tb) / len(ta | tb) >= 0.5 and len(ta & tb) >= 2:
out.append(("set", f"'{a}' and '{b}' overlap; merge them or make the difference obvious"))
size = len(json.dumps([model_view(t) for t in tools]))
out.append(("info", f"definitions are about {size // 4} tokens, sent with every model call"))
return out
def cmd_lint(args) -> None:
for set_name in (["mirror", "designed"] if args.set == "both" else [args.set]):
tools = SETS[set_name][0]
per_tool = {t["name"]: lint_tool(t) for t in tools}
set_findings = lint_set(tools)
problems = sum(len(v) for v in per_tool.values()) + sum(1 for r, _ in set_findings if r != "info")
print(f"== {set_name}: {len(tools)} tools, {problems} finding(s)")
for name, findings in per_tool.items():
mark = "ok " if not findings else f"{len(findings)}x "
print(f" {mark:4}{name}")
if args.set != "both":
for rule, msg in findings:
print(f" [{rule}] {msg}")
for rule, msg in set_findings:
print(f" [{rule}] {msg}")
print()
if args.set == "both":
print("Add --set mirror or --set designed to see every finding.")
# ---------- respond: what the model would read back ----------
def show(set_name: str, tool: str, args: dict) -> None:
result = SETS[set_name][1](tool, args)
text = json.dumps(result)
shown = text if len(text) <= 300 else text[:300] + f"... ({len(text) - 300} more characters)"
print(f" {tool}({json.dumps(args)})\n -> {len(text)} characters: {shown}")
def cmd_respond(args) -> None:
print("1. Read order 4711 (blocked by the credit check)")
show("mirror", "get_order", {"id": "4711"})
show("designed", "sales_order_get_block_summary", {"sales_order": "4711", "detail": "concise"})
print("\n2. A slightly wrong order number")
show("mirror", "get_order", {"id": "SO-4711"})
show("designed", "sales_order_get_block_summary", {"sales_order": "SO-4711", "detail": "concise"})
print("\n3. Is customer 10023 over the credit limit, with order 4711?")
show("mirror", "get_credit", {"id": "10023"})
show("designed", "customer_get_credit_exposure", {"customer": "10023", "sales_order": "4711"})
print("\n4. Which orders does customer 10023 have?")
show("mirror", "query", {"entity": "A_SalesOrder", "filter": "SoldToParty eq '10023'"})
show("designed", "sales_order_search", {"customer": "10023", "blocked_only": False, "limit": None})
print("\n5. Release order 4711, twice")
show("mirror", "update_order", {"id": "4711", "fields": {"TotalCreditCheckStatus": ""}})
why = "Exposure 52,500 EUR is 5% over the 50,000 EUR limit; customer pays on time."
show("designed", "release_request_create", {"sales_order": "4711", "justification": why})
show("designed", "release_request_create", {"sales_order": "4711", "justification": why})
print("\nThe mirror set returns codes, whole records and bare status numbers. The designed set returns "
"what the next decision needs, and errors that say how to recover.")
# ---------- select: does the model pick the right tool? ----------
# Each case: a request, and the tools that would be right in each set (None: the set has no right tool).
CASES = [
("Why is order 4711 blocked?",
{"mirror": ["get_order", "get_order_status"], "designed": ["sales_order_get_block_summary"]}),
("How far over its credit limit is customer 10023?",
{"mirror": ["get_credit"], "designed": ["customer_get_credit_exposure"]}),
("Which orders of customer 10051 are blocked?",
{"mirror": ["query"], "designed": ["sales_order_search"]}),
("What address data is missing for customer 10051?",
{"mirror": ["get_bp"], "designed": ["customer_get_address_gaps"]}),
("Please get order 4711 released, the customer always pays.",
{"mirror": None, "designed": ["release_request_create"]}),
("Is order 4725 blocked for pricing or for credit?",
{"mirror": ["get_order", "get_order_status"], "designed": ["sales_order_get_block_summary"]}),
("How many orders does customer 10023 have?",
{"mirror": ["query"], "designed": ["sales_order_search"]}),
("Does customer 10023 need the head of finance to approve?",
{"mirror": ["get_credit"], "designed": ["customer_get_credit_exposure"]}),
("Order 4723 is incomplete. What is missing for its customer?",
{"mirror": ["get_order", "get_bp"], "designed": ["sales_order_get_block_summary", "customer_get_address_gaps"]}),
("Unblock order 4723 now that the postal code is fixed.",
{"mirror": None, "designed": ["release_request_create"]}),
]
STOP = set("a an the is are of for to and or it its this that what how does do in on with by be i me my "
"now please get has have any one only when whether there their them you your use".split())
def words(text: str) -> set:
out = set()
for w in re.findall(r"[a-z]+", text.lower()):
if w in STOP:
continue
out.add(w[:5]) # compare first five letters, so "orders" matches "order" and "released" "release"
return out
def sample_pick(question: str, tools: list) -> str:
"""A crude stand-in for a model: pick the tool whose name and description share the most words with
the request. Ties go to the tool listed first. It knows nothing about meaning; it only shows that
the words in a definition are what selection has to go on."""
q = words(question)
scores = [(len(q & words(t["name"].replace("_", " ") + " " + t["description"])), -i, t["name"])
for i, t in enumerate(tools)]
return max(scores)[2]
def real_picker(model: str, tools: list):
"""Ask a real model through SAP's orchestration service which tool it would call first."""
from dotenv import load_dotenv
load_dotenv()
names = ["AICORE_CLIENT_ID", "AICORE_CLIENT_SECRET", "AICORE_AUTH_URL", "AICORE_BASE_URL",
"AICORE_RESOURCE_GROUP"]
missing = [n for n in names if not os.environ.get(n)]
if missing:
sys.exit("Missing in .env: " + ", ".join(missing) + ". See 'Set up for Unit 5', Step 5. "
"Or use --sample to try without an account.")
from gen_ai_hub.orchestration_v2 import (FunctionObject, FunctionTool, LLMModelDetails, ModuleConfig,
OrchestrationConfig, OrchestrationService,
PromptTemplatingModuleConfig, SystemMessage, Template, UserMessage)
strict = all(t["parameters"].get("additionalProperties") is False for t in tools)
fts = [FunctionTool(function=FunctionObject(name=t["name"], description=t["description"],
parameters=t["parameters"], strict=strict)) for t in tools]
template = Template(template=[SystemMessage(content="You help SAP order-to-cash clerks with blocked sales "
"orders. Call the one tool that best fits the request."),
UserMessage(content="{{?question}}")], tools=fts)
config = OrchestrationConfig(modules=ModuleConfig(prompt_templating=PromptTemplatingModuleConfig(
prompt=template, model=LLMModelDetails(name=model, params={"temperature": 0}, timeout=60, max_retries=1))))
service = OrchestrationService(config=config)
def pick(question: str) -> str:
message = service.run(placeholder_values={"question": question}).final_result.choices[0].message
return message.tool_calls[0].function.name if message.tool_calls else "(no tool)"
return pick, service
def cmd_select(args) -> None:
print("Which tool is picked first for each request?" + (" [sample: word-matching stand-in]" if args.sample
else f" [model: {args.model}]") + "\n")
totals = {}
for set_name in ("mirror", "designed"):
tools = SETS[set_name][0]
service = None
if args.sample:
pick = lambda q, tools=tools: sample_pick(q, tools)
else:
pick, service = real_picker(args.model, [model_view(t) for t in tools])
right = 0
print(f"== {set_name}")
try:
for question, expected in CASES:
ok_names = expected[set_name]
try:
chosen = pick(question)
except Exception as error:
sys.exit(f"The call failed: {type(error).__name__}: {str(error)[:400]}")
good = ok_names is not None and chosen in ok_names
right += good
note = "" if ok_names is not None else " (this set has no safe tool for it)"
print(f" {'right' if good else 'WRONG'} {chosen:<33} <- {question}{note}")
finally:
if service is not None:
service.close_http_connection()
totals[set_name] = right
print()
for set_name, right in totals.items():
print(f"{set_name:<9} {right}/{len(CASES)} right")
if args.sample:
print("\n[sample] A word-matching stand-in, not a model. Run with --model for a real result.")
def main() -> None:
parser = argparse.ArgumentParser(description="Tool design for agents: lint, respond, select.")
sub = parser.add_subparsers(dest="command", required=True)
p = sub.add_parser("lint", help="check tool definitions against design rules")
p.add_argument("--set", choices=["both", "mirror", "designed"], default="both")
sub.add_parser("respond", help="compare what each set returns")
p = sub.add_parser("select", help="test which tool is picked for ten requests")
p.add_argument("--sample", action="store_true", help="use the word-matching stand-in (no account)")
p.add_argument("--model", default="gpt-4o-mini", help="model name from your catalog")
args = parser.parse_args()
{"lint": cmd_lint, "respond": cmd_respond, "select": cmd_select}[args.command](args)
if __name__ == "__main__":
main()
A linter is a program that checks code or definitions against rules without running them. Run:
python unit09/tool_design.py lint
You should see:
== mirror: 7 tools, 57 finding(s)
7x get_order
6x get_order_details
6x get_order_status
7x get_bp
7x get_credit
10x query
11x update_order
[set] 'get_order' and 'get_order_details' overlap; merge them or make the difference obvious
[set] 'get_order' and 'get_order_status' overlap; merge them or make the difference obvious
[set] 'get_order_details' and 'get_order_status' overlap; merge them or make the difference obvious
[info] definitions are about 296 tokens, sent with every model call
== designed: 5 tools, 0 finding(s)
ok sales_order_get_block_summary
ok customer_get_credit_exposure
ok customer_get_address_gaps
ok sales_order_search
ok release_request_create
[info] definitions are about 853 tokens, sent with every model call
Add --set mirror or --set designed to see every finding.
Now see the findings for the mirrored set:
python unit09/tool_design.py lint --set mirror
The last tool shows why mirrored write tools are dangerous:
11x update_order
[name] name the object and the job, e.g. sales_order_get_block_summary
[description] 1 sentence(s); say what it does, when to use it, what it returns
[description] no 'Use it when/for ...': the model can't tell when to pick it
[parameter] 'id' has no description
[parameter] 'id' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'
[parameter] 'fields' has no description
[parameter] 'fields' is ambiguous; say what it is, e.g. 'sales_order' or 'customer'
[parameter] 'fields' is a free-form object; list the allowed fields
[strict] set additionalProperties to false
[safety] declare readOnlyHint; when it is missing, MCP's default says the tool may write
[safety] a tool that writes must say a person approves, and your code must enforce it
A clean lint does not prove a tool set works. It catches the easy mistakes so the test in Step 5 can focus on behavior.
You will see five pairs. The first two look like this (results longer than 300 characters are cut, with a count of what was left out):
1. Read order 4711 (blocked by the credit check)
get_order({"id": "4711"})
-> 810 characters: {"__metadata": {"type": "API_SALES_ORDER_SRV.A_SalesOrderType"}, "SalesOrder": "4711", "SalesOrderType": "OR", "SalesOrganization": "1010", "DistributionChannel": "10", "OrganizationDivision": "00", "SoldToParty": "10023", "CreationDate": "2026-10-01", "CreatedByUser": "CB9980000046", "PurchaseOrder... (510 more characters)
sales_order_get_block_summary({"sales_order": "4711", "detail": "concise"})
-> 190 characters: {"sales_order": "4711", "customer": "10023", "customer_name": "Nordhafen Tools GmbH", "net_value": 1800.0, "currency": "EUR", "blocked": true, "block_reason": "Blocked by the credit check."}
2. A slightly wrong order number
get_order({"id": "SO-4711"})
-> 28 characters: {"error": "400 Bad Request"}
sales_order_get_block_summary({"sales_order": "SO-4711", "detail": "concise"})
-> 131 characters: {"error": "sales_order must be digits only, for example '4711'. You sent \"SO-4711\". Remove any prefix or spaces and call again."}
Look for three more things in the rest of the output:
Pair 3: the designed credit tool returns "exposure": 52500.0, "percent_over_limit": 5.0 and "approver": "CREDIT_MANAGER". Your code did the arithmetic and applied the rule; the mirrored tool returns two bare numbers.
Pair 4: the mirrored query returns about 5,700 characters for one customer; the designed search returns about 600, with "total": 7, "showing": 5 and a note on how to narrow.
Pair 5: calling release_request_create twice returns the same "request_id": "RR-4711". That is what idempotent means: a repeated call does no extra harm.
#Step 5: Run the tool-selection test with the stand-in
python unit09/tool_design.py select --sample
You should see:
Which tool is picked first for each request? [sample: word-matching stand-in]
== mirror
right get_order <- Why is order 4711 blocked?
right get_credit <- How far over its credit limit is customer 10023?
WRONG get_order <- Which orders of customer 10051 are blocked?
WRONG get_order <- What address data is missing for customer 10051?
WRONG get_order <- Please get order 4711 released, the customer always pays. (this set has no safe tool for it)
right get_order <- Is order 4725 blocked for pricing or for credit?
WRONG get_order <- How many orders does customer 10023 have?
WRONG get_order <- Does customer 10023 need the head of finance to approve?
right get_order <- Order 4723 is incomplete. What is missing for its customer?
WRONG get_order <- Unblock order 4723 now that the postal code is fixed. (this set has no safe tool for it)
== designed
right sales_order_get_block_summary <- Why is order 4711 blocked?
...
right release_request_create <- Unblock order 4723 now that the postal code is fixed.
mirror 4/10 right
designed 10/10 right
[sample] A word-matching stand-in, not a model. Run with --model for a real result.
Note the two release requests. The mirrored set has no safe tool for them: its only candidate, update_order, would change the order directly. The test counts both as misses on purpose.
#Step 6 (optional): Run the test with a real model
This step needs your SAP AI Core keys in .env, as set up in Unit 5. Replace MODEL_NAME with a model from your catalog that supports tool calling:
The output has the same shape as Step 5, with [model: MODEL_NAME] in the header. Compare the two scores, and look at which requests each set got wrong. The mirrored set's two release requests always count as misses, because it has no safe tool for them. Run it twice: a model's choices can vary between runs, which is one reason to keep the test.
SAP's Python SDK documentation (Orchestration Service V2) offers three ways to define a tool:
The @function_tool() decorator: the function's signature and docstring describe the tool to the model. That makes the docstring your description, so write it with the same care, three or four sentences, when to use it.
FunctionTool with a FunctionObject of name, description, parameters and strict, as the lab's real_picker does. FunctionTool.from_function(..., strict=True) builds one from a Python function.
A plain JSON Schema dictionary, useful when the tool runs somewhere else.
The model's choice arrives in response.final_result.choices[0].message.tool_calls; your code runs the tool and returns a ToolChatMessage with the tool_call_id. The SDK has no built-in agentic loop, as Agents from first principles showed, so every design decision here, including validation, limits and approval, sits in your code. Orchestration modules such as data masking and content filtering, from SAP Generative AI Hub and the orchestration service, apply to each call of the loop.
SAP's Joule Studio course lists six tool types for custom Joule agents: Joule skills, document grounding, a calculator, human in the loop, other Joule agents, and MCP servers. Its three principles line up with this topic:
SAP's principle
In this topic
It must be a validated capability
Test a tool on its own before an agent sees it; run the selection test after
Define its purpose in plain business language
Decision 3: what, returns, when, limits
Instructions as a tool usage guide for a new team member
Selection strategy, sequencing, source priority and limits in the agent's instructions
The course's examples of good instructions name tools in the instructions: "use the GetOrderStatus tool" for delivery status, and a discount tool that must not be used for VIP customers, with a hand-over to a human sales manager. That last example is a business rule. Put it in the instructions so the agent can explain it, and also enforce it in the tool's code, as customer_get_credit_exposure does with the approval rule. Building Joule agents is covered later in Unit 9.
The mirrored records in the lab use the real field list of A_SalesOrder from API_SALES_ORDER_SRV, as in SAP's own sample record: SalesOrder, SoldToParty, TotalNetAmount, DeliveryBlockReason, TotalCreditCheckStatus, OverallSDProcessStatus and more than 20 others. A designed tool calls the same released API, as in Calling your first SAP API, and then:
selects the fields the job needs, for example with OData $select;
turns codes into texts using your system's configuration, because block reasons and status values are configured per system;
adds what the clerk would look up next, such as the customer's name.
Wrapping SAP APIs as tools with the user's identity, authorizations and audit logging is covered later in Unit 9.
When you offer tools through MCP, the same rules hold, and annotations become part of the definition. The server in Set up for Unit 9 already marked its tools with ToolAnnotations(read_only_hint=True) and returned errors with ToolError, which the client receives as an error result. This topic used the 2025-06-18 version of the MCP specification's tools page; the protocol topic later in Unit 9 covers the current version.
Tool definitions travel with each orchestration call, so longer definitions make every step of the loop a larger request in the generative AI hub; Set up for Unit 5 covers access and plans. For Joule Studio, access and pricing were changing as of October 2026; check SAP's current terms before planning.
JSON Schema in your code, any vendor's tool calling API
FunctionTool, @function_tool() or a schema in the orchestration SDK
Descriptions
Yours to write and test
Yours to write too: docstrings in the SDK, the description form in Joule Studio
Validation and business rules
Your tool code
Your tool code, or the governed logic inside a Joule skill
Approval before changes
A draft-and-approve tool and your own approval step
Joule Studio's human-in-the-loop tool for Joule agents
Reuse across agents
An MCP server
MCP servers as a Joule Studio tool type
Testing tool selection
A selection test like the lab's, in your harness
Check what Joule Studio's own testing offers; keep your own cases either way
A rule of thumb: whichever runtime runs the agent, the tool set is your design. Neither SAP nor a framework can write a description that knows your clerks' jobs.
Security and SAP authorizations. A tool runs with some identity. Prefer the calling user's own SAP authorizations for read tools, so a tool can never return more than the user could see. Validate every argument in code, as the MCP specification requires of servers. Unit 11 covers agent permissions in depth.
Errors that don't leak. Helpful is not the same as revealing. "Does not exist or you may not see it" helps the model without telling an unauthorized user that an order exists.
Writes. No tool changes an SAP record without a person's approval, enforced in code. Prefer tools that create a request over tools that change the record, make them idempotent, and log who asked and what was proposed.
Annotations are not controls. Declare them so clients can ask for confirmation. Don't rely on them from servers you don't trust, and never instead of authorizations.
Prompt injection through results. Any text a tool returns, such as an order note, can carry instructions. Keep free text out of results unless the job needs it, and label it as data.
Evaluation and change control. A description change is a behavior change. Version tool definitions with the agent's instructions, and re-run the selection test before each release. Keep held-out cases you don't tune against.
Cost. Count definition tokens and result sizes per call. Offer only the tools a task needs, cap list results, and keep a detail switch for the rare case that needs more.
Operations. Log every tool call with its arguments, result size, duration, errors and tool-definition version. Rising error rates on one tool usually mean its description or its error messages need work.
Clean core. Tools call released SAP APIs from BTP. Code-to-text translation and business rules live in the tool layer, not in custom code inside S/4HANA.
You will add a sixth tool for pricing disputes, two test requests for it, and a short design note. The note feeds the next Unit 9 topics, where tools run in multi-step agents and over real SAP APIs.
Open unit09/tool_design.py and find the line RELEASE_REQUESTS = {}.
Just above it, add made-up price data for order 4725:
Find {"name": "release_request_create", inside DESIGNED. Just above that line, at the same indentation, add the new tool:
{"name": "sales_order_get_price_conditions",
"description": "Read the price conditions of one sales order: list price, discount and net price. "
"Use it when an order is blocked for pricing or the customer disputes a price. "
"It only reads; it changes nothing.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["sales_order"],
"properties": {"sales_order": {"type": "string",
"description": "Sales order number, digits only."}}},
"annotations": READ},
In run_designed, find the last line, return {"error": f"Unknown tool {name!r}. .... Just above it, at the same indentation as the other if name == lines, add:
if name == "sales_order_get_price_conditions":
number = args.get("sales_order")
if not isinstance(number, str) or not number.isdigit():
return bad_number("sales_order", number)
return PRICES.get(number) or {"error": f"No price conditions found for sales order {number}."}
In CASES, just before the closing ], add two requests:
("What discount did order 4725 get?",
{"mirror": ["get_order_details"], "designed": ["sales_order_get_price_conditions"]}),
("The customer disputes the price on order 4725. What was agreed?",
{"mirror": ["get_order_details"], "designed": ["sales_order_get_price_conditions"]}),
Run the lint for the designed set:
python unit09/tool_design.py lint --set designed
You should see 6 tools, 0 finding(s) and the definitions growing to about 970 tokens. If the new tool has findings, fix its definition until it has none.
Run the selection test:
python unit09/tool_design.py select --sample
You should see designed 12/12 right. If you have keys, also run it with --model MODEL_NAME.
Now break it on purpose: change the new tool's description to "Get prices.", run steps 6 and 7 again, and note what changed. Then undo the change.
In unit09, create tool_design_notes.md with these headings:
Scores: the mirror and designed scores, with --sample and, if you ran it, with a real model.
The broken description: what the lint and the test showed in step 8.
Results: for the price tool, which fields you return and which you leave out, and why.
Writes: which of the six tools may change anything, and how a person approves it.
Open questions: at least two, such as which SAP API the price tool would call and which identity it would use.
Done when:lint --set designed shows 6 tools, 0 finding(s), select --sample shows designed 12/12 right, and tool_design_notes.md covers scores, the broken description, results, writes and at least two open questions.
Pick one answer for each question. The explanation appears after you choose.
1Why does Anthropic's guidance say to avoid tools that merely wrap API endpoints?
Answer: B. One tool per endpoint gives the model near-twin choices and whole-record results. Consolidated tools named for a job, like search instead of list-everything, make the right choice obvious and keep results small.
2Claude's documentation suggests grouping related actions into one tool with an action parameter. Where does this topic stop consolidating for SAP data, and why?
Answer: D. Fewer tools reduce selection errors, but a write needs its own authorization check, approval and audit line. Separate write tools are easy to enforce and easy to leave out of a read-only agent.
3In the lab, why does customer_get_credit_exposure compute the exposure and the approver in code?
Answer: A. OpenAI's guide says to offload work to your code, and a rule that decides who approves must hold every time. The tool returns the exposure, the percentage and the approver, and the model explains them.
4A tool must accept "no limit given". How do you keep its schema ready for strict mode?
Answer: B. OpenAI's strict mode needs every field in required and additionalProperties: false. Optional values are expressed as null, as in the lab's "type": ["integer", "null"].
5Your search tool finds 400 orders for a customer. What should it return?
Answer: C. Anthropic recommends sensible defaults for pagination and truncation, and a note that steers the agent to narrower searches. The lab's search returns 5 of 7 with a note to set blocked_only or raise limit.
6An agent keeps calling get_order with "SO-4711" and getting "400 Bad Request". What is the best fix?
Answer: D. The error text is the only guidance the model gets for its next try. Say what was wrong, what is expected and give an example, and put the format in the parameter description so the first call is right.
7An MCP server from a third party marks its delete_items tool with readOnlyHint: true. What follows?
Answer: B. The MCP specification says clients must treat annotations as untrusted unless they come from trusted servers, and hints are not enforcement. Note also the default when the hint is missing is false, meaning the tool may write.
8The clerks ask the agent to release order 4711. Which tool design does this topic recommend?
Answer: C. A draft request lets the agent prepare the case while a person decides, and idempotency means a repeated call creates no second request. A prompt warning is not a control, and an update tool still changes SAP data without approval.
Ask a question
Testing: only staff see this
Stuck on something in this layer? Ask it here. Questions are answered in the order they arrive, and the answer appears under My questions.
Sign in (free) to ask a question. You can ask anonymously.
Sources
Writing effective tools for AI agents, using AI agents (Anthropic Engineering, 11 September 2025)— tools as a contract between deterministic systems and non-deterministic agents; don't merely wrap API endpoints; consolidate (search_contacts instead of list_contacts); namespacing; meaningful names over cryptic identifiers; response_format concise or detailed; pagination, filtering, truncation with defaults; actionable errors; describe tools as to a new hire; unambiguous parameter names such as user_id; evaluation metrics and held-out test sets
Define tools (Claude Platform documentation)— descriptions of at least 3 to 4 sentences, the most important factor in tool performance; say what, when, parameters and caveats; group related actions with an action parameter; namespacing; name regex; return high-signal fields and stable identifiers; input_examples; strict mode
Function calling (OpenAI API documentation)— describe purpose and each parameter; the intern test; enums and object structure to prevent invalid states; don't make the model fill arguments you already know; combine functions always called in sequence; aim for fewer than 20 functions at the start of a turn (soft); strict mode needs additionalProperties false, all fields required, null for optional
Tools (Model Context Protocol specification, version 2025-06-18)— tool fields name, title, description, inputSchema, outputSchema, annotations; annotations untrusted unless from trusted servers; protocol errors vs tool execution errors with isError; servers must validate inputs, apply access controls, rate limit and sanitize outputs; a human in the loop able to deny invocations
Orchestration Service V2 API (SAP Cloud SDK for AI, Python)— tools from the @function_tool() decorator (signature and docstring describe the tool), FunctionTool with FunctionObject (name, description, parameters, strict), or a JSON Schema dictionary; tool_calls in the response; ToolChatMessage; no built-in abstraction for the agentic loop
Leveraging tools and Model Context Protocol (MCP) for custom Joule agents (SAP Learning)— tool types Joule skills, document grounding, calculator, human in the loop, Joule agents and MCP; description quality determines agent accuracy; three principles (validated capability, purpose in plain business language, instructions as a tool usage guide for a new team member); tool selection, sequencing, data source priority and limitations