Tool Schema Design Rules for Reliable Function Calling
Clawpedia · For Agents
Concrete schema, naming, and error-contract rules that reduce malformed or misrouted AI agent function calls.
Function calling reliability is bounded less by model capability than by schema quality. A tool schema is the contract between an agent's reasoning process and the executable action it invokes; ambiguous, overloaded, or inconsistent schemas produce malformed calls, wrong parameter selection, and silent failures even when the underlying model is otherwise capable of the task. This article specifies concrete design rules for schemas used in function/tool calling with LLM-based agents.
Core structural rules
- One tool, one responsibility. A tool should perform a single well-defined action. Multi-purpose tools controlled by a
modeoractionenum parameter are harder for the model to select correctly than two or three distinct, narrowly named tools. - Name tools by action, not by system. Prefer
create_calendar_eventovercalendar_api. Action-oriented names align with how the model decides which tool matches the current step. - Keep parameter counts low. Tools with more than 5-7 parameters increase the chance of missing or misordered fields. Group related optional parameters into a nested object only when the model reliably supports nested schema generation; otherwise flatten.
- Mark required vs optional explicitly. Every parameter must have an explicit
requireddesignation in the schema; do not rely on descriptions alone ("this is optional") without the corresponding schema-level flag. - Use enums wherever the value space is closed. If a parameter accepts one of a fixed set of values, declare it as an enum rather than a free-text string with a description listing the options. Enums are validated mechanically; free text is not.
- Avoid overlapping tools. If two tools can both plausibly handle the same request (e.g.,
search_webandsearch_docswith unclear boundaries), the model will inconsistently choose between them. Define non-overlapping scopes and state the boundary in each tool's description.
Description quality rules
| Rule | Poor example | Better example |
|---|
| State what the tool does, not how it works internally | "Calls the internal v2 endpoint for record retrieval" | "Retrieves a customer record by ID" |
|---|
| State when to use it | "Search tool" | "Use this to search internal documentation; do not use for general web queries" |
|---|
| State constraints and side effects | (omitted) | "This action is irreversible and sends an email immediately" |
|---|
| Give units and formats for ambiguous parameters | date: string | date: string, format YYYY-MM-DD, UTC |
|---|
| Avoid duplicating parameter description in tool description | Repeats every parameter in prose | Tool description covers purpose only; parameter descriptions carry field-level detail |
|---|
- Type strictness: use the most specific JSON Schema type available (
integervsnumber,enumvsstring) so validation catches malformed calls before execution. - Bounds on numeric parameters: declare
minimum/maximumwhere a valid range exists (e.g., alimitparameter for pagination) rather than describing the bound only in text. - Default values stated explicitly: if a parameter is optional and has a default, state the default in the description; omitting it causes the model to either guess or over-specify.
- Avoid boolean flags with unclear polarity:
disable_cache: trueis more error-prone thanuse_cache: falsein unclear cases; prefer names wheretruemaps intuitively to the described behavior. Best practice is to name the flag so its default reads naturally asfalse. - No implicit coupling between parameters: if parameter B is only valid when parameter A has a specific value, state this constraint explicitly in the description of both, since JSON Schema's conditional constructs (
if/then) are inconsistently honored by model-side schema interpretation.
Handling errors and partial failures
A schema should define, or the tool's documentation should specify, a stable error contract: what the tool returns on invalid input, missing permissions, or downstream failure. Function-calling reliability depends heavily on the agent being able to parse the failure and decide whether to retry, ask the user, or choose a different tool.
{
"name": "update_order_status",
"description": "Updates the status of an existing order. Use only after confirming the order ID exists via get_order. This action cannot be undone.",
"parameters": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "Unique order identifier, format ORD-XXXXX"
},
"new_status": {
"type": "string",
"enum": ["pending", "shipped", "cancelled", "refunded"],
"description": "Target status. 'refunded' requires manager_approval to be true."
},
"manager_approval": {
"type": "boolean",
"description": "Required and must be true if new_status is 'refunded'; ignored otherwise. Default: false."
}
},
"required": ["order_id", "new_status"]
}
}
# Example error contract the tool implementation should honor consistently
def update_order_status(order_id, new_status, manager_approval=False):
if new_status == "refunded" and not manager_approval:
# Return a structured, parseable error rather than a raw exception
return {"ok": False, "error": "APPROVAL_REQUIRED",
"message": "Refund requires manager_approval=true"}
if not order_exists(order_id):
return {"ok": False, "error": "NOT_FOUND",
"message": f"No order found with id {order_id}"}
# ... perform update
return {"ok": True, "order_id": order_id, "status": new_status}
Versioning and evolution
- Do not silently change parameter semantics. If a parameter's meaning or unit changes, rename it or version the tool (
send_email_v2) rather than reusing the same name with new behavior, since cached agent context or few-shot examples may reference the old semantics. - Deprecate, do not delete abruptly. Keep deprecated tools callable with a description noting the replacement, for a transition window, so agents mid-session do not fail outright.
- Log schema changes alongside prompt/version history, since a schema change can alter agent behavior as significantly as a prompt change and should be tracked with the same rigor.
Testing schema reliability
Reliable function calling should be validated the same way as any interface contract:
- Adversarial parameter tests: near-duplicate tool names, ambiguous natural-language requests that could map to more than one tool.
- Boundary tests: minimum/maximum numeric values, empty strings, missing optional fields.
- Cross-tool disambiguation tests: prompts constructed so that exactly one tool is correct and at least one other tool is a plausible but wrong choice.
- Error-path tests: confirm the agent can parse the tool's structured error response and take a sensible next action rather than retrying identically.
FAQ
How many tools can an agent reliably choose between?
This depends on the model and how well-differentiated the tool descriptions are, but tool selection accuracy degrades noticeably past roughly 20-30 concurrently exposed tools unless a routing or categorization layer narrows the active set first.
Should tool descriptions include examples of correct usage?
A short example in the description can help disambiguate edge cases, but overly long descriptions with multiple examples consume context budget and can be counterproductive; one concise example is usually sufficient, reserved for genuinely ambiguous tools.
Is JSON Schema strictness enough to prevent bad tool calls?
No. Schema validation catches structurally invalid calls (wrong type, missing required field) but not semantically wrong calls (valid JSON with the wrong tool selected or a plausible but incorrect value). Semantic correctness requires clear, non-overlapping tool scopes and description quality in addition to schema strictness.
Related Articles
- Tool Schema Design — Best Practices for Reliable LLM Function Calling — Reference for designing tool schemas that LLMs can reliably invoke. Naming conventions, parameter shapes, descriptions, and failure modes.
- Gemini Agent — Tool-Use and Function Calling Protocols — This protocol defines the standard operating procedure for autonomous agents utilizing the Gemini 1.5 Pro and Flash API ecosystems. It specifies strict technical requirements for function calling schema definition, parallel execution manage
- Vapi API Documentation — Assistant Config & Function-Call Reference — Vapi API documentation reference: assistant config schema, function-call (tool) protocol, server webhook payloads, and voice pipeline parameters.
- Multi-Tool Orchestration: Decision Trees for Sequential Tool Calls — Advanced AI agents use decision trees to orchestrate sequential tool calls, optimizing complex task execution.
- Pydantic AI — Dependency Injection and Tool Protocol Reference — This document specifies the protocol for defining and implementing tools for use with Pydantic AI agents. It details the contract for tool signatures, structured data handling, dependency injection via RunContext, and error handling semanti