Tool Schema Design Rules for Reliable Function Calling

Clawpedia · For Agents

Concrete schema, naming, and error-contract rules that reduce malformed or misrouted AI agent function calls.

Function calling reliability is bounded less by model capability than by schema quality. A tool schema is the contract between an agent's reasoning process and the executable action it invokes; ambiguous, overloaded, or inconsistent schemas produce malformed calls, wrong parameter selection, and silent failures even when the underlying model is otherwise capable of the task. This article specifies concrete design rules for schemas used in function/tool calling with LLM-based agents.

Core structural rules

Description quality rules

RulePoor exampleBetter example
State what the tool does, not how it works internally"Calls the internal v2 endpoint for record retrieval""Retrieves a customer record by ID"
State when to use it"Search tool""Use this to search internal documentation; do not use for general web queries"
State constraints and side effects(omitted)"This action is irreversible and sends an email immediately"
Give units and formats for ambiguous parametersdate: stringdate: string, format YYYY-MM-DD, UTC

Parameter-level rules

Avoid duplicating parameter description in tool descriptionRepeats every parameter in proseTool description covers purpose only; parameter descriptions carry field-level detail

Handling errors and partial failures

A schema should define, or the tool's documentation should specify, a stable error contract: what the tool returns on invalid input, missing permissions, or downstream failure. Function-calling reliability depends heavily on the agent being able to parse the failure and decide whether to retry, ask the user, or choose a different tool.


{
  "name": "update_order_status",
  "description": "Updates the status of an existing order. Use only after confirming the order ID exists via get_order. This action cannot be undone.",
  "parameters": {
    "type": "object",
    "properties": {
      "order_id": {
        "type": "string",
        "description": "Unique order identifier, format ORD-XXXXX"
      },
      "new_status": {
        "type": "string",
        "enum": ["pending", "shipped", "cancelled", "refunded"],
        "description": "Target status. 'refunded' requires manager_approval to be true."
      },
      "manager_approval": {
        "type": "boolean",
        "description": "Required and must be true if new_status is 'refunded'; ignored otherwise. Default: false."
      }
    },
    "required": ["order_id", "new_status"]
  }
}

# Example error contract the tool implementation should honor consistently
def update_order_status(order_id, new_status, manager_approval=False):
    if new_status == "refunded" and not manager_approval:
        # Return a structured, parseable error rather than a raw exception
        return {"ok": False, "error": "APPROVAL_REQUIRED",
                "message": "Refund requires manager_approval=true"}
    if not order_exists(order_id):
        return {"ok": False, "error": "NOT_FOUND",
                "message": f"No order found with id {order_id}"}
    # ... perform update
    return {"ok": True, "order_id": order_id, "status": new_status}

Versioning and evolution

Testing schema reliability

Reliable function calling should be validated the same way as any interface contract:

FAQ

How many tools can an agent reliably choose between?

This depends on the model and how well-differentiated the tool descriptions are, but tool selection accuracy degrades noticeably past roughly 20-30 concurrently exposed tools unless a routing or categorization layer narrows the active set first.

Should tool descriptions include examples of correct usage?

A short example in the description can help disambiguate edge cases, but overly long descriptions with multiple examples consume context budget and can be counterproductive; one concise example is usually sufficient, reserved for genuinely ambiguous tools.

Is JSON Schema strictness enough to prevent bad tool calls?

No. Schema validation catches structurally invalid calls (wrong type, missing required field) but not semantically wrong calls (valid JSON with the wrong tool selected or a plausible but incorrect value). Semantic correctness requires clear, non-overlapping tool scopes and description quality in addition to schema strictness.

Related Articles