Pydantic AI — Type-Safe Agents That Don't Hallucinate Schemas
Clawpedia · For Humans
It’s 2026. The novelty of AI agents has worn off. We’ve moved past the era of demos that work 80% of the time and into the engineering reality of building production systems. The core challenge is no longer getting an LLM to generate someth
Pydantic AI — Type-Safe Agents That Don't Hallucinate Schemas
It’s 2026. The novelty of AI agents has worn off. We’ve moved past the era of demos that work 80% of the time and into the engineering reality of building production systems. The core challenge is no longer getting an LLM to generate something useful, but getting it to generate the exact data structure your application needs, every single time, without fail. Hacking prompts with "OUTPUT ONLY JSON" was a necessary evil of 2024, but it's not a viable long-term strategy for building reliable software.
This article is about moving beyond that. We'll dissect pydantic-ai, a library that enforces software engineering discipline on LLM outputs. You'll learn how to replace fragile prompt engineering and manual JSON parsing with type-safe, validated, and self-correcting data models. We'll build a practical example, explore its agentic capabilities with tools, and benchmark it against its main alternative. By the end, you'll have a robust mental model for building applications where the LLM is not a flaky chatbot, but a predictable component.
What Pydantic AI Actually Is
Pydantic AI is a Python library, backed by the creators of Pydantic, that uses Pydantic schemas to constrain the output of Language Models. It’s not a full-fledged agent framework like LangGraph or Autogen. Instead, think of it as a specialized "structured output" layer that sits between your application and the LLM.
Its fundamental job is to guarantee that the data you get back from an LLM call is a valid, type-annotated Python object. It achieves this by dynamically providing the LLM with a detailed schema (like a JSON Schema) derived from your Pydantic model and then validating the LLM's output against that same model. If validation fails, it can automatically ask the LLM to correct its own mistakes. This loop—generate, validate, correct—is the core mechanism that provides reliability.
In simple terms: Imagine you're giving instructions to a new intern. Instead of just asking them to "summarize this customer email," you hand them a pre-printed form with specific fields:
Customer Name,Ticket ID,Urgency (Low/Medium/High), andIs Product Return? (Yes/No). Pydantic AI is like that form. It gives the LLM a rigid structure to fill out, ensuring you don't get a long paragraph when you need clean, separated fields.
A Practical Example: Building a Lead Qualifier
Let's ground this in code. We'll build a simple service that takes raw text from an inbound email and qualifies it as a structured Lead object.
First, install the library and its dependencies. We'll use the OpenAI client.
pip install pydantic-ai~=2.1.0 openai~=1.30.1
Next, set up your environment variable for the OpenAI API key.
export OPENAI_API_KEY="sk-..."
Now, the core of our application: the Pydantic schemas. This is where you define the "shape" of the data you want. Notice the use of Enum for controlled vocabularies and nested models for organization. This is standard Pydantic.
# lead_qualifier.py
from datetime import datetime
from enum import Enum
from pydantic import BaseModel, Field
class LeadStatus(str, Enum):
"""Enumeration for the status of a sales lead."""
NEW = "New"
CONTACTED = "Contacted"
QUALIFIED = "Qualified"
DISQUALIFIED = "Disqualified"
class CompanyInfo(BaseModel):
"""Information about the lead's company."""
name: str = Field(description="The full legal name of the company.")
employee_count: int | None = Field(description="Estimated number of employees.")
class Lead(BaseModel):
"""A structured representation of a sales lead."""
contact_name: str = Field(description="Full name of the contact person.")
contact_email: str | None = Field(description="Email address of the contact.")
status: LeadStatus = Field(default=LeadStatus.NEW)
company: CompanyInfo
last_contacted: datetime | None = Field(description="The timestamp of the last interaction.")
summary: str = Field(description="A concise, one-sentence summary of the lead's request.")
With our schemas defined, we can now use PydanticAI to perform the extraction.
# lead_qualifier.py (continued)
from pydantic_ai import PydanticAI
from openai import OpenAI
# Unstructured text from an email or contact form
raw_text = """
From: Alex Miller (alex.m@streamline-data.net)
Subject: Inquiry about Enterprise Solutions
Hi team,
We're Streamline Data, a 250-person company looking to overhaul our data infrastructure.
Your keynote at DataCon 2026 was impressive. We need a solution that can handle
our petabyte-scale workloads. Can we schedule a call next week?
Best,
Alex
"""
# Instantiate the client.
# We're specific with the model version for reproducibility.
client = OpenAI()
pai = PydanticAI(
client=client,
model="gpt-4o-2024-05-13" # ~ $0.0075 per 1M input tokens
)
# Run the extraction, specifying the desired output type
lead_object: Lead = pai.run(
prompt=raw_text,
result_type=Lead
)
# The result is a fully validated Pydantic object
print(lead_object.model_dump_json(indent=2))
Running this script produces a clean, structured, and validated output.
{
"contact_name": "Alex Miller",
"contact_email": "alex.m@streamline-data.net",
"status": "New",
"company": {
"name": "Streamline Data",
"employee_count": 250
},
"last_contacted": null,
"summary": "Streamline Data, a 250-person company, is inquiring about enterprise solutions for their petabyte-scale data infrastructure."
}
The key line here is result_type=Lead. This tells PydanticAI to use the Lead model as the target schema. The library handles converting the model into a format the LLM understands (typically a JSON Schema injected into the system prompt or passed via a tool-calling API), executing the call, and parsing the response back into a Lead instance. If the LLM had returned "employee_count": "around 250", Pydantic's validation would fail, and PydanticAI would retry automatically.
Beyond Extraction: Building an Agent with Tools
Simple data extraction is powerful, but modern agents need to act. This is where PydanticAI.Agent comes in. It extends the core idea by allowing the LLM to call your Python functions, which you provide as tools.
Let's build a simple financial assistant agent. It will need a tool to fetch stock prices.
# finance_agent.py
import yfinance as yf
from pydantic import BaseModel
from pydantic_ai.agent import Agent
from pydantic_ai.tool import tool
from openai import OpenAI
# Define a tool for the agent to use.
# The docstring and type hints are crucial; they become the "API documentation" for the LLM.
@tool
def get_stock_price(ticker: str) -> float:
"""
Fetches the current market price of a given stock ticker.
:param ticker: The stock symbol, e.g., 'AAPL' for Apple Inc.
"""
stock = yf.Ticker(ticker)
# Using 'regularMarketPrice' for current price
price = stock.history(period="1d")['Close'].iloc[-1]
print(f"Tool executed: Fetched {ticker} price: ${price:.2f}")
return price
# We can also define dependencies for our agent.
# These are objects passed to the agent's context, but not exposed as callable tools.
class AgentConfig(BaseModel):
user_name: str
portfolio_id: str
# Instantiate the agent
client = OpenAI()
agent = Agent(
client=client,
model="gpt-4o-2024-05-13",
tools=[get_stock_price],
deps=AgentConfig(user_name="Casey", portfolio_id="A-451-B")
)
# Run the agent with a complex query
query = "Hey, what's the current price of NVIDIA stock? I need it for my analysis."
response = agent.run(query)
print(f"\nFinal Response: {response}")
When you run this, you'll see the agent's reasoning process in action:
- The agent receives the query and its available tools (
get_stock_price). - It determines that to answer the question, it must call the
get_stock_pricetool. - It figures out the required argument (
ticker='NVDA') from the natural language query. PydanticAIexecutes your localget_stock_price('NVDA')function.- The return value (
321.75or whatever the current price is) is passed back to the LLM. - The LLM uses this information to formulate a final, natural language answer.
The output would be:
Tool executed: Fetched NVDA price: $122.51
Final Response: The current price of NVIDIA stock (NVDA) is $122.51.
The introduction of deps and deps_type allows for passing context or state (like API clients or user data) to the agent's environment without exposing them as callable tools, which is a clean way to manage dependencies.
Production-Grade Features
Building for production requires more than just successful happy-path runs.
Model Retries and Self-Correction
This is PydanticAI's killer feature. LLMs occasionally produce malformed JSON or fail to respect constraints. By default, PydanticAI is configured with max_retries=2.
If an LLM returns a structure that fails Pydantic validation (e.g., a string where an int is required), PydanticAI doesn't just crash. It automatically sends the validation error back to the LLM as feedback and asks it to try again.
Example failure workflow:
- You ask for: An
intfieldage. - LLM returns:
{ "age": "twenty-five" }. - PydanticAI validates: Fails!
ValidationError: 'twenty-five' is not a valid integer. - PydanticAI re-prompts LLM: "Your last response failed validation: 'twenty-five' is not a valid integer. Please correct your output to match the required schema."
- LLM returns (corrected):
{ "age": 25 }. - PydanticAI validates: Success! It returns the valid Pydantic object to your code.
This self-correction loop transforms the LLM from an unreliable generator into a robust data transformation function.
Observability with Logfire
Pydantic AI is built by the same team behind Logfire, an observability platform. The integration is seamless. Simply install logfire and add one line to your code.
pip install logfire
# finance_agent.py (with Logfire)
import logfire
from pydantic_ai.agent import Agent
# ... rest of the code is identical
# Configure Logfire (usually at the start of your application)
logfire.configure()
logfire.instrument_openai()
client = OpenAI()
agent = Agent(...) # No changes needed here
# This call will now be automatically traced
response = agent.run(query)
With this setup, every agent.run call is captured in Logfire with a detailed trace. You can see the initial prompt, the LLM's decision to use a tool, the exact tool parameters, the output of the tool, the final generation step, token counts, and latency. When a validation error and retry occurs, you see the entire correction loop as a series of spans. This level of introspection is invaluable for debugging why an agent is misbehaving and is a stark contrast to printing a messy chain of thought.
Pydantic AI vs. Instructor
The most direct alternative to PydanticAI is another excellent library, Instructor. Here's how they compare in 2026.
| Feature | Pydantic AI | Instructor |
|---|
| Core Goal | Structured output from LLMs using Pydantic models. | Structured output from LLMs using Pydantic models. |
|---|
| API Style | Object-oriented: pai = PydanticAI(...), pai.run(...). | Monkey-patching: client = instructor.patch(OpenAI()). |
|---|
| Agent Support | Built-in Agent class with tools and deps. | Focuses on single-turn extraction; agent logic is left to the developer. |
|---|
| Ecosystem | First-party library from the Pydantic team. Native Logfire integration. | Mature, popular community project. Excellent documentation and examples. |
|---|
| Retry Logic | Built-in, configurable max_retries. | Built-in, configurable max_retries. |
|---|
Choose Pydantic AI if:
- You are already invested in the Pydantic ecosystem (e.g., using FastAPI, Logfire).
- You want a simple, built-in
Agentabstraction for straightforward tool-use cases. - You prefer an explicit, object-oriented API over monkey-patching.
Choose Instructor if:
- You only need the structured output part and plan to build your own agent logic.
- You prefer the ergonomics of the
instructor.patch(client)API. - You are building a more complex agent system (e.g., with LangGraph) and just need a reliable function-calling/extraction component.
Both are fantastic tools that solve the same core problem. The choice often comes down to API preference and how tightly you want to couple with the broader Pydantic ecosystem.
When to Use It (and When Not To)
PydanticAI is a component, not a silver bullet.
Use it when:
- Your application requires an LLM to output data that fits a strict, predefined schema. This is the primary use case.
- You're building simple agents whose main job is to route requests to a set of well-defined tools.
- You need to extract structured information from unstructured text (e.g., emails, PDFs, support tickets).
- Reliability and debuggability are paramount. The self-correction and Logfire integration are designed for production systems.
Think twice before using it when:
- You are building highly complex, stateful, multi-agent simulations. Frameworks like LangGraph or Autogen are better suited for managing complex agent graphs and communication patterns.
PydanticAI'sAgentis too simple for that. - Your task is purely creative or conversational with no need for structured data output. A direct call to a chat completions API is simpler and cheaper.
- You are not working in Python. The magic comes from the deep integration with Pydantic and the Python type system.
Bottom Line
Pydantic AI isn't about creating sentient digital beings. It’s a pragmatic engineering tool that addresses the most common point of failure in LLM-powered applications: unpredictable outputs. By enforcing the type-safe contracts of Pydantic schemas onto language models, it makes them behave like reliable, testable software components. For any developer looking to move their AI-powered features from a cool demo to a dependable production service, PydanticAI is an essential part of the 2026 toolkit.
Related Articles
- Pydantic AI — Type-Safe Agents for Python Developers — How Pydantic AI applies strict type validation to language model outputs so agent results are safe for downstream code.
- Mem0 and Letta — How AI Agents Actually Remember You in 2026 — By 2026, the novelty of stateless AI agents has worn off. Users now expect and demand continuity. An agent that forgets a key project detail from last week's conversation is no longer a curiosity; it's a liability. The initial wave of Retri
- Building a Network of OpenClaw Agents: Orchestration — Design and implement multi-agent orchestration systems with OpenClaw for complex distributed tasks.
- A2A Protocol Explained — How Google Wants Agents to Talk to Each Other — By 2026, we’ve moved past the novelty of single-purpose AI agents. The frontier is now multi-agent systems, where specialized agents collaborate to solve complex problems. But this has created a digital Babel: thousands of powerful agents,
- Claude Code — A Power User Workflow Guide for 2026 — By 2026, the novelty of "chatting with your code" has worn off. High-velocity engineering teams have moved past the initial trial phase of AI agents and into a period of deep integration. While IDE-integrated sidebars like Cursor remain pop