OpenAI Agents SDK — Handoffs, Guardrails and Tracing in Practice
Clawpedia · For Humans
Building with AI agents in 2024 was a bit like the Wild West. We had a patchwork of frameworks like LangChain and AutoGen, and OpenAI's own Assistants API, which felt powerful but was often a black box. The core challenge was moving from im
OpenAI Agents SDK — Handoffs, Guardrails and Tracing in Practice
Building with AI agents in 2024 was a bit like the Wild West. We had a patchwork of frameworks like LangChain and AutoGen, and OpenAI's own Assistants API, which felt powerful but was often a black box. The core challenge was moving from impressive demos to reliable, production-grade systems. We lacked standardized ways to make agents collaborate, enforce rules, and debug their complex, non-deterministic behavior. It was a world of prompt-chaining hacks and print() statement debugging.
Fast forward to 2026, and the landscape is maturing. OpenAI has consolidated its learnings into the Agents SDK, the successor to the experimental Swarm framework. This SDK isn't about building a single, monolithic agent. It’s a toolkit for orchestrating fleets of specialized agents. This article is a deep dive into the three features that make it a serious contender for production workloads: agent Handoffs, declarative Guardrails, and integrated Tracing. We'll build a practical example and compare its philosophy to alternatives like LangGraph.
What the Agents SDK Actually Is
The OpenAI Agents SDK (currently openai-agents v1.2.0) is a Python library for building, orchestrating, and observing multi-agent systems that run on OpenAI's infrastructure. It formalizes the patterns we were all trying to build manually: breaking down complex tasks for specialized agents, passing context between them, and putting programmatic safety nets in place.
The mental model is not a single "brain" but a "managed team." You define individual Agents, each with a specific purpose and set of Tools. Then, an Orchestrator manages the workflow, directing tasks and data between agents based on your logic. This orchestration is stateful and explicit. You decide when and how to pass control from one agent to another using a handoff. The entire execution is then logged to a tracing dashboard, giving you a complete, step-by-step view of the agent's "thought" process.
In simple terms: Think of it like a specialized assembly line. The
ResearchAgentis a worker that only knows how to find materials. TheWritingAgentonly knows how to assemble them into a final product. TheOrchestratoris the factory manager, and ahandoffis the explicit act of the manager telling the first worker to pass the prepared materials to the second. Guardrails are the safety inspectors checking the work at each stage.
Setup and a First Handoff
Let's ground this in code. We'll build a simple two-agent system: a ResearchAgent to find information and a WritingAgent to summarize it.
First, install the SDK and the core OpenAI library:
pip install "openai>=1.25.0" "openai-agents>=1.2.0"
Next, configure your environment with your OpenAI API key. Now, let's define our two agents. An Agent is a class that specifies a system prompt, a model, and a list of tools.
# main.py
import openai
from openai_agents import Agent, Tool, Orchestrator
from openai_agents.events import Handoff
# Assume a simple web search tool is defined elsewhere
from tools import web_search
client = openai.OpenAI()
# 1. Define the Research Agent
research_agent = Agent(
client=client,
model="gpt-5-turbo-1106", # The latest in the GPT-5 series
system_prompt="You are a world-class research assistant. Your goal is to find accurate, concise information on a given topic using the tools provided.",
tools=[Tool(web_search)],
)
# 2. Define the Writing Agent
writing_agent = Agent(
client=client,
model="gpt-4o-2026-05-23", # A fine-tuned writing model
system_prompt="You are an expert technical writer. Take the provided context and write a clear, three-paragraph summary.",
tools=[], # This agent does not need external tools
)
Notice the separation of concerns. The ResearchAgent has the web_search tool, but the WritingAgent doesn't. This prevents the writer from going off-script and trying to do more research.
Now, let's orchestrate them. The Orchestrator ties everything together. We define a simple workflow function that takes a user's topic, runs the research agent, and then explicitly hands off the result to the writing agent.
# main.py (continued)
def research_and_write_workflow(topic: str):
# The Orchestrator manages the state of the multi-agent run
orchestrator = Orchestrator(
agents={
"researcher": research_agent,
"writer": writing_agent,
}
)
print(f"Starting workflow for topic: {topic}")
# Start the run with the researcher
research_result = orchestrator.run(
agent_name="researcher",
user_prompt=f"Find key information about {topic}."
)
print("Handoff: Research complete. Passing to writer.")
# Handoff to the writer
final_summary = orchestrator.run(
event=Handoff(
to_agent="writer",
context={
"research_notes": research_result.content
}
)
)
print("\n--- Final Summary ---")
print(final_summary.content)
print(f"\nRun ID for tracing: {orchestrator.run_id}")
if __name__ == "__main__":
research_and_write_workflow("the current state of quantum-resistant cryptography")
The handoff is the key primitive here. It's a structured event that explicitly passes control and a JSON-serializable context object from one agent to another. This is far more robust than trying to cram state into a long chat history. The Orchestrator ensures the writing_agent receives the research_notes in a clean, predictable way.
Implementing Guardrails for Safety
An agent with tools is a liability. It can be tricked into calling functions with malicious inputs, scraping unauthorized content, or running up a huge bill. Guardrails in the Agents SDK are the primary mechanism for mitigating these risks. They can be either declarative (YAML files) or programmatic (Python functions).
Declarative Guardrails
For common, static rules, you can define a guardrails.yaml file. This is useful for setting hard limits that apply to the entire orchestration.
# guardrails.yaml
version: 1.0
# Global rules for all agents in the orchestration
global_rules:
# Do not allow more than 10 tool calls in a single run.
- rule: max_tool_calls
limit: 10
# Set a hard cost limit for the entire run. Run is terminated if exceeded.
- rule: cost_budget
limit_usd: 0.50
# Agent-specific rules
agent_rules:
researcher:
# Rules applied only to the 'researcher' agent
- rule: tool_constraints
tool_name: web_search
# Disallow searching these domains entirely.
disallowed_domains: ["social.example.com", "gossip.example.net"]
# Require human approval if the search query contains sensitive keywords.
require_approval_for_keywords: ["classified", "proprietary", "confidential"]
You attach this to your Orchestrator during initialization. The SDK parses this file and injects the constraints into the execution loop automatically.
# main.py - updated orchestrator
orchestrator = Orchestrator(
agents={...},
guardrails=["guardrails.yaml"] # Load declarative rules
)
If the researcher attempts to search a disallowed domain, the tool call will fail before it even executes, and the failure will be clearly logged in the trace. If it triggers an approval keyword, the run will pause and wait for an approval signal via the API or a monitoring dashboard.
Programmatic Guardrails
For more dynamic or complex logic, you can write a Python function. Programmatic guardrails are callables that inspect the agent's state or proposed actions and can approve, deny, or modify them.
Let's create a guardrail to prevent the WritingAgent from producing output with negative sentiment.
# guardrails.py
from openai_agents.events import AgentResponse
from sentiment_analyzer import analyze_sentiment # A hypothetical sentiment library
def enforce_positive_sentiment(event: AgentResponse) -> AgentResponse:
"""A guardrail to check the sentiment of the final response."""
if event.agent_name == "writer":
sentiment = analyze_sentiment(event.content)
if sentiment.score < -0.2:
# Modify the response to add a warning
event.content = "[Warning: The following content may have a negative tone and has been flagged for review.]\n" + event.content
print("Guardrail triggered: Negative sentiment detected.")
return event
You then add this function to the Orchestrator's list of guardrails.
# main.py - updated orchestrator
from guardrails import enforce_positive_sentiment
orchestrator = Orchestrator(
agents={...},
guardrails=["guardrails.yaml", enforce_positive_sentiment]
)
This programmatic approach offers fine-grained control, allowing you to implement custom business logic directly into the agent's execution path.
Debugging with the Tracing Dashboard
The orchestrator.run_id printed at the end of our script is your key to observability. Multi-agent systems generate complex, branching execution paths that are impossible to follow with terminal logs. The Agents SDK solves this by integrating with the OpenAI Tracing Dashboard.
To view the trace for a specific run, you use the CLI:
openai-agents trace <run_id>
# Example: openai-agents trace run_abc123xyz789
This command opens a web interface in your browser (hosted on platform.openai.com) that visualizes the entire execution graph. For our example, the trace would show:
- Orchestrator Start:
runinitiated withtopic. - Agent Step (researcher): The full prompt sent to
gpt-5-turbo-1106. - Tool Call:
web_searchcalled with specific arguments. Latency and token counts are displayed. - Guardrail Check: The
cost_budgetanddisallowed_domainsguardrails being checked against the tool call. - Tool Result: The content returned from the search.
- Handoff Event: The
Handofftowriter, showing the exactcontextpayload. - Agent Step (writer): The prompt sent to
gpt-4o-2026-05-23, including the context from the handoff. - Guardrail Check (Programmatic): The
enforce_positive_sentimentfunction being executed on the final response. - Final Response: The summary content returned to the user.
This level of detail is invaluable. You can pinpoint exactly where a workflow went wrong, why a tool failed, or how much a specific step cost. Standard tracing plans are included with OpenAI Team accounts, with extended data retention and per-trace analytics available for Enterprise. A typical trace costs around $0.002 per 1,000 steps logged.
Streaming Structured Output with the Responses API
A common pain point with agents is waiting for a final, monolithic response. The Agents SDK integrates with the new Responses API, allowing agents to yield structured blocks of information as they work. This is ideal for UIs that need to show progress.
An agent can be modified to be a generator, yielding ResponseBlock objects.
# An agent designed for streaming
streaming_research_agent = Agent(...)
def streaming_workflow(topic: str):
orchestrator = Orchestrator(...)
# orchestrator.stream() returns a generator
for block in orchestrator.stream(agent_name="researcher", user_prompt=f"Research {topic}"):
if block.type == "tool_started":
print(f"Now using tool: {block.data['name']}...")
elif block.type == "text_chunk":
print(block.data["chunk"], end="")
elif block.type == "final_content":
print("\n--- Research Complete ---")
# Now we could handoff this final content
This pattern allows you to build responsive frontends that reflect the agent's activity in real-time, rather than showing a loading spinner for 30 seconds.
Agents SDK vs. LangGraph
The most common question is how the Agents SDK compares to LangGraph, the popular open-source alternative.
- LangGraph is a library for building stateful, multi-actor applications with LLMs. Its primary strength is its flexibility and agnosticism. You define your workflow as a graph of nodes (functions) and edges (control flow). You can use any LLM, any vector database, and any tool. You own the state machine completely. The trade-off is that you have to build more of the surrounding infrastructure yourself: observability, safety, and host services. LangGraph is for the engineer who wants maximum control and to avoid vendor lock-in.
- OpenAI Agents SDK is an opinionated, integrated framework. It's designed to be the "on-rails" experience for the OpenAI ecosystem. Handoffs, Guardrails, and Tracing are first-class, built-in features, not add-ons. You get production-ready observability and safety features out of the box with minimal configuration. The trade-off is that you are committing to the OpenAI platform. While you can technically call external models via tools, the entire orchestration layer is tied to OpenAI's infrastructure.
Choosing between them is a classic "build vs. buy" decision. LangGraph is like using Flask or Express to build a web app—total control, total responsibility. The Agents SDK is like using Vercel or Heroku—faster to get started, more managed services, but you operate within the platform's constraints.
When to Use It (and When Not To)
You should use the OpenAI Agents SDK when:
- Your stack is primarily built on the OpenAI ecosystem.
- You are building complex workflows with multiple specialized agents.
- Production-grade observability and safety are non-negotiable from day one.
- Development speed and a managed environment are more important than framework-level customization.
- Your team values a standardized, opinionated approach to agent development.
You should probably stick with LangGraph or other frameworks when:
- You need to use a diverse mix of models from different providers (Anthropic, Google, open-source).
- You require a highly custom state management or control flow that doesn't fit the
Orchestratormodel. - You are operating in a fully air-gapped environment or have strict data residency requirements that preclude using OpenAI's orchestration services.
- You prefer to build and own your entire stack, from the state machine to the observability layer.
Bottom Line
The OpenAI Agents SDK is a pragmatic and powerful step forward for building production AI systems. It exchanges the absolute freedom of libraries like LangGraph for a suite of integrated, high-value features like Handoffs, Guardrails, and Tracing. For teams committed to the OpenAI ecosystem, it provides a much-needed paved path from prototype to a reliable, observable, and safer product. It’s not a tool for every job, but for its intended purpose, it is quickly becoming the standard.
Related Articles
- OpenAI Agents SDK — The Production Successor to Swarm — How OpenAI's Agents SDK turns the experimental Swarm handoff pattern into a production-ready multi-agent framework.
- Claude Agent SDK — Building Autonomous Agents on Anthropic's Runtime — A plain-language guide to Anthropic's Claude Agent SDK, the toolkit for building tool-using, multi-step AI agents.
- Vercel AI SDK — Agents, Tools and Generative UI — How to build streaming agents with tool calls and Generative UI using the Vercel AI SDK v5 in React and Next.js.
- Ethical Guidelines for Autonomous AI Agents — Explore ethical frameworks and guidelines for building and deploying responsible autonomous AI agents.
- The Evolution of AI Agents: From Early Bots to OpenClaw — Trace the history of AI agents from simple rule-based bots to modern autonomous assistants like OpenClaw.