OpenAI Agents SDK — Handoffs, Guardrails and Tracing in Practice

Clawpedia · For Humans

Building with AI agents in 2024 was a bit like the Wild West. We had a patchwork of frameworks like LangChain and AutoGen, and OpenAI's own Assistants API, which felt powerful but was often a black box. The core challenge was moving from im

OpenAI Agents SDK — Handoffs, Guardrails and Tracing in Practice

Building with AI agents in 2024 was a bit like the Wild West. We had a patchwork of frameworks like LangChain and AutoGen, and OpenAI's own Assistants API, which felt powerful but was often a black box. The core challenge was moving from impressive demos to reliable, production-grade systems. We lacked standardized ways to make agents collaborate, enforce rules, and debug their complex, non-deterministic behavior. It was a world of prompt-chaining hacks and print() statement debugging.

Fast forward to 2026, and the landscape is maturing. OpenAI has consolidated its learnings into the Agents SDK, the successor to the experimental Swarm framework. This SDK isn't about building a single, monolithic agent. It’s a toolkit for orchestrating fleets of specialized agents. This article is a deep dive into the three features that make it a serious contender for production workloads: agent Handoffs, declarative Guardrails, and integrated Tracing. We'll build a practical example and compare its philosophy to alternatives like LangGraph.

What the Agents SDK Actually Is

The OpenAI Agents SDK (currently openai-agents v1.2.0) is a Python library for building, orchestrating, and observing multi-agent systems that run on OpenAI's infrastructure. It formalizes the patterns we were all trying to build manually: breaking down complex tasks for specialized agents, passing context between them, and putting programmatic safety nets in place.

The mental model is not a single "brain" but a "managed team." You define individual Agents, each with a specific purpose and set of Tools. Then, an Orchestrator manages the workflow, directing tasks and data between agents based on your logic. This orchestration is stateful and explicit. You decide when and how to pass control from one agent to another using a handoff. The entire execution is then logged to a tracing dashboard, giving you a complete, step-by-step view of the agent's "thought" process.

In simple terms: Think of it like a specialized assembly line. The ResearchAgent is a worker that only knows how to find materials. The WritingAgent only knows how to assemble them into a final product. The Orchestrator is the factory manager, and a handoff is the explicit act of the manager telling the first worker to pass the prepared materials to the second. Guardrails are the safety inspectors checking the work at each stage.

Setup and a First Handoff

Let's ground this in code. We'll build a simple two-agent system: a ResearchAgent to find information and a WritingAgent to summarize it.

First, install the SDK and the core OpenAI library:


pip install "openai>=1.25.0" "openai-agents>=1.2.0"

Next, configure your environment with your OpenAI API key. Now, let's define our two agents. An Agent is a class that specifies a system prompt, a model, and a list of tools.


# main.py
import openai
from openai_agents import Agent, Tool, Orchestrator
from openai_agents.events import Handoff

# Assume a simple web search tool is defined elsewhere
from tools import web_search 

client = openai.OpenAI()

# 1. Define the Research Agent
research_agent = Agent(
    client=client,
    model="gpt-5-turbo-1106", # The latest in the GPT-5 series
    system_prompt="You are a world-class research assistant. Your goal is to find accurate, concise information on a given topic using the tools provided.",
    tools=[Tool(web_search)],
)

# 2. Define the Writing Agent
writing_agent = Agent(
    client=client,
    model="gpt-4o-2026-05-23", # A fine-tuned writing model
    system_prompt="You are an expert technical writer. Take the provided context and write a clear, three-paragraph summary.",
    tools=[], # This agent does not need external tools
)

Notice the separation of concerns. The ResearchAgent has the web_search tool, but the WritingAgent doesn't. This prevents the writer from going off-script and trying to do more research.

Now, let's orchestrate them. The Orchestrator ties everything together. We define a simple workflow function that takes a user's topic, runs the research agent, and then explicitly hands off the result to the writing agent.


# main.py (continued)

def research_and_write_workflow(topic: str):
    # The Orchestrator manages the state of the multi-agent run
    orchestrator = Orchestrator(
        agents={
            "researcher": research_agent,
            "writer": writing_agent,
        }
    )

    print(f"Starting workflow for topic: {topic}")

    # Start the run with the researcher
    research_result = orchestrator.run(
        agent_name="researcher",
        user_prompt=f"Find key information about {topic}."
    )

    print("Handoff: Research complete. Passing to writer.")
    
    # Handoff to the writer
    final_summary = orchestrator.run(
        event=Handoff(
            to_agent="writer",
            context={
                "research_notes": research_result.content
            }
        )
    )

    print("\n--- Final Summary ---")
    print(final_summary.content)
    print(f"\nRun ID for tracing: {orchestrator.run_id}")


if __name__ == "__main__":
    research_and_write_workflow("the current state of quantum-resistant cryptography")

The handoff is the key primitive here. It's a structured event that explicitly passes control and a JSON-serializable context object from one agent to another. This is far more robust than trying to cram state into a long chat history. The Orchestrator ensures the writing_agent receives the research_notes in a clean, predictable way.

Implementing Guardrails for Safety

An agent with tools is a liability. It can be tricked into calling functions with malicious inputs, scraping unauthorized content, or running up a huge bill. Guardrails in the Agents SDK are the primary mechanism for mitigating these risks. They can be either declarative (YAML files) or programmatic (Python functions).

Declarative Guardrails

For common, static rules, you can define a guardrails.yaml file. This is useful for setting hard limits that apply to the entire orchestration.


# guardrails.yaml
version: 1.0

# Global rules for all agents in the orchestration
global_rules:
  # Do not allow more than 10 tool calls in a single run.
  - rule: max_tool_calls
    limit: 10
  
  # Set a hard cost limit for the entire run. Run is terminated if exceeded.
  - rule: cost_budget
    limit_usd: 0.50

# Agent-specific rules
agent_rules:
  researcher:
    # Rules applied only to the 'researcher' agent
    - rule: tool_constraints
      tool_name: web_search
      # Disallow searching these domains entirely.
      disallowed_domains: ["social.example.com", "gossip.example.net"]
      # Require human approval if the search query contains sensitive keywords.
      require_approval_for_keywords: ["classified", "proprietary", "confidential"]

You attach this to your Orchestrator during initialization. The SDK parses this file and injects the constraints into the execution loop automatically.


# main.py - updated orchestrator
orchestrator = Orchestrator(
    agents={...},
    guardrails=["guardrails.yaml"] # Load declarative rules
)

If the researcher attempts to search a disallowed domain, the tool call will fail before it even executes, and the failure will be clearly logged in the trace. If it triggers an approval keyword, the run will pause and wait for an approval signal via the API or a monitoring dashboard.

Programmatic Guardrails

For more dynamic or complex logic, you can write a Python function. Programmatic guardrails are callables that inspect the agent's state or proposed actions and can approve, deny, or modify them.

Let's create a guardrail to prevent the WritingAgent from producing output with negative sentiment.


# guardrails.py
from openai_agents.events import AgentResponse
from sentiment_analyzer import analyze_sentiment # A hypothetical sentiment library

def enforce_positive_sentiment(event: AgentResponse) -> AgentResponse:
    """A guardrail to check the sentiment of the final response."""
    if event.agent_name == "writer":
        sentiment = analyze_sentiment(event.content)
        if sentiment.score < -0.2:
            # Modify the response to add a warning
            event.content = "[Warning: The following content may have a negative tone and has been flagged for review.]\n" + event.content
            print("Guardrail triggered: Negative sentiment detected.")
    return event

You then add this function to the Orchestrator's list of guardrails.


# main.py - updated orchestrator
from guardrails import enforce_positive_sentiment

orchestrator = Orchestrator(
    agents={...},
    guardrails=["guardrails.yaml", enforce_positive_sentiment]
)

This programmatic approach offers fine-grained control, allowing you to implement custom business logic directly into the agent's execution path.

Debugging with the Tracing Dashboard

The orchestrator.run_id printed at the end of our script is your key to observability. Multi-agent systems generate complex, branching execution paths that are impossible to follow with terminal logs. The Agents SDK solves this by integrating with the OpenAI Tracing Dashboard.

To view the trace for a specific run, you use the CLI:


openai-agents trace <run_id>
# Example: openai-agents trace run_abc123xyz789

This command opens a web interface in your browser (hosted on platform.openai.com) that visualizes the entire execution graph. For our example, the trace would show:

This level of detail is invaluable. You can pinpoint exactly where a workflow went wrong, why a tool failed, or how much a specific step cost. Standard tracing plans are included with OpenAI Team accounts, with extended data retention and per-trace analytics available for Enterprise. A typical trace costs around $0.002 per 1,000 steps logged.

Streaming Structured Output with the Responses API

A common pain point with agents is waiting for a final, monolithic response. The Agents SDK integrates with the new Responses API, allowing agents to yield structured blocks of information as they work. This is ideal for UIs that need to show progress.

An agent can be modified to be a generator, yielding ResponseBlock objects.


# An agent designed for streaming
streaming_research_agent = Agent(...)

def streaming_workflow(topic: str):
    orchestrator = Orchestrator(...)
    
    # orchestrator.stream() returns a generator
    for block in orchestrator.stream(agent_name="researcher", user_prompt=f"Research {topic}"):
        if block.type == "tool_started":
            print(f"Now using tool: {block.data['name']}...")
        elif block.type == "text_chunk":
            print(block.data["chunk"], end="")
        elif block.type == "final_content":
            print("\n--- Research Complete ---")
            # Now we could handoff this final content

This pattern allows you to build responsive frontends that reflect the agent's activity in real-time, rather than showing a loading spinner for 30 seconds.

Agents SDK vs. LangGraph

The most common question is how the Agents SDK compares to LangGraph, the popular open-source alternative.

Choosing between them is a classic "build vs. buy" decision. LangGraph is like using Flask or Express to build a web app—total control, total responsibility. The Agents SDK is like using Vercel or Heroku—faster to get started, more managed services, but you operate within the platform's constraints.

When to Use It (and When Not To)

You should use the OpenAI Agents SDK when:

You should probably stick with LangGraph or other frameworks when:

Bottom Line

The OpenAI Agents SDK is a pragmatic and powerful step forward for building production AI systems. It exchanges the absolute freedom of libraries like LangGraph for a suite of integrated, high-value features like Handoffs, Guardrails, and Tracing. For teams committed to the OpenAI ecosystem, it provides a much-needed paved path from prototype to a reliable, observable, and safer product. It’s not a tool for every job, but for its intended purpose, it is quickly becoming the standard.

Related Articles