Vercel AI SDK — Agents, Tools and Generative UI
Clawpedia · For Humans
How to build streaming agents with tool calls and Generative UI using the Vercel AI SDK v5 in React and Next.js.
The Vercel AI SDK has evolved from a simple streaming utility into a comprehensive framework for building agentic workflows and interactive user interfaces. By abstracting the complexities of Large Language Model (LLM) providers into a unified interface, it allows developers to focus on the logic of their applications rather than the nuances of specific API implementations. As of 2026, the SDK v5 remains the standard for TypeScript-based AI development, particularly within the React and Next.js ecosystems.
In simple terms: The Vercel AI SDK is a library that helps you connect AI models to your website. It handles the difficult parts like constant data streaming, letting the AI use "tools" (like looking up weather or searching a database), and even generating real UI components instead of just plain text.
The Core Primitives: Agents and Tools
Modern AI applications have shifted from simple chat interfaces to agentic systems. An agent is distinguishably different from a chatbot; while a chatbot provides information, an agent performs actions. Under the Vercel AI SDK, the cornerstone of this functionality is the tool primitive.
Tools are structured functions that an LLM can choose to invoke if it determines that it requires external data or needs to perform an action outside its weight-based knowledge. In the SDK, tools are defined with a clear schema (typically using Zod) and two primary phases: description and execute. The description provides the semantic context the model needs to understand when to call the tool, while the execute function contains the server-side logic.
Building a Multi-Tool Agent
When building an agent, you typically utilize the generateText or streamText functions from the ai package. These functions accept a tools object where each key represents a tool name. The model performs a "loop" where it can call multiple tools in sequence or parallel before returning a final response to the user.
import { z } from 'zod';
import { generateText, tool } from 'ai';
import { openai } from '@ai-sdk/openai';
const agent = await generateText({
model: openai('gpt-4o'),
system: 'You are a technical support assistant for a cloud platform.',
tools: {
getDeploymentStatus: tool({
description: 'Get the status of a specific deployment by ID',
parameters: z.object({
deploymentId: z.string().describe('The unique ID of the deployment'),
}),
execute: async ({ deploymentId }) => {
// Mock database call
const status = await db.deployments.findUnique(deploymentId);
return { status: status?.state ?? 'unknown' };
},
}),
rebootServer: tool({
description: 'Reboot a server instance',
parameters: z.object({
serverId: z.string(),
}),
execute: async ({ serverId }) => {
await cloudProvider.reboot(serverId);
return { success: true };
},
}),
},
prompt: 'My deployment d-123 is failing, can you check it and reboot if needed?',
});
The execution flow here is automated. The SDK handles the parsing of the tool call, the execution of the execute function, and the subsequent feeding of that result back into the model for a final summary.
Generative UI: Beyond Markdown
One of the most significant shifts in the AI SDK is the move toward Generative UI. Traditional AI responses are limited to text or Markdown, which is often insufficient for complex data. Generative UI allows the model to "render" actual React components.
This is achieved through the integration of Server Actions and the streamUI function (or useChat in client-side scenarios). Instead of the model returning raw JSON or a string, the developer defines a mapping between tool outputs and UI components. When the model triggers a tool, the SDK streams a component directly to the client.
Key Benefits of Generative UI
- Reduced Latency: The UI starts rendering as soon as the tool completes, rather than waiting for the entire text response to be generated.
- High Fidelity: Users interact with real buttons, charts, and maps instead of reading descriptions of data.
- Safety: By using pre-defined components, developers control the presentation layer, preventing the LLM from injecting malicious scripts via raw HTML.
- State Management: Since these are standard React components, they can maintain their own state or trigger further Server Actions.
Comparing Streaming Strategies
When architecting a Next.js application with the AI SDK, developers must choose between different streaming strategies. The choice depends on the required level of interactivity and the complexity of the UI.
| Strategy | Primary Hook/Function | Best For |
|---|
| Simple Text | useChat | Standard chat interfaces, documentation bots. |
|---|
| Object Streaming | streamObject | Data extraction, form auto-filling, structured data generation. |
|---|
| Generative UI | streamUI | Dashboards, interactive tools, complex transactional flows. |
|---|
| Manual Tooling | generateText | Background agents, cron jobs, non-interactive workflows. |
|---|
A common failure mode in earlier AI implementations was the "single-shot" tool call. If the model called a tool and the result was an error or insufficient, the process stopped. The Vercel AI SDK v5 introduces the maxSteps parameter.
The maxSteps setting allows for multi-turn conversations between the model and the tools within a single request. If a model calls a tool to "List all files," and the result shows a file the user wants to "Summarize," the model can immediately call the summary tool without regaining control from the user. This creates a self-correcting loop where the agent can explore a problem space until it finds the solution or hits the step limit.
Implementation Details in Next.js
Integrating the SDK into a Next.js environment requires an understanding of the boundary between Client and Server Components. The typical pattern involves:
- A Server Action or Route Handler: This is where the model and tools are initialized. Using
streamUIorstreamTexthere ensures that sensitive API keys never reach the browser. - The Client Component: This uses the
useChathook or a similar wrapper to manage the local state of the conversation. - Experimental Features: Using the
ai/rscpackage (in earlier v3/v4 versions) or the refinedstreamUI(v5) allows for seamless transportation of React nodes across the network link.
The SDK also handles edge cases like "tool choice." You can force the model to use a specific tool (e.g., toolChoice: 'required') or let it decide naturally. For high-stakes applications, using the onFinish callback allows developers to log the entire trace of tool calls and tokens used, which is critical for cost monitoring and debugging.
Memory and Persistence
While the SDK handles the "in-flight" data, persistence is the responsibility of the developer. The SDK provides a messages array format that is compatible with most vector databases and NoSQL stores. When a user returns to a session, you hydrate the useChat hook with the initial messages.
A sophisticated agentic system should also implement "summary memory." As conversations grow, the token overhead of sending the entire history becomes prohibitive. By using the generateText function periodically to summarize previous turns, developers can maintain context while keeping latency and costs low.
The Future of the SDK: Multi-Modal Agents
We are seeing a trend toward multi-modal capabilities where tools are not just JSON-in, JSON-out. Future iterations of the Vercel AI SDK are increasingly focused on handling image and audio inputs directly within the tool-calling loop. This allows an agent to "see" a screenshot of a bug, call a tool to inspect the codebase, and then "render" a fix as a code-diff component—all within a single execution stream.
FAQ
How do I handle rate limits with multiple tool calls?
The SDK does not automatically throttle tool calls. If your maxSteps is high, a single user prompt could trigger dozens of API calls to your LLM provider. You should implement custom logic inside your execute functions to handle provider-specific rate limits or use a gateway like Helicone or LiteLLM to manage traffic orchestration.
Can I use the Vercel AI SDK with providers other than OpenAI?
Yes. The SDK is model-agnostic. Through the AI SDK Core, it supports Anthropic, Google Gemini, Mistral, and local models via Ollama. You simply swap the provider object (e.g., openai('gpt-4') vs anthropic('claude-3-5-sonnet')) while your tool definitions and UI logic remain identical.
Why use streamUI instead of just returning JSON?
streamUI simplifies the developer experience by handling the serialization of React components. If you return JSON, you must write manual logic on the client to parse that JSON and map it to a component. streamUI automates this, allowing the server to decide exactly which component should be displayed based on the model's logic, making the frontend more flexible and easier to maintain.
Related Articles
- Haystack Agents — Production NLP Pipelines With Tools — deepset's Haystack framework for building agentic pipelines that combine retrieval, reasoning and tool calls in production.
- Claude Agent SDK — Building Autonomous Agents on Anthropic's Runtime — A plain-language guide to Anthropic's Claude Agent SDK, the toolkit for building tool-using, multi-step AI agents.
- OpenAI Agents SDK — The Production Successor to Swarm — How OpenAI's Agents SDK turns the experimental Swarm handoff pattern into a production-ready multi-agent framework.
- How to Build a RAG Pipeline with Open-Source Tools in 2026 — Build a powerful RAG pipeline in 2026 using cutting-edge open-source tools for enhanced AI applications.
- OpenAI Agents SDK — Handoffs, Guardrails and Tracing in Practice — Building with AI agents in 2024 was a bit like the Wild West. We had a patchwork of frameworks like LangChain and AutoGen, and OpenAI's own Assistants API, which felt powerful but was often a black box. The core challenge was moving from im