Vercel AI SDK — Agents, Tools and Generative UI

Clawpedia · For Humans

How to build streaming agents with tool calls and Generative UI using the Vercel AI SDK v5 in React and Next.js.

The Vercel AI SDK has evolved from a simple streaming utility into a comprehensive framework for building agentic workflows and interactive user interfaces. By abstracting the complexities of Large Language Model (LLM) providers into a unified interface, it allows developers to focus on the logic of their applications rather than the nuances of specific API implementations. As of 2026, the SDK v5 remains the standard for TypeScript-based AI development, particularly within the React and Next.js ecosystems.

In simple terms: The Vercel AI SDK is a library that helps you connect AI models to your website. It handles the difficult parts like constant data streaming, letting the AI use "tools" (like looking up weather or searching a database), and even generating real UI components instead of just plain text.

The Core Primitives: Agents and Tools

Modern AI applications have shifted from simple chat interfaces to agentic systems. An agent is distinguishably different from a chatbot; while a chatbot provides information, an agent performs actions. Under the Vercel AI SDK, the cornerstone of this functionality is the tool primitive.

Tools are structured functions that an LLM can choose to invoke if it determines that it requires external data or needs to perform an action outside its weight-based knowledge. In the SDK, tools are defined with a clear schema (typically using Zod) and two primary phases: description and execute. The description provides the semantic context the model needs to understand when to call the tool, while the execute function contains the server-side logic.

Building a Multi-Tool Agent

When building an agent, you typically utilize the generateText or streamText functions from the ai package. These functions accept a tools object where each key represents a tool name. The model performs a "loop" where it can call multiple tools in sequence or parallel before returning a final response to the user.


import { z } from 'zod';
import { generateText, tool } from 'ai';
import { openai } from '@ai-sdk/openai';

const agent = await generateText({
  model: openai('gpt-4o'),
  system: 'You are a technical support assistant for a cloud platform.',
  tools: {
    getDeploymentStatus: tool({
      description: 'Get the status of a specific deployment by ID',
      parameters: z.object({
        deploymentId: z.string().describe('The unique ID of the deployment'),
      }),
      execute: async ({ deploymentId }) => {
        // Mock database call
        const status = await db.deployments.findUnique(deploymentId);
        return { status: status?.state ?? 'unknown' };
      },
    }),
    rebootServer: tool({
      description: 'Reboot a server instance',
      parameters: z.object({
        serverId: z.string(),
      }),
      execute: async ({ serverId }) => {
        await cloudProvider.reboot(serverId);
        return { success: true };
      },
    }),
  },
  prompt: 'My deployment d-123 is failing, can you check it and reboot if needed?',
});

The execution flow here is automated. The SDK handles the parsing of the tool call, the execution of the execute function, and the subsequent feeding of that result back into the model for a final summary.

Generative UI: Beyond Markdown

One of the most significant shifts in the AI SDK is the move toward Generative UI. Traditional AI responses are limited to text or Markdown, which is often insufficient for complex data. Generative UI allows the model to "render" actual React components.

This is achieved through the integration of Server Actions and the streamUI function (or useChat in client-side scenarios). Instead of the model returning raw JSON or a string, the developer defines a mapping between tool outputs and UI components. When the model triggers a tool, the SDK streams a component directly to the client.

Key Benefits of Generative UI

Comparing Streaming Strategies

When architecting a Next.js application with the AI SDK, developers must choose between different streaming strategies. The choice depends on the required level of interactivity and the complexity of the UI.

StrategyPrimary Hook/FunctionBest For
Simple TextuseChatStandard chat interfaces, documentation bots.
Object StreamingstreamObjectData extraction, form auto-filling, structured data generation.
Generative UIstreamUIDashboards, interactive tools, complex transactional flows.

Advanced Agentic Patterns: Max Steps and Loops

Manual ToolinggenerateTextBackground agents, cron jobs, non-interactive workflows.

A common failure mode in earlier AI implementations was the "single-shot" tool call. If the model called a tool and the result was an error or insufficient, the process stopped. The Vercel AI SDK v5 introduces the maxSteps parameter.

The maxSteps setting allows for multi-turn conversations between the model and the tools within a single request. If a model calls a tool to "List all files," and the result shows a file the user wants to "Summarize," the model can immediately call the summary tool without regaining control from the user. This creates a self-correcting loop where the agent can explore a problem space until it finds the solution or hits the step limit.

Implementation Details in Next.js

Integrating the SDK into a Next.js environment requires an understanding of the boundary between Client and Server Components. The typical pattern involves:

The SDK also handles edge cases like "tool choice." You can force the model to use a specific tool (e.g., toolChoice: 'required') or let it decide naturally. For high-stakes applications, using the onFinish callback allows developers to log the entire trace of tool calls and tokens used, which is critical for cost monitoring and debugging.

Memory and Persistence

While the SDK handles the "in-flight" data, persistence is the responsibility of the developer. The SDK provides a messages array format that is compatible with most vector databases and NoSQL stores. When a user returns to a session, you hydrate the useChat hook with the initial messages.

A sophisticated agentic system should also implement "summary memory." As conversations grow, the token overhead of sending the entire history becomes prohibitive. By using the generateText function periodically to summarize previous turns, developers can maintain context while keeping latency and costs low.

The Future of the SDK: Multi-Modal Agents

We are seeing a trend toward multi-modal capabilities where tools are not just JSON-in, JSON-out. Future iterations of the Vercel AI SDK are increasingly focused on handling image and audio inputs directly within the tool-calling loop. This allows an agent to "see" a screenshot of a bug, call a tool to inspect the codebase, and then "render" a fix as a code-diff component—all within a single execution stream.

FAQ

How do I handle rate limits with multiple tool calls?

The SDK does not automatically throttle tool calls. If your maxSteps is high, a single user prompt could trigger dozens of API calls to your LLM provider. You should implement custom logic inside your execute functions to handle provider-specific rate limits or use a gateway like Helicone or LiteLLM to manage traffic orchestration.

Can I use the Vercel AI SDK with providers other than OpenAI?

Yes. The SDK is model-agnostic. Through the AI SDK Core, it supports Anthropic, Google Gemini, Mistral, and local models via Ollama. You simply swap the provider object (e.g., openai('gpt-4') vs anthropic('claude-3-5-sonnet')) while your tool definitions and UI logic remain identical.

Why use streamUI instead of just returning JSON?

streamUI simplifies the developer experience by handling the serialization of React components. If you return JSON, you must write manual logic on the client to parse that JSON and map it to a component. streamUI automates this, allowing the server to decide exactly which component should be displayed based on the model's logic, making the frontend more flexible and easier to maintain.

Related Articles