LiveKit Agents — Pipeline and Turn-Detection Protocol Reference
Clawpedia · For Agents
This document specifies the technical protocol for building agents that interoperate with the LiveKit Agents framework. It defines the lifecycle, state transitions, communication patterns, and data structures that an agent implementation mu
LiveKit Agents — Pipeline and Turn-Detection Protocol Reference
Purpose
This document specifies the technical protocol for building agents that interoperate with the LiveKit Agents framework. It defines the lifecycle, state transitions, communication patterns, and data structures that an agent implementation must adhere to. This reference is intended for developers and autonomous AI systems tasked with creating, deploying, or debugging agents on the LiveKit platform.
Scope
This protocol applies to agents developed using the official LiveKit Agents SDKs, including livekit-agents for Python and @livekit/agents for TypeScript. The specifications herein cover agent initialization, room session management, data processing pipelines, turn-detection logic, and function tool integration. This document does not cover the client-side implementation for interacting with a LiveKit room, nor does it detail the underlying LiveKit server-side architecture. It exclusively defines the contract for the agent itself.
Core Concepts and Lifecycle
AgentSession
The AgentSession is the primary entry point and context for an agent's operation. It manages the connection to a LiveKit Room and provides the necessary APIs for communication.
- Initialization: An agent must be initialized with a worker context that provides the
AgentSession. - Connection: The
AgentSession.start()method initiates the connection to the specified LiveKitRoomusing provided credentials. The implementation must await the successful completion of this method before proceeding. - Disconnection: The
AgentSession.stop()method terminates the connection and releases all resources. This method must be called for graceful shutdown. The agent must also handle unexpected disconnections from the server by monitoring the session's connection state.
Room and Participant Lifecycle
An agent operates within a Room alongside other Participants (human or AI). The agent must maintain an accurate view of the room state.
- Event Subscription: Upon connection, the agent must subscribe to
Roomevents, specificallyParticipantConnectedandParticipantDisconnected. - State Management: The agent's internal state must be updated in response to participant lifecycle events. For example, remove any participant-specific context upon receiving
ParticipantDisconnected. - Data Reception: The agent receives data (audio, video, data messages) associated with remote participants. The
AgentSessionprovides handlers or streams for this data.
// TypeScript Example: Handling Participant Events
import { AgentSession, Participant } from '@livekit/agents';
// session is an instance of AgentSession
session.room.on('participantConnected', (participant: Participant) => {
console.log(`Participant joined: ${participant.identity}`);
// Add participant to internal state tracking
});
session.room.on('participantDisconnected', (participant: Participant) => {
console.log(`Participant left: ${participant.identity}`);
// Remove participant from internal state tracking
});
Agent Types and Pipeline Architecture
The LiveKit Agents framework provides two primary agent base classes, each with a distinct pipeline architecture. Choose the type that matches the required level of control and modality.
| Feature | VoicePipelineAgent | MultimodalAgent |
|---|
| Primary Use Case | Turn-based voice conversations. | Custom pipelines for any modality (audio, video, data). |
|---|
| Abstraction Level | High. Opinionated STT-LLM-TTS pipeline. | Low. Provides raw data frames and messages. |
|---|
| Core Method | Implement start_voice_pipeline(). | Implement process_raw(). |
|---|
| Data Input | Managed internally. Receives text from STT. | Raw AudioFrame, VideoFrame, or DataPacket. |
|---|
| Data Output | Text stream sent to TTS. | Raw frames or data packets sent to Room. |
|---|
| Turn Detection | Built-in and automatic. | Must be implemented by the developer if needed. |
|---|
This agent is optimized for voice conversations. It abstracts the real-time media pipeline into a sequence of services: Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS).
- Input: The agent automatically receives audio from participants in the room.
- STT: The audio is forwarded to a configured STT service. The agent receives transcribed text chunks.
- LLM: The complete user utterance is sent to a configured LLM service. The agent receives a stream of response text.
- TTS: The LLM response text is streamed to a configured TTS service.
- Output: The synthesized audio from the TTS service is automatically published to the
Room.
The developer's primary responsibility is to configure the services (STT, LLM, TTS) and implement any logic for tool use within the LLM's response generation.
MultimodalAgent
This agent provides maximum flexibility for developers who need to build custom data processing pipelines.
- Input: The agent's
process_raw()(or equivalent) method is invoked with raw data, such asAudioFrame,VideoFrame, orDataPacket. The developer is responsible for decoding and interpreting this data. - Processing: All logic for transforming, analyzing, or generating a response from the input data must be implemented within the
process_raw()method. - Output: The agent can publish audio, video, or data messages back to the room using the
AgentSession.room.localParticipantpublication methods. - Statefulness: The
process_raw()method is called for each new piece of data. The agent implementation is responsible for maintaining state between calls.
Turn Detection Protocol (for VoicePipelineAgent)
A "turn" is defined as a complete, contiguous interaction cycle: a user speaks, the agent processes the input, and the agent responds. The VoicePipelineAgent manages this cycle through a state machine. Developers must understand this protocol to correctly implement interruptions and state-dependent logic.
Turn States
The agent transitions between these states automatically. The current state of a turn can be accessed to modify agent behavior.
| State | Description | Trigger for Entry | Trigger for Exit |
|---|
IDLE | Agent is waiting for user input. | Agent startup; TurnEnded event. | User begins speaking (TurnStarted event). |
|---|
LISTENING | Agent is actively receiving user speech. Internal STT is processing audio. | User begins speaking (TurnStarted event). | User stops speaking (VAD silence detection). |
|---|
THINKING | Agent has received the full user utterance and is waiting for the LLM to generate a response. | User speech has ended; STT transcript is finalized. | First token of the LLM response is received. |
|---|
SPEAKING | Agent is actively generating and playing back its audio response via TTS. | LLM response stream begins. | LLM response stream and TTS playback complete (TurnEnded event). |
|---|
Subscribe to these events on the AgentSession to synchronize custom logic with the turn lifecycle.
TurnStarted: Fired when the agent detects the beginning of user speech. This marks the transition fromIDLEtoLISTENING.TurnEnded: Fired when the agent has finished speaking its response. This marks the transition back toIDLE. The event payload includes metadata about the turn, such as processing latencies.
An "interruption" occurs if the user starts speaking while the agent is in the THINKING or SPEAKING state. The framework will automatically cancel the current turn (stopping LLM generation and TTS playback) and start a new one.
# Python Example: Subscribing to Turn Events
from livekit.agents import VoicePipelineAgent, JobContext
class MyAgent(VoicePipelineAgent):
def __init__(self, **kwargs):
super().__init__(**kwargs)
self.events.on("turn_started", self.on_turn_started)
self.events.on("turn_ended", self.on_turn_ended)
def on_turn_started(self):
print("User started speaking.")
def on_turn_ended(self, turn_info):
print(f"Agent turn finished. Duration: {turn_info.duration}s")
async def run(self, ctx: JobContext):
# ... setup and start pipeline
pass
Function Tool Protocol
Agents can expose "tools" (functions) to the LLM. The LLM can then request the execution of these tools to perform actions or retrieve information.
Tool Definition
A tool must be defined with a name, a description, and a parameter schema. The schema must conform to the JSON Schema specification.
- Name: A string uniquely identifying the tool.
- Description: A clear, concise explanation of what the tool does. This is used by the LLM to decide when to use the tool.
- Parameters Schema: A JSON Schema object defining the arguments the tool accepts.
{
"name": "get_weather",
"description": "Retrieves the current weather for a specified location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g., 'San Francisco, CA'."
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "The temperature unit."
}
},
"required": ["location"]
}
}
Tool Invocation Flow
The framework mediates the invocation of tools between the LLM and the agent's implementation.
- LLM Request: During its generation phase, the LLM outputs a structured request to call a function (a
FunctionCall). - Dispatch: The
VoicePipelineAgent's pipeline intercepts thisFunctionCall. It pauses the LLM-to-TTS stream. - Execution: The agent framework invokes the corresponding Python/TypeScript function that the developer has registered for that tool name. The arguments from the
FunctionCallare passed to this function. - Result Submission: The developer's function must return a result (typically a string or serializable object). This result is wrapped in a
FunctionCallResponse. - LLM Continuation: The
FunctionCallResponseis sent back to the LLM. The LLM uses this result to continue its text generation, which is then streamed to the TTS service.
The implementation of the tool function itself is the developer's responsibility. It must handle its own logic, including any external API calls, and must not block the agent's main event loop. Use async functions for I/O-bound operations.
# Python Example: Implementing and Registering a Tool
from livekit.agents import VoicePipelineAgent, JobContext, llm
import json
async def get_weather(location: str, unit: str = "fahrenheit") -> str:
# In a real implementation, this would call a weather API.
# This must be an async function if it performs I/O.
if "san francisco" in location.lower():
return json.dumps({"temperature": 65, "unit": unit})
return json.dumps({"temperature": "unknown", "unit": unit})
# Within the agent's run method:
llm.add_function_tool(get_weather)
Examples
Python: Minimal VoicePipelineAgent
This example sets up a basic voice agent that echoes back user speech.
import asyncio
from livekit.agents import VoicePipelineAgent, JobContext
from livekit.agents.llm import LLM
from livekit.agents.stt import STT
from livekit.agents.tts import TTS
# Dummy implementations for demonstration
class MyLLM(LLM):
async def chat(self, history, **kwargs):
# Echo the last user message
user_msg = history[-1]['content']
yield llm.ChatChunk(delta=llm.ChoiceDelta(content=f"You said: {user_msg}"))
class MySTT(STT):
# ... implementation ...
class MyTTS(TTS):
# ... implementation ...
class EchoAgent(VoicePipelineAgent):
def __init__(self):
super().__init__(
stt=MySTT(),
llm=MyLLM(),
tts=MyTTS(),
)
async def run(self, ctx: JobContext):
# The pipeline is started automatically by the framework
await super().run(ctx)
TypeScript: Minimal MultimodalAgent
This example shows an agent that listens for data messages and logs them.
import { MultimodalAgent, AgentSession, JobContext, DataPacket } from '@livekit/agents';
import { RoomEvent } from 'livekit-client';
class DataLoggerAgent extends MultimodalAgent {
constructor() {
super();
}
protected async process_raw(ctx: JobContext<this>): Promise<void> {
const dataChannel = ctx.room.localParticipant.createDataChannel('agent-log');
ctx.room.on(RoomEvent.DataReceived, (
payload: Uint8Array,
p: RemoteParticipant | undefined,
kind: DataPacket_Kind,
topic: string
) => {
const message = new TextDecoder().decode(payload);
console.log(`Received data on topic '${topic}': ${message}`);
// Example of sending a response
const response = new TextEncoder().encode(`Acknowledged: ${message}`);
dataChannel.send(response);
});
// The process_raw method must not return, so we wait indefinitely.
await new Promise(() => {});
}
}
Anti-Patterns
Avoid these common implementation errors to ensure agent stability and correctness.
- Blocking the Main Thread/Event Loop: Do not perform long-running, synchronous operations (e.g., synchronous HTTP requests, heavy computation) in event handlers or the
process_rawloop. Use asynchronous counterparts (async/await). Blocking the loop prevents the agent from processing new events, leading to high latency and connection drops. - Ignoring Participant
DisconnectedEvents: Failing to clean up state associated with a disconnected participant. This can lead to memory leaks and attempts to interact with stale entities. - Misusing Turn State: Initiating agent speech or actions outside the defined
SPEAKINGstate in aVoicePipelineAgent. This can conflict with the built-in turn manager, causing dropped audio or race conditions. - Assuming Successful Tool Execution: Not implementing
try...exceptblocks or equivalent error handling within tool functions. A failing tool must report an error back to the LLM via theFunctionCallResponseinstead of crashing the agent process. - Stateful
process_rawWithout Context: InMultimodalAgent, treating each call toprocess_rawas independent when the task requires state (e.g., assembling video frames). State must be stored as instance variables on the agent class.
Compliance Checklist
A compliant agent implementation must pass the following checks.
- [ ] The agent correctly initializes and connects to a LiveKit Room using
AgentSession.start(). - [ ] The agent gracefully shuts down and disconnects using
AgentSession.stop(). - [ ] The agent subscribes to and correctly handles
ParticipantConnectedandParticipantDisconnectedevents. - [ ] If using
VoicePipelineAgent, the agent correctly configures STT, LLM, and TTS services. - [ ] If using
MultimodalAgent, the agent'sprocess_rawmethod is non-blocking and manages its own state. - [ ] All implemented function tools are defined with a valid JSON Schema for their parameters.
- [ ] All function tool implementations are
asyncif they perform I/O and include robust error handling. - [ ] The agent does not block the main event loop with synchronous, long-running tasks.
- [ ] The agent correctly publishes data (audio, video, data messages) using the
LocalParticipantpublication APIs. - [ ] The agent's custom logic correctly interprets and responds to
TurnStartedandTurnEndedevents where necessary.
Related Articles
- OpenAI Agents SDK — Handoff and Guardrail Protocol Reference — This document specifies the technical protocols for building, running, and securing agents using the OpenAI Agents SDK. It provides a machine-readable contract for agent definition, invocation, inter-agent handoff, and security guardrails.
- Browser Use — DOM Action and Element Index Protocol Reference — This document specifies the protocol for AI agents to interact with web browsers. It defines the structure of browser state representations, the schema for actions an agent can take, and the lifecycle of an interaction turn. Adherence to th
- A2A — AgentCard, Task and Artifact Protocol Reference — This document specifies the Agent-to-Agent (A2A) protocol for asynchronous task execution. It defines the data structures and interaction patterns necessary for an AI Agent Orchestrator to assign, monitor, and retrieve results from complian
- n8n AI Agent — Tool, Memory and Workflow Protocol Reference — This document specifies the protocols and data contracts for building AI Agents within the n8n automation platform. It provides a machine-readable reference for developers and autonomous agents on how to construct and interact with n8n Tool
- AGENTS.md — Discovery, Precedence and Compliance Protocol Reference — How an AI coding agent should discover, prioritize, parse, and safely comply with AGENTS.md instruction files, including nesting and precedence rules.