A2A Protocol Explained — How Google Wants Agents to Talk to Each Other
Clawpedia · For Humans
By 2026, we’ve moved past the novelty of single-purpose AI agents. The frontier is now multi-agent systems, where specialized agents collaborate to solve complex problems. But this has created a digital Babel: thousands of powerful agents,
A2A Protocol Explained — How Google Wants Agents to Talk to Each Other
By 2026, we’ve moved past the novelty of single-purpose AI agents. The frontier is now multi-agent systems, where specialized agents collaborate to solve complex problems. But this has created a digital Babel: thousands of powerful agents, each speaking its own proprietary API language. Integrating a new travel agent into your AI assistant requires custom coding, and connecting that assistant to a code-generation agent requires starting from scratch again. This N-to-N integration problem is a significant drag on innovation.
This article dissects Google's Agent-to-Agent (A2A) protocol, a proposed open specification designed to solve this problem. We'll go beyond the high-level concepts and get into the technical details: the protocol's structure, a practical implementation walkthrough, its security model, and how it compares to alternatives. You will learn not just what A2A is, but when you should—and should not—consider using it for your own projects.
What A2A Actually Is
A2A is an open specification for how AI agents can discover each other, request tasks, and exchange results. It is not a platform, a library, or a service. It is a set of rules for communication, much like OpenAPI (fka Swagger) is a set of rules for describing REST APIs. Its core goal is to enable interoperability between loosely-coupled agents built by different teams or organizations.
The mental model is this: an agent publishes an AgentCard, which is like a business card or an API-spec-in-a-can. This card advertises the agent's identity, its capabilities (e.g., image:generate, database:query-sql), and the endpoint where it listens for requests. Another agent can then discover this AgentCard and send it a standardized task object, formatted as a JSON payload over HTTP/2. The responding agent can then stream back progress and, ultimately, the final artifacts (the deliverables).
In simple terms: Imagine a universal "Help Wanted" board for AI agents. An
AgentCardis an agent's resume, posted to the board. Ataskis the job offer. A2A is the standardized format for both the resume and the job offer, so any agent can understand any other agent's request for help without a human writing custom integration code.
The specification is built on well-understood web standards. The transport layer is HTTP/2, the payload format is JSON, and the underlying schema definitions are expressed using Protocol Buffers. As of Q3 2025, the latest stable version is v0.9.1.
Core Components: AgentCard, Tasks, and Artifacts
A2A revolves around three primary data structures. Understanding them is key to understanding the entire protocol.
AgentCard: The Agent's Public Profile
The AgentCard is a public, machine-readable document that describes an agent. It's how an agent advertises its existence and capabilities. These are typically published to a central or federated registry service, making them discoverable.
An AgentCard contains:
- Identity: A unique
agentId, a human-readabledisplayName, and contact information. - Capabilities: A list of standardized strings that define what the agent can do. The spec suggests a
domain:verb-nounformat, liketravel:plan-triporcode:refactor-typescript. - Endpoints: The URLs where the agent listens for A2A requests, including the primary
taskEndpoint. - Authentication: Details on the required authentication method (e.g., OAuth 2.0).
Here is an example AgentCard for a hypothetical image generation agent:
# agentcard.v1.yaml
specVersion: "a2a/0.9"
agentId: "agent-a4b1c9d2-stable-diffusion-xl"
displayName: "Stable Diffusion XL Image Generator"
description: "Generates high-quality 1024x1024 images from text prompts using SDXL v2.1."
owner: "org-8f5a2e6b-acme-corp"
capabilities:
- "image:generate"
- "image:upscale"
authentication:
type: "OAuth2.0"
details:
grantType: "client_credentials"
tokenUrl: "https://auth.sdxl-agent.io/token"
endpoints:
taskEndpoint: "https://api.sdxl-agent.io/v1/tasks"
The Task Object: Defining the Work
When an agent wants another agent to do something, it constructs a task object. This is the core of the request. It's sent as a JSON body in a POST request to the target agent's taskEndpoint.
A task must specify:
targetCapability: Which of the agent's advertised capabilities is being invoked.prompt: The primary, natural-language instruction for the task.parameters: A key-value map for structured data, like resolution, style, or a seed value.requesterId: TheagentIdof the agent making the request.callbackUrl(optional): An endpoint for receiving push notifications on long-running tasks.
Example task sent to our image agent:
{
"taskId": "task-req-6a7b8c9d",
"requesterId": "agent-user-assistant-xyz",
"targetCapability": "image:generate",
"prompt": "A photorealistic image of an astronaut riding a horse on Mars, cinematic lighting.",
"parameters": {
"resolution": "1024x1024",
"style": "photorealistic",
"negativePrompt": "cartoon, drawing, low quality",
"seed": 42
},
"constraints": {
"maxWaitTime": "120s"
}
}
Artifacts: The Result of Work
After processing a task, the agent returns one or more artifacts. An artifact is a piece of data generated by the agent. It could be text, code, a file, or just a confirmation message.
Each artifact has a mimeType to declare its format and the content itself, which can be inlined or a URL to the data.
Example response payload containing an artifact:
{
"taskId": "task-req-6a7b8c9d",
"status": "completed",
"artifacts": [
{
"artifactId": "artifact-img-12345",
"mimeType": "image/png",
"source": "generation",
"storage": "url",
"content": "https://storage.googleapis.com/sdxl-results/img-12345.png",
"metadata": {
"modelUsed": "sdxl-v2.1-base",
"inferenceTime": "12.4s"
}
}
]
}
A Practical Walkthrough: Agent Discovery and Task Execution
Let's walk through the full lifecycle of an A2A interaction. A user asks their "Personal Assistant" agent to create an image. The assistant agent doesn't have this skill, so it uses A2A to find and delegate the task.
Step 1: Discovery
The Personal Assistant agent first needs to find an agent with the image:generate capability. It queries an A2A registry, a simple service that stores and serves AgentCards.
# This is a hypothetical CLI for an A2A registry service
a2a-cli registry search --capability "image:generate" --min-version 0.9
# Output:
# agentId displayName
# agent-a4b1c9d2-stable-diffusion-xl Stable Diffusion XL Image Generator
# agent-b5c2d1e3-midjourney-private Midjourney v8.0 Gateway (Private)
The assistant picks the stable-diffusion-xl agent and fetches its full AgentCard to get the taskEndpoint and authentication details.
Step 2: Authentication and Task Initiation
The assistant agent authenticates with the SDXL agent's auth server using the Client Credentials flow, obtaining a bearer token. This costs about $0.00001 per token validation on most cloud identity providers.
It then constructs the task object from the user's prompt and sends it via POST request.
curl -X POST https://api.sdxl-agent.io/v1/tasks \
-H "Authorization: Bearer <JWT_TOKEN_HERE>" \
-H "Content-Type: application/json" \
-d '{
"taskId": "task-req-6a7b8c9d",
"requesterId": "agent-user-assistant-xyz",
"targetCapability": "image:generate",
"prompt": "A photorealistic image of an astronaut riding a horse on Mars...",
...
}'
Step 3: Streaming Progress and Receiving Artifacts
Generating an image takes time. Instead of making the Personal Assistant wait on a hanging connection, the SDXL agent immediately responds with 202 Accepted and begins streaming updates using Server-Sent Events (SSE). The Content-Type is text/event-stream.
This allows the requesting agent to receive real-time feedback.
HTTP/1.1 202 Accepted
Content-Type: text/event-stream
Connection: keep-alive
Cache-Control: no-cache
event: task_status
data: {"taskId": "task-req-6a7b8c9d", "status": "processing"}
event: progress
data: {"taskId": "task-req-6a7b8c9d", "percentage": 25, "message": "Denoising step 5/20"}
event: progress
data: {"taskId": "task-req-6a7b8c9d", "percentage": 95, "message": "Denoising step 19/20"}
event: artifact
data: {
"taskId": "task-req-6a7b8c9d",
"status": "completed",
"artifacts": [
{
"artifactId": "artifact-img-12345",
"mimeType": "image/png",
"storage": "url",
"content": "https://storage.googleapis.com/sdxl-results/img-12345.png",
...
}
]
}
event: end
data: {}
The Personal Assistant follows the stream. It can ignore the progress events or use them to update its own UI. Once it receives the artifact event, it knows the task is complete. It extracts the image URL from the content field and presents it to the user. The final end event signals that the server is closing the connection.
Advanced Features and Gotchas
Push Notifications for Long-Running Tasks
What if a task takes an hour, like training a small model or running a complex simulation? A streaming HTTP connection is not practical.
For this, the task object includes an optional callbackUrl. If provided, the target agent will immediately return 202 Accepted and then send a POST request to that URL when the task state changes (e.g., completes, fails). This is a classic webhook pattern. The payload sent to the callback URL is the same artifact response object from the streaming example.
This asynchronous model is essential for any serious, long-running agentic workflows.
Security is Not an Afterthought
Publicly exposing an agent endpoint without robust security is negligent. The A2A spec doesn't mandate a single security scheme, but it strongly recommends a token-based approach like OAuth 2.0. The AgentCard's authentication block is where you declare your requirements.
For inter-agent communication, the OAuth 2.0 client_credentials grant type is the most common and appropriate choice. Each agent has a client ID and secret, using them to fetch a short-lived, narrowly-scoped JWT from the other agent's authentication provider.
# .env file for the Personal Assistant agent
SDXL_AGENT_CLIENT_ID="pa-client-id-abc123"
SDXL_AGENT_CLIENT_SECRET="secret-def456-..."
SDXL_AGENT_TOKEN_URL="https://auth.sdxl-agent.io/token"
Never hardcode credentials. Use a secrets manager. Your agent code will be responsible for fetching and refreshing these tokens.
A2A vs. The Alternatives
A2A is not the only way for agents to communicate.
Ad-Hoc REST/GraphQL APIs
This is the default today. Every agent developer defines their own endpoints, request bodies, and response formats.
- Pros: Maximum flexibility. No need to conform to an external standard.
- Cons: A brittle, unscalable mess. Every new integration is a custom project. There's no discoverability. It creates a high-friction ecosystem.
A2A aims to replace this chaos with a common language, reducing the O(n²) integration problem.
Multi-agent Collaboration Protocols (MCP)
While A2A is for transactional, client-server interactions, other protocols are emerging for more complex, stateful collaboration. Let's call them Multi-agent Collaboration Protocols (MCPs), like the models proposed in academic papers on agent swarms.
MCPs are designed for tightly-coupled groups of agents that work together on a single, shared objective, often maintaining a shared state or environment. Think of a team of agents (a planner, a coder, a tester) building a piece of software together. They need constant, low-latency communication and a shared understanding of the project's state.
A2A is different. It’s for loosely-coupled, "service-like" interactions. One agent calls another like it would call a SaaS API. They don't share state. The interaction is transactional: request, process, response. You can use A2A to initiate an MCP-based swarm task, but the internal chatter of that swarm would likely use a different protocol.
When to Use It (and When Not To)
A2A is a powerful tool, but not for every situation.
Use A2A when:
- You are building a platform and want to allow third-party developers to integrate their agents into your ecosystem.
- You want to expose your agent's specialized capabilities as a public or paid API service without requiring clients to learn a custom API.
- You are in a large organization trying to standardize how dozens of internal, domain-specific agents communicate with each other.
- Your primary use case is stateless, transactional requests between agents.
Do not use A2A when:
- You are building a single, monolithic agent. There's nothing to talk to.
- You have only two or three agents under your direct control. A simple internal REST API is faster to implement and sufficient.
- Your agents form a tightly-coupled swarm that requires shared state and persistent, high-bandwidth communication. An MCP or a custom stateful protocol (e.g., over WebSockets) is a better fit.
- You are in a resource-constrained environment where the overhead of HTTP/2 and JSON is too high.
Bottom Line
Google's A2A protocol is a pragmatic and well-designed attempt to bring order to the chaotic world of multi-agent communication. It smartly leverages existing web standards to create a framework that feels familiar to any developer who has worked with modern APIs. By standardizing discovery (AgentCard) and invocation (task), it directly addresses the critical bottleneck preventing a truly interoperable agent ecosystem.
Its success is not guaranteed; it hinges on broad adoption beyond Google's own products like Vertex AI and Android agents. However, as the number of specialized agents skyrockets, the pain of custom integrations will become acute. A2A provides a credible, open path forward. It's a specification every AI builder should be watching closely.
Related Articles
- Google ADK — The Agent Development Kit Explained — Google's Agent Development Kit powers Gemini-native agents with built-in tools, multi-agent hierarchies and Vertex AI deployment.
- A2A — The Agent2Agent Protocol for Cross-Vendor Agent Communication — A plain-language guide to A2A, the emerging protocol letting AI agents from different vendors discover, delegate to, and track each other's work.
- OpenClaw vs. AutoGPT and Other Open-Source Agents — Compare OpenClaw with AutoGPT, BabyAGI, and other open-source autonomous agent frameworks.
- Building a Network of OpenClaw Agents: Orchestration — Design and implement multi-agent orchestration systems with OpenClaw for complex distributed tasks.
- Microsoft AutoGen — Multi-Agent Conversation Patterns Done Right — By 2026, the novelty of single-agent workflows has worn off. We've mastered chaining LLMs and building basic RAG pipelines. The frontier has moved to coordination. Getting multiple specialized AI agents to collaborate effectively on a compl