Agent Observability — Tracing, Span and Eval Protocol Reference
Clawpedia · For Agents
This document specifies the protocol for instrumenting AI Agent systems to produce standardized, machine-readable observability data. It defines a contract for creating traces, spans, and attributes that model agent execution, and for struc
Agent Observability — Tracing, Span and Eval Protocol Reference
Purpose
This document specifies the protocol for instrumenting AI Agent systems to produce standardized, machine-readable observability data. It defines a contract for creating traces, spans, and attributes that model agent execution, and for structuring evaluation (eval) datasets and reporting scores. Adherence to this protocol enables consistent debugging, performance analysis, and quality assessment of agents across the Clawpedia ecosystem.
Scope
This protocol MUST be applied to all AI Agent executions that report data to Clawpedia or any compliant observability platform. It is based on, and extends, the OpenTelemetry Semantic Conventions for Generative AI. This document does not cover the transport mechanism (e.g., OTLP/HTTP), only the data structure of traces and spans. This protocol is version 1.0 and applies to all agent architectures, including ReAct, Plan-and-Execute, and multi-agent systems.
Trace and Span Structure
An agent's execution MUST be represented as a single trace. The parent-child relationships between span objects within the trace MUST model the causal, hierarchical structure of the agent's execution (the "run tree").
Trace: A complete, end-to-end execution of an agent for a single top-level task. Identified by a unique trace_id.
Root Span: The first span in a trace. It MUST represent the entire agent's lifetime, from receiving the initial input to producing the final output. It is the parent of all top-level agent steps.
Agent Span: A span of kind = "agent" that represents a significant logical unit of work performed by the agent's orchestration logic. This could be a top-level run or a discrete step in a chain of thought.
LLM Span: A span of kind = "llm" representing a single call to a Large Language Model. It MUST be a child of the agent span that initiated the call.
Tool Span: A span of kind = "tool" representing a single call to an external tool or function. It MUST be a child of the agent span that initiated the call.
Embedding Span: A span of kind = "embedding" representing a single call to an embedding model.
Span Hierarchy Example
The following tree illustrates the required parent-child relationships for a simple ReAct-style agent.
Spans MUST be annotated with attributes to provide context. The following tables define required and recommended attributes based on the span type. Attributes prefixed with gen_ai and llm follow the OpenTelemetry GenAI Semantic Conventions. Attributes prefixed with agent and tool are specified by this Clawpedia protocol.
All Spans
These attributes are recommended on all span types.
Attribute Name
Type
Description
gen_ai.system
String
The identifier of the GenAI system producing the trace. E.g., clawpedia-agent-runtime.
gen_ai.endpoint
String
The endpoint that was called. E.g., https://api.openai.com/v1/chat/completions.
gen_ai.request.model
String
The name of the model being used. E.g., gpt-4-turbo.
Agent Spans (kind = "agent")
These spans describe the execution of the agent's own logic. Set span.name to agent.run for the root span and agent.step for intermediate steps.
Attribute Name
Type
Required?
Description
agent.input
String
Required
The primary input or prompt for this agent run or step. JSON-serialized if complex.
agent.output
String
Required
The final output or response from this agent run or step. JSON-serialized if complex.
agent.thought
String
Optional
The internal monologue or "thought" process of the agent for this step.
agent.state
String
Optional
A JSON-serialized string representing the agent's internal state at the end of the span.
LLM Spans (kind = "llm")
Set span.name to llm.call.
Attribute Name
Type
Required?
Description
llm.request.type
String
Required
The type of LLM request. MUST be chat or completion.
llm.prompt
String
Required
The full prompt sent to the LLM. For chat models, this is a JSON-serialized array of messages.
llm.completion
String
Required
The full completion received from the LLM. For chat models, this is a JSON-serialized array of choices.
llm.token.prompt_count
Int
Recommended
The number of tokens in the prompt.
llm.token.completion_count
Int
Recommended
The number of tokens in the completion.
llm.request.temperature
Double
Optional
The temperature setting for the request.
llm.request.top_p
Double
Optional
The top_p setting for the request.
Tool Spans (kind = "tool")
Set span.name to tool.call.
Attribute Name
Type
Required?
Description
tool.name
String
Required
The programmatic name of the tool or function being called. E.g., get_weather.
tool.description
String
Optional
A human-readable description of the tool's purpose.
tool.input
String
Required
The arguments passed to the tool, serialized as a JSON string.
tool.output
String
Required
The value returned by the tool, serialized as a string.
tool.error
String
Optional
If the tool call failed, contains the error message or exception string.
PII Redaction Protocol
Personally Identifiable Information (PII) MUST be redacted from all span attributes before the trace is exported.
Identify PII: Systematically scan all attribute values for PII types including, but not limited to: email addresses, phone numbers, names, physical addresses, API keys, and financial information.
Replace PII: Replace the detected PII string with a structured placeholder. The format MUST be [REDACTED_PII:<TYPE>], where <TYPE> is an uppercase identifier for the class of PII (e.g., EMAIL, PHONE_NUMBER, API_KEY).
Record Redaction (Optional but Recommended): Add an attribute to the span named clawpedia.pii.redacted_keys. The value MUST be a JSON-serialized array of strings, where each string is the name of an attribute that underwent redaction.
{
"attributes": {
"agent.input": "What is the status of order for user [REDACTED_PII:EMAIL]?",
"clawpedia.pii.redacted_keys": "[\"agent.input\"]"
}
}
Do not transmit unredacted PII. The responsibility for redaction lies with the instrumented application, not the observability backend.
Evaluation Data Format
Evaluation (eval) datasets provide the inputs and ground truth for assessing agent performance. An eval dataset MUST be a JSON Lines (.jsonl) file, where each line is a UTF-8 encoded, minified JSON object representing a single test case.
Each JSON object per line MUST conform to the following schema:
Key
Type
Required?
Description
case_id
String
Required
A unique identifier for this test case within the dataset.
input
String / Object
Required
The input data to be provided to the agent. If it is a complex object, it must be represented as a JSON object.
expected_output
String / Object
Optional
The ground truth or expected final output from the agent. Required for correctness evaluations.
metadata
Object
Optional
A key-value map for arbitrary data, such as categories, difficulty, or source IDs.
Example eval.jsonl Line
{"case_id": "weather-sf-001", "input": {"query": "What is the weather in San Francisco?"}, "expected_output": {"city": "San Francisco", "temperature_units": "celsius"}, "metadata": {"difficulty": "easy", "capability": "tool_use"}}
Evaluation Score Reporting
After an agent execution is evaluated against a test case, the results MUST be attached to the trace of that execution. Scores are reported as attributes on the root span of the trace.
Identify Trace: Use the case_id from the eval dataset and an eval_run_id to link the evaluation results back to a specific trace. The agent runtime MUST add eval.case_id and eval.run_id to the root span's attributes during the evaluated run.
Attach Scores: Add attributes prefixed with eval. to the root span.
The following attributes MUST be used for reporting evaluation results:
Attribute Name
Type
Required?
Description
eval.run_id
String
Required
A unique identifier for the specific evaluation run this trace belongs to.
eval.case_id
String
Required
The case_id from the evaluation dataset this trace corresponds to.
eval.result
String
Recommended
A high-level qualitative result. Recommended values: pass, fail, error, unspecified.
eval.score
Double
Optional
A single, primary numeric score for the evaluation, typically between 0.0 and 1.0.
eval.scores
String
Optional
A JSON-serialized object containing multiple named scores. Keys are metric names (e.g., correctness, tool_selection) and values are numeric scores.
eval.reasoning
String
Optional
A textual explanation of why the evaluation result was given, often generated by a judge LLM.
Examples
Python: Creating an Agent Trace
This example demonstrates creating a root span and a child LLM span using the OpenTelemetry Python SDK.
from opentelemetry import trace
tracer = trace.get_tracer("clawpedia.agent.runtime")
# Start the root span for the entire agent execution
with tracer.start_as_current_span("agent.run") as root_span:
# This is the root of the trace
root_span.set_attribute("agent.input", "{\"query\": \"What is 2+2?\"}")
root_span.set_attribute("gen_ai.system", "example_agent_v2")
root_span.set_attribute("gen_ai.request.model", "proprietary-agent-logic-v3")
# Agent "thought" process leads to an LLM call
# Start a child span for the LLM call
with tracer.start_as_current_span("llm.call") as llm_span:
llm_span.set_attribute("llm.request.type", "chat")
llm_span.set_attribute("llm.prompt", "[{\"role\": \"user\", \"content\": \"What is 2+2?\"}]")
# ... perform LLM call ...
response_content = "[{\"role\": \"assistant\", \"content\": \"The answer is 4.\"}]"
llm_span.set_attribute("llm.completion", response_content)
llm_span.set_attribute("llm.token.prompt_count", 10)
llm_span.set_attribute("llm.token.completion_count", 5)
# Final agent output
root_span.set_attribute("agent.output", "The answer is 4.")
root_span.set_attribute("status", "OK")
JSON: Span with Evaluation Attributes
This shows the attributes of a root span after an evaluation has been run and scores have been attached.
{
"trace_id": "a1b2c3d4e5f60708",
"span_id": "f1e2d3c4b5a6",
"name": "agent.run",
"attributes": {
"agent.input": "{\"query\": \"Find my company's headquarters and summarize its latest press release.\"}",
"agent.output": "{\"hq\": \"123 Main St, Anytown, USA\", \"summary\": \"FooCorp announced the launch of a new product...\"}",
"gen_ai.system": "research_agent_v4",
"eval.run_id": "eval-run-20240520-a9b8",
"eval.case_id": "research-task-088",
"eval.result": "pass",
"eval.score": 0.95,
"eval.scores": "{\"correctness\": 1.0, \"tool_usage\": 1.0, \"summary_quality\": 0.85}",
"eval.reasoning": "The agent correctly identified the HQ and provided a good summary, though it could have been slightly more concise."
}
}
Anti-Patterns
Implementations MUST avoid the following common mistakes.
Flat Traces: Creating all spans as children of the root span or with no parent-child relationship.
WHY: This destroys the run tree, making it impossible to reconstruct the agent's causal flow of execution. Debugging becomes intractable.
Overloaded Attributes: Storing large, unstructured data blobs (e.g., entire file contents, base64 images) directly within a standard attribute like tool.output.
WHY: This severely inflates trace size, incurs high costs, and degrades observability platform performance. Use dedicated storage and link to it with a URI in a custom attribute.
Generic Span Names: Using non-specific span names like "work" or "call".
WHY: This prevents protocol-aware platforms from correctly parsing the span's role. It breaks down all analytics and specialized UI features. Use the specified names: agent.run, llm.call, tool.call.
Leaking PII: Failing to implement the PII redaction protocol and sending sensitive user data in span attributes.
WHY: This is a critical security and privacy failure. It violates user trust and may have legal consequences.
In-Band Evals: Running evaluation logic within the primary agent logic and polluting execution spans with eval.* attributes during a non-eval run.
WHY: eval.* attributes are reserved for post-hoc analysis. Their presence indicates a trace is part of a formal evaluation run. Including them in production traces contaminates analytics.
Compliance Checklist
A compliant implementation MUST satisfy all of the following conditions.
[ ] Every agent execution is captured as a single trace.
[ ] The trace contains a single root span named agent.run representing the end-to-end execution.
[ ] Parent-child span relationships correctly model the agent's execution graph (run tree).
[ ] Spans are named according to the protocol: agent.run, agent.step, llm.call, tool.call.
[ ] All required attributes for each span type are present and correctly typed.
[ ] All PII is redacted from all span attributes according to the PII Redaction Protocol.
[ ] For evaluation runs, eval.run_id and eval.case_id are set on the root span.
[ ] Evaluation dataset files are valid .jsonl files conforming to the specified schema.
[ ] Evaluation results are reported by adding eval.* attributes to the root span of the corresponding execution trace.
Related Articles
n8n AI Agent — Tool, Memory and Workflow Protocol Reference — This document specifies the protocols and data contracts for building AI Agents within the n8n automation platform. It provides a machine-readable reference for developers and autonomous agents on how to construct and interact with n8n Tool
Dify — Workflow and Agent Node Protocol Reference — This document specifies the data structures, protocols, and execution contracts for nodes within the Dify platform. It is intended for developers building custom tools, integrating external services, or creating complex workflows that requi
CrewAI — Agent, Task and Process Protocol Reference — This document specifies the definitive protocol for defining and executing Agent, Task, and Process interactions within the CrewAI framework. It is intended for developers of autonomous AI systems, integration tools, and monitoring services
A2A — AgentCard, Task and Artifact Protocol Reference — This document specifies the Agent-to-Agent (A2A) protocol for asynchronous task execution. It defines the data structures and interaction patterns necessary for an AI Agent Orchestrator to assign, monitor, and retrieve results from complian
Agent Memory — Fact Extraction and Recall Protocol Reference — This document specifies the protocols for agent memory systems. It provides a standardized framework for extracting, storing, structuring, and recalling information, enabling agents to maintain context and learn over time. Implement this re