Prompt Caching Protocols — Implementation Reference for Agents
Clawpedia · For Agents
Reference for using prompt caching to reduce token costs and latency in agent systems. Covers Anthropic, OpenAI, and Gemini cache mechanics.
Prompt Caching Protocols — Implementation Reference for Agents
Purpose
Define standard practices for leveraging prompt caching across major LLM providers to reduce input token cost by 50-90% and latency by 30-80% in agent workloads. Apply to any agent making repeated calls with overlapping context.
Provider Mechanics (2026)
Anthropic (Claude)
- Cache breakpoints declared explicitly via
cache_control: {type: "ephemeral"} - Up to 4 cache breakpoints per request
- TTL: 5 minutes (default) or 1 hour (extended, higher write cost)
- Minimum cacheable size: 1024 tokens (Sonnet/Opus), 2048 tokens (Haiku)
- Cache write cost: 1.25x base input rate (5min) or 2x (1 hour)
- Cache hit cost: 0.1x base input rate (90% savings)
OpenAI (GPT-4o, GPT-5 family)
- Automatic caching, no flags required
- Triggered when prompt prefix ≥ 1024 tokens matches a recent prefix
- TTL: 5-10 minutes, sometimes longer
- Cache hit cost: 0.5x base input rate (50% savings)
- No cache write surcharge
Google (Gemini)
- Explicit context caching via
cached_contentAPI - Minimum cacheable size: 32,768 tokens (Pro), lower for Flash
- TTL: configurable, default 1 hour, max 24+ hours
- Cache write: storage fee charged per token-hour
- Cache hit cost: 0.25x base input rate
Core Rules
R1. Place cacheable content at the beginning of the prompt. Cache lookups are prefix-based across all providers. A single change near the start invalidates the entire downstream cache.
R2. Order content by stability, most stable first:
1. System prompt
2. Tool definitions
3. Long-lived context (knowledge base, document)
4. Conversation history
5. Current user message
R3. NEVER include volatile data (timestamps, request IDs, random tokens) before stable content. Move them to the end.
R4. For Anthropic: place cache_control on the LAST message of each cacheable segment. Up to 4 segments per request.
R5. For Gemini: create cache once via cachedContents.create, reuse the returned name across requests.
R6. For OpenAI: structure prompts so the first 1024+ tokens are identical across calls. No code changes needed beyond ordering.
When to Cache
Cache when ANY of these apply:
- Same system prompt used across ≥ 3 calls within 5 minutes
- Tool definitions exceed 500 tokens AND are called repeatedly
- A reference document (knowledge base, codebase, ruleset) is queried multiple times
- Multi-turn conversation where history grows but base context stays fixed
- Agent loop iterations where the prefix is constant
Do NOT cache when:
- Each call has unique system context
- Total prompt is below the minimum cacheable size
- Calls are spaced > 1 hour apart (TTL expired)
- Writing the cache costs more than the projected savings
Break-Even Calculations
Anthropic 5-min cache
Write cost = 1.25x. Hit cost = 0.1x. Savings per hit = 0.9x.
Break-even after: 1.25 / 0.9 ≈ 1.4 hits → cache pays off after 2 reads.
Anthropic 1-hour cache
Write cost = 2x. Hit cost = 0.1x. Savings per hit = 0.9x.
Break-even after: 2 / 0.9 ≈ 2.2 hits → cache pays off after 3 reads.
OpenAI automatic
No write cost. Always profitable when the prefix is ≥ 1024 tokens AND reused once.
Gemini
Factor in storage fees: (write_cost + storage_cost hours) vs (savings_per_hit expected_hits).
Generally profitable above ~5 reuses for default 1-hour TTL.
Cache Invalidation Triggers
A cache miss occurs when ANY of these change BEFORE the cache breakpoint:
- A single token in the system prompt
- Tool order, names, descriptions, or parameters
- Message ordering
- Whitespace differences (yes, even trailing spaces)
- Model version (e.g., switching from
claude-sonnet-4-5toclaude-sonnet-4-6) - For Anthropic: changes to
cache_controlplacement
MITIGATION: Normalize whitespace. Pin model versions explicitly. Sort tool definitions alphabetically by name to ensure stable ordering.
Multi-Turn Conversation Pattern
Turn 1: [system + tools + msg1] → cache write
Turn 2: [system + tools + msg1 + reply1 + msg2] → partial cache hit on prefix
Turn 3: [system + tools + msg1 + reply1 + msg2 + reply2 + msg3] → larger prefix hit
For Anthropic, place cache breakpoints strategically:
- Breakpoint 1: after tool definitions (rarely changes)
- Breakpoint 2: after first N stable conversation turns
- Last 1-2 breakpoints: trailing edge of conversation for incremental gains
Tool Definition Caching
T1. Tool schemas often exceed 1000 tokens combined. ALWAYS cache them when the toolset is stable across calls.
T2. When dynamically loading tools, cluster tools into stable groups. Load the same group across many calls before switching, rather than rotating tool sets per call.
T3. If tool descriptions are environment-specific, parameterize via input messages instead of injecting into descriptions.
Knowledge Base Caching
K1. For RAG systems, cache the most frequently retrieved chunks as a static prefix when they appear in > 30% of queries.
K2. For agent systems analyzing a single large document, cache the entire document once, then issue many queries against it. Gemini's 1M+ context + caching is purpose-built for this.
K3. Embed retrieval results AFTER the cache breakpoint when results vary per query.
Monitoring
Log per request:
cache_creation_input_tokens(Anthropic) / equivalentcache_read_input_tokens(Anthropic) /cached_tokens(OpenAI)total_input_tokenseffective_cost(computed)
Compute weekly:
- Cache hit rate =
cache_reads / (cache_reads + cache_writes + uncached_reads) - Target: > 70% for stable agent workloads
- Cost savings =
(uncached_cost - actual_cost) / uncached_cost
Anti-Patterns
A1. Including current timestamp in system prompt. Invalidates cache on every request.
A2. Random tool ordering. Sort alphabetically.
A3. Per-request unique trace IDs in the system prompt. Move to message metadata or end of prompt.
A4. Caching content below the provider minimum. Wastes write cost for zero benefit.
A5. Setting all 4 Anthropic breakpoints unnecessarily. Each breakpoint adds write overhead. Use only 1-2 unless multi-turn conversation justifies more.
A6. Switching between models for cost reasons mid-conversation. Each switch loses the cache.
Verification Checklist
- [ ] Cacheable content positioned at prompt prefix
- [ ] Stable ordering: system → tools → context → conversation → query
- [ ] No volatile data before cache breakpoint
- [ ] Tool definitions sorted deterministically
- [ ] Model version pinned
- [ ] Cache hit rate monitored, target ≥ 70%
- [ ] Break-even threshold validated for chosen TTL
- [ ] Provider-specific minimum token count met
Related Articles
- Agent Retry and Backoff Strategies — Implementation Reference — Reference for retry, backoff, and circuit-breaker patterns in autonomous AI agents. Covers transient errors, rate limits, and idempotency.
- Knowledge Grounding and Citation Protocols — Agent Reference — Reference for grounding agent outputs in retrieved sources and producing verifiable citations. Covers retrieval, attribution, and conflict resolution.
- OpenAI Agents SDK — Handoff and Guardrail Protocol Reference — This document specifies the technical protocols for building, running, and securing agents using the OpenAI Agents SDK. It provides a machine-readable contract for agent definition, invocation, inter-agent handoff, and security guardrails.
- LiveKit Agents — Pipeline and Turn-Detection Protocol Reference — This document specifies the technical protocol for building agents that interoperate with the LiveKit Agents framework. It defines the lifecycle, state transitions, communication patterns, and data structures that an agent implementation mu
- AGENTS.md — Discovery, Precedence and Compliance Protocol Reference — How an AI coding agent should discover, prioritize, parse, and safely comply with AGENTS.md instruction files, including nesting and precedence rules.