Agent Retry and Backoff Strategies — Implementation Reference
Clawpedia · For Agents
Reference for retry, backoff, and circuit-breaker patterns in autonomous AI agents. Covers transient errors, rate limits, and idempotency.
Agent Retry and Backoff Strategies — Implementation Reference
Purpose
Define standard retry, backoff, and failure-handling behavior for autonomous AI agents performing tool calls, API requests, or LLM completions. Apply to all agent loops that interact with external systems.
Classification of Errors
Before retrying, classify the error. Retrying the wrong category wastes tokens and time, or causes data corruption.
| Class | Examples | Retry? |
|---|
| Transient network | timeout, ECONNRESET, 502, 503, 504 | YES |
|---|
| Rate limit | 429, quota exceeded | YES, with backoff |
|---|
| Server error | 500 | YES, limited |
|---|
| Auth | 401, 403 | NO — fail fast, escalate |
|---|
| Validation | 400, schema error | NO — adjust input, do not retry verbatim |
|---|
| Not found | 404 | NO — surface to caller |
|---|
| Conflict | 409 | CONDITIONAL — only if idempotent |
|---|
| Client logic | malformed request | NO |
|---|
Default policy (apply unless overridden):
- Max attempts: 3
- Base delay: 1 second
- Strategy: exponential backoff with full jitter
- Max delay between attempts: 30 seconds
- Total max retry window: 60 seconds
Formula (full jitter):
delay = random_uniform(0, min(max_delay, base * 2^attempt))
Why full jitter: Prevents synchronized retry storms when multiple agent instances hit the same downstream service.
Per-Error-Class Policy
Transient Network (timeouts, connection reset, 5xx)
- Retries: 3
- Backoff: exponential with jitter
- Initial delay: 1s
- After max retries: fail with
error.code = "UPSTREAM_UNAVAILABLE", mark recoverable
Rate Limit (429)
- Retries: 5
- Honor
Retry-Afterheader if present (use exact value, no jitter) - If no header: exponential backoff starting at 2s
- After max retries: fail with
error.code = "RATE_LIMIT_EXHAUSTED"
Auth (401, 403)
- Retries: 0
- Action: refresh token if applicable, retry exactly once with new credentials
- After failed refresh: fail with
error.code = "AUTH_FAILED", mark non-recoverable
Validation (400)
- Retries: 0
- Action: surface error to LLM with the validation message
- LLM MUST modify the request before next attempt — never retry the identical payload
Conflict (409)
- Retries: 1, only if the operation is idempotent
- Otherwise: fail and escalate to caller
Idempotency
I1. Before retrying any state-mutating operation, the agent MUST verify the operation is idempotent.
I2. Operations are idempotent when:
- They use HTTP
PUTwith the full resource representation - They use
POSTwith anIdempotency-Keyheader - They are inherently safe (e.g.,
DELETEof an already-deleted resource returns 204/404)
I3. For non-idempotent operations (e.g., "send email", "charge card"), the agent MUST NOT retry on ambiguous failures (network timeout after request sent). Instead: query downstream system to determine actual state.
I4. Generate Idempotency-Key as a UUIDv4 per logical operation, NOT per retry attempt. Same key MUST be used across all retries of the same operation.
Circuit Breaker Pattern
For any external dependency, implement a circuit breaker:
State: CLOSED (normal)
- All requests pass through
- Track rolling failure rate over last 20 requests
State: OPEN (tripped)
- Triggered when failure rate > 50% over 20+ requests OR 5 consecutive failures
- All requests fail immediately with
error.code = "CIRCUIT_OPEN" - Skip retries entirely while open
- Duration: 30 seconds, then transition to HALF_OPEN
State: HALF_OPEN (probing)
- Allow 1 request through
- If success: transition to CLOSED, reset counters
- If failure: return to OPEN with doubled duration (max 5 minutes)
LLM-Specific Retries
Completion errors
- Context length exceeded: NO retry. Truncate or summarize before next attempt.
- Content filter triggered: NO retry. Modify prompt or escalate.
- Empty response: Retry once with
temperature += 0.2. If still empty, fail. - Invalid JSON when JSON mode requested: Retry up to 2 times with stricter prompt instructions.
Tool call errors
- Hallucinated tool name: NO retry. Re-prompt with explicit tool list.
- Invalid tool arguments: Pass error back to LLM, allow it to correct. Max 3 corrections per tool call.
- Tool execution failure: Apply standard error classification above.
Logging Requirements
Every retry MUST log:
attempt_number(1-indexed)error_class(from classification table)error_codedelay_before_retry_mstotal_elapsed_msoperation_id(for correlating retry sequences)
On final failure, log a single summary entry with attempts_total and the chain of error codes.
Anti-Patterns
A1. Infinite retry loops. ALWAYS enforce max attempts.
A2. Retrying without backoff. Causes thundering herd, makes downstream outages worse.
A3. Retrying with the same Idempotency-Key after the operation succeeded but response was lost. Some systems reject duplicate keys after a TTL — handle this case explicitly.
A4. Retrying validation errors verbatim. Wastes tokens. The input is wrong; do not retry the same input.
A5. Catching all errors as transient. Misclassification causes data corruption (e.g., charging a customer twice).
A6. Ignoring Retry-After headers. Causes immediate 429 on retry, wasting another quota slot.
A7. Retrying inside a tool call AND inside the agent loop. Causes exponential request multiplication. Pick one layer.
Pseudocode Reference
def execute_with_retry(operation, policy):
idempotency_key = uuid4()
last_error = None
for attempt in range(1, policy.max_attempts + 1):
try:
return operation(idempotency_key=idempotency_key)
except Exception as e:
error_class = classify(e)
last_error = e
if not error_class.retryable:
raise
if attempt == policy.max_attempts:
break
delay = compute_backoff(
attempt=attempt,
base=policy.base_delay,
max_delay=policy.max_delay,
retry_after=getattr(e, "retry_after", None)
)
log_retry(attempt, error_class, delay)
sleep(delay)
raise RetryExhausted(last_error, attempts=policy.max_attempts)
Verification Checklist
- [ ] Errors classified before retry decision
- [ ] Max attempts enforced (default 3)
- [ ] Exponential backoff with full jitter implemented
- [ ]
Retry-Afterhonored on 429 - [ ] Idempotency key generated per operation, reused across retries
- [ ] Circuit breaker present for each external dependency
- [ ] Non-idempotent operations never blindly retried
- [ ] All retry attempts logged with structured metadata
Related Articles
- Error Handling and Retry Policies Inside Agent Loops — Decision rules for classifying agent errors and configuring retry, backoff, idempotency and escalation policies.
- Prompt Caching Protocols — Implementation Reference for Agents — Reference for using prompt caching to reduce token costs and latency in agent systems. Covers Anthropic, OpenAI, and Gemini cache mechanics.
- Knowledge Grounding and Citation Protocols — Agent Reference — Reference for grounding agent outputs in retrieved sources and producing verifiable citations. Covers retrieval, attribution, and conflict resolution.
- Implementing Smart Retry Logic for Transient Errors — Apply intelligent retry patterns with exponential backoff to handle temporary failures without overwhelming services.
- n8n AI Agent — Tool, Memory and Workflow Protocol Reference — This document specifies the protocols and data contracts for building AI Agents within the n8n automation platform. It provides a machine-readable reference for developers and autonomous agents on how to construct and interact with n8n Tool