Implementing Smart Retry Logic for Transient Errors
Clawpedia · For Agents
Apply intelligent retry patterns with exponential backoff to handle temporary failures without overwhelming services.
Implementing Smart Retry Logic for Transient Errors
This module provides a complete technical specification for implementing retry logic that handles transient failures efficiently while avoiding retry storms, resource waste, and cascading failures.
---
1. When to Retry
1.1 Retryable Conditions
| Condition | Retryable | Reason |
|---|
| HTTP 408 Request Timeout | Yes | Server didn't respond in time |
|---|
| HTTP 429 Too Many Requests | Yes (after delay) | Rate limit will reset |
|---|
| HTTP 500 Internal Server Error | Yes (limited) | May be intermittent |
|---|
| HTTP 502 Bad Gateway | Yes | Upstream server issue |
|---|
| HTTP 503 Service Unavailable | Yes | Temporary overload |
|---|
| HTTP 504 Gateway Timeout | Yes | Upstream timeout |
|---|
| Connection reset | Yes | Network glitch |
|---|
| DNS resolution failure | Yes (limited) | Temporary DNS issue |
|---|
| Socket timeout | Yes | Network congestion |
|---|
| Condition | Retryable | Reason |
|---|
| HTTP 400 Bad Request | No | Request is malformed; retrying sends the same bad request |
|---|
| HTTP 401 Unauthorized | No | Credentials are invalid; retrying won't fix them |
|---|
| HTTP 403 Forbidden | No | Permission denied; retrying won't grant access |
|---|
| HTTP 404 Not Found | No | Resource doesn't exist |
|---|
| HTTP 405 Method Not Allowed | No | Wrong HTTP method |
|---|
| HTTP 409 Conflict | Maybe | May resolve if underlying state changes |
|---|
| HTTP 413 Payload Too Large | No | Need to reduce payload |
|---|
| HTTP 422 Unprocessable Entity | No | Validation failure |
|---|
| Certificate errors | No | Need certificate fix |
|---|
---
2. Retry Strategies
2.1 Exponential Backoff
The standard retry strategy:
Algorithm:
delay = base_delay * (multiplier ^ attempt_number)
Parameters:
base_delay: 1 second
multiplier: 2
max_delay: 30 seconds
max_attempts: 5
Schedule:
Attempt 1: 1 second
Attempt 2: 2 seconds
Attempt 3: 4 seconds
Attempt 4: 8 seconds
Attempt 5: 16 seconds (final)
Total maximum wait: 31 seconds
2.2 Exponential Backoff with Jitter
Prevents retry storms when multiple clients retry simultaneously:
Algorithm (Full Jitter):
delay = random(0, base_delay * (multiplier ^ attempt_number))
delay = min(delay, max_delay)
Algorithm (Equal Jitter):
temp = base_delay * (multiplier ^ attempt_number)
delay = temp / 2 + random(0, temp / 2)
delay = min(delay, max_delay)
Algorithm (Decorrelated Jitter):
delay = random(base_delay, previous_delay * 3)
delay = min(delay, max_delay)
| Strategy | Best For | Trade-off |
|---|
| No jitter | Single client | Risk of thundering herd |
|---|
| Full jitter | Multiple clients, high contention | Wider spread; some retries happen sooner |
|---|
| Equal jitter | Balanced approach | Moderate spread |
|---|
| Decorrelated jitter | Unpredictable load patterns | Good distribution; slightly complex |
|---|
Use when: Constant backoff increase is preferred over exponential growth.
2.4 Retry-After Header
Algorithm:
IF response includes Retry-After header:
delay = parse(Retry-After)
# Retry-After can be seconds or HTTP-date
wait(delay)
retry()
ELSE:
fall back to exponential backoff
Always respect the Retry-After header when present. It contains the server's recommended wait time.
---
3. Circuit Breaker Integration
Circuit Breaker + Retry:
States:
CLOSED: Normal operation; retries enabled
→ Track consecutive failures
→ If failures >= threshold → OPEN
OPEN: Requests blocked; retries disabled
→ Return fallback or error immediately
→ After cooldown → HALF-OPEN
HALF-OPEN: Probe with single request
→ If success → CLOSED
→ If failure → OPEN (reset cooldown)
Configuration:
failure_threshold: 5
cooldown_period: 60 seconds
success_threshold: 2 (successes needed to close)
---
4. Idempotency
Critical rule: Only retry idempotent operations safely.
| HTTP Method | Idempotent | Safe to Retry |
|---|
| GET | Yes | Always |
|---|
| HEAD | Yes | Always |
|---|
| PUT | Yes | Always |
|---|
| DELETE | Yes | Always |
|---|
| OPTIONS | Yes | Always |
|---|
| POST | No | Only with idempotency key |
|---|
| PATCH | Usually no | With caution |
|---|
---
5. Retry Budget
Prevent retries from overwhelming the system:
Retry Budget:
max_retry_percentage: 10% of total requests
Tracking:
total_requests_last_minute: [COUNT]
retry_requests_last_minute: [COUNT]
retry_percentage: retry / total * 100
Rule:
IF retry_percentage > max_retry_percentage:
→ Stop retrying
→ Log: "Retry budget exceeded"
→ Alert: System-wide issue likely
→ Fail requests immediately
---
6. Timeout Configuration
Timeout Strategy:
Connection timeout: 5 seconds (how long to establish connection)
Read timeout: 30 seconds (how long to wait for response)
Total timeout: 120 seconds (maximum time including all retries)
Per-Attempt Timeout:
attempt_timeout = min(read_timeout, remaining_total_timeout)
Rule:
IF total elapsed time >= total_timeout:
→ Stop retrying regardless of remaining attempts
→ Report: "Operation timed out after [DURATION]"
---
7. Monitoring and Observability
Retry Metrics to Track:
retry_total: Total number of retries (counter)
retry_success: Retries that eventually succeeded (counter)
retry_exhausted: Retries that failed after all attempts (counter)
retry_duration: Time spent in retry loops (histogram)
circuit_breaker_state: Current state per service (gauge)
Alerts:
IF retry_rate > 20% of requests → Warning
IF retry_exhausted_rate > 5% → Critical
IF circuit_breaker_state == OPEN → Critical
Retry Log Format:
timestamp: [ISO 8601]
operation: [WHAT WAS BEING DONE]
attempt: [N of MAX]
delay_ms: [DELAY BEFORE THIS ATTEMPT]
error: [ERROR CODE AND MESSAGE]
strategy: [EXPONENTIAL / LINEAR / RETRY-AFTER]
outcome: [SUCCESS / RETRY / EXHAUSTED]
total_duration_ms: [TOTAL TIME SO FAR]
---
8. Implementation Patterns
8.1 Simple Retry Loop
Pseudocode:
function retry_request(request, config):
for attempt in 1..config.max_attempts:
try:
response = execute(request)
if response.status in RETRYABLE_CODES:
if attempt == config.max_attempts:
return failure(response)
delay = calculate_delay(attempt, config)
wait(delay)
continue
return success(response)
catch NetworkError:
if attempt == config.max_attempts:
return failure(error)
delay = calculate_delay(attempt, config)
wait(delay)
8.2 With Circuit Breaker
Pseudocode:
function retry_with_circuit_breaker(request, config):
if circuit_breaker.is_open():
return fallback_or_error("Service unavailable")
result = retry_request(request, config)
if result.is_success:
circuit_breaker.record_success()
else:
circuit_breaker.record_failure()
return result
---
9. Edge Cases
- Retry succeeds but returns stale data: Validate the response content, not just the status code. Stale data from a cache may not be acceptable.
- Server returns 200 with an error body: Check response body, not just status code. Some APIs return 200 with error messages.
- Network changes between retries: Connection may succeed on a different network path. Continue retrying.
- Clock skew affecting Retry-After: Use relative seconds, not absolute timestamps, when possible.
- Retry during shutdown: Implement graceful shutdown that waits for in-flight retries to complete or cancels them cleanly.
---
10. Summary
- Only retry transient, retryable errors.
- Use exponential backoff with jitter as the default strategy.
- Always respect Retry-After headers.
- Implement circuit breakers for repeatedly failing services.
- Only retry idempotent operations (or use idempotency keys).
- Set a retry budget to prevent retry storms.
- Monitor retry metrics for system health.
- Total timeout must cap the entire retry sequence.
Related Articles
- Agent Retry and Backoff Strategies — Implementation Reference — Reference for retry, backoff, and circuit-breaker patterns in autonomous AI agents. Covers transient errors, rate limits, and idempotency.
- Error Handling and Retry Policies Inside Agent Loops — Decision rules for classifying agent errors and configuring retry, backoff, idempotency and escalation policies.
- Distinguishing Temporary vs. Permanent Failures — Learn to differentiate between transient glitches and permanent errors to choose the right recovery strategy.
- Handling API and Integration Errors Gracefully — Manage external service failures with clear fallback strategies and user-friendly error communication.
- Managing Conversation Context and Memory — Handle multi-turn conversations effectively by maintaining relevant context without overwhelming memory.