Manage external service failures with clear fallback strategies and user-friendly error communication.
Handling API and Integration Errors Gracefully
This module provides a comprehensive framework for managing errors from external APIs and third-party integrations. Graceful error handling means the user never sees a raw stack trace—only clear information and actionable options.
---
1. Error Interception Layer
Error Handling Pipeline:
API Call
→ Response received
→ Step 1: Check HTTP status code
→ Step 2: Parse error response body
→ Step 3: Classify the error
→ Step 4: Determine recovery strategy
→ Step 5: Execute recovery or report to user
→ Step 6: Log the event
---
2. Error Classification Matrix
HTTP Status
Category
Retryable
User Action Required
400
Client error: bad request
No
Fix request parameters
401
Auth: invalid credentials
No
Re-authenticate
403
Auth: insufficient permissions
No
Request access
404
Resource not found
No
Verify resource exists
408
Timeout
Yes
Retry with longer timeout
409
Conflict
Maybe
Resolve conflicting state
413
Payload too large
No
Reduce payload size
422
Validation error
No
Fix input data
429
Rate limited
Yes (after delay)
Wait and retry
500
Server error
Yes (limited)
Report if persists
502
Bad gateway
Yes
Retry after delay
503
Service unavailable
Yes
Retry with backoff
504
Gateway timeout
Yes
Retry with longer timeout
---
3. Recovery Strategies
3.1 Automatic Retry
For transient errors (429, 500, 502, 503, 504):
Retry Configuration:
max_retries: 3
initial_delay: 1 second
backoff_multiplier: 2
max_delay: 30 seconds
jitter: random 0-500ms added to each delay
Retry Schedule:
Attempt 1: immediate
Attempt 2: 1s + jitter
Attempt 3: 2s + jitter
Attempt 4: 4s + jitter (final)
3.2 Rate Limit Handling
Rate Limit Protocol:
1. Check for Retry-After header
→ If present: wait the specified duration
→ If absent: use exponential backoff
2. Check for X-RateLimit-Remaining header
→ If available: preemptively slow down before hitting zero
3. Log rate limit events for future request scheduling
4. Report to user if rate limits significantly delay the task
3.3 Circuit Breaker
For repeatedly failing integrations:
Circuit Breaker States:
CLOSED (normal operation):
→ Track failure count
→ If failures >= threshold → Switch to OPEN
OPEN (blocking requests):
→ Return cached data or fallback
→ After cooldown period → Switch to HALF-OPEN
HALF-OPEN (testing):
→ Allow one request through
→ If success → Switch to CLOSED
→ If failure → Switch to OPEN
Configuration:
failure_threshold: 5 consecutive failures
cooldown_period: 60 seconds
half_open_max_requests: 1
3.4 Fallback Strategies
Strategy
When to Use
Implementation
Cached response
Recent cache available
Return cached data with "as of [TIME]" notice
Degraded mode
Non-critical feature fails
Continue without the feature; notify user
Alternative API
Backup service available
Transparently switch; notify user
Manual override
All automated recovery fails
Present options to user
Queue for later
Non-urgent task
Save request; retry when service recovers
---
4. Error Response Parsing
4.1 Common API Error Formats
Format 1 (Standard):
{ "error": { "code": "RATE_LIMIT", "message": "Too many requests" } }
Format 2 (Simple):
{ "error": "Invalid parameter: email" }
Format 3 (Detailed):
{ "errors": [{ "field": "email", "code": "invalid", "message": "Not a valid email" }] }
Format 4 (HTML):
<html><body><h1>503 Service Unavailable</h1></body></html>
4.2 Parsing Protocol
Parsing Steps:
1. Check Content-Type header
2. If JSON → Parse and extract error message/code
3. If HTML → Extract text content from error page
4. If empty → Use HTTP status code for classification
5. If unparseable → Log raw response; use status code
---
5. User Communication
5.1 Error Report Template
Template:
"The [SERVICE NAME] integration encountered an issue.
What happened: [PLAIN ENGLISH DESCRIPTION]
Impact: [WHAT THIS MEANS FOR THE USER'S TASK]
What I'm doing:
- [AUTOMATIC RECOVERY ACTION, IF ANY]
What you can do:
1. [OPTION 1]
2. [OPTION 2]
Technical detail (for support): [ERROR CODE]"
5.2 Context-Specific Templates
Authentication Error:
"I can't connect to [SERVICE] — your login credentials have expired.
To fix: Re-authenticate with [SERVICE] and try again."
Rate Limit:
"[SERVICE] is limiting our request rate.
I'll automatically retry in [TIME]. No action needed from you."
Service Down:
"[SERVICE] is currently unavailable.
I'll retry periodically. In the meantime:
- [WHAT CAN BE DONE WITHOUT THE SERVICE]
- [WHEN TO TRY AGAIN]"
---
6. Integration-Specific Guidelines
6.1 REST APIs
Issue
Handling
Pagination errors
Verify cursor/offset; retry from last successful page
Handling Ambiguous User Requests Gracefully — Protocols for detecting ambiguity in user prompts and resolving it through clarification, inference, or safe default behavior.
Effective Error Handling and Uncertainty Recognition — A comprehensive guide for AI agents on recognizing uncertainty, handling errors gracefully, and avoiding the fabrication of facts when knowledge is insufficient.