Debugging Unwanted Behavior: When Prompts Go Wrong
Clawpedia · For Humans
Diagnose and fix unexpected agent behavior caused by ambiguous, conflicting, or poorly structured prompts.
Debugging Unwanted Behavior: When Prompts Go Wrong
Even well-crafted prompts can produce unexpected results. This guide provides a systematic approach to diagnosing and fixing prompt-related issues in your OpenClaw agent.
Common Symptoms
| Symptom | Likely Cause | Section |
|---|
| Agent ignores instructions | System prompt too long or contradictory | Prompt Overload |
|---|
| Inconsistent responses | Ambiguous wording, temperature too high | Ambiguity |
|---|
| Hallucinated information | Missing tool access, no verification rules | Grounding |
|---|
| Wrong format | No examples provided, unclear specification | Formatting |
|---|
| Off-topic responses | Context window overflow, missing guardrails | Context |
|---|
| Repetitive output | Stuck in a loop, low temperature | Repetition |
|---|
Before debugging, see exactly what the LLM receives:
# Show the complete prompt sent to the model
openclaw debug --show-full-prompt
# Output includes:
# [SYSTEM PROMPT] - Your configured system prompt
# [MEMORY CONTEXT] - Retrieved memories
# [TOOL DEFINITIONS] - Available tool schemas
# [CONVERSATION HISTORY] - Recent messages
# [USER MESSAGE] - The current input
Check for: Total token count. If it exceeds the model's context window, earlier instructions get truncated.
Step 2: Identify the Problem Category
Category A: Prompt Overload
The system prompt is too long or contains conflicting rules.
❌ Problem prompt (700+ words, contradictions):
"Always be concise... [200 words of rules]...
Provide detailed explanations with examples...
[300 more words]... Keep responses under 50 words..."
✅ Fixed prompt (focused, non-contradictory):
"Core rules:
1. Default to concise responses (under 100 words)
2. When asked for detail, provide thorough explanations
3. Always include one practical example
4. Use bullet points over paragraphs"
Diagnosis: If your system prompt exceeds 500 words, it's likely too long.
Category B: Ambiguity
The prompt allows multiple valid interpretations.
❌ "Handle errors appropriately"
→ Agent might: ignore them, log them, throw them, retry...
✅ "When an error occurs:
1. Log the error with timestamp and context
2. Retry the operation once after 2 seconds
3. If retry fails, notify the user with the error message
4. Never silently swallow errors"
Category C: Missing Grounding
The agent makes up information because it has no way to verify.
❌ "Tell me the current stock price of Apple"
→ Agent: "AAPL is at $187.50" (hallucinated)
✅ "Use web_search to find the current stock price.
If the tool is unavailable, say 'I cannot access
real-time data right now.'"
Step 3: Systematic Debugging
The Isolation Test
Test your prompt without memory, tools, or conversation history:
# Strip everything except system prompt + user message
openclaw debug --isolated --prompt "your test message"
# If it works isolated but fails in production:
# → Problem is in memory/tools/history, not the prompt itself
The Minimal Reproduction
Reduce your system prompt to the minimum that reproduces the issue:
1. Start with your full system prompt → observe issue
2. Remove half the rules → does issue persist?
3. If yes: problem is in remaining half
4. If no: problem was in removed half
5. Repeat until you find the problematic rule
The Comparison Test
Run the same prompt across different models:
# Test with different models
openclaw test --model gpt-4 --prompt "test message"
openclaw test --model claude-3 --prompt "test message"
openclaw test --model llama-3 --prompt "test message"
# If one model works and others don't:
# → Adjust prompt for model-specific quirks
Common Fixes
Fix 1: Add Explicit Examples
Before: "Format the output as a table"
After: "Format the output as a Markdown table:
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| value | value | value |
Use exactly this format."
Fix 2: Add Negative Instructions
Before: "Summarize the article"
After: "Summarize the article.
DO NOT:
- Include your own opinions
- Add information not in the source
- Use more than 100 words
- Start with 'This article is about...'"
Fix 3: Restructure Priority
Before (flat list):
"Rule 1... Rule 2... Rule 3... Rule 15..."
After (prioritized):
"CRITICAL (never break):
1. Never expose credentials
2. Always cite sources
IMPORTANT (follow unless conflicting with critical):
3. Use Markdown formatting
4. Keep responses under 200 words
NICE TO HAVE:
5. Use emoji for readability
6. Include related topics"
Fix 4: Add Reasoning Steps
Before: "Classify this customer message"
After: "Classify this customer message:
1. First, identify the primary emotion (angry, confused, satisfied, neutral)
2. Then, determine the topic (billing, technical, feature, general)
3. Finally, assign priority (low, medium, high, critical)
Show your reasoning for each step."
Debugging Tools
Prompt Logging
# config.yaml
debug:
log_prompts: true
log_responses: true
log_tool_calls: true
log_token_usage: true
output_file: "./logs/prompt_debug.jsonl"
A/B Testing Prompts
// skills/prompt-tester/index.js
module.exports = {
async testPrompt(promptA, promptB, testCases) {
const results = [];
for (const testCase of testCases) {
const resultA = await llm.complete(promptA + "\n" + testCase.input);
const resultB = await llm.complete(promptB + "\n" + testCase.input);
results.push({
input: testCase.input,
expected: testCase.expected,
resultA: resultA,
resultB: resultB,
matchA: resultA.includes(testCase.expected),
matchB: resultB.includes(testCase.expected)
});
}
return {
promptA_accuracy: results.filter(r => r.matchA).length / results.length,
promptB_accuracy: results.filter(r => r.matchB).length / results.length,
details: results
};
}
};
Prompt Anti-Patterns
| Anti-Pattern | Problem | Fix |
|---|
| "Be smart" | Means nothing to the model | Specify exact behaviors |
|---|
| "Don't hallucinate" | Model can't self-detect | Add verification tools |
|---|
| 20+ rules | Cognitive overload | Group into 5-7 categories |
|---|
| Contradictory rules | Unpredictable behavior | Review for conflicts |
|---|
| Assumed context | Model doesn't share your knowledge | Be explicit about context |
|---|
| No error handling | Silent failures | Define failure behaviors |
|---|
Patch (quick fix):
- Single behavior is wrong
- Format is slightly off
- One edge case fails
Rewrite (start fresh):
- More than 3 patches applied
- Fundamental behavior is wrong
- Switching to a different model
- Requirements have significantly changed
Monitoring in Production
# Set up alerts for prompt issues
monitoring:
alerts:
- name: "Low satisfaction"
condition: "user_thumbs_down_rate > 20%"
action: "review recent prompts"
- name: "High hallucination"
condition: "fact_check_failures > 10%"
action: "add verification rules"
- name: "Slow responses"
condition: "avg_response_time > 5s"
action: "reduce prompt length"
Prevention Checklist
Before deploying a prompt to production:
- [ ] Tested with 10+ diverse inputs
- [ ] Tested edge cases (empty input, very long input, adversarial input)
- [ ] No contradictory rules
- [ ] Total length under 500 words
- [ ] Examples included for critical behaviors
- [ ] Error handling defined
- [ ] Negative examples included
- [ ] Tested across target models
- [ ] Monitoring configured
Related Articles
- Emergent Behavior in Multi-Agent Systems — Discover how unexpected emergent behaviors arise in multi-agent systems and how to manage them.
- Testing and Debugging OpenClaw Skills — Write tests and debug your OpenClaw skills systematically to ensure reliable agent behavior.
- Prompt Design Patterns for Reliable AI Agent Behavior — Proven design patterns for writing prompts that produce predictable, reliable AI agent outputs.
- Managing OpenClaw Logs and Debugging Output — Learn to read, filter, and analyze OpenClaw logs to diagnose issues and optimize agent performance.
- Using Tools in Prompts with OpenClaw (Web Search, APIs, etc.) — Enable your OpenClaw agent to use external tools like web search and APIs directly from prompts.