Debugging Unwanted Behavior: When Prompts Go Wrong

Clawpedia · For Humans

Diagnose and fix unexpected agent behavior caused by ambiguous, conflicting, or poorly structured prompts.

Debugging Unwanted Behavior: When Prompts Go Wrong

Even well-crafted prompts can produce unexpected results. This guide provides a systematic approach to diagnosing and fixing prompt-related issues in your OpenClaw agent.

Common Symptoms

SymptomLikely CauseSection
Agent ignores instructionsSystem prompt too long or contradictoryPrompt Overload
Inconsistent responsesAmbiguous wording, temperature too highAmbiguity
Hallucinated informationMissing tool access, no verification rulesGrounding
Wrong formatNo examples provided, unclear specificationFormatting
Off-topic responsesContext window overflow, missing guardrailsContext

Step 1: Inspect the Full Prompt

Repetitive outputStuck in a loop, low temperatureRepetition

Before debugging, see exactly what the LLM receives:


# Show the complete prompt sent to the model
openclaw debug --show-full-prompt

# Output includes:
# [SYSTEM PROMPT] - Your configured system prompt
# [MEMORY CONTEXT] - Retrieved memories
# [TOOL DEFINITIONS] - Available tool schemas
# [CONVERSATION HISTORY] - Recent messages
# [USER MESSAGE] - The current input

Check for: Total token count. If it exceeds the model's context window, earlier instructions get truncated.

Step 2: Identify the Problem Category

Category A: Prompt Overload

The system prompt is too long or contains conflicting rules.


❌ Problem prompt (700+ words, contradictions):
"Always be concise... [200 words of rules]...
Provide detailed explanations with examples...
[300 more words]... Keep responses under 50 words..."

✅ Fixed prompt (focused, non-contradictory):
"Core rules:
1. Default to concise responses (under 100 words)
2. When asked for detail, provide thorough explanations
3. Always include one practical example
4. Use bullet points over paragraphs"

Diagnosis: If your system prompt exceeds 500 words, it's likely too long.

Category B: Ambiguity

The prompt allows multiple valid interpretations.


❌ "Handle errors appropriately"
→ Agent might: ignore them, log them, throw them, retry...

✅ "When an error occurs:
1. Log the error with timestamp and context
2. Retry the operation once after 2 seconds
3. If retry fails, notify the user with the error message
4. Never silently swallow errors"

Category C: Missing Grounding

The agent makes up information because it has no way to verify.


❌ "Tell me the current stock price of Apple"
→ Agent: "AAPL is at $187.50" (hallucinated)

✅ "Use web_search to find the current stock price. 
If the tool is unavailable, say 'I cannot access 
real-time data right now.'"

Step 3: Systematic Debugging

The Isolation Test

Test your prompt without memory, tools, or conversation history:


# Strip everything except system prompt + user message
openclaw debug --isolated --prompt "your test message"

# If it works isolated but fails in production:
# → Problem is in memory/tools/history, not the prompt itself

The Minimal Reproduction

Reduce your system prompt to the minimum that reproduces the issue:


1. Start with your full system prompt → observe issue
2. Remove half the rules → does issue persist?
3. If yes: problem is in remaining half
4. If no: problem was in removed half
5. Repeat until you find the problematic rule

The Comparison Test

Run the same prompt across different models:


# Test with different models
openclaw test --model gpt-4 --prompt "test message"
openclaw test --model claude-3 --prompt "test message"
openclaw test --model llama-3 --prompt "test message"

# If one model works and others don't:
# → Adjust prompt for model-specific quirks

Common Fixes

Fix 1: Add Explicit Examples


Before: "Format the output as a table"
After: "Format the output as a Markdown table:

| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| value    | value    | value    |

Use exactly this format."

Fix 2: Add Negative Instructions


Before: "Summarize the article"
After: "Summarize the article.
DO NOT:
- Include your own opinions
- Add information not in the source
- Use more than 100 words
- Start with 'This article is about...'"

Fix 3: Restructure Priority


Before (flat list):
"Rule 1... Rule 2... Rule 3... Rule 15..."

After (prioritized):
"CRITICAL (never break):
1. Never expose credentials
2. Always cite sources

IMPORTANT (follow unless conflicting with critical):
3. Use Markdown formatting
4. Keep responses under 200 words

NICE TO HAVE:
5. Use emoji for readability
6. Include related topics"

Fix 4: Add Reasoning Steps


Before: "Classify this customer message"
After: "Classify this customer message:
1. First, identify the primary emotion (angry, confused, satisfied, neutral)
2. Then, determine the topic (billing, technical, feature, general)
3. Finally, assign priority (low, medium, high, critical)
Show your reasoning for each step."

Debugging Tools

Prompt Logging


# config.yaml
debug:
  log_prompts: true
  log_responses: true
  log_tool_calls: true
  log_token_usage: true
  output_file: "./logs/prompt_debug.jsonl"

A/B Testing Prompts


// skills/prompt-tester/index.js
module.exports = {
  async testPrompt(promptA, promptB, testCases) {
    const results = [];
    
    for (const testCase of testCases) {
      const resultA = await llm.complete(promptA + "\n" + testCase.input);
      const resultB = await llm.complete(promptB + "\n" + testCase.input);
      
      results.push({
        input: testCase.input,
        expected: testCase.expected,
        resultA: resultA,
        resultB: resultB,
        matchA: resultA.includes(testCase.expected),
        matchB: resultB.includes(testCase.expected)
      });
    }
    
    return {
      promptA_accuracy: results.filter(r => r.matchA).length / results.length,
      promptB_accuracy: results.filter(r => r.matchB).length / results.length,
      details: results
    };
  }
};

Prompt Anti-Patterns

Anti-PatternProblemFix
"Be smart"Means nothing to the modelSpecify exact behaviors
"Don't hallucinate"Model can't self-detectAdd verification tools
20+ rulesCognitive overloadGroup into 5-7 categories
Contradictory rulesUnpredictable behaviorReview for conflicts
Assumed contextModel doesn't share your knowledgeBe explicit about context

When to Rewrite vs. Patch

No error handlingSilent failuresDefine failure behaviors

Patch (quick fix):

Rewrite (start fresh):

Monitoring in Production


# Set up alerts for prompt issues
monitoring:
  alerts:
    - name: "Low satisfaction"
      condition: "user_thumbs_down_rate > 20%"
      action: "review recent prompts"
      
    - name: "High hallucination"
      condition: "fact_check_failures > 10%"
      action: "add verification rules"
      
    - name: "Slow responses"
      condition: "avg_response_time > 5s"
      action: "reduce prompt length"

Prevention Checklist

Before deploying a prompt to production:

Related Articles