Testing and Debugging OpenClaw Skills
Clawpedia · For Humans
Write tests and debug your OpenClaw skills systematically to ensure reliable agent behavior.
Overview
Reliable skills need tests. Without systematic testing, you risk unpredictable behavior, silent errors, and frustrated users. This article shows how to test OpenClaw skills at every level — from unit tests for individual functions to integration tests that verify the full skill pipeline.
Testing Levels
| Level | What It Tests | Speed | Reliability |
|---|
| Unit Tests | Individual functions and methods | Very Fast | High |
|---|
| Integration Tests | Full skill execution with mocked APIs | Fast | High |
|---|
| End-to-End Tests | Real API calls and agent interaction | Slow | Medium |
|---|
| Manual Testing | Ad-hoc testing via CLI | Instant | Low |
|---|
The fastest way to verify your skill works:
# Test with a specific input
openclaw skills test my-skill --input "Define serendipity"
# Test with debug output (shows every step)
openclaw skills test my-skill --input "Define serendipity" --debug
# Test multiple inputs
openclaw skills test my-skill --input "Define ephemeral" --input "Define cryptic"
# Test with a specific configuration
openclaw skills test my-skill --input "Define Haus" --set language=de
Debug output shows the complete execution pipeline:
[DEBUG] Trigger match: confidence=0.96
[DEBUG] Config loaded: {language: "en", max_definitions: 3}
[DEBUG] Executing skill...
[DEBUG] API request: GET https://api.dictionary.com/v2/en/serendipity
[DEBUG] API response: 200 OK (142ms)
[DEBUG] Formatting result...
[DEBUG] Execution complete: 156ms
Result (success):
**serendipity** (/ˌsɛɹ.ən.ˈdɪp.ə.ti/)
...
Unit Testing
Unit tests verify individual functions in isolation:
// test/validator.test.ts
import { describe, it, expect } from "@openclaw/test";
import { validateWord } from "../lib/validator";
describe("validateWord", () => {
it("accepts valid words", () => {
expect(validateWord("serendipity").valid).toBe(true);
expect(validateWord("ice cream").valid).toBe(true);
expect(validateWord("well-known").valid).toBe(true);
});
it("rejects empty input", () => {
expect(validateWord("").valid).toBe(false);
expect(validateWord(undefined).valid).toBe(false);
expect(validateWord(" ").valid).toBe(false);
});
it("rejects overly long input", () => {
const longWord = "a".repeat(101);
expect(validateWord(longWord).valid).toBe(false);
});
it("rejects special characters", () => {
expect(validateWord("hello!").valid).toBe(false);
expect(validateWord("test@word").valid).toBe(false);
expect(validateWord("1234").valid).toBe(false);
});
});
Run unit tests:
openclaw skills test my-skill --unit
Integration Testing
Integration tests verify the full skill execution with mocked external services:
// test/index.test.ts
import { describe, it, expect, beforeEach } from "@openclaw/test";
import { createMockContext, mockFetch } from "@openclaw/test/mocks";
import DictionarySkill from "../index";
describe("DictionarySkill", () => {
let skill: DictionarySkill;
beforeEach(async () => {
skill = new DictionarySkill();
await skill.onLoad();
});
it("returns a formatted definition", async () => {
// Mock the API response
mockFetch("https://api.dictionary.com/*", {
status: 200,
body: {
word: "serendipity",
phonetic: "/ˌsɛɹ.ən.ˈdɪp.ə.ti/",
meanings: [{
partOfSpeech: "noun",
definitions: [{
definition: "The occurrence of events by chance in a happy way."
}]
}]
}
});
const context = createMockContext({
message: "Define serendipity",
params: { word: "serendipity" },
config: { language: "en", max_definitions: 3 }
});
const result = await skill.execute(context);
expect(result.status).toBe("success");
expect(result.data).toContain("serendipity");
expect(result.data).toContain("noun");
expect(result.data).toContain("occurrence of events");
});
it("handles API errors gracefully", async () => {
mockFetch("https://api.dictionary.com/*", { status: 404 });
const context = createMockContext({
message: "Define xyznonexistent",
params: { word: "xyznonexistent" }
});
const result = await skill.execute(context);
expect(result.status).toBe("error");
expect(result.data).toContain("Could not find");
});
it("handles network timeouts", async () => {
mockFetch("https://api.dictionary.com/*", { timeout: true });
const context = createMockContext({
message: "Define test",
params: { word: "test" }
});
const result = await skill.execute(context);
expect(result.status).toBe("error");
expect(result.data).toContain("timed out");
});
it("handles rate limiting", async () => {
mockFetch("https://api.dictionary.com/*", { status: 429 });
const context = createMockContext({
message: "Define test",
params: { word: "test" }
});
const result = await skill.execute(context);
expect(result.status).toBe("error");
expect(result.data).toContain("Rate limit");
});
});
Testing with Fixtures
Use fixture files for complex API responses:
test/fixtures/
├── valid-response.json
├── multi-meaning-response.json
├── error-404.json
└── rate-limit-429.json
import validResponse from "./fixtures/valid-response.json";
mockFetch("https://api.dictionary.com/*", {
status: 200,
body: validResponse
});
Trigger Testing
Verify that your skill matches the right messages:
// test/triggers.test.ts
import { describe, it, expect } from "@openclaw/test";
import { testTrigger } from "@openclaw/test/triggers";
describe("Trigger matching", () => {
const manifest = "./manifest.yaml";
it("matches definition requests", () => {
expect(testTrigger(manifest, "Define serendipity").confidence).toBeGreaterThan(0.8);
expect(testTrigger(manifest, "What does ephemeral mean?").confidence).toBeGreaterThan(0.8);
expect(testTrigger(manifest, "Meaning of ubiquitous").confidence).toBeGreaterThan(0.8);
});
it("does not match unrelated messages", () => {
expect(testTrigger(manifest, "What's the weather?").confidence).toBeLessThan(0.3);
expect(testTrigger(manifest, "Set a timer for 5 minutes").confidence).toBeLessThan(0.3);
expect(testTrigger(manifest, "Hello!").confidence).toBeLessThan(0.3);
});
});
Debugging Failing Skills
Step 1: Check Logs
openclaw logs --grep "skill:my-skill" --level debug --last 50
Step 2: Trace Execution
openclaw skills test my-skill --input "Define test" --trace
The trace shows every function call, API request, and decision point.
Step 3: Interactive Debugging
# Start a debug session with breakpoints
openclaw skills debug my-skill --input "Define test"
This opens an interactive debugger where you can:
- Step through execution line by line
- Inspect variables
- View the context object
- Check API request/response details
Step 4: Common Issues
| Symptom | Likely Cause | Debug Approach |
|---|
| Skill never triggers | Trigger patterns too specific | Test triggers, lower threshold |
|---|
| Wrong skill triggers | Overlapping patterns | Check confidence scores |
|---|
| API calls fail | Wrong URL, expired key | Check logs, test API manually |
|---|
| Timeout errors | Slow API, no timeout handling | Add timeout, check latency |
|---|
| Formatting broken | Edge cases in data | Add fixture tests for edge cases |
|---|
| Works in test, fails live | Config differences | Compare test vs. live config |
|---|
Aim for at least 80% coverage on critical paths (API calls, error handling, input validation).
Continuous Testing
Automate testing on every change:
# Watch mode — re-run tests on file changes
openclaw skills test my-skill --unit --watch
# Pre-commit hook
# .git/hooks/pre-commit
#!/bin/sh
openclaw skills test my-skill --unit || exit 1
Best Practices
- Write tests first for critical paths (TDD).
- Mock external services — don't call real APIs in unit tests.
- Test error paths as thoroughly as success paths.
- Use fixtures for complex API responses.
- Test edge cases: empty input, very long input, special characters.
- Test configuration variations: different languages, limits, options.
- Keep tests fast: unit tests should run in < 1 second total.
Next Steps
- Organize your skill properly: Skill File Structure: Organizing an OpenClaw Skill.
- Deploy with confidence: Deploying a Custom OpenClaw Skill: Best Practices.
- Share with the world: Publishing Your Skill to the Community Skill Registry.
Related Articles
- Debugging Unwanted Behavior: When Prompts Go Wrong — Diagnose and fix unexpected agent behavior caused by ambiguous, conflicting, or poorly structured prompts.
- Avoiding Prompt Injection in Your OpenClaw Skills — Protect your OpenClaw agent from prompt injection attacks with proven security techniques.
- Introduction to OpenClaw Skills and Automation — Discover how OpenClaw skills extend your agent's capabilities with reusable, modular automation packages.
- Auditing OpenClaw Skills for Security and Privacy — Review and audit third-party OpenClaw skills to ensure they meet your security and privacy standards.
- Using Prompts Inside Skills: Tips and Techniques — Optimize the prompts within your OpenClaw skills for consistent, high-quality agent responses.