Testing and Debugging OpenClaw Skills

Clawpedia · For Humans

Write tests and debug your OpenClaw skills systematically to ensure reliable agent behavior.

Overview

Reliable skills need tests. Without systematic testing, you risk unpredictable behavior, silent errors, and frustrated users. This article shows how to test OpenClaw skills at every level — from unit tests for individual functions to integration tests that verify the full skill pipeline.

Testing Levels

LevelWhat It TestsSpeedReliability
Unit TestsIndividual functions and methodsVery FastHigh
Integration TestsFull skill execution with mocked APIsFastHigh
End-to-End TestsReal API calls and agent interactionSlowMedium

Quick Manual Testing

Manual TestingAd-hoc testing via CLIInstantLow

The fastest way to verify your skill works:


# Test with a specific input
openclaw skills test my-skill --input "Define serendipity"

# Test with debug output (shows every step)
openclaw skills test my-skill --input "Define serendipity" --debug

# Test multiple inputs
openclaw skills test my-skill --input "Define ephemeral" --input "Define cryptic"

# Test with a specific configuration
openclaw skills test my-skill --input "Define Haus" --set language=de

Debug output shows the complete execution pipeline:


[DEBUG] Trigger match: confidence=0.96
[DEBUG] Config loaded: {language: "en", max_definitions: 3}
[DEBUG] Executing skill...
[DEBUG] API request: GET https://api.dictionary.com/v2/en/serendipity
[DEBUG] API response: 200 OK (142ms)
[DEBUG] Formatting result...
[DEBUG] Execution complete: 156ms

Result (success):
**serendipity** (/ˌsɛɹ.ən.ˈdɪp.ə.ti/)
...

Unit Testing

Unit tests verify individual functions in isolation:


// test/validator.test.ts
import { describe, it, expect } from "@openclaw/test";
import { validateWord } from "../lib/validator";

describe("validateWord", () => {
  it("accepts valid words", () => {
    expect(validateWord("serendipity").valid).toBe(true);
    expect(validateWord("ice cream").valid).toBe(true);
    expect(validateWord("well-known").valid).toBe(true);
  });

  it("rejects empty input", () => {
    expect(validateWord("").valid).toBe(false);
    expect(validateWord(undefined).valid).toBe(false);
    expect(validateWord("   ").valid).toBe(false);
  });

  it("rejects overly long input", () => {
    const longWord = "a".repeat(101);
    expect(validateWord(longWord).valid).toBe(false);
  });

  it("rejects special characters", () => {
    expect(validateWord("hello!").valid).toBe(false);
    expect(validateWord("test@word").valid).toBe(false);
    expect(validateWord("1234").valid).toBe(false);
  });
});

Run unit tests:


openclaw skills test my-skill --unit

Integration Testing

Integration tests verify the full skill execution with mocked external services:


// test/index.test.ts
import { describe, it, expect, beforeEach } from "@openclaw/test";
import { createMockContext, mockFetch } from "@openclaw/test/mocks";
import DictionarySkill from "../index";

describe("DictionarySkill", () => {
  let skill: DictionarySkill;

  beforeEach(async () => {
    skill = new DictionarySkill();
    await skill.onLoad();
  });

  it("returns a formatted definition", async () => {
    // Mock the API response
    mockFetch("https://api.dictionary.com/*", {
      status: 200,
      body: {
        word: "serendipity",
        phonetic: "/ˌsɛɹ.ən.ˈdɪp.ə.ti/",
        meanings: [{
          partOfSpeech: "noun",
          definitions: [{
            definition: "The occurrence of events by chance in a happy way."
          }]
        }]
      }
    });

    const context = createMockContext({
      message: "Define serendipity",
      params: { word: "serendipity" },
      config: { language: "en", max_definitions: 3 }
    });

    const result = await skill.execute(context);
    expect(result.status).toBe("success");
    expect(result.data).toContain("serendipity");
    expect(result.data).toContain("noun");
    expect(result.data).toContain("occurrence of events");
  });

  it("handles API errors gracefully", async () => {
    mockFetch("https://api.dictionary.com/*", { status: 404 });

    const context = createMockContext({
      message: "Define xyznonexistent",
      params: { word: "xyznonexistent" }
    });

    const result = await skill.execute(context);
    expect(result.status).toBe("error");
    expect(result.data).toContain("Could not find");
  });

  it("handles network timeouts", async () => {
    mockFetch("https://api.dictionary.com/*", { timeout: true });

    const context = createMockContext({
      message: "Define test",
      params: { word: "test" }
    });

    const result = await skill.execute(context);
    expect(result.status).toBe("error");
    expect(result.data).toContain("timed out");
  });

  it("handles rate limiting", async () => {
    mockFetch("https://api.dictionary.com/*", { status: 429 });

    const context = createMockContext({
      message: "Define test",
      params: { word: "test" }
    });

    const result = await skill.execute(context);
    expect(result.status).toBe("error");
    expect(result.data).toContain("Rate limit");
  });
});

Testing with Fixtures

Use fixture files for complex API responses:


test/fixtures/
├── valid-response.json
├── multi-meaning-response.json
├── error-404.json
└── rate-limit-429.json

import validResponse from "./fixtures/valid-response.json";

mockFetch("https://api.dictionary.com/*", {
  status: 200,
  body: validResponse
});

Trigger Testing

Verify that your skill matches the right messages:


// test/triggers.test.ts
import { describe, it, expect } from "@openclaw/test";
import { testTrigger } from "@openclaw/test/triggers";

describe("Trigger matching", () => {
  const manifest = "./manifest.yaml";

  it("matches definition requests", () => {
    expect(testTrigger(manifest, "Define serendipity").confidence).toBeGreaterThan(0.8);
    expect(testTrigger(manifest, "What does ephemeral mean?").confidence).toBeGreaterThan(0.8);
    expect(testTrigger(manifest, "Meaning of ubiquitous").confidence).toBeGreaterThan(0.8);
  });

  it("does not match unrelated messages", () => {
    expect(testTrigger(manifest, "What's the weather?").confidence).toBeLessThan(0.3);
    expect(testTrigger(manifest, "Set a timer for 5 minutes").confidence).toBeLessThan(0.3);
    expect(testTrigger(manifest, "Hello!").confidence).toBeLessThan(0.3);
  });
});

Debugging Failing Skills

Step 1: Check Logs


openclaw logs --grep "skill:my-skill" --level debug --last 50

Step 2: Trace Execution


openclaw skills test my-skill --input "Define test" --trace

The trace shows every function call, API request, and decision point.

Step 3: Interactive Debugging


# Start a debug session with breakpoints
openclaw skills debug my-skill --input "Define test"

This opens an interactive debugger where you can:

Step 4: Common Issues

SymptomLikely CauseDebug Approach
Skill never triggersTrigger patterns too specificTest triggers, lower threshold
Wrong skill triggersOverlapping patternsCheck confidence scores
API calls failWrong URL, expired keyCheck logs, test API manually
Timeout errorsSlow API, no timeout handlingAdd timeout, check latency
Formatting brokenEdge cases in dataAdd fixture tests for edge cases

Test Coverage


# Generate coverage report
openclaw skills test my-skill --unit --coverage

# Output:
# File              Stmts   Branch  Funcs   Lines
# index.ts          92%     85%     100%    92%
# lib/api.ts        88%     75%     100%    88%
# lib/formatter.ts  100%    100%    100%    100%
# lib/validator.ts  100%    95%     100%    100%
# Total             94%     87%     100%    94%
Works in test, fails liveConfig differencesCompare test vs. live config

Aim for at least 80% coverage on critical paths (API calls, error handling, input validation).

Continuous Testing

Automate testing on every change:


# Watch mode — re-run tests on file changes
openclaw skills test my-skill --unit --watch

# Pre-commit hook
# .git/hooks/pre-commit
#!/bin/sh
openclaw skills test my-skill --unit || exit 1

Best Practices

Next Steps

Related Articles