Performance Tuning: Speeding Up OpenClaw

Clawpedia · For Humans

Optimize response times and resource usage for a faster, more efficient OpenClaw experience.

Performance Tuning: Speeding Up OpenClaw

A fast agent is a useful agent. This guide covers practical optimizations to reduce latency, lower costs, and improve the overall responsiveness of your OpenClaw deployment.

Where Time Goes

A typical OpenClaw request breaks down as follows:

PhaseTimePercentage
Message parsing5ms1%
Memory retrieval50-200ms10%
Prompt assembly10ms1%
LLM API call500-3000ms70-80%
Tool execution100-2000ms10-20%
Response delivery20ms2%

The LLM API call dominates. Optimization strategy should focus there first.

Level 1: LLM Optimization

Choose the Right Model

ModelLatencyQualityCost
GPT-4 Turbo2-5sExcellentHigh
GPT-3.5 Turbo0.5-1.5sGoodLow
Claude 3 Haiku0.3-1sGoodLow
Llama 3 (local)0.5-2sGoodFree

Model Routing

Mistral 7B (local)0.3-1sModerateFree

Route simple queries to fast models, complex ones to powerful models:


// lib/model-router.js
function selectModel(message, context) {
  const complexity = estimateComplexity(message);
  
  if (complexity === "simple") {
    // Quick factual questions, greetings, status checks
    return "gpt-3.5-turbo";
  } else if (complexity === "moderate") {
    // Summaries, basic analysis, formatting
    return "gpt-4-turbo";
  } else {
    // Complex reasoning, code generation, multi-step planning
    return "gpt-4";
  }
}

function estimateComplexity(message) {
  const wordCount = message.split(" ").length;
  const hasCodeRequest = /code|function|implement|debug/i.test(message);
  const hasAnalysis = /analyze|compare|evaluate|design/i.test(message);
  
  if (wordCount < 10 && !hasCodeRequest && !hasAnalysis) return "simple";
  if (hasCodeRequest || hasAnalysis) return "complex";
  return "moderate";
}

Streaming Responses

Show output as it is generated:


// Enable streaming for faster perceived response time
const stream = await openai.chat.completions.create({
  model: "gpt-4-turbo",
  messages: messages,
  stream: true,
});

let fullResponse = "";
for await (const chunk of stream) {
  const content = chunk.choices[0]?.delta?.content || "";
  fullResponse += content;
  // Send partial response to user immediately
  await platform.sendTypingIndicator();
  if (fullResponse.length % 100 === 0) {
    await platform.updateMessage(fullResponse);
  }
}

Level 2: Prompt Optimization

Reduce Token Count


Before (142 tokens):
"You are a helpful, professional, knowledgeable assistant 
that always provides detailed, comprehensive, thorough 
answers to every question asked by the user in a friendly 
and approachable manner..."

After (48 tokens):
"You are a professional assistant. Be concise and helpful. 
Use Markdown formatting. Cite sources when relevant."

Rule of thumb: Every 1,000 tokens saved = 0.5-1s faster response.

Compress Memory Context


// Instead of injecting raw memories
// Before: 2000 tokens of memory context
const rawMemories = await memory.search(query, { limit: 10 });

// After: 500 tokens of compressed context
const memories = await memory.search(query, { limit: 5 });
const compressed = memories.map(m => 
  `[${m.date}] ${m.summary || m.content.substring(0, 100)}`
).join("
");

Level 3: Caching

Response Cache


// Cache identical or near-identical queries
const cache = new Map();

async function getCachedResponse(prompt, options = {}) {
  const key = hashPrompt(prompt);
  const cached = cache.get(key);
  
  if (cached && Date.now() - cached.timestamp < options.ttl) {
    return cached.response;
  }
  
  const response = await llm.complete(prompt);
  cache.set(key, { response, timestamp: Date.now() });
  return response;
}

Embedding Cache


// Cache vector embeddings to avoid recomputing
const embeddingCache = new LRUCache({ max: 10000 });

async function getEmbedding(text) {
  const key = hashText(text);
  if (embeddingCache.has(key)) return embeddingCache.get(key);
  
  const embedding = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: text,
  });
  
  embeddingCache.set(key, embedding.data[0].embedding);
  return embedding.data[0].embedding;
}

Level 4: Infrastructure

Connection Pooling


# config.yaml
performance:
  database:
    pool_size: 20
    idle_timeout: 30000
  redis:
    pool_size: 10
  http:
    keep_alive: true
    max_sockets: 50

Local Model Acceleration


# Use GPU-accelerated local models
local_model:
  backend: "ollama"
  model: "llama3:8b"
  gpu_layers: 35      # Offload to GPU
  context_size: 8192
  batch_size: 512      # Larger batches for throughput
  threads: 8           # CPU threads for non-GPU work

Level 5: Async Processing


// Process non-urgent tasks asynchronously
async function handleMessage(message) {
  // Immediate response for simple queries
  const quickResponse = await getQuickResponse(message);
  if (quickResponse) return quickResponse;
  
  // For complex tasks, acknowledge and process async
  await platform.send("Working on it...");
  
  // Process in background
  const result = await processComplex(message);
  await platform.send(result);
}

Benchmarking


# Built-in benchmark tool
openclaw benchmark --messages 100 --concurrent 10

# Output:
# Messages processed: 100
# Average latency: 1.2s
# P50: 0.9s
# P95: 2.8s
# P99: 4.1s
# Tokens/second: 45
# Cache hit rate: 32%
# Errors: 0

Optimization Checklist

OptimizationImpactEffortPriority
Model routingHighMedium1
Response streamingHighLow2
Reduce prompt tokensMediumLow3
Response cachingMediumMedium4
Embedding cacheLowLow5
Connection poolingMediumLow6
Local modelsHighHigh7

Monitoring Performance

Async processingMediumMedium8

Track these metrics to identify bottlenecks:


metrics:
  latency:
    - llm_first_token_ms     # Time to first token
    - llm_total_ms           # Total LLM time
    - memory_retrieval_ms    # Vector search time
    - tool_execution_ms      # External tool time
    - total_response_ms      # End-to-end time
    
  throughput:
    - messages_per_minute
    - tokens_per_second
    - cache_hit_rate
    
  cost:
    - tokens_per_message
    - cost_per_message
    - cost_per_user_per_day

Quick Wins

Related Articles