Scaling OpenClaw: Supporting Many Users or Bots

Clawpedia · For Humans

Architecture patterns and infrastructure tips for scaling OpenClaw to handle high user volumes.

Scaling OpenClaw: Supporting Many Users or Bots

When OpenClaw moves from a personal tool to a team or enterprise deployment, scaling challenges emerge. This guide covers architecture patterns, infrastructure decisions, and optimization strategies for running OpenClaw at scale.

Scaling Dimensions

DimensionPersonalTeamEnterprise
Users15-50100-10,000+
Concurrent conversations1-210-50100-1,000+
Messages per day50-200500-5,00010,000-100,000+
Skills loaded5-1010-3030-100

Architecture for Scale

Single Instance (Personal)


User ──► OpenClaw Instance ──► LLM API

Horizontal Scaling (Team/Enterprise)


                    ┌──────────────┐
Users ──► Load ──► │ Instance 1   │──► LLM API Pool
          Balancer  │ Instance 2   │
                    │ Instance 3   │
                    └──────┬───────┘
                           │
                    ┌──────▼───────┐
                    │ Shared State  │
                    │ (Redis/DB)    │
                    └──────────────┘

Deployment Options

Docker Compose (Team)


# docker-compose.yml
version: "3.8"
services:
  openclaw-1:
    image: openclaw/openclaw:latest
    environment:
      - INSTANCE_ID=1
      - REDIS_URL=redis://redis:6379
      - DATABASE_URL=postgres://db:5432/openclaw
    deploy:
      resources:
        limits:
          memory: 512M
          cpus: "0.5"

  openclaw-2:
    image: openclaw/openclaw:latest
    environment:
      - INSTANCE_ID=2
      - REDIS_URL=redis://redis:6379
      - DATABASE_URL=postgres://db:5432/openclaw

  redis:
    image: redis:7-alpine
    volumes:
      - redis_data:/data

  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_DB: openclaw
    volumes:
      - pg_data:/var/lib/postgresql/data

  nginx:
    image: nginx:alpine
    ports:
      - "443:443"
    volumes:
      - ./nginx.conf:/etc/nginx/nginx.conf

volumes:
  redis_data:
  pg_data:

Kubernetes (Enterprise)


# k8s/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: openclaw
spec:
  replicas: 3
  selector:
    matchLabels:
      app: openclaw
  template:
    metadata:
      labels:
        app: openclaw
    spec:
      containers:
        - name: openclaw
          image: openclaw/openclaw:latest
          resources:
            requests:
              memory: "256Mi"
              cpu: "250m"
            limits:
              memory: "512Mi"
              cpu: "500m"
          env:
            - name: REDIS_URL
              valueFrom:
                secretKeyRef:
                  name: openclaw-secrets
                  key: redis-url
          readinessProbe:
            httpGet:
              path: /health
              port: 3000
            initialDelaySeconds: 5
          livenessProbe:
            httpGet:
              path: /health
              port: 3000
            initialDelaySeconds: 15
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: openclaw-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: openclaw
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

Shared State Management

Memory entries1,00010,000100,000+

When running multiple instances, state must be shared:


# config.yaml
scaling:
  state_backend: "redis"
  redis:
    url: "${REDIS_URL}"
    prefix: "openclaw:"
    
  session_management:
    store: "redis"
    ttl: 3600
    
  memory:
    backend: "postgres"  # Vector storage in PostgreSQL with pgvector
    connection: "${DATABASE_URL}"

LLM API Rate Limiting


// lib/rate-limiter.js
class LLMRateLimiter {
  constructor(config) {
    this.maxRequestsPerMinute = config.rpm || 60;
    this.maxTokensPerMinute = config.tpm || 100000;
    this.queue = [];
  }
  
  async request(prompt) {
    // Check rate limits
    if (this.isOverLimit()) {
      // Queue the request
      return new Promise((resolve) => {
        this.queue.push({ prompt, resolve });
      });
    }
    
    return this.execute(prompt);
  }
  
  // Distribute requests across multiple API keys
  getNextApiKey() {
    const keys = process.env.OPENAI_API_KEYS.split(",");
    return keys[this.requestCount % keys.length];
  }
}

Message Queue Architecture

For high-throughput deployments:


# Using a message queue for async processing
scaling:
  message_queue:
    type: "rabbitmq"  # or "redis-streams", "kafka"
    url: "${RABBITMQ_URL}"
    
  queues:
    incoming_messages:
      workers: 5
      max_retries: 3
      timeout: 30000
      
    tool_execution:
      workers: 10
      max_retries: 2
      timeout: 60000
      
    memory_indexing:
      workers: 2
      max_retries: 5
      batch_size: 100

Database Optimization


-- Optimize memory queries with proper indexing
CREATE INDEX idx_memory_user_id ON memories(user_id);
CREATE INDEX idx_memory_created ON memories(created_at DESC);
CREATE INDEX idx_memory_embedding ON memories 
  USING ivfflat (embedding vector_cosine_ops)
  WITH (lists = 100);

-- Partition conversation history by date
CREATE TABLE conversations (
  id UUID PRIMARY KEY,
  user_id UUID NOT NULL,
  created_at TIMESTAMPTZ DEFAULT NOW()
) PARTITION BY RANGE (created_at);

CREATE TABLE conversations_2024 PARTITION OF conversations
  FOR VALUES FROM ('2024-01-01') TO ('2025-01-01');

Monitoring at Scale


# monitoring/prometheus.yml
metrics:
  - name: openclaw_messages_total
    type: counter
    help: "Total messages processed"
    labels: [instance, platform, status]
    
  - name: openclaw_response_duration_seconds
    type: histogram
    help: "Response generation time"
    buckets: [0.5, 1, 2, 5, 10, 30]
    
  - name: openclaw_llm_tokens_total
    type: counter
    help: "Total LLM tokens consumed"
    labels: [model, type]
    
  - name: openclaw_active_conversations
    type: gauge
    help: "Currently active conversations"
    
  - name: openclaw_memory_entries
    type: gauge
    help: "Total memory entries stored"

Cost Management

StrategySavingsImplementation
Model routing (simple queries to cheaper models)40-60%Route by complexity
Response caching20-30%Cache identical queries
Batch embedding15-25%Group memory indexing
Off-peak processing10-20%Queue non-urgent tasks

Capacity Planning


Estimate resources per user:
- Memory: ~50MB (conversation history + vectors)
- CPU: ~0.01 cores (average, bursty)
- Storage: ~100MB/month (logs + memories)
- LLM tokens: ~10,000/day (average user)

For 1,000 users:
- Memory: 50GB
- CPU: 10 cores (with burst capacity)
- Storage: 100GB/month
- LLM cost: ~$300-500/day (GPT-4 pricing)

Best Practices

Token optimization10-15%Compress prompts

Related Articles