Running OpenClaw with Local GPU-Powered Models

Clawpedia · For Humans

Set up and optimize local GPU inference for running open-source models with OpenClaw.

Running OpenClaw with Local GPU-Powered Models

Running LLMs locally gives you complete privacy, zero API costs, and offline capability. This guide covers setting up GPU-accelerated models for OpenClaw using Ollama, llama.cpp, and other local inference engines.

Why Run Locally?

BenefitDetails
PrivacyData never leaves your machine
CostNo per-token charges after hardware investment
LatencyNo network roundtrip, potentially faster
OfflineWorks without internet connection

Hardware Requirements

ControlFine-tune and customize models freely
Model SizeVRAM NeededRecommended GPUPerformance
7-8B params6-8 GBRTX 3060/406030-50 tokens/s
13B params10-12 GBRTX 3080/407020-35 tokens/s
30-34B params20-24 GBRTX 3090/409010-20 tokens/s
70B params40-48 GB2x RTX 3090 or A1005-15 tokens/s

Setup with Ollama

70B (quantized Q4)24-32 GBRTX 40908-12 tokens/s

Ollama is the easiest way to run local models with OpenClaw.

Installation


# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh

# Verify
ollama --version

# Pull a model
ollama pull llama3:8b        # 4.7 GB, fast
ollama pull llama3:70b       # 40 GB, powerful
ollama pull mistral:7b       # 4.1 GB, efficient
ollama pull codellama:13b    # 7.4 GB, code-focused

Configure OpenClaw


# config.yaml
agent:
  model:
    provider: "ollama"
    name: "llama3:8b"
    url: "http://localhost:11434"
    options:
      temperature: 0.7
      num_ctx: 8192        # Context window
      num_gpu: 99          # Offload all layers to GPU
      num_thread: 8        # CPU threads for non-GPU work

Setup with llama.cpp

For maximum performance and control:


# Build llama.cpp with CUDA support
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release

# Download a GGUF model
# From HuggingFace: Meta-Llama-3-8B-Instruct-Q5_K_M.gguf

# Start the server
./build/bin/llama-server 
  --model ./models/llama-3-8b-instruct.Q5_K_M.gguf 
  --port 8080 
  --n-gpu-layers 35 
  --ctx-size 8192 
  --threads 8

# OpenClaw config for llama.cpp
agent:
  model:
    provider: "openai-compatible"
    name: "llama3"
    url: "http://localhost:8080/v1"
    api_key: "not-needed"

Quantization Guide

Quantization reduces model size and memory usage at the cost of some quality:

QuantizationSize ReductionQuality LossRecommended For
Q8_0~50%MinimalBest quality, enough VRAM
Q6_K~57%Very smallGood balance
Q5_K_M~63%SmallMost users
Q4_K_M~70%ModerateLimited VRAM
Q3_K_M~75%NotableVery limited VRAM

# Download specific quantization
ollama pull llama3:8b-instruct-q5_K_M

Model Selection for Local Use

Q2_K~82%SignificantNot recommended
Use CaseRecommended ModelSizeWhy
General assistantLlama 3 8B4.7 GBBest quality/size ratio
Code assistanceCodeLlama 13B7.4 GBTrained on code
Creative writingMistral 7B4.1 GBGood creative output
MultilingualQwen 2 7B4.5 GBStrong multilingual
AnalysisLlama 3 70B (Q4)24 GBDeep reasoning

GPU Monitoring


# NVIDIA GPU monitoring
nvidia-smi --query-gpu=utilization.gpu,memory.used,memory.total,temperature.gpu --format=csv -l 1

# Watch during inference
watch -n 1 nvidia-smi

# In OpenClaw logs
openclaw debug --show-gpu-stats

Performance Optimization

Batch Processing


# Process multiple requests in batches
local_model:
  batching:
    enabled: true
    max_batch_size: 4
    max_wait_ms: 100

KV Cache Optimization


# Reuse context for conversations
local_model:
  kv_cache:
    enabled: true
    max_entries: 10
    shared_prefix: true  # Share system prompt cache across users

Model Preloading


# Keep model loaded in memory (Ollama)
curl http://localhost:11434/api/generate -d '{"model": "llama3:8b", "keep_alive": -1}'

Hybrid Setup: Local + Cloud


# Use local for simple tasks, cloud for complex ones
models:
  default:
    provider: "ollama"
    name: "llama3:8b"
    
  complex:
    provider: "openai"
    name: "gpt-4-turbo"
    
routing:
  rules:
    - condition: "message.length < 100 AND no_tool_calls"
      model: "default"
    - condition: "requires_code_generation"
      model: "complex"
    - condition: "cloud_api_unavailable"
      model: "default"  # Fallback to local

Apple Silicon (M1/M2/M3)

Privacy-criticalPhi-3 Mini2.3 GBRuns on CPU

Apple Silicon Macs are excellent for local LLMs:


# Ollama automatically uses Metal (GPU) on macOS
ollama pull llama3:8b

# Performance on Apple Silicon:
# M1 Pro (16GB): ~25 tokens/s with Llama 3 8B
# M2 Max (32GB): ~35 tokens/s with Llama 3 8B
# M3 Max (48GB): ~20 tokens/s with Llama 3 70B Q4

Troubleshooting

IssueSolution
"Out of memory"Use smaller quantization (Q4 instead of Q8)
Slow generationIncrease GPU layers, check VRAM usage
Poor qualityTry larger model or higher quantization
Model not loadingCheck available VRAM, close other GPU apps

Cost Comparison


Cloud API (GPT-4, moderate usage):
- 10,000 messages/month
- Average 1,000 tokens/message
- Cost: ~$300/month

Local (RTX 4090 setup):
- Hardware: $1,600 (one-time)
- Electricity: ~$20/month
- Break-even: ~5 months
- After break-even: essentially free

Best Practices

Connection refusedVerify Ollama/llama.cpp server is running

Related Articles