Set up and optimize local GPU inference for running open-source models with OpenClaw.
Running OpenClaw with Local GPU-Powered Models
Running LLMs locally gives you complete privacy, zero API costs, and offline capability. This guide covers setting up GPU-accelerated models for OpenClaw using Ollama, llama.cpp, and other local inference engines.
Why Run Locally?
Benefit
Details
Privacy
Data never leaves your machine
Cost
No per-token charges after hardware investment
Latency
No network roundtrip, potentially faster
Offline
Works without internet connection
Control
Fine-tune and customize models freely
Hardware Requirements
Model Size
VRAM Needed
Recommended GPU
Performance
7-8B params
6-8 GB
RTX 3060/4060
30-50 tokens/s
13B params
10-12 GB
RTX 3080/4070
20-35 tokens/s
30-34B params
20-24 GB
RTX 3090/4090
10-20 tokens/s
70B params
40-48 GB
2x RTX 3090 or A100
5-15 tokens/s
70B (quantized Q4)
24-32 GB
RTX 4090
8-12 tokens/s
Setup with Ollama
Ollama is the easiest way to run local models with OpenClaw.
Installation
# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh
# Verify
ollama --version
# Pull a model
ollama pull llama3:8b # 4.7 GB, fast
ollama pull llama3:70b # 40 GB, powerful
ollama pull mistral:7b # 4.1 GB, efficient
ollama pull codellama:13b # 7.4 GB, code-focused
Configure OpenClaw
# config.yaml
agent:
model:
provider: "ollama"
name: "llama3:8b"
url: "http://localhost:11434"
options:
temperature: 0.7
num_ctx: 8192 # Context window
num_gpu: 99 # Offload all layers to GPU
num_thread: 8 # CPU threads for non-GPU work
Setup with llama.cpp
For maximum performance and control:
# Build llama.cpp with CUDA support
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release
# Download a GGUF model
# From HuggingFace: Meta-Llama-3-8B-Instruct-Q5_K_M.gguf
# Start the server
./build/bin/llama-server
--model ./models/llama-3-8b-instruct.Q5_K_M.gguf
--port 8080
--n-gpu-layers 35
--ctx-size 8192
--threads 8
# Process multiple requests in batches
local_model:
batching:
enabled: true
max_batch_size: 4
max_wait_ms: 100
KV Cache Optimization
# Reuse context for conversations
local_model:
kv_cache:
enabled: true
max_entries: 10
shared_prefix: true # Share system prompt cache across users
Model Preloading
# Keep model loaded in memory (Ollama)
curl http://localhost:11434/api/generate -d '{"model": "llama3:8b", "keep_alive": -1}'
Hybrid Setup: Local + Cloud
# Use local for simple tasks, cloud for complex ones
models:
default:
provider: "ollama"
name: "llama3:8b"
complex:
provider: "openai"
name: "gpt-4-turbo"
routing:
rules:
- condition: "message.length < 100 AND no_tool_calls"
model: "default"
- condition: "requires_code_generation"
model: "complex"
- condition: "cloud_api_unavailable"
model: "default" # Fallback to local
Apple Silicon (M1/M2/M3)
Apple Silicon Macs are excellent for local LLMs:
# Ollama automatically uses Metal (GPU) on macOS
ollama pull llama3:8b
# Performance on Apple Silicon:
# M1 Pro (16GB): ~25 tokens/s with Llama 3 8B
# M2 Max (32GB): ~35 tokens/s with Llama 3 8B
# M3 Max (48GB): ~20 tokens/s with Llama 3 70B Q4
Troubleshooting
Issue
Solution
"Out of memory"
Use smaller quantization (Q4 instead of Q8)
Slow generation
Increase GPU layers, check VRAM usage
Poor quality
Try larger model or higher quantization
Model not loading
Check available VRAM, close other GPU apps
Connection refused
Verify Ollama/llama.cpp server is running
Cost Comparison
Cloud API (GPT-4, moderate usage):
- 10,000 messages/month
- Average 1,000 tokens/message
- Cost: ~$300/month
Local (RTX 4090 setup):
- Hardware: $1,600 (one-time)
- Electricity: ~$20/month
- Break-even: ~5 months
- After break-even: essentially free
Best Practices
Start with Ollama — easiest setup, good performance
Use Q5_K_M quantization — best quality/size balance
Monitor GPU temperature — sustained load can cause throttling
Keep models updated — new quantization methods improve regularly
Test quality — compare local vs cloud on your specific tasks
Have a cloud fallback — for when local hardware is insufficient
Consider fine-tuning — local models can be customized for your domain
What AI models can I use with OpenClaw? — Complete list of supported AI models and providers compatible with OpenClaw, from GPT to open-source alternatives.