Can OpenClaw use local language models (like LLaMA or Ollama)?

Clawpedia · For Humans

Run OpenClaw with locally hosted models using LLaMA, Ollama, or other self-hosted inference solutions.

Can OpenClaw Use Local Language Models (Like LLaMA or Ollama)?

Absolutely. One of OpenClaw's core strengths is its ability to run entirely offline using local language models. This means your conversations never leave your machine, giving you complete privacy and zero API costs.

This guide covers how to set up local models, which hardware you need, and how to get the best performance.

---

Why Use Local Models?

BenefitDetails
PrivacyNo data sent to external servers
CostZero ongoing API fees
SpeedNo network latency (if hardware is sufficient)
OfflineWorks without internet
CustomizationFine-tune models for your specific use case

Tradeoffs:

---

Supported Local Model Providers

ProviderDescriptionSupported Models
OllamaMost popular, easiest setupLLaMA 3, Mistral, Phi-3, Gemma, CodeLlama
llama.cppDirect C++ inferenceAny GGUF model
LM StudioGUI-based model managerThousands of models
vLLMHigh-performance servingMost HuggingFace models
LocalAIOpenAI-compatible APIVarious architectures

---

Setting Up with Ollama (Recommended)

Ollama is the easiest way to run local models with OpenClaw.

Step 1: Install Ollama


# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.ai/install.sh | sh

# Windows
# Download from https://ollama.ai/download

Step 2: Pull a Model


# General-purpose (recommended starter)
ollama pull llama3.1

# Smaller, faster model
ollama pull phi3

# Coding-focused
ollama pull codellama

# Large, most capable
ollama pull llama3.1:70b

Step 3: Configure OpenClaw


openclaw config set provider ollama
openclaw config set model llama3.1

Step 4: Test


openclaw chat "Tell me a joke"

---

Hardware Requirements

RAM Requirements by Model Size

Model SizeRAM NeededExample Models
1–3B4 GBPhi-3 Mini, TinyLlama
7–8B8 GBLLaMA 3 8B, Mistral 7B
13B16 GBLLaMA 2 13B, CodeLlama 13B
34B32 GBCodeLlama 34B
70B48–64 GBLLaMA 3 70B

Tip: On Apple Silicon Macs, unified memory makes even 13B models run smoothly on 16GB machines.

GPU Acceleration

PlatformGPU Support
NVIDIAFull CUDA support (best performance)
Apple SiliconMetal acceleration (excellent)
AMDROCm support (Linux only)
CPU onlyWorks but slower

---

Model Recommendations

For General Use


# Best quality-to-speed ratio
ollama pull llama3.1

# Fastest responses
ollama pull phi3

For Coding


# Dedicated coding model
ollama pull codellama:13b

# General model with strong coding
ollama pull llama3.1

For Limited Hardware


# Tiny but capable (2GB RAM)
ollama pull phi3:mini

# Good balance for 8GB machines
ollama pull mistral

---

Using LM Studio

LM Studio provides a graphical interface for managing models:


openclaw config set provider openai-compatible
openclaw config set provider.base_url http://localhost:1234/v1
openclaw config set model local-model

---

Using llama.cpp Directly

For maximum control:


# Clone and build
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make -j

# Download a GGUF model
wget https://huggingface.co/TheBloke/Llama-2-7B-GGUF/resolve/main/llama-2-7b.Q4_K_M.gguf

# Start the server
./server -m llama-2-7b.Q4_K_M.gguf -c 4096 --port 8080

# Configure OpenClaw
openclaw config set provider openai-compatible
openclaw config set provider.base_url http://localhost:8080/v1

---

Quantization Explained

Local models come in different quantization levels that trade quality for size:

QuantizationQualitySize (7B model)Speed
F16Best~14 GBSlowest
Q8_0Near-perfect~7 GBFast
Q5_K_MVery good~5 GBFaster
Q4_K_MGood~4 GBFast
Q3_K_MAcceptable~3 GBFastest
Q2_KDegraded~2.5 GBFastest

Recommendation: Use Q4_K_M or Q5_K_M for the best balance of quality and performance.

---

Hybrid Mode: Local + Cloud

You can configure OpenClaw to use local models by default and fall back to cloud models for complex tasks:


# ~/.openclaw/config.yaml
provider: ollama
model: llama3.1

fallback:
  provider: openai
  model: gpt-4o
  trigger: complexity  # Use cloud for complex queries

Or switch manually:


# Quick switch in chat
/model gpt-4o      # Switch to cloud
/model llama3.1    # Switch back to local

---

Performance Tuning


# Increase context window
openclaw config set context.max_tokens 8192

# Set number of CPU threads
ollama set OLLAMA_NUM_THREADS 8

# Enable GPU layers (NVIDIA)
ollama set OLLAMA_NUM_GPU 999

# Limit memory usage
ollama set OLLAMA_MAX_LOADED_MODELS 1

---

Troubleshooting

"Model not found"


ollama list          # Check installed models
ollama pull llama3.1  # Re-download

Slow Responses

Out of Memory

---

Summary

OpenClaw fully supports local language models through Ollama, llama.cpp, LM Studio, and other providers. Local models give you complete privacy at zero cost, with the tradeoff of requiring decent hardware. Start with Ollama and LLaMA 3 for the easiest setup, and use hybrid mode to get the best of both local and cloud models.

Related Articles