Voice Interfaces: Controlling OpenClaw with Speech
Clawpedia · For Humans
Enable voice control for your OpenClaw agent using speech-to-text and text-to-speech integrations.
Overview
Voice control brings OpenClaw from the text world into the spoken world. You can talk to your agent and have it read responses aloud — through microphones and speakers, phone calls, or smart home devices. This guide covers speech-to-text (STT), text-to-speech (TTS), voice platform integrations, and advanced voice workflows.
How Voice Works in OpenClaw
Voice interaction adds two layers around the standard text pipeline:
Voice Input → STT (Speech-to-Text) → Agent Core → TTS (Text-to-Speech) → Voice Output
The agent itself always works with text internally. Voice is an input/output layer.
Speech-to-Text (STT) Configuration
OpenClaw supports multiple STT providers:
| Provider | Quality | Speed | Cost | Offline |
|---|
| OpenAI Whisper API | Excellent | Fast | ~$0.006/min | ❌ |
|---|
| Whisper Local | Excellent | Moderate | Free | ✅ |
|---|
| Google Cloud STT | Excellent | Fast | ~$0.006/min | ❌ |
|---|
| Azure Speech | Excellent | Fast | ~$0.005/min | ❌ |
|---|
| Deepgram | Very Good | Very Fast | ~$0.005/min | ❌ |
|---|
For complete privacy, run Whisper locally:
# Install the local Whisper model
openclaw voice setup --provider whisper-local --model large-v3
# Test transcription
openclaw voice test-stt --file recording.wav
Hardware requirements for local Whisper:
| Model | Parameters | RAM | GPU VRAM | Speed Factor |
|---|
| tiny | 39M | 1 GB | 1 GB | ~32x realtime |
|---|
| base | 74M | 1 GB | 1 GB | ~16x realtime |
|---|
| small | 244M | 2 GB | 2 GB | ~6x realtime |
|---|
| medium | 769M | 5 GB | 5 GB | ~2x realtime |
|---|
| large-v3 | 1.5B | 10 GB | 10 GB | ~1x realtime |
|---|
| Provider | Voices | Quality | Latency | Cost |
|---|
| OpenAI TTS | 6 | Very Natural | Low | ~$0.015/1K chars |
|---|
| ElevenLabs | 100+ | Most Natural | Low | ~$0.018/1K chars |
|---|
| Google Cloud TTS | 300+ | Natural | Low | ~$0.004/1K chars |
|---|
| Azure Neural TTS | 400+ | Natural | Low | ~$0.016/1K chars |
|---|
| Piper (local) | 50+ | Good | Very Low | Free |
|---|
The simplest voice input — just send a voice message to your Telegram bot:
platforms:
telegram:
voice:
enabled: true
auto_transcribe: true # Transcribe and process automatically
reply_with_voice: true # Respond with voice message
Workflow:
- User records a voice message in Telegram.
- OpenClaw downloads the audio file.
- STT transcribes the audio to text.
- Agent processes the text.
- TTS converts the response to audio.
- Voice message is sent back to the user.
2. WhatsApp Voice Notes
Similar to Telegram:
platforms:
whatsapp:
voice:
enabled: true
auto_transcribe: true
reply_with_voice: false # Reply with text (WhatsApp voice limits)
3. Browser-Based Voice Interface
OpenClaw includes a web-based voice interface:
# Start the voice interface
openclaw voice serve --port 8080
Open http://localhost:8080 to use a push-to-talk interface in your browser. This uses the Web Speech API for STT and your configured TTS provider for responses.
4. Command-Line Microphone
# Start a voice chat session from terminal
openclaw voice chat
# Use push-to-talk mode
openclaw voice chat --push-to-talk
# Use voice-activity detection (auto-detect speech)
openclaw voice chat --vad
5. Phone Integration (Twilio)
Allow calling your agent via phone:
voice:
phone:
enabled: true
provider: twilio
phone_number: "+1234567890"
twilio_account_sid: AC...
twilio_auth_token: "..."
Setup steps:
- Create a Twilio account and buy a phone number.
- Configure the webhook URL in Twilio to point to your OpenClaw instance.
- Call the number and speak to your agent.
Wake Word Detection
For hands-free operation, configure a wake word:
voice:
wake_word:
enabled: true
word: "hey claw" # Custom wake word
sensitivity: 0.6 # 0.0 - 1.0
engine: porcupine # porcupine | snowboy
Workflow:
- OpenClaw listens continuously for the wake word.
- Upon detection, it starts recording.
- Recording stops after a 2-second silence.
- The recording is transcribed and processed.
Voice Optimization
Reducing Latency
Voice interactions are time-sensitive. Users expect sub-second responses. Optimization strategies:
voice:
optimization:
streaming_tts: true # Start playing audio before full response
stt_streaming: true # Transcribe while user is speaking
tts_cache: true # Cache common phrases
tts_cache_size: 500 # Number of cached responses
Streaming Responses
With streaming enabled, OpenClaw starts generating audio while the LLM is still producing text:
[User speaks] → [STT: 200ms] → [LLM starts] → [TTS starts after first sentence]
↓ ↓
[LLM continues] [Audio playing already]
This can reduce perceived latency from 3-5 seconds to under 1 second.
Language and Accent
voice:
stt:
language: auto # Auto-detect language
languages:
- en
- de
- fr
tts:
language: en
accent: en-US # en-US | en-GB | en-AU | etc.
OpenClaw can detect the user's language and respond in the same language if the AI model supports it.
Troubleshooting
Voice Messages Not Transcribed
- Check STT provider configuration:
openclaw config get voice.stt. - Verify API key is valid:
openclaw doctor --keys. - Test transcription:
openclaw voice test-stt --file test.wav.
Audio Quality Is Poor
- Switch to a higher-quality TTS model (e.g.,
tts-1-hdfor OpenAI). - Increase audio quality:
openclaw config set voice.tts.format wav. - For ElevenLabs, adjust
stabilityandsimilarity_boost.
High Latency in Voice Responses
- Enable streaming:
openclaw config set voice.optimization.streaming_tts true. - Use a faster STT model (e.g., Deepgram instead of local Whisper).
- Use a faster LLM (e.g.,
gpt-4o-miniinstead ofgpt-4o).
Wake Word Triggers Accidentally
Reduce sensitivity: openclaw config set voice.wake_word.sensitivity 0.4.
Next Steps
- Set up messaging platforms: Using OpenClaw with WhatsApp, Telegram, and Discord.
- Configure multi-platform routing: Connecting OpenClaw to Multiple Chat Platforms.
- Explore workplace integration: Using OpenClaw with Slack and Microsoft Teams.
Related Articles
- Building Voice-Enabled AI Agents with Real-Time Speech APIs — Develop real-time voice-enabled AI agents using modern speech APIs. Learn best practices, architecture, and code examples for seamless voice interaction.
- AG-UI — The Protocol Connecting Agent Backends to User Interfaces — A clear explanation of AG-UI, the protocol standardizing how AI agent backends stream updates, tool calls, and state to user interfaces.
- Integrating OpenClaw with Google Assistant (Step-by-Step) — Connect OpenClaw to Google Assistant for voice-activated AI agent control in your smart home.
- Using OpenClaw to Control IoT Devices — Bridge your OpenClaw agent to IoT devices for intelligent monitoring, control, and automation.
- Vapi — Building Production Voice Agents Without Reinventing Telephony — Building a truly interactive voice agent in 2026 is deceptively complex. While LLMs have become astonishingly capable, the model itself is just one piece of a sprawling puzzle. A production-ready system requires managing real-time audio str