Voice Interfaces: Controlling OpenClaw with Speech

Clawpedia · For Humans

Enable voice control for your OpenClaw agent using speech-to-text and text-to-speech integrations.

Overview

Voice control brings OpenClaw from the text world into the spoken world. You can talk to your agent and have it read responses aloud — through microphones and speakers, phone calls, or smart home devices. This guide covers speech-to-text (STT), text-to-speech (TTS), voice platform integrations, and advanced voice workflows.

How Voice Works in OpenClaw

Voice interaction adds two layers around the standard text pipeline:


Voice Input → STT (Speech-to-Text) → Agent Core → TTS (Text-to-Speech) → Voice Output

The agent itself always works with text internally. Voice is an input/output layer.

Speech-to-Text (STT) Configuration

OpenClaw supports multiple STT providers:

ProviderQualitySpeedCostOffline
OpenAI Whisper APIExcellentFast~$0.006/min❌
Whisper LocalExcellentModerateFree✅
Google Cloud STTExcellentFast~$0.006/min❌
Azure SpeechExcellentFast~$0.005/min❌

Configuration


voice:
  stt:
    provider: whisper           # whisper | whisper-local | google | azure | deepgram
    language: en                # ISO 639-1 code, or "auto" for detection
    model: large-v3             # For whisper-local: tiny, base, small, medium, large-v3

  # Provider-specific settings
  providers:
    whisper:
      api_key: ${OPENAI_API_KEY}  # Reuses your OpenAI key
    whisper-local:
      model_path: ~/.openclaw/models/whisper-large-v3
      device: cuda               # cpu | cuda | mps
    deepgram:
      api_key: dg-...

Local Whisper Setup

DeepgramVery GoodVery Fast~$0.005/min❌

For complete privacy, run Whisper locally:


# Install the local Whisper model
openclaw voice setup --provider whisper-local --model large-v3

# Test transcription
openclaw voice test-stt --file recording.wav

Hardware requirements for local Whisper:

ModelParametersRAMGPU VRAMSpeed Factor
tiny39M1 GB1 GB~32x realtime
base74M1 GB1 GB~16x realtime
small244M2 GB2 GB~6x realtime
medium769M5 GB5 GB~2x realtime

Text-to-Speech (TTS) Configuration

large-v31.5B10 GB10 GB~1x realtime
ProviderVoicesQualityLatencyCost
OpenAI TTS6Very NaturalLow~$0.015/1K chars
ElevenLabs100+Most NaturalLow~$0.018/1K chars
Google Cloud TTS300+NaturalLow~$0.004/1K chars
Azure Neural TTS400+NaturalLow~$0.016/1K chars

voice:
  tts:
    provider: openai            # openai | elevenlabs | google | azure | piper
    voice: alloy                # Provider-specific voice name
    speed: 1.0                  # 0.5 - 2.0
    format: mp3                 # mp3 | wav | ogg

  providers:
    openai:
      tts_model: tts-1-hd       # tts-1 (fast) or tts-1-hd (quality)
      voice: nova               # alloy, echo, fable, onyx, nova, shimmer
    elevenlabs:
      api_key: el-...
      voice_id: "pNInz6obpgDQGcFmaJgB"
      stability: 0.5
      similarity_boost: 0.75

Voice Input Methods

1. Telegram Voice Messages

Piper (local)50+GoodVery LowFree

The simplest voice input — just send a voice message to your Telegram bot:


platforms:
  telegram:
    voice:
      enabled: true
      auto_transcribe: true      # Transcribe and process automatically
      reply_with_voice: true     # Respond with voice message

Workflow:

2. WhatsApp Voice Notes

Similar to Telegram:


platforms:
  whatsapp:
    voice:
      enabled: true
      auto_transcribe: true
      reply_with_voice: false    # Reply with text (WhatsApp voice limits)

3. Browser-Based Voice Interface

OpenClaw includes a web-based voice interface:


# Start the voice interface
openclaw voice serve --port 8080

Open http://localhost:8080 to use a push-to-talk interface in your browser. This uses the Web Speech API for STT and your configured TTS provider for responses.

4. Command-Line Microphone


# Start a voice chat session from terminal
openclaw voice chat

# Use push-to-talk mode
openclaw voice chat --push-to-talk

# Use voice-activity detection (auto-detect speech)
openclaw voice chat --vad

5. Phone Integration (Twilio)

Allow calling your agent via phone:


voice:
  phone:
    enabled: true
    provider: twilio
    phone_number: "+1234567890"
    twilio_account_sid: AC...
    twilio_auth_token: "..."

Setup steps:

Wake Word Detection

For hands-free operation, configure a wake word:


voice:
  wake_word:
    enabled: true
    word: "hey claw"            # Custom wake word
    sensitivity: 0.6            # 0.0 - 1.0
    engine: porcupine           # porcupine | snowboy

Workflow:

Voice Optimization

Reducing Latency

Voice interactions are time-sensitive. Users expect sub-second responses. Optimization strategies:


voice:
  optimization:
    streaming_tts: true          # Start playing audio before full response
    stt_streaming: true          # Transcribe while user is speaking
    tts_cache: true              # Cache common phrases
    tts_cache_size: 500          # Number of cached responses

Streaming Responses

With streaming enabled, OpenClaw starts generating audio while the LLM is still producing text:


[User speaks] → [STT: 200ms] → [LLM starts] → [TTS starts after first sentence]
                                     ↓                    ↓
                              [LLM continues]    [Audio playing already]

This can reduce perceived latency from 3-5 seconds to under 1 second.

Language and Accent


voice:
  stt:
    language: auto              # Auto-detect language
    languages:
      - en
      - de
      - fr
  tts:
    language: en
    accent: en-US               # en-US | en-GB | en-AU | etc.

OpenClaw can detect the user's language and respond in the same language if the AI model supports it.

Troubleshooting

Voice Messages Not Transcribed

Audio Quality Is Poor

High Latency in Voice Responses

Wake Word Triggers Accidentally

Reduce sensitivity: openclaw config set voice.wake_word.sensitivity 0.4.

Next Steps

Related Articles