Appearance
ElevenLabs
ElevenLabs gives the agent a voice: any voice in your ElevenLabs account, including ones you cloned or designed. Use it in the cascade (with any brain, Claude included), or run the whole voice conversation on ElevenLabs Agents.
Key: create one in your ElevenLabs settings (elevenlabs.io/app/settings/api-keys). The key needs to read your voices. Without permission to list models, the server offers the well-known ones (eleven_flash_v2_5, eleven_turbo_v2_5, eleven_multilingual_v2).
What it's used for
| Use | Supported |
|---|---|
| Text to speech (cascade) | Yes: eleven_flash_v2_5 when no model is chosen |
| Speech to text (cascade) | Yes: Scribe v2 Realtime (scribe_v2_realtime) |
| ElevenLabs Agents | Yes (voice.mode: elevenlabs_agent) |
| Text replies (the brain) | No: connect Anthropic, OpenAI or Gemini too |
Voices
The voice list is your account's voices: your own (cloned, designed, professional) first, then the rest, with ElevenLabs' default voices last. Each shows its gender, accent, language and a preview. In the cascade, voice.cascade.tts.speed sets the speed; style isn't used by ElevenLabs.
ElevenLabs Agents
With voice.mode: elevenlabs_agent, voice conversations run on ElevenLabs' agent platform:
- Set up for you. The first time someone starts voice, the server creates an ElevenLabs agent for the bot (named "Wireface:" and the agent's name), and updates it when the voice, TTS model, LLM or language changes. If you delete it in the ElevenLabs dashboard, a new one is made.
- Same instructions. Each conversation sends the bot's full instructions, so the agent behaves as it does in text.
- Same tools. The bot's tools (built-in, HTTP, MCP and page tools) are registered as client tools that run on your Wireface server: ElevenLabs asks, your server runs them and answers. Tools that ask the visitor first still do.
- Choose its LLM with
voice.elevenlabsAgent.llm(an ElevenLabs model id; empty for ElevenLabs' default), and its voice and TTS model withvoiceandttsModel. - No pictures. ElevenLabs Agents can't take images. When the agent looks through the camera, a vision model describes the frame in words for it:
vision.proxy, or the bot's brain when that is empty. - The agent's conversation length limit is set from
voice.maxSessionMinutes.