Appearance
Voice
With voice on, the chat shows a microphone button. Visitors talk; the face speaks the answers, with its lips moving to the audio. Voice and text are one conversation: visitors can switch at any time, everything said is written into the transcript, and a message typed during a voice conversation is answered aloud.
Turning it on
In the bot's Voice tab: turn voice on, pick an engine and its connection, model and voice (the lists come from your connected keys; voices have a preview button), then publish. The microphone button appears once the engine is fully set up; until then the admin panel lists what's missing.
Engines
| Engine | Setting | Providers | Notes |
|---|---|---|---|
| OpenAI Realtime | voice.mode: realtime with an OpenAI connection | OpenAI | Speech to speech, fast. Turn detection is semantic_vad (waits for the visitor to finish their thought) or server_vad (answers after a pause). Sees camera pictures. |
| Gemini Live | voice.mode: realtime with a Gemini connection | Google Gemini | Speech to speech. Takes a live camera view (continuous camera mode). |
| Cascade | voice.mode: cascade | Speech to text: OpenAI or ElevenLabs Scribe. Brain: the bot's text model (Anthropic, OpenAI or Gemini). Voice: ElevenLabs, OpenAI or Gemini | The bot's own brain answers, with all its tools and knowledge, then each sentence is spoken as soon as it's written. Its brain sees each turn's camera frame. A little slower to answer; any brain (Claude included) with any voice. |
| ElevenLabs Agents | voice.mode: elevenlabs_agent | ElevenLabs | Uses ElevenLabs' agent platform with an ElevenLabs voice. |
ElevenLabs Agents. The server creates an ElevenLabs agent for the bot (named "Wireface:" and the agent's name) the first time voice starts, and updates it when the voice settings change; you don't set anything up in ElevenLabs. The bot's instructions are sent with each conversation, and its tools (built-in, HTTP, MCP and page tools) run as client tools on your Wireface server, so it behaves like the other engines. It can't take pictures: camera frames are described in words by a vision model (vision.proxy, or the bot's brain when that is empty).
Whichever engine you pick, the agent's instructions tell it to speak in short, natural sentences, and (unless tools.builtins.express is off) it changes the face's expression through a silent express tool.
Tools in voice. Every engine gets the bot's tools: built-in, HTTP, MCP and the page tools your page registered. Tools that ask the visitor first show their question in the chat while the voice conversation waits.
Talking over the agent
With voice.bargeIn on (the default), the visitor can interrupt: the agent stops mid-sentence, and from then on it remembers only the part of its reply the visitor actually heard.
Push-to-talk
In noisy places, push-to-talk is more reliable: the visitor holds a button (or, with the button focused, the space bar) while speaking. voice.pushToTalk:
allowed(default): voice is hands-free; your page can start push-to-talk withWirefaceChat.startVoice({ mode: 'ptt' });only: the microphone button starts push-to-talk;off: hands-free only.
Pressing the button also stops the agent if it is talking.
Starting voice from your page
js
document.querySelector('#talk').addEventListener('click', () => WirefaceChat.startVoice());startVoice() opens the chat and starts voice. Call it from a click (the browser asks for the microphone). If the chat window is still loading, the request waits until it's ready. startVoice({ camera: true }) also offers the camera: the visitor sees the consent prompt first. stopVoice() ends voice.
Browser rules
- HTTPS. Browsers only give the microphone to secure pages: your site and the chat server both need HTTPS (
localhostcounts as secure for testing). - Permission. The browser asks the visitor the first time. If they block it, the chat says so and emits
mic_denied; they can allow it again from the address bar. APermissions-Policyheader on your site can also block it (see CSP). - Sound needs a click. Browsers only play sound after the visitor interacts. The chat's microphone button always counts. The chat window is allowed to play sound (its iframe has
allow="autoplay"), so a click on your own button usually counts too; if voice doesn't start that way in a browser you support, point visitors to the chat's microphone button. - Echo. The chat asks the browser for echo cancellation and noise suppression. Headphones still work best.
Limits
| Setting | Default | What happens |
|---|---|---|
voice.maxSessionMinutes | 15 | Voice ends after this long; the chat carries on in text. |
voice.idleSeconds | 120 | Voice ends after this long with nobody speaking. |
security.caps.concurrentVoiceSessions | 5 | More voice conversations at once than this get voice_capacity. |
security.caps.dailyVoiceMinutes | none | Minutes of audio (both ways) per UTC day; voice stops at the cap (spend_cap_reached). |
security.caps.dailyUsd | none | The bot's estimated spend per UTC day; it stops voice and replies alike. |
When someone from your team takes over the conversation, voice ends and the chat continues in text.
Opening hours with behavior.hours.scope: bot stop the agent answering typed messages outside hours, but they don't hold back a voice conversation in this version.
Greeting
With welcome.speakGreeting on, the agent says hello first when a conversation starts with voice.
Events
The voice event reports starting, listening, thinking, speaking, stopped and error. The underlying messages (voice.ready, voice.state, transcripts and audio frames) are in the protocol.
Next: Vision and webcam, or the details of each provider.