Appearance
OpenAI
One OpenAI key covers everything: text replies, realtime voice, and both halves of the cascade.
Key: create one at platform.openai.com/api-keys.
What it's used for
| Use | Supported |
|---|---|
| Text replies (the brain) | Yes: GPT models and reasoning models |
| Images and camera frames | Yes, except gpt-3.5, o1-mini and o3-mini (the admin panel warns) |
| Tools | Yes |
| Realtime voice | Yes: the Realtime API (voice.mode: realtime with an OpenAI connection) |
| Speech to text (cascade) | Yes: realtime transcription, gpt-4o-mini-transcribe by default |
| Text to speech (cascade) | Yes: gpt-4o-mini-tts by default |
Models
The server reads your key's model list and sorts it by use. New bots start with the newest gpt-N model for text and gpt-realtime for realtime voice (when your key has them).
- Reasoning models get
brain.effortas their reasoning effort, and keep their reasoning from one turn to the next.brain.temperatureonly applies to non-reasoning models.
Realtime voice
- Turn detection (
voice.realtime.turnDetection):semantic_vad(the default) waits until the visitor seems to have finished their thought, witheagernessfromlowtohigh(orauto);server_vadanswers after about half a second of silence. - Voices:
marinandcedar(recommended),alloy,ash,ballad,coral,echo,sage,shimmerandverse. Preview them from the admin panel. - Camera: sees frames with each turn, or when the agent uses
look_at_camera. It doesn't take a live video stream (use Gemini Live for that). - Interrupted replies are cut where the visitor stopped hearing them, so the model's memory matches what was said.
Text to speech
All the realtime voices plus fable, nova and onyx. voice.cascade.tts.speed sets the speed, and voice.cascade.tts.style is passed as instructions ("warm and unhurried").