Appearance
Google Gemini
Gemini models for text, the Live API for voice (with a live camera view), and Gemini's text to speech.
Key: create one in Google AI Studio (aistudio.google.com/apikey).
What it's used for
| Use | Supported |
|---|---|
| Text replies (the brain) | Yes |
| Images and camera frames | Yes |
| Tools | Yes |
| Realtime voice | Yes: the Live API (voice.mode: realtime with a Gemini connection) |
| Live camera video | Yes: with vision.webcam.mode: continuous, Gemini Live gets a frame about once a second while the picture changes |
| Speech to text (cascade) | No: use OpenAI or ElevenLabs |
| Text to speech (cascade) | Yes: gemini-3.8-flash-lite-tts when no model is chosen |
Models
The server reads your key's model list. New bots start with the newest gemini-N-flash model for text, and a Live model for voice.
- On Gemini 3 and later,
brain.effortsets the thinking level:mediumandhighmean high,minimalandlowmean low. Older models don't get a thinking level. brain.temperatureis passed through when set.
Voices
Gemini's fixed set of named voices (Puck, Charon, Kore, Fenrir, Aoede, Leda, Orus, Zephyr and more), with a description of each in the admin panel's list. In the cascade, voice.cascade.tts.style is passed to the voice as a style prompt; speed isn't supported by Gemini.