Appearance
Vision and webcam
The agent can see two ways: images visitors send, and, if they agree, their camera. Both need a brain that can see images (Claude, Gemini and most OpenAI models can; the admin panel warns when the chosen model can't, and the features stay off).
Image uploads
On by default (vision.uploads). Visitors attach images with the paperclip button, by pasting, or by dropping them on the chat:
- JPEG, PNG, GIF and WebP. The chat shrinks photos to at most 1568 pixels before sending;
- up to
maxPerMessageimages per message (4) andmaxMBeach (10 MB); - the server decodes and re-encodes every image, which strips EXIF data (including GPS position) and anything else hidden in the file;
- uploads count against
security.rateLimits.uploadsPerHour(30 per visitor).
The images become part of the conversation: the agent sees them with the message, and your team sees them in the transcript.
The camera
Off by default. Turn it on in the bot's Vision tab (vision.webcam.enabled). The chat then shows a camera button. Clicking it first shows your consent text (vision.webcam.consentText, with an optional privacy link); only if the visitor clicks Allow does the browser ask for the camera. A small self-view shows them what the agent sees, with a "Camera on" label.
The camera turns off when the visitor clicks the button again, when the chat closes, or when the page is hidden (frames pause while the tab isn't visible).
Modes
vision.webcam.mode decides when the agent sees the camera:
| Mode | The agent sees |
|---|---|
per_turn (default) | The latest frame (from the last 15 seconds) with each message the visitor sends. In voice, a frame every few seconds while the picture changes. |
continuous | A live view, for engines that take video (Gemini Live): a frame about once a second while the picture changes. Other engines get per_turn. |
on_demand | Nothing, until the agent decides to look with the look_at_camera tool. |
The chat sends frames only when the picture has changed (or every few seconds otherwise), at most vision.webcam.maxFps a second and maxWidth pixels wide. In a voice conversation on ElevenLabs Agents, which can't take pictures, the camera works on_demand. (The cascade's brain gets each turn's frame like text chat does.)
look_at_camera
With the camera enabled, the agent has a look_at_camera tool (turn it off under tools.builtins). It asks the chat for a fresh frame and looks at it, so the agent can say "let me look" when the visitor holds something up. If the camera is off, the agent is told so and can ask the visitor to turn it on.
For ElevenLabs Agents, which can't take images, a vision model describes the frame in words for it: vision.proxy, or the bot's brain when that is empty.
What the agent won't do
The agent's built-in rules tell it to describe and help with what is shown, never to identify people from their face, and never to guess sensitive traits (age, ethnicity, health, religion and the like) from how they look.
How long frames are kept
- Camera frames the agent sees are stored with the conversation while it is going on. Unless you turn on
vision.webcam.persistFrames, they are deleted after the conversation ends (within about two hours). - Your team can only see camera frames in transcripts when
vision.webcam.shareWithHumanAgentsis on. - Uploaded images are kept with the conversation, and deleted with it (see Privacy). Images a visitor attached but never sent are deleted after a day.
In the protocol
Frames travel as binary WebSocket frames (JPEG) after vision.start, paced by the server's vision.state. See Widget protocol.