Appearance
Widget protocol
The chat window talks to the server over one WebSocket, wss://<host>/v1/widget/ws?bot=<public id>. This page describes protocol version 1. You don't need it to embed the chat; it is here for debugging (the browser's network tab shows every message) and for anyone building on the server.
Only pages on the chat server's own origin can open the socket: the server checks the Origin header (unless NODE_ENV is production, it also accepts localhost and 127.0.0.1 on any port). Your website never talks to it directly; widget.js puts the chat window, which does, in an iframe.
- Text frames are JSON objects with a
tfield (the message type). Up to 256 KB each. - Binary frames carry audio and camera pictures, with an 8-byte header (see Binary frames).
Connecting
- Open the socket. Within 5 seconds, send
hellowithprotocol: 1. Any other first message, or none, closes the socket with 4401. - The server checks the bot is live, that the page in
hello.page.urlis in the bot's allowed origins, and (when the bot requires it) the visitor's identity. It answerssession.ready, or a fatalerrorand a close code (see close codes). - Keep
session.ready.visitorToken: it is who the visitor is for this bot. Send it in the nexthello.
The conversation starts with the visitor's first chat.send (or voice, a lead form or a hand-off request): the server sends conversation.started with its id and a resumeToken.
Resuming
Stored messages carry seq, increasing by one within a conversation. To pick up after a reload or a dropped connection, send hello with visitorToken, conversationId, resumeToken and lastSeq (the highest seq you have). The server answers session.ready with resumed: true, then history.sync with the messages after lastSeq. A conversation resumes while it is open and (for one the AI is handling) its last message is within the bot's behavior.persistence.resumeWindowHours; otherwise the visitor starts a new one.
The server pings every 20 seconds at the WebSocket level and drops connections that don't answer. The ping message is for measuring round trips.
Client messages
hello
The first message, within 5 seconds of connecting. With conversationId (and its resumeToken) it resumes that conversation; without one the server picks up the visitor's latest open conversation within behavior.persistence.resumeWindowHours. lastSeq asks only for messages after that sequence number; without it, a bot with behavior.persistence.showHistory: false sends no history (the conversation still carries on). Facts in context reach the agent as a note when they differ from the last ones it got. This server doesn't use tz or caps yet.
| Field | Type | Limits and notes |
|---|---|---|
protocol | number | |
visitorToken? | string | up to 200 chars |
conversationId? | string | 1 to 100 chars |
resumeToken? | string | up to 200 chars |
lastSeq? | number | at least 0 |
identity? | Identity | |
page? | PageInfo | |
locale? | string | up to 35 chars |
tz? | string | up to 64 chars |
caps? | { audioIn: boolean; audioOut: boolean; webcam: boolean } | |
clientTools? | ClientToolDecl[] | up to 64 items |
context? | Record<string, Json> | |
preview? | string | up to 2000 chars. A server-signed admin preview token: the session uses the bot's draft config. |
conversation.new
Leave the current conversation and start fresh. The old one is closed, unless it held no more than the welcome and one message. The server answers with a new session.ready.
No fields.
chat.send
A visitor message. clientMsgId (your own id) comes back on the stored message so you can match it. attachments are ids from the upload endpoint; the bot's vision.uploads.maxPerMessage applies. Text longer than security.maxMessageChars is cut.
| Field | Type | Limits and notes |
|---|---|---|
clientMsgId | string | 1 to 64 chars |
text | string | up to 20000 chars |
attachments? | string[] | up to 10 items, each up to 100 chars |
chat.typing
The visitor is typing (shown to your team in the admin panel).
| Field | Type | Limits and notes |
|---|---|---|
active | boolean |
chat.cancel
Stop the reply being written now. responseId is ignored: the reply in progress stops.
| Field | Type | Limits and notes |
|---|---|---|
responseId? | string | 1 to 100 chars |
context.update
The page changed (page), new facts for the agent (vars, merged into earlier ones), or the visitor's language changed (locale: the server remembers it and, with identity.languagePolicy: match_visitor, tells the agent to answer in it). Each reaches the agent as a note.
| Field | Type | Limits and notes |
|---|---|---|
page? | PageInfo | |
vars? | Record<string, Json> | |
locale? | string | 2 to 35 chars |
identify
The visitor signed in after hello. See Identity verification.
| Field | Type | Limits and notes |
|---|---|---|
identity | Identity |
voice.start
Start voice: conversation (hands-free) or ptt (push-to-talk). The server answers voice.state connecting, then voice.ready, or an error. The server doesn't use camera: the chat window turns the camera on itself, after the visitor agrees.
| Field | Type | Limits and notes |
|---|---|---|
mode | VoiceMode | |
camera? | boolean |
voice.stop
End voice.
No fields.
voice.ptt
Push-to-talk pressed (down) or released (up). Pressing also stops the agent talking.
| Field | Type | Limits and notes |
|---|---|---|
state | 'down' | 'up' |
voice.interrupt
Stop the agent talking now.
No fields.
input.vad
A hint that the visitor is speaking. Accepted and ignored by this server.
| Field | Type | Limits and notes |
|---|---|---|
speaking | boolean |
playback.mark
How many milliseconds of utterance utt have played (the widget sends it every 250 ms while the agent talks). If the visitor cuts in, the stored reply is trimmed to what was heard.
| Field | Type | Limits and notes |
|---|---|---|
utt | number | at least 0 |
ms | number | at least 0 |
playback.done
Utterance utt finished playing.
| Field | Type | Limits and notes |
|---|---|---|
utt | number | at least 0 |
ms | number | at least 0 |
playback.cleared
The widget stopped playing utt early (after agent.audio.clear, or push-to-talk).
| Field | Type | Limits and notes |
|---|---|---|
utt | number | at least 0 |
ms | number | at least 0 |
vision.start
The visitor agreed to the camera (consent must be true) and it is on. The bot's vision.webcam.mode decides how frames are used, not the mode sent here.
| Field | Type | Limits and notes |
|---|---|---|
mode | VisionMode | |
consent | true | must be true |
width | number | 16 to 4096 |
height | number | 16 to 4096 |
vision.stop
The camera is off.
No fields.
tool.result
The page's answer to a tool.call: result (JSON) when ok, else error.
| Field | Type | Limits and notes |
|---|---|---|
callId | string | 1 to 100 chars |
ok | boolean | |
result? | Json | |
error? | string | up to 2000 chars |
tool.confirm.reply
The visitor's answer to a tool.confirm: allowed true to run the tool, false not to. The first answer from any of the conversation's tabs counts.
| Field | Type | Limits and notes |
|---|---|---|
callId | string | 1 to 100 chars |
allowed | boolean |
tools.register
Replace the list of tools this page offers (up to 64). The bot's tools.client settings decide which ones the agent may use.
| Field | Type | Limits and notes |
|---|---|---|
tools | ClientToolDecl[] | up to 64 items |
feedback
Thumbs on one of the agent's messages: 1, -1, or 0 to clear, with an optional comment.
| Field | Type | Limits and notes |
|---|---|---|
messageId | string | 1 to 100 chars |
value | -1 | 0 | 1 | |
comment? | string | up to 2000 chars |
csat
A 1 to 5 rating for the conversation, with an optional comment.
| Field | Type | Limits and notes |
|---|---|---|
score | number | 1 to 5 |
comment? | string | up to 2000 chars |
lead.submit
The lead form. Only keys in the bot's behavior.leadCapture.fields are kept.
| Field | Type | Limits and notes |
|---|---|---|
fields | Record<string, string> |
handoff.request
Ask for a person from the team (the chat menu's "Talk to a person").
| Field | Type | Limits and notes |
|---|---|---|
reason? | string | up to 1000 chars |
ping
Answered with pong and the same ts.
| Field | Type | Limits and notes |
|---|---|---|
ts | number |
Server messages
session.ready
The answer to hello (and to conversation.new). conversationId is null until the visitor's first message. Keep visitorToken and resumeToken to resume later. bot is the bot's public config (below), status the conversation's state, preview whether this is a draft preview.
| Field | Type | Notes |
|---|---|---|
conversationId | string | null | |
visitorId | string | |
visitorToken | string | |
resumeToken | string | null | |
resumed | boolean | |
status | ConversationStatus | |
bot | PublicBotConfig | |
limits | { maxMessageChars: number; maxUploadMB: number } | |
preview | boolean |
history.sync
The messages of a resumed conversation, after fromSeq (the lastSeq you sent, or 0), up to toSeq. Empty when the bot hides history from returning visitors (behavior.persistence.showHistory: false) and no lastSeq was sent.
| Field | Type | Notes |
|---|---|---|
messages | WireMessage[] | |
fromSeq | number | |
toSeq | number |
conversation.started
The visitor's first message created a conversation. Keep the new resumeToken.
| Field | Type | Notes |
|---|---|---|
conversationId | string | |
resumeToken | string |
message.created
A stored message: the visitor's own (with its clientMsgId, also sent to their other tabs), one from a person on the team, or a system line such as "Ana joined the chat" (role system). The agent's replies arrive as agent.response.* instead.
| Field | Type | Notes |
|---|---|---|
message | WireMessage |
message.updated
A stored message changed (its feedback).
| Field | Type | Notes |
|---|---|---|
id | string | |
patch | Partial<Pick<WireMessage, 'text' | 'status' | 'feedback' | 'citations'>> |
agent.response.start
The agent started a reply. messageId is the id it will be stored under.
| Field | Type | Notes |
|---|---|---|
responseId | string | |
messageId | string | |
modality | Modality |
agent.text.delta
More of the reply. In voice, utt ties the words to the audio.
| Field | Type | Notes |
|---|---|---|
responseId | string | |
delta | string | |
utt? | number | |
atMs? | number |
agent.response.end
The reply finished. text is the whole reply (mood tags removed), status is complete, interrupted or error, and seq its sequence number (0 if nothing was stored).
| Field | Type | Notes |
|---|---|---|
responseId | string | |
messageId | string | |
seq | number | |
status | 'complete' | 'interrupted' | 'error' | |
text | string |
agent.status
What the agent is doing: thinking, tool (with the tool's name and a label such as "Looking that up"), speaking, listening or idle.
| Field | Type | Notes |
|---|---|---|
state | AgentState | |
tool? | { name: string; label: string } |
agent.citations
Knowledge-base sources for a reply (up to 5), after its agent.response.end, when behavior.citations is on.
| Field | Type | Notes |
|---|---|---|
responseId | string | |
sources | Citation[] |
face.express
A brief expression for the face (from an [[express:...]] tag).
| Field | Type | Notes |
|---|---|---|
name | string | |
seconds? | number |
face.mood
The face's mood (from a [[mood:...]] tag, or the voice engine's express tool).
| Field | Type | Notes |
|---|---|---|
name | string | null | |
intensity? | number |
voice.ready
Voice is live. Send microphone audio as binary frames: PCM16 mono at in.rate (16000 or 24000 Hz, depending on the engine), in.frameMs per frame. Agent audio arrives at 24 kHz.
| Field | Type | Notes |
|---|---|---|
voiceSessionId | string | |
engine | string | |
in | { rate: 16000 | 24000; codec: 'pcm16'; frameMs: number } | |
out | { rate: 24000; codec: 'pcm16' } | |
bargeIn | boolean | |
ptt | boolean |
voice.state
connecting, live, reconnecting (the connection to the voice provider dropped and is coming back) or ended (with a reason such as visitor, idle, time_limit, cap, handoff or error: ...).
| Field | Type | Notes |
|---|---|---|
state | 'connecting' | 'live' | 'reconnecting' | 'ended' | |
reason? | string |
user.transcript
What the visitor said: partial text while they speak, then final.
| Field | Type | Notes |
|---|---|---|
itemId | string | |
text | string | |
final | boolean |
speech.started
The visitor started speaking.
No fields.
speech.stopped
The visitor stopped speaking (the agent is about to answer).
No fields.
agent.audio.start
Agent audio for utterance utt begins; binary frames follow.
| Field | Type | Notes |
|---|---|---|
utt | number | |
responseId | string |
agent.audio.end
No more audio for utt (let what is buffered finish).
| Field | Type | Notes |
|---|---|---|
utt | number |
agent.audio.clear
Stop playing utt now: the visitor talked over it (barge_in) or stopped it (cancel).
| Field | Type | Notes |
|---|---|---|
utt | number | |
reason | 'barge_in' | 'cancel' | 'handoff' |
vision.state
Whether the camera is in use and how: the mode the server chose, frames a second to send (0: only when asked), the largest width, and the JPEG quality.
| Field | Type | Notes |
|---|---|---|
active | boolean | |
mode | VisionMode | |
fps | number | |
maxWidth | number | |
quality | number |
vision.request
Send a camera frame now, with requestId as the frame's stream number (the agent used look_at_camera).
| Field | Type | Notes |
|---|---|---|
requestId | number | |
reason? | string |
tool.call
Run a page tool and answer with tool.result within timeoutMs. args were checked against the tool's schema.
| Field | Type | Notes |
|---|---|---|
callId | string | |
name | string | |
args | Json | |
timeoutMs | number |
tool.cancel
Withdraws a tool.call (the server gave up on it, and the page tool's signal is aborted) or a tool.confirm (it was answered in another tab, timed out, or the reply was stopped).
| Field | Type | Notes |
|---|---|---|
callId | string |
tool.confirm
A tool marked "ask first" wants to run: show question and the arguments in details, then answer with tool.confirm.reply. Every tab of the conversation gets it. No answer within 2 minutes counts as no.
| Field | Type | Notes |
|---|---|---|
callId | string | |
name | string | |
question | string | |
details | Array<{ label: string; value: string }> |
tool.status
Not sent by this server version (it sends agent.status with state tool).
| Field | Type | Notes |
|---|---|---|
callId | string | |
name | string | |
label | string | |
status | 'started' | 'succeeded' | 'failed' |
handoff.state
Hand-off to a person: pending, active (with the person's name), ended or unavailable (with the bot's offline text in message).
| Field | Type | Notes |
|---|---|---|
status | 'pending' | 'active' | 'ended' | 'unavailable' | |
agent? | Author | |
message? | string |
human.typing
A person from the team is typing.
| Field | Type | Notes |
|---|---|---|
active | boolean | |
agentName? | string |
lead.form
Show the lead form with these fields. reason is offline when no one from the team is available.
| Field | Type | Notes |
|---|---|---|
fields | LeadField[] | |
reason? | string |
conversation.ended
The conversation was closed (the agent said goodbye, or the team closed it). csat says whether to ask for a rating.
| Field | Type | Notes |
|---|---|---|
reason | string | |
csat | boolean |
config.updated
The bot was republished. The widget reconnects to load the new look and settings; the conversation carries on with the version it started on.
No fields.
notice
A note to show the visitor (in previews: why a reply failed).
| Field | Type | Notes |
|---|---|---|
level | 'info' | 'warn' | |
text | string |
error
code (see Errors), a message, fatal (the server closes the socket next) and, when rate limited, retryAfterMs.
| Field | Type | Notes |
|---|---|---|
code | ErrorCode | |
message | string | |
fatal | boolean | |
retryAfterMs? | number |
pong
The answer to ping.
| Field | Type | Notes |
|---|---|---|
ts | number |
Binary frames
Every binary frame starts with an 8-byte little-endian header:
| Offset | Size | Field | |
|---|---|---|---|
| 0 | u8 | kind | What the frame is (below) |
| 1 | u8 | codec | How the payload is encoded (below) |
| 2 | u16 | stream | Depends on the kind |
| 4 | u32 | seq | A counter within the stream |
| 8 | ... | payload |
| Kind | Name | |
|---|---|---|
0x01 | MIC | client -> server: microphone PCM16 at voice.ready.in.rate, ~20 ms per frame. Stream: 0 (the server ignores it). Codec: PCM16. |
0x02 | AGENT_AUDIO | server -> client: agent PCM16 at 24 kHz, <= 100 ms per frame. Stream: the utterance id (utt). Codec: PCM16, at most 100 ms (4,800 bytes) per frame. |
0x03 | CAMERA | client -> server: a webcam frame. Stream: 0 for periodic frames, or the requestId of a vision.request. Codec: JPEG. |
0x04 | MOUTH | server -> client: mouth packets (reserved). Reserved: not sent in this version. |
| Codec | Name | |
|---|---|---|
0x00 | PCM16 | Signed 16-bit little-endian mono samples. |
0x01 | OPUS | Reserved. |
0x10 | JPEG | A JPEG image. |
0x11 | WEBP | Reserved. |
0x20 | JSON | Reserved. |
The largest payloads the server accepts from a widget:
| Kind | Max |
|---|---|
MIC | 16 KB |
CAMERA | 512 KB |
Send microphone audio only after voice.ready, and camera frames only after vision.state says active: the server drops frames it isn't expecting.
Types
The types the messages refer to, as written in packages/protocol/src/types.ts.
Json
ts
export type Json = string | number | boolean | null | Json[] | { [key: string]: Json };JsonSchema
ts
export type JsonSchema = {
type?: string | string[];
description?: string;
properties?: Record<string, JsonSchema>;
required?: string[];
enum?: Json[];
items?: JsonSchema;
[key: string]: unknown;
};Role
ts
export type Role = 'user' | 'assistant' | 'human_agent' | 'system';Modality
ts
export type Modality = 'text' | 'voice';ConversationStatus
ts
export type ConversationStatus = 'ai' | 'handoff_pending' | 'human' | 'closed';VisionMode
ts
export type VisionMode = 'continuous' | 'per_turn' | 'on_demand';VoiceMode
ts
export type VoiceMode = 'conversation' | 'ptt';Corner
ts
export type Corner = 'bottom-right' | 'bottom-left' | 'top-right' | 'top-left';Layout
ts
export type Layout = 'panel' | 'drawer' | 'floating' | 'inline';ThemeMode
ts
export type ThemeMode = 'light' | 'dark' | 'auto';AttachmentRef
ts
export interface AttachmentRef {
id: string;
url: string;
mime: string;
w?: number;
h?: number;
origin?: 'upload' | 'webcam' | 'tool';
}Citation
ts
export interface Citation {
title: string;
url?: string;
}Author
ts
export interface Author {
name: string;
avatarUrl?: string;
}WireMessage
ts
/** A message as the widget sees it. */
export interface WireMessage {
id: string;
seq: number;
role: Role;
text: string;
attachments: AttachmentRef[];
modality: Modality;
createdAt: number;
author?: Author;
clientMsgId?: string;
status?: 'complete' | 'interrupted' | 'error';
feedback?: -1 | 0 | 1;
citations?: Citation[];
}PageInfo
ts
export interface PageInfo {
url: string;
title?: string;
referrer?: string;
}ClientToolDecl
ts
export interface ClientToolDecl {
name: string;
description: string;
parameters: JsonSchema;
confirm?: boolean | string;
}Identity
ts
export interface Identity {
userId: string;
/** HMAC-SHA256(identity secret, userId), hex, computed on the customer's server. */
userHash?: string;
name?: string;
email?: string;
traits?: Record<string, string | number | boolean | null>;
}LeadField
ts
export interface LeadField {
key: string;
label: string;
type: 'text' | 'email' | 'tel' | 'textarea' | 'select';
required: boolean;
options?: string[];
}FaceSkinRef
ts
export type FaceSkinRef =
| { kind: 'catalogue'; id: string }
| { kind: 'custom'; image: string; landmarks: string; teeth?: string }
| null;LauncherConfig
ts
export interface LauncherConfig {
type: 'bubble' | 'tab' | 'none';
position: Corner;
edge: 'right' | 'left' | 'bottom';
align: 'start' | 'center' | 'end';
label: string;
icon: 'face' | 'chat' | 'avatar';
offset: { x: number; y: number };
size: number;
hideOnMobile: boolean;
}AppearanceConfig
ts
export interface AppearanceConfig {
layout: Layout;
launcher: LauncherConfig;
theme: { mode: ThemeMode; accent: string; radius: number; font: 'system' | 'inter' | 'serif' | 'mono' };
panel: { width: number; height: number };
drawer: { width: number; side: 'auto' | 'left' | 'right'; modal: boolean };
floating: { width: number; height: number; composer: boolean };
mobile: { fullscreen: boolean };
face: { show: boolean; size: 'small' | 'medium' | 'large'; textReplies: 'still' | 'mime' | 'speak' };
captions: boolean;
zIndex: number;
}GreetingConfig
ts
export interface GreetingConfig {
enabled: boolean;
text: string;
delayMs: number;
frequency: 'session' | 'visitor' | 'always';
pages: { include: string[]; exclude: string[] };
mobile: boolean;
}PublicBotConfig
ts
export interface PublicBotConfig {
id: string;
version: number;
name: string;
agent: { name: string; role?: string; avatarUrl?: string };
face: { config: Record<string, unknown>; skin: FaceSkinRef; defaultMood: string | null; thumb?: string };
appearance: AppearanceConfig;
texts: {
welcome: string;
suggestedPrompts: string[];
placeholder: string;
title: string;
subtitle: string;
offline: string;
strings: Record<string, string>;
};
greeting: GreetingConfig;
features: {
voice: { enabled: boolean; modes: VoiceMode[]; engine: string | null };
webcam: { enabled: boolean; mode: VisionMode; consentText: string; privacyUrl?: string };
uploads: { enabled: boolean; maxPerMessage: number; maxMB: number };
handoff: boolean;
feedback: boolean;
csat: boolean;
leadForm: { enabled: boolean; when: 'start' | 'tool' | 'before_handoff' | 'offline'; fields: LeadField[] } | null;
newConversation: boolean;
transcriptDownload: boolean;
/** typed replies are read aloud (textReplies 'speak', with a speech voice set up) */
spokenReplies: boolean;
};
availability: { online: boolean; nextOpenAt: string | null };
branding: { poweredBy: boolean };
locale: { default: string };
}AgentState
ts
export type AgentState = 'thinking' | 'tool' | 'speaking' | 'listening' | 'idle';