Skip to content

Widget protocol ​

The chat window talks to the server over one WebSocket, wss://<host>/v1/widget/ws?bot=<public id>. This page describes protocol version 1. You don't need it to embed the chat; it is here for debugging (the browser's network tab shows every message) and for anyone building on the server.

Only pages on the chat server's own origin can open the socket: the server checks the Origin header (unless NODE_ENV is production, it also accepts localhost and 127.0.0.1 on any port). Your website never talks to it directly; widget.js puts the chat window, which does, in an iframe.

  • Text frames are JSON objects with a t field (the message type). Up to 256 KB each.
  • Binary frames carry audio and camera pictures, with an 8-byte header (see Binary frames).

Connecting ​

  1. Open the socket. Within 5 seconds, send hello with protocol: 1. Any other first message, or none, closes the socket with 4401.
  2. The server checks the bot is live, that the page in hello.page.url is in the bot's allowed origins, and (when the bot requires it) the visitor's identity. It answers session.ready, or a fatal error and a close code (see close codes).
  3. Keep session.ready.visitorToken: it is who the visitor is for this bot. Send it in the next hello.

The conversation starts with the visitor's first chat.send (or voice, a lead form or a hand-off request): the server sends conversation.started with its id and a resumeToken.

Resuming ​

Stored messages carry seq, increasing by one within a conversation. To pick up after a reload or a dropped connection, send hello with visitorToken, conversationId, resumeToken and lastSeq (the highest seq you have). The server answers session.ready with resumed: true, then history.sync with the messages after lastSeq. A conversation resumes while it is open and (for one the AI is handling) its last message is within the bot's behavior.persistence.resumeWindowHours; otherwise the visitor starts a new one.

The server pings every 20 seconds at the WebSocket level and drops connections that don't answer. The ping message is for measuring round trips.

Client messages ​

hello ​

The first message, within 5 seconds of connecting. With conversationId (and its resumeToken) it resumes that conversation; without one the server picks up the visitor's latest open conversation within behavior.persistence.resumeWindowHours. lastSeq asks only for messages after that sequence number; without it, a bot with behavior.persistence.showHistory: false sends no history (the conversation still carries on). Facts in context reach the agent as a note when they differ from the last ones it got. This server doesn't use tz or caps yet.

FieldTypeLimits and notes
protocolnumber
visitorToken?stringup to 200 chars
conversationId?string1 to 100 chars
resumeToken?stringup to 200 chars
lastSeq?numberat least 0
identity?Identity
page?PageInfo
locale?stringup to 35 chars
tz?stringup to 64 chars
caps?{ audioIn: boolean; audioOut: boolean; webcam: boolean }
clientTools?ClientToolDecl[]up to 64 items
context?Record<string, Json>
preview?stringup to 2000 chars. A server-signed admin preview token: the session uses the bot's draft config.

conversation.new ​

Leave the current conversation and start fresh. The old one is closed, unless it held no more than the welcome and one message. The server answers with a new session.ready.

No fields.

chat.send ​

A visitor message. clientMsgId (your own id) comes back on the stored message so you can match it. attachments are ids from the upload endpoint; the bot's vision.uploads.maxPerMessage applies. Text longer than security.maxMessageChars is cut.

FieldTypeLimits and notes
clientMsgIdstring1 to 64 chars
textstringup to 20000 chars
attachments?string[]up to 10 items, each up to 100 chars

chat.typing ​

The visitor is typing (shown to your team in the admin panel).

FieldTypeLimits and notes
activeboolean

chat.cancel ​

Stop the reply being written now. responseId is ignored: the reply in progress stops.

FieldTypeLimits and notes
responseId?string1 to 100 chars

context.update ​

The page changed (page), new facts for the agent (vars, merged into earlier ones), or the visitor's language changed (locale: the server remembers it and, with identity.languagePolicy: match_visitor, tells the agent to answer in it). Each reaches the agent as a note.

FieldTypeLimits and notes
page?PageInfo
vars?Record<string, Json>
locale?string2 to 35 chars

identify ​

The visitor signed in after hello. See Identity verification.

FieldTypeLimits and notes
identityIdentity

voice.start ​

Start voice: conversation (hands-free) or ptt (push-to-talk). The server answers voice.state connecting, then voice.ready, or an error. The server doesn't use camera: the chat window turns the camera on itself, after the visitor agrees.

FieldTypeLimits and notes
modeVoiceMode
camera?boolean

voice.stop ​

End voice.

No fields.

voice.ptt ​

Push-to-talk pressed (down) or released (up). Pressing also stops the agent talking.

FieldTypeLimits and notes
state'down' | 'up'

voice.interrupt ​

Stop the agent talking now.

No fields.

input.vad ​

A hint that the visitor is speaking. Accepted and ignored by this server.

FieldTypeLimits and notes
speakingboolean

playback.mark ​

How many milliseconds of utterance utt have played (the widget sends it every 250 ms while the agent talks). If the visitor cuts in, the stored reply is trimmed to what was heard.

FieldTypeLimits and notes
uttnumberat least 0
msnumberat least 0

playback.done ​

Utterance utt finished playing.

FieldTypeLimits and notes
uttnumberat least 0
msnumberat least 0

playback.cleared ​

The widget stopped playing utt early (after agent.audio.clear, or push-to-talk).

FieldTypeLimits and notes
uttnumberat least 0
msnumberat least 0

vision.start ​

The visitor agreed to the camera (consent must be true) and it is on. The bot's vision.webcam.mode decides how frames are used, not the mode sent here.

FieldTypeLimits and notes
modeVisionMode
consenttruemust be true
widthnumber16 to 4096
heightnumber16 to 4096

vision.stop ​

The camera is off.

No fields.

tool.result ​

The page's answer to a tool.call: result (JSON) when ok, else error.

FieldTypeLimits and notes
callIdstring1 to 100 chars
okboolean
result?Json
error?stringup to 2000 chars

tool.confirm.reply ​

The visitor's answer to a tool.confirm: allowed true to run the tool, false not to. The first answer from any of the conversation's tabs counts.

FieldTypeLimits and notes
callIdstring1 to 100 chars
allowedboolean

tools.register ​

Replace the list of tools this page offers (up to 64). The bot's tools.client settings decide which ones the agent may use.

FieldTypeLimits and notes
toolsClientToolDecl[]up to 64 items

feedback ​

Thumbs on one of the agent's messages: 1, -1, or 0 to clear, with an optional comment.

FieldTypeLimits and notes
messageIdstring1 to 100 chars
value-1 | 0 | 1
comment?stringup to 2000 chars

csat ​

A 1 to 5 rating for the conversation, with an optional comment.

FieldTypeLimits and notes
scorenumber1 to 5
comment?stringup to 2000 chars

lead.submit ​

The lead form. Only keys in the bot's behavior.leadCapture.fields are kept.

FieldTypeLimits and notes
fieldsRecord<string, string>

handoff.request ​

Ask for a person from the team (the chat menu's "Talk to a person").

FieldTypeLimits and notes
reason?stringup to 1000 chars

ping ​

Answered with pong and the same ts.

FieldTypeLimits and notes
tsnumber

Server messages ​

session.ready ​

The answer to hello (and to conversation.new). conversationId is null until the visitor's first message. Keep visitorToken and resumeToken to resume later. bot is the bot's public config (below), status the conversation's state, preview whether this is a draft preview.

FieldTypeNotes
conversationIdstring | null
visitorIdstring
visitorTokenstring
resumeTokenstring | null
resumedboolean
statusConversationStatus
botPublicBotConfig
limits{ maxMessageChars: number; maxUploadMB: number }
previewboolean

history.sync ​

The messages of a resumed conversation, after fromSeq (the lastSeq you sent, or 0), up to toSeq. Empty when the bot hides history from returning visitors (behavior.persistence.showHistory: false) and no lastSeq was sent.

FieldTypeNotes
messagesWireMessage[]
fromSeqnumber
toSeqnumber

conversation.started ​

The visitor's first message created a conversation. Keep the new resumeToken.

FieldTypeNotes
conversationIdstring
resumeTokenstring

message.created ​

A stored message: the visitor's own (with its clientMsgId, also sent to their other tabs), one from a person on the team, or a system line such as "Ana joined the chat" (role system). The agent's replies arrive as agent.response.* instead.

FieldTypeNotes
messageWireMessage

message.updated ​

A stored message changed (its feedback).

FieldTypeNotes
idstring
patchPartial<Pick<WireMessage, 'text' | 'status' | 'feedback' | 'citations'>>

agent.response.start ​

The agent started a reply. messageId is the id it will be stored under.

FieldTypeNotes
responseIdstring
messageIdstring
modalityModality

agent.text.delta ​

More of the reply. In voice, utt ties the words to the audio.

FieldTypeNotes
responseIdstring
deltastring
utt?number
atMs?number

agent.response.end ​

The reply finished. text is the whole reply (mood tags removed), status is complete, interrupted or error, and seq its sequence number (0 if nothing was stored).

FieldTypeNotes
responseIdstring
messageIdstring
seqnumber
status'complete' | 'interrupted' | 'error'
textstring

agent.status ​

What the agent is doing: thinking, tool (with the tool's name and a label such as "Looking that up"), speaking, listening or idle.

FieldTypeNotes
stateAgentState
tool?{ name: string; label: string }

agent.citations ​

Knowledge-base sources for a reply (up to 5), after its agent.response.end, when behavior.citations is on.

FieldTypeNotes
responseIdstring
sourcesCitation[]

face.express ​

A brief expression for the face (from an [[express:...]] tag).

FieldTypeNotes
namestring
seconds?number

face.mood ​

The face's mood (from a [[mood:...]] tag, or the voice engine's express tool).

FieldTypeNotes
namestring | null
intensity?number

voice.ready ​

Voice is live. Send microphone audio as binary frames: PCM16 mono at in.rate (16000 or 24000 Hz, depending on the engine), in.frameMs per frame. Agent audio arrives at 24 kHz.

FieldTypeNotes
voiceSessionIdstring
enginestring
in{ rate: 16000 | 24000; codec: 'pcm16'; frameMs: number }
out{ rate: 24000; codec: 'pcm16' }
bargeInboolean
pttboolean

voice.state ​

connecting, live, reconnecting (the connection to the voice provider dropped and is coming back) or ended (with a reason such as visitor, idle, time_limit, cap, handoff or error: ...).

FieldTypeNotes
state'connecting' | 'live' | 'reconnecting' | 'ended'
reason?string

user.transcript ​

What the visitor said: partial text while they speak, then final.

FieldTypeNotes
itemIdstring
textstring
finalboolean

speech.started ​

The visitor started speaking.

No fields.

speech.stopped ​

The visitor stopped speaking (the agent is about to answer).

No fields.

agent.audio.start ​

Agent audio for utterance utt begins; binary frames follow.

FieldTypeNotes
uttnumber
responseIdstring

agent.audio.end ​

No more audio for utt (let what is buffered finish).

FieldTypeNotes
uttnumber

agent.audio.clear ​

Stop playing utt now: the visitor talked over it (barge_in) or stopped it (cancel).

FieldTypeNotes
uttnumber
reason'barge_in' | 'cancel' | 'handoff'

vision.state ​

Whether the camera is in use and how: the mode the server chose, frames a second to send (0: only when asked), the largest width, and the JPEG quality.

FieldTypeNotes
activeboolean
modeVisionMode
fpsnumber
maxWidthnumber
qualitynumber

vision.request ​

Send a camera frame now, with requestId as the frame's stream number (the agent used look_at_camera).

FieldTypeNotes
requestIdnumber
reason?string

tool.call ​

Run a page tool and answer with tool.result within timeoutMs. args were checked against the tool's schema.

FieldTypeNotes
callIdstring
namestring
argsJson
timeoutMsnumber

tool.cancel ​

Withdraws a tool.call (the server gave up on it, and the page tool's signal is aborted) or a tool.confirm (it was answered in another tab, timed out, or the reply was stopped).

FieldTypeNotes
callIdstring

tool.confirm ​

A tool marked "ask first" wants to run: show question and the arguments in details, then answer with tool.confirm.reply. Every tab of the conversation gets it. No answer within 2 minutes counts as no.

FieldTypeNotes
callIdstring
namestring
questionstring
detailsArray<{ label: string; value: string }>

tool.status ​

Not sent by this server version (it sends agent.status with state tool).

FieldTypeNotes
callIdstring
namestring
labelstring
status'started' | 'succeeded' | 'failed'

handoff.state ​

Hand-off to a person: pending, active (with the person's name), ended or unavailable (with the bot's offline text in message).

FieldTypeNotes
status'pending' | 'active' | 'ended' | 'unavailable'
agent?Author
message?string

human.typing ​

A person from the team is typing.

FieldTypeNotes
activeboolean
agentName?string

lead.form ​

Show the lead form with these fields. reason is offline when no one from the team is available.

FieldTypeNotes
fieldsLeadField[]
reason?string

conversation.ended ​

The conversation was closed (the agent said goodbye, or the team closed it). csat says whether to ask for a rating.

FieldTypeNotes
reasonstring
csatboolean

config.updated ​

The bot was republished. The widget reconnects to load the new look and settings; the conversation carries on with the version it started on.

No fields.

notice ​

A note to show the visitor (in previews: why a reply failed).

FieldTypeNotes
level'info' | 'warn'
textstring

error ​

code (see Errors), a message, fatal (the server closes the socket next) and, when rate limited, retryAfterMs.

FieldTypeNotes
codeErrorCode
messagestring
fatalboolean
retryAfterMs?number

pong ​

The answer to ping.

FieldTypeNotes
tsnumber

Binary frames ​

Every binary frame starts with an 8-byte little-endian header:

OffsetSizeField
0u8kindWhat the frame is (below)
1u8codecHow the payload is encoded (below)
2u16streamDepends on the kind
4u32seqA counter within the stream
8...payload
KindName
0x01MICclient -> server: microphone PCM16 at voice.ready.in.rate, ~20 ms per frame. Stream: 0 (the server ignores it). Codec: PCM16.
0x02AGENT_AUDIOserver -> client: agent PCM16 at 24 kHz, <= 100 ms per frame. Stream: the utterance id (utt). Codec: PCM16, at most 100 ms (4,800 bytes) per frame.
0x03CAMERAclient -> server: a webcam frame. Stream: 0 for periodic frames, or the requestId of a vision.request. Codec: JPEG.
0x04MOUTHserver -> client: mouth packets (reserved). Reserved: not sent in this version.
CodecName
0x00PCM16Signed 16-bit little-endian mono samples.
0x01OPUSReserved.
0x10JPEGA JPEG image.
0x11WEBPReserved.
0x20JSONReserved.

The largest payloads the server accepts from a widget:

KindMax
MIC16 KB
CAMERA512 KB

Send microphone audio only after voice.ready, and camera frames only after vision.state says active: the server drops frames it isn't expecting.

Types ​

The types the messages refer to, as written in packages/protocol/src/types.ts.

Json ​

ts
export type Json = string | number | boolean | null | Json[] | { [key: string]: Json };

JsonSchema ​

ts
export type JsonSchema = {
  type?: string | string[];
  description?: string;
  properties?: Record<string, JsonSchema>;
  required?: string[];
  enum?: Json[];
  items?: JsonSchema;
  [key: string]: unknown;
};

Role ​

ts
export type Role = 'user' | 'assistant' | 'human_agent' | 'system';

Modality ​

ts
export type Modality = 'text' | 'voice';

ConversationStatus ​

ts
export type ConversationStatus = 'ai' | 'handoff_pending' | 'human' | 'closed';

VisionMode ​

ts
export type VisionMode = 'continuous' | 'per_turn' | 'on_demand';

VoiceMode ​

ts
export type VoiceMode = 'conversation' | 'ptt';

Corner ​

ts
export type Corner = 'bottom-right' | 'bottom-left' | 'top-right' | 'top-left';

Layout ​

ts
export type Layout = 'panel' | 'drawer' | 'floating' | 'inline';

ThemeMode ​

ts
export type ThemeMode = 'light' | 'dark' | 'auto';

AttachmentRef ​

ts
export interface AttachmentRef {
  id: string;
  url: string;
  mime: string;
  w?: number;
  h?: number;
  origin?: 'upload' | 'webcam' | 'tool';
}

Citation ​

ts
export interface Citation {
  title: string;
  url?: string;
}

Author ​

ts
export interface Author {
  name: string;
  avatarUrl?: string;
}

WireMessage ​

ts
/** A message as the widget sees it. */
export interface WireMessage {
  id: string;
  seq: number;
  role: Role;
  text: string;
  attachments: AttachmentRef[];
  modality: Modality;
  createdAt: number;
  author?: Author;
  clientMsgId?: string;
  status?: 'complete' | 'interrupted' | 'error';
  feedback?: -1 | 0 | 1;
  citations?: Citation[];
}

PageInfo ​

ts
export interface PageInfo {
  url: string;
  title?: string;
  referrer?: string;
}

ClientToolDecl ​

ts
export interface ClientToolDecl {
  name: string;
  description: string;
  parameters: JsonSchema;
  confirm?: boolean | string;
}

Identity ​

ts
export interface Identity {
  userId: string;
  /** HMAC-SHA256(identity secret, userId), hex, computed on the customer's server. */
  userHash?: string;
  name?: string;
  email?: string;
  traits?: Record<string, string | number | boolean | null>;
}

LeadField ​

ts
export interface LeadField {
  key: string;
  label: string;
  type: 'text' | 'email' | 'tel' | 'textarea' | 'select';
  required: boolean;
  options?: string[];
}

FaceSkinRef ​

ts
export type FaceSkinRef =
  | { kind: 'catalogue'; id: string }
  | { kind: 'custom'; image: string; landmarks: string; teeth?: string }
  | null;

LauncherConfig ​

ts
export interface LauncherConfig {
  type: 'bubble' | 'tab' | 'none';
  position: Corner;
  edge: 'right' | 'left' | 'bottom';
  align: 'start' | 'center' | 'end';
  label: string;
  icon: 'face' | 'chat' | 'avatar';
  offset: { x: number; y: number };
  size: number;
  hideOnMobile: boolean;
}

AppearanceConfig ​

ts
export interface AppearanceConfig {
  layout: Layout;
  launcher: LauncherConfig;
  theme: { mode: ThemeMode; accent: string; radius: number; font: 'system' | 'inter' | 'serif' | 'mono' };
  panel: { width: number; height: number };
  drawer: { width: number; side: 'auto' | 'left' | 'right'; modal: boolean };
  floating: { width: number; height: number; composer: boolean };
  mobile: { fullscreen: boolean };
  face: { show: boolean; size: 'small' | 'medium' | 'large'; textReplies: 'still' | 'mime' | 'speak' };
  captions: boolean;
  zIndex: number;
}

GreetingConfig ​

ts
export interface GreetingConfig {
  enabled: boolean;
  text: string;
  delayMs: number;
  frequency: 'session' | 'visitor' | 'always';
  pages: { include: string[]; exclude: string[] };
  mobile: boolean;
}

PublicBotConfig ​

ts
export interface PublicBotConfig {
  id: string;
  version: number;
  name: string;
  agent: { name: string; role?: string; avatarUrl?: string };
  face: { config: Record<string, unknown>; skin: FaceSkinRef; defaultMood: string | null; thumb?: string };
  appearance: AppearanceConfig;
  texts: {
    welcome: string;
    suggestedPrompts: string[];
    placeholder: string;
    title: string;
    subtitle: string;
    offline: string;
    strings: Record<string, string>;
  };
  greeting: GreetingConfig;
  features: {
    voice: { enabled: boolean; modes: VoiceMode[]; engine: string | null };
    webcam: { enabled: boolean; mode: VisionMode; consentText: string; privacyUrl?: string };
    uploads: { enabled: boolean; maxPerMessage: number; maxMB: number };
    handoff: boolean;
    feedback: boolean;
    csat: boolean;
    leadForm: { enabled: boolean; when: 'start' | 'tool' | 'before_handoff' | 'offline'; fields: LeadField[] } | null;
    newConversation: boolean;
    transcriptDownload: boolean;
    /** typed replies are read aloud (textReplies 'speak', with a speech voice set up) */
    spokenReplies: boolean;
  };
  availability: { online: boolean; nextOpenAt: string | null };
  branding: { poweredBy: boolean };
  locale: { default: string };
}

AgentState ​

ts
export type AgentState = 'thinking' | 'tool' | 'speaking' | 'listening' | 'idle';

Wireface Chat 0.1.0. These docs are served by your own server.