Appearance
Protocol and engines
The widget protocol
The chat frame and the server talk over one WebSocket, /v1/widget/ws?bot=pk_.... Every message, its fields and the binary frame layout are in the Widget protocol reference, which is generated from the code. In brief:
- Control messages are JSON objects with a
tfield. The frame's first message ishello, withprotocol: 1(PROTOCOL_VERSION); a widget on another version getsprotocol_unsupportedand is asked to reload. - Stored messages carry a per-conversation
seq. A widget that reconnects sends itslastSeqand gets what it missed inhistory.sync. - A reply streams as
agent.response.start, any number ofagent.text.delta, thenagent.response.end, all with the sameresponseId. - Audio and camera pictures are binary frames: an 8-byte little-endian header (kind, codec, stream, sequence number), then PCM16 or an encoded image.
The definitions live in packages/protocol/src:
| File | Holds |
|---|---|
types.ts | ClientMessage and ServerMessage, the shared types, PROTOCOL_VERSION |
schemas.ts | The zod schemas the server checks every client message with (parseClientMessage()); server only |
binary.ts | FrameKind, Codec, encodeFrame() and decodeFrame() |
errors.ts | Error codes and WebSocket close codes |
port.ts | The messages between the loader and the frame, over their MessagePort (PORT_PROTOCOL) |
To add a client message: its type in types.ts, its schema in schemas.ts, a case in WidgetConnection.dispatch() (gateway/gateway.ts) or a handler that a module registers in conn.handlers, the frame's side in packages/widget/src/frame/chat.ts, and a note in docs/scripts/notes/protocol.ts. A server message needs the type, the frame's handling and the note. The docs generator fails until each is there (see Testing).
Two kinds of compatibility matter. The frame is always the one the current server serves, so the frame and the server move together, and a frame left open across an upgrade reconnects and is told to reload if protocol changed. A loader, though, can be much older: pages pin widget@x.y.z.js (see Released loaders), and a pinned loader opens the current server's frame. Keep changes to the loader-to-frame messages in port.ts backwards compatible.
Text engines
A text engine is one provider's chat model. It implements TextEngine from apps/server/src/engines/types.ts:
ts
interface TextEngine {
readonly provider: ProviderId;
step(req: TextRequest, ev: StepEvents, signal: AbortSignal): Promise<StepResult>;
}step() makes one streamed model call. The agent loop around it, with tool calls, results and the round limit, is runTurn() in engines/turn-runner.ts, the same for every provider. The rules for a step():
- Input.
reqhas the model, the system prompt, the tools (ToolSpec: name, description, JSON Schema), the history asCanonicalMessage[],effort,maxOutputTokens,temperatureandimages, a loader that gives an attachment's bytes.toTurns()turns the history into the alternating user and assistant turns every provider wants;annotate()writes notes as<context>blocks and a team member's messages as theirs. - Streaming. Report text as it arrives with
ev.text(delta), and the start of each tool call withev.toolStart(), so the widget can show the tool's status. - Output. Return the assistant's
parts(text andtool_callparts), themodelactually used (it can differ after a fallback),stop(end,tool_use,max_tokensorrefusal),usage, andnative: the provider's own blocks for this turn.nativeis stored with the message and replayed verbatim the next time the history goes to the same provider, which keeps reasoning signatures and similar provider data valid. Replay it only whenm.native.provideris yours. - Errors. Throw an
EngineError(message, fatal, code)with a message a visitor may see.fatalmeans retrying won't help (a bad key, no credit), and the session drops the engine so the next reply opens a new one. Turn HTTP and SDK errors into aProviderErrorwithclassify()ortoProviderError()(providers/http.ts,providers/adapters.ts). Ifsignalis aborted, let the error through.
The three text engines are providers/anthropic/text-engine.ts (the Messages API), providers/openai/text-engine.ts (the Responses API) and providers/gemini/text-engine.ts (streamGenerateContent). createTextEngine() in engines/factory.ts picks one for a connection's provider.
Voice engines
A voice engine is a live, speech-to-speech conversation. It implements VoiceEngine from apps/server/src/engines/voice.ts (not types.ts), and VoiceSession (gateway/voice-session.ts) drives it:
ts
interface VoiceEngine {
readonly kind: VoiceEngineKind; // 'openai_realtime' | 'gemini_live' | 'cascade' | 'elevenlabs_agent'
readonly inputRate: 16000 | 24000;
readonly caps: { images: boolean; video: boolean; truncate: boolean };
on<K extends keyof VoiceEvents>(event: K, fn: VoiceEvents[K]): () => void;
start(init: VoiceInit): Promise<void>;
pushAudio(pcm: Buffer): void;
pushImage(img: { data: Buffer; mime: string }, why: 'continuous' | 'turn' | 'tool'): void;
sendText(text: string): void;
addNote(text: string): void;
cancelResponse(): void;
truncate(responseId: string, heardMs: number): void;
submitToolResult(callId: string, name: string, result: { isError: boolean; text: string }, silent: boolean): void;
ptt(state: 'down' | 'up'): void;
close(reason: string): void;
}- Audio. In: PCM16 mono at
inputRate. Out: always PCM16 mono at 24 kHz, whatever the provider sends. start(init)gets the instructions, tools, history, model, voice, turn detection settings, push-to-talk, language and whether to greet first. Resolve once the session is configured, and emitready.- Events (
VoiceEvents):speechStarted,speechStopped,userTranscript(partial, then final),responseStarted,audio,responseText,responseEnded(withcompleted,interruptedorfailed, and usage),toolCall,reconnecting,errorandclosed. The session turns them into protocol messages, stores the messages, and runs tool calls with the conversation'sToolRunner. capstells the session what the engine can take: camera pictures (images), a live video feed (video), and whether it can cut a reply at the point the visitor stopped hearing it (truncate). An engine that can't see gets tool images described in words.- Helpers in
engines/voice.ts:Emitterfor the events,connectWs()(opens a provider WebSocket, and turns an HTTP refusal into aProviderError),engineError()andvoiceInstructions().
The cascade (engines/cascade-engine.ts) implements the same interface without a speech-to-speech model: it connects a speech-to-text stream (SttStream in providers/speech.ts), answers each final transcript with the bot's text brain through ConversationSession.reply('voice', ...), and speaks the reply a sentence at a time with SpokenReply. See Voice for how a voice session runs.
The history rules
Conversations and messages are stored by ConversationsService (apps/server/src/services/conversations.service.ts). What a model sees is built from them on every reply.
- Append-only.
append()gives each message the conversation's nextseq, in a transaction, and nothing is reordered or deleted while a conversation lives (whole conversations are deleted, by the team, retention or an erase request).updateMessage()changes only a few fields (text,parts,native,status,heardTextand feedback), to fill in a voice message's transcript, mark a status or record feedback. - The system prompt is fixed. It is built once, when the conversation starts, and stored on the conversation (
conversations.system_prompt). Anything learned later (a new page, page context, a sign-in, voice on or off, knowledge for the next reply) is appended as anotemessage that only the model sees. That keeps provider caches warm and replayednativeblocks valid. A conversation also keeps the bot version it started with. - Visibility.
publicmessages are seen by the visitor and the model,internalones by the team only, andmodelones (notes and tool results) by the model only.history()leaves outinternalandsystem_eventmessages and failed assistant messages, and an interrupted voice reply becomes the part the visitor heard (heard_text), without itsnativeblocks. - Compaction. When the history outgrows
brain.compactAtTokens(a rough estimate: characters / 4, plus 1,600 per image), the session asks the brain for a summary and appends it as a note withengine: 'compaction'andmetrics.compactedFromSeq. From then onhistory()starts with the latest summary, followed by the messages from thatseq. Nothing is deleted. Helpers are inagent/compaction.ts; what visitors and operators see is in Long conversations. - Turn order.
inTurnOrder()puts the rows in the order a model needs. Providers want each tool call followed directly by its results, and a reply's steps together; but a message that arrives while a reply is being written is stored when it arrives, which can be between a tool call and its result, or before the reply's own rows (they are stored as each step completes). Such a message moves to just after the reply it arrived during, and the earlier history stays as it was sent. - Messages sent mid-reply. A message sent while a reply is being written (once that reply has read the history) is stored with the reply's id in
messages.during_response_id.inTurnOrder()always puts it after that reply, so the model sees it as the latest, unanswered message, and the session's queue answers it with the next reply. The tests are inapps/server/test/text.test.ts(inTurnOrder) andtext-chat.test.ts(messages sent mid-reply).
Adding a provider or an engine
The four providers are wired through the same places. Follow one of them (OpenAI touches every kind) and change these, in this order:
- The catalogue.
packages/shared/src/providers.ts: the id inProviderIdandPROVIDER_IDS, and its entry inPROVIDERS(name, the kinds of model it offers, where to get a key, a key prefix if it has one, a one-line blurb). The REST API's provider enum and the admin panel's connect dialog and pickers read these. The type checker then points at everyRecord<ProviderId, ...>that needs the new id, such as the admin panel's monograms inapps/admin/src/pages/providers/provider-parts.tsx. - Addresses and keys.
apps/server/src/config/env.ts:PROVIDER_BASE_URL_<NAME>(andPROVIDER_WS_URL_<NAME>for WebSockets), so tests and proxies can point it elsewhere, and<NAME>_API_KEYfor the first-run import, with notes for each indocs/scripts/notes/env.ts. Add the URLs toProviderUrlsandproviderUrls()inproviders/adapters.ts, and the key toimportFromEnv()inproviders/providers.service.ts. - Connection validation and catalogues. An adapter in
providers/adapters.ts, registered increateAdapters():validate(key)checks the key and returns what it can be used for (throw aProviderError, viaproviderJson()orclassify(), so a bad key, no credit and an outage are told apart);models(key)sorts each model into aModelKind(text,realtime,stt,tts,embedding) and marks one recommended per kind (markRecommended());voices(key)lists voices.ProvidersServicecaches both catalogues for a day per key. - A text engine.
providers/<name>/text-engine.tsimplementingTextEngine(see above), acaseincreateTextEngine()(engines/factory.ts), and the provider inBotsService.withDefaults()(the brains a new bot is given, in order) andcanSee()(services/bots.service.ts). - A voice engine. A class implementing
VoiceEngine. Its kind goes inVoiceEngineKind, which is declared in bothengines/voice.tsandservices/bots.service.ts. Then:BotsService.voiceEngine()(which engine a config uses), the publish checks inBotsService.check()(which connections eachvoice.modeaccepts), the construction inVoiceSession.begin()and the usage provider inVoiceSession.ended(). A newvoice.modealso goes inpackages/shared/src/bot-config.tsand the admin panel's voice tab. - Speech for the cascade. Text to speech is a branch in
speak()inproviders/speech.ts, with the provider in theTtsSettingsandTtsRequestunions (engines/spoken-reply.ts,providers/speech.ts) and inBotsService.replyVoice(). Speech to text is anSttStreamclass inproviders/speech.ts, chosen inCascadeEngine.start(). - Voices. If the provider has no endpoint that lists voices, keep a fixed list in
providers/voices.ts, as OpenAI and Gemini do. Add a branch toProvidersService.voicePreview()so the admin panel can play samples. - Prices. Approximate list prices in
PRICESandSPEECH_PRICES(services/usage.service.ts). A model with no price costs 0, so daily spend caps can't see it. - The mock provider. Routes in
packages/mock-providers/src/index.ts: the models list, the key check, and the streaming endpoint in the provider's own event format, answering withmockReply()from a turn extracted from the request (seeanthropicTurn()and the others). WebSocket APIs go insrc/realtime.ts. Add the URLs toMockProviders.urlsandmockEnv(), and to thePROVIDER_*variablese2e/global-setup.tssets by hand. - Tests. Extend
publishedBot()'sprovideroption and default models (apps/server/test/helpers.ts), then cover the key check (good,badandbrokekeys), the catalogues, a streamed reply with a tool call (see "works with OpenAI and Gemini brains too" intext-chat.test.ts), and voice invoice.test.tsif it has a voice engine.
Each provider also has a page under Providers in these docs.
Tools
The agent's tools are in apps/server/src/agent/tools. ToolsService (tools.service.ts) gives each conversation a ToolRunner over the tools its bot allows, and the same runner serves text and voice. There are four kinds:
| Kind | Where | Notes |
|---|---|---|
| Built-in | builtinSpecs() and builtin() in tools.service.ts | search_knowledge, capture_lead, handoff_to_human, end_conversation and look_at_camera, each switched by tools.builtins and its feature. In text, face expressions are [[mood:...]] tags; voice engines get the silent express tool instead (EXPRESS_TOOL in gateway/voice-session.ts). |
| HTTP | http-tool.ts | A request the admin describes, with placeholders filled from the arguments, stored secrets and the conversation. Sent through lib/ssrf.ts, which refuses private addresses unless they are allowed. |
| MCP | mcp-pool.ts | One client per server, connected on first use and kept: Streamable HTTP, then the older SSE transport, and stdio only with ALLOW_STDIO_MCP. A server's tools are cached on its row and offered as mcp_<slug>_<tool>. |
| Client (page) | client() in tools.service.ts, Gateway.callClientTool() | Tools the page registered with the JavaScript API. The server sends tool.call to the tab that has the tool and waits for tool.result or the timeout. |
Every call goes through the runner's run(), which:
- checks the arguments against the tool's JSON Schema (Ajv), and gives the model the problem if they don't fit;
- asks the visitor first when the tool is marked so (
Gateway.confirmTool(), which showstool.confirmin every tab; no answer in two minutes counts as no, and nobody there to ask means it isn't done); - stops the tool after
tools.toolTimeoutMs; - wraps results from HTTP, MCP and page tools in
<untrusted>(untrusted()inagent/prompt.ts), cut totools.maxResultChars, so the model treats them as information, not instructions; - records the call in the
tool_callstable.
It never throws: a failure comes back as an isError result, which the model reads. With knowledge in auto mode, text replies get the best matches as a note before each reply instead of the search_knowledge tool, while voice engines, which can't be given knowledge up front, still get the tool.
To add a built-in tool: its switch in tools.builtins (packages/shared/src/bot-config.ts, with a note in docs/scripts/notes/config.ts), its name in BUILTIN, its spec in builtinSpecs(), its case in builtin(), and its status label in toolLabel() (gateway/session.ts). What operators see of tools is in Tools and MCP and Client tools.