Skip to content

Protocol and engines ​

The widget protocol ​

The chat frame and the server talk over one WebSocket, /v1/widget/ws?bot=pk_.... Every message, its fields and the binary frame layout are in the Widget protocol reference, which is generated from the code. In brief:

  • Control messages are JSON objects with a t field. The frame's first message is hello, with protocol: 1 (PROTOCOL_VERSION); a widget on another version gets protocol_unsupported and is asked to reload.
  • Stored messages carry a per-conversation seq. A widget that reconnects sends its lastSeq and gets what it missed in history.sync.
  • A reply streams as agent.response.start, any number of agent.text.delta, then agent.response.end, all with the same responseId.
  • Audio and camera pictures are binary frames: an 8-byte little-endian header (kind, codec, stream, sequence number), then PCM16 or an encoded image.

The definitions live in packages/protocol/src:

FileHolds
types.tsClientMessage and ServerMessage, the shared types, PROTOCOL_VERSION
schemas.tsThe zod schemas the server checks every client message with (parseClientMessage()); server only
binary.tsFrameKind, Codec, encodeFrame() and decodeFrame()
errors.tsError codes and WebSocket close codes
port.tsThe messages between the loader and the frame, over their MessagePort (PORT_PROTOCOL)

To add a client message: its type in types.ts, its schema in schemas.ts, a case in WidgetConnection.dispatch() (gateway/gateway.ts) or a handler that a module registers in conn.handlers, the frame's side in packages/widget/src/frame/chat.ts, and a note in docs/scripts/notes/protocol.ts. A server message needs the type, the frame's handling and the note. The docs generator fails until each is there (see Testing).

Two kinds of compatibility matter. The frame is always the one the current server serves, so the frame and the server move together, and a frame left open across an upgrade reconnects and is told to reload if protocol changed. A loader, though, can be much older: pages pin widget@x.y.z.js (see Released loaders), and a pinned loader opens the current server's frame. Keep changes to the loader-to-frame messages in port.ts backwards compatible.

Text engines ​

A text engine is one provider's chat model. It implements TextEngine from apps/server/src/engines/types.ts:

ts
interface TextEngine {
  readonly provider: ProviderId;
  step(req: TextRequest, ev: StepEvents, signal: AbortSignal): Promise<StepResult>;
}

step() makes one streamed model call. The agent loop around it, with tool calls, results and the round limit, is runTurn() in engines/turn-runner.ts, the same for every provider. The rules for a step():

  • Input. req has the model, the system prompt, the tools (ToolSpec: name, description, JSON Schema), the history as CanonicalMessage[], effort, maxOutputTokens, temperature and images, a loader that gives an attachment's bytes. toTurns() turns the history into the alternating user and assistant turns every provider wants; annotate() writes notes as <context> blocks and a team member's messages as theirs.
  • Streaming. Report text as it arrives with ev.text(delta), and the start of each tool call with ev.toolStart(), so the widget can show the tool's status.
  • Output. Return the assistant's parts (text and tool_call parts), the model actually used (it can differ after a fallback), stop (end, tool_use, max_tokens or refusal), usage, and native: the provider's own blocks for this turn. native is stored with the message and replayed verbatim the next time the history goes to the same provider, which keeps reasoning signatures and similar provider data valid. Replay it only when m.native.provider is yours.
  • Errors. Throw an EngineError(message, fatal, code) with a message a visitor may see. fatal means retrying won't help (a bad key, no credit), and the session drops the engine so the next reply opens a new one. Turn HTTP and SDK errors into a ProviderError with classify() or toProviderError() (providers/http.ts, providers/adapters.ts). If signal is aborted, let the error through.

The three text engines are providers/anthropic/text-engine.ts (the Messages API), providers/openai/text-engine.ts (the Responses API) and providers/gemini/text-engine.ts (streamGenerateContent). createTextEngine() in engines/factory.ts picks one for a connection's provider.

Voice engines ​

A voice engine is a live, speech-to-speech conversation. It implements VoiceEngine from apps/server/src/engines/voice.ts (not types.ts), and VoiceSession (gateway/voice-session.ts) drives it:

ts
interface VoiceEngine {
  readonly kind: VoiceEngineKind; // 'openai_realtime' | 'gemini_live' | 'cascade' | 'elevenlabs_agent'
  readonly inputRate: 16000 | 24000;
  readonly caps: { images: boolean; video: boolean; truncate: boolean };
  on<K extends keyof VoiceEvents>(event: K, fn: VoiceEvents[K]): () => void;
  start(init: VoiceInit): Promise<void>;
  pushAudio(pcm: Buffer): void;
  pushImage(img: { data: Buffer; mime: string }, why: 'continuous' | 'turn' | 'tool'): void;
  sendText(text: string): void;
  addNote(text: string): void;
  cancelResponse(): void;
  truncate(responseId: string, heardMs: number): void;
  submitToolResult(callId: string, name: string, result: { isError: boolean; text: string }, silent: boolean): void;
  ptt(state: 'down' | 'up'): void;
  close(reason: string): void;
}
  • Audio. In: PCM16 mono at inputRate. Out: always PCM16 mono at 24 kHz, whatever the provider sends.
  • start(init) gets the instructions, tools, history, model, voice, turn detection settings, push-to-talk, language and whether to greet first. Resolve once the session is configured, and emit ready.
  • Events (VoiceEvents): speechStarted, speechStopped, userTranscript (partial, then final), responseStarted, audio, responseText, responseEnded (with completed, interrupted or failed, and usage), toolCall, reconnecting, error and closed. The session turns them into protocol messages, stores the messages, and runs tool calls with the conversation's ToolRunner.
  • caps tells the session what the engine can take: camera pictures (images), a live video feed (video), and whether it can cut a reply at the point the visitor stopped hearing it (truncate). An engine that can't see gets tool images described in words.
  • Helpers in engines/voice.ts: Emitter for the events, connectWs() (opens a provider WebSocket, and turns an HTTP refusal into a ProviderError), engineError() and voiceInstructions().

The cascade (engines/cascade-engine.ts) implements the same interface without a speech-to-speech model: it connects a speech-to-text stream (SttStream in providers/speech.ts), answers each final transcript with the bot's text brain through ConversationSession.reply('voice', ...), and speaks the reply a sentence at a time with SpokenReply. See Voice for how a voice session runs.

The history rules ​

Conversations and messages are stored by ConversationsService (apps/server/src/services/conversations.service.ts). What a model sees is built from them on every reply.

  • Append-only. append() gives each message the conversation's next seq, in a transaction, and nothing is reordered or deleted while a conversation lives (whole conversations are deleted, by the team, retention or an erase request). updateMessage() changes only a few fields (text, parts, native, status, heardText and feedback), to fill in a voice message's transcript, mark a status or record feedback.
  • The system prompt is fixed. It is built once, when the conversation starts, and stored on the conversation (conversations.system_prompt). Anything learned later (a new page, page context, a sign-in, voice on or off, knowledge for the next reply) is appended as a note message that only the model sees. That keeps provider caches warm and replayed native blocks valid. A conversation also keeps the bot version it started with.
  • Visibility. public messages are seen by the visitor and the model, internal ones by the team only, and model ones (notes and tool results) by the model only. history() leaves out internal and system_event messages and failed assistant messages, and an interrupted voice reply becomes the part the visitor heard (heard_text), without its native blocks.
  • Compaction. When the history outgrows brain.compactAtTokens (a rough estimate: characters / 4, plus 1,600 per image), the session asks the brain for a summary and appends it as a note with engine: 'compaction' and metrics.compactedFromSeq. From then on history() starts with the latest summary, followed by the messages from that seq. Nothing is deleted. Helpers are in agent/compaction.ts; what visitors and operators see is in Long conversations.
  • Turn order. inTurnOrder() puts the rows in the order a model needs. Providers want each tool call followed directly by its results, and a reply's steps together; but a message that arrives while a reply is being written is stored when it arrives, which can be between a tool call and its result, or before the reply's own rows (they are stored as each step completes). Such a message moves to just after the reply it arrived during, and the earlier history stays as it was sent.
  • Messages sent mid-reply. A message sent while a reply is being written (once that reply has read the history) is stored with the reply's id in messages.during_response_id. inTurnOrder() always puts it after that reply, so the model sees it as the latest, unanswered message, and the session's queue answers it with the next reply. The tests are in apps/server/test/text.test.ts (inTurnOrder) and text-chat.test.ts (messages sent mid-reply).

Adding a provider or an engine ​

The four providers are wired through the same places. Follow one of them (OpenAI touches every kind) and change these, in this order:

  1. The catalogue. packages/shared/src/providers.ts: the id in ProviderId and PROVIDER_IDS, and its entry in PROVIDERS (name, the kinds of model it offers, where to get a key, a key prefix if it has one, a one-line blurb). The REST API's provider enum and the admin panel's connect dialog and pickers read these. The type checker then points at every Record<ProviderId, ...> that needs the new id, such as the admin panel's monograms in apps/admin/src/pages/providers/provider-parts.tsx.
  2. Addresses and keys. apps/server/src/config/env.ts: PROVIDER_BASE_URL_<NAME> (and PROVIDER_WS_URL_<NAME> for WebSockets), so tests and proxies can point it elsewhere, and <NAME>_API_KEY for the first-run import, with notes for each in docs/scripts/notes/env.ts. Add the URLs to ProviderUrls and providerUrls() in providers/adapters.ts, and the key to importFromEnv() in providers/providers.service.ts.
  3. Connection validation and catalogues. An adapter in providers/adapters.ts, registered in createAdapters(): validate(key) checks the key and returns what it can be used for (throw a ProviderError, via providerJson() or classify(), so a bad key, no credit and an outage are told apart); models(key) sorts each model into a ModelKind (text, realtime, stt, tts, embedding) and marks one recommended per kind (markRecommended()); voices(key) lists voices. ProvidersService caches both catalogues for a day per key.
  4. A text engine. providers/<name>/text-engine.ts implementing TextEngine (see above), a case in createTextEngine() (engines/factory.ts), and the provider in BotsService.withDefaults() (the brains a new bot is given, in order) and canSee() (services/bots.service.ts).
  5. A voice engine. A class implementing VoiceEngine. Its kind goes in VoiceEngineKind, which is declared in both engines/voice.ts and services/bots.service.ts. Then: BotsService.voiceEngine() (which engine a config uses), the publish checks in BotsService.check() (which connections each voice.mode accepts), the construction in VoiceSession.begin() and the usage provider in VoiceSession.ended(). A new voice.mode also goes in packages/shared/src/bot-config.ts and the admin panel's voice tab.
  6. Speech for the cascade. Text to speech is a branch in speak() in providers/speech.ts, with the provider in the TtsSettings and TtsRequest unions (engines/spoken-reply.ts, providers/speech.ts) and in BotsService.replyVoice(). Speech to text is an SttStream class in providers/speech.ts, chosen in CascadeEngine.start().
  7. Voices. If the provider has no endpoint that lists voices, keep a fixed list in providers/voices.ts, as OpenAI and Gemini do. Add a branch to ProvidersService.voicePreview() so the admin panel can play samples.
  8. Prices. Approximate list prices in PRICES and SPEECH_PRICES (services/usage.service.ts). A model with no price costs 0, so daily spend caps can't see it.
  9. The mock provider. Routes in packages/mock-providers/src/index.ts: the models list, the key check, and the streaming endpoint in the provider's own event format, answering with mockReply() from a turn extracted from the request (see anthropicTurn() and the others). WebSocket APIs go in src/realtime.ts. Add the URLs to MockProviders.urls and mockEnv(), and to the PROVIDER_* variables e2e/global-setup.ts sets by hand.
  10. Tests. Extend publishedBot()'s provider option and default models (apps/server/test/helpers.ts), then cover the key check (good, bad and broke keys), the catalogues, a streamed reply with a tool call (see "works with OpenAI and Gemini brains too" in text-chat.test.ts), and voice in voice.test.ts if it has a voice engine.

Each provider also has a page under Providers in these docs.

Tools ​

The agent's tools are in apps/server/src/agent/tools. ToolsService (tools.service.ts) gives each conversation a ToolRunner over the tools its bot allows, and the same runner serves text and voice. There are four kinds:

KindWhereNotes
Built-inbuiltinSpecs() and builtin() in tools.service.tssearch_knowledge, capture_lead, handoff_to_human, end_conversation and look_at_camera, each switched by tools.builtins and its feature. In text, face expressions are [[mood:...]] tags; voice engines get the silent express tool instead (EXPRESS_TOOL in gateway/voice-session.ts).
HTTPhttp-tool.tsA request the admin describes, with placeholders filled from the arguments, stored secrets and the conversation. Sent through lib/ssrf.ts, which refuses private addresses unless they are allowed.
MCPmcp-pool.tsOne client per server, connected on first use and kept: Streamable HTTP, then the older SSE transport, and stdio only with ALLOW_STDIO_MCP. A server's tools are cached on its row and offered as mcp_<slug>_<tool>.
Client (page)client() in tools.service.ts, Gateway.callClientTool()Tools the page registered with the JavaScript API. The server sends tool.call to the tab that has the tool and waits for tool.result or the timeout.

Every call goes through the runner's run(), which:

  • checks the arguments against the tool's JSON Schema (Ajv), and gives the model the problem if they don't fit;
  • asks the visitor first when the tool is marked so (Gateway.confirmTool(), which shows tool.confirm in every tab; no answer in two minutes counts as no, and nobody there to ask means it isn't done);
  • stops the tool after tools.toolTimeoutMs;
  • wraps results from HTTP, MCP and page tools in <untrusted> (untrusted() in agent/prompt.ts), cut to tools.maxResultChars, so the model treats them as information, not instructions;
  • records the call in the tool_calls table.

It never throws: a failure comes back as an isError result, which the model reads. With knowledge in auto mode, text replies get the best matches as a note before each reply instead of the search_knowledge tool, while voice engines, which can't be given knowledge up front, still get the tool.

To add a built-in tool: its switch in tools.builtins (packages/shared/src/bot-config.ts, with a note in docs/scripts/notes/config.ts), its name in BUILTIN, its spec in builtinSpecs(), its case in builtin(), and its status label in toolLabel() (gateway/session.ts). What operators see of tools is in Tools and MCP and Client tools.

Wireface Chat 0.1.0. These docs are served by your own server.