Skip to content

Architecture ​

These pages are for people working on Wireface Chat itself. To add the chat to a website instead, start with the Introduction and the embedding pages.

The big picture ​

text
Host page (any website)
  widget.js               the loader (packages/widget/src/loader): launcher in a shadow root, the JavaScript API,
                          page tools
     |  MessageChannel port (packages/protocol/src/port.ts)
     v
  iframe /frame/pk_...    the chat frame app (packages/widget/src/frame, Preact): messages, the face, microphone,
                          camera
     |
     |  WebSocket /v1/widget/ws   JSON control messages, plus binary audio and camera frames
     v
Fastify server (apps/server)      one Node.js process
  gateway/          a WidgetConnection per socket, a ConversationSession per live conversation, a VoiceSession
  engines/          runTurn() (the agent loop), the TextEngine and VoiceEngine interfaces, the cascade voice
  providers/        Anthropic, OpenAI, Gemini, ElevenLabs: key checks, catalogues, text and voice engines, speech
  agent/            the system prompt, compaction, tools (built-in, HTTP, MCP, page tools)
  knowledge/        knowledge sources, chunking, SQLite FTS5 search
  services/         bots, conversations, visitors, handoff, usage, vision, webhooks, analytics, workspace
  http/routes/      the REST API at /api/v1, and /api/v1/live (the admin panel's WebSocket)
  db/, storage/     SQLite through Drizzle (DATA_DIR/wireface.db), files (DATA_DIR/files)
     |                                    |
     |  HTTPS, SSE, WebSockets            |  also serves
     v                                    v
  AI providers, MCP servers,            /admin/          the admin SPA (apps/admin, React)
  HTTP tools                            /docs/           these docs (docs, VitePress)
                                        /core/<hash>/    the face engine (vendor/wireface-core)
                                        /widget.js, /frame/assets/   the widget build (packages/widget/dist)

The browser never talks to a provider. It speaks one protocol to the chat server whichever provider is behind it, and provider keys never leave the server. The full list of what the server serves is in the Introduction.

The server starts in apps/server/src/main.ts: it reads the environment (config/env.ts), builds the base context (context.ts: env, database, sealer, blob store, event bus, logger, rate limits), creates the services (services/index.ts), re-encrypts secrets if the master key is being rotated (db/rekey.ts) and calls buildApp() in app.ts. buildApp(ctx) registers the Fastify plugins, the REST routes, the widget routes and the static files. The tests call it too, with their own context.

Live state is kept in memory: open conversations and voice sessions, pending tool calls and confirmations, rate limits, and the event bus (events/bus.ts) that feeds the admin live view, webhooks and analytics. That is why each database is served by one process.

Packages ​

PathPackageWhat it does
apps/server@wireface/serverThe Fastify server: widget gateway, voice relay, provider engines, tools, knowledge, the admin REST API and the static files. Built by tsdown into one ESM file, dist/main.js, with the workspace packages inlined.
apps/admin@wireface/adminThe admin panel: a React SPA (TanStack Router and Query, Tailwind, Radix) built by Vite into apps/admin/dist, which the server serves at /admin/. It uses the REST API and the /api/v1/live socket.
packages/protocol@wireface/protocolThe wire types. types.ts: the widget WebSocket messages. schemas.ts: zod schemas for what widgets send (only the server imports it, so zod stays out of the widget). binary.ts: audio and camera frames. port.ts: loader to frame messages. errors.ts: error and close codes.
packages/shared@wireface/sharedThe bot configuration schema (bot-config.ts; parseBotConfig({}) is a complete bot), the public config a widget gets (public-config.ts), the provider catalogue (providers.ts), bot templates and prefixed ids.
packages/widget@wireface/widgetwidget.js (src/loader, a dependency-free IIFE) and the chat frame (src/frame, Preact with signals). releases/ keeps every released loader.
packages/chat@wireface/chatThe npm package: a typed ESM loader for widget.js, and React components and hooks. The only package published to npm.
packages/face-kit@wireface/face-kitThe face for the frame and the admin panel: loads Wireface Core at runtime from /core/<hash>/, applies face settings, mimes lip sync for text replies, and draws a 2D wireframe head (wire-head.ts) when there is no WebGL2.
packages/markdown@wireface/markdownSafe Markdown for chat bubbles: marked's lexer, rendered through any h() (Preact's or React's), never as HTML.
packages/snippets@wireface/snippetsEmbed code in every documented form (script tag, iframe, npm, React, Next.js, Vue, GTM, identity hashes), shared by the admin panel, the docs and the tests. Pure string building.
packages/brand@wireface/brandBrand tokens (tokens.css) and logos from wireface.dev. No code.
packages/mock-providers@wireface/mock-providersFake Anthropic, OpenAI, Gemini and ElevenLabs APIs (HTTP, SSE and realtime WebSockets) for tests and offline development. See Testing.
vendor/wireface-core@wireface/coreThe Wireface face engine (plain ES modules on WebGL2), copied from ../core by pnpm sync-core. The server serves it at /core/<contentHash>/ and imports its config validation and expression names. Never edited here: see Wireface Core.
docs@wireface/docsThese docs (VitePress). Part of the reference is generated from the code: see Testing.
e2e@wireface/e2ePlaywright tests against a real server, the mock providers and a fixture shop site.
examples(not a package)Runnable embedding examples, served by examples/serve.mjs.

The private packages are used as TypeScript source: their exports point at src/*.ts, so there is nothing to build between them. tsx runs them in the server during development, Vite bundles them into the widget and the admin panel, and tsdown inlines them into the server build and the npm package.

A typed message ​

What happens between the visitor pressing Enter and the reply appearing. Paths are under apps/server/src unless they say otherwise.

  1. The frame. packages/widget/src/frame/chat.ts sends chat.send, with a clientMsgId of its own, on the socket it opened to /v1/widget/ws?bot=pk_....
  2. The socket. http/routes/widget.ts accepts the upgrade only from the server's own origin (the frame's), and outside production also from localhost pages, then hands it to Gateway.accept(). The first message must be hello, within 5 seconds: WidgetConnection.hello() in gateway/gateway.ts resolves the published bot (or a signed preview of the draft), checks the page against the bot's allowed origins, checks identity, finds a conversation to resume and answers session.ready. Every text frame is validated with parseClientMessage() from @wireface/protocol/schemas.
  3. The gateway. WidgetConnection.chatSend() applies the rate limits and, on the visitor's first message, creates the conversation (ensureSession()): the system prompt is built by agent/prompt.ts and stored with the conversation, conversation.started goes to the widget, and the welcome message is stored. Then it calls ConversationSession.userText().
  4. The session. gateway/session.ts stores the visitor's message (claiming uploaded images, adding the camera frame in per-turn mode) and broadcasts message.created to every tab of the conversation. It then queues a reply, unless a person from the team has the chat, the bot is outside its hours, or a voice session is live (the voice engine answers instead). Replies run one at a time; messages that arrive meanwhile share the next one.
  5. The reply. ConversationSession.reply() sends agent.response.start and agent.status: thinking, opens the text engine for the bot's brain connection (engines/factory.ts, with the key decrypted by ProvidersService.key()), runs knowledge retrieval (beforeReply, which adds a context note), compacts the history if it is too long, and reads it with ConversationsService.history().
  6. The agent loop. runTurn() in engines/turn-runner.ts calls the engine's step(). If the step asks for tools, they run through the conversation's ToolRunner (agent/tools/tools.service.ts), their results are stored and sent back, and the loop goes round again, up to brain.maxToolRounds. Each finished step is stored as a message as soon as it completes, with the provider's own blocks kept in native.
  7. The provider. The engine (providers/<provider>/text-engine.ts) turns the provider-neutral history into the provider's format, makes one streamed call and reports each piece of text with ev.text(delta).
  8. Back to the widget. Each delta passes through TagStripper (lib/text.ts), which turns [[mood:happy]] tags into face.mood and face.express messages, then goes to every tab as agent.text.delta and to the event bus as message.delta (the admin live view). At the end the usage is recorded and the widget gets agent.response.end (with the stored message's id and seq), agent.citations when knowledge was used, and agent.status: idle.

The admin panel sees the same conversation through the event bus: message.created, message.delta and agent.status are forwarded to its /api/v1/live socket, and the webhooks service turns events into deliveries.

Voice ​

Voice uses the same conversation, history and tools as text. What it adds is a VoiceSession (gateway/voice-session.ts) between one widget and a VoiceEngine (engines/voice.ts):

  1. voice.start reaches VoiceSession.start(), which picks the engine with BotsService.voiceEngine() (from voice.mode and the connection's provider) and checks the bot's hours, its daily cap and security.caps.concurrentVoiceSessions.
  2. begin() builds the instructions (voiceInstructions() around the same system prompt) and the tools (the silent express tool, then every tool the bot has, knowledge search included), and starts the engine with the conversation's history. The widget gets voice.ready with the input rate (16 or 24 kHz) and the output format (24 kHz PCM16).
  3. The frame streams the microphone as binary MIC frames, which go to the engine's pushAudio(). Turn detection happens at the provider, or in the cascade's speech to text.
  4. Engine events become protocol messages: speech.started and speech.stopped, user.transcript, agent.audio.start, binary AGENT_AUDIO frames of 100 ms, agent.text.delta and agent.audio.end. When the visitor stops speaking, an empty user message is stored straight away and filled in when the transcript arrives, so it stays ahead of the reply.
  5. Barge-in. Speech while the agent is playing sends agent.audio.clear and calls engine.truncate(responseId, heardMs), with heardMs from the widget's playback.mark messages. The reply is stored as interrupted, with the part that was heard in heard_text.
EngineCodeHow it answers
openai_realtimeproviders/openai/realtime-engine.tsSpeech to speech at the provider (24 kHz in). Tool calls come back as toolCall events; the session runs them with the same ToolRunner and answers with submitToolResult().
gemini_liveproviders/gemini/live-engine.tsThe same, at 16 kHz, and it takes a live camera view.
elevenlabs_agentproviders/elevenlabs/agent-engine.tsAn ElevenLabs agent that ensureAgent() creates and keeps in sync; every tool is a client tool run by this server.
cascadeengines/cascade-engine.tsSpeech to text (OpenAI or ElevenLabs Scribe, providers/speech.ts), then the text path above, then text to speech a sentence at a time (SpokenReply, engines/spoken-reply.ts).

The realtime engines run the model at the provider, and VoiceSession.ended() stores their replies. The cascade calls ConversationSession.reply('voice', speak) instead: the bot's own text brain answers through runTurn(), with its usual tools and knowledge, and stores and streams the reply exactly as for text, while the voice session only relays the audio. Text typed during a voice session goes to the engine's sendText() and is answered aloud. Typed replies read aloud (appearance.face.textReplies: 'speak') reuse SpokenReply on the text path, with no voice session. For the visitor's side of all this, see Voice.

Where data lives ​

WhatWhereCode
The databaseDATA_DIR/wireface.db: SQLite through Drizzle (drizzle-orm/better-sqlite3), in WAL mode with foreign keys ondb/schema.ts, db/client.ts
Migrationsapps/server/drizzle/*.sql with drizzle/meta, applied by openDb() on every startsee Database migrations
FilesDATA_DIR/files, behind the BlobStore interface: keys like ws/<workspace id>/att/<id>.jpg (images), ws/<workspace id>/faces/... (custom faces), ws/<workspace id>/kb/<id> (knowledge files) and cache/previews/... (voice samples)storage/blob-store.ts
SecretsProvider keys, tool secrets, webhook and identity secrets and MCP headers, sealed with AES-256-GCM in their database columnslib/crypto.ts (Sealer), db/rekey.ts
The master keyWIREFACE_MASTER_KEY, or DATA_DIR/master.key, generated on first runcontext.ts
Live stateMemory only: sessions, voice sessions, rate limits, the event busgateway/, lib/rate-limit.ts, events/bus.ts

A few conventions hold across the schema: ids are prefixed text (cnv_..., msg_..., see packages/shared/src/ids.ts), times are epoch milliseconds, and JSON is stored as text. Every row that belongs to a customer has a workspace_id, and the services filter by it. better-sqlite3 is synchronous, so a transaction is a plain function: tx(db, () => ...) in db/client.ts. Full-text search (messages_fts for the admin's search, kb_chunks_fts for knowledge retrieval) uses FTS5 tables kept in step by triggers, from the hand-written drizzle/0001_fts.sql.

What an operator backs up, and how, is in Backups.

Wireface Chat 0.1.0. These docs are served by your own server.