Appearance
Architecture
These pages are for people working on Wireface Chat itself. To add the chat to a website instead, start with the Introduction and the embedding pages.
The big picture
text
Host page (any website)
widget.js the loader (packages/widget/src/loader): launcher in a shadow root, the JavaScript API,
page tools
| MessageChannel port (packages/protocol/src/port.ts)
v
iframe /frame/pk_... the chat frame app (packages/widget/src/frame, Preact): messages, the face, microphone,
camera
|
| WebSocket /v1/widget/ws JSON control messages, plus binary audio and camera frames
v
Fastify server (apps/server) one Node.js process
gateway/ a WidgetConnection per socket, a ConversationSession per live conversation, a VoiceSession
engines/ runTurn() (the agent loop), the TextEngine and VoiceEngine interfaces, the cascade voice
providers/ Anthropic, OpenAI, Gemini, ElevenLabs: key checks, catalogues, text and voice engines, speech
agent/ the system prompt, compaction, tools (built-in, HTTP, MCP, page tools)
knowledge/ knowledge sources, chunking, SQLite FTS5 search
services/ bots, conversations, visitors, handoff, usage, vision, webhooks, analytics, workspace
http/routes/ the REST API at /api/v1, and /api/v1/live (the admin panel's WebSocket)
db/, storage/ SQLite through Drizzle (DATA_DIR/wireface.db), files (DATA_DIR/files)
| |
| HTTPS, SSE, WebSockets | also serves
v v
AI providers, MCP servers, /admin/ the admin SPA (apps/admin, React)
HTTP tools /docs/ these docs (docs, VitePress)
/core/<hash>/ the face engine (vendor/wireface-core)
/widget.js, /frame/assets/ the widget build (packages/widget/dist)The browser never talks to a provider. It speaks one protocol to the chat server whichever provider is behind it, and provider keys never leave the server. The full list of what the server serves is in the Introduction.
The server starts in apps/server/src/main.ts: it reads the environment (config/env.ts), builds the base context (context.ts: env, database, sealer, blob store, event bus, logger, rate limits), creates the services (services/index.ts), re-encrypts secrets if the master key is being rotated (db/rekey.ts) and calls buildApp() in app.ts. buildApp(ctx) registers the Fastify plugins, the REST routes, the widget routes and the static files. The tests call it too, with their own context.
Live state is kept in memory: open conversations and voice sessions, pending tool calls and confirmations, rate limits, and the event bus (events/bus.ts) that feeds the admin live view, webhooks and analytics. That is why each database is served by one process.
Packages
| Path | Package | What it does |
|---|---|---|
apps/server | @wireface/server | The Fastify server: widget gateway, voice relay, provider engines, tools, knowledge, the admin REST API and the static files. Built by tsdown into one ESM file, dist/main.js, with the workspace packages inlined. |
apps/admin | @wireface/admin | The admin panel: a React SPA (TanStack Router and Query, Tailwind, Radix) built by Vite into apps/admin/dist, which the server serves at /admin/. It uses the REST API and the /api/v1/live socket. |
packages/protocol | @wireface/protocol | The wire types. types.ts: the widget WebSocket messages. schemas.ts: zod schemas for what widgets send (only the server imports it, so zod stays out of the widget). binary.ts: audio and camera frames. port.ts: loader to frame messages. errors.ts: error and close codes. |
packages/shared | @wireface/shared | The bot configuration schema (bot-config.ts; parseBotConfig({}) is a complete bot), the public config a widget gets (public-config.ts), the provider catalogue (providers.ts), bot templates and prefixed ids. |
packages/widget | @wireface/widget | widget.js (src/loader, a dependency-free IIFE) and the chat frame (src/frame, Preact with signals). releases/ keeps every released loader. |
packages/chat | @wireface/chat | The npm package: a typed ESM loader for widget.js, and React components and hooks. The only package published to npm. |
packages/face-kit | @wireface/face-kit | The face for the frame and the admin panel: loads Wireface Core at runtime from /core/<hash>/, applies face settings, mimes lip sync for text replies, and draws a 2D wireframe head (wire-head.ts) when there is no WebGL2. |
packages/markdown | @wireface/markdown | Safe Markdown for chat bubbles: marked's lexer, rendered through any h() (Preact's or React's), never as HTML. |
packages/snippets | @wireface/snippets | Embed code in every documented form (script tag, iframe, npm, React, Next.js, Vue, GTM, identity hashes), shared by the admin panel, the docs and the tests. Pure string building. |
packages/brand | @wireface/brand | Brand tokens (tokens.css) and logos from wireface.dev. No code. |
packages/mock-providers | @wireface/mock-providers | Fake Anthropic, OpenAI, Gemini and ElevenLabs APIs (HTTP, SSE and realtime WebSockets) for tests and offline development. See Testing. |
vendor/wireface-core | @wireface/core | The Wireface face engine (plain ES modules on WebGL2), copied from ../core by pnpm sync-core. The server serves it at /core/<contentHash>/ and imports its config validation and expression names. Never edited here: see Wireface Core. |
docs | @wireface/docs | These docs (VitePress). Part of the reference is generated from the code: see Testing. |
e2e | @wireface/e2e | Playwright tests against a real server, the mock providers and a fixture shop site. |
examples | (not a package) | Runnable embedding examples, served by examples/serve.mjs. |
The private packages are used as TypeScript source: their exports point at src/*.ts, so there is nothing to build between them. tsx runs them in the server during development, Vite bundles them into the widget and the admin panel, and tsdown inlines them into the server build and the npm package.
A typed message
What happens between the visitor pressing Enter and the reply appearing. Paths are under apps/server/src unless they say otherwise.
- The frame.
packages/widget/src/frame/chat.tssendschat.send, with aclientMsgIdof its own, on the socket it opened to/v1/widget/ws?bot=pk_.... - The socket.
http/routes/widget.tsaccepts the upgrade only from the server's own origin (the frame's), and outside production also fromlocalhostpages, then hands it toGateway.accept(). The first message must behello, within 5 seconds:WidgetConnection.hello()ingateway/gateway.tsresolves the published bot (or a signed preview of the draft), checks the page against the bot's allowed origins, checks identity, finds a conversation to resume and answerssession.ready. Every text frame is validated withparseClientMessage()from@wireface/protocol/schemas. - The gateway.
WidgetConnection.chatSend()applies the rate limits and, on the visitor's first message, creates the conversation (ensureSession()): the system prompt is built byagent/prompt.tsand stored with the conversation,conversation.startedgoes to the widget, and the welcome message is stored. Then it callsConversationSession.userText(). - The session.
gateway/session.tsstores the visitor's message (claiming uploaded images, adding the camera frame in per-turn mode) and broadcastsmessage.createdto every tab of the conversation. It then queues a reply, unless a person from the team has the chat, the bot is outside its hours, or a voice session is live (the voice engine answers instead). Replies run one at a time; messages that arrive meanwhile share the next one. - The reply.
ConversationSession.reply()sendsagent.response.startandagent.status: thinking, opens the text engine for the bot's brain connection (engines/factory.ts, with the key decrypted byProvidersService.key()), runs knowledge retrieval (beforeReply, which adds a context note), compacts the history if it is too long, and reads it withConversationsService.history(). - The agent loop.
runTurn()inengines/turn-runner.tscalls the engine'sstep(). If the step asks for tools, they run through the conversation'sToolRunner(agent/tools/tools.service.ts), their results are stored and sent back, and the loop goes round again, up tobrain.maxToolRounds. Each finished step is stored as a message as soon as it completes, with the provider's own blocks kept innative. - The provider. The engine (
providers/<provider>/text-engine.ts) turns the provider-neutral history into the provider's format, makes one streamed call and reports each piece of text withev.text(delta). - Back to the widget. Each delta passes through
TagStripper(lib/text.ts), which turns[[mood:happy]]tags intoface.moodandface.expressmessages, then goes to every tab asagent.text.deltaand to the event bus asmessage.delta(the admin live view). At the end the usage is recorded and the widget getsagent.response.end(with the stored message's id andseq),agent.citationswhen knowledge was used, andagent.status: idle.
The admin panel sees the same conversation through the event bus: message.created, message.delta and agent.status are forwarded to its /api/v1/live socket, and the webhooks service turns events into deliveries.
Voice
Voice uses the same conversation, history and tools as text. What it adds is a VoiceSession (gateway/voice-session.ts) between one widget and a VoiceEngine (engines/voice.ts):
voice.startreachesVoiceSession.start(), which picks the engine withBotsService.voiceEngine()(fromvoice.modeand the connection's provider) and checks the bot's hours, its daily cap andsecurity.caps.concurrentVoiceSessions.begin()builds the instructions (voiceInstructions()around the same system prompt) and the tools (the silentexpresstool, then every tool the bot has, knowledge search included), and starts the engine with the conversation's history. The widget getsvoice.readywith the input rate (16 or 24 kHz) and the output format (24 kHz PCM16).- The frame streams the microphone as binary
MICframes, which go to the engine'spushAudio(). Turn detection happens at the provider, or in the cascade's speech to text. - Engine events become protocol messages:
speech.startedandspeech.stopped,user.transcript,agent.audio.start, binaryAGENT_AUDIOframes of 100 ms,agent.text.deltaandagent.audio.end. When the visitor stops speaking, an empty user message is stored straight away and filled in when the transcript arrives, so it stays ahead of the reply. - Barge-in. Speech while the agent is playing sends
agent.audio.clearand callsengine.truncate(responseId, heardMs), withheardMsfrom the widget'splayback.markmessages. The reply is stored asinterrupted, with the part that was heard inheard_text.
| Engine | Code | How it answers |
|---|---|---|
openai_realtime | providers/openai/realtime-engine.ts | Speech to speech at the provider (24 kHz in). Tool calls come back as toolCall events; the session runs them with the same ToolRunner and answers with submitToolResult(). |
gemini_live | providers/gemini/live-engine.ts | The same, at 16 kHz, and it takes a live camera view. |
elevenlabs_agent | providers/elevenlabs/agent-engine.ts | An ElevenLabs agent that ensureAgent() creates and keeps in sync; every tool is a client tool run by this server. |
cascade | engines/cascade-engine.ts | Speech to text (OpenAI or ElevenLabs Scribe, providers/speech.ts), then the text path above, then text to speech a sentence at a time (SpokenReply, engines/spoken-reply.ts). |
The realtime engines run the model at the provider, and VoiceSession.ended() stores their replies. The cascade calls ConversationSession.reply('voice', speak) instead: the bot's own text brain answers through runTurn(), with its usual tools and knowledge, and stores and streams the reply exactly as for text, while the voice session only relays the audio. Text typed during a voice session goes to the engine's sendText() and is answered aloud. Typed replies read aloud (appearance.face.textReplies: 'speak') reuse SpokenReply on the text path, with no voice session. For the visitor's side of all this, see Voice.
Where data lives
| What | Where | Code |
|---|---|---|
| The database | DATA_DIR/wireface.db: SQLite through Drizzle (drizzle-orm/better-sqlite3), in WAL mode with foreign keys on | db/schema.ts, db/client.ts |
| Migrations | apps/server/drizzle/*.sql with drizzle/meta, applied by openDb() on every start | see Database migrations |
| Files | DATA_DIR/files, behind the BlobStore interface: keys like ws/<workspace id>/att/<id>.jpg (images), ws/<workspace id>/faces/... (custom faces), ws/<workspace id>/kb/<id> (knowledge files) and cache/previews/... (voice samples) | storage/blob-store.ts |
| Secrets | Provider keys, tool secrets, webhook and identity secrets and MCP headers, sealed with AES-256-GCM in their database columns | lib/crypto.ts (Sealer), db/rekey.ts |
| The master key | WIREFACE_MASTER_KEY, or DATA_DIR/master.key, generated on first run | context.ts |
| Live state | Memory only: sessions, voice sessions, rate limits, the event bus | gateway/, lib/rate-limit.ts, events/bus.ts |
A few conventions hold across the schema: ids are prefixed text (cnv_..., msg_..., see packages/shared/src/ids.ts), times are epoch milliseconds, and JSON is stored as text. Every row that belongs to a customer has a workspace_id, and the services filter by it. better-sqlite3 is synchronous, so a transaction is a plain function: tx(db, () => ...) in db/client.ts. Full-text search (messages_fts for the admin's search, kb_chunks_fts for knowledge retrieval) uses FTS5 tables kept in step by triggers, from the hand-written drizzle/0001_fts.sql.
What an operator backs up, and how, is in Backups.