Appearance
Testing
| Command | What it runs |
|---|---|
pnpm test | Unit and integration tests (Vitest), from the repo root |
pnpm e2e | Browser tests (Playwright) against a real server |
pnpm -r typecheck | TypeScript in every package that has a typecheck script (the same as pnpm typecheck) |
npx biome check . | Formatting and lint (the same as pnpm lint) |
pnpm --filter @wireface/docs gen | The generated reference pages, which fail when the code and their notes drift apart |
Unit and integration tests
sh
pnpm test # everything, once
pnpm test text-chat # only test files whose path contains "text-chat"
pnpm test:watch # rerun on changeAlways run Vitest from the repo root. The root vitest.config.ts collects packages/*/test/**/*.test.{ts,tsx} and apps/*/test/**/*.test.{ts,tsx}, runs them in forked processes and sets a 15 second test timeout. Running Vitest inside a package (pnpm --filter @wireface/server exec vitest) skips that config, so tests fall back to Vitest's 5 second default and the slower ones, such as the origin rejection test, time out.
The server tests (apps/server/test) start a real server against the mock providers: setup and auth, providers, text chat, voice, tools and MCP, knowledge, vision, identity, origins, SSRF, spoken replies, workspaces and more. The packages with tests of their own are protocol, shared, markdown, snippets and chat.
The server test helpers
apps/server/test/helpers.ts has what most server tests need:
| Helper | What it does |
|---|---|
startTestApp(overrides?) | Starts the mock providers on a random port, then a real server (buildApp()) on another, with an in-memory SQLite database (openDb(':memory:')), a temporary DATA_DIR for files, a random master key, SETUP_TOKEN=test-setup-token, silent logs, and every provider pointed at the mocks (mockEnv()). overrides are more environment variables. It returns { app, ctx, mock, base, api(), close() }. |
t.api(method, url, body?) | A REST call through app.inject(), keeping the session cookie and CSRF token between calls. |
setupOwner(t) | Runs first-run setup and signs in as the owner. |
publishedBot(t, opts?) | Connects a provider if there isn't one (opts.provider: Anthropic by default, or OpenAI or Gemini), creates a bot from the support template, sets its brain and allowedOrigins: ['https://shop.test'], applies opts.patch and publishes it. Returns { id, publicId, connectionId }. |
TestWidget.open(t, publicId) | A widget WebSocket that records everything the server sends: wait(type, test?), of(type), send(), sendBinary(), clear(). |
helloWidget(t, publicId, hello?) | Opens a TestWidget, sends hello from https://shop.test/products, and waits for session.ready. Returns { w, ready }. |
chat(w, text, attachments?) | Sends chat.send and waits for the next agent.response.end. |
A test file usually shares one server:
ts
import { afterAll, beforeAll, expect, it } from 'vitest';
import { chat, helloWidget, publishedBot, setupOwner, startTestApp, type TestApp } from './helpers.ts';
let t: TestApp;
beforeAll(async () => {
t = await startTestApp();
await setupOwner(t);
});
afterAll(async () => t?.close());
it('answers a typed message', async () => {
const bot = await publishedBot(t);
const { w } = await helloWidget(t, bot.publicId);
const end = await chat(w, 'hello there');
expect(end.text).toContain('You said: hello there');
w.close();
});t.mock.requests holds every request the mocks received (route, body and headers), for asserting on what the server sent a provider. t.ctx is the server's context, for reaching services directly.
The mock providers
packages/mock-providers is one HTTP server on 127.0.0.1 that answers like the four providers: routes under /anthropic, /openai/v1, /gemini and /elevenlabs, with SSE streams in each provider's own format, and WebSockets for OpenAI Realtime (and its transcription sessions), Gemini Live, ElevenLabs Scribe and ElevenLabs Agents. In code, startMockProviders(port) starts it and mockEnv(mock) gives the PROVIDER_* variables that point a server at it. pnpm --filter @wireface/mock-providers start runs it on its own (see Local setup).
Any API key works, except a key containing bad, which is rejected (401, or Gemini's 400 API_KEY_INVALID), and one containing broke, which is out of credit (429 insufficient_quota).
The mock model (src/brain.ts) answers the last user turn deterministically, so tests can assert on the reply:
| The visitor's message | The mock model |
|---|---|
contains please fail | An HTTP 500 (the visitor sees the bot's apology) |
contains please refuse | A refusal (the visitor sees the bot's refusal message) |
use <tool> with {json} | A call to that tool, if the request offers it (by its exact name, or a name ending in _<tool>, such as an MCP tool), with the JSON as its arguments. After the result: The <tool> tool said: <result>. |
| has an image | I can see the image you sent. |
| anything else | You said: <text>. How else can I help? |
text
use capture_lead with {"name": "Sam", "email": "sam@example.com"}Replies start with a [[mood:...]] tag and stream in small pieces, so tags split across deltas are exercised too. Text to speech returns a tone (about 60 ms per character, 3 seconds at most). The realtime mocks hear "speech" in any loud audio (200 ms of it starts a turn, 300 ms of quiet ends it), and what the visitor "said" is mockVoice.transcript (hello from voice unless a test sets it).
src/showcase.ts is the exception: a bot whose system prompt names Bloom & Co. gets scripted answers instead, from the assistant of a florist (Sunday delivery, anniversary flowers, a wedding call-back through capture_lead, a handoff). They exist for the docs screenshots; every other bot gets the usual replies, so no test depends on them. Change the script and e2e/tests/docs-screens.spec.ts together.
Browser tests
The Playwright tests in e2e/ drive Chromium against a real server. Before the first run, build and install the browser:
sh
pnpm build # the e2e server serves the built widget and admin panel
pnpm --filter @wireface/e2e exec playwright install chromiumThen:
sh
pnpm e2e # every test, in the desktop and mobile projects
pnpm e2e voice # only files matching "voice"
pnpm --filter @wireface/e2e test widget --project desktop
E2E_DEBUG=1 pnpm e2e # with the server's and the mocks' outpute2e/global-setup.ts prepares a whole world before the tests:
- It starts the mock providers on port 18899.
- It starts the server from source on
127.0.0.1:18800, withNODE_ENV=production, a temporaryDATA_DIR, a fixed setup token and master key, everyPROVIDER_*variable pointing at the mocks, the real provider keys blanked, and wireface.dev sign-in turned on so the sign-in page offers it. - It serves the fixture shop site from
e2e/fixturesonhttp://localhost:18900, another origin than the server, as a customer's site is. The pages'{{HOST}}and{{BOT}}are filled in, and?bot=<key>picks the bot. - Through the REST API it runs setup, connects Anthropic and OpenAI (both mocked), and publishes a bot for each case:
panel,drawer,floating,inline,tools(page tools),closed(another site's origin only),vision,speaker(typed replies read aloud),voice(OpenAI Realtime) andbloom(the screenshots). - It writes
e2e/.state.json: the server and shop addresses, each bot's public id by key, the owner's cookie and CSRF token, and the process ids. Tests read it, andglobal-teardown.tsstops the processes. The file is ignored by Git and by Biome.
The tests run one at a time, with a 45 second timeout each. Chromium gets a fake microphone (e2e/fixtures/speech.wav, half a second of tone and a second of silence, looped; it is generated when missing) and a fake camera. The desktop project runs everything not tagged @mobile; mobile (a Pixel 7) runs only the tests tagged @mobile. A failing test keeps a trace and a screenshot in e2e/test-results.
If a run is killed before teardown, its processes can outlive it and keep ports 18800, 18899 and 18900. The next run then fails in global setup, because the old server answers. Stop them first (their ids are in e2e/.state.json).
The docs screenshots
The screenshots in these docs, and the website's copies, come from e2e/tests/docs-screens.spec.ts, which drives the bloom bot through the scripted answers of showcase.ts. It only runs when asked:
sh
pnpm build
DOCS_SCREENS=1 pnpm --filter @wireface/e2e test docs-screens --project desktoppowershell
pnpm build
$env:DOCS_SCREENS = '1'; pnpm --filter @wireface/e2e test docs-screens --project desktopIt writes the raw captures to e2e/screenshots/docs, the docs' PNGs (at most 1600 px wide) to docs/public/screenshots, and WebP copies for the website (960 and 1500 px) to ../site/assets/chat when the website's repo is next to this one. SCREENSHOTS=1 runs screens.spec.ts instead, which captures each layout to e2e/screenshots for a quick visual check.
Type checking and Biome
sh
pnpm -r typecheck # tsc -p tsconfig.json in each package
npx biome check . # formatting and lint for the whole repo
npx biome check --write . # the same, applying safe fixes and formattingThe Biome settings are in biome.json: two spaces, 120 columns, single quotes, semicolons and trailing commas. It skips vendor/, build output, packages/widget/releases, e2e/.state.json, Playwright's output and generated *.gen.ts files.
After a widget build, pnpm --filter @wireface/widget size checks the size budgets: 12 KB gzipped for the loader and 48 KB for the frame's entry.
The docs reference generator
Part of the Reference is generated from the code by docs/scripts/gen.ts:
sh
pnpm --filter @wireface/docs genpnpm docs:dev and pnpm docs:build run it first, and so does the Docker image build.
| Page | Generated from |
|---|---|
reference/javascript-api.md, reference/events.md | packages/widget/src/loader/types.ts and loader/index.ts |
reference/configuration.md | packages/shared/src/bot-config.ts (the zod schema's types and defaults) |
reference/errors.md | packages/protocol/src/errors.ts |
reference/protocol.md | packages/protocol/src/types.ts, schemas.ts and binary.ts |
reference/environment.md | apps/server/src/config/env.ts |
reference/strings.md | packages/widget/src/frame/i18n.ts |
reference/rest-api.md | apps/server/src/http/routes/*.ts and auth/rbac.ts |
The words come from docs/scripts/notes/*.ts (config, env, errors, js-api and protocol). The generator fails, listing each problem, when something in the code has no note, or a note is left for something that's gone. It also fails when a client message has no runtime schema in schemas.ts, a REST route has no summary or a tag missing from TAG_TITLES in gen-rest.ts, or a bot setting whose default gen-config.ts describes in words (DEFAULT_WORDS) gets a new default. So a change that adds an environment variable, a bot setting, an error code, a protocol message or a JavaScript API member also needs its note. Don't edit the generated pages: they say so at the top. Change the source or the note and run the generator again.
Flaky runs
Some failures come from the machine, not the code. On some Windows machines loopback connections stall for seconds now and then (Vitest has reported connect ETIMEDOUT 127.0.0.1), and the widget waits 6 seconds before it drops a stalled WebSocket and tries again, which is long for a test. Rerun a failing e2e test once, on its own, before blaming the code. A test that fails twice is a real failure.