Unlighthouse v1 Architecture
Unlighthouse v1 Architecture
Unlighthouse v1 is a site-wide Lighthouse scanner built on a ports-and-adapters core. The CLI and MCP server use the core engine directly. The Cloudflare app reuses the same contracts, lifecycle, route-audit, retention, and storage modules inside a native Workflow, while keeping Worker bindings and deployment policy in the app.
What changed from v0
- v0 ran a single Node-only pipeline: chrome-launcher + puppeteer-cluster + a
history.dbSQLite file. - v1 splits that pipeline into four interchangeable ports and one event bus. Runtime-neutral scan modules are shared with a Cloudflare Workflow backed by PSI or a Lighthouse Container connected to Browser Rendering, D1, and R2.
createUnlighthouse/UnlighthouseContext/useLogger()/ the v0history.dbare gone.createUnlighthouseCore(opts)is the only factory; everything reads and writes through theStorageport.- Node ≥ 24.13.1 baseline on every published package (unlocks
node:sqlite, nativefetch,require(esm)).
Ports
SeedSource[] ──► Crawler ──► Auditor (often AuditorRouter)
│ ▲ │
│ └─core.run() └── audits each URL
│
▼
Storage (rows + blobs)
│
└──── @unlighthouse/core/api ──► UI
│
(Hookable bus: scan:*, assert:*, compare:*, quota:*, log)A port qualifies for v1 when it has at least two real adapters today, or a treeshake invariant forces an import boundary. Single-adapter seams (Policy, BrowserPool, Broadcaster, RateLimiter) stay inline shapes until a second adapter ships.
SeedSource — pure URL producers
interface SeedSource {
seeds(): AsyncIterable<Seed>
}
interface Seed { url: string; from: 'sitemap' | 'manual' | 'route-def' }Shipped adapters: seeds/{sitemap, manual, route-definitions, fuse}. All pure HTTP except route-definitions (Node fs scan of Nuxt/Next page files). manual accepts an opaque urls: string[] and bypasses URL discovery — the answer to SPA / hash-routing / dotted-slug scans.
Crawler — drives the seed-to-audit loop
interface Crawler {
run(opts: {
seeds: SeedSource
audit: (url: string, ctx: CrawlCtx) => Promise<void>
allows?: (url: string) => boolean
crawlDelayMs?: number
signal?: AbortSignal
}): AsyncIterable<CrawlEvent>
pause?(): Promise<void>
resume?(): Promise<void>
state?(): CrawlerState
}Two adapters ship under @unlighthouse/core/crawlers/*:
| Adapter | Use case | Pause/resume | Worker-safe |
|---|---|---|---|
crawlers/parallel-map | Finite seed list, no discovery (PSI/CrUX bulk runs) | yes | yes |
crawlers/crawlee | Sitemap + manual + link-discovery URL graph walk | yes | no (Node-only) |
The Cloudflare app does not pretend its durable, step-indexed Workflow is a Crawler. It owns bounded same-origin discovery and delegates each route audit through a service binding. The Worker-safe @unlighthouse/core/runtime and @unlighthouse/core/crawlers/parallel-map entrypoints cannot pull in Crawlee or other Node-only dependencies; CI enforces that import boundary.
Auditor — produces a Lighthouse report for one URL
interface Auditor {
audit(url: string, page?: Page, opts?: { signal?: AbortSignal }): Promise<LighthouseReport>
readonly capabilities: AuditorCapabilities
}
interface AuditorCapabilities {
reliablePerfScores: boolean
reliableFieldData: boolean
supportsThrottling: boolean
categories: Category[]
}Shipped adapters: auditors/{local, cdp-connect, psi, crux, dataforseo, mock}.
local— spawns Chrome locally (chrome-launcher + Puppeteer).reliablePerfScores: true.cdp-connect— connects to any remote Chrome over a CDP WebSocket. Covers Cloudflare Browser Run, browserless.io, and self-hosted puppeteer-cluster in one adapter.reliablePerfScores: false(network RTT contaminates LCP/TBT).psi/crux/dataforseo— fetch precomputed reports.crux.capabilities.reliableFieldData = true; the others are lab/synthetic.mock— fixture-driven for tests.
AuditorRouter (composition adapter, also an Auditor) takes a single pick function. Helpers ship as composable functions — roundRobinPick, weightedPick, rateLimitedPick, fallbackPick, predicatePick — so new routing strategies don't need a PR to core. The PSI 25k/day quota cliff is closed by fallbackPick({ primary: rateLimitedPick(...), onQuotaExceeded: 'local' }).
Storage — rows and blobs
interface Storage {
sites: SiteRepository
scans: ScanRepository
routes: RouteRepository
reports: ReportRepository
comparisons: ComparisonRepository
packRuns: PackRunRepository
blobs: BlobStore
}Rows use storage/drizzle; blobs use the BlobStore contract, with storage/unstorage-blobs for portable local and object-storage drivers. A pure storage/memory adapter is used in tests. The Cloudflare adapter binds the same storage contract to D1 and R2.
Adapters per host
| Host | Crawler | Auditor | Rows | Blobs |
|---|---|---|---|---|
Local CLI (unlighthouse) | crawlee | local (chrome-launcher) | drizzle + better-sqlite3 | unstorage-blobs + fs |
MCP server (unlighthouse mcp) | parallel-map or crawlee | inherits CLI auditor | inherits CLI storage | inherits CLI storage |
Cloudflare app (apps/cloudflare) | bounded Workflow discovery | PSI or Container Lighthouse, optional CrUX | drizzle + D1 | BlobStore + R2 |
| SaaS / multi-tenant | host-owned orchestration or parallel-map | routeAuditors({ psi, crux, local }) | drizzle (sqlite today, postgres v1.x) | R2 / S3 via unstorage |
The dashboard, MCP, and HTTP API are projections of one command registry — same shapes, same types, regardless of host.
Multi-site and the device matrix
scan.start.input.device: Device | Device[]. One scan can fan out across mobile and desktop in a single matrix; ExtractedMetrics carries a device column and ScanRoute's primary key is (scanId, url, device). Closes LHCI #138 (63👍, open 6+ years). See decision D-029.
Multi-site is the dashboard concern: each scan.start accepts an optional site override (PR #345) so one Worker / one CLI install can scan many hostnames without spinning up separate configs. The dashboard groups runs by (site, branch) and surfaces the device-matrix as side-by-side columns.
Packs
A Pack is the v1 unit of curated, multi-audit output. It picks a problem class (Core Web Vitals, image optimisation, JS bundle health, a11y quick wins), declares which auditors it needs, and ships a reconciler that joins raw audit signals into a typed PackReport with prioritised fixes. Packs let agents and humans get an answer instead of a flat dump of Lighthouse audits.
interface Pack<TReport = unknown> {
name: string // 'cwv', 'images', '@unlighthouse-pack/geo'
version: string
description: string
requires: AuditorRequirement[]
reconciler: (ctx: PackReconcileCtx) => Promise<TReport>
reportSchema: z.ZodType<TReport>
}Built-in packs at v1.0 (under @unlighthouse/core/packs/*):
| Pack | What it answers |
|---|---|
overview | The scan.summary default — top-N regressions across all categories, sub-1KB output. |
cwv | Core Web Vitals (LCP / CLS / INP) breakdown with field-vs-lab gap, lab + CrUX joined. |
images | Per-route image findings (modern formats, dimensions, lazy-loading) with byte savings. |
js-bundle | Heaviest JS chunks, unused JS, third-party blocking. |
a11y-quick-wins | High-impact, low-effort accessibility fixes prioritised by route count. |
seo-basics | Title / meta / canonical / heading / robots issues with redirect-chain detection. |
crux | Field data joined per-route (PR #354 — separate auditor capability, surfaced lab vs field gap). |
insights | Lighthouse insight audits (*-insight) ranked site-wide by total + worst-route metric savings, with a priority order. |
agentic-browsing | Lighthouse 13 agentic-browsing readiness: WebMCP tool/form/schema coverage, llms.txt presence, agent-accessibility pass rate. |
Third-party packs publish as @unlighthouse-pack/<name> on npm. Auditors flagged kind: 'custom' only run when an active pack requires them, so installing a heavy GEO pack doesn't tax scans that don't use it.
The Nuxt-aware reference pack is first-party but opt-in. Import it from @unlighthouse/core/packs/nuxt and register nuxtPack with the host or core factory; it is intentionally absent from builtInPacks.
Packages
Seven published packages at v1.0:
| Package | Role |
|---|---|
@unlighthouse/contracts | Types + Zod schemas. Single peer dep on zod ≥ 4. Imported by every other package. |
@unlighthouse/core | The engine. Adapters live as subpath exports (./seeds, ./crawlers, ./auditors, ./storage, ./packs). Owns the local Tinypool Lighthouse worker implementation. |
@unlighthouse/cloudflare | Reusable Worker adapters and runtime classes on explicit subpaths: auditors, D1/R2 storage, seeds, Durable Objects, and the scan Workflow. The root export is storage-only. |
@unlighthouse/lighthouse-container | Generic OCI image + h3 Node server that drives real Lighthouse against an externally managed Chrome CDP endpoint. |
@unlighthouse/mcp | MCP server preset. Exposes the command registry over stdio for Claude Code, Cursor, and other MCP clients. |
@unlighthouse/ui | Static UI bundle. Types-only on @unlighthouse/contracts + @unlighthouse/core/api/types — no runtime dep on the engine. |
unlighthouse | CLI preset (the npm-name umbrella). Owns bins (unlighthouse, unlighthouse-ci, unlighthouse-mcp), c12 config loader, listhen, host-specific imperative rules. |
The maintained Cloudflare deployment lives in apps/cloudflare/, not inside the adapter package. The app owns the Worker entrypoint, bindings, migrations, auditor tier policy, assets, and deployment runbook; @unlighthouse/cloudflare stays reusable and environment-agnostic within the Cloudflare runtime.
The dependency graph is acyclic:
contracts (peer: zod)
↑
└── core (peer: contracts; drizzle-orm + unstorage)
↑
├── ui [types-only on core/api]
├── lighthouse-container [Node/OCI image]
└── mcp, cloudflare, unlighthouse [hosts / adapters]CI enforces a treeshake invariant per host bundle (Worker HTTP-only, Worker CF rendering, Worker memory-only, UI typecheck, MCP). A Worker bundle that accidentally pulls in lighthouse, chrome-launcher, better-sqlite3, or listhen fails the build.
Where to go next
- Self-Host on Cloudflare — end-to-end walkthrough of the maintained
apps/cloudflaredeployment. - MCP for AI Agents — drive Unlighthouse from Claude Code / Cursor.
- CLI — the local-host story.
- API Reference — command registry, types, and config schema.