---
title: "Unlighthouse v1 Architecture"
description: "How Unlighthouse v1 is structured: runtime-neutral scan modules, explicit host adapters, and separate Node and Cloudflare orchestration."
canonical_url: "https://unlighthouse.dev/v1/architecture"
last_updated: "2026-10-03T08:07:19.058Z"
---

# Unlighthouse v1 Architecture

Unlighthouse v1 is a site-wide Lighthouse scanner built on a **ports-and-adapters** core. The CLI and MCP server use the core engine directly. The Cloudflare app reuses the same contracts, lifecycle, route-audit, retention, and storage modules inside a native Workflow, while keeping Worker bindings and deployment policy in the app.

<!-- TODO screenshot: high-level architecture diagram (CLI vs Worker, same core) -->

## What changed from v0

- v0 ran a single Node-only pipeline: chrome-launcher + puppeteer-cluster + a `history.db` SQLite file.
- v1 splits that pipeline into four interchangeable ports and one event bus. Runtime-neutral scan modules are shared with a Cloudflare Workflow backed by PSI or a Lighthouse Container connected to Browser Rendering, D1, and R2.
- `createUnlighthouse` / `UnlighthouseContext` / `useLogger()` / the v0 `history.db` are gone. `createUnlighthouseCore(opts)` is the only factory; everything reads and writes through the `Storage` port.
- Node ≥ 24.13.1 baseline on every published package (unlocks `node:sqlite`, native `fetch`, `require(esm)`).

## Ports

```
SeedSource[]  ──►  Crawler  ──►  Auditor (often AuditorRouter)
                       │  ▲              │
                       │  └─core.run()   └── audits each URL
                       │
                       ▼
                    Storage (rows + blobs)
                       │
                       └──── @unlighthouse/core/api ──► UI
                                  │
               (Hookable bus: scan:*, assert:*, compare:*, quota:*, log)
```

A port qualifies for v1 when it has at least two real adapters today, or a treeshake invariant forces an import boundary. Single-adapter seams (Policy, BrowserPool, Broadcaster, RateLimiter) stay inline shapes until a second adapter ships.

### `SeedSource` — pure URL producers

```ts
interface SeedSource {
  seeds(): AsyncIterable<Seed>
}
interface Seed { url: string; from: 'sitemap' | 'manual' | 'route-def' }
```

Shipped adapters: `seeds/{sitemap, manual, route-definitions, fuse}`. All pure HTTP except `route-definitions` (Node `fs` scan of Nuxt/Next page files). `manual` accepts an opaque `urls: string[]` and bypasses URL discovery — the answer to SPA / hash-routing / dotted-slug scans.

### `Crawler` — drives the seed-to-audit loop

```ts
interface Crawler {
  run(opts: {
    seeds: SeedSource
    audit: (url: string, ctx: CrawlCtx) => Promise<void>
    allows?: (url: string) => boolean
    crawlDelayMs?: number
    signal?: AbortSignal
  }): AsyncIterable<CrawlEvent>
  pause?(): Promise<void>
  resume?(): Promise<void>
  state?(): CrawlerState
}
```

Two adapters ship under `@unlighthouse/core/crawlers/*`:

| Adapter                 | Use case                                            | Pause/resume | Worker-safe    |
| ----------------------- | --------------------------------------------------- | ------------ | -------------- |
| `crawlers/parallel-map` | Finite seed list, no discovery (PSI/CrUX bulk runs) | yes          | yes            |
| `crawlers/crawlee`      | Sitemap + manual + link-discovery URL graph walk    | yes          | no (Node-only) |

The Cloudflare app does not pretend its durable, step-indexed Workflow is a `Crawler`. It owns bounded same-origin discovery and delegates each route audit through a service binding. The Worker-safe `@unlighthouse/core/runtime` and `@unlighthouse/core/crawlers/parallel-map` entrypoints cannot pull in Crawlee or other Node-only dependencies; CI enforces that import boundary.

### `Auditor` — produces a Lighthouse report for one URL

```ts
interface Auditor {
  audit(url: string, page?: Page, opts?: { signal?: AbortSignal }): Promise<LighthouseReport>
  readonly capabilities: AuditorCapabilities
}
interface AuditorCapabilities {
  reliablePerfScores: boolean
  reliableFieldData:  boolean
  supportsThrottling: boolean
  categories: Category[]
}
```

Shipped adapters: `auditors/{local, cdp-connect, psi, crux, dataforseo, mock}`.

- **`local`** — spawns Chrome locally (chrome-launcher + Puppeteer). `reliablePerfScores: true`.
- **`cdp-connect`** — connects to any remote Chrome over a CDP WebSocket. Covers Cloudflare Browser Run, [browserless.io](http://browserless.io), and self-hosted puppeteer-cluster in one adapter. `reliablePerfScores: false` (network RTT contaminates LCP/TBT).
- **`psi`** / **`crux`** / **`dataforseo`** — fetch precomputed reports. `crux.capabilities.reliableFieldData = true`; the others are lab/synthetic.
- **`mock`** — fixture-driven for tests.

`AuditorRouter` (composition adapter, also an `Auditor`) takes a single `pick` function. Helpers ship as composable functions — `roundRobinPick`, `weightedPick`, `rateLimitedPick`, `fallbackPick`, `predicatePick` — so new routing strategies don't need a PR to `core`. The PSI 25k/day quota cliff is closed by `fallbackPick({ primary: rateLimitedPick(...), onQuotaExceeded: 'local' })`.

### `Storage` — rows and blobs

```ts
interface Storage {
  sites: SiteRepository
  scans: ScanRepository
  routes: RouteRepository
  reports: ReportRepository
  comparisons: ComparisonRepository
  packRuns: PackRunRepository
  blobs: BlobStore
}
```

Rows use `storage/drizzle`; blobs use the `BlobStore` contract, with `storage/unstorage-blobs` for portable local and object-storage drivers. A pure `storage/memory` adapter is used in tests. The Cloudflare adapter binds the same storage contract to D1 and R2.

## Adapters per host

| Host                               | Crawler                                    | Auditor                                    | Rows                                    | Blobs                  |
| ---------------------------------- | ------------------------------------------ | ------------------------------------------ | --------------------------------------- | ---------------------- |
| Local CLI (`unlighthouse`)         | `crawlee`                                  | `local` (chrome-launcher)                  | `drizzle` + `better-sqlite3`            | `unstorage-blobs` + fs |
| MCP server (`unlighthouse mcp`)    | `parallel-map` or `crawlee`                | inherits CLI auditor                       | inherits CLI storage                    | inherits CLI storage   |
| Cloudflare app (`apps/cloudflare`) | bounded Workflow discovery                 | PSI or Container Lighthouse, optional CrUX | `drizzle` + D1                          | `BlobStore` + R2       |
| SaaS / multi-tenant                | host-owned orchestration or `parallel-map` | `routeAuditors({ psi, crux, local })`      | `drizzle` (sqlite today, postgres v1.x) | R2 / S3 via unstorage  |

The dashboard, MCP, and HTTP API are projections of one command registry — same shapes, same types, regardless of host.

## Multi-site and the device matrix

`scan.start.input.device: Device | Device[]`. One scan can fan out across mobile and desktop in a single matrix; `ExtractedMetrics` carries a `device` column and `ScanRoute`'s primary key is `(scanId, url, device)`. Closes [LHCI #138](https://github.com/GoogleChrome/lighthouse-ci/issues/138) (63👍, open 6+ years). See decision **D-029**.

Multi-site is the dashboard concern: each `scan.start` accepts an optional `site` override (PR #345) so one Worker / one CLI install can scan many hostnames without spinning up separate configs. The dashboard groups runs by `(site, branch)` and surfaces the device-matrix as side-by-side columns.

<!-- TODO screenshot: multi-site dashboard with mobile + desktop columns -->

## Packs

A **Pack** is the v1 unit of curated, multi-audit output. It picks a problem class (Core Web Vitals, image optimisation, JS bundle health, a11y quick wins), declares which auditors it needs, and ships a reconciler that joins raw audit signals into a typed `PackReport` with prioritised fixes. Packs let agents and humans get an *answer* instead of a flat dump of Lighthouse audits.

```ts
interface Pack<TReport = unknown> {
  name: string                                   // 'cwv', 'images', '@unlighthouse-pack/geo'
  version: string
  description: string
  requires: AuditorRequirement[]
  reconciler: (ctx: PackReconcileCtx) => Promise<TReport>
  reportSchema: z.ZodType<TReport>
}
```

Built-in packs at v1.0 (under `@unlighthouse/core/packs/*`):

| Pack               | What it answers                                                                                                                 |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------- |
| `overview`         | The `scan.summary` default — top-N regressions across all categories, sub-1KB output.                                           |
| `cwv`              | Core Web Vitals (LCP / CLS / INP) breakdown with field-vs-lab gap, lab + CrUX joined.                                           |
| `images`           | Per-route image findings (modern formats, dimensions, lazy-loading) with byte savings.                                          |
| `js-bundle`        | Heaviest JS chunks, unused JS, third-party blocking.                                                                            |
| `a11y-quick-wins`  | High-impact, low-effort accessibility fixes prioritised by route count.                                                         |
| `seo-basics`       | Title / meta / canonical / heading / robots issues with redirect-chain detection.                                               |
| `crux`             | Field data joined per-route (PR #354 — separate auditor capability, surfaced lab vs field gap).                                 |
| `insights`         | Lighthouse insight audits (`*-insight`) ranked site-wide by total + worst-route metric savings, with a priority order.          |
| `agentic-browsing` | Lighthouse 13 agentic-browsing readiness: WebMCP tool/form/schema coverage, `llms.txt` presence, agent-accessibility pass rate. |

Third-party packs publish as `@unlighthouse-pack/<name>` on npm. Auditors flagged `kind: 'custom'` only run when an active pack requires them, so installing a heavy GEO pack doesn't tax scans that don't use it.

The Nuxt-aware reference pack is first-party but opt-in. Import it from `@unlighthouse/core/packs/nuxt` and register `nuxtPack` with the host or core factory; it is intentionally absent from `builtInPacks`.

## Packages

Seven published packages at v1.0:

| Package                              | Role                                                                                                                                                                        |
| ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `@unlighthouse/contracts`            | Types + Zod schemas. Single peer dep on `zod ≥ 4`. Imported by every other package.                                                                                         |
| `@unlighthouse/core`                 | The engine. Adapters live as subpath exports (`./seeds`, `./crawlers`, `./auditors`, `./storage`, `./packs`). Owns the local Tinypool Lighthouse worker implementation.     |
| `@unlighthouse/cloudflare`           | Reusable Worker adapters and runtime classes on explicit subpaths: auditors, D1/R2 storage, seeds, Durable Objects, and the scan Workflow. The root export is storage-only. |
| `@unlighthouse/lighthouse-container` | Generic OCI image + h3 Node server that drives real Lighthouse against an externally managed Chrome CDP endpoint.                                                           |
| `@unlighthouse/mcp`                  | MCP server preset. Exposes the command registry over stdio for Claude Code, Cursor, and other MCP clients.                                                                  |
| `@unlighthouse/ui`                   | Static UI bundle. Types-only on `@unlighthouse/contracts` + `@unlighthouse/core/api/types` — no runtime dep on the engine.                                                  |
| `unlighthouse`                       | CLI preset (the npm-name umbrella). Owns bins (`unlighthouse`, `unlighthouse-ci`, `unlighthouse-mcp`), c12 config loader, listhen, host-specific imperative rules.          |

The maintained Cloudflare deployment lives in `apps/cloudflare/`, not inside the adapter package. The app owns the Worker entrypoint, bindings, migrations, auditor tier policy, assets, and deployment runbook; `@unlighthouse/cloudflare` stays reusable and environment-agnostic within the Cloudflare runtime.

The dependency graph is acyclic:

```
contracts            (peer: zod)
   ↑
   └── core              (peer: contracts; drizzle-orm + unstorage)
          ↑
          ├── ui                                    [types-only on core/api]
          ├── lighthouse-container                  [Node/OCI image]
          └── mcp, cloudflare, unlighthouse         [hosts / adapters]
```

CI enforces a treeshake invariant per host bundle (Worker HTTP-only, Worker CF rendering, Worker memory-only, UI typecheck, MCP). A Worker bundle that accidentally pulls in `lighthouse`, `chrome-launcher`, `better-sqlite3`, or `listhen` fails the build.

## Where to go next

- [Self-Host on Cloudflare](/self-host-cloudflare) — end-to-end walkthrough of the maintained `apps/cloudflare` deployment.
- [MCP for AI Agents](/integrations/mcp) — drive Unlighthouse from Claude Code / Cursor.
- [CLI](/integrations/cli) — the local-host story.
- [API Reference](/api-doc) — command registry, types, and config schema.

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
