Skip to main content
V1

Unlighthouse v1 Architecture

Unlighthouse v1 Architecture

Unlighthouse v1 is a site-wide Lighthouse scanner built on a ports-and-adapters core. The CLI and MCP server use the core engine directly. The Cloudflare app reuses the same contracts, lifecycle, route-audit, retention, and storage modules inside a native Workflow, while keeping Worker bindings and deployment policy in the app.

What changed from v0

  • v0 ran a single Node-only pipeline: chrome-launcher + puppeteer-cluster + a history.db SQLite file.
  • v1 splits that pipeline into four interchangeable ports and one event bus. Runtime-neutral scan modules are shared with a Cloudflare Workflow backed by PSI or a Lighthouse Container connected to Browser Rendering, D1, and R2.
  • createUnlighthouse / UnlighthouseContext / useLogger() / the v0 history.db are gone. createUnlighthouseCore(opts) is the only factory; everything reads and writes through the Storage port.
  • Node ≥ 24.13.1 baseline on every published package (unlocks node:sqlite, native fetch, require(esm)).

Ports

   SeedSource[]  ──►  Crawler  ──►  Auditor (often AuditorRouter)
                       │  ▲              │
                       │  └─core.run()   └── audits each URL
                       │
                       ▼
                    Storage (rows + blobs)
                       │
                       └──── @unlighthouse/core/api ──► UI
                                  │
               (Hookable bus: scan:*, assert:*, compare:*, quota:*, log)

A port qualifies for v1 when it has at least two real adapters today, or a treeshake invariant forces an import boundary. Single-adapter seams (Policy, BrowserPool, Broadcaster, RateLimiter) stay inline shapes until a second adapter ships.

SeedSource — pure URL producers

interface SeedSource {
  seeds(): AsyncIterable<Seed>
}
interface Seed { url: string; from: 'sitemap' | 'manual' | 'route-def' }

Shipped adapters: seeds/{sitemap, manual, route-definitions, fuse}. All pure HTTP except route-definitions (Node fs scan of Nuxt/Next page files). manual accepts an opaque urls: string[] and bypasses URL discovery — the answer to SPA / hash-routing / dotted-slug scans.

Crawler — drives the seed-to-audit loop

interface Crawler {
  run(opts: {
    seeds: SeedSource
    audit: (url: string, ctx: CrawlCtx) => Promise<void>
    allows?: (url: string) => boolean
    crawlDelayMs?: number
    signal?: AbortSignal
  }): AsyncIterable<CrawlEvent>
  pause?(): Promise<void>
  resume?(): Promise<void>
  state?(): CrawlerState
}

Two adapters ship under @unlighthouse/core/crawlers/*:

AdapterUse casePause/resumeWorker-safe
crawlers/parallel-mapFinite seed list, no discovery (PSI/CrUX bulk runs)yesyes
crawlers/crawleeSitemap + manual + link-discovery URL graph walkyesno (Node-only)

The Cloudflare app does not pretend its durable, step-indexed Workflow is a Crawler. It owns bounded same-origin discovery and delegates each route audit through a service binding. The Worker-safe @unlighthouse/core/runtime and @unlighthouse/core/crawlers/parallel-map entrypoints cannot pull in Crawlee or other Node-only dependencies; CI enforces that import boundary.

Auditor — produces a Lighthouse report for one URL

interface Auditor {
  audit(url: string, page?: Page, opts?: { signal?: AbortSignal }): Promise<LighthouseReport>
  readonly capabilities: AuditorCapabilities
}
interface AuditorCapabilities {
  reliablePerfScores: boolean
  reliableFieldData:  boolean
  supportsThrottling: boolean
  categories: Category[]
}

Shipped adapters: auditors/{local, cdp-connect, psi, crux, dataforseo, mock}.

  • local — spawns Chrome locally (chrome-launcher + Puppeteer). reliablePerfScores: true.
  • cdp-connect — connects to any remote Chrome over a CDP WebSocket. Covers Cloudflare Browser Run, browserless.io, and self-hosted puppeteer-cluster in one adapter. reliablePerfScores: false (network RTT contaminates LCP/TBT).
  • psi / crux / dataforseo — fetch precomputed reports. crux.capabilities.reliableFieldData = true; the others are lab/synthetic.
  • mock — fixture-driven for tests.

AuditorRouter (composition adapter, also an Auditor) takes a single pick function. Helpers ship as composable functions — roundRobinPick, weightedPick, rateLimitedPick, fallbackPick, predicatePick — so new routing strategies don't need a PR to core. The PSI 25k/day quota cliff is closed by fallbackPick({ primary: rateLimitedPick(...), onQuotaExceeded: 'local' }).

Storage — rows and blobs

interface Storage {
  sites: SiteRepository
  scans: ScanRepository
  routes: RouteRepository
  reports: ReportRepository
  comparisons: ComparisonRepository
  packRuns: PackRunRepository
  blobs: BlobStore
}

Rows use storage/drizzle; blobs use the BlobStore contract, with storage/unstorage-blobs for portable local and object-storage drivers. A pure storage/memory adapter is used in tests. The Cloudflare adapter binds the same storage contract to D1 and R2.

Adapters per host

HostCrawlerAuditorRowsBlobs
Local CLI (unlighthouse)crawleelocal (chrome-launcher)drizzle + better-sqlite3unstorage-blobs + fs
MCP server (unlighthouse mcp)parallel-map or crawleeinherits CLI auditorinherits CLI storageinherits CLI storage
Cloudflare app (apps/cloudflare)bounded Workflow discoveryPSI or Container Lighthouse, optional CrUXdrizzle + D1BlobStore + R2
SaaS / multi-tenanthost-owned orchestration or parallel-maprouteAuditors({ psi, crux, local })drizzle (sqlite today, postgres v1.x)R2 / S3 via unstorage

The dashboard, MCP, and HTTP API are projections of one command registry — same shapes, same types, regardless of host.

Multi-site and the device matrix

scan.start.input.device: Device | Device[]. One scan can fan out across mobile and desktop in a single matrix; ExtractedMetrics carries a device column and ScanRoute's primary key is (scanId, url, device). Closes LHCI #138 (63👍, open 6+ years). See decision D-029.

Multi-site is the dashboard concern: each scan.start accepts an optional site override (PR #345) so one Worker / one CLI install can scan many hostnames without spinning up separate configs. The dashboard groups runs by (site, branch) and surfaces the device-matrix as side-by-side columns.

Packs

A Pack is the v1 unit of curated, multi-audit output. It picks a problem class (Core Web Vitals, image optimisation, JS bundle health, a11y quick wins), declares which auditors it needs, and ships a reconciler that joins raw audit signals into a typed PackReport with prioritised fixes. Packs let agents and humans get an answer instead of a flat dump of Lighthouse audits.

interface Pack<TReport = unknown> {
  name: string                                   // 'cwv', 'images', '@unlighthouse-pack/geo'
  version: string
  description: string
  requires: AuditorRequirement[]
  reconciler: (ctx: PackReconcileCtx) => Promise<TReport>
  reportSchema: z.ZodType<TReport>
}

Built-in packs at v1.0 (under @unlighthouse/core/packs/*):

PackWhat it answers
overviewThe scan.summary default — top-N regressions across all categories, sub-1KB output.
cwvCore Web Vitals (LCP / CLS / INP) breakdown with field-vs-lab gap, lab + CrUX joined.
imagesPer-route image findings (modern formats, dimensions, lazy-loading) with byte savings.
js-bundleHeaviest JS chunks, unused JS, third-party blocking.
a11y-quick-winsHigh-impact, low-effort accessibility fixes prioritised by route count.
seo-basicsTitle / meta / canonical / heading / robots issues with redirect-chain detection.
cruxField data joined per-route (PR #354 — separate auditor capability, surfaced lab vs field gap).
insightsLighthouse insight audits (*-insight) ranked site-wide by total + worst-route metric savings, with a priority order.
agentic-browsingLighthouse 13 agentic-browsing readiness: WebMCP tool/form/schema coverage, llms.txt presence, agent-accessibility pass rate.

Third-party packs publish as @unlighthouse-pack/<name> on npm. Auditors flagged kind: 'custom' only run when an active pack requires them, so installing a heavy GEO pack doesn't tax scans that don't use it.

The Nuxt-aware reference pack is first-party but opt-in. Import it from @unlighthouse/core/packs/nuxt and register nuxtPack with the host or core factory; it is intentionally absent from builtInPacks.

Packages

Seven published packages at v1.0:

PackageRole
@unlighthouse/contractsTypes + Zod schemas. Single peer dep on zod ≥ 4. Imported by every other package.
@unlighthouse/coreThe engine. Adapters live as subpath exports (./seeds, ./crawlers, ./auditors, ./storage, ./packs). Owns the local Tinypool Lighthouse worker implementation.
@unlighthouse/cloudflareReusable Worker adapters and runtime classes on explicit subpaths: auditors, D1/R2 storage, seeds, Durable Objects, and the scan Workflow. The root export is storage-only.
@unlighthouse/lighthouse-containerGeneric OCI image + h3 Node server that drives real Lighthouse against an externally managed Chrome CDP endpoint.
@unlighthouse/mcpMCP server preset. Exposes the command registry over stdio for Claude Code, Cursor, and other MCP clients.
@unlighthouse/uiStatic UI bundle. Types-only on @unlighthouse/contracts + @unlighthouse/core/api/types — no runtime dep on the engine.
unlighthouseCLI preset (the npm-name umbrella). Owns bins (unlighthouse, unlighthouse-ci, unlighthouse-mcp), c12 config loader, listhen, host-specific imperative rules.

The maintained Cloudflare deployment lives in apps/cloudflare/, not inside the adapter package. The app owns the Worker entrypoint, bindings, migrations, auditor tier policy, assets, and deployment runbook; @unlighthouse/cloudflare stays reusable and environment-agnostic within the Cloudflare runtime.

The dependency graph is acyclic:

contracts            (peer: zod)
   ↑
   └── core              (peer: contracts; drizzle-orm + unstorage)
          ↑
          ├── ui                                    [types-only on core/api]
          ├── lighthouse-container                  [Node/OCI image]
          └── mcp, cloudflare, unlighthouse         [hosts / adapters]

CI enforces a treeshake invariant per host bundle (Worker HTTP-only, Worker CF rendering, Worker memory-only, UI typecheck, MCP). A Worker bundle that accidentally pulls in lighthouse, chrome-launcher, better-sqlite3, or listhen fails the build.

Where to go next

Did this page help you?
Anything that could be done better? :)
Help us improve this page. You can edit this page on GitHub or provide anonymous feedback below.