Skip to main content
Integrations

Self-hosting Unlighthouse

Self-hosting Unlighthouse

The default unlighthouse CLI is built for local-first use: an SQLite file beside the working directory, blobs on local disk, the HTTP API wide open on localhost. That story breaks the moment you want a team to share a dashboard, archive scan history across runs, or run scheduled audits from a hosted runner.

This guide walks the env vars you wire to flip Unlighthouse from "my laptop" mode to "a service my team uses."

What changes vs. the CLI default

ConcernCLI defaultSelf-hosted
Row storelocal SQLite filelocal SQLite, libSQL/Turso
Blob storelocal filesystemlocal fs, S3 / R2 / MinIO, memory
AuthnoneBearer token (single shared admin token)
CORS* (loopback assumption)allowlist of dashboard origins
ProxynoneX-Forwarded-For honoured
Rate limitnonetoken-bucket per (token | IP)
ShutdownabruptSIGTERM drain with in-flight scan cancel

Everything is opt-in via env vars. Set nothing and the CLI behaves exactly as before.

Storage

Picking a row store

# Default — local SQLite file (no env needed).
UNLIGHTHOUSE_DB_URL=file:/var/lib/unlighthouse/db.sqlite

# Bare path = file: shorthand
UNLIGHTHOUSE_DB_URL=/var/lib/unlighthouse/db.sqlite

# libSQL / Turso (remote SQLite). Recommended for hosted setups —
# managed backups, branchable databases, sqlite-compat schema.
UNLIGHTHOUSE_DB_URL='libsql://your-db.turso.io?authToken=eyJ...'
# Or pass the token separately so it doesn't end up in `ps`/logs:
UNLIGHTHOUSE_DB_URL='libsql://your-db.turso.io'
UNLIGHTHOUSE_DB_AUTH_TOKEN=eyJ...

Postgres and MySQL are reserved schemes — the parser rejects them today with a clear error. The current init SQL uses SQLite syntax (unixepoch(), AUTOINCREMENT, etc.) and would need a dialect fork; an issue tracks the work.

Schema bumps after deploy: the runtime-migrations module (applyMigrations) uses better-sqlite3's sync introspection (PRAGMA table_info + prepared statements). It does not run against libSQL. Fresh databases get the latest schema via INIT_SQL_STATEMENTS; an existing libSQL DB at an older schema needs a manual upgrade. We log a warning at boot when you connect to libSQL so you know.

Picking a blob store

LHR JSONs are megabytes each and screenshots are MBs more — keep them out of your row store. Defaults to local fs; swap to S3 / R2 / MinIO with one var.

# Default — local files (no env needed).
UNLIGHTHOUSE_BLOBS_DRIVER=fs
UNLIGHTHOUSE_BLOBS_BASE=/var/lib/unlighthouse/blobs

# S3 / S3-compatible (R2, MinIO, Backblaze B2). Bucket required;
# region defaults to 'auto', endpoint derived from region for real
# AWS or set explicitly for everyone else.
UNLIGHTHOUSE_BLOBS_DRIVER=s3
UNLIGHTHOUSE_BLOBS_S3_BUCKET=my-unlighthouse-blobs
UNLIGHTHOUSE_BLOBS_S3_REGION=us-east-1
UNLIGHTHOUSE_BLOBS_S3_ENDPOINT=https://<uid>.r2.cloudflarestorage.com   # R2
# UNLIGHTHOUSE_BLOBS_S3_ENDPOINT not needed for real AWS S3
UNLIGHTHOUSE_BLOBS_S3_ACCESS_KEY_ID=AKIA...
UNLIGHTHOUSE_BLOBS_S3_SECRET_ACCESS_KEY=...

# Memory — for tests / ephemeral CI runs that don't need persistence.
UNLIGHTHOUSE_BLOBS_DRIVER=memory

The S3 credentials fall through to the AWS SDK's standard chain (~/.aws, instance profile, env vars) when you don't pass explicit keys — useful on AWS infra where the runtime already has the permissions it needs.

Auth

A single shared admin token, set via env. When unset, the API is wide open exactly as the CLI default. When set, every /api/* request must carry Authorization: Bearer <token>.

# Generate a high-entropy token. <16 chars logs a warn at boot.
UNLIGHTHOUSE_API_TOKEN=$(openssl rand -hex 32)

What's exempt from the gate:

  • OPTIONS preflight (otherwise CORS breaks)
  • /api/health and /api/ready (monitoring probes shouldn't carry the secret)
  • Connections from 127.0.0.1 / ::1 when UNLIGHTHOUSE_LOCAL_BYPASS=1 — lets an operator SSH in and curl without exporting the token to every shell session.

The token is compared via crypto.timingSafeEqual on equal-length buffers; a length mismatch short-circuits without doing the compare so timing doesn't leak the expected length.

This is a single shared admin token, not a per-user token. Anyone who has it can read every scan, mutate every site, delete every history row. For team setups that need per-user identity, revocation, or read-only "build tokens" for CI-only upload, watch the issue tracker — that landed in our spec but ships in a later release.

CORS

Default depends on whether you've set UNLIGHTHOUSE_API_TOKEN:

  • No token → CORS is open (*). The CLI dev story: you may run the dashboard from localhost, a tailnet tunnel, a VPN — without auth there's nothing meaningful for CORS to protect, and a strict allowlist would break the "show me this on my phone via tailnet" workflow.
  • Token set, env unset → localhost allowlist. The token is the barrier; CORS narrows blast radius if it leaks. Set UNLIGHTHOUSE_CORS_ORIGINS explicitly when you expose beyond localhost.

Override with a comma-separated list of dashboard origins for hosted use:

UNLIGHTHOUSE_CORS_ORIGINS=https://unlighthouse.acme.com,https://staging.acme.com

* is accepted as a sentinel for "open" but logs a loud warning when paired with UNLIGHTHOUSE_API_TOKEN — the combination means a successful XSS on any tab can read your API. Pin specific origins instead.

Allowed origins are echoed back exact-match (the Access-Control-Allow-Origin header reflects the request Origin rather than *); a Vary: Origin header is emitted so CDNs / shared caches don't cross-pollute responses.

Behind a reverse proxy

Almost every hosted setup terminates TLS at a proxy (nginx, Cloudflare, fly.io's edge, Railway's router) and forwards plain HTTP to Unlighthouse. Two switches matter:

# Tell Unlighthouse to honour X-Forwarded-For. Without this, every
# request looks like it came from the proxy IP — auth bypass and
# rate-limiting both lose the real client identity.
UNLIGHTHOUSE_TRUST_PROXY=1
Do not combine UNLIGHTHOUSE_TRUST_PROXY=1 with UNLIGHTHOUSE_LOCAL_BYPASS=1 in a hosted deploy. The proxy itself is on the loopback, so every request via the proxy looks "local" and the auth gate goes silent for the whole internet. A loud warn logs at boot when you do this, but the start isn't blocked.

The leftmost entry in X-Forwarded-For is treated as the originator (per the standard); intermediate proxy hops are ignored. Trust-proxy is a binary toggle — there's no proxy-IP allowlist today, so only enable when you actually own the proxy.

Rate limit

In-memory token bucket per (token | IP). Default 120 req/min is comfortable for a dashboard that's polling and chatty without inviting runaway loops.

# Cap at 60 requests/minute per bucket.
UNLIGHTHOUSE_RATE_LIMIT=60

# Disable entirely.
UNLIGHTHOUSE_RATE_LIMIT=0

429 responses include Retry-After (seconds till the next token) plus X-RateLimit-Limit and X-RateLimit-Remaining on every response so well-behaved clients self-throttle.

/health, /ready, and OPTIONS preflights are exempt — monitoring probes shouldn't trip the limit.

The bucket is in-process. Two replicas behind a load balancer each get their own 120/min budget; a determined caller can multiply the cap by the replica count. For horizontal scale, swap for a Redis-backed limiter (the surface area is one Map — the swap is mechanical) — issue tracked.

Graceful shutdown

SIGTERM (k8s, systemd, fly.io, Railway) and SIGINT (Ctrl+C) both trigger a drain:

  1. Stop accepting new connections (listening socket closed, load balancer routes elsewhere).
  2. Cancel any in-flight scan (releases Chrome workers cleanly so the next boot doesn't have to reap zombies).
  3. Best-effort .close() on the row-store DB handle.
  4. process.exit(0).

Cap with:

# Seconds. Default 10 — long enough for an audit currently inside
# lighthouse() to wind down, short enough that platform-level
# escalation (k8s grace=30s, systemd TimeoutStopSec=90s) doesn't
# fire.
UNLIGHTHOUSE_SHUTDOWN_TIMEOUT=10

A second SIGTERM/SIGINT during a drain forces immediate exit(1) — matches the convention nuxt / next / vite use for twice-pressed Ctrl+C.

Example deployments

Single-VM with local SQLite + filesystem blobs

For small teams or staging environments. Pair with a process supervisor (systemd, pm2) that restarts on failure.

UNLIGHTHOUSE_API_TOKEN=$(openssl rand -hex 32)
UNLIGHTHOUSE_CORS_ORIGINS=https://lh.your-team.dev
UNLIGHTHOUSE_TRUST_PROXY=1
UNLIGHTHOUSE_DB_URL=file:/var/lib/unlighthouse/db.sqlite
UNLIGHTHOUSE_BLOBS_BASE=/var/lib/unlighthouse/blobs

Run behind nginx terminating TLS:

server {
  listen 443 ssl http2;
  server_name lh.your-team.dev;
  # ... ssl_certificate / ssl_certificate_key ...

  location / {
    proxy_pass http://127.0.0.1:5678;
    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-For $remote_addr;
    proxy_set_header X-Forwarded-Proto $scheme;
    # WebSocket upgrade for the live scan stream:
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection 'upgrade';
  }
}

Hosted with Turso + Cloudflare R2

For full hosted operation. Storage durability + multi-region reads come from Turso; blob storage scales independently on R2.

UNLIGHTHOUSE_API_TOKEN=$(openssl rand -hex 32)
UNLIGHTHOUSE_CORS_ORIGINS=https://unlighthouse.your-team.dev
UNLIGHTHOUSE_TRUST_PROXY=1

UNLIGHTHOUSE_DB_URL=libsql://your-team-unlighthouse.turso.io
UNLIGHTHOUSE_DB_AUTH_TOKEN=eyJ...

UNLIGHTHOUSE_BLOBS_DRIVER=s3
UNLIGHTHOUSE_BLOBS_S3_BUCKET=your-team-unlighthouse-blobs
UNLIGHTHOUSE_BLOBS_S3_ENDPOINT=https://<uid>.r2.cloudflarestorage.com
UNLIGHTHOUSE_BLOBS_S3_ACCESS_KEY_ID=$R2_KEY
UNLIGHTHOUSE_BLOBS_S3_SECRET_ACCESS_KEY=$R2_SECRET
UNLIGHTHOUSE_BLOBS_S3_REGION=auto

Docker compose (sketch)

services:
  unlighthouse:
    image: node:24
    working_dir: /app
    command: ["sh", "-c", "pnpm install && pnpm cli"]
    environment:
      UNLIGHTHOUSE_API_TOKEN: ${UNLIGHTHOUSE_API_TOKEN}
      UNLIGHTHOUSE_CORS_ORIGINS: https://lh.your-team.dev
      UNLIGHTHOUSE_TRUST_PROXY: "1"
      UNLIGHTHOUSE_DB_URL: file:/data/db.sqlite
      UNLIGHTHOUSE_BLOBS_BASE: /data/blobs
      CHROME_PATH: /usr/bin/google-chrome
      CHROME_FLAGS: "--no-sandbox --disable-setuid-sandbox"
    ports:
      - "5678:5678"
    volumes:
      - unlighthouse-data:/data
    stop_grace_period: 15s

volumes:
  unlighthouse-data:

(You'll need a base image with Chrome installed — the official browserless/chrome image or a custom Dockerfile copying google-chrome-stable into Node base.)

What's not here yet

Tier-2 work that's spec'd but not implemented:

  • Per-user identity / role split (build vs admin tokens) — today's single token is admin-only. CI-only "upload" tokens with no read/delete rights would let you safely embed credentials in build pipelines.
  • Multi-tenant routing — the HandlerCtx.tenant field is defined but unused. A multi-tenant host would resolve the token to a tenant and construct a tenant-scoped storage adapter; not wired yet.
  • Postgres / MySQL adapters — the init SQL would need a dialect fork.
  • Redis-backed rate limit — current limiter is single-process.
  • Audit logging — every mutation should land in an append-only log; today they don't.

If any of these blocks your deploy, open an issue. They're all mechanical; the order they ship in is driven by what teams actually run into.

Did this page help you?
Anything that could be done better? :)
Help us improve this page. You can edit this page on GitHub or provide anonymous feedback below.