Skip to main content
Guides

URL Discovery

The v1 host combines explicit URLs, sitemaps, and route definitions. The crawler follows links from server-rendered HTML when link following is enabled.

Sitemap URLs

import { defineUnlighthouseConfig } from 'unlighthouse/config'

export default defineUnlighthouseConfig({
  site: 'https://example.com',
  scanner: { sitemap: ['https://example.com/sitemap.xml'], maxRoutes: 100 },
})

Set scanner.sitemap: false to disable sitemap discovery. The v1 pipeline does not read legacy robots.txt sitemap or exclusion rules. Provide sitemap URLs explicitly and configure scanner.exclude for excluded paths.

Explicit URLs

import { defineUnlighthouseConfig } from 'unlighthouse/config'

export default defineUnlighthouseConfig({
  site: 'https://example.com',
  urls: ['/', '/about', '/blog/example'],
})

Explicit URLs stop link following. If you use a URL provider function, also supply site.

Set scanner.crawler: false to audit only seed URLs. Use scanner.maxRoutes to limit a scan that follows links.

The crawler does not render the application in a browser. Use a sitemap or explicit URLs for routes whose links appear after JavaScript runs. Lighthouse still audits each page in Chrome.

Route names

Configure Route Definitions to match URLs to framework page templates. Dynamic templates need concrete URLs from another seed source or the crawler.

Did this page help you?
Anything that could be done better? :)
Help us improve this page. You can edit this page on GitHub or provide anonymous feedback below.