---
title: "URL Discovery"
description: "URL Discovery for the Unlighthouse v1 beta."
canonical_url: "https://unlighthouse.dev/v1/guide/guides/url-discovery"
last_updated: "2026-10-03T08:07:17.315Z"
---

The v1 host combines explicit URLs, sitemaps, and route definitions.
The crawler follows links from server-rendered HTML when link following is enabled.

## Sitemap URLs

```ts
import { defineUnlighthouseConfig } from 'unlighthouse/config'

export default defineUnlighthouseConfig({
  site: 'https://example.com',
  scanner: { sitemap: ['https://example.com/sitemap.xml'], maxRoutes: 100 },
})
```

Set `scanner.sitemap: false` to disable sitemap discovery.
The v1 pipeline does not read legacy robots.txt sitemap or exclusion rules.
Provide sitemap URLs explicitly and configure `scanner.exclude` for excluded paths.

## Explicit URLs

```ts
import { defineUnlighthouseConfig } from 'unlighthouse/config'

export default defineUnlighthouseConfig({
  site: 'https://example.com',
  urls: ['/', '/about', '/blog/example'],
})
```

Explicit URLs stop link following.
If you use a URL provider function, also supply `site`.

## Stop link following

Set `scanner.crawler: false` to audit only seed URLs.
Use `scanner.maxRoutes` to limit a scan that follows links.

## JavaScript-rendered links

The crawler does not render the application in a browser.
Use a sitemap or explicit URLs for routes whose links appear after JavaScript runs.
Lighthouse still audits each page in Chrome.

## Route names

Configure [Route Definitions](/guide/guides/route-definitions) to match URLs to framework page templates.
Dynamic templates need concrete URLs from another seed source or the crawler.

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
