mapping-urls
>-
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
<!--
AI-RULEZ :: GENERATED FILE — DO NOT EDIT
Content-Hash: blake3:f350881f7fc3fae75a1b3aec431fddcb0cd9fd62539b5f6ade1493d4b15a076b
Source-Hash: blake3:b5689383a3914da8e3cc6ad3614356ae7f98c90cf99cb1b4b47be38456ff7a7c
Schema-Version: v1
-->
Mapping URLs
crawlberg map <url> discovers the URLs a site exposes without rendering or
extracting any page content. It reads sitemap.xml (including nested
sitemaps), then falls back to link extraction from the seed page. Use it to
plan a crawl, audit a site's surface, or feed a URL list into another tool.
Quick recipe
crawlberg map https://example.com --limit 500 --search docs --format markdown
Markdown output prints one URL per line — convenient to pipe into a file or a
follow-up crawl. JSON output (default) returns a structured MapResult.
Flag surface
| Flag | Default | Purpose |
| ---------------------- | ------- | ------------------------------------------------------------- |
| --limit | — | Maximum number of URLs to return. Unbounded if unset. |
| --search | — | Case-insensitive substring filter on discovered URLs. |
| --respect-robots-txt | off | Honour robots.txt. Pass it for any third-party host. |
| --format | json | json (full MapResult) or markdown (one URL per line). |
| --timeout | 30000 | Per-request timeout in ms. |
| --browser-mode | auto | auto, always, never — see the headless-fallback skill. |
| --browser-endpoint | — | External CDP ws:// URL. |
| --config | — | Inline JSON or @file.json for the full CrawlConfig. |
map takes a single seed URL positionally. There is no --depth or
--max-pages here — those bound a crawl, not a map. Scope is the seed host's
sitemaps plus links found on the seed page; bound the result with --limit and
narrow it with --search.
How discovery works
- Fetch and parse
sitemap.xml, following nested<sitemapindex>entries. - If no sitemap (or a thin one), extract links from the seed page's HTML.
- Apply the
--searchsubstring filter (case-insensitive), then--limit.
No page bodies are rendered, so a map of hundreds of URLs returns in seconds —
far cheaper than crawling. In --browser-mode auto the seed fetch still falls
back to headless Chrome if the seed page is a JS shell that hides its links;
pass --browser-mode never to keep it static-only.
Output
Markdown mode
https://example.com/
https://example.com/docs/
https://example.com/docs/getting-started
https://example.com/blog/post-one
JSON mode
Top-level MapResult with a urls array; each entry carries the discovered
url. Read result.urls[i].url for each string when scripting.
crawlberg map https://example.com --format json | jq -r '.urls[].url'
Common patterns
Discover then crawl a subsection
crawlberg map https://example.com --search /docs/ --format markdown > urls.txt
Feed the filtered list into a bounded crawl, or scrape individual entries.
Audit a third-party site politely
crawlberg map https://unknown.example --respect-robots-txt --limit 200
When to reach for crawl instead
If the user needs the page content (Markdown, metadata, tables) rather than
just the URL list, use crawlberg crawl — see the crawling-a-site skill.
Reach for map first when the goal is enumeration, planning, or seeding.
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/mapping-urls/SKILL.md