跳到主要内容
知仓学习社ZHICANG

crawling-a-site

>-

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少1 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Crawling a site

Reach for kreuzcrawl crawl when one URL is not enough — the user wants

the docs site, the blog, the marketing pages, or the whole domain.

Quick recipe

kreuzcrawl crawl https://example.com \
  --depth 3 \
  --max-pages 200 \
  --concurrent 8 \
  --rate-limit 250 \
  --stay-on-domain \
  --respect-robots-txt \
  --format markdown

Defaults you should usually override:

  • --depth 2 is shallow — set it explicitly.
  • --max-pages is unbounded by default; cap it for any unknown site.
  • --concurrent 10 is aggressive for small hosts; drop to 4-8 for

third-party sites.

Flag surface

| Flag | Default | Purpose |

| ----------------------- | ------- | ---------------------------------------------------------- |

| --depth, -d | 2 | Maximum hop count from the seed URL. |

| --max-pages, -n | — | Hard cap on pages fetched. Set this on any unknown site. |

| --concurrent, -c | 10 | Parallel in-flight requests. |

| --rate-limit | 200 | Milliseconds between requests to the same origin. |

| --stay-on-domain | off | Skip links that leave the seed domain. |

| --respect-robots-txt | off | Honour robots.txt. Pass it for any third-party host. |

| --proxy | — | HTTP, HTTPS, or SOCKS5 proxy URL. |

| --user-agent | — | Override the request UA. Be honest. |

| --timeout | 30000 | Per-request timeout in ms. |

| --browser-mode | auto | auto, always, never — see the headless-fallback skill. |

| --browser-endpoint | — | External CDP ws:// URL. |

| --format | json | json or markdown. |

| --config | — | Inline JSON or @file.json for the full CrawlConfig. |

Multiple seed URLs are accepted positionally — the engine fans out with

batch_crawl and aggregates results.

When to pick which flags

Docs sites you own

kreuzcrawl crawl https://docs.example.com \
  --depth 5 --max-pages 1000 --concurrent 16 --rate-limit 100 \
  --stay-on-domain --format markdown > docs.md

Higher concurrency and lower rate limits are fine on infrastructure you

control.

Third-party sites

kreuzcrawl crawl https://blog.unknown.example \
  --depth 2 --max-pages 50 --concurrent 4 --rate-limit 500 \
  --stay-on-domain --respect-robots-txt --format markdown

Stay shallow, cap pages, throttle hard, obey robots.

Multi-seed batch

kreuzcrawl crawl \
  https://example.com/blog \
  https://example.com/docs \
  https://example.com/pricing \
  --depth 2 --max-pages 100 --stay-on-domain --format json

JSON output for batch is an array of { seed_url, result } entries — each

result is a full crawl payload or { error: ... }.

Output

Markdown mode

---
URL: https://example.com/page-one
---
# Page One

… markdown content …

---
URL: https://example.com/page-two
---
…

JSON mode

Top-level CrawlResult with pages: [...]. Each page carries the rendered

Markdown plus metadata, links, images, JSON-LD, and HTTP response info. Read

result.pages[i].markdown.content for the Markdown string.

Politeness checklist

  • Pass --respect-robots-txt on every third-party crawl.
  • Cap --max-pages — a runaway BFS can issue tens of thousands of requests.
  • Bump --rate-limit for hosts that show signs of stress (5xx, slowdowns).
  • Identify yourself via --user-agent kreuzcrawl (contact@example.com).

Common pitfalls

  • No pages returned. The seed page may be JS-only — the engine falls

back to headless automatically in --browser-mode auto, but never mode

will silently produce an empty crawl. Re-run with --browser-mode always

or check the headless-fallback skill.

  • Crawl leaves the domain. Pass --stay-on-domain. Combine with

allow_subdomains: true in --config JSON to include subdomains.

  • Slow crawl. The default rate limit is 200 ms per origin — multiple

seed URLs on the same host still share the bucket. Spread seeds across

hosts or raise --concurrent for unrelated origins.

  • Memory growth. Each page carries full Markdown plus structured data.

Stream JSON output to a file rather than holding it in memory; set

--max-pages aggressively if downstream cannot keep up.

When to reach for map instead

If the user only needs the list of URLs (sitemap analysis, link planning,

seeding another tool), use kreuzcrawl map <url> — it skips rendering and

returns a flat MapResult with hundreds of URLs in seconds.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/kreuzberg-dev/plugins/plugins/kreuzcrawl/skills/crawling-a-site/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能

同名技能的其他版本

有 2 个不同仓库或目录里都有叫 crawling-a-site 的技能。它们内容并不相同,别混用: