跳到主要内容
知仓学习社ZHICANG

headless-fallback

>-

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少1 个文本文件,4 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Headless fallback

Some pages are unscrapable without a real browser — SPA shells, infinite

scroll, Cloudflare interstitials, JS-rendered article bodies. Kreuzcrawl

ships with an optional headless-Chrome backend driven by chromiumoxide.

Modes

--browser-mode auto    # default — try static first, fall back to browser on JS/WAF
--browser-mode always  # skip static, go straight to browser
--browser-mode never   # static only, fail closed

auto (default)

The engine fetches statically, then inspects the response. It launches

headless Chrome and re-fetches when it sees:

  • WAF responses from one of 8 detected vendors (Cloudflare, Akamai, AWS WAF,

Imperva, DataDome, PerimeterX, Sucuri, F5).

  • SPA shells: <noscript> warnings, near-empty <body> with heavy JS.
  • Heuristic JS-render-required signals.

This is the right default. The browser only spins up when needed.

always

Skip the static probe entirely. Use when:

  • The user already told you the page needs JS.
  • You are scraping a site you know is React/Vue/Svelte SPA.
  • You need <script>-emitted state that never lands in static HTML.
kreuzcrawl scrape https://spa.example.com --browser-mode always --format markdown

never

Static only — the browser path is disabled. Use when:

  • You are in a hot loop where a stray Chrome launch would blow the budget.
  • You are running in a sandbox without a Chrome binary.
  • The user explicitly wants only static fetches.

In never mode, JS-only pages return empty/stub content. Inspect

markdown.content and markdown.warnings before treating the result as

final.

Symptoms that point to headless

In --browser-mode never or when you suspect the auto detector missed a

signal:

  • markdown.content is short, nav-only, or just a loading message.
  • status_code is 200 but metadata.headings is empty on a page that

clearly has headings.

  • markdown.warnings mentions JS-render-required or WAF detection.
  • 403/406/503 with WAF response headers (server: cloudflare,

cf-mitigated, x-amz-cf-id, set-cookie: __cf_bm=…).

Re-run with --browser-mode always. If that succeeds, leave it set for

that host.

External CDP endpoint

Point at an already-running Chrome (Browserless, Steel, your own) instead

of launching locally:

kreuzcrawl scrape https://example.com \
  --browser-mode always \
  --browser-endpoint ws://browser.internal:9222/devtools/browser/<id> \
  --format markdown

The endpoint must be a WebSocket URL — ws:// or wss://. The CLI

rejects anything else with a clear error.

Use external CDP when:

  • You are running in containers or CI without a local Chrome.
  • You want a shared, warm browser pool across many crawl jobs.
  • You need browser-side residential proxies or stealth configuration the

local Chrome cannot provide.

Performance cost

Headless Chrome is expensive relative to a static fetch:

  • Cold start: 1-3 seconds the first time it launches.
  • Per-page overhead: 500 ms-2 s for NetworkIdle wait, plus the page's own

JS load time.

  • Memory: each tab takes 100-300 MB; long crawls should bound

--concurrent.

Mitigations:

  • Stay in --browser-mode auto — the engine only pays the cost when it

needs to.

  • Use --browser-endpoint to share one warm browser across jobs.
  • Drop --concurrent when you know the crawl will route through Chrome.

Wait strategies

Pass via --config JSON when you need control:

kreuzcrawl scrape https://example.com --browser-mode always \
  --config '{"browser":{"wait_strategy":{"type":"Selector","selector":".article-body"}}}'

Supported strategies:

  • NetworkIdle (default) — wait until the network goes quiet.
  • Selector — wait until a CSS selector resolves.
  • Fixed — wait a fixed duration.

extra_wait adds milliseconds on top of the wait strategy if the page

keeps loading content after the primary signal.

Persistent profiles

kreuzcrawl scrape https://app.example.com --browser-mode always \
  --config '{"browser_profile":"prod","save_browser_profile":true}'

Profile names are path-traversal-validated. Use them to keep cookies,

localStorage, and login state across runs without re-authenticating.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/kreuzberg-dev/plugins/plugins/kreuzcrawl/skills/headless-fallback/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能

同名技能的其他版本

有 2 个不同仓库或目录里都有叫 headless-fallback 的技能。它们内容并不相同,别混用: