跳到主要内容
知仓学习社ZHICANG

headless-fallback

>-

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少1 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

<!--

AI-RULEZ :: GENERATED FILE — DO NOT EDIT

Content-Hash: blake3:9ce8ead27ee0f0b76d36ce4a455a5d17936077d6582e737c9b9a7ba46819b686

Source-Hash: blake3:b5689383a3914da8e3cc6ad3614356ae7f98c90cf99cb1b4b47be38456ff7a7c

Schema-Version: v1

-->

Headless fallback

Some pages are unscrapable without a real browser — SPA shells, infinite

scroll, Cloudflare interstitials, JS-rendered article bodies. Crawlberg

ships with an optional headless-Chrome backend driven by chromiumoxide.

Modes

--browser-mode auto    # default — try static first, fall back to browser on JS/WAF
--browser-mode always  # skip static, go straight to browser
--browser-mode never   # static only, fail closed

auto (default)

The engine fetches statically, then inspects the response. It launches

headless Chrome and re-fetches when it sees:

  • WAF responses from one of 8 detected vendor fingerprints (Cloudflare,

Akamai, AWS WAF, Imperva, DataDome, PerimeterX, F5, plus a generic

catch-all).

  • SPA shells: <noscript> warnings, near-empty <body> with heavy JS.
  • Heuristic JS-render-required signals.

This is the right default. The browser only spins up when needed.

always

Skip the static probe entirely. Use when:

  • The user already told you the page needs JS.
  • You are scraping a site you know is React/Vue/Svelte SPA.
  • You need <script>-emitted state that never lands in static HTML.
crawlberg scrape https://spa.example.com --browser-mode always --format markdown

never

Static only — the browser path is disabled. Use when:

  • You are in a hot loop where a stray Chrome launch would blow the budget.
  • You are running in a sandbox without a Chrome binary.
  • The user explicitly wants only static fetches.

In never mode, JS-only pages return empty/stub content. Inspect

markdown.content and markdown.warnings before treating the result as

final.

Symptoms that point to headless

In --browser-mode never or when you suspect the auto detector missed a

signal:

  • markdown.content is short, nav-only, or just a loading message.
  • status_code is 200 but metadata.headings is empty on a page that

clearly has headings.

  • markdown.warnings mentions JS-render-required or WAF detection.
  • 403/406/503 with WAF response headers (server: cloudflare,

cf-mitigated, x-amz-cf-id, set-cookie: __cf_bm=…).

Re-run with --browser-mode always. If that succeeds, leave it set for

that host.

External CDP endpoint

Point at an already-running Chrome (Browserless, Steel, your own) instead

of launching locally:

crawlberg scrape https://example.com \
  --browser-mode always \
  --browser-endpoint ws://browser.internal:9222/devtools/browser/<id> \
  --format markdown

The endpoint must be a WebSocket URL — ws:// or wss://. The CLI

rejects anything else with a clear error.

Use external CDP when:

  • You are running in containers or CI without a local Chrome.
  • You want a shared, warm browser pool across many crawl jobs.
  • You need browser-side residential proxies or stealth configuration the

local Chrome cannot provide.

Performance cost

Headless Chrome is expensive relative to a static fetch:

  • Cold start: 1-3 seconds the first time it launches.
  • Per-page overhead: 500 ms-2 s for NetworkIdle wait, plus the page's own

JS load time.

  • Memory: each tab takes 100-300 MB; long crawls should bound

--concurrent.

Mitigations:

  • Stay in --browser-mode auto — the engine only pays the cost when it

needs to.

  • Use --browser-endpoint to share one warm browser across jobs.
  • Drop --concurrent when you know the crawl will route through Chrome.

Wait strategies

Pass via --config JSON when you need control:

crawlberg scrape https://example.com --browser-mode always \
  --config '{"browser":{"wait":"selector","wait_selector":".article-body"}}'

Supported strategies (the browser.wait field is a string enum; pair

"selector" with a sibling wait_selector):

  • network_idle (default) — wait until the network goes quiet.
  • selector — wait until the CSS selector in wait_selector resolves.
  • fixed — wait a fixed duration.

extra_wait adds milliseconds on top of the wait strategy if the page

keeps loading content after the primary signal.

Persistent profiles

crawlberg scrape https://app.example.com --browser-mode always \
  --config '{"browser_profile":"prod","save_browser_profile":true}'

Profile names are path-traversal-validated. Use them to keep cookies,

localStorage, and login state across runs without re-authenticating.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/kreuzberg-dev/plugins/plugins/crawlberg/skills/headless-fallback/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能

同名技能的其他版本

有 2 个不同仓库或目录里都有叫 headless-fallback 的技能。它们内容并不相同,别混用: