跳到主要内容
知仓学习社ZHICANG

ultimate-browsing

Renders, drives, and screenshots web pages: JS-rendered sources, clicks and forms, persistent logins, WAF-blocked hosts (platform-native readers, st…

执行命令读凭据联网写文件读文件严重 0 · 高危 16code-yeongyu/oh-my-openagent

它会碰到什么

扫了多少53 个文本文件,236 KB
它会碰到什么执行命令读凭据联网写文件读文件
命中总数182 处
命中统计严重 0 · 高 16 · 中 15 · 低 39

这个仓库里自带 1 个测试样本文件(有些技能仓会放故意的恶意样本做演示),它们不计入上面的能力与命中。

逐条看命中(16 条严重或高危)
  • engine/executor.py:73exec-spawn
    proc = subprocess.run(
  • engine/fetch_chain.py:170cred-envread
    _jmin = int(os.environ.get("INSANE_JITTER_MS_MIN", "150"))
  • engine/fetch_chain.py:171cred-envread
    _jmax = int(os.environ.get("INSANE_JITTER_MS_MAX", "400"))
  • engine/surrogate.py:117cred-envread
    if required_env and not os.environ.get(str(required_env)):
  • engine/tests/test_playwright_stealth.py:54cred-envread
    fs.writeFileSync(process.env.STEALTH_RECEIPT, JSON.stringify(events));
  • engine/tests/test_playwright_stealth.py:64cred-envread
    env["STEALTH_RECEIPT"] = str(receipt)
  • engine/tests/test_playwright_stealth.py:67exec-spawn
    result = subprocess.run(
  • engine/tests/test_playwright_templates.py:30cred-envread
    if (process.env.PW_FAKE_SELECTOR_FAIL === '1') {
  • engine/tests/test_playwright_templates.py:43cred-envread
    if (process.env.PW_FAKE_CLOSE_FAIL === '1') {
  • engine/tests/test_playwright_templates.py:136exec-spawn
    return subprocess.run(
  • engine/tests/test_surrogate.py:134cred-envread
    os.environ["OMOB_TEST_TOKEN"] = "secret"
  • scripts/cookie_crypto.py:49exec-spawn
    result = subprocess.run(
  • scripts/cookie_paths.py:68cred-envread
    return Path(os.environ.get("XDG_CONFIG_HOME", str(home / ".config")))
  • scripts/cookie_paths.py:72cred-envread
    return Path(os.environ.get("LOCALAPPDATA", str(home / "AppData" / "Local")))
  • scripts/extract_cookies.py:248exec-spawn
    proc = subprocess.run(
  • scripts/tests/test_extract_cookies.py:232exec-spawn
    with patch("extract_cookies.subprocess.run") as run:

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Ultimate Browsing

Web access for everything a plain fetch cannot finish: a page that renders in JS, a click or a form, a screenshot, a login that must persist across pages, or a host that blocks generic fetchers (WAF / 403 / Cloudflare). Start at the cheapest tier that can do the job and climb only when it cannot:

Tier 1 — insane-search (headless extraction + WAF bypass) -> Tier 1.5 — agent-reach (platform-native APIs, esp. Chinese platforms) -> Tier 2 — a real browser: 2a Bun.WebView, 2b a local-Chrome playwright-core script from js eval for Chrome semantics, stealth, trace, or auth.

PHASE 0 — ROUTE FIRST (MANDATORY)

User request
  |
  +- extract text/data from a URL --------------------- TIER 1  insane-search
  +- URL blocked / 403 / Cloudflare / WAF ------------- TIER 1  insane-search
  +- YouTube/Vimeo/TikTok subtitles or metadata ------- TIER 1  insane-search (yt-dlp)
  +- read an article / blog / Reddit / HN / arXiv ----- TIER 1  insane-search
  |
  +- Chinese platform (xhs/douyin/weibo/bilibili/v2ex/wechat)  TIER 1.5 agent-reach
  +- podcast transcript / stock forum ----------------- TIER 1.5 agent-reach
  +- Twitter feed / LinkedIn profile / GitHub via CLI - TIER 1.5 agent-reach
  |
  +- Tier 1/1.5 returned empty or partial ------------- TIER 2  2a kernel browser -> 2b stealth
  +- click / fill form / scroll / interact ------------ TIER 2  2a kernel browser -> 2b stealth
  +- screenshot / render / play video ----------------- TIER 2  2a kernel browser -> 2b stealth
  +- login session across pages / inject cookies ------ TIER 2  2b Chrome stealth (profile + cookies)
  +- test web app / QA / dogfood ---------------------- TIER 2  2a kernel browser -> 2b stealth
  |
  +- simple search query ------------------------------ NOT this skill (use web-search)

Read the matching reference before acting: [references/insane-search/README.md](references/insane-search/README.md), [references/agent-reach/README.md](references/agent-reach/README.md), or [references/chrome-stealth.md](references/chrome-stealth.md).

Tier 1 — insane-search (headless extraction)

When: content extraction, blocked-URL bypass, media metadata — no browser UI needed.

Why first: ~10x faster than a browser, no process spin-up; handles most "fetch this blocked page" requests via curl_cffi TLS impersonation, yt-dlp (1858 sites), official public APIs, mobile URL transforms, Phase-2.5 surrogate archives (Wayback / archive.today snapshots, provenance-tagged — see [references/insane-search/cache-archive.md](references/insane-search/cache-archive.md)), a key-gated Jina Reader (JINA_API_KEY), and a Playwright real-Chrome fallback. The engine lives inside this skill at engine/ and is invoked as a module. Surrogate results are dated COPIES: a result whose provenance is snapshot must be reported with its snapshot_timestamp, never presented as the live page.

# Core command — auto-detects WAF, runs the full fetch grid (run from the skill dir):
python3 -m engine "https://example.com/blocked-page"
#   add --selector "<CSS>" for positive-proof validation, --device auto|desktop|mobile,
#   --trace to inspect every attempt, --json for machine-readable output.

# YouTube subtitles / metadata (no browser):
yt-dlp --write-sub --write-auto-sub --sub-lang "en,ko" --skip-download -o "/tmp/%(id)s" "<URL>"

# Reddit / HN / Bluesky / arXiv etc. use official public endpoints — see the Phase 0 index in
# references/insane-search/README.md (Twitter syndication, Reddit .json, HN Firebase, ...).

The full engine harness (rules R1-R7, the Phase 0 official-API index, the no-site-name rule, and the references/insane-search/*.md deep-dives for TLS, Playwright routing, Naver, media, etc.) is in [references/insane-search/README.md](references/insane-search/README.md). Read it before tuning the engine or adding a WAF profile.

Escalate to Tier 1.5 or Tier 2 when

  • The target is a Chinese / social platform with a native reader -> Tier 1.5.
  • insane-search returns empty/partial, or the page needs JS interaction, a screenshot, a persistent login, or media playback -> Tier 2.

Tier 1.5 — agent-reach (platform-native readers)

When: the target is a platform with a first-class API/CLI that beats generic fetching — especially Chinese platforms that stealth browsers still cannot reach cleanly. Several channels are zero-config (Douyin, V2EX, Reddit, RSS, YouTube); others need a one-time auth you supply via environment variables if you have access (JINA_API_KEY for Jina Reader — anonymous access is dead, see references/insane-search/jina.md; TWITTER_* for X; a transcription key for podcasts).

| Category | Platforms | Entry |

|---|---|---|

| social | xhs (Xiaohongshu), douyin, weibo, bilibili, V2EX, Reddit, Twitter/X | [references/agent-reach/social.md](references/agent-reach/social.md) |

| web | Jina Reader, WeChat articles, RSS | [references/agent-reach/web.md](references/agent-reach/web.md) |

| video | YouTube, Bilibili, podcast transcripts, Douyin video | [references/agent-reach/video.md](references/agent-reach/video.md) |

| career | LinkedIn | [references/agent-reach/career.md](references/agent-reach/career.md) |

| dev | GitHub (gh CLI) | [references/agent-reach/dev.md](references/agent-reach/dev.md) |

| search | Exa AI | [references/agent-reach/search.md](references/agent-reach/search.md) |

mcporter call 'douyin.parse_douyin_video_info(url: "<URL>")'   # douyin, zero-config
curl -s "https://r.jina.ai/https://weibo.com/<uid>/<pid>"      # weibo via Jina
yt-dlp --dump-json "<bilibili-url>"                            # Bilibili (overseas: add --cookies-from-browser)
curl -s "https://www.v2ex.com/api/topics/hot.json"            # V2EX public API

Routing table, per-platform auth (set TWITTER_* env vars, gh auth login, a transcription key — only if you have access), rate-limit notes, and known version quirks are in [references/agent-reach/README.md](references/agent-reach/README.md).

Tier 2 — a real browser (real interaction)

When: real interaction is needed (clicks, forms, screenshots, video, persistent login), or Tier 1/1.5 failed.

Tier 2a — kernel browser (default)

Use new Bun.WebView() from the js-eval kernel on Bun >= 1.4: macOS defaults to system WebKit; Linux/Windows need installed Chrome/Chromium/Edge. WebView is headless, WebKit has no CDP, and type() emits no keyboard events. For other kernels or when those differences matter, use 2b.

const view = new Bun.WebView({ width: 1280, height: 800 })
try {
  await view.navigate(url)
  const title = await view.evaluate("document.title")
  await Bun.write(pngPath, await view.screenshot())
} finally {
  view[Symbol.dispose]()
}

Use 2b for real-Chrome semantics, stealth, trace, authenticated profiles, or a page the kernel browser cannot reach.

Tier 2b — Chrome stealth (blocked or logged-in pages)

WRITE a playwright-core script and run it from js eval against installed local Chrome: chromium.launch({ channel: "chrome" }), or launchPersistentContext on a task-owned profile. For authenticated state, CLONE the user's profile first (rsync -a <profile>/ <tmp-clone>/); NEVER launch against or clear cookies/cache/site data from the live profile. Codex: prefer browser:control-in-app-browser for ordinary page control.

Keep the engine's Playwright templates for script-based extraction. Stealth is optional: the user installs playwright-extra + puppeteer-extra-plugin-stealth once in the engine directory and the script wraps the playwright-core browser type. Setup, persistent-context arguments, screenshots, and cleanup are in [references/chrome-stealth.md](references/chrome-stealth.md). A stealth flag is not proof of access: inspect the rendered result and report challenges that remain.

Cookie login (cross-platform)

scripts/extract_cookies.py reads cookies from a local Chromium-family or Firefox-family browser and optionally injects them into the running CDP session. It resolves browser profile paths and decrypts cookie values per-OS (macOS Keychain, Linux libsecret, Windows DPAPI):

# Extract cookies to a file:
mkdir -p ~/.local/state/omo-cookies
python3 scripts/extract_cookies.py --browser chrome --domain youtube.com --output ~/.local/state/omo-cookies/youtube.cookies.json
# Extract and inject into the running CDP session:
python3 scripts/extract_cookies.py --browser chrome --domain youtube.com --inject --cdp 9242

Cookie export files are written with owner-only 0600 permissions. Do not place live auth cookies in shared temp directories or commit them to a repo. Cookie injection sends values to CDP over stdin rather than argv. Cookies apply on next navigation — reload after injecting. Google services use fingerprint-bound tokens that may not transfer across browser profiles. Full detail in [references/chrome-stealth.md](references/chrome-stealth.md).

Reference docs

| File | When to read |

|------|-------------|

| [references/insane-search/README.md](references/insane-search/README.md) | Tier-1 engine harness (R1-R7, Phase 0 API index, no-site-name rule) + its *.md deep-dives |

| [references/agent-reach/README.md](references/agent-reach/README.md) | Tier-1.5 routing table, platform auth, per-category *.md |

| [references/chrome-stealth.md](references/chrome-stealth.md) | Tier-2 playwright-core scripts, optional stealth setup, cloned profiles, cookie login |

Environment variables

# agent-reach auth: set the channel-specific env vars from each tool's docs only if you have access
# insane-search needs no env vars — it auto-installs deps on first run

Anti-patterns

  • Do NOT launch Chrome stealth for plain text extraction — use Tier 1.
  • Use stealth plugins only in an explicitly installed script environment, not injected into WebView.
  • Close every WebView/browser context when done and remove only task-owned profile clones.
  • Do NOT inject cookies without reloading the page.
  • Do NOT hardcode site domains/selectors into engine/** or waf_profiles.yaml — runtime hints only (see the no-site-name rule in the insane-search reference).

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 69,100
本站分层T1
该仓技能数62
原文件路径packages/shared-skills/skills/ultimate-browsing/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 62 个技能