跳到主要内容
知仓学习社ZHICANG

llm-crawler-access-check

Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copi…

不碰外部(只输出文字)无严重或高危命中davepoon/buildwithclaude

它会碰到什么

扫了多少1 个文本文件,6 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

AI crawler access check

One wrong line in robots.txt removes a site from AI answers completely, and no

amount of content work can compensate. This check takes under a minute and

should run before any other AI-visibility work.

Scope

Read https://<domain>/robots.txt and nothing else. Do not crawl the site, do

not attempt to access disallowed paths, and do not bypass any access control.

This is a read of one public file.

Procedure

1. Fetch

Fetch https://<domain>/robots.txt.

  • 404 or empty - everything is allowed by default. Say so; that is a valid

and often correct configuration. Stop and report.

  • Non-200 other than 404, or unreachable - report the status code and stop.

Do not guess at contents.

  • Served as HTML (a soft 404 returning the site's error page) - flag it.

Crawlers may parse it as garbage. This is itself a finding.

2. Resolve each agent

For each agent below, apply standard robots.txt matching: the most specific

User-agent group that names the agent wins, and * applies only when no group

names it. Within the winning group, the longest matching path rule wins, and

Allow beats Disallow on an equal-length match.

| Agent | Operator | Purpose | What blocking it actually costs |

| --- | --- | --- | --- |

| OAI-SearchBot | OpenAI | search index | citations in ChatGPT Search |

| ChatGPT-User | OpenAI | live fetch during a chat | the model cannot open your page when a user asks about it |

| GPTBot | OpenAI | training | background model knowledge, not search citations |

| PerplexityBot | Perplexity | search index | Perplexity citations |

| Perplexity-User | Perplexity | live fetch during a query | live page reads |

| ClaudeBot | Anthropic | index and training | Anthropic-side retrieval |

| Googlebot | Google | main index | AI Overviews and AI Mode, plus normal search |

| Google-Extended | Google | Gemini grounding and training | Gemini grounding only - not AI Overviews |

| Bingbot | Microsoft | Bing index | Microsoft Copilot, which rides the Bing index |

| Applebot | Apple | index | Apple search surfaces |

| Applebot-Extended | Apple | training | Apple Intelligence training only |

| CCBot | Common Crawl | open crawl corpus | an input to many downstream models |

Crawler names change and new ones appear. Before finalizing, check each

operator's own published crawler documentation for agents added or renamed since

this list was written, and include them. State which list you used.

3. Report

Produce a table with one row per agent and exactly these columns:

Agent | Verdict (ALLOWED / BLOCKED / PARTIAL) | Rule responsible | Impact

  • Rule responsible must quote the literal line from robots.txt, or say

no matching rule - allowed by default. Never state a verdict without the

line that produced it.

  • PARTIAL means important paths are disallowed while the site root is

allowed. Name the disallowed paths.

Then give:

  • Verdict - one sentence: is this site reachable by AI answer engines, or not?
  • What to change - the exact robots.txt lines to add, remove, or edit,

as a code block the user can paste. If nothing needs to change, say that

plainly rather than inventing work.

  • What this check did not cover - robots.txt is only the first gate.

Server-side blocking by WAF, CDN bot rules, IP reputation, or Cloudflare bot

management can block a crawler that robots.txt allows, and none of that is

visible in this file. Say so every time.

Three mistakes this check exists to catch

  1. **Blocking GPTBot to opt out of training, and assuming that is the whole

story.** It is not. OAI-SearchBot governs whether a site can be cited in

ChatGPT Search, and it is a separate agent with a separate rule. Blocking one

does not block the other, in either direction.

  1. Blocking Google-Extended to stay out of AI Overviews. It does not do

that. AI Overviews and AI Mode are built on the normal Googlebot index.

Blocking Google-Extended opts out of Gemini grounding and training and has

no effect on AI Overviews. To leave AI Overviews, the mechanism is the

nosnippet, max-snippet, or data-nosnippet family, and it costs normal

search snippets too. Say that tradeoff out loud rather than letting the user

discover it later.

  1. **A blanket User-agent: * / Disallow: / inherited from a staging config,

a bot-mitigation template, or a security hardening guide.** This is common

and almost always unintentional on a production marketing site.

If the user asks whether they should block AI crawlers

Do not answer with a recommendation. Lay out the tradeoff and let them decide:

allowing search crawlers is what makes citation possible, allowing training

crawlers affects model knowledge but not citation, and the two decisions are

independent. Publishers with a licensing position and companies that want to be

recommended by AI assistants land in different places, and both are legitimate.


About

Maintained by MaxAEO — <https://maxaeo.ai> — which works on AI answer-engine

visibility. The crawler matrix used here is kept current against each

operator's own published crawler documentation; where an agent has no official

documentation, this skill says so rather than guessing.

This check is free, read-only, and runs on one public file. It does not require

an account, an API key, or any paid service.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 3,470
本站分层T1
该仓技能数381
原文件路径plugins/all-skills/skills/llm-crawler-access-check/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 381 个技能