Free tool · No sign-up
Which AI crawlers can read your site?
Most site owners have never checked, and a good number have blocked the wrong ones. This reads your robots.txt and tells you not just what is allowed, but what each decision actually costs you.
The distinction that matters
Training crawlers and answer crawlers are not the same thing
Blocking GPTBot keeps your content out of OpenAI's training data. It does not stop ChatGPT citing you — that is OAI-SearchBot. Plenty of sites have blocked the first believing they did the second, and quietly removed themselves from nothing at all. Others blocked the answer crawler and wondered why they vanished from AI results.
The same confusion surrounds Google-Extended, which governs Gemini training and grounding and has no bearing whatsoever on your Google Search ranking. Googlebot is the one that matters there, and it is also the crawler behind AI Overviews.
Most tools in this space print a list of user agents and a tick or a cross. The column worth reading is the one that says what happens next.
While we are here
On llms.txt
There is a lot of advice telling you to add an llms.txt file. The measured
evidence does not support it: across more than 500 million AI bot visits in a 90-day
window, 408 fetched the file. Google's 2026 guidance says plainly it is
not needed for AI Overviews or AI Mode, and no major model provider has committed to
reading it.
It does do real work as a routing map for AI coding agents — Cursor, Claude Code, Copilot and similar. That is a decent reason to have one if you publish developer documentation. It is not a reason to expect more AI search visibility, and we would rather say so than sell you a file.
AI crawlers · Questions
What people get wrong about this
What does this tool check?
It reads your robots.txt and works out whether each major AI crawler is allowed to fetch your site, using the same grouping and longest-match rules crawlers use. It then explains what blocking each one actually does, which is the part most owners get wrong.
Does blocking GPTBot stop ChatGPT from citing my site?
No, and this is the single most common mistake. GPTBot collects training data. The crawler that decides whether ChatGPT can cite you in search results is OAI-SearchBot, and ChatGPT-User is the one that fetches a page when someone asks about it directly. Blocking GPTBot protects your content from training while leaving you visible in answers — which is what most businesses actually want.
Does blocking Google-Extended hurt my Google ranking?
No. Google-Extended controls Gemini training and grounding only. It has no effect on Google Search ranking at all. Googlebot is the crawler behind both classic Search and AI Overviews, so blocking that one would remove you from Google entirely.
Should I create an llms.txt file?
Probably not for AI search. Across a 90-day window of more than 500 million AI bot visits, only 408 fetched llms.txt — GPTBot, ClaudeBot, PerplexityBot and Google-Extended overwhelmingly skip it and crawl HTML directly. Google stated in 2026 that it is not needed for AI Overviews or AI Mode, and no major model provider has committed to using it. It does have a genuine use as a routing map for AI coding agents such as Cursor, Claude Code and Copilot. Just do not expect it to change your visibility in AI answers.
I have no robots.txt. Is that bad?
Not in itself. No robots.txt means nothing is disallowed, so every crawler is allowed. That is a permissive default rather than a broken one. It only becomes a problem if you intended to restrict something.
Should I block AI training crawlers?
It depends on what your content is worth to you and it is a business decision rather than a technical default. Publishers with original research often block training while allowing answer crawlers. A services business usually wants maximum visibility and blocks nothing. What matters is that the decision is deliberate — most sites have simply never looked.
Want this set deliberately?
Which AI systems may read your content, and which may train on it, is a business decision. Most sites have simply never made it. We will audit what you are allowing today and set it to match what you actually want.