How robots.txt rules are matched
A crawler follows only the group whose User-agent line best matches its name, and falls back to the User-agent: * group only when no specific group exists. That surprises people: if you add a GPTBot group, GPTBot ignores everything in your * group.
Inside the chosen group, the longest matching path wins. If an Allow and a Disallow rule match with the same length, Allow wins. An asterisk matches any characters and a dollar sign anchors the end of the URL. This checker applies those rules exactly, so what you see is what Googlebot and well-behaved AI crawlers do.
AI training vs AI search crawlers
AI companies now run separate crawlers for separate jobs. GPTBot, ClaudeBot and Google-Extended collect content for model training. OAI-SearchBot, Claude-SearchBot and PerplexityBot build the indexes that AI answers cite. ChatGPT-User and Claude-User fetch a page live when a person asks about it.
Blocking all of them keeps you out of AI answers entirely. Many sites choose the middle path: block training crawlers, allow search and user-triggered ones so the brand still shows up and gets cited. The generator's default preset does exactly that.
Common robots.txt mistakes
The most expensive one is a staging file shipped to production with Disallow: / under User-agent: *, which removes the whole site from search. Others: blocking CSS and JavaScript that Google needs to render pages, using robots.txt to hide pages that are linked elsewhere (use noindex instead), and forgetting the Sitemap line.
Frequently asked questions
How do I check if my robots.txt blocks a page?
Paste your robots.txt, enter the URL path and pick a crawler. The checker shows allowed or blocked and the exact rule that decided it.
Should I block GPTBot?
Block GPTBot if you do not want your content used for OpenAI model training. Blocking it does not remove you from ChatGPT search, which uses OAI-SearchBot and ChatGPT-User.
Does robots.txt remove a page from Google?
No. It stops crawling, not indexing. A blocked page can still appear in results if other sites link to it. Use a noindex tag to keep a page out of the index.
Does Google support crawl-delay?
No. Googlebot ignores crawl-delay. Some other crawlers respect it.