A navy building with a closed teal gate on one wing and an open archway on the docs wing, where small geometric agents walk inside.
This content was generated using AI.

Blocking AI crawlers has become a default setting. In its July 2025 announcement on AI crawler permissions, Cloudflare said more than one million customers had chosen to block AI crawlers since the option appeared in September 2024, and that every new domain would now be asked whether to allow them.

Most of those decisions were made to protect marketing pages, blogs and paid content. The docs usually sit on the same domain and inherit the same rules, without anyone deciding they should. That is a problem, because documentation is the one part of your site you actively want AI agents to read. When a customer asks Claude Code, Cursor or Copilot to integrate your API, the agent fetches your docs. If it gets a block or a challenge page, it doesn’t raise an error the developer notices. It falls back on what it already knows, and writes code against whatever version of your API it saw in training.

Why so many sites now block AI crawlers

The tooling makes blocking easy. Cloudflare’s managed robots.txt disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. Its AI bot policies, which replace the legacy “Block AI bots” setting (deprecated on 15 September 2026), offer “Block (on all pages)” as a choice for Search, Agent and Training bots.

“All pages” includes /docs. And “Agent” covers bots doing real-time work for a user, such as chat bots, which is where a developer’s coding assistant fits. For marketing content, blocking is often the right call. For docs, it deserves a separate decision.

Not all AI bots are the same bot

This is the part that is easy to miss. AI companies run several bots with different jobs.

Anthropic documents three agents: ClaudeBot collects content that could contribute to model training, Claude-User accesses websites when individuals ask Claude questions, and Claude-SearchBot improves search result quality. Anthropic is explicit that disabling Claude-User “prevents our system from retrieving your content in response to a user query”.

OpenAI splits its crawlers the same way: GPTBot for training, OAI-SearchBot for ChatGPT’s search features, and ChatGPT-User for user actions. For ChatGPT-User, OpenAI notes that “because these actions are initiated by a user, robots.txt rules may not apply”.

So a rule against training bots is a content decision you can make deliberately. A blanket rule, or a copied “block all AI” list that includes Claude-User or ChatGPT-User, blocks the developer’s own agent at the moment it is trying to use your product. There is a second trade-off worth weighing: blocking training bots from your docs means future models learn less about your current API. Some teams will accept that. It should be a choice, not a side effect.

Also remember what robots.txt is. Google describes it as a way to manage crawler traffic that “is not a mechanism for keeping a web page out of Google”, and Cloudflare’s own docs say “robots.txt compliance is voluntary”.

Edge bot controls are the bigger hazard

robots.txt is at least readable. Edge controls are silent. Bot challenges, WAF rules and features like Cloudflare’s AI Labyrinth, which adds invisible links with nofollow tags so that non-compliant crawlers “will be stuck in a maze of never-ending links”, change what an automated client receives. The agent gets a 403, or a challenge page returned with a 200, and no docs.

We hit this on our own site. Google’s OAuth brand verification for weesholapara.com kept failing with “your home page is behind a login page”. The cause was Cloudflare bot controls we had enabled, AI Labyrinth among them: Google’s checker was served a challenge page and read it as a login. Disabling those controls fixed it. If Google’s own verifier couldn’t tell a challenge page from a login wall, a coding agent won’t either.

The Agent-Friendly Documentation Spec puts it bluntly: a docs site that returns “login pages, 401/403 responses, or SSO redirects” is “completely opaque to agents”, which then fall back on training data or secondary sources. It also notes that bot protection usually catches documentation “as collateral” without a deliberate decision.

If some docs must stay gated, the spec lists alternatives agents can still use: a public llms.txt, public markdown or an API endpoint, docs bundled with your SDK, or an MCP server with authentication handled server-side. We cover two of these in llms.txt for developer documentation and documentation MCP servers.

Stable URLs: what agents do with a moved page

Reachable also means reachable at the address the agent has. Agents arrive with URLs from training data, old llms.txt files and other people’s blog posts, and many of those addresses are stale.

Vercel’s December 2024 analysis, The rise of the AI crawler, found ChatGPT spent 34.82% of its fetches on 404 pages and Claude 34.16%, against 8.22% for Googlebot. ChatGPT also spent 14.36% of fetches following redirects, against 1.49% for Googlebot. That is crawler data from late 2024, not coding agent data, but the pattern is clear: AI clients hit dead and moved URLs far more often than search crawlers do.

Three rules follow.

Moved page: a 301 on the same host. Google uses a 301 as a “strong” signal that the redirect target should be processed. The spec’s redirect-behavior check passes same-host HTTP redirects and warns on cross-host ones, because Claude Code, for example, doesn’t automatically follow cross-host redirects. It fails JavaScript redirects outright, “because agents don’t execute JavaScript”. Moving docs from example.com/docs to docs.example.com is a cross-host redirect, so update links to point at the final address rather than relying on the hop.

Gone page: a real 404 or 410. Google defines a soft 404 as a page telling the user it doesn’t exist while returning a 200 (success) status. Its fix is 404 or 410 for removed content and 301 for relocated content.

Soft 404s hurt agents more than people. In the spec’s words, soft 404s “are worse than real 404s for agents. The agent sees a 200 and tries to extract information from the error page content rather than recognizing the page doesn’t exist.” A person reads “Page not found” and searches. An agent may summarise your 404 page’s navigation links as the answer.

How to check it

Forward this to whoever runs your docs.

Reachable, as an agent would see it:

curl -sI -A "Claude-User" https://docs.example.com/quickstart
curl -s https://example.com/robots.txt

Expect a 200 and real HTML, not a challenge page. In robots.txt, look for rules naming Claude-User or ChatGPT-User, or a User-agent: * block disallowing /docs. A single curl won’t reproduce every edge rule, and the spec warns that behavioural enforcement can pass a spot check and still challenge real multi-page agent sessions. So also ask a coding agent to fetch a few docs pages and quote a specific sentence from each.

Soft 404:

curl -s -o /dev/null -w "%{http_code}\n" https://docs.example.com/this-page-does-not-exist

Anything other than 404 or 410 is a soft 404.

Stable URLs: keep a plain text file of retired URLs (old paths, entries from your last llms.txt, links from popular external posts) and run:

lychee --files-from retired-urls.txt

The lychee CLI follows up to 10 redirects by default (--max-redirects) and can output json or markdown reports for CI. To run it on every pull request, see five CI workflows for documentation.

Decide per section, not per domain

Our view: protect what you sell, and open what helps people use what you sell. In Cloudflare terms that usually means scoping AI bot rules away from the docs path or hostname, keeping user-initiated agents allowed, and making the training question a deliberate one. The broader case for writing docs that serve both readers is in writing docs for humans and AI agents.

Get the full checklist

Reachable and Stable URLs are two of the 16 checks in our agent-ready docs checklist. Both sit in the Find layer and both are rated Critical. Get the free checklist.

Frequently asked questions

  • Yes. Anthropic and OpenAI run separate bots for separate jobs. ClaudeBot and GPTBot collect content for training, while Claude-User and ChatGPT-User fetch pages when a person asks a question. A rule against training bots is a content decision you can make deliberately. A blanket "block all AI" list also stops the developer's own agent. Keep in mind that blocking training bots from your docs means future models learn less about your current API.

  • Fetch a docs page with curl using the Claude-User user agent and expect a 200 with real HTML, not a challenge page. Check robots.txt for rules naming Claude-User or ChatGPT-User, or a User-agent: * block disallowing /docs. A single curl won't reproduce every edge rule, and behavioural enforcement can pass a spot check, so also ask a coding agent to fetch a few docs pages and quote a specific sentence from each.

  • A moved page should return a 301 on the same host, because some agents, Claude Code for example, don't automatically follow cross-host redirects. A removed page should return a real 404 or 410. Avoid soft 404s, where a "not found" page comes back with a 200 status. An agent sees the 200 and tries to extract information from the error page, and may summarise its navigation links as the answer.