A teal scanning beam passing over stacked navy document pages, ticking their surfaces, while a coral flaw sits inside one page out of reach
This content was generated using AI.

If a docs platform has sent you an AI agent readiness score, or you’ve run one yourself, you are probably wondering how much weight to give the grade. Our view: run it, fix what it finds, and read the grade for what it is. It tells you whether agents can fetch and read your docs. It says nothing about whether what they read is correct.

This is also where our agent-ready docs checklist starts. Its opening step is a baseline scan with npx afdocs check <docs URL>, and most of the checklist’s FIND and READ checks map to things these scanners test.

Three scanners, one spec

Three free scanners get most of the attention:

They look like competing products, but they share a foundation: the open Agent-Friendly Documentation Spec, authored by “Dachary Carey + community contributors” (SPEC.md). Mintlify says its score “is built on the Agent-Friendly Documentation Spec, an open standard that defines what good documentation looks like for agents.” Fern’s page says it is “Built on the open-source Agent-Friendly Docs Spec”. The spec’s README calls AFDocs “a companion CLI tool and Node.js library that implements this spec”, and the AFDocs README describes it as “Powering Agent Score by Fern.”

The spec’s scope matters. It targets “coding agents that fetch documentation during real-time development workflows”, naming Claude Code, Cursor and GitHub Copilot. It explicitly does not target training crawlers, answer engines such as Perplexity or Google AI Overviews, or RAG pipelines that pre-index docs (SPEC.md). So a good score does not mean your docs will appear in AI search answers. That isn’t what it measures.

What the checks actually test

The spec groups its checks into seven categories (agent-docs-spec README). Mintlify’s post frames each one as a question, which is a useful way to read them:

  • Content discoverability. “Can agents find your docs?” Most of the llms.txt checks live here: does the file exist, is it valid, is it small enough, do its links resolve and point to markdown. If you want an llms.txt checker, this category is it. We cover the file itself in llms.txt for developer documentation.
  • Markdown availability. “Can agents get clean Markdown?” Do .md URLs work, and does the server honour content negotiation via Accept headers.
  • Page size and truncation risk. “Are pages a reasonable size or will agents truncate them?” Is content rendered server-side, how big is the page, and where does real content start after the navigation chrome. The spec passes pages under 50,000 characters and fails those over 100,000 characters, which “will be truncated by Claude Code and likely all other platforms” (SPEC.md). Its own example: “A 4.56MB file means the agent sees roughly 2% of it.”
  • Content structure. “Is the content well-formed for agents?” Do tabs serialise into something sensible, are code fences closed (markdown-code-fence-validity), are links portable.
  • URL stability. “Are URLs going to help or hinder agents?” Soft 404s and redirect behaviour.
  • Observability. “Do resources stay accurate over time?” llms.txt coverage of the site, parity between HTML and markdown versions, cache headers.
  • Authentication. “Can agents reach your docs?” Is content behind a login, is there an alternative path, does bot protection block automated fetching.

Why the numbers don’t match

Compare tools and you’ll see different counts. As of writing, the spec repository is at draft v0.6.0, dated 2026-09-13, and defines “28 checks across 7 categories” (agent-docs-spec README). AFDocs says it implements that version. Fern’s page says 22 checks. Mintlify’s April 27, 2026 post says 29, and explains: “We added checks on top based on how we observe agents using documentation.”

The spec is a draft that has changed between versions, and vendors update on their own schedules. The practical point is that roughly two dozen checks, in the same seven areas, sit behind every one of these scores.

What an A grade does not tell you

The spec is candid about its limits. In its own words, it “focuses on meeting the technical constraints of agent platforms (truncation limits, content negotiation, discovery); it does not consider qualitative evaluation of content” (agent-docs-spec README).

That leaves a lot of room. A page can pass every size and markdown check while its procedure skips a permission the user needs. A code sample can sit in a perfectly closed fence and still be missing its imports. An OpenAPI spec can be published, linked and fetchable while describing an API that changed last quarter. An agent has no way to tell any of these apart from good docs. It will follow them confidently.

Fern labels grade A “Agent-Ready”. We’d read that more narrowly: ready to be read, not yet shown to be trustworthy. Mintlify makes a fair point in the other direction: “A low score isn’t a failure. Most documentation wasn’t written with agents in mind, because agents weren’t a primary audience until recently.” The mirror is also true. A high score isn’t a pass.

This is how the scanners line up with our checklist. The FIND and READ layers (access, stable URLs, llms.txt, rendering, markdown, page size, tabs) are largely what these tools test. The UNDERSTAND layer (nothing left to guess, complete code samples, consistent terminology, frontmatter) and the TRUST layer (procedures still work, the spec matches reality, tested samples, version labels) need human review or tests against the product. For the API side of that work, see what agent-ready API docs require. If you want agents to query your docs directly rather than fetch pages, documentation MCP servers are the next step.

How to run a baseline

This is the part to forward to an engineer. From the AFDocs README:

npx afdocs check https://docs.example.com --format scorecard
  • The scorecard gives an overall score and letter grade, per-category scores, and per-check pass, warn and fail results with suggested fixes.
  • Run it before you change anything so you have a baseline, then again after each fix.
  • If you track the score in CI, pin the version. The README warns that while the tool is in early development (0.x), “Check IDs, CLI flags, and output formats may change between minor versions.”

Mintlify and Fern offer the same kind of scan through a web form if you’d rather not run a CLI.

Run it, then do the work it can’t see

Run the score, fix what it finds, then turn to what it cannot see: whether the content is complete, unambiguous and still true. That is the argument of our pillar post, writing docs for humans and AI agents: the same qualities serve both readers.

Get the full checklist

This baseline scan is where our agent-ready docs checklist starts, before its four layers: Find, Read, Understand and Trust. The scanners cover much of the first two; the remaining checks cover what scanners miss, across all 16 checks. Get the free checklist.

Frequently asked questions

  • No. The Agent-Friendly Documentation Spec behind these scores targets coding agents that fetch documentation during real-time development workflows, naming Claude Code, Cursor and GitHub Copilot. It explicitly does not target training crawlers, answer engines such as Perplexity or Google AI Overviews, or RAG pipelines that pre-index docs. A good score tells you agents can fetch and read your docs, not that they will surface in AI search.

  • The shared spec is a draft that has changed between versions, and vendors update on their own schedules. Draft v0.6.0, dated 13 September 2026, defines 28 checks across seven categories, and AFDocs says it implements that version. Fern's page says 22 checks. Mintlify's April 2026 post says 29, because it added checks based on how it observes agents using documentation. All sit in the same seven areas.

  • The spec says it focuses on technical constraints such as truncation limits, content negotiation and discovery, and does not consider qualitative evaluation of content. So a page can pass every check while its procedure skips a needed permission, a code sample can sit in a closed fence but lack its imports, and a published OpenAPI spec can describe an API that has since changed. Those need human review or tests against the product.