Document pages on a teal conveyor pass four navy checkpoint gates, with one page set aside under a small coral flag
This content was generated using AI.

The useful question about AI generated documentation isn’t “can AI write docs?” It is “which parts can we hand over, and where does a person still sign off?” A new benchmark, DoGBench, gives a better answer than vendor pitches or dismissals do. Its paper describes it as “the first benchmark for generating and maintaining real user-facing software documentation” (DoGBench on arXiv, submitted 30 September 2026).

What DoGBench tests

The benchmark has 292 items drawn from open-source projects including Helm, PostHog and Mautic, according to the arXiv abstract. For each item, the DoGBench site says the agent “received a repository as it existed before the change and a triggering signal, such as a code pull request or a reported documentation gap.” It then has to decide whether the docs need updating and, if so, produce an acceptable patch in a single attempt.

The detail that matters most: 87 of the 292 items need no docs change at all, and 205 do (dogbench.ai). An agent that edits everything gets penalised. Patches are scored with rubrics the site says were “validated with project maintainers”: 3,273 criteria across the 205 items that need a change, 798 of them P0, the critical tier.

How to read the scores

The headline number is easy to misread, so here is how the scoring works, all from dogbench.ai:

  • A critical failure caps a patch at 60. In the site’s words, “A critical failure caps a patch’s score at 60 out of 100”, and failing any P0 criterion triggers the cap. A fluent patch with one critical error cannot score well.
  • The combined score is a harmonic mean. It combines delivered patch quality and abstention recall (correctly leaving docs alone), so strength on one can’t hide weakness on the other.
  • The best combined score in the paper is 47.3 out of 100. That is the top result among the seven agents the paper evaluated, on the 117-item held-out split (82 items needing an update, 35 needing none). It is a combined score, not a percentage of correct patches.
  • The highest P0-clean delivery rate was 39.0%. A patch is P0-clean when it has no critical failures. The Promptless launch post gives this as the best among the paper’s seven agents, and dogbench.ai puts it at 32 of the 82 held-out tasks that required an update. If someone asks “how often did the best agent ship a patch with no critical defect?”, this is the number.

The site’s leaderboard also shows 54.8, labelled the “Highest composite score held by a cloud agent”. That is a leaderboard entry, not one of the paper’s seven, and leaderboards move, so treat it as the current top of the board rather than the paper’s result.

Where agents fail

The authors audited 1,267 submissions. According to dogbench.ai, 45.5% “had a task-completion gap”, 36.6% “contained technical inaccuracies” and 32.5% “omitted part of the central concept or reference information.” Fabricated content appeared in 6.1% of submissions.

That last figure is worth dwelling on. The main problem isn’t hallucination. It is incomplete or wrong work that reads well, which is harder to catch in a quick review.

The launch post describes the patterns behind those failures: “In 36.0% of submissions, the agent described an interface without checking how readers use it. In 33.1%, the agent missed decisive evidence and filled the gap with a plausible assumption. In 30.1%, the agent stopped at the first plausible page and left other affected pages stale.” The arXiv abstract reports the same three figures.

Then there is the first decision: does this change need a docs update at all? Across the seven agents, correct abstention “ranged from 28.6% to 85.7% of the 35 items that needed no change” (Promptless launch post). Some agents edit docs that should have been left alone most of the time.

Who made it, and why that matters

DoGBench was built by Promptless (dogbench.ai), which sells a product that, per its site, wakes up “When a PR opens, a slack thread about docs starts, or a DOC ticket gets created” and decides whether that warrants a doc update. The launch post states that “The paper’s ranking doesn’t include our agent.” Readers should weigh the results knowing the affiliation, and knowing the benchmark is new.

Where to put humans in an AI docs workflow

Each failure mode maps to a review gate. Our view, based on the numbers above:

The “does this need a docs change?” decision. A human owner triages, or at least approves, before any patch is drafted. Abstention is where agents varied most, from 28.6% to 85.7%.

Completeness across related pages. The reviewer checks the other pages a change touches, not just the diff. That is the 30.1% of submissions that stopped at the first plausible page.

Accuracy against the product. Someone who knows the product, or a test that runs the documented steps, verifies the patch. 36.6% of audited submissions contained technical inaccuracies.

Reader use. The reviewer asks one question: can a user complete the task with this? That targets the 45.5% task-completion gap and the 36.0% of submissions that described an interface without checking how readers use it.

Everything outside those gates (drafting, formatting, first-pass wording, opening the PR) is a reasonable candidate for automation. Our post on how to automate documentation with CI workflows covers those parts. The trigger problem isn’t unique to docs either: internal release notes can’t be generated from your commits for the same reason, because the change signal alone doesn’t carry enough context.

The launch post reaches a similar conclusion under the heading “Expert review remains necessary”, noting that documentation owners know who the readers are and what they are trying to do, context a code change alone may not reveal. If you publish AI-assisted docs, being open about where that review happens builds more trust than pretending it doesn’t, a point we make in building trust in AI-assisted content. And since agents increasingly read your docs as well as write them, the same gaps matter on both sides, as we argue in writing docs for humans and AI agents.

Get the full checklist

These failures show up in the Understand and Trust layers of our agent-ready docs checklist, for example “Nothing left to guess” (Understand, Critical, checked by human review) and “Procedures still work” (Trust, Critical). This is one of the reasons behind the checklist: 16 checks across four layers. Get the free checklist.

Frequently asked questions

  • In the DoGBench paper, the best combined score among the seven agents evaluated was 47.3 out of 100 on the 117-item held-out split. That is a harmonic mean of patch quality and correctly leaving docs alone, not a percentage of correct patches. The highest P0-clean delivery rate, meaning patches with no critical failure, was 39.0%, or 32 of the 82 held-out tasks that required an update.

  • An audit of 1,267 DoGBench submissions found 45.5% had a task-completion gap, 36.6% contained technical inaccuracies and 32.5% omitted part of the central concept or reference information. Fabricated content appeared in only 6.1%. The main risk is incomplete or wrong work that reads well, such as stopping at the first plausible page and leaving other affected pages stale, which is harder to catch in a quick review.

  • We'd put people at four gates: deciding whether a change needs a docs update at all, since correct abstention ranged from 28.6% to 85.7% across agents; checking completeness across related pages; verifying accuracy against the product, by an expert or a test that runs the steps; and asking whether a user can complete the task. Drafting, formatting, first-pass wording and opening the PR are reasonable candidates for automation.