Teal stepping stones cross a navy ravine with small flags marking each stone, and one gap in the path is bridged by a coral plank
This content was generated using AI.

This post is about writing procedures for software: the how-to guides, setup pages and runbooks in your product docs. If you are after a standard operating procedure template for a factory floor or an ISO audit, this is the wrong page.

Most advice on writing procedures is about form: numbered steps, one action per step, imperative verbs. That advice is still right. It just assumes a reader who repairs what the page leaves out. A human fills a gap from experience, asks a colleague, or notices the screen doesn’t match the screenshot and stops. An AI agent does none of that. It fills the gap with a guess, carries on, and reports success.

Why procedures break differently for agents

A human reader is a gap-filling machine with a fallback: ask someone. An agent’s fallback is a plausible assumption. So the details writers leave out because “everyone knows” (which role you need, which version the steps apply to, what happens if you skip an option, what success looks like) become the points where an agent goes wrong without telling you.

Writing these down isn’t writing for robots. The same sentences help the new hire on their first week, which is the argument we make in writing docs for humans and AI agents.

Evidence: agents assume, even when they are the writer

The clearest recent data comes from a benchmark of agents writing docs, not reading them. DoGBench, built by Promptless, gives an agent a repository as it existed before a change and a trigger such as a code pull request, then asks it to decide whether the documentation needs updating and to write the patch.

When the authors audited the submissions, the Promptless launch post reports that “In 33.1%, the agent missed decisive evidence and filled the gap with a plausible assumption”, and “In 36.0% of submissions, the agent described an interface without checking how readers use it.” The same two figures appear in the DoGBench paper’s abstract on arXiv.

To be clear about the limits: DoGBench did not measure agents following procedures. Our inference is narrower. Filling missing information with a plausible assumption is how these systems behave when evidence is absent. When an agent is the reader and your procedure leaves out the role or the default, expect the same behaviour.

The four things to make explicit

Access

State the account type, role or permission, and any API key scope the task needs. Google’s developer style guide on procedures puts it as: “Ensure that the reader has the information that they need in order to prepare for the task ahead of time.” An agent with the wrong role will hit an error at step 3 and, unless told what that error means, may try workarounds rather than stop.

Versions

Say which product, CLI or API version the steps apply to, and how to check it. Flags get renamed and screens move. An agent can’t tell an outdated page from a broken install unless the page names the version.

Defaults

For every optional flag or field, say what happens if it is left alone. Diataxis on how-to guides recommends conditional imperatives: “If you want x, do y. To achieve w, do z.” The default is the other half of that condition. Leave it out and the agent picks one, or worse, accepts whatever the tool does and assumes that was the intent. (More on the framework in our Diataxis post.)

Expected results

Say what success looks like after key steps and at the end. Google’s guide says “State the action first and the result second.” The Microsoft Writing Style Guide adds: “Make sure that customers know where the action should take place before you describe the action.” For an agent, the stated result is often the only signal it has for whether to continue or stop.

Before and after

The product here is invented: a fictional CLI called shipctl. The task is creating an API key.

Before:

## Create an API key

Go to Settings and create an API key. Add it to your
environment and run the deploy.

After:

## Create an API key

Before you begin:
- You need the Admin role on the workspace. Members can view keys but not create them.
- You need shipctl 2.4 or later. Check with `shipctl --version`.

1. In a terminal, run `shipctl keys create --name deploy`.
   Keys expire after 90 days unless you pass `--expires never`.
   The command prints a key ID and the secret once.
   If you see `403 forbidden`, your role is not Admin.
2. In your shell profile, set `SHIPCTL_KEY` to the secret.
3. Run `shipctl deploy`. The command ends with `Deploy complete`.

From the “before” version, an agent has to guess at least four things: whether its account is allowed to create keys, whether “Settings” means the web console or a CLI command, how long the key lasts, and what a successful deploy looks like. Each guess is a sentence the writer didn’t write.

How to check it

A self-review for each procedure page, short enough to forward:

[ ] Says who can do this (role, plan, permission)
[ ] Says which version it applies to, and how to check
[ ] For every optional input, says what happens by default
[ ] After each step that changes state, says what you should see
[ ] Says what to do if the result does not match

Then test it. Give an agent only the page (no web access), ask it to complete or plan the task, and ask it to list every point where it had to guess. Each guess is a missing sentence. This speeds up the human review but doesn’t replace it: someone who knows the product still decides which guesses matter. Whether the steps still run against the current product is a separate check, covered by automated procedure testing.

This page-level discipline sits inside the broader habits in our post on technical writing best practices when AI makes writing cheap. The cost of skipping it shows up downstream, often as a support bot giving wrong answers with complete confidence.

Get the full checklist

This is the “Nothing left to guess” check, in the Understand layer and rated Critical. It is the only one of the 16 checks in our agent-ready docs checklist whose tool is human review, because no scanner can tell you that step 3 silently needs admin rights. Get the free checklist.

Frequently asked questions

  • A human reader fills gaps from experience, asks a colleague, or stops when the screen doesn't match the screenshot. An AI agent does none of that. It fills the gap with a plausible assumption, carries on, and reports success. So the details writers leave out because everyone knows them, such as the role needed, the version the steps apply to, or what success looks like, become the points where an agent goes wrong without telling you.

  • Give an agent only the page, with no web access, ask it to complete or plan the task, and ask it to list every point where it had to guess. Each guess is a missing sentence. This speeds up human review but doesn't replace it: someone who knows the product still decides which guesses matter. Checking that the steps still run against the current product is a separate job, handled by automated procedure testing.

  • Yes. For every optional flag or field, say what happens if it is left alone. Diataxis recommends conditional imperatives such as "If you want x, do y", and the default is the other half of that condition. Leave it out and an agent picks a value itself, or accepts whatever the tool does and assumes that was the intent. One clear sentence about the default removes that guess entirely.