Most advice on documentation code samples stops at “keep them short and clear”. That advice is fine as far as it goes. The harder problems are that a short sample is often incomplete, and that a complete sample quietly stops working a few releases later because nothing runs it.
The snippet that only works if you already know the answer
You have seen this sample. It calls client.create(...) and that is the whole block. No import, no install line, no version, no auth setup. A teammate reads it and fills the gaps without thinking, because they already know which package the client comes from and how the token gets set.
An AI agent reading the same block does something different. It copies the snippet literally and guesses the rest, often from training data that may be a few releases behind. It will happily invent an import path or a package name that looks right.
How often do models get that guess wrong? A study by Spracklen et al., We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs, analysed 576,000 code samples from 16 popular LLMs. The rate of hallucinated packages was “at least 5.2%” for commercial models and “21.7%” for open-source models, with “205,474 unique examples of hallucinated package names”. To be precise about what that shows: the study measures models generating code, not agents reading documentation. The narrower point it supports is enough. When a model has to supply a dependency itself, it gets it wrong often enough to matter. So don’t make it supply one.
Missing context in a sample was always a minor annoyance for humans. For agents it is a defect. (For the wider picture of what agents need from your docs, see agent-ready API docs: what they require.)
What “complete” means
Google’s Technical Writing course on sample code sets a reasonable bar. Good samples “build without errors” and “perform the task it claims to perform”, and the docs should provide “all information necessary to run the sample code, including any dependencies and setup”.
In practice, a complete sample has:
- the imports
- the install or dependency line, with a version
- required config or environment variables, with obvious placeholders
- the expected output, where it helps someone confirm it worked
This pulls against brevity, and sometimes brevity wins. Google’s style guide on code samples allows omissions but sets a rule for them: “Indicate omitted code by using a comment in the syntax of the language of your code sample. Don’t use three dots or the ellipsis character.” It also says that if a block contains an omission, don’t format it as click-to-copy. That is the honest rule. Leaving code out is fine when the omission is explicit, so nobody, human or agent, mistakes a fragment for a runnable program.
The language tag: one word, two readers
The cheapest fix in this whole post is the word after the opening fence. markdownlint’s rule MD040 (fenced-code-language) flags any fenced code block with no language specified. The rule’s own rationale is rendering: “Specifying a language improves content rendering by using the correct syntax highlighting”.
Our view is that the same word does more work for an agent. Knowing whether a block is bash, Python, JSON or console output changes what it does with it: run it, import it, parse it, or treat it as an example of output. MD040 accepts text for plain output, and its allowed_languages option lets you restrict tags to a list your team agrees on.
A related structural check: the Agent Docs Spec includes a code fence validity check, because “An unclosed code fence causes everything after it to be interpreted as code rather than prose.”
Why samples break after launch
A sample is a claim about the product, and the product moves. Google’s course is blunt about it: “Always test your sample code. Over time, systems change and your sample code may break.”
The rustdoc book states the purpose of documentation tests just as plainly: to “make sure that examples within your documentation are up to date and working”. Review won’t keep that true. Running the samples will. This is a CI problem, and you already have CI (our guide to five CI workflows for documentation covers where this fits).
Two ways to keep samples tested
Neither is better in general. Pick whichever matches how your docs are built.
Include samples from tested code. The sample lives in a real file in the repo, compiled and tested with everything else, and the docs pull it in. PyMdown’s Snippets extension is a concrete example. Mark a section in the source file with # --8<-- [start:func] and # --8<-- [end:func], then include it in the docs with --8<-- "example.py:func". Turn on check_paths and the build fails if a referenced snippet file can’t be found. Other doc stacks have equivalent include features.
Run the docs as tests. Python’s doctest “searches for pieces of text that look like interactive Python sessions, and then executes those sessions to verify that they work exactly as shown”, and it can run against a plain text file. Rust runs the examples in doc comments with cargo test --doc.
One caution. rustdoc lets you hide lines from the rendered output by starting them with # . They still compile in the test, but readers don’t see them. That keeps the test honest while handing readers, and agents, a sample that looks incomplete, which is the problem we started with. Our recommendation: show the imports even when the tooling lets you hide them. Hide boilerplate if you must, never the lines that say where things come from.
For testing whole procedures (UI steps, CLI sequences) rather than individual snippets, see our docs-as-tests guide.
Live links on every PR
Samples sit among links: SDK repos, reference pages, downloads. A sample that works but points to a dead install page still fails the reader. The lychee-action runs lychee in GitHub Actions to “check links in Markdown, HTML, and text files”. It is a small addition to a pipeline you already run.
How to check it
Something to forward to an engineer. Three pieces, no full workflow file.
Enable MD040 with an agreed tag list in .markdownlint.json:
{
"MD040": {
"allowed_languages": ["bash", "python", "json", "yaml", "text"]
}
}
Add the link checker as a step, adapted from the action’s README:
- name: Link checker
uses: lycheeverse/lychee-action@v2
with:
args: --no-progress './**/*.md'
fail: true
Set fail: true explicitly. The README’s own example uses fail: false, which reports broken links without blocking the merge. Cache results in .lycheecache to reduce rate limiting, and list known-flaky URLs in .lycheeignore.
Run doc tests for your language in the same job:
python -m doctest -v docs/quickstart.md
(or cargo test --doc for Rust). To make these required checks that block a merge, follow the same pattern as our Vale merge gate guide. For broader API documentation practice, see our REST API documentation best practices.
Get the full checklist
This guide covers two of the 16 checks in our agent-ready docs checklist: “Complete code samples” in the Understand layer (Critical) and “Tested samples, live links” in the Trust layer (Important). Get the free checklist.
References
- We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (Spracklen et al.)
- Google Technical Writing Two, Sample code
- Google developer documentation style guide, Code samples
- markdownlint rule MD040 (fenced-code-language)
- Agent Docs Spec (SPEC.md)
- rustdoc, Documentation tests
- PyMdown Extensions, Snippets
- Python doctest documentation
- lychee-action (GitHub Action)