Several differently shaped navy tags flow into a funnel and come out as one teal tag, while a single coral tag is set aside
This content was generated using AI.

Picture a product where the getting-started guide talks about a “workspace”, the API reference calls the same thing a “project”, and a screenshot caption calls it a “space”. A human reader usually works out that these are one object, and occasionally guesses wrong. An AI agent or a retrieval system has no such instinct. It may treat them as three different things, blend them into a muddled answer, or miss the right page entirely because the user asked about a “workspace” and the page says “project”.

That is why we treat terminology consistency as a functional requirement, not a matter of polish. What an English teacher would praise as elegant variation is a bug in technical documentation. The good news is that the fix is cheap and mechanical.

Why synonyms break agents and retrieval

Search people have a name for this problem: vocabulary mismatch. Wikipedia’s summary of the vocabulary mismatch research cites Furnas et al. (1987), who found that “on average 80% of the times different people (experts in the same field) will name the same thing differently”. The same article cites Zhao and Callan (2010): “an average query term fails to appear in 30-40% of the documents that are relevant to the user query”.

Retrieval-augmented AI inherits the problem. In Embeddings aren’t magic on Towards Data Science, a user asks a RAG assistant about the rule on contractor overtime. The assistant replies “I couldn’t find that information.” The answer was in the document all along, under “non-employee labor”. The author’s fix is not a better embedding model but a curated keyword dictionary that maps one term to the other.

Your users will always bring their own words. You can’t control that. What you can control is how many words your own docs use for each concept. Every extra synonym splits a concept across pages, so a user’s word matches some of them and misses the rest. A controlled vocabulary (one approved term per concept, with known synonyms mapped to it) shrinks that gap from your side. If your support bot is giving wrong answers, this is one of the first places to look; we cover the broader pattern in why your AI support bot gives wrong answers.

Style guides already agree; they just don’t enforce

Nobody argues against this rule. The Microsoft Writing Style Guide on simple words and concise sentences says it in one line: “Use one term consistently to represent one concept.”

Google’s Technical Writing One course on words puts it more vividly: “If you rename a term in the middle of a document, your ideas won’t compile (in your users’ heads).” It also allows a sensible exception. Introduce the short form once, as in “Protocol Buffers (or protobufs for short)”, and then use it consistently.

So the guidance exists. The gap is that it lives in a style guide nobody opens while writing. A rule that depends on someone remembering to check will drift the moment a deadline arrives.

The glossary is the source of truth

Start with a documentation glossary that records, for each concept:

  • the one approved term
  • the synonyms you reject
  • a one-line definition

Keep it in the repository next to the docs, so a change to a term goes through the same review as a change to the content. The linter configuration is then maintained alongside the glossary (or generated from it), so the list people read and the list the check enforces never disagree.

The glossary also forces the useful conversation. If marketing renames “project” to “workspace”, someone has to decide whether “project” is now banned, and the glossary is where that decision gets written down.

How to check it

Vale, the prose linter, gives you two mechanisms. If you haven’t set it up yet, start with installing and configuring Vale.

Vocabularies. A Vale vocabulary is two plain-text files stored in <StylesPath>/config/vocabularies/<name>/, as the Vale vocabularies documentation describes. Entries in accept.txt feed the built-in Vale.Terms rule, which makes sure occurrences exactly match the entry (useful for the capitalisation of product names). Entries in reject.txt feed Vale.Avoid, which flags every occurrence as an error. Entries are case-sensitive regular expressions.

Substitution rules. A reject list says a word is wrong. A Vale substitution rule also says what to write instead, which is what you want for synonyms. A minimal rule, saved as something like styles/YourStyle/Terminology.yml:

extends: substitution
message: "Use '%s' instead of '%s' (see the glossary)."
level: error
ignorecase: true
swap:
  team space: workspace
  shared space: workspace
  log in to: sign in to

The swap keys are the rejected terms and the values are the approved ones. Per the Vale substitution check documentation, a message can carry one or two %s specifiers, as in its example “Consider using ‘%s’ instead of ‘%s’.”, and ignorecase makes all matches case-insensitive. Keys can be regular expressions, and a value can offer several alternatives separated by |.

Then run Vale on every pull request so the wrong term is flagged before it merges. Our guide to Vale as a GitHub Actions merge gate has the setup.

Rolling it out without drowning in warnings

The quickest way to kill a linter is to make it noisy. A few habits help:

  • Start small. Pick the handful of terms (five to ten is plenty) that cause real confusion. Support tickets and your docs search logs are good places to find them.
  • Warn before you fail. Ship new rules at level: warning, fix the backlog, then promote them to error.
  • Prefer phrases over single words. “Project” or “space” may be a rejected synonym in one sentence and a perfectly legitimate word in the next. Swap on specific phrases such as “team space” rather than generic nouns, or the rule will flag correct text and people will learn to ignore it.

Over-broad rules are the honest counter-argument to all of this. A narrow rule set that people trust does more for consistency than a sweeping one they route around.

Get the full checklist

This is one of 16 checks in our agent-ready docs checklist: “One term per concept”, in the UNDERSTAND layer, priority Important. The checklist covers how agents find, read, understand and trust your documentation. Get the free checklist.

Frequently asked questions

  • An AI agent or retrieval system lacks a human reader's instinct for connecting synonyms. It may treat "workspace", "project" and "space" as three different things, blend them into a muddled answer, or miss the right page because the user's word isn't the one on the page. Every extra synonym splits a concept across pages, so a user's term matches some and misses the rest. One approved term per concept shrinks that gap from your side.

  • Entries in a vocabulary's reject.txt feed Vale's built-in Vale.Avoid rule, which flags every occurrence as an error but doesn't say what to write instead. A substitution rule also names the replacement: its swap keys are the rejected terms and its values are the approved ones, so the message can tell the writer exactly which word to use. For synonyms, a substitution rule is usually what you want.

  • Start with the five to ten terms that cause real confusion; support tickets and docs search logs are good places to find them. Ship new rules at warning level, fix the backlog, then promote them to error. Swap on specific phrases such as "team space" rather than generic nouns like "space", which may be legitimate elsewhere. A narrow rule set people trust does more for consistency than a sweeping one they route around.