Microsemantics is the set of sentence-level rules that make a page’s facts cheap for search and AI engines to extract. This article defines microsemantics, then sets out its nine rules: present indicative modality, subordinate text, first-word sequences, triple verbalisation, bold answer terms, numeric precision, post-noun instances, conditional clause placement, and one value per fact. It then lists the formatting codes a brief uses to request each answer shape, a checklist for editing a draft, and the questions writers ask.
What microsemantics is
Microsemantics is the discipline of word choice, word order and modality inside a sentence, applied so that each sentence yields one extractable fact.
Language models parse prose by word distance and probability. A sentence that states a fact in subject, predicate and object order costs little to extract. A sentence that hedges, delays or wraps the fact in a pronoun costs more, and qualitative re-ranking prefers the cheaper candidate. Microsemantics is where the ranking trade-off between responsiveness and cost is decided, one sentence at a time.
Present indicative modality and no modal verbs
Present indicative modality is the rule that every factual sentence states what is, with no modal verbs: will, should, may, might or could.
| Version | Sentence |
|---|---|
| Speculative | If you are building a topical map, you might want to start with a keyword export, and it should probably cover your main services. |
| Indicative | A topical map starts from a query network and covers every attribute of the central entity. |
Subordinate text: the first sentence under a heading
Subordinate text is the first sentence under a heading, and it carries the highest passage extraction weight on the page.
Subordinate text answers the heading at once, with no introduction. A heading that asks what a query network is opens with “A query network is” and the definition. A heading that asks a price opens with the price. Every heading on this site follows the rule, which is why each section here opens with a definition.
First-word sequences
A first-word sequence is the opening three to five words of a paragraph, and it carries a concrete entity and attribute, never a pronoun or a transition.
| Version | Sentence |
|---|---|
| Pronoun padding | Furthermore, it is also worth remembering that these need to be reviewed before launch. |
| Entity first | Content briefs pass a microsemantic review before launch. |
Triples and data-to-text verbalisation
A triple is a subject, predicate and object statement, and data-to-text verbalisation is the act of writing each fact so the triple reads straight off the sentence.
Search engines decompose documents into subject, predicate and object triples to populate knowledge graphs. Retrieval-augmented generation systems in AI search lift the same triples into answers. A sentence written as its triple needs no interpretation:
- Unstructured: “PropXpro grew a lot after we rebuilt the site, roughly two hundred pages in total.”
- Triple: ⟨PropXpro, has architecture size, 216 pages⟩
- Verbalised: “PropXpro runs on a 216-page semantic content network.”
Bold answer terms
A bold answer term is the exact value a sentence answers with, set in bold so the eye and the parser find the answer, not the question.
Bold marks the value, never the query term or the entity name. “A new domain launches 15 to 20 quality nodes” bolds the answer. Bolding “quality nodes” alone repeats the question back to the reader.
Numeric precision
Numeric precision is the rule that every quantity appears as an exact number with its unit, never as a vague quantifier.
“Many pages” becomes “216 pages”. “A fast turnaround” becomes “five business days”. Each exact number is also one value sitewide: the same fact carries the same number on every page that states it.
Post-noun instances
A post-noun instance is a concrete example placed directly after the plural noun it belongs to.
“AI answer engines such as Google AI Overviews, ChatGPT and Perplexity cite extractable passages” names the instances where the class appears. “Many AI tools cite pages” names none, and the parser learns nothing about which engines.
Conditional clause placement
Conditional clause placement is the rule that the core fact comes first and the condition (if, when, because) comes second.
| Version | Sentence |
|---|---|
| Condition first | When a domain is older than 12 months, the launch batch grows. |
| Fact first | The launch batch grows to 30 to 50 quality nodes when the domain is older than 12 months. |
One value per fact
One value per fact is the sitewide rule that each entity, attribute and value pair carries a single value on every page.
Two pages that state different page counts for the same project contradict each other, and a contradiction loses the citation for both. Repeated facts reuse the exact same n-gram as well. Synonym rotation reads as variety to a person and as a second, unconfirmed fact to a parser.
Microsemantic formatting codes
Microsemantic formatting codes are the shorthand a content brief uses to request a specific answer shape for each heading.
| Code | Answer shape | Limit |
|---|---|---|
| FS | Featured snippet answer, value in bold within the first five words | 40 words or 320 characters |
| PAA | People Also Ask answer, one declarative sentence | 30 words or 220 characters |
| DEF | Definition: [entity] is a [class] that [differentiator] | 35 words |
| LST | List preceded by a full definition sentence ending in a colon | 3 to 7 items |
| EXA | Exact definitive answer with every qualifier and metric | 30 words by default |
| ANCR | Anchor text matching the target page’s title tag vocabulary | 2 to 5 words, 100 words from the previous link |
Where these codes sit inside a brief is set out in how to write a content brief.
A microsemantic editing checklist
A microsemantic edit checks every draft against eight questions, in this order:
- Does the first sentence under each heading answer the heading?
- Does any sentence contain will, should, may, might or could?
- Does any paragraph open with a pronoun or a transition word?
- Does each sentence state one fact as a readable triple?
- Is every answer value in bold, and nothing else?
- Is every quantity an exact number with its unit?
- Do instances follow their plural nouns, and conditions follow their facts?
- Does every fact match its value on every other page?
Sentence rules only pay off when every page of the network follows them, and the copywriting service applies them to every brief. → Microsemantic copywriting
Microsemantics questions
Does microsemantic writing read as robotic?
Microsemantic writing reads as plain and direct. Rhythm comes from varied sentence length and concrete examples, not from hedges and transitions.
Is “can” a modal verb under these rules?
“Can” states capability, not speculation, so the rule allows it sparingly. Will, should, may, might and could stay out of factual sentences.
How long is a sentence under microsemantic rules?
A sentence under microsemantic rules runs 22 words at most. Longer sentences split into two facts.
Do microsemantics matter for AI search?
Microsemantics matter more for AI search, because answer engines lift individual sentences. A sentence that carries its own triple is quoted intact.