# AI Slop After the Em Dash

*We banned the em dash and the model stopped using it. The slop moved into sentence shapes, then section headers, then last week into the titles of a slide deck whose paragraphs passed every rule I had.*

> For a year the em dash was the tell. Ban it and the text looked human. Then the tells moved into sentence shapes, section headers, invented frameworks, and this month into slide titles. This post traces why models keep producing the same shapes after each ban, what a word count over my own 25 posts turned up, which single check has held across every layer so far, and the setup I now run so my agents apply it to decks, diagrams, tables and Slack messages as well as prose.

Published: 2026-09-14 · Reading time: 15 min · Tags: agents, claude-code, writing, ai-slop, product-management, memory
Canonical URL: https://huytieu.com/blog/ai-slop-after-the-em-dash/
Author: Huy Tieu (huytieu.com)

---

<!-- slop-ok: this post quotes AI-slop tells as exhibits -->

Two years ago you could spot AI writing by the em dash. Three per paragraph, one in every heading, a rhythm crutch no human editor would let through. I put "no em dashes" in my rules file, the dashes disappeared, and for a while the text read like a person wrote it.

The colon came next. "The detail that makes it work: a separate agent grades it." Same job the dash was doing, one character over. After that came "It's not X. It's Y.", then headings called "Why this matters" over two paragraphs, then three-pillar frameworks nobody asked for. Each time I banned a form, the form vanished and something with the same function appeared one level up.

<style>
:root{--sf-ground:#fbfbfa;--sf-ink:#16181d;--sf-muted:#6b7280;--sf-quiet:#9aa1ad;--sf-line:#e3e5e9;--sf-strong:#c8ccd3;--sf-accent:#d97757;--sf-sans:ui-sans-serif,-apple-system,"Segoe UI",Inter,sans-serif;--sf-mono:ui-monospace,SFMono-Regular,"SF Mono",Menlo,monospace}
:root[data-theme="dark"]{--sf-ground:#171b22;--sf-ink:#f1f5f9;--sf-muted:#8b95a5;--sf-quiet:#5b6472;--sf-line:#262c36;--sf-strong:#3a4250;--sf-accent:#e0885f}
.slopfig{margin:34px 0;padding:22px 22px 16px;border:1px solid var(--sf-line);border-radius:4px;background:var(--sf-ground);color:var(--sf-ink)}
.slopfig figcaption{margin-top:16px;padding-top:10px;border-top:1px solid var(--sf-line);font:500 9px/1.5 var(--sf-mono);letter-spacing:.13em;text-transform:uppercase;color:var(--sf-muted)}
.slopfig__eyebrow{font:500 9px/1 var(--sf-mono);letter-spacing:.13em;text-transform:uppercase;color:var(--sf-quiet);margin-bottom:14px}
.slopfig b{font-weight:520}
/* ladder */
.sf-rung{display:grid;grid-template-columns:74px 1fr;gap:14px;align-items:start;padding:9px 0;border-top:1px dashed var(--sf-line)}
.sf-rung:first-of-type{border-top:0}
.sf-rung__lvl{font:500 9px/1.7 var(--sf-mono);letter-spacing:.1em;text-transform:uppercase;color:var(--sf-quiet)}
.sf-rung__spec{font:400 13px/1.6 var(--sf-mono);color:var(--sf-muted);word-break:break-word}
.sf-rung__name{font:520 14px/1.4 var(--sf-sans);color:var(--sf-ink);margin-bottom:2px}
.sf-rung--now .sf-rung__lvl,.sf-rung--now .sf-rung__name{color:var(--sf-accent)}
.sf-rung--now .sf-rung__spec{color:var(--sf-ink)}
/* forces */
.sf-forces{display:grid;grid-template-columns:1fr;gap:10px}
.sf-force{display:block}
.sf-force__box{border:1px solid var(--sf-line);border-radius:3px;padding:10px 12px}
.sf-force__t{font:520 13px/1.35 var(--sf-sans)}
.sf-force__d{font:400 12px/1.5 var(--sf-sans);color:var(--sf-muted);margin-top:3px}
.sf-forces{position:relative}
.sf-force__arm{display:none}
.sf-sink{margin-top:12px;position:relative;border:1px solid var(--sf-accent);border-radius:3px;padding:12px 14px;text-align:center}
.sf-sink::before{content:"";position:absolute;top:-13px;left:50%;width:1px;height:12px;background:var(--sf-strong)}
.sf-sink__t{font:520 14px/1.3 var(--sf-sans);color:var(--sf-accent)}
.sf-sink__d{font:400 12px/1.5 var(--sf-sans);color:var(--sf-muted);margin-top:4px}
/* split */
.sf-split{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:4px}
.sf-branch{border:1px solid var(--sf-line);border-radius:3px;padding:12px 14px}
.sf-branch--off{border-style:dashed;border-color:var(--sf-accent)}
.sf-branch__h{font:500 9px/1 var(--sf-mono);letter-spacing:.11em;text-transform:uppercase;color:var(--sf-quiet);margin-bottom:8px}
.sf-branch--off .sf-branch__h{color:var(--sf-accent)}
.sf-branch__l{font:400 12px/1.75 var(--sf-mono);color:var(--sf-muted)}
.sf-branch--off .sf-branch__l{color:var(--sf-ink)}
.sf-stem{text-align:center;font:400 11px/1 var(--sf-mono);color:var(--sf-quiet);padding:8px 0 10px}
/* bars */
.sf-bar{display:grid;grid-template-columns:172px 1fr 46px;gap:10px;align-items:center;padding:3px 0}
.sf-bar__l{font:400 12px/1.4 var(--sf-sans);color:var(--sf-muted);text-align:right}
.sf-bar__t{height:9px;background:var(--sf-strong);border-radius:1px}
.sf-bar__t--a{background:var(--sf-accent)}
.sf-bar__t--g{background:transparent;border:1px dashed var(--sf-strong);height:9px;box-sizing:border-box}
.sf-bar__v{font:400 11px/1.4 var(--sf-mono);color:var(--sf-quiet)}
.sf-bar--hi .sf-bar__l{color:var(--sf-ink)}
.sf-key{display:flex;gap:16px;flex-wrap:wrap;margin-top:12px;font:400 10px/1 var(--sf-mono);letter-spacing:.06em;text-transform:uppercase;color:var(--sf-quiet)}
.sf-key span{display:flex;align-items:center;gap:6px}
.sf-key i{width:16px;height:8px;display:inline-block;background:var(--sf-strong)}
.sf-key i.a{background:var(--sf-accent)}
.sf-key i.g{background:transparent;border:1px dashed var(--sf-strong)}
/* funnel */
.sf-funnel{display:grid;gap:6px}
.sf-step{border:1px solid var(--sf-line);border-radius:2px;padding:9px 12px;display:flex;justify-content:space-between;align-items:baseline;gap:12px}
.sf-step--a{border-color:var(--sf-accent)}
.sf-step__l{font:400 13px/1.4 var(--sf-sans);color:var(--sf-muted)}
.sf-step--a .sf-step__l{color:var(--sf-ink)}
.sf-step__n{font:500 15px/1 var(--sf-mono);color:var(--sf-ink)}
.sf-step--a .sf-step__n{color:var(--sf-accent)}
/* specimens */
.sf-spec{display:grid;grid-template-columns:1fr 1fr;gap:0;border:1px solid var(--sf-line);border-radius:4px;overflow:hidden;margin:12px 0;background:var(--sf-ground);color:var(--sf-ink)}
.sf-spec__side{padding:11px 13px}
.sf-spec__side+.sf-spec__side{border-left:1px solid var(--sf-line)}
.sf-spec__h{font:500 9px/1 var(--sf-mono);letter-spacing:.11em;text-transform:uppercase;color:var(--sf-quiet);margin-bottom:7px}
.sf-spec__side--y .sf-spec__h{color:var(--sf-accent)}
.sf-spec__t{font:400 13px/1.6 var(--sf-sans);color:var(--sf-ink)}
.sf-spec__side--n .sf-spec__t{color:var(--sf-muted)}
.sf-spec__why{font:400 11px/1.5 var(--sf-sans);color:var(--sf-quiet);margin-top:6px}
@media (max-width:620px){.sf-split,.sf-spec{grid-template-columns:1fr}.sf-spec__side+.sf-spec__side{border-left:0;border-top:1px solid var(--sf-line)}.sf-bar{grid-template-columns:120px 1fr 42px}.sf-bar__l{font-size:11px}}
</style>

<figure class="slopfig" role="img" aria-label="Four levels of AI-slop tells, each appearing after the one below it was banned">
  <div class="slopfig__eyebrow">Ban a form, the function relocates</div>
  <div class="sf-rung"><div class="sf-rung__lvl">2024<br>Word</div><div><div class="sf-rung__name">Punctuation and diction</div><div class="sf-rung__spec">— &nbsp;·&nbsp; delve &nbsp;·&nbsp; robust &nbsp;·&nbsp; seamless</div></div></div>
  <div class="sf-rung"><div class="sf-rung__lvl">2025<br>Sentence</div><div><div class="sf-rung__name">Shapes that carry the same rhythm</div><div class="sf-rung__spec">"It's not X. It's Y." &nbsp;·&nbsp; "The best part: it learns."</div></div></div>
  <div class="sf-rung"><div class="sf-rung__lvl">2026<br>Structure</div><div><div class="sf-rung__name">Scaffolding over thin content</div><div class="sf-rung__spec">## Why this matters &nbsp;·&nbsp; ## The bottom line &nbsp;·&nbsp; three pillars, five gates</div></div></div>
  <div class="sf-rung sf-rung--now"><div class="sf-rung__lvl">Now<br>Medium</div><div><div class="sf-rung__name">Surfaces no rule named</div><div class="sf-rung__spec">| One harness, four swappable sides |<br>[ Four walls, seven accounts ] &nbsp;·&nbsp; chart legends, commit messages, alt text</div></div></div>
  <figcaption>Each ban removed a form and left the function, which reappeared one level up. The em dash was the first place the problem was visible.</figcaption>
</figure>

Last week it reached a place I had not thought to look. I asked my agent to build a slide deck about an AI test runner from my notes. The paragraphs passed every rule I had. The slide titles were "One harness, four swappable sides", "Four walls, seven accounts", "Five calls to make this month", headlines in the slot where a title belongs. Every rule I had was written for prose, and the agent had, reasonably, treated a slide title as layout.

## Training corpus and preference tuning

For a year I treated slop as a list of habits. The model uses "delve", you ban "delve", it stops using "delve", and the sentence that needed a filler word finds another one, because the model is producing text that resembles a finished, thorough, confident answer and that is what it was rewarded for.

Most text about work on the internet is marketing copy, consulting decks, listicles, LinkedIn posts, and documentation written to a template. In that corpus a title is a headline, a section has a "Why this matters", an argument has three parts, and a conclusion restates the opening. When a model reaches for the most probable shape of a slide title, it reaches for "Four walls, seven accounts".

Preference tuning then rewards the look of completeness. When people rate two answers, the one with headers, bullets, a summary and a confident closing line tends to win, because it looks like more work was done. So the model learns structure as a proxy for quality. A paragraph of reasoning scores lower than the same reasoning wrapped as "the three-layer model", so the wrapper appears even when the source material contains no layers, and answers end with "the bottom line" because answers that end that way rate as more finished than answers that stop when the information stops.

And words cost the model nothing. A human writer runs out of energy; the sentences that survive are the ones you cared enough to type. A model fills every slot the template offers, so two paragraphs get a header, a single point gets a bulleted list, and a table gets a "So what" column.


<figure class="slopfig" role="img" aria-label="Three training pressures converging on one output shape">
  <div class="slopfig__eyebrow">Three pressures, one output shape</div>
  <div class="sf-forces">
    <div class="sf-force"><div class="sf-force__box"><div class="sf-force__t">The training median is business prose</div><div class="sf-force__d">Marketing copy, consulting decks, listicles, templated docs. Where titles are most numerous, a title is a headline.</div></div></div>
    <div class="sf-force"><div class="sf-force__box"><div class="sf-force__t">Preference tuning rewards the look of completeness</div><div class="sf-force__d">Raters pick the answer with headers, bullets and a closing line. Structure becomes a proxy for quality, so reasoning arrives wrapped in a framework the source never had.</div></div></div>
    <div class="sf-force"><div class="sf-force__box"><div class="sf-force__t">Words cost the model nothing</div><div class="sf-force__d">A human writer runs out of energy, which filters sentences. A model fills every slot the template offers.</div></div></div>
  </div>
  <div class="sf-sink"><div class="sf-sink__t">Text that resembles a finished, thorough, confident answer</div><div class="sf-sink__d">Remove any one surface form and the three pressures are still pointing the same way.</div></div>
  <figcaption>No single ban moves the target, because the target is the appearance of completeness rather than any particular word.</figcaption>
</figure>

Under those conditions a ban moves the behaviour to the nearest unbanned surface: from a word to a sentence shape, from a sentence shape to section structure, from prose structure to whatever medium has no rule yet.

There is data on this now. [A Reddit user pulled 89,239 posts](https://www.reddit.com/r/ClaudeAI/comments/1ucpw87/i_pulled_90000_reddit_posts_about_what_makes/) from 47 subreddits about spotting AI writing, filtered to 7,984 on topic, and hand-audited 600 of them for what readers actually cite as the tell. The em dash leads at 7.1 percent of audited posts. Second is uniform sentence rhythm at 4.0 percent, third "not just X, it's Y" at 2.8, then the five-paragraph essay shape and sycophancy, both at 2.5. The words everyone bans, "delve" and its cousins, sit at 1.3 percent. The keyword scanner in the same study ranked "however", "thus" and "hence" as the top match, at 6.3 percent of posts, and readers cited them as a tell zero times. Rhythm and emptiness, the tells readers cite most after the dash, cannot be keyword-matched at all.


<figure class="slopfig" role="img" aria-label="Bar chart of AI-writing tells ranked by how often readers cite them">
  <div class="slopfig__eyebrow">What readers name, out of 600 hand-audited posts</div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">Em dash</div><div class="sf-bar__t sf-bar__t--a" style="width:100.0%"></div><div class="sf-bar__v">7.1%</div></div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">Uniform sentence rhythm</div><div class="sf-bar__t sf-bar__t--g" style="width:56.3%"></div><div class="sf-bar__v">4.0%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"not just X, it's Y"</div><div class="sf-bar__t" style="width:39.4%"></div><div class="sf-bar__v">2.8%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Five-paragraph essay shape</div><div class="sf-bar__t" style="width:35.2%"></div><div class="sf-bar__v">2.5%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Sycophancy ("great question")</div><div class="sf-bar__t sf-bar__t--g" style="width:35.2%"></div><div class="sf-bar__v">2.5%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"dive in" / "deep dive"</div><div class="sf-bar__t" style="width:28.2%"></div><div class="sf-bar__v">2.0%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Bullet lists where prose belongs</div><div class="sf-bar__t" style="width:23.9%"></div><div class="sf-bar__v">1.7%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Diction memes ("delve")</div><div class="sf-bar__t" style="width:18.3%"></div><div class="sf-bar__v">1.3%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Rule-of-three triads</div><div class="sf-bar__t" style="width:16.9%"></div><div class="sf-bar__v">1.2%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"in today's fast-paced world"</div><div class="sf-bar__t" style="width:9.9%"></div><div class="sf-bar__v">0.7%</div></div>
  <div class="sf-bar"><div class="sf-bar__l">Empty phrasing</div><div class="sf-bar__t sf-bar__t--g" style="width:9.9%"></div><div class="sf-bar__v">0.7%</div></div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">"however" / "thus" / "hence"</div><div class="sf-bar__t" style="width:1.0%"></div><div class="sf-bar__v">0%</div></div>
  <div class="sf-key"><span><i class="a"></i>top tell</span><span><i></i>keyword-matchable</span><span><i class="g"></i>no scanner can see it</span></div>
  <figcaption>"However", "thus" and "hence" are the top keyword match in the same corpus at 6.3% of posts, and readers cite them as a tell zero times. The cheap signal and the cited signal point in different directions.</figcaption>
</figure>

One commenter on that thread had run the same experiment I had, with Claude. He banned constructions one at a time and watched the model comply with each ban while keeping the underlying behaviour: ban the em dash and it switched to ", and" and semicolons. After two dozen bans he found the driver was over-explanation and restatement, and a single instruction, "be airy, don't over-explain", did more than the whole list. The [study's repository](https://github.com/JCarterJohnson/vibecoded-design-tells) is public and MIT licensed.

## Slide titles

My rules file says headings identify subject matter, never rhetorical function. It lists banned heading shapes. The agent had it in context and still wrote "One harness, four swappable sides" into the slide-title slot. When I asked why, the answer was that the paragraphs had been written in a writing step, where the rules fired, and the slide titles had been written in a layout step, as strings that look like slide titles look, where the rules did not fire because that step was not, in the agent's own accounting, writing.


<figure class="slopfig" role="img" aria-label="One rules file, two agent steps, only one of them checked">
  <div class="slopfig__eyebrow">Where the rule fired and where it did not</div>
  <div class="sf-stem">one rules file, loaded in context</div>
  <div class="sf-split">
    <div class="sf-branch">
      <div class="sf-branch__h">Step the agent calls writing &nbsp;→&nbsp; checked</div>
      <div class="sf-branch__l">paragraphs<br>bullets<br>speaker notes</div>
    </div>
    <div class="sf-branch sf-branch--off">
      <div class="sf-branch__h">Step the agent calls layout &nbsp;→&nbsp; unchecked</div>
      <div class="sf-branch__l">slide titles, kickers<br>tile and card labels<br>table headers, row keys<br>diagram node labels<br>chart titles, legends<br>button copy, alt text<br>commit messages<br>memory-file descriptions</div>
    </div>
  </div>
  <figcaption>Nobody exempted titles. The exemption fell out of how the task was split, and every surface the rules did not name sat on the unchecked side.</figcaption>
</figure>

Every medium I had not explicitly named was sitting on that unchecked side: diagram node labels, table headers, chart legends, Slack drafts, commit messages, the one-line descriptions on memory files.

## Information gain as the test

The test that has survived every surface change so far is marginal information gain: after each unit of output, ask what the reader now knows that they did not know before it, and delete the unit if the answer is nothing.

The same question removes "delve" (a shorter word carries the same information), "It's not X. It's Y." (X was never on the table), "Why this matters" as a heading (it names the prose's job and says nothing about the subject), and a three-pillar framework whose pillars were not in the source. It works at the medium level too. A reader glancing at the slide index learns nothing from "Four walls, seven accounts" and learns exactly when to jump from "Common challenges across customers". Here is the full set from that deck, before and after, with the check applied to each title:

| Was | Now |
|---|---|
| What runs today | Current AI Test Runner workflow |
| Who uses the runner, and the wall each one hits | Customers using the AI Test Runner and their use cases |
| Four walls, seven accounts | Common challenges across customers |
| What stands between the runner and production use | Current challenges |
| One harness, four swappable sides | New runner architecture |
| Why hard apps are the niche | AUT knowledge base: how it is built with a customer |
| Where the same engine can live | Two paths: local file-based vs online integrated |
| One harness, two shells, and a bridge | Recommendation: ship the local path first, keep True as the system of record |
| Five calls to make this month | Decisions needed |
| Where the facts come from | Sources |

The right column is duller, and "Sources" tells a reader scanning the index what is on the slide without opening it.

The same check catches what the bans miss inside the content. The deck had a slide body that said "the walls are the pattern" and a tile that said "hard apps are hard for everyone; we are the ones with a slot for the answer". Neither adds a fact the surrounding text lacks, and both sit where a conclusion would go. A ban list names repeated patterns and cannot see a one-off empty sentence.

The test also settles the argument about single words. My first draft of the new rule tried to protect words like "robust" by giving an example where the word was needed: "a robust retry on OTP fetch". My reviewer asked why the retry needed "robust". It did not, and "a retry that survives a mail-slot timeout" tells the reader what it does, where "robust" tells them the author wanted the sentence to sound engineered. The same applies to "clean", "plain", "simple", "elegant". Remove the adjective, and if the reader would do nothing differently, it was decoration. The exception is a term of art, "robust statistics", where the word is the name of the thing.

## Word counts across my own archive

Once I stopped trusting the ban list I counted. The twenty-five posts published here before this one, all of them drafted with an agent and edited by me, about 2,158 sentences.

My first pass used `grep -o` and reported "ship or shipped: 124 uses". That was wrong, and wrong in a way worth keeping: `grep -o` counts substrings, so it had been scoring "relationship", "shipping" and "ships" as instances of my tic. The tool I ended up writing counts on word boundaries and reports a per-file rate, which is the number that matters. The real figure is 51 across 25 posts, about two a post.

<figure class="slopfig" role="img" aria-label="Bar chart of over-used words across 25 posts">
  <div class="slopfig__eyebrow">Uses across 25 posts, word-boundary counts</div>
  <div class="sf-bar"><div class="sf-bar__l">just</div><div class="sf-bar__t" style="width:100.0%"></div><div class="sf-bar__v">56</div></div>
  <div class="sf-bar"><div class="sf-bar__l">actually</div><div class="sf-bar__t" style="width:94.6%"></div><div class="sf-bar__v">53</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"the whole"</div><div class="sf-bar__t" style="width:71.4%"></div><div class="sf-bar__v">40</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"the real"</div><div class="sf-bar__t" style="width:53.6%"></div><div class="sf-bar__v">30</div></div>
  <div class="sf-bar"><div class="sf-bar__l">"here's"</div><div class="sf-bar__t" style="width:50.0%"></div><div class="sf-bar__v">28</div></div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">ship</div><div class="sf-bar__t sf-bar__t--a" style="width:48.2%"></div><div class="sf-bar__v">27</div></div>
  <div class="sf-bar"><div class="sf-bar__l">genuinely</div><div class="sf-bar__t" style="width:44.6%"></div><div class="sf-bar__v">25</div></div>
  <div class="sf-bar"><div class="sf-bar__l">small</div><div class="sf-bar__t" style="width:44.6%"></div><div class="sf-bar__v">25</div></div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">shipped</div><div class="sf-bar__t sf-bar__t--a" style="width:42.9%"></div><div class="sf-bar__v">24</div></div>
  <div class="sf-bar"><div class="sf-bar__l">cheap</div><div class="sf-bar__t" style="width:32.1%"></div><div class="sf-bar__v">18</div></div>
  <div class="sf-bar"><div class="sf-bar__l">quietly</div><div class="sf-bar__t" style="width:30.4%"></div><div class="sf-bar__v">17</div></div>
  <div class="sf-bar"><div class="sf-bar__l">plain</div><div class="sf-bar__t" style="width:23.2%"></div><div class="sf-bar__v">13</div></div>
  <div class="sf-bar"><div class="sf-bar__l">clean</div><div class="sf-bar__t" style="width:23.2%"></div><div class="sf-bar__v">13</div></div>
  <div class="sf-bar"><div class="sf-bar__l">simple</div><div class="sf-bar__t" style="width:21.4%"></div><div class="sf-bar__v">12</div></div>
  <div class="sf-bar"><div class="sf-bar__l">delve</div><div class="sf-bar__t" style="width:14.3%"></div><div class="sf-bar__v">8</div></div>
  <div class="sf-bar sf-bar--hi"><div class="sf-bar__l">honestly</div><div class="sf-bar__t sf-bar__t--a" style="width:12.5%"></div><div class="sf-bar__v">7</div></div>
  <div class="sf-bar"><div class="sf-bar__l">robust</div><div class="sf-bar__t" style="width:12.5%"></div><div class="sf-bar__v">7</div></div>
  <div class="sf-key"><span><i class="a"></i>words someone had already told me about</span><span><i></i>the rest, found by counting</span></div>
  <figcaption>Not one of these is a banned word. Each is a word I reach for, counted often enough that the agent now reaches for it on my behalf.</figcaption>
</figure>

The shapes were the louder finding.

<figure class="slopfig" role="img" aria-label="One opener measured at three scales across the archive">
  <div class="slopfig__eyebrow">One opener, three scales</div>
  <div class="sf-funnel">
    <div class="sf-step"><div class="sf-step__l">Sentences in the archive</div><div class="sf-step__n">2,158</div></div>
    <div class="sf-step"><div class="sf-step__l">Sentences opening with "The"</div><div class="sf-step__n">357</div></div>
    <div class="sf-step sf-step--a"><div class="sf-step__l">Headings opening with "The" &nbsp;("The shot", "The wedge")</div><div class="sf-step__n">53</div></div>
    <div class="sf-step"><div class="sf-step__l">Closing lines of the form "That is the point."</div><div class="sf-step__n">21</div></div>
  </div>
  <figcaption>Fifty-three headings begin with "The", the subheading shape strangers name as an AI tell. The model did not bring it. It read 25 posts of mine and handed it back.</figcaption>
</figure>

Behind "The" in the heading census sit "What" at 22 and "Why" at 7. The narrative heading I borrowed from a writer I admire, "Then the build lied to me", appears seven times across two posts, and the agent has since treated it as the house pattern.

This is specific to working with one agent over time. Besides the median of the internet, the model samples the median of me, from the posts and rules and memory files it has in context, and it regresses toward my most frequent choices because those are the highest-probability tokens in my corpus. "Ship" was a word I used for software releases. In the agent's hands it became the verb for publishing a post, saving a file, sending a Slack message and deploying a service. "Honest" was a word I used once when I meant it, and it came back as "the honest part:" at the top of paragraphs where nothing dishonest was on offer.

The census is now part of my pre-publish scan. It prints only what sits above the per-file rate, and I decide whether each one is my voice or my mode.

## Specimens

The patterns are easier to refuse next to their repair. Every line on the left came out of a draft of mine, most of them out of drafts of this post.

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Number pairing</div><div class="sf-spec__t">One harness, four swappable sides</div><div class="sf-spec__why">Two small integers joined by a comma read as structure and carry none.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Names its subject</div><div class="sf-spec__t">New runner architecture</div><div class="sf-spec__why">A reader scanning the index knows which slide to open.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Question as title</div><div class="sf-spec__t">Why hard apps are the niche</div><div class="sf-spec__why">Withholds the answer so the slide can deliver it. Fine in a magazine, a beat wasted on slide eight.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Names its subject</div><div class="sf-spec__t">How the knowledge base is built with a customer</div><div class="sf-spec__why">Same slide, findable from the index.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">"X, not Y"</div><div class="sf-spec__t">The rule now reads as a test, not a permission.</div><div class="sf-spec__why">Nobody proposed "permission". The contrast is invented to give the sentence a shape.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">States the thing</div><div class="sf-spec__t">Remove the adjective. If the reader would do nothing differently, it was decoration.</div><div class="sf-spec__why">The rule itself, with no straw man standing in front of it.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Verdict kicker</div><div class="sf-spec__t">The right column is duller. That is the point.</div><div class="sf-spec__why">A closing line placed for cadence, adding no fact to the paragraph.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Carries the fact</div><div class="sf-spec__t">"Sources" tells a reader scanning the index what is on the slide without opening it.</div><div class="sf-spec__why">The same claim with its reason attached.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Filler adjective</div><div class="sf-spec__t">a robust retry on OTP fetch</div><div class="sf-spec__why">A retry is a retry. "Robust" says the author wanted the sentence to sound engineered.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Names the behaviour</div><div class="sf-spec__t">a retry that survives a mail-slot timeout</div><div class="sf-spec__why">Now the reader knows what it does. Keep "robust" only where it is the term of art.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Empty content under a fine heading</div><div class="sf-spec__t">The walls are the pattern.</div><div class="sf-spec__why">No ban list names this. It is not a repeated shape, only a sentence with nothing in it.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Adds a fact</div><div class="sf-spec__t">Three of seven accounts cannot reach a cloud-only runner at all.</div><div class="sf-spec__why">Same slot in the paragraph, now carrying a count the reader can act on.</div></div>
</div>

<div class="sf-spec">
  <div class="sf-spec__side sf-spec__side--n"><div class="sf-spec__h">Declared count</div><div class="sf-spec__t">Three causes sit under this. [...] The deck showed a fourth cause.</div><div class="sf-spec__why">The taxonomy is invented, then has to be patched when a fourth thing arrives.</div></div>
  <div class="sf-spec__side sf-spec__side--y"><div class="sf-spec__h">Just the causes</div><div class="sf-spec__t">Most text about work on the internet is business prose. Preference tuning then rewards the look of completeness. And words cost the model nothing.</div><div class="sf-spec__why">Each paragraph carries a mechanism. Nobody has to count them.</div></div>
</div>

## Current setup

My rules file now opens with the target: optimize for information gain over apparent completeness, and compose as finding, evidence, reasoning, decision rather than principle, framework, exposition, takeaway. The ban list is still there, as examples of the target being missed.

After the deck the agent saved a rule that lists the places the quality bar applies: slide titles, kickers, tile and card labels, table headers and row keys, diagram node and edge labels, chart titles and legends, button copy, subtitles, Slack, email, Jira, PR bodies, commit messages, code comments, memory files, and the content under all of them.

The banned shapes carry verbatim examples from my own outputs. "One harness, four swappable sides" is in the memory file as the example of number pairing, "The wedge" as metaphor in place of the noun, "Why hard apps are the niche" as question-as-title, "The fix was one line." as the verdict sentence. A rule that quotes the exact string the agent produced last Tuesday is harder to route around than an abstract one.

Rules in a file are advisory, and the deck showed what advisory gets you, so the last step was a scanner that runs whether or not anyone remembers it. Before the agent writes a file, publishes a page, or sends a Slack or Confluence message, the script reads the outgoing text and refuses the call if it contains a hard tell, returning the hit list so the rewrite is targeted. The same script runs on the agent's chat replies. Hard tells refuse on one hit: the em dash, honesty framing, "not X, it's Y" and its trailing cousin "X, not Y", headings that name their own rhetorical job, "That is the point" closers, sycophantic openers like "great question". Filler adjectives are counted and refused at three in one piece, because three is where use turns into decoration. Text inside quotes or code formatting is skipped, so a post like this one, which quotes the tells as exhibits, can still be written.

Both halves are now skills in [COG](https://github.com/huytieu/COG-second-brain), the public version of my agent setup, so they work without my vault:

- **[slop-gate](https://github.com/huytieu/COG-second-brain/tree/main/skills/slop-gate)** is the scanner. `python3 scan.py DRAFT.md` prints the tells and exits non-zero, so it runs as a CLI, a CI step, or a Claude Code hook with `--hook`. The pattern list, the quoted-span exemption and the `slop-ok` escape hatch are in the skill.
- **[voice-baseline](https://github.com/huytieu/COG-second-brain/tree/main/skills/voice-baseline)** is the census. Point it at a directory of your own writing and it reports the per-file rate of every word you over-use, plus the shapes: verdict endings, kicker closers, heading first words. `--check DRAFT.md` prints only what sits above your own baseline. Every number in the section above came out of it, including the correction to my grep.

The [no-ai-slop](https://github.com/huytieu/COG-second-brain/tree/main/skills/no-ai-slop) skill they sit beside picked up the reader-cited ranking, the list of media, and the voice-mode caps in the same release.

The hook fired on its own author within the hour. My agent's reply describing the hook listed the banned words in a sentence, and the gate refused the reply until the list was moved into quotes. A few messages later I caught it writing "reads as a test, not a permission", the trailing contrast form, which the first version of the script did not cover, and the pattern went in.

A model from a different family reviews what the lead agent writes. Same-family reviewers share the author's training distribution and therefore the author's sense of what a finished answer looks like. On this post, the reviewer found 57 problems in a draft I thought was finished, 6 of them the exact verdict sentence the post complains about.

Each catch becomes a memory file the agent reads at the start of the next session. The deck produced three of them: titles name the subject, the bar applies to every medium, the bar applies to content as well as labels.

## Alt text, issue descriptions, memory summaries

The drive to resemble a finished answer is what preference tuning selects for, and every lab is doing preference tuning, so the forms will keep moving to whatever surface I have not written a rule for yet. My guess for the next one is somewhere I do not read carefully: alt text on images, the descriptions on GitHub issues, the one-line summaries my agents write for their own memory files. The hook will not catch those until someone names the shape, so the first one through will be a human reader, probably my tech lead, who reads the deck I send him and tells me which slide he could not find.
