Fable 5.1, Opus 5 and Opus 5.5 on a Month of My Own Work
I measured 114 Claude Code sessions of my own work to see what Opus 5.5 changed, including how often it breaks the anti-slop rules I set in September.
Last Thursday I noticed my agent had stopped using em dashes. I had a hook that blocked them, a rule in every config file, and a whole post about how banning them pushes the slop somewhere else. Then Opus 5.5 arrived and the dash count dropped to nearly zero by itself: none in chat, 0.6 per 10k words in files.
In that post I ended up with a rules file, a hook that refuses any write or reply containing a hard tell, and a guess that preference tuning would keep moving the slop to whatever I hadn't written a rule for. My first reaction to 5.5 was to rewrite my anti-slop skill for the new model. My second was to wonder what else had changed that I hadn't noticed, because a model that gets past a punctuation tic has probably changed in ways that matter more.
So I counted. Everything I do at work runs through Claude Code on one Mac: Jira and GitHub hygiene, slide decks, meeting notes, customer POC docs, the odd side project. Every session leaves a JSONL transcript in ~/.claude/projects/. Each assistant message records which model wrote it. A month of those files covers three models on the same person doing the same kind of work.
model dates in my logs lead sessions assistant msgs
------------- --------------------- ------------- --------------
Fable 5.1 Sep 03 - Sep 24 51 3,270
Opus 5 Aug 28 - Sep 23 44 3,417
Opus 5.5 Sep 24 - Sep 28 19 1,895
Opus 5.5 has four days of data. Keep that in mind for every number below. Four days is enough to see big shifts and nowhere near enough to trust small ones, so I'll flag which is which.
How I counted
A Python script walks every transcript, dedupes repeated lines, attributes each response to its model, and matches each tool call to its result so errors land on the model that made the call. A session belongs to a model when that model wrote at least 80% of its responses. The counts stop at 11:00 on September 28, Vietnam time. I left out the session in which I wrote this post, since it quotes every tell I was counting.
A few definitions carry most of the weight:
- Human prompt: a message I typed. Hook feedback, teammate messages from other sessions, and skill preambles are excluded. My first pass counted the slop hook's "rewrite this" messages as corrections from me, which made Opus 5 look seven times worse than it was.
- Active time: the gaps between consecutive transcript events, summed, dropping any gap over ten minutes so that lunch and meetings don't count.
- Tool error: a tool result flagged
is_error. That covers failed shell commands, edits whose target text didn't match, and calls a hook denied.
None of this measures quality. It measures how much work a model does for each thing I ask, how often that work fails, and what the text looks like. Whether the output was right is something I only know from using it, and I'll get to that.
More work per request, in less time
This was the largest change, and the one I hadn't noticed while using the model.
tool calls per active minutes tool error
human prompt per human prompt rate
Fable 5.1 15.2 ######## 11.4 ########### 9.8% ##########
Opus 5 18.8 ########## 9.3 ######### 6.0% ######
Opus 5.5 24.2 ############ 7.9 ######## 5.3% #####
For each thing I ask, Opus 5.5 makes about 60% more tool calls than Fable 5.1 and finishes in 30% less time. Its tool error rate is a little over half of Fable's. Shell errors follow the same order: 11.7% for Fable, 6.5% for Opus 5, 5.8% for Opus 5.5.
I counted prompts per session to check whether I simply asked 5.5 for smaller things. The counts are nearly flat: 4.4, 4.0 and 4.1. I opened the same number of turns with each model, and 5.5 did more inside each one.
I think the gap in time comes from how little 5.5 stalls. Fable 5.1 reasons well and spends a lot of wall-clock doing it, one in ten of its tool calls fails, and each failure costs a round trip to diagnose. Fable also runs into session limits. My notes from August record a 404-account enrichment fan-out under a Fable lead that burned 6.5M subagent tokens and finished 106 accounts before the limit stopped it. Since then my routing rules pin every fan-out to Sonnet.
I tried to line up matched tasks across models and learned something about my own method. Sessions don't hold one task each. The Opus 5 session that opened with "move tickets from the Sep 17 release to Sep 23" made 78 tool calls, and I first read that as 78 calls to move tickets. When I opened the transcript, the move took six: list milestones, find gh, list the open issues, move them, confirm. The other 72 were a meeting transcription, a Kai memory ticket, and a design review I had piled into the same window. The Opus 5.5 session that created this week's release and moved the unfinished work took five.
On a small, well-defined task the two models behave the same. The gap in the averages comes from the long turns: the ones where I paste a Slack thread and a spreadsheet and say "see what is still missing", and the model has to decide on its own what to read, what to change, and when it's done.
A bigger working set
Opus 5.5 carries more context and writes more per message.
context per msg output tokens msgs with a
(tokens) per msg thinking block
Fable 5.1 244k 805 54%
Opus 5 207k 712 52%
Opus 5.5 334k 949 68%
Context per response is 37% above Fable's and 61% above Opus 5's, and two thirds of its responses start with a thinking block. Part of that is session length and part is how much it reads before it acts. My guess is this explains some of the lower error rate: when the model has read the file, the schema or the milestone first, its first write is more likely to land. I can't prove that from transcripts alone.
Delegation
My CLAUDE.md has a delegation cap: don't spawn a subagent for work the lead can finish in a few tool calls. I wrote it because Opus 5 over-delegated. By count, Opus 5 was the most restrained of the three.
Agent spawns per 100 human prompts
Fable 5.1 52 ##########################
Opus 5 39 ###################
Opus 5.5 51 #########################
words per text block in a subagent's output
Fable 5.1 11 #
Opus 5 31 ###
Opus 5.5 152 ###############
Opus 5.5 spawns about as often as Fable. What changed is what comes back. Fable ran only three subagent transcripts in this window, so its row in the second chart is thin, but the Opus comparison has 211 and 28 subagent sessions behind it. Opus 5 subagents returned terse status lines. Opus 5.5 subagents return full reports, five times longer per block, and read far more before writing: Read is 26% of their tool calls against 14% for Opus 5 subagents.
I haven't decided whether that's good. The reports are useful, and the time-per-prompt numbers say the delegation isn't slowing anything down. My Worker Output Rule asks workers to write anything long to a file and return a path, and a 152-word average per block suggests some of them are skipping that. I need to check before blaming the model.
Parallel tool calls in one response, the thing Fable does best, went the other way: 14.5% of Fable's tool-calling responses fire several tools at once, against 4.0% for Opus 5 and 6.1% for Opus 5.5. Fable still orchestrates wider.
Against the September rules
The September post ended with a gate: a hook that reads Markdown and HTML file writes, published pages, Slack and Confluence sends, and the last chat reply of each turn, and refuses the call on one hit of a hard tell. The hard tells are em dashes, honesty framing, "not X, it's Y" and "X, not Y", headings that name their own rhetorical job, "The <Noun>" headings, emoji headings, "That is the point" closers, sycophancy, throat-clearing, and recap endings. Filler words are a softer rule: one or two pass, three or more in one piece get refused. It is the most direct way I have to ask whether a model follows my rules, so I replayed every response through the same script.
One confound first. The rules file and hook were rewritten on September 14, the day of that post, and the filler-word check was tuned again on the 17th. Opus 5 ran from August 28, so half its month was under older rules. The table below starts on September 15, which puts all three models under the same rules file and the same hook. The transcript keeps the model's first attempt, before the hook refused it, so this measures what the model wrote, not what I let through.
share of outputs the slop gate would refuse
Sep 15 - Sep 28, same rules file and hook for all three
chat replies Markdown files
refused (n) refused (n)
Fable 5.1 2.4% (331) 42% (12)
Opus 5 30.5% (620) 42% (38)
Opus 5.5 3.2% (593) 18% (34)
Against Opus 5, the gap is large: 30.5% of Opus 5 chat replies broke at least one hard rule, and 3.2% of Opus 5.5 replies did. Almost all of the Opus 5 failures were two rules, the em dash in 23.6% of replies and "X, not Y" in 13.2%. Opus 5.5 wrote zero em dashes in 593 replies.
Against Fable, the chat numbers are level. Fable was already at 2.4%, and I had been blaming "the model" for tells that came almost entirely from Opus 5. Files are where 5.5 pulls ahead of both: 18% of its Markdown writes would have been refused, against 42% for the other two. The file counts are small, 12 for Fable, so I'd hold that one loosely.
per-rule hit rate, % of outputs, Sep 15 - Sep 28
chat files
Fable Opus5 Opus5.5 Fable Opus5 Opus5.5
em dash 0.0 23.6 0.0 16.7 23.7 2.9
"X, not Y" 2.4 13.2 3.2 33.3 36.8 17.6
honesty framing 0.0 0.2 0.0 0.0 0.0 0.0
3+ filler words 0.0 0.2 0.2 0.0 0.0 0.0
emoji heading 0.0 0.0 0.0 0.0 0.0 2.9
ends on a question 0.6 2.6 1.4 0.0 0.0 2.9
uses "ship" 0.9 4.2 1.7 0.0 21.1 23.5
Several rules from September never fired for any model in this window: rhetorical headings, "The <Noun>" headings, verdict kickers, fake-profound closers, sycophancy, throat-clearing, and recap endings. Those tells were real in the older transcripts I quoted in September. By mid-September the rules file had them spelled out with examples from my own outputs, and all three models stopped. They don't separate the models any more.
What Opus 5.5 still does is "X, not Y", in 3% of replies and 18% of files. About half of those hits are precise disambiguation, like "measured after the change, not after it". My rules ban the shape outright, and after reading these I'm less sure they should. The other open item is "ship": my rule reserves it for software releases, and 24% of 5.5's files use it. Some of those are about releases. I haven't hand-checked how many.
Punctuation moved in a direction my rules didn't cover.
per 10k words, Markdown the agent wrote to files, whole month
em dash "X, not Y" semicolons parentheses
Fable 5.1 18.1 8.3 133 215
Opus 5 38.3 4.8 75 102
Opus 5.5 0.6 1.8 101 176
Opus 5.5 writes about one parenthetical every 57 words in files, dates and counts and qualifiers packed into brackets, and chains clauses with semicolons, mostly inside table cells. Its chat replies are built from bullets, 144 bullet lines per 10k words against 39 for Opus 5. This is the move the September post predicted: take the dash away and the same compression shows up in the next mark over. On dashes and parentheses Opus 5.5 looks more like Fable than like Opus 5, though its semicolon rate sits between the two.
I updated my anti-slop skill for this today. It tells the model that the September tells are mostly handled by the hook, and to check its own output for parenthetical stuffing, semicolon chains, bullet-and-bold replies, and "So the plan is:" colon lead-ins. The gate itself stays as it is. A hard block on parentheses would fire on every date and ticket ID.
What didn't change, or went the wrong way
Human corrections, meaning prompts where I start with "no", "wrong", "why did you" and the like, rose from 1.8 per 100 prompts on Fable to 2.9 on Opus 5 and 3.8 on Opus 5.5. For Opus 5.5 the pattern matched three messages. Reading them, one was a duplicate of another, which leaves two corrections. One was scope: I asked for release hygiene and it started adjusting demo videos when I only wanted a line in the Confluence release note. The other was style: it labelled workflows "WF1" and "WF2" in a doc, which nobody but the two of us could follow. Neither is a failure I'd pin on the model, and two is not a trend.
Interrupts, where I hit Escape mid-turn, were 3.8 per 100 prompts on 5.5 against well under one on the others. That's three events in two sessions, too few to read anything into.
Changes to my setup
before (Sep 17) after (Sep 28)
--------------------------------- ---------------------------------
lead: Fable 5 for fresh sessions lead: Opus 5.5
fan-out: Sonnet, pinned fan-out: Sonnet, pinned
review: Codex CLI review: Codex CLI
slop focus: dashes, contrasts slop focus: parens, semicolons,
bullet replies
For October I'm running Opus 5.5 as my default lead. The case is the first chart: more done per request, fewer failed calls, less time. Fable 5.1 stays for the rare job that needs wide parallel fan-out managed by the lead, where it's still the strongest of the three. It also has the tightest limits, so it no longer gets the everyday work.
I'll rerun the same script at the end of October with a month of 5.5 data. The four-day numbers could shrink once the novelty wears off and I start handing it the tasks I used to save for Fable. Or they could hold. The script, model-compare.py, is in my vault, and the transcripts accumulate whether I look at them or not.