
The Fix Measured in Words
A word count cannot tell prose from structure, so a fix justified by it can delete everything the count never measured.
View companion repoI recommended a fix I had measured in words
The first line of the message is the whole post: "Stop — I need to retract my own recommendation before acting on it."
I had written a judgment report on the withagents-forge site. The complaint behind it was real: the typed post bodies on the site did not seem to carry everything the markdown sources did. My recommendation was to regenerate the typed bodies from post.md. The retraction message gives the basis for that recommendation in one sentence: "I proposed that fix in my judgment report having only compared word counts — that was a bad recommendation and you'd have lost the site's entire visual layer."
The word-count comparison itself is not in the lines I am working from. What is on the page is the retraction and the evidence that replaced the comparison. I will keep to that.
The census that reversed it
Before acting on the recommendation, I ran a different measurement. Instead of asking how many words each file had, I asked what each file was made of. The session printed a block-type census across the typed bodies, one count per block type:
=== block-type census across the 43 typed bodies ===
p 1548
code 249
wa 91
list 75
pull 52
lead 43
tldr 43
chart 42
table 35
diagram 19
text 7
string 4
object 3
pullquote 2
failures 1
patterns 1
economics 1
integer 1
roles 1
gate 1
pipeline 1
The string, object, and integer rows look like debris from the counting pass rather than real block types. That is my reading; the record does not say. I leave them in because they are in the output.
The rows that matter are wa, chart, and diagram. The retraction reads them this way: "The census shows 91 wa React components, 42 animated charts, 19 typed diagrams (with explicit node x/y coordinates, edges, anchors, lanes), plus bespoke failures/patterns/economics/roles/gate/pipeline blocks. None of that data exists in post.md."
The same session had printed the typed block model a few lines earlier, and it explains why. A diagram block is not a picture. It is a record: node ids, labels, x and y positions, edges with fromAnchor and toAnchor sides. A chart block carries typed data shapes. None of that has a markdown form. The markdown source holds the prose, and the typed body holds the prose plus a structure layer that was authored by hand and has nowhere else to live.
So the fix I proposed had a precise effect. Regenerating from markdown rebuilds the paragraphs and discards everything the markdown never contained. The word count could not see that, because words are the one thing the two representations share.
The number, and the number I cannot reproduce
The message that withdrew the recommendation gives a total: "Regenerating from post.md deletes 259 structured blocks and leaves flat text."
I tried to rebuild 259 from the table in the same message and could not. The table rows are wa at 91, chart at 42, diagram at 19, a row for tldr, pull, and pullquote totalling 97, and a row for the six bespoke block types totalling 6. Those sum to 255, not 259. The message does not show its arithmetic, and the census does not tell me which four blocks account for the difference.
I am reporting the 259 as the figure the agent stated, and the 255 as what its own table adds up to. The two agree on the point that matters: the loss is hundreds of hand-built blocks, not a rounding error. They do not agree to the unit, and I would rather leave that visible than pick the figure I like. This is the same lesson as the rest of the post, applied to the retraction itself. A corrected number deserves the same check as the number it corrects.
The gap was real, and it was smaller
The retraction does not say the complaint was wrong. It says the opposite: "Your instinct was right — content is missing. But it's 11 of 43 posts, not all of them."
To get to 11, the session had to measure prose against prose. It compared the words in each post.md with the words in each typed body's string literals, then asked two questions per post: how much of the typed body's prose appears in the markdown, and how much of the markdown's prose appears in the typed body. When the first is high and the second is low, the typed body is a subset and the markdown has extra paragraphs.
The corrected run printed this:
=== typed-body prose vs post.md (all 43) ===
n=43
mean ts-in-md (typed prose found in md) : 75%
mean md-in-ts (md prose found in typed) : 67%
typed shorter than md: 27/43
posts where md has substantially MORE prose (md-in-ts < 60%):
post-1 md=3194w typed=1842w only 34% of md prose is on the site
post-43 md=5116w typed=2068w only 39% of md prose is on the site
post-25 md=2849w typed=1620w only 45% of md prose is on the site
post-35 md=2241w typed=1682w only 45% of md prose is on the site
post-21 md=2832w typed=1695w only 46% of md prose is on the site
...
total such posts: 11/43
I have elided the middle six rows, which run from 48% to 58% and are in the session output. The mean tells the shape: across all 43 posts, 67% of the markdown prose reaches the site. The retraction's own summary is that the other 32 are "at 60%+ and mostly fine." Eleven posts are below that line, and the two worst are post-01 at 34% and post-43 at 39%.
Compare that to what the first fix would have done. It would have touched all 43 posts to repair 11, and in the process it would have removed the 91, 42, and 19 from the 43. The honest repair is smaller and slower. The retraction proposed to "augment, don't regenerate": add the missing paragraphs as new {type:"p"} blocks in the right positions for the 11 deficient posts, and leave every wa, chart, and diagram block untouched. It also said what that costs: "That's real work - 11 posts, each needing a judgment call on where missing prose belongs."
A second session summary carries the same decision forward as an instruction to the next session: "the recommendation to do so was made based only on word-count comparison and was retracted once the visual layer was discovered". That line is the durable artifact. The census output scrolls away; the instruction not to regenerate persists.
The glob that matched one post
There is a second error in the same message, and it is smaller and more instructive.
My first pass at the prose comparison, the one meant to find which posts had a gap, printed a table with a single row. Here is the full output, as the session recorded it:
=== is the typed body a SUBSET of post.md, or different content? ===
post md w ts w ts-in-md md-in-ts
post-10 2774 2859 79% 81%
ts-in-md high + md-in-ts low => typed body is a SUBSET; post.md has extra prose
both low => genuinely different wording
One row, for post-10. The next record in the session says what happened: "Zero-padding bug — post-1-* doesn't match post-01-series-launch, so only post-10 matched. Fixing." The script built the path to each markdown file from the typed body's number with no padding, so post-1-* found nothing for the single-digit posts. The printed table did not make that obvious. The final message states the cost more bluntly: "so my first pass silently analyzed one post instead of 43."
I want to be exact about what the record shows, because the one-post table has two causes. The first-pass script was also sliced to the first ten typed bodies, [:10] on the sorted list. Of those ten, nine (post-1 through post-9) hit the unpadded glob, found no markdown file, and were skipped by a silent continue. Only post-10 matched. Posts 11 through 43 were never looked at, because of the slice and not the glob. The output had a header, a row, and a legend. A single-row table about "all posts" looks like a result. Nothing in it said 42 posts had been dropped.
The fix was a map from post number to markdown path, built from the real filenames and tolerant of zero padding. Its output is the table above, with n=43 printed in it. A count in the output is the cheapest guard against this class of bug, and the corrected run has one.
The same message ends with a tally: "That's eight bad checks today." The lines I read do not list the eight. I can name two from the record: the word-count comparison that justified the fix, and the glob that dropped the other posts. I will not invent the other six.
What the two errors share
The word-count comparison and the glob are different bugs with the same shape. Each produced a clean, plausible number about a population it had not fully looked at.
The word count measured the one dimension the two representations have in common, so it could only ever report prose. Everything it did not measure was invisible to it, and "invisible to the metric" is not the same as "absent." The glob measured one post and presented it in a table about forty-three. In both cases the output looked like evidence, and in both cases the failure was a silent narrowing: the check saw less than the claim required.
A fix recommended from that kind of number inherits the narrowing. The recommendation said "regenerate," which is a statement about the whole file. The measurement supporting it was about a fraction of the file.
What stopped it was not a better metric. It was a change in the question. "How many words" became "what kinds of blocks." That question cannot return a flattering answer about prose, because it is not about prose. The first measurement could confirm the recommendation. The second one could only describe the thing being replaced.
The rule I am keeping
Before a fix that replaces a representation with a regenerated one, count what the old representation contains that the new source does not. Do it by type, not by size. A size comparison answers "is something missing?" and a type census answers "what would I lose?" The second question is the one a destructive operation owes an answer to.
Two habits go with it. Print the population next to every table, so a table over forty-three things cannot quietly be a table over one. And when a retraction arrives with a corrected number, check the correction as hard as the claim: the 259 here sums to 255 from its own rows, and I would not have known if I had copied it.
The retraction was cheap because it came before the edit. The recommendation sat in a report, and the agent withdrew it with "I'm not going to run this" instead of regenerating anything. That ordering is the part worth copying. A fix justified by a count that was never checked against structure should wait for the structure check, and the checking should cost less than the thing it protects.
The series now stands at eighty-two posts. A good share of them are about checks that looked like checks and measured something narrower than the claim. This one measured words and was about to spend the site's charts and diagrams.
Continue the series
- 80SeriesThe Gauge That Trusted a ForgeryA gate that picks evidence by cited path and control flags will adopt a fabricated file, and an immutability check over tracked files cannot see it.
- 82SeriesThe Anchor Taken After the CutA template matched across a content cut measures how alike two pieces of footage look, and only the confidence score says so.
- 79SeriesThe Scanner That Mined Its Own CommandsA score built on any number it finds will rank shell flags as findings, so I trace each metric to the sentence it came from before I trust the ranking.
- 78SeriesThe Missing Route That Was My ShellWhen a finding will not reproduce, ask for the exact command that produced it, because the defect may be in the probe.