
The Scanner That Mined Its Own Commands
A score built on any number it finds will rank shell flags as findings, so I trace each metric to the sentence it came from before I trust the ranking.
View companion repoThe run that exited 0
The mining run finished with no errors, and none of the eleven candidates it queued was usable. That is the agent's own summary of the run, and it is the sentence I most want on the page: "The mining run finished with no errors, but none of the 11 candidates it added to the queue is usable as a post. Every headline metric is a stray number pulled from log noise, not a result from the work."
This post is about the tool that picks the topics for this series. The candidate miner reads my Claude Code session logs, clusters what it finds, scores the clusters, and writes the top ones into a queue of candidate posts. So this entry is self-referential: the pipeline that proposes topics for the field journal produced a batch of candidates that were all empty, and the exit status gave no hint of it.
The run's own log line was clean:
mine_complete run_id=20261002T203123Z-mine candidates=11 reason=ok tier=1
The agent's report adds that it scanned 1,115 sessions, scored 53 clusters, queued the top 20% at tier 1, took 15.8 seconds, exited 0, and wrote nothing to stderr. By every signal a script emits, that is a good run. The scores on the queued candidates ran from 0.82 to 0.92, which reads as strong.
What the numbers actually were
The candidate schema requires a headline metric. The agent traced each of the eleven back to the text it was extracted from. This is the table it produced, verbatim:
| Metric | Where the number really came from |
|---|---|
| 260 | a hook's `"durationMs":260` |
| 120 | the milliseconds in a timestamp (`06:46:43.120Z`) |
| 400 | an `API Error: 400` status code |
| 256 | `sha256sum` |
| 100 | the middle of a base64 blob |
| 12.0 (×2) | the version tag `v1.12.0` |
| 1838, 0416 | plan folder names (`260903-1838-…`, `260906-0416-…`) |
| 200 (×2) | a frame range (`f000100`–`f000200`) and a stray `200` in a prompt |
Eleven numbers, and none is a result. A hook's latency, the 256 in sha256sum, an HTTP status, a version string, a clock fragment, a folder name. The extractor had a regular expression that took any number of three or more digits. It did not ask what the number counted.
The same report records two more facts about those eleven candidates. All of them had zero evidence anchors and a TODO placeholder thesis. None came from the blog-series or WithAgents projects. Those are the fields a human would look at first, and the queue accepted the rows anyway.
Why the score said otherwise
The scores were high because the scorer rewarded the noise. The agent's explanation was specific: the scores of 0.82 to 0.92 "look strong only because 'headline metric strength' gives these noise numbers full marks."
One component of the score asked whether a cluster had a headline metric, and how strongly. A number was a number. 260 from a hook payload got the same credit as a measured result would have. The queue takes the top 20% of the scored clusters, so a bad field changes which rows rise, not how many.
This was not the first time. The agent's report says an earlier scorer fix, commit b0b7cd0 from the 14th of September, stopped the miner treating dates as metrics. This run showed what that fix left behind: durations, status codes, hashes, version strings, plan-folder timestamps and base64 text. The report adds that the extra-wide 1000-hour window pulled in more of all of them. That is the agent's account of the cause; the record does not show the window change tested in isolation.
The second bug was in the citations
Wrong metrics would have been caught by a reader eventually. The part that would have survived a skim was the citation. A candidate must cite a session line that states its metric. The citation lookup searched the raw JSONL line for the bare digits. So 260 was "confirmed" by finding 260 inside "durationMs":260, and the quote attached to the candidate was a fragment of a JSON blob.
The agent summed it up in the later report:
The miner picked up any 3-digit-or-longer number, including ones inside shell
commands and agent instructions, with nothing to stop it matching the middle of
an identifier. All 11 values from the first run were noise, for example
`head -200`, `shasum -a 256`, `git rev-parse v1.12.0`, a plan-folder path and
`API Error: 400`. Separately, the citation lookup searched the raw log line for
the bare digits, which is how `260` got cited from `"durationMs":260`.
So the validation step and the extraction step shared the same blindness. The extractor said "260 is a metric", and the verifier agreed because the string 260 was present. A check that can only find what the extractor already matched is not a second opinion.
The fix, and the leak I caught in it
The rewrite is commit 90e1a24, "fix(mining): extract real metrics and mine only on-project sessions." The core is a regular expression that refuses to see a number unless something is attached to it:
_METRIC_RE = re.compile(
r"(?<![\w.\-/:#@=$^\[])"
r"(?P<num>\d{1,3}(?:,\d{3})+|\d+\.\d+|\d{3,})"
r"(?:(?P<pct>%)|(?P<mult>x)(?!\w)|[ \t]+(?P<unit>[A-Za-z][A-Za-z-]*[A-Za-z])(?![\w/-]))"
)
# HTTP/exit statuses read as "400 messages" or "500 errors" -- never a finding.
_STATUS_CONTEXT_RE = re.compile(r"(?:error|status|http|code|returned|exit)\W{0,3}$", re.I)
The lookbehind rejects numbers glued to a path, a version, a flag, a key or a timestamp. The tail demands a percent sign, an x, or a unit word. A status-context pattern drops 400 when it follows "error". The commit message lists the other rules: stopword units are rejected, as are 0% and 100%, because those read as pass or progress states. Metrics may come only from prose text blocks, not bash commands or agent prompts, and corroboration across sessions requires the same number and the same unit. Citations now quote the sentence that states the metric.
Midway through writing this, the agent found a hole in its own pattern. It said so plainly:
One leak: `shasum -a 256 plugins/proofpunk` treated `plugins/…` as the unit. A
unit word that runs straight into `/` is a path, not a noun, so I'll reject that
case.
That is the same 256 from the noise table, coming back as a number with a "unit". In the committed pattern the trailing (?![\w/-]) is what refuses a unit that runs into a slash. The session shows the test. Before the change, the check printed noise leaking: [('shasum -a 256 plugins/proofpunk', [('256 plugin', 10)])]. After it, 'shasum -a 256 plugins/proofpunk' mapped to [], while 'found 256 plugins.' still came out as ['256 plugin'], so a real count of plugins is still accepted.
The commit also adds a constraint that has nothing to do with numbers. The miner now reads only project directories whose names contain blog-series or withagents, because off-project candidates were being queued and then batch-rejected. The agent's report says none of the eleven first-run candidates came from the blog-series or WithAgents projects, so a cleaner extractor alone would still have filled the queue with posts about the wrong project.
The re-run
The agent re-ran the miner under the same 1000-hour window. Its table:
| | Before | After |
|---|---|---|
| Clusters with a confirmed metric | 19 of 53 | 3 of 53 |
| Queued PRDs with a headline metric | 11, all noise | 2, both real |
| Top score | 0.924 | 0.755 |
Nineteen clusters had counted as confirmed before the extractor fix and three after, in run 20261002T233629Z-mine. The top score fell from 0.924 to 0.755, and that drop is the honest number: the old top score was the highest-ranked piece of noise. The re-run scanned 1,082 sessions with no errors, against 1,115 in the first. The record I read does not explain the difference, and I am not going to guess at it.
The two surviving metrics were "192 endpoints" from awesome-list-site and "243 session digests" from yt-shorts-detector. The agent graded the second one itself: its second citation quotes a workflow description that was pasted into the session, not an independent result, "so it's weaker evidence than it looks." The other nine candidates were queued with no headline metric, which the PRD schema allows. An empty field is a correct output when nothing was measured.
The agent moved the eleven noise PRDs into .candidates/rejected/ with a reason note rather than deleting them. They shared topic identifiers with the new run, and the duplicate check would otherwise have kept the bad versions. It also said what had not changed: all eleven new candidates were about other projects, so the publishing step still had nothing on-project to work on. The change was not committed at that point. A later run, 20261002T234450Z-mine, added the on-project filter, read 42 transcripts from the blog-series and withagents-forge project directories, scored 6 clusters and queued 2 candidates at 0.646 and 0.649, neither with a confirmed metric. Both changes were then committed together as 90e1a24. The 0.755 and 3-of-53 figures above belong to the earlier run, which had no on-project filter.
What I would not claim
I did not measure how many of the earlier series topics were chosen from a noise metric. The record covers this one run, so I will not extend the finding backward. What I can say is that the same miner produced earlier queues, and an earlier fix (b0b7cd0) had already shown that the metric component misfired on dates. That is a pattern of the metric field being weak, and it is not a count of bad posts.
I also have no test in the record that proves the new pattern keeps out every shape of identifier. It rejected the cases in the table and the one leak the agent found. The next noise shape is unknown to me by definition.
The rule
A score is only as good as the least checked field feeding it. Here that field was a number the extractor accepted on shape alone, and the verifier accepted on the same shape.
Three habits come out of it. When a pipeline exits 0 and reports a count, open one output row and trace one value back to its source text before reading the ranking. When a verifier looks for a value, make it look for the claim the value is supposed to support, not the string. And make the empty answer legal: a field that cannot be left blank will be filled with whatever is nearest, and the nearest thing in a log is usually a command.
The tool I trusted to find the stories could not tell a result from a flag, and I only learned that by opening its output.
Continue the series
- 78SeriesThe Missing Route That Was My ShellWhen a finding will not reproduce, ask for the exact command that produced it, because the defect may be in the probe.
- 80SeriesThe Gauge That Trusted a ForgeryA gate that picks evidence by cited path and control flags will adopt a fabricated file, and an immutability check over tracked files cannot see it.
- 77SeriesThe Deploy That Changed NothingA deploy command's exit code and SUCCESS banner describe the control plane, and only the running service's own version endpoint says what is actually serving.
- 81SeriesThe Fix Measured in WordsA word count cannot tell prose from structure, so a fix justified by it can delete everything the count never measured.