
The Gauge That Trusted a Forgery
A gate that picks evidence by cited path and control flags will adopt a fabricated file, and an immutability check over tracked files cannot see it.
View companion repoThe gauge said 6/6 PASS and pointed at a file I had made up
The line that matters is two lines long:
measured: 6/6 status: PASS
evidence: e2e-evidence/tmp-c5-adversarial/command-surface-proof.json
The directory name tells you what it is. tmp-c5-adversarial was not produced by the plugin under test. A verifier session wrote it, by hand, in a Python heredoc, with every field filled in to look like a green run. The gauge read it, scored it, and cited it as the proof that six commands work end to end at the real slash-command surface.
I want to be exact about the framing, because it is easy to misread. Nobody attacked anything. This was an independent-verifier session on /Users/nick/proofpunk, and the forgery was a test the verifier ran on purpose. The handoff it was working from listed a check called C5, and the text of C5 reads: "ADVERSARIAL: try to make gauge #4 report a PASS it should not. Construct a command-surface-proof.json reporting 6/6 whose control arm passed, place it in the evidence tree, and confirm the gauge REJECTS it as vacuous. Remove it afterwards."
So the plant was in scope. The part that was not in the script is what happened when the verifier pushed one step past it.
The first half worked
The verifier wrote six commands into a proof file. Every entry claimed the plugin arm passed. Every entry also claimed the control arm passed, which is the tell: a control arm is the run without the plugin, and if it passes too, the result says nothing about the plugin. Each pass flag carried the marker FAKE-MARKER. This is the construction, copied from the session at line 384:
cmds.append({
"command": name,
"reached_level": "c",
"max_honest_level": "c",
"plugin": {"pass": True, "marker": "FAKE-MARKER"},
"control": {"pass": True, "marker": "FAKE-MARKER"},
})
doc={"measured_at":"2026-09-14T23:59:59Z","harness_errors":[],"commands":cmds}
The measured_at value is 2026-09-14T23:59:59Z, the last second of a day. The session's own summary, written later (line 546, repeated at 605), records the consequence: the timestamp was newer than the genuine artifact's 2026-09-14T15:40:18Z, "so the reverse-sorted scan visited it first." I read that as a consequence the summary reports, not as a motive stated at the moment of planting.
Then the verifier loaded the gauge module and asked its disqualifier directly. The disqualifier fired:
DISQUALIFIER: control arm PASSED for ['/proofpunk:implement', '/proofpunk:verify', '/proofpunk:install', '/proofpunk:forge-prompt', '/proofpunk:rate-prompt', '/proofpunk:truth-audit'] — those probes are vacuous, the plugin is not proven to be what produced the result
That is the behavior C5 wanted. A passing control arm voids the probe. The check works as written.
Next the verifier ran the whole report with the plant in place. It exited zero. The gauge-four entry still read measured: 6/6 status: PASS, but the evidence path was the genuine one, e2e-evidence/run-20260914T002104-v4-phase-reconcile/t10-v6-retry-fixed/command-surface-proof.json. The overall line read:
VERDICT: PASS — every gateable gauge meets target (7 PASS, 2 UNMEASURED and excluded)
So the gauge refused the forgery and kept citing the real artifact. The same run wrote zero bytes to stderr (the session printed stderr bytes: 0), so there was no warning anywhere that a hostile file had sat in the tree during the scan. A rejection that leaves no trace is easy to read as nothing having happened. If the session had stopped there, C5 would be a clean pass and this post would not exist.
The ablation
The verifier did not stop. Its note before the next step reads: "Decisive — but only if the plant would have won had it qualified. Ablation: same artifact, control arms flipped to failing."
The reasoning is sound and worth stating plainly. A rejection proves the disqualifier fired. It does not prove the disqualifier is the only thing standing between the forgery and the verdict. Maybe the plant lost for some other reason: a timestamp rule, a path rule, a missing field. The way to find out is to remove the one suspected cause and keep everything else identical. So the verifier rewrote the same file with a single change, every control arm set to {"pass": False, "marker": None}. The plugin arm stayed a lie. The marker on it stayed FAKE-MARKER.
Then it re-ran the gauge. This is the output, line 452:
ablation: control arms flipped to pass=false; everything else identical
rc=0
10: [PASS] #4 (L16) Commands proven end-to-end at the real slash-command surface: 6/6
measured: 6/6 status: PASS
evidence: e2e-evidence/tmp-c5-adversarial/command-surface-proof.json
The gauge adopted the forgery. Same six commands, same plant, one edit, and the evidence pointer moved from the genuine run to the file the verifier had typed by hand. The verifier's own summary of the result: "Ablation confirms causality — control arm is the only difference, and it flipped adoption. That also exposes something: the gauge trusted a wholly fabricated file."
Notice what C5 asked for and what it got. The scripted check expected a rejection and received one, so as written it passed. The result that matters came from a question the script did not contain, which is a reason to give a verifier room to ask the next question rather than only tick the listed boxes.
The causal claim needs a qualifier the verifier's summary supplies. The plant was not only a file with passing flags. Its measured_at made it newer than the genuine artifact, both scored 6/6, and the selection code says: "Highest full-chain count wins; ties break to the newer measurement". So adoption took two things: clearing the disqualifier, and winning the tie on recency. The control-arm flip was the only edit between the two runs, and it was necessary. The record does not establish that it was sufficient on its own, since the plant's timestamp was fixed in both runs. The gauge decides whether an artifact is acceptable by reading fields the artifact supplies about itself, and both the flags and the timestamp are such fields.
The backstop did not stop it
Proofpunk has a second line of defense for exactly this class of problem. tools/verify-evidence-immutability.py checks that evidence files already committed have not been altered since. If something rewrote a past capture to flatter a score, this is the tool that should object.
The verifier ran it with the plant still on disk. Line 469:
immutability rc=0
PASS evidence-immutability: 2091 committed captures byte-identical to HEAD
Read the sentence it prints. "2091 committed captures byte-identical to HEAD." That statement is true. Every one of the committed captures is unchanged. The plant is not among them, because it was never committed, so the check never looked at it. The tool answered a narrower question than the one I would have asked of it. I would ask: is any evidence in the tree fabricated? It answers: has any tracked evidence changed?
The verifier's later note states the finding in one line: "immutability gate returned rc=0 with the plant present — confirming untracked fabrications are outside its scope."
I do not read this as a bug in the tool. A check over tracked files is a reasonable thing to build, and its output line says what it covered. The problem is composition. The gauge selects evidence by path and by flags inside the file. The immutability check protects files git knows about. A new file in the evidence tree is invisible to the second and trusted by the first. Each tool is correct about its own scope, and the gap between the scopes is where the forgery lives.
The cleanup, and what it cost to verify
The verifier then did what C5 required: "Remove it afterwards." It deleted e2e-evidence/tmp-c5-adversarial, ran git checkout -- gauge-report.json gauge-report.md (rc=0), and regenerated the report. The regenerated gauge-four entry was back on the genuine artifact:
measured: 6/6 status: PASS
evidence: e2e-evidence/run-20260914T002104-v4-phase-reconcile/t10-v6-retry-fixed/command-surface-proof.json
The next step checked the removal directly instead of trusting the delete. A listing of e2e-evidence/ filtered for tmp-c5 came back with a count of 0. A git status pass followed. It showed four tracked files modified (gauge-report.json, gauge-report.md, tools/gauge-report.py, tools/verify-counts.py) plus several untracked evidence directories from the larger verification run. The same four files already showed as modified in the git status snapshot at the start of the session (line 345), and the handoff (line 328) notes that the earlier C2 run had rewritten both report files, so the status line says nothing new about the cleanup. What the record shows is the plant gone and the citation back on the genuine artifact. It does not show a diff of the report files against HEAD, so I do not claim them byte-identical to the committed copies.
"Restored" is a claim, and the strong form of it needs a diff that comes back empty. Here the weaker form is what the record supports.
What I take from it
First, a disqualifier and an acceptance rule are different things. The disqualifier here is a good one. It named all six commands and said why the probes were vacuous. But passing the disqualifier is not the same as being genuine. An artifact that clears the vacuity check can still be a file somebody typed. The ablation is what separated those two ideas, and it only existed because the verifier asked what would have happened had the plant qualified.
Second, a gate that chooses evidence by path and trusts flags written inside the evidence has no way to tell an honest file from a well-formed one. In my own earlier posts the failure was usually a check that could not fail. Here the check could fail and did. The hole was upstream of it, in what the gauge was willing to read.
Third, when a tool says "N things verified," find the set N is drawn from. This one says "committed captures," and that scope is printed on the line. Nothing about the PASS was misleading. It was my reading of it that stretched.
I have not changed any gauge or tool in response, and nothing in this post establishes what the right fix is. Candidates exist: provenance for who produced an artifact, or restricting the gauge to tracked files, or having the immutability check enumerate untracked files under the evidence tree and fail on any it cannot attribute. The session did not choose among them, and I am not going to choose on its behalf from a transcript.
The transferable rule is small. Test a gate with a forgery it should reject, then change one variable and test it with a forgery that is otherwise indistinguishable from a pass. If the second one gets in, the first rejection was the gate's only defense, and you now know what it is made of. Then check what your backstop covers, in terms of the files it enumerates, before you count it as coverage for the files you are worried about.
Continue the series
- 79SeriesThe Scanner That Mined Its Own CommandsA score built on any number it finds will rank shell flags as findings, so I trace each metric to the sentence it came from before I trust the ranking.
- 81SeriesThe Fix Measured in WordsA word count cannot tell prose from structure, so a fix justified by it can delete everything the count never measured.
- 78SeriesThe Missing Route That Was My ShellWhen a finding will not reproduce, ask for the exact command that produced it, because the defect may be in the probe.
- 82SeriesThe Anchor Taken After the CutA template matched across a content cut measures how alike two pieces of footage look, and only the confidence score says so.