how to audit an argument you didn't make
Part one was about why. This part is about how, written so that you could do it yourself, to this argument or any other one.
Most arguments about the mathematics of evolution end the same way: two people who are sure they’re doing math, doing language at each other. The audit’s job was to move the fight back into a place where it can be settled. That place isn’t a better essay. It’s a procedure, written down before anyone knew the results, that a stranger can run again and get the same numbers.
Here it is on one card. The rest of this piece is what each line means and what it caught.
- Write the rules before you have results.
- Prove the instrument on the textbook first.
- Read everything, verbatim.
- One claim, one file.
- Three verdicts, not one.
- Build the tree so the gaps show.
- Write the prediction down, then let it fail.
- Review every check three times.
- Publish everything that reruns.
1. write the rules before you have results
The rules went into the repo on day one, before a single check had run. Three verdicts per claim. Pinned definitions for the words each side uses differently: fixation versus fixed difference, base pair versus mutation event, census size versus effective size, how long one sweep takes versus how many run at once. Every number has to trace to a parameters file, a cited quote or a stated derivation, and a lint script fails the build when one doesn’t. Forward simulations decide disputed points; shortcuts get validated before anyone trusts them.
Rules written after the results are just a description of what you wanted to find. Rules written before are the only kind that can tell you no.
2. prove the instrument on the textbook first
Before the simulation code was allowed near anyone’s claim, it had to reproduce results nobody disputes. Neutral fixation probability should be 1/2N: the simulation gave 0.00989 against 0.01, and 0.00246 against 0.0025. Conditional fixation time, Kimura’s fixation formula, the neutral substitution rate at equilibrium: all matched.
If those had failed, nothing downstream would mean anything. It’s the same reason you run the test suite on a clean checkout before you blame someone else’s commit. Every check in the repo still starts from baseline_textbook.py.
3. read everything, verbatim
The first pass was a fast survey: three AI agents reading through summarizing tools. That pass leaned toward the mainstream view, and the project wrote that down. So none of its numbers were allowed to count. The second pass replaced every summary with a raw local copy, a bibliography entry and quotes machine-checked as exact matches against the source text.
The corpus ended up as 154 of Day’s blog posts, 32 Zenodo records, 37 primary papers cited by either side, and 48 critic and ally sources. Full texts stay out of the repo for copyright reasons. The repo keeps links and hashes, so you can fetch the original and confirm the quote yourself. One book was never bought, so every claim that appears only there is tagged secondhand.
4. one claim, one file
Every mathematical or empirical assertion became its own file, 193 of them. Each holds the verbatim quote with a locator and a date, a formal statement, the stated and unstated assumptions, the other side’s response, a prediction, and three empty verdict slots.
The split was 107 from Day, 46 from critics, 16 from allies and 24 from the literature. That’s uneven because Day wrote far more quantitative material than anyone answering him. The rule was equal scrutiny per claim, not equal numbers of claims.
One file shows why this matters. Day’s ratio of substitution rate to mutation rate appears across his work as N/Nₑ (19 to 46 for mammals), as 0.743, as about 0.5, as 32.3 and as 800,000. Some of those say the molecular clock runs fast and some say it runs slow. On 2026-08-27 he conceded that Nₑ doesn’t enter the identity at all, and the 32.3 figure appeared five weeks later without a derivation. Put in one table with dates, the version becomes the unit of analysis. A verdict on a January paper may not apply to an October blog post, and the reverse holds too.
5. three verdicts, not one
“Right” and “wrong” are too coarse. Every claim gets three separate answers:
- Internal: does the conclusion follow from the author’s own premises?
- Fidelity: does the cited source actually say what the author says it does?
- External: are the premises realistic?
Simulation can settle the first two. Only data settles the third. Splitting them is what lets the audit say several true things about one claim at once. Day’s formula for an “empty pipe” is exact for its premise, so internal holds, even though the premise doesn’t describe a real ancestral population, so external fails. A paper Day cites for a selection coefficient really does give that number, but for negative selection, not beneficial, so fidelity fails: the source was misread. On the other side, one critic’s bacterial mutation rate of 1e-11 sits below the measured 8.9e-11, and another’s “38 million matches 35 million” counts the same differences twice.
Of the three columns, external is the one with the most “contested” verdicts. That’s not a failure of the method. It’s the method telling you where the real disagreement lives.
6. build the tree so the gaps show
The claims hang on one argument tree: Day’s root claim, eight branches beneath it, and thirty nodes marked as the ones the root actually depends on. The tree isn’t drawn by hand. It’s generated from the claim files, so it can’t drift away from them, and the lint fails on orphans and missing sources.
That generation step is where neutrality can be audited. Every argument has to attach to a node, so an argument nobody made shows up as an empty slot. The clearest example: no critic in the corpus engaged the one piece of published literature that answers the cost-of-selection argument on its own terms. The empty slot made that visible, and the audit records it as a gap on the critics’ side.
The latest round went further. Every argument, on both sides, is now written out as numbered premises and a conclusion. The 238 objections between them are each typed: does it attack a premise, the step from premises to conclusion, or the conclusion itself? A standard computation from argumentation theory then reports which arguments survive.
7. write the prediction down, then let it fail
Every check records its prediction before the run. The point isn’t to be right. The point is to make it possible to be visibly wrong.
That happened. Day estimates that about 2.3% of founder events produce a hypermutator. The audit found that the size of the number holds (exact models give 2.0 to 2.75%) and that the route to it doesn’t. Then the audit’s own simulation landed at 2.75% under the literal reading of Day’s model, outside the band the audit had pre-registered. Its own falsifier fired, and the result went into the record as is. An audit that only ever confirms its predictions isn’t checking anything.
8. review every check three times
No check counts until it passes three reviews: one for correctness, one steelmanning Day, one steelmanning the critics.
The reviews earned their keep. After one round, eight major problems were found and fixed. In one model the mutation supply was inflated; it was fixed, re-run, and one of the audit’s own claims (“within 15%”) was withdrawn. Random seeds weren’t reproducible; fixed and re-run. An overclaim about ancient DNA was withdrawn outright. In another round the audit withdrew its own earlier objection to one of Day’s figures, because the steelman showed the objection was wrong. In another, a claim file that said “no response located” was corrected to credit a critic who had worked the same problem to the same answer six weeks earlier.
The reviews also fixed a rule that matters more than any single number: credit goes only to whoever made the argument. If the audit or the literature found something, it isn’t credited to the critics, because no critic said it.
9. publish everything that reruns
The repo is public. Every check runs with a fixed seed, so the same command gives the same numbers on your machine as on mine. The figures in the README are drawn only from committed outputs, so you can regenerate them and diff the result. Corrections from either side go in as issues, with a locator.
That’s the whole credential. Not a degree, not a journal, not me. If you can rerun it, you don’t have to trust me. If you can’t, you shouldn’t.
what the recipe doesn’t fix
The method is only as good as its honesty about its own limits, so here they are, straight from the project’s briefing:
- AI reviewing AI. Nearly all of the reading, extraction, simulation and reviewing was done by Claude agents I directed. The steelman reviews are AI reviewing AI, and Day’s papers list a Claude co-author too.
- No outside expert yet. Nothing has been peer-reviewed, and no human population geneticist has checked the simulations.
- Tested regimes only. Several results hold at a population of 1,000, a selection coefficient of 0.01 and one closed population. “Holds in the tested regime” is weaker than “holds.”
- No right of reply. No author on either side has been contacted yet.
- A moving target. Day’s numbers change between versions, and a verdict on one may not survive the next.
None of these is hidden, and each one is a place where a reader can push. That’s the design: the recipe doesn’t ask you to trust the cook. It hands you the kitchen.
Part one of the audit: prove it with the code. Next: where Day is right, the points where the argument’s math holds, written straight. Written against evo-sim on the research branch at 9a2148b. Every verdict is provisional until the audit’s final round.