Eleven mistakes in one day, none of them arithmetic

On 29 July we ran a full statistical audit of the strategy. By the end of it we had made eleven mistakes. Not one was a calculation error. Every figure was computed correctly, and eleven conclusions drawn from those figures were wrong.

That is the uncomfortable kind. A calculation error announces itself eventually: something does not reconcile, a total is off, a test fails. A wrong conclusion drawn from a correct number looks exactly like a right one.

What they had in common

Sorted afterwards, the eleven fell into four groups, and each group is a question we had failed to ask.

Where did this number come from, and has it been checked against anything? Several errors came from figures pulled out of a saved grid or a previous run and never re-derived. One of them was computed on a series that measured only part of the account, and it had been sitting there for a day and a half, quietly feeding a section of a report.

Did we read the whole series, or only the part that agreed with us? One conclusion rested on three data points out of seven. The other four were in the same file, and they pointed the other way.

Is this already decided somewhere? Three of the eleven were false alarms about the live system, and the answer to each was in our own documents. Raising an alarm feels like diligence. Raising an alarm about something already settled is noise with a serious face on.

What mechanism would have to be true for this explanation to hold, and did we check it? Two errors were confident causal stories with the wrong quantity underneath. In one, the story would have been refuted by a number we had already computed and not looked at.

What we changed

We wrote the four questions down and made them a gate: nothing gets said until each has an answer, or until the gap is stated out loud as an unchecked assumption. There is no partial credit here. Either the source has been verified or the sentence carries a note saying it has not.

Then we did the more useful thing, which was to notice that a checklist protects you only when you remember you need it. That morning we were sure we had verified our sources, and wrote so. So wherever a check could be turned into code, it was: anchors that raise rather than warn, a harness that refuses to accept a figure computed twice by different paths if the two disagree.

Checks of the first kind, "is this number what I think it is", turn into code almost always. Checks of the fourth kind, "is my causal story sound", almost never. That asymmetry is worth knowing about yourself.

Why this is published

Because a research log that contains only the days when the work went well is marketing.

The eleven mistakes cost us a day and a half and one section of a report that had to be withdrawn and recomputed. What they bought was the process we now use, which is worth considerably more than the day.