A result we could not explain, and what we did with it

We buy the asset on a schedule. The idea we were testing was to interrupt that schedule: while the market is rising, hold the purchases aside, sell them when the trend turns, and buy back more of the asset lower down. The whole point is denominated in coins rather than in dollars, so the question is simply whether the round trip ends with more of the asset than doing nothing.

The first version used the trend measure we already run. It returned half a percent, and a second number explained why it was not more. At the moment the trend turned, the selling price was below the average price we had paid during that same rise, and it was below it in sixty three percent of the cases. The measure reacts to a move that has already happened, so we would have been selling after the drop rather than before it.

That is a mechanism, and it pointed somewhere specific: if the delay is what costs us, a faster measure should help. So we tested a family of low lag filters, chosen beforehand on how quickly and how closely they agreed with a centred reference that is allowed to see the future. Such a reference is useless in real time and ideal as a yardstick. We fixed the choice before looking at any returns, because selecting a filter by its profit is how you end up measuring your own search rather than the market.

One of them returned six percent more of the asset. That was the moment to be careful rather than pleased, because the same filter scored worse on the diagnostic that mattered: it sold below its own average purchase price sixty nine percent of the time, against sixty three for the slower one. The profit and the explanation were pointing in opposite directions, and only one of them can be right.

So we took the filter's own sequence of regimes, kept every one of their durations exactly as they were, and shuffled their order. That destroys any relationship with price while preserving everything else about the shape of the signal. Then we ran the same overlay on three hundred of those shuffles.

The median shuffle lost four and a half percent. The best gained twenty eight. The ninety fifth percentile of pure chance sat at plus twelve and a half percent, and our real filter, with its six, did not reach it. Thirteen percent of random orderings did at least as well as the thing we had built.

The overlay is not in the account and will not be. What we kept is the rule that produced this page: a number without a mechanism is not a finding, however much we would like it to be. We had two positive results that week, and neither survived the question of why.