Why a good backtest is not evidence

Every backtest you are ever shown is the survivor of a search. Someone tried one moving average, then another; a wider stop, then a narrower one; a filter on, then off. The curve you end up looking at is the best of those attempts, and how many attempts there were is not something the curve can tell you.

That is not a detail of presentation. It decides whether the curve means anything at all, and it cannot be recovered afterwards from the curve itself. Two identical charts, one from a single idea tested once and one from the best of a thousand, are indistinguishable on the page.

The arithmetic, made uncomfortable

Take twenty strategies with no edge whatsoever. Not weak strategies: strategies whose true expected return is exactly zero, built from noise. Test each one at the conventional 5% level.

On average, one of them passes.

Not because it works. Because you asked twenty times and 5% of twenty is one. The strategy that passes will have a p-value under 0.05, an equity curve that rises, and a story you can tell about why it works. Everything a real edge would produce, produced by nothing.

Now change one word. Stop calling them twenty strategies and call them one strategy with a dial on it: twenty settings of a stop, or a target, or a holding period. Nothing has changed mathematically. The best setting will look convincing for exactly the same empty reason, except that now it comes with a coherent narrative, because a dial has a direction and you can explain why the winning setting is the sensible one.

Run a hundred settings and the best of them will look very good indeed.

Nobody has to cheat

The word for this is data snooping, and the important thing about it is that it is not a failure of integrity. It is the default outcome of ordinary, careful, honest work.

Consider how research actually proceeds. You have an idea. You test it. The result is promising but not convincing, so you refine it: a different exit, a filter for the obviously bad periods, a longer lookback. Each refinement is motivated by something you learned from the last test, which is to say, by the data. Eventually you have a version that works.

At no point did anyone act in bad faith. At no point did anyone try a thousand random combinations hoping one would stick. And yet the final version is the survivor of a search whose size nobody wrote down, selected using information from the very data it is now being validated on.

This is why "we did not curve-fit, we reasoned from first principles" is not a defence. The question is not whether you were thoughtful. The question is how many things you looked at before deciding what to keep.

The p-value you were given answers the wrong question

A p-value describes the probability of seeing a result at least this extreme if nothing real is going on. Fine. But which result, out of what?

The number that gets reported is computed as though the winning configuration were the only thing anyone ever tried. It answers: "if I had tested exactly this one strategy, once, how surprising would this be?"

You did not test exactly one strategy once. You tested a set and reported the maximum. The correct question is about the maximum of the set:

If none of these had any real edge, how often would the best of them look at least this good?

That question has an answer and there are established ways to compute it. They all share one requirement, which is the reason they are so often skipped: you have to state the size of your search honestly, and the correction gets harsher the more you tried.

"But I only tried a few things"

The objection is reasonable, and three things compound against it.

The abandoned ideas count. A strategy you tested for an afternoon and dropped is a trial. It came from the same data and it competed for the same slot. A log that records only the survivors has kept the wrong half: the discarded ideas are the part that sets the size of the search, and once they are gone the size cannot be reconstructed.

The choices that do not look like parameters count. Which year to start from. Whether to exclude the crash. How to treat the outlier that one Tuesday. Which data vendor. What to do about the exchange that delisted. None of these look like a dial and every one of them behaves like one, because each was decided after seeing what it did to the result. Simmons, Nelson and Simonsohn named these researcher degrees of freedom in 2011, and showed that a handful of them, exercised in combination on data with nothing in it, is enough to produce a statistically significant result at will.

Reusing a held-out sample counts. Which brings us to the cure that is not one.

Out of sample is not the cure people think it is

The standard defence is to hold back a slice of history, develop on the rest, and test once on the slice you kept. It is a good practice and it is genuinely useful.

It is also good for exactly one use.

The moment a strategy disappoints on the held-out slice and you go back to adjust it, that slice has joined the search. Test again and you are no longer validating out of sample; you are optimising on a smaller sample with extra ceremony. Repeat it a few times and the held-out data has been consumed, without there being a moment anyone could point to as the one where it happened.

The honest version requires counting those rounds as trials, which requires having written them down at the time.

There is no technique that fixes this after the fact. There is only a research log kept honestly while the work is happening.

What a plateau tells you that a peak never will

Suppose you have done all of the above properly and the corrected number still passes. There is one more thing to look at, and it is often more informative than the number.

Look at the whole surface, not the winner.

If the result comes from a single sharp peak surrounded by mediocrity, that is what luck produces. Luck is local. A configuration that works only at one precise setting, with neighbours that do not work, is describing an accident of this particular history, and small changes in market behaviour will step off it.

If most of the region is positive and the surface is broad and flat, something structural is going on. Your specific setting matters less than the fact that a whole neighbourhood works, which also means you are not depending on having picked the right point.

So the questions are: how many of your configurations were positive? How many had a negative mean? And where in the ranking does the one you actually trade sit?

That last question is the sharpest instrument in the box. A manager trading the top of their own grid has told you which number they optimised, whatever else they say. A manager trading somewhere in the middle, who can explain what non-performance reason put them there, has told you something quite different.

What this means in practice

Four habits follow, and none of them is technical.

Declare the set before you look at the answer. The correction is computed over the set you actually searched, so the set has to exist as a written thing, not as a memory of what you happened to try.

Count everything, including the failures and the choices that were not parameters. The count will be larger than you expect and larger than you would like.

Publish the search alongside the result. This is not our idea: López de Prado states it as the third law of backtesting, that every result must appear together with all the trials that produced it, because without them the probability of a false discovery cannot be assessed. Taken seriously it is a formatting rule rather than a principle. If a multi-year curve appears without the size of the search that produced it, the reader has been handed the numerator and asked to assume the denominator.

And treat a passing number as the beginning of the argument rather than the end of it. It means the result is hard to explain by luck. It does not mean the result will persist, which is a different question, answered by different evidence, over a longer period than anyone wants.

There is no version of this that ends with certainty. There is only the difference between a result whose search has been priced and one whose search is unknown, and that difference is most of what separates research from a sales document.