Five questions to ask of any backtest

These are ordered by how hard they are to answer dishonestly rather than by importance. The first can be answered falsely but not comfortably, because the natural follow-up is a list. The last is the most technical and the only one that requires work to have been done in advance rather than recalled.

None of them is a trap. Each asks for a fact the manager either has or does not, and every one of the five is easy for someone who has already done the work.

1. How many configurations did you try in total, including the ones you abandoned?

A good answer is a number, given without hesitation, and it includes the ideas that were dropped after an afternoon. It usually comes with a research log, because the only way to know the number is to have written things down while they were happening.

A bad answer is a range, or a redirection to how principled the process was. "We reasoned from first principles rather than curve-fitting" does not answer the question; it answers a different one, about intent. The size of the search is a fact about what happened, not about how carefully anyone was thinking.

"We only tested a handful" is not reassuring on its own. It is reassuring when followed by the handful.

2. Where does the live configuration rank in that set?

A good answer is a middling rank with a reason that has nothing to do with performance: robustness across neighbouring settings, operational simplicity, a constraint imposed by the venue, a decision made before the results were in.

First place needs a very good story. The top of a grid is where luck accumulates, so a manager trading their own maximum has told you which number they optimised, whatever else they say. It is not disqualifying. It does mean the burden of explanation sits with them.

The follow-up worth asking: how much worse is the median configuration than the one you trade? If the answer is "not much", you are looking at a plateau and the specific choice matters little. If the answer is "enormously", you are looking at a peak, and peaks are what randomness produces.

3. What is the benchmark, and is it what the investor would otherwise have held?

A good answer names the actual alternative and is comfortable when that alternative is embarrassing. If the money would have sat in an index, the index is the benchmark, even when the index won.

A bad answer names the convention for the category. Absolute-return funds benchmarked to cash while their investors would have been invested; strategies reported in one currency while their investors count in another; anything measured against zero when the capital was never going to be idle.

The follow-up: what fraction of capital was actually deployed on average? A return on money that was never at risk is a different number from a return on money entrusted, and the gap is a decision rather than a result.

4. What are the skew and kurtosis of the return series?

A good answer is two numbers, stated in one sentence, without being asked twice.

A bad answer is a Sharpe ratio. A manager who has not computed the moments has not looked at their own tail, because the tail is precisely what those numbers describe and the ratio is precisely where it is hidden.

This question does double duty. It tells you the shape of the thing you are being offered, and it tells you whether the person offering it thinks in distributions or in averages. The second is often the more useful signal.

If the strategy is built on many small gains and rare large losses, then the interesting number is not the average outcome but the worst single one, and whether it arrived during the period on offer. A negative skew figure is how you find out that it is that kind of strategy.

5. How many of those trials were independent?

The best answer is that they measured the correlation across their candidates and know the effective count. It requires having correlated the return series of the candidates against each other, which is a piece of work that has to have been done rather than remembered.

"The question has not come up" is an honest answer, and it tells you which way the correction is wrong. Using the raw count on a correlated grid is the conservative error, so a result that passed under it passed the harder version. The answer to be careful with is the reverse: a set that was never fully written down, where the count is too low and no correction can recover it.

The answer to be wary of is a large number quoted proudly, because a large search raises the bar the speaker has to clear. If ten thousand combinations were tested and the reported significance did not move, one of the two numbers is not doing its job.

When you cannot get answers

Sometimes the material simply is not available: a fund that will not discuss its process, a track record without a research log behind it, a manager who inherited the strategy.

That is not automatically disqualifying, but it changes what the record can support. Without the size of the search you cannot distinguish a real edge from the best of an unknown number of attempts, so the record becomes weak evidence about the future regardless of how good it looks.

The honest position in that case is to size the position as though the evidence were weak, because it is. What you must not do is treat an unanswerable question as an answered one because the curve was pretty.

What to do with the answers

None of the five produces a verdict on its own. Together they tell you something more useful than a verdict: whether the person in front of you has already asked themselves the hard questions.

A manager who has priced their own search, knows their own moments, and can explain why they are not trading their own maximum has done the work whether or not the strategy eventually stops working. A manager who finds these questions surprising has not, and the strategy working so far is not evidence that they have.

Turning them on yourself

The five are easier to ask than to answer, and the useful exercise is answering them about your own work before anyone else does.

Two of them are uncomfortable in a specific way. The count of trials requires a log you may not have kept, and starting one now does not recover the trials you have already spent. The rank question requires checking whether the configuration you chose was chosen partly because it looked good, and the check is easy: name the non-performance reason for the choice. If there isn't one, performance was the reason.

The point of asking yourself first is not self-flagellation. It is that the answers change what you do next: a strategy whose search you cannot price is a strategy you should size smaller, and knowing that early is cheaper than discovering it later.