You have twenty versions of a strategy and one of them looks good. The naive p-value describes the odds of that result appearing by chance if you had only ever tried that one version. You didn't. So the number answers a question nobody asked.
Peter Hansen's Superior Predictive Ability test, published in 2005, fixes the question rather than the arithmetic. Instead of asking about your winner, it asks about your search.
If none of the candidates in this set had any real edge, how often would the best of them still look at least this good?
Everything else is machinery for answering that honestly.
Read this part carefully. It is what the p-value is a statement about, and getting it wrong changes what the number means without changing the number.
The null is not "this strategy has no edge". It is: none of the models in this set beats the benchmark. Rejecting it means at least one of them does.
Two things follow, and both are practical.
The set has to be declared before you look. You cannot run the test on the configuration you liked, after the fact, and call it corrected. The correction is computed over the set you actually searched, so that set has to exist as a written thing rather than as a memory of what you happened to try. A test run on the winner alone is the naive p-value wearing a costume.
Rejection is a weaker claim than it sounds, and that is a feature. You have not shown that your chosen configuration works. You have shown that the set contains something real. Which one is a separate question, and an honest write-up says so.
The test needs to know what luck looks like on histories like yours. You have exactly one history, so it manufactures thousands of synthetic ones from it and counts how often chance produces a winner as good as the one you have.
The obvious way to manufacture them is to shuffle your returns. The obvious way is wrong.
Financial returns are not independent draws. Volatility clusters: calm weeks follow calm weeks and violent days arrive together. Trends persist for a while and then do not. A naive shuffle destroys exactly that structure, and it produces synthetic histories that are far tamer than anything the market has ever done. Test against tame histories and everything looks significant.
The standard answer is the stationary bootstrap of Politis and Romano, which resamples blocks of consecutive observations rather than single days, with block lengths drawn at random so no single choice of block dominates the answer.
That leaves you one parameter to defend: the average block length. Report it, try several, and show the answer holds. If the p-value moves materially when blocks go from five days to twenty, the number is describing your resampling scheme rather than your strategy, and it should not be published without that caveat.
The earlier test in this family is White's Reality Check, and the differences between the two are not academic: they can move a decision.
Studentising. Hansen scales each candidate by its own variability instead of comparing raw average returns. Without that, a wildly volatile configuration with a big average looks stronger than a steady one with a smaller average, when the steady one is usually the better bet and certainly the more repeatable. The correction is the same idea that makes a t-statistic more useful than a raw difference of means.
Excluding hopeless candidates from the null. This one is subtle and it matters more. Under Reality Check, every candidate you tried sits in the null distribution, including the ones that lost badly. A pile of obviously terrible configurations drags the distribution down and makes your winner look better by contrast. Hansen allows candidates that could not plausibly beat the benchmark to be excluded, which raises the bar back to where it belongs.
The practical consequence: padding your grid with variants you never seriously considered makes your result look better under the older test. Under Hansen's it does not. If someone reports a "bootstrap p-value" without saying which variant they ran, ask, because the answer tells you whether padding would have helped them.
A reported number should come with enough to reproduce it. Four things, and none is optional:
The set. How many candidates, and what were they? "We tested our parameters" is not an answer. The size of the set is the whole point of the correction.
The replications. A few hundred is not enough for a stable tail estimate. Thousands is normal. A number computed on too few resamples wobbles, and the direction it wobbles is whichever way the author stopped.
The block lengths and the seeds. Both, along with evidence that the answer is stable across them.
The implementation. Use a reference implementation where one exists. We wrote our own before checking it against the reference, and ours returned the friendlier p-value of the two. One comparison is not a law about home-grown code, but it is a cheap check with an obvious asymmetry: an error in your own implementation that makes the number harsher gets found immediately, and one that makes it kinder does not.
If those four are missing, the p-value is decoration. It is not that the author is lying; it is that nobody, including them, can tell what the number means.
It is a correction for search, and that is all it is. Three things it leaves untouched.
It says nothing about the future. A result that survives the correction is hard to explain by luck across the period you tested. Whether the effect persists is a different question, answered by different evidence, over a longer horizon than anyone would like.
It cannot correct for trials you did not declare. Ideas abandoned before they reached the grid, choices made after seeing the data, a held-out sample reused three times: none of that is in the set, so none of it is priced. The test is only as honest as the set you feed it.
It does not tell you which candidate to trade. The rejection is about the set. Picking the winner out of it is a separate decision, and picking the maximum is usually the wrong one, because the maximum is where luck accumulates.
Which is why the number is the beginning of the argument rather than the end of it. The shape of the surface behind it, and where your own configuration sits on that surface, will tell you more than the p-value ever does.