Three hundred settings can be worth one trial

Every correction for multiple testing takes a number N: how many things you tried. The count of configurations is the number closest to hand, and it is not the number the correction wants.

The correction wants the number of independent trials. If your configurations are correlated with each other, three hundred rows are not three hundred experiments. They are one experiment, slightly perturbed, three hundred times.

The formula, and what it does

For a set of candidates that are equally correlated with each other at ρ, the effective count is

`` N_eff = N / (1 + (N - 1) · ρ) ``

which is the standard effective-sample-size expression for equicorrelated variables. Put numbers in it and the result is more severe than intuition suggests:

average pairwise ρN = 100N = 300N = 1000
0.501.981.992.00
0.701.421.431.43
0.801.251.251.25
0.901.111.111.11

Read across a row rather than down a column. Once correlation is high, the effective count barely depends on how many rows you have. A thousand configurations at ρ = 0.8 are worth 1.25 independent trials, and so are a hundred. Adding rows to a correlated grid does not add evidence; it adds rows.

The formula is a simplification, because real grids are not perfectly equicorrelated. It is the right order of magnitude, and the direction of the simplification is not in your favour: a grid with a few genuinely different members and many near-duplicates behaves closer to the near-duplicates than the average correlation suggests.

Why a parameter grid is correlated by construction

Take a signal and put dials on it: where to take profit, how long to hold, how often to enter. Every combination of dial settings is one row.

Every row trades the same signal at the same moments. They enter together, exit at slightly different points, and their daily return series move together. Changing an exit threshold by half a percent does not ask a new question of the data. It asks the same question with a slightly different microphone.

That is a structural property of how grids are built, not an empirical claim about anyone's research. It is why measuring ρ on a grid usually lands in the range where the table above is brutal.

The contrast makes it concrete: twenty exit thresholds on one market share one signal and one set of entry moments. Twenty markets tested with one rule do not, because the return series have different drivers. The second set buys evidence. The first buys rows.

How to measure it on your own set

Take the daily return series of every candidate, not their summary statistics. Compute the average pairwise correlation. Put it in the formula.

Two details change the answer. Correlate returns rather than equity curves, because curves share a common drift that inflates the correlation for a reason that has nothing to do with strategy similarity. And use the full history rather than a recent window, because correlation between related strategies rises in stressed periods, which is exactly when the independence you assumed would matter.

Which number to publish

Two defensible choices, pointing in opposite directions.

The effective count is the technically correct input and produces the friendlier answer, because a smaller N means a lower bar. The raw count is conservative and harder to argue with.

Publish the conservative number and show the effective count beside it. A reader who wants the generous reading can take it; a reader who does not is not being led anywhere. The arrangement worth avoiding is the friendly number in the headline and the conservative one in a footnote.

The general habit: when a methodological choice moves your result, publish both ends and say which one you rely on. It costs a sentence and removes the whole category of suspicion that you picked the convenient one.

Why a large search is not a boast

"We tested ten thousand combinations" is offered as evidence of thoroughness. Under any correction for multiple testing it raises the bar the speaker has to clear.

If the correction was applied properly, ten thousand attempts made the result harder to achieve. If it was not applied, the sentence has told you that instead. Either way the number works against the claim it was meant to support.

The version that means something names the effective count and where the live configuration sits among the candidates: this many settings, this many independent by the measure above, and the one we trade ranks here.

What it changes about building a grid

Two consequences follow for the research itself, and both are visible in the table.

A wider grid costs correction and buys almost nothing. Adding a fourth value to a dial adds rows and raises N in the correction, while N_eff moves by a rounding error.

A genuinely different dimension is worth more than another dial. A second asset, a structurally different exit, a different holding regime: these raise the correction too, and unlike extra dial values they raise N_eff with it.

Which gives you a test to apply before running anything: if the thing you are about to add would be highly correlated with what you already have, you are about to pay for information you will not receive.