Trading Edge

The Number Nobody Writes Down: How Many Ideas You Tested

Test twenty condition combinations against a 5% threshold and roughly one will look significant by chance alone. Almost no trader records how many they tried, which makes it the most important missing field in the whole journal.

E

EdgeFlow

·Updated

A trader posts in the group chat: "Found it. My win rate goes from 48% to 67% when I only take the setup on trend days in London."

Twenty people congratulate him. Nobody asks the one question that decides whether the number means anything at all.

How many combinations did you check before that one?

He doesn't know. He never wrote it down. And without it, that 67% is uninterpretable. Not wrong, not fake, just impossible to read. It could be a genuine discovery, or the loudest of the forty things he tried on a Sunday afternoon. From the outside those two look identical, and from the inside they feel identical too.

Every journal on the market has fields for date, instrument, direction, risk, result, tags, screenshot, notes. Not one has a field for how many ideas you tested before you kept this one. That missing field is the subject of this post.

Quick vocabulary, then we get to the numbers

I'll keep this short because the concepts are simpler than their names.

R is one unit of risk. Risk $150, make $300, that's +2R.

Expectancy is your average result per trade in R. If you're fuzzy on what separates a setup from an actual trading edge, that one's worth reading first.

n is just the number of trades in that slice. When you see n = 26 further down, it means twenty-six trades and nothing more clever than that.

Data snooping (also called data mining, or p-hacking when people are being unkind) is what happens when you search through data for a pattern and then report the pattern as if you'd predicted it in advance.

Multiple comparisons is the technical name for the specific problem: the more separate things you test, the more likely at least one looks good by pure chance.

And significance, in the way most people use it, means: if there were nothing real here, how often would a result this good show up anyway? Less than one time in twenty is the customary bar. Hold on to that. That bar was built for people running one test.

The coin-flip version

Take twenty coins. Flip each one ten times.

At least one of those coins will probably come up 8 heads or better. That coin isn't special and it isn't lucky. It's simply what twenty coins do.

Now the trading version. You have 200 trades logged and no edge whatsoever. Pure noise. You slice that history twenty different ways and check each slice against the one-in-twenty bar. On average, one slice clears it. That happens because you looked twenty times, not because anything was there: one-in-twenty events turn up roughly once in twenty tries.

Here's the same idea as a probability. If each test has a 5% chance of a false alarm, the chance of getting at least one false alarm across several independent tests is:

Tests runChance at least one looks "significant" by luck
15%
523%
1040%
2064%
5092%

At fifty tests you're more likely than not to find a false winner even if your strategy is worthless. Fifty isn't a wild number for someone who has journalled for two years and pokes at the data most weekends. I have no survey to point at, only my own count. I passed fifty a long time ago and never wrote a single one of them down.

How fast the count explodes

This is where people underestimate themselves badly.

Say you tag six conditions on your trades. Session, higher-timeframe direction, volatility state, entry type, day of week, whether you took partials. Six tags. Feels modest.

Test each one on its own: 6 tests.

Test every pair of them: 15 more.

That's 21 tests already, and you haven't touched triples. Eight tags puts you at 36. Ten tags gets you 55. What you can check grows much faster than what you track, which is why testing condition combinations needs a plan rather than a Sunday afternoon. Twenty-one tests against a one-in-twenty bar means you should expect roughly one false winner, on average, out of a journal with no edge in it at all.

A worked example with real numbers

Let's put actual figures on it.

You have 200 trades. Baseline win rate 48%, roughly break-even after costs. Frustrating, familiar.

You slice by those six tags, singles and pairs, 21 combinations. One comes back beautiful: 20 winners out of 30 trades. That's 67%, against a 48% baseline.

How surprising is that on its own? Getting 20 or more winners out of 30 when the true rate is 48% happens about 3 times in 100 by chance. That clears the usual one-in-twenty bar. If it were the only thing you'd ever tested, you'd have something worth taking seriously.

But it wasn't the only thing. It was one of 21.

The plain-language correction is this: multiply your surprise by the number of places you looked.

3.07% × 21 ≈ 64%.

Sixty-four percent. That's closer to two chances in three than to a coin flip. What looked like a 1-in-33 event turns out to be an ordinary thing to stumble on when you check 21 slices of a break-even journal. You'd almost be unlucky to walk away without one.

Flip it around if you prefer thresholds to multiplication. With 21 tests, the honest bar isn't 5% anymore. It's 5% ÷ 21, which is about 0.24%. The bar got twenty-one times stricter, and your finding missed it by a factor of about thirteen.

That's it. That's the whole adjustment. Multiply the surprise by the number of tests, or divide the threshold by the number of tests. You can do either in your head at the desk. It's called the Bonferroni correction and it's the bluntest tool in the box, which is exactly why it's the one worth carrying around.

The number is always bigger than you think

Now the uncomfortable part. When traders do try to count, they undercount, and usually by a lot.

Your real hypothesis count includes:

  • The slices you clicked through and closed because they looked bad. Those were tests.
  • The tweaks. Break-even at 1R versus 1.5R versus never is three tests, not one.
  • The version of this same hunt you ran in March and abandoned.
  • The ideas you tested by eye. Scrolling the list, noticing Fridays look rough, moving on. That's a test, you just never wrote down the result.
  • The ideas you inherited. When a mentor says the London open filter works, he searched too. You took his winner without his count.

Statisticians call that last bit the garden of forking paths. Every choice about how to cut the data, including the ones that felt obvious, was a fork you could have taken differently.

So your honest number is never precise. Mine isn't either. Write down the count you can defend and admit the real one is higher. "At least 21" beats a blank field, and it beats "I only tested the one."

What the field actually looks like

You don't need software for this. A text file works. One block per hunt:

2026-03-08 — question: does BE-at-1R help or hurt?
combinations checked: 8 (4 setups × 2 BE rules)
best result: Setup B without BE, +0.34R vs +0.05R baseline, n = 26 trades
adjusted: 8 × the raw surprise. Not significant.
also checked: [the other 7, results kept in file]
status: UNVERIFIED — needs 25 forward trades before it becomes a rule

Four things make that block work.

The question is recorded before the answer, so you can't quietly reframe it later.

The count is there, so the result can be interpreted by you in six months or by anyone you show it to.

The losers are kept. Everybody skips this and it's the part that matters most. Delete the 7 combinations that didn't work and the survivor starts looking like a discovery. Knowing how many you ran is what lets you correct honestly later. That thinking shaped how we built EdgeFlow: the engine scores every combination it can assemble from your tags and reports how many it scored, ranks what comes back by a cautious lower bound rather than the flattering headline expectancy, and shows the near-misses and the money-losing combinations next to the winners. The best result arrives with company.

The status says UNVERIFIED, and it stays that way until the idea has met trades it was not built on. That word is doing real work in the block, so don't upgrade it early.

What the number does and doesn't do

A high hypothesis count doesn't kill your idea. That's the misreading I want to head off.

It tells you what the idea currently is: a candidate, not a conclusion. It sets the height of the bar the idea has to clear next, and it clears that bar with trades it has never seen. Split your history chronologically, find the pattern on the first chunk, test it on the second, and if it survives, keep going forward. That's what in-sample versus out-of-sample means in practice. New data doesn't care how many times you looked at the old data.

The multiplication trick is rough, and I'd rather say so than pretend. Your 21 tests aren't independent. "London session" and "London session plus trending" overlap heavily, so the honest correction sits somewhere gentler than 21, and nobody can hand you the exact number. The direction is still right, and a rough correction you actually do beats a precise one you never get round to. If it dies at 21, it would have died at 12 too.

The wider pattern is the one covered in overfitting your own trading journal: searching is fine, reporting only the winner is the problem. The count is the receipt for the search. And the anatomy of an overfit setup walks through one that fooled its owner for months, if you want to see a false winner from the inside.

Start counting from today

You can't reconstruct your past count. Don't try, you'll flatter yourself.

Start from now. Next time you open the journal to go hunting, write the date, the question, and a tally mark for every combination you check, including the ones you throw away in two seconds. Do that for a month and you'll have a number that changes how you read your own results more than any new metric could.

My guess is that most people land somewhere in the dozens, and that one or two favourite rules won't survive the correction. I have no survey to point at, only my own count and the ones traders have shown me. If that's how it goes for you, treat it as the first honest read of your own testing rather than a verdict on your trading. It feeds straight into the wider question of whether you're holding a real edge or a lucky streak.

If you'd rather not keep the tally by hand, EdgeFlow's edge discovery counts the combinations for you and ranks what comes back by the cautious lower bound instead of the best-looking number.

Frequently asked questions

What is data snooping in trading?

Searching your trade history for a pattern and then presenting it as though you'd predicted it beforehand. The search itself is fine and necessary. The problem is reporting the winner without mentioning the search that produced it.

How many tests is too many?

There's no cutoff. More tests just raise the bar each individual result has to clear. Twenty tests is fine if you correct for twenty and verify on unseen trades. Twenty tests is a disaster if you report the best one as a finding.

Can I just use a bigger sample instead of counting?

A bigger sample helps but doesn't replace the count. More data makes each test more reliable. It does nothing about the fact that you ran forty of them and kept the best.

Is the multiply-by-N correction statistically exact?

No. It assumes your tests are independent, and slices of the same journal never are, so it over-corrects. Use it as a sanity check, not as proof, and let forward data do the real work.

Continue reading

Related articles