Trading Edge

How Many Trades Before a Setup Means Anything?

There is no universal 100-trade threshold. The number you need scales with how large the effect is and how many ideas you tested first. How to derive your own minimum, and what statistical significance actually means for a trader.

E

EdgeFlow

Ask how many trades you need before you can trust a setup and the answer comes back fast. A hundred. Sometimes thirty. Sometimes two hundred, if the person answering wants to sound careful.

None of those numbers were derived from anything. They are round. That is mostly why they are popular.

The honest version is that the number is an output, not an input. You work it out. And two things drive it: how big the effect is that you are trying to see, and how many other ideas you already tested against the same data.

Let me make both of those concrete, with real arithmetic.

First, what are you even counting?

Before you count trades you need a fixed thing to count. Same rules, same conditions, same management. Change your stop placement halfway through and you have two samples of thirty, not one of sixty. That is the whole subject of what a trading edge actually is.

Assume that part is done. Now the counting problem.

Why you need any number at all

Flip a fair coin a hundred times. You expect fifty heads. You get 55/45 fairly often. 58/42 is not rare. Nothing is wrong with the coin. That is just what randomness looks like from close up.

Your trades behave the same way. Twelve wins out of twenty tells you almost nothing about the setup. It tells you that you took twenty trades.

So the real question is not "how many trades do I need." It is: how many trades before a result this good stops being explainable by luck? That is the same question approached from a different side in real edge or lucky streak, and the answer depends entirely on how good "this good" is.

Driver one: how big is the edge

Big effects are easy to spot. Small ones hide in the noise.

Take a normal setup. You risk 1R, you target 2R, you win 40% of the time. R just means one unit of your risk, so a 2R winner pays twice what a loser costs.

Expectancy = (0.40 × 2R) − (0.60 × 1R) = +0.20R per trade

Now look at the individual trades that produced that average. They are all either +2R or −1R. Nothing sits at +0.20R. The spread of those outcomes, the standard deviation, works out to about 1.47R.

Sit with that for a second. The trade-to-trade noise is roughly seven times the size of the average edge you are trying to measure. Seven times.

That ratio is the entire game. And the sample size you need scales with the square of it. Halve your edge and you need four times the trades. That squaring is why "a few dozen should be enough" is so badly wrong for the small edges most of us actually have.

A rough working formula:

n ≈ 8 × (standard deviation ÷ edge)²

For our example: 8 × (1.47 ÷ 0.20)² = 8 × 54 = about 430 trades.

Four hundred and thirty, for a setup that genuinely makes +0.2R every time you take it.

Where does the 8 come from? Two choices baked together: a 5% tolerance for calling a dead setup alive, and roughly an 80% chance of catching the edge when it really is there. Settle for a coin-flip chance of detecting it and the constant drops to about 4, halving everything below. Neither setting is a law of nature.

The table (use it, don't worship it)

Assuming a standard deviation around 1.5R, which is typical for a 1:2 style strategy:

Real edge per tradeTrades needed to see it reliably
+1.00R~20
+0.75R~30
+0.50R~70
+0.30R~200
+0.20R~450
+0.10R~1,800
+0.05R~7,000

A +1R row is rare enough in discretionary retail trading that I would treat it as a reason to re-check the data before believing it. And in the journals I have looked at, the genuine edges tend to land somewhere between +0.05R and +0.30R. That is my experience talking, not a published number, so hold it loosely. But if it is even roughly right, the honest requirement for most traders sits in the hundreds, not at a hundred.

One adjustment matters more than anything else here: that 1.5R is not your number. If you let winners run and occasionally book +8R, your standard deviation might be 2.5R, which nearly triples every figure in the table. If you scalp fixed 1:1 targets it might be 1.0R and everything gets cheaper. Pull your own R-multiples and compute the spread. Two minutes in a spreadsheet.

Driver two: how many ideas you tested first

This is the part nobody adjusts for, and it is the part that quietly wrecks journals.

The standard 5% significance bar means: if this setup were completely worthless, there would still be a 5% chance of a result this good turning up by accident. One in twenty. That is fine when you test one idea.

Test twenty combinations and you should expect one of them to clear the bar on noise alone. Not "might." Expect.

Most of us have done this without realising it. You open the journal on a Sunday. Sort by session. Sort by day of week. Check whether the higher timeframe was aligned. Split the first hour from the rest. That is a dozen tests before lunch, and then one cell reads 68% and it feels like a discovery.

This has a name: multiplicity. The more separate things you test, the more of them get lucky, and the higher the bar has to be before any single winner means anything. Correcting for it is not brutal on its own. Roughly:

  • one hypothesis: baseline
  • about five ideas tested: multiply the required sample by ~1.5
  • about twenty: multiply by ~2
  • about a hundred: multiply by ~2.5

It grows slowly, which is the good news. The bad news is you almost always tested more than you think. Every slice you eyeballed and dismissed still counts, because it still had a chance to come up lucky.

And there is a second cost that hurts more than the multiplier.

Slicing splits your sample. Say you have 200 trades at +0.15R overall. You cut by session (three) and by higher-timeframe alignment (two). Six buckets, roughly 33 trades each. One bucket prints +0.6R and looks brilliant. But to trust a +0.6R claim found by searching, you would want somewhere north of 60 trades in that specific bucket. You have 33. Your sample got smaller and your bar got higher at the same moment.

That +0.6R is a hypothesis, not a finding. Treat it as something to test going forward, not something to trade tomorrow. There is a longer treatment of how this goes wrong in overfitting your own trading journal.

What "statistically significant" actually means

Worth stating plainly, because the phrase gets thrown around badly.

Significant means: a result like this would be unlikely if there were no edge at all. That is the whole claim.

It does not mean the edge is big. Feed a large enough sample in and a +0.02R edge comes back "significant" while still being worthless after commission and slippage.

It does not mean the edge will continue. The test looks backwards. Markets do not.

And it does not mean there is a 95% chance you are right. That is not what the number says, though nearly everyone reads it that way.

The reverse trips people up too. Not significant does not mean no edge. Usually it means not enough trades. You failed to prove something, which is different from disproving it. That distinction, and a handful of related ways small samples mislead, is covered in small sample traps in trading statistics.

A better question than yes or no

Here is my actual opinion, and it is the one thing I would keep from this article: stop asking whether a setup is significant. Ask what range your true edge plausibly sits in.

Worked example. 200 trades, average +0.25R, standard deviation 1.5R.

One term first. Standard error is just how far off your measured average is likely to be. You take the spread of your trades and divide it by the square root of how many trades you have. More trades, less wobble. It is a square root because noise partly cancels itself out as trades pile up, which is why the fourth hundred trades sharpens the picture far less than the first hundred did.

Standard error = 1.5 ÷ √200 = 0.106

95% band ≈ 0.25 ± (1.96 × 0.106) = +0.04R to +0.46R

Read that sentence out loud: my edge is probably real, but it might be a fifth of what I think, or nearly double. That is a far more useful thing to know than a yes/no verdict, and it tells you exactly how much to bet. Plan around the bottom of the band, not the middle. The same reasoning applied to win rate specifically is in the win rate confidence interval.

That lower bound is also the rule we ended up hard-coding into EdgeFlow. A condition combination does not get called a strong edge unless the bottom of its expectancy band clears zero, and combinations are ranked by that cautious lower number rather than the flattering headline average. When there is not enough data, it says so instead of inventing a verdict.

What to do when you don't have 400 trades

Almost nobody does, at least not per setup. Some things that genuinely help:

Test fewer, bigger ideas. Write the hypothesis down before you open the data. One sentence. "Trades where I waited for the retest do better." Then check that one thing. The multiplicity penalty vanishes when you actually only tested one idea.

Stop hunting tiny filters. A +0.05R improvement needs thousands of trades to confirm. You will never get there. Chase effects big enough to be visible in the sample you can realistically build.

Split by time, not by luck. Use the earlier 70% of your trades to find candidate rules, then check them on the most recent 30% you have not touched. Much harder to pass by accident than re-sorting the same rows.

Pool where the logic is the same. If the setup runs on the same mechanism across three instruments, that may be one sample of 180 rather than three of 60. Be honest about whether it really is the same thing.

Trade it small while it is unproven. Uncertainty is a position sizing input, not a reason to sit out. Nobody can promise a setup will hold up. You can promise yourself it will not cost much if it doesn't.

So what number should you use?

If you want something to carry around:

  • 30 trades tells you whether you can follow your own rules. It says nothing about edge.
  • 100 trades detects a large edge, 0.5R and up. Most of us do not have one of those.
  • 300 to 500 is the honest range for a normal edge, tested as a single pre-committed hypothesis.
  • 1,000+ if you are chasing small effects or slicing your journal into many buckets.

So no, 100 is not the answer. For most of the questions traders actually ask their journal, it is too low by a factor of three or four. That is uncomfortable, and it is still true.

The upside of knowing this: you stop retiring setups after fifteen losers, and you stop promoting them after fifteen winners. Both of those are noise. Neither is news.

Frequently asked questions

Is 100 trades enough to validate a strategy?

Only if the edge is large, around +0.5R per trade or better, and you tested exactly one hypothesis. For a typical +0.2R edge you need roughly four times that.

Does a higher win rate reduce the sample I need?

Not directly. What matters is your average edge relative to the spread of your outcomes. A high win rate with tiny winners can have worse signal-to-noise than a 35% strategy with large ones.

How do I calculate the standard deviation of my R-multiples?

Export your trades, put each R result in a column, and use the standard deviation function in any spreadsheet. Divide it by your average R, square that, multiply by 8.

What if my setup is significant but barely profitable?

Then it passed the wrong test. Significance only asks whether the edge is above zero. Ask instead whether the lower end of the range still beats your costs and your time.

If you want to see this applied to your own trades rather than a worked example, that is roughly what EdgeFlow's edge analysis is built to do.

Continue reading

Related articles