Expectancy

One or Two Trades Are Probably Carrying Your Expectancy

Remove your single best trade and recompute. If your average R collapses, your expectancy was a story about one trade rather than a property of a strategy. A robustness check any trader can run on their journal today.

E

EdgeFlow

Most of us check expectancy once, see a positive number, and move on. Fair enough. It's the number the whole journal builds toward.

Here's a cheap way to find out whether that number means anything.

Open your journal. Find your single best trade. Delete it. Recompute.

If the average barely moves, good sign. If it falls off a cliff, you weren't measuring a strategy. You were measuring one trade that happened to go your way.

Almost nobody runs this check. Most expectancy write-ups stop at the formula, hand you a positive number and call it a day. Even once you have the formula and how to read it in R down, the formula is the easy part. Whether the number survives contact with reality is the part that decides how much money you're allowed to risk on it.

First, the two words you need

R-multiple just means "how many times my initial risk did this trade make or lose?" Risk 100 dollars, make 250, that's +2.5R. Risk 100, hit your stop, that's −1R. Measuring in R lets you compare a small position from March with a big one from July without the position sizes muddying everything.

Expectancy is your average R per trade. Add up every R result, divide by the number of trades. Forty trades that total +12R gives you +0.30R per trade. That's it. No formula gymnastics needed if you're already recording outcomes in R.

Now the interesting question: is that +0.30R a property of your process, or a property of one Tuesday in April?

The test, with actual numbers

Imagine a journal with 60 trades. Total result: +18R. So expectancy is 18 ÷ 60 = +0.30R per trade.

That looks strong. On 1% risk, you're clearing roughly 0.3% of the account per trade before costs. Most traders would happily size up on that.

Now find the best trade. Say it's a +14R runner. You caught a gap-and-go (price jumps at the open and just keeps going), left it alone, and it ran all day.

Take it out. You now have 59 trades totalling +4R.

4 ÷ 59 = +0.07R per trade.

Your edge just dropped by more than three quarters because one row left the spreadsheet. At +0.07R, after spread and commissions, you're somewhere between break-even and slightly negative. That's not a strategy you size up on. That's a strategy you keep testing.

And nothing about your trading changed. The only thing that changed was the removal of a single observation. If one observation can do that, the number was describing the observation, not the process behind it.

What a robust edge looks like under the same test

Take a second journal. Also 60 trades, also +18R, also +0.30R expectancy. Identical headline number.

But here the best trade is +4.2R. Remove it and you're left with 59 trades totalling +13.8R, which is +0.23R per trade.

Down a bit. Still clearly positive. Still a plan.

Two traders, same expectancy on paper, completely different confidence in what happens next. The first one is holding a lottery ticket she already cashed. The second one is holding something that behaves like a repeatable process.

This is the whole point of the exercise. Expectancy on its own is a single number, and single numbers hide their own fragility.

Mean versus median, and what the gap actually means

Second thing to look at, and it costs you one more calculation.

Sort every trade by R, smallest to largest, and take the middle value. That's your median R. The mean is the average. The median is the typical trade.

Back to the first journal. Suppose 24 of those 60 trades were winners, so a 40% win rate, and most losers hit the full stop at −1R. Line them up and the middle trade is a loser. So:

  • mean R: +0.30R
  • median R: −1R

That gap is not automatically a problem. Plenty of real, durable strategies live exactly there. Trend following has a negative median and makes its money in the tail. If you take small, frequent losses and let a few winners run, a losing median is the design, not a bug.

The gap tells you where the money comes from. Your job is to then ask a harder question: how many of those tail trades does my sample actually contain?

One +14R winner in 60 trades is not evidence that +14R winners arrive at a rate you can count on. It's evidence that one happened. To claim the tail is real, you need to have produced several of them, under different conditions, following the same rules.

That's the difference between "my edge is asymmetric" and "I got lucky once and built a story around it." Both look identical in a mean.

Also drop your worst trade

Run the mirror version. Remove the single biggest loser and recompute.

In our first journal the worst trade was −2.5R, a stop that got jumped through on news. Remove it: 59 trades, +20.5R, which is +0.35R. Barely moves.

That asymmetry is the tell. Removing the best trade destroys the edge. Removing the worst trade changes almost nothing. Your upside is concentrated in one event and your downside is spread evenly across everything. In practice that means your losses are reliable and your wins are not, which is a rough combination to size.

The healthiest pattern is when both removals move the number by a similar, modest amount. Nothing is carrying the result on its own.

The concentration check

If you want one more angle, this one is my favourite because it's brutally clear.

Sort your trades by R, best at the top. Then count how many trades it takes to reach half your total profit.

Say a journal has 80 trades and +24R total. Half of that is +12R. If your top 3 trades already cover it, then 3.75% of your sample produced half the result. Everything else, all 77 remaining trades, contributed the other half at roughly +0.16R each.

Compare that to a journal where it takes 18 trades to reach half the profit. Same total, entirely different structure. In the second case the edge is spread across the sample, which is what "repeatable" actually looks like in data.

There's no magic threshold here and I'm not going to invent one. But when two or three trades out of eighty carry half your P&L, you should treat your expectancy as a hypothesis rather than a measurement.

This kind of thing is worth checking on the sub-groups too, not just the whole journal. Filter down to "London session, HTF (higher timeframe) aligned" and you might have 22 trades with a beautiful +0.55R expectancy, of which one trade contributes +9R. Slicing your data into smaller buckets makes outlier dependence much more likely, not less, which is one of the small sample traps that turns a promising filter into a wasted quarter.

Profit factor and win rate have the same problem

Worth saying, because traders often reach for a second metric to confirm the first.

Profit factor is gross profit divided by gross loss. That +14R winner sits right in the numerator. It inflates profit factor exactly the way it inflates expectancy, so it isn't independent confirmation of anything. It's the same distortion wearing a different hat, and it's one of several reasons profit factor is a weaker number than it looks.

Win rate has a different weakness. It ignores size entirely, so an outlier can't distort it. But it comes with its own uncertainty band that most traders never look at. Forty percent on 60 trades is genuinely compatible with a true rate anywhere from roughly the high twenties to the low fifties, which is why the confidence interval around your win rate belongs next to the win rate itself.

Neither metric rescues you here. You have to look at the distribution.

So what do you do when your edge fails the test?

Not "throw the strategy away." That's the wrong reaction and it's how people abandon things that were fine.

Start with the honest one. Did that outlier come from your rules, or from an accident?

Go read the trade. Really read it. Did you hold it because your plan said hold, or because you were on a plane and couldn't close it? Did you let it run past your normal target because a rule told you to, or because it felt too good to touch?

If the +14R came from behaviour that isn't written down anywhere in your plan, you can't count it. You didn't execute the strategy; you executed something else that happened to work. Either write the rule down and start collecting evidence for it, or accept that your real expectancy is the +0.07R version.

Sometimes the answer is genuinely encouraging. You find that your plan does allow runners, you just rarely hold them. Then the problem isn't the strategy, it's that your management is chopping the tail off the distribution the edge depends on. Management is part of the edge, and a strategy that only pays in the tail cannot survive a habit of taking profits at +1.5R.

And then the boring answer: keep trading it small and collect more data. Nobody can promise you the outlier repeats. More trades is the only thing that settles it. What you shouldn't do is size up on a number you already know is fragile.

My own rule, and you're free to disagree: size on the trimmed number, plan for the full one. Risk as if the best trade never happened. Build your targets and your patience around the possibility that it does.

This is also the reason EdgeFlow scores discovered combinations on stability, not just on average result. A combination whose entire edge sits on one trade isn't a plan you can size up on, so it shouldn't rank above a steadier one that looks slightly worse on the headline number. Ranking by the cautious lower bound rather than the flattering average is less exciting and a lot more survivable.

What this test won't tell you

It won't tell you the strategy is dead. A fragile expectancy on 60 trades is mostly a statement about 60 trades.

It won't tell you the outlier can't happen again. Some strategies are genuinely built on rare, large wins, and for those the test will always look ugly until the sample gets big enough to hold several of them.

And it won't tell you what your true expectancy is. Nothing will, exactly. What it tells you is how much weight your current number can carry, which is a smaller question but a much more useful one.

Ten minutes. One deleted row. You'll know more about your edge than the formula ever told you.

If you want the whole picture behind this, start with the expectancy fundamentals and then come back and run the test on your own journal.

Continue reading

Related articles