Case Study

Anatomy of an Overfit Setup: A Worked Example

A step-by-step walkthrough using illustrative numbers: 180 trades, eleven tags, and the four filters that turned a break-even record into a fictional 78% win rate over nine trades. Every individual step looks entirely reasonable.

E

EdgeFlow

Everyone nods along when you tell them overfitting is bad. Then they go home and do it, because in the moment it never feels like overfitting. It feels like finally understanding your own trading.

So let's slow one down and watch it frame by frame. One journal, eleven tags, four filtering decisions. Every decision has a reason you'd accept from a friend at a trading meetup. The result at the end is nonsense.

And you still won't be able to point at the step where it went wrong. That's the interesting part.

First, where these numbers come from

They come from my head.

I made them up to be realistic, and I'm saying so up front because a post about traders fooling themselves with numbers shouldn't open with numbers you can't check. This is not customer data. EdgeFlow has no aggregate dataset to draw on, and I'm not inventing one. The arithmetic below is real, every step is one I've watched traders take, and you can run the same process on your own journal tonight.

The journal

Call him Sam, a composite of about four traders I know, one of whom is me.

Sam has 180 trades logged over roughly eleven months, everything measured in R, which just means multiples of what you risked. Risk $150 and make $300, that's +2R. Lose the lot, that's −1R.

His baseline: +0.05R per trade over 180 trades, win rate 41%.

About +9R for eleven months of work. After commissions and slippage, call it zero. It's the number that starts the whole search.

He tags every trade with eleven things:

  • Setup: higher-timeframe (HTF) trend state, level type, trigger type
  • Environment: session, volatility bucket, day of week
  • Execution: entry timing, entry timeframe, risk size
  • Management: break-even moved or not, target type

Eleven tags isn't an excessive journal. That's a decent one. Hold that number, we come back to it.

Filter one: drop the trades that weren't really my strategy

Twenty-two of the 180 are tagged as rule breaks. No plan, chased entry, or the classic: a second trade thirty seconds after a loss because it felt owed to him.

His reasoning: I'm trying to measure whether my strategy works. Those trades weren't my strategy. They were me having a bad afternoon.

Honestly? Fair. If you're evaluating a process, trades that ignored the process measure something else.

180 → 158 trades. +0.15R per trade.

Three times the baseline, and he hasn't touched the strategy yet. His biggest problem might be discipline rather than method. Most of us skate past that and keep clicking.

Filter two: only the sessions I actually trade properly

Sam is a London and New York open trader. But 52 of the remaining trades are Asia session and New York afternoon, mostly from a spring stretch when he was trying to be more active.

His reasoning: I don't trade those hours anymore. Including them tells me about a version of me that no longer exists.

Also fair. Arguably it's just data hygiene: if a rule already changed, older trades describe history rather than the current system.

158 → 106 trades. +0.31R per trade.

Twice as good again. And notice how clean the story is now. He was diluting a solid London approach with tired late-day trades. Coherent, flattering, possibly even true. Nothing so far has proved it.

Filter three: only when the higher timeframe agreed

Now the actual strategy question. He splits what's left by higher-timeframe trend state and keeps the trades where the HTF was trending his way. Ranging and counter-trend go.

His reasoning: my whole method is continuation. Taking it into a range is a different trade with a different distribution. Of course I should measure them separately.

That one is more than fair, it's textbook. Setups behave differently across regimes, and lumping them together hides real information. Separating them is most of what serious work on confluence combinations consists of.

106 → 47 trades. +0.55R per trade.

Half an R per trade. If that held up he'd have a very good system. We're also down to roughly a quarter of the trades we started with, and nobody has said a word about it.

Filter four: the one that looks smartest of all

Last step, and it's the one I'd have been proudest of in his shoes.

He notices he moves his stop to break-even the moment a trade goes a bit his way. Of the 47 left, only 9 were allowed to run to a structural target, meaning he left them alone until price reached the next level marked on his chart. The other 38 got the break-even treatment, and those 38 average +0.06R. Basically nothing.

His reasoning: I'm strangling my winners. The management is eating the edge.

The most defensible filter of the four. Moving to break-even early genuinely does cap winners while doing nothing about losers, and "my exits are the problem" is one of the more common conclusions I hear traders reach. He may well be right.

47 → 9 trades. Win rate 78%. +2.62R per trade.

Nine is small enough to just list:

  • Winners: +1.9, +2.2, +2.4, +3.1, +3.4, +4.0, +8.6
  • Losers: −1.0, −1.0

Winners total +25.6R, losers −2.0R, net +23.6R over 9 trades, so +2.62R each. The maths is correct. Nothing on the screen is lying to him.

The whole thing on one page

StepFilterTrades leftR per trade
0Everything180+0.05R
1Rule-following trades only158+0.15R
2London and NY open only106+0.31R
3HTF trending, aligned47+0.55R
4Ran to target, no break-even9+2.62R

Read the third column downward and ignore everything else. 180, 158, 106, 47, 9.

That column is the story. The fourth column is what he wrote in his notebook.

It's why the trade count behind every condition is on screen in EdgeFlow, and why the remaining sample updates as you exclude one. Watching a sample shrink while you click lands very differently from arriving at 9 and reading the number beside it.

Three things that break it

One trade is carrying a third of it

Take out the +8.6R, the one where a data release happened to run his way. Eight trades left, +15.0R, +1.88R each. Still looks great, so the problem isn't only the outlier.

Try something gentler. Take the two smallest winners, +1.9 and +2.2, and imagine they'd gone the other way. Both were probably within a few ticks of the stop at some point. New numbers: +1.94R per trade, five winners from nine, so 56%.

Two ordinary trades flipping and the headline win rate drops 22 points. That's not an edge with some noise on it. That's noise with an edge painted on.

78% doesn't mean 78%

A win rate from 9 trades carries enormous uncertainty. Run the confidence interval on a win rate honestly — that's the range of true win rates your handful of trades would be a normal outcome for — and 7 from 9 is consistent with anywhere from about 45% to 94%.

Call that a pedantic footnote if you like. It's the actual meaning of the number. He wrote "78% win rate." What the data supports is "somewhere between a coin flip and outstanding, can't tell which yet." One of those changes your position sizing. The other one shouldn't.

He ran more tests than he thinks

Eleven tags. Every tag he can filter on is a test, so that's 11. Every pair is 55. Every triple is 165. Every group of four is 330. Search all combinations up to four filters deep and you've run 561 tests. At a normal 5% threshold, roughly 28 come back looking good on luck alone, even if not one of his tags means anything about the market.

And 561 is charitable: it counts each tag once, when most have three or four values he could have picked. Session alone is four choices.

He opened maybe fifteen of those doors and remembered one. That's the mechanism, and counting your hypotheses before trusting any result is the only defence I know of that actually works.

So which step was the mistake?

None of them. That's the whole reason I wrote this out.

Reread the four reasons. Excluding rule breaks when measuring a strategy, excluding hours you no longer trade, separating trending from ranging, checking whether early break-even caps your winners. Every one of those is correct, and the last is probably the most valuable question in the list.

The error lives in the sequence. Four defensible decisions taken one after another, each one chosen after seeing what the previous one produced, with no record of the branches he didn't take and no fresh data at the end. Every step was a fair question. Together they were a search.

And a search always finds something. Blame the arithmetic for that, not Sam. Give me 180 coin flips and eleven ways to slice them and I'll hand you a subset with a 78% "win rate" every time. The journal version of curve-fitting is harder to spot than the backtest version precisely because each individual click is so reasonable.

There's a quieter problem underneath it. With eleven tags per trade his sample was fragmented before he started, and four tags is usually enough to fragment a journal past the point where subsets mean much. He never had 180 trades of evidence about any specific combination. He had 180 trades and a filing system fine enough to slice them into crumbs.

What he should actually do with this

Not delete it. Filtering isn't pointless, it's how hypotheses get generated. The mistake is treating the output of a search as a conclusion instead of the start of a test.

Cut it back to one condition. Test break-even management against everything else across all 158 rule-following trades, not the 9 survivors of a four-filter stack. Working out which single condition carries the result beats hunting for the perfect stack, and it's a much shorter afternoon.

Report both numbers. "+0.15R when I follow my rules, +0.05R including the days I don't" is the honest pair. Excluding rule breaks is legitimate as its own analysis, but the number with your bad afternoons in it is the one your account pays out.

Write the rule down precisely, then leave the old data alone. Precise enough that a stranger could apply it to the next 40 trades without asking a question. Vague rules get quietly bent to fit whatever happens next.

Split by time, not by preference. Discover on the older trades, verify on the newer ones, in date order. In-sample versus out-of-sample for traders covers why shuffling randomly instead lets the future leak backwards into the test.

Expect the effect to shrink. If he lets winners run on the next 30-odd qualifying trades and comes back with 48% and +0.2R, that isn't a failure. That's a real improvement on +0.05R, and roughly what a genuine effect looks like once the luck is out. It only feels like a letdown against a number that was never real.

Nobody can promise the shrunk version stays positive. Sometimes there's nothing there and the flat baseline was the truth all along. Ranking candidates by a cautious lower bound instead of the prettiest number in the search is how we try to keep that honest, and the edge discovery walkthrough covers how it fits together.

The uncomfortable summary is short. Sam didn't find an edge. He found nine trades and a story, and the story was good enough that he'd have sized up on it.

His earliest warning was free: the trade count. If there's a chain in your own journal ending under 30 trades, read the counts on the way down before you read anything else. There's a free interactive demo on the EdgeFlow homepage if you'd rather watch it happen on someone else's numbers first.

Continue reading

Related articles