Trading Journal

Every Filter You Add Quietly Changes the Question

Filter your journal to A+ setups, London session, with HTF bias, and 380 trades becomes 11. The number still displays confidently, but it now answers a far weaker question. How to watch sample attrition happen in real time.

E

EdgeFlow

You open your journal on a Sunday. 380 trades in there. The overall number is fine. Not exciting, not embarrassing. Fine.

So you start clicking.

Setup grade: A+ only. Session: London. Higher-timeframe bias: aligned.

And there it is. A beautiful number. Way better than your overall number. You sit back and think, okay, that's the edge, I've been diluting it with garbage for a year.

Most of us have done exactly this. What almost nobody does in that moment is look at the small grey text next to the chart, the one saying how many trades are actually left.

Eleven.

What actually happened when you clicked

Let's put real numbers on it so it isn't abstract.

Quick vocabulary first, because I'm going to use it the whole way through. R is just a way of measuring a trade in units of what you risked. If you risked $200 and made $400, that's +2R. Lost the full $200, that's −1R. It lets you compare trades of different sizes. Expectancy is the average R per trade across a group of trades. Positive expectancy means that group, historically, made money per trade on average.

Imagine the filter chain goes like this:

Filter appliedTrades leftAverage R per trade
Everything380+0.18R
Setup grade = A+143+0.31R
...and London session52+0.44R
...and HTF bias aligned11+1.36R

Look at that bottom row. +1.36R per trade. If that were real and repeatable, you'd have a genuinely excellent system.

Now let's open up those eleven trades, because eleven is small enough to just list.

Winners: +1.2, +1.8, +2.1, +2.4, +2.6, +3.0, +5.7 Losers: −1.0, −1.0, −1.0, −0.8

Add it up: +18.8R from winners, −3.8R from losers, net +15.0R. Divide by 11 and you get +1.36R. Win rate 7 out of 11, so 64%. The math is correct. The screen is not lying to you.

Here's the problem. Take out that one +5.7R trade, the one where you happened to be on the right side of a news candle. Now you have +9.3R over 10 trades. +0.93R. One trade was carrying nearly a third of the headline.

Now something gentler. Take the two smallest winners, the +1.2 and the +1.8, and imagine they'd gone the other way. Both stopped out at −1R instead. Not a wild scenario. Those were probably a few ticks from your stop at some point.

New total: +10.0R over 11 trades. +0.91R. Win rate drops to 5 out of 11, so 45%.

Two trades. Two ordinary trades going the other way, which they easily could have, and your "edge" fell by a third.

Do the same experiment on the 380-trade sample and flipping two trades moves the average by roughly one hundredth of an R. You wouldn't even notice. That's the whole thing in one comparison. At 380 trades, one trade is 0.26% of your evidence. At 11, one trade is 9%.

Three clicks, four completely different questions

The part I find more interesting than the sample size is what the filters did to the question.

Before you clicked anything, the number on screen answered this: does the way I trade, across everything I actually do, make money?

That's a strong question. Boring, but strong. It's the one your account balance answers too.

Click "A+ only" and the question becomes: does the way I trade make money on the trades I personally graded as A+?

Add London: ...and only in one session.

Add HTF bias aligned: ...and only when I'd already decided the higher timeframe agreed with me.

Four different questions. The dashboard displayed all four in exactly the same font, in exactly the same box, with exactly the same confidence. Nothing on screen said "heads up, you're now answering a much narrower question with much less evidence."

And that last question isn't really a question about the future at all. It's a question about eleven specific mornings. You can describe what happened on those mornings with total accuracy. You cannot generalise from them.

The A+ filter is the sneaky one

I want to single out the setup-grade filter, because it causes a problem the other two don't.

London session is objective. The clock says what it says. HTF bias is mostly objective, assuming you wrote it down before entering and didn't revise it later.

"A+" is a grade you gave yourself. And the honest question is: when did you give it?

If you graded the setup before entry, fine, that's a real prior belief and it's testable. But a lot of us grade after the trade closes, or half-consciously while it's running. Once that happens the grade is partly an outcome, not a condition. Trades that worked feel like they were A+. Trades that chopped around and stopped you out feel like B setups you shouldn't have taken.

Filter to A+ and of course the win rate goes up. You've filtered to trades that partly won. That's not an edge, that's a definition eating its own tail.

Check this in your own journal before you trust any grade-based filter. Are your A+ grades timestamped with the entry, or added later? If later, the filter is contaminated and needs rebuilding from objective pieces: structure, level, trigger, session, volatility. Those you can defend. A letter grade you assigned while feeling good, you can't. That's part of why it pays to be deliberate about what you track in a trading journal rather than logging a vibe and a screenshot.

Your filter bar is a search engine for flattering subsets

Here's the uncomfortable structural point.

Say your journal has 3 setup grades, 4 sessions, 2 HTF states, 5 entry models, and 3 volatility buckets. Nothing extreme, that's a fairly normal journal. The number of subsets you can build by combining those runs into the hundreds. Add a date range and it's effectively unlimited.

Somewhere in there, a combination looks fantastic. There is always one. That would still be true if your trades were literally coin flips. Randomness clusters. Give me enough ways to slice a random dataset and I'll hand you a slice with a 70% win rate and a tidy equity curve.

The filter bar doesn't know you're searching. It just answers. And because you only stop clicking when you find something good, the process has a built-in bias toward stopping on noise. Nobody stops on the combination that looks mediocre.

This is the same trap as running lots of combinations deliberately, which I've written about in testing condition combinations. And the fragmentation problem gets worse fast once you're tagging heavily. Four tags on a trade sounds thorough, but it slices your sample into pieces the moment you start querying it, and four tags is often enough to fragment your sample past the point of usefulness.

The habit that fixes most of this

It's small and slightly annoying and it works.

Read the trade count before you read the performance number.

That's it. Train your eye to go to the count first. Not the win rate, not the R, not the equity curve. The count. If it says 11, you're still allowed to look at the rest, but you look at it the way you'd look at a rumour. Interesting, unverified, possibly nothing.

A second habit, slightly harder: before you touch a filter, say out loud what you expect to find. "I think London beats New York for this setup." Then click, and check the count. If you were wrong, that's information. If you had no prediction, you weren't testing anything. You were browsing.

This is honestly the main reason we built the remaining trade count to update live on every condition toggle in EdgeFlow, sitting right next to the toggles instead of buried in a corner. You watch 380 fall to 143 fall to 52 fall to 11 as you click. The cost of each filter is visible at the moment you pay it, not discovered three conclusions later.

So how many trades do you actually need?

Nobody can give you a clean number, and anyone who does is selling something. It depends on your win rate, how spread out your outcomes are, and how many things you've already tested. A high-win-rate scalping approach stabilises faster than a 35%-win-rate swing approach living off occasional +5R runners.

What I can tell you is the rough feel of it. Eleven is a story. Fifty is a hint. A couple hundred starts being an argument. The reasoning behind those ranges is in how many trades to validate a setup, including why the honest answer is a range rather than a single number.

That range idea matters more than people think. "My edge is +1.36R" is almost never a true statement. "My edge is probably somewhere between −0.2R and +2.5R, and eleven trades can't narrow it further" is true, and it's the kind of statement you can act on. Usually the action is: don't bet the account on it yet.

What filtering is genuinely good for

I'm not telling you to stop filtering. Filtering is how you find things. I'm telling you that filtering produces hypotheses, not conclusions, and the difference is everything.

A few things that keep it useful:

Filter one condition at a time against your full baseline, rather than stacking three at once. One filter on 380 trades leaves you something you can read. Three leaves you eleven. If you want to know which single condition is doing the work, that's a different and much better-posed question, and isolating which condition carries your edge is a more productive use of an afternoon than hunting for the perfect stack.

Split your data by time. Discover on the older portion, verify on the newer portion, and keep it chronological so you're never testing an idea on the same trades that suggested it. It's also why EdgeFlow ranks findings by the cautious lower bound rather than the flattering headline number, and says plainly when there isn't enough data to answer. A tool that refuses to answer beats one that answers confidently from eleven trades.

And when a thin combination looks great, don't discard it. Write it down as an open question, tag it consistently going forward, and check back in three months with trades that didn't exist when you formed the idea. That's slow, and it's the only honest version of this. The edge discovery approach walks through that process properly.

The eleven trades might turn out to be real. They might. But right now they're eleven trades wearing the costume of a discovery, and the only thing standing between you and a badly-sized position is whether you looked at the count before you looked at the number.

Common questions

Why do my trading stats change so much when I filter?

Because each filter removes trades, and with fewer trades every remaining trade has more influence on the average. Going from 380 trades to 11 means one outcome swings your numbers by roughly 9% instead of 0.26%.

Is a small filtered sample useless?

Not useless, just weak evidence. Fine as a hypothesis to test going forward. Not fine as a reason to change position sizing or abandon a setup today.

Should I stop using filters in my journal?

No. Use them to generate ideas, one condition at a time, and always check the remaining trade count before reading any performance figure.

Why does filtering to A+ setups always look good?

Often because the grade was assigned with some knowledge of the outcome. Rebuild the filter from objective conditions recorded before entry and the effect usually shrinks.

If you want to see how much your own numbers move as the sample shrinks, open your journal and click through a filter chain while watching only the count. It's a five-minute exercise, and it changes how you read every stat afterwards.

Continue reading

Related articles