Two traders spend the same Sunday afternoon doing the same thing.
Both open a journal with a couple of hundred trades in it. Both start combining conditions. Higher timeframe aligned plus a sweep — price running through an obvious high or low and then turning straight back. London session plus expanding volatility. Both find something that looks good. Both change how they trade on Monday.
One of them did research. The other fooled himself.
From the inside, on that Sunday, the two feel identical. Same clicking, same little jolt when a number comes back better than expected. The difference isn't what they found. It's whether they wrote anything down before they started looking.
The two answers you usually get
Ask around about combination testing and you get one of two responses.
The first is that it's basically impossible, so don't bother. That's where the "cap your tags at four" advice comes from, and it sits on a real observation: four tags is often enough to fragment a journal past the point where any slice returns something readable. True enough. But the conclusion doesn't follow. The answer to a sample-size problem isn't to stop asking questions. It's to ask fewer, better ones and know what each one costs you.
The second answer is worse. More confluence is better, stack them up, three conditions beat two. Every condition you add makes the setup rarer and the evidence behind it thinner. At some point you're not filtering for quality, you're filtering for a handful of specific mornings that happened to go your way.
Here's the one opinion I'll push here. Combination testing is at the same time the correct way to find an edge and the most efficient way to manufacture a fake one. Both, at once. The whole job is a protocol that keeps the first without buying the second.
Two words, quickly
R measures a trade in units of what you risked. Risk $200, make $300, that's +1.5R. Get stopped out, that's −1R. It lets you put a small trade and a big one on the same scale.
Expectancy is your average R per trade across a group of trades. An expectancy of +0.15R means that group historically returned 0.15 times your risk per trade, on average.
Why the instinct is right
Before the warnings, the case for doing this at all. I don't want to talk you out of it.
Conditions rarely carry a fixed meaning. A liquidity sweep isn't a good thing or a bad thing. It's a good thing in some contexts and a coin flip in others, and measuring it on its own blends those contexts into one flat, uninformative number.
Imagine you check sweeps across 71 trades and get +0.21R. Mildly positive, nothing to write home about. Now split that set: 41 of those trades had the higher timeframe pointing the same way, and they run at +0.38R. The other 30 didn't, and they run at −0.02R. Same 71 trades, same +0.21R average. The blended figure told you almost nothing. The split told you something you could actually trade.
That's the honest case for combinations, and it's why confluence behaves like an interaction rather than a checklist. Averaging across context hides the edge. So you have to look.
You just have to look in a way that leaves a paper trail.
The protocol
Four steps. None of them need software, all of them fit in a text file, and the setup takes about twenty minutes before you're allowed to touch the data.
1. Name the combinations before you look
Write the list first. Every combination you intend to check, on paper, before you open the journal.
This is the step everyone skips, and it does most of the work. Once the list exists you can't quietly expand it. You can't check twelve things, find one winner, and later remember it as "I had a hunch about London and sweeps." The list is a receipt against your own memory, and your memory is the thing lying to you here. Not the data.
Keep it short. Six to twelve combinations is plenty for one session. If you catch yourself writing twenty-five, you've stopped forming hypotheses and started fishing.
Be specific, too. "Trending" is not a condition, it's a feeling. "Daily close above the 20 EMA at the time of entry" is a condition. If you couldn't have tagged it the same way six months ago, it isn't testable yet.
2. Count them
Write the number at the top of the page. Ten, twelve, whatever it is.
That number sets the bar every result has to clear. Rough rule: multiply how surprising your best result is by how many places you looked. Something that would turn up by luck 3 times in 100 becomes, across ten tests, roughly a 30% chance of appearing somewhere in your list through pure noise. That's not a discovery. That's the expected outcome.
Your real count is always higher than your written one. The slices you clicked through and closed. The idea you eyeballed last March. The filter your mentor handed you, which came out of his own search that you never saw. The number nobody writes down covers why the honest figure is usually two or three times what people admit to.
Write down the count you can defend, then assume it's low.
3. Watch the sample shrink
Record the trade count for every combination next to its result, in the same row.
Not afterwards. Same row, same moment, so the two numbers are impossible to read separately. Most dashboards put the performance figure in 32-point type and the sample size in grey 11-point off to one side, which is exactly backwards. Every filter you add quietly changes the question you're answering, and the trade count is the only thing on screen telling you how far it has drifted.
4. Keep the failures
When a combination doesn't work, it stays in the file. Struck through, marked dead, whatever you like. It stays.
Delete the losers and the survivor stops looking like a survivor and starts looking like a discovery. That's the moment the self-deception locks in, and it happens quietly, usually inside a week. Three months later you genuinely remember testing one idea that worked, because the evidence of the other nine is gone.
Running it end to end
Here's the protocol on made-up but realistic numbers.
Imagine 240 trades. Baseline expectancy +0.08R, win rate 44%. Roughly break-even after costs, which is where a lot of serious traders live for a while.
Four conditions, all recorded before entry: higher timeframe aligned, London session, sweep before entry, expanding volatility. Four singles and six pairs. Ten combinations, written down before opening the file.
Those four are stand-ins. The protocol doesn't care whether your conditions are sweeps, moving-average states, earnings dates or the size of the spread at entry. It only cares that you tagged them the same way every time.
| Combination | Trades | Expectancy |
|---|---|---|
| Baseline (everything) | 240 | +0.08R |
| HTF aligned | 128 | +0.19R |
| London | 96 | +0.02R |
| Sweep before entry | 71 | +0.21R |
| Expanding volatility | 110 | +0.14R |
| HTF + London | 54 | +0.26R |
| HTF + sweep | 41 | +0.38R |
| HTF + volatility | 62 | +0.22R |
| London + sweep | 26 | +0.55R |
| London + volatility | 44 | +0.05R |
| Sweep + volatility | 33 | +0.41R |
Your eye goes straight to London + sweep. +0.55R, nearly seven times baseline. That's the row that gets screenshotted into a group chat.
It's also the row with 26 trades. You started with 240 and you're now reasoning about 11% of your history. One outlier in there moves the average a long way, and at 26 trades you can't tell an outlier from a feature. It's the best of ten, which is precisely what noise produces.
Now look at the row nobody screenshots. HTF aligned: 128 trades, +0.19R. Half your sample, a modest lift, and something you could plausibly lean on.
Then notice what the table says that no single row does. Every pair above +0.35R does it on fewer than 45 trades. The one condition with a real sample behind it — HTF aligned, 128 trades — shows the most modest lift of the lot. That inversion is the whole lesson: the rows that look strongest are strongest partly because they are smallest.
London is worth reading properly too. On its own it's +0.02R. Paired with expanding volatility it's +0.05R across 44 trades, so it isn't a condition that comes alive whenever you give it a partner. Its one impressive number is the 26-trade row you just spent two paragraphs learning not to trust. That's what a dead condition looks like once you slice it enough ways: eventually one thin slice comes back hot. Working out which condition actually carries your edge usually beats collecting more of them.
The same reading tells you something structural. HTF alignment is behaving like a requirement, sweep like a supporting condition, London like neither. Sorting your conditions into required, supporting and avoid is a more durable result from an afternoon like this than any single winning row.
What you do with a survivor
Nothing, for now. That's the hard part.
A combination that survives the protocol is a candidate, not a rule. It gets written down with its trade count, its hypothesis count, the date, and the word UNVERIFIED next to it. Then it gets tested on trades that didn't exist when you found it. Either split your history by date and check whether the older portion's pattern holds in the newer one, or just keep tagging consistently and come back in three months. New data doesn't care how many combinations you tried on the old data, which is why it's the only clean answer to a big hypothesis count.
This is roughly the plain-English version of what EdgeFlow's edge discovery automates: surface the combinations that failed and the ones still short on data alongside the winners rather than only the winner, rank candidates by a cautious lower bound instead of the flattering headline number, split discovery from verification chronologically. No magic in any of it. Run it in a spreadsheet by hand at least once, because that's what makes the cost of each extra combination feel real.
Where this stays uncertain
I'd rather be straight about the limits.
The multiply-by-N correction is crude. Your ten combinations overlap heavily, so the true adjustment is gentler than multiplying by ten, and nobody can hand you the exact figure for your journal. The direction is reliable even when the size isn't.
There's no clean trade count that makes a combination safe, either. It depends on your win rate, how spread out your outcomes are, and how many things you tested. Twenty-six trades is a story. Sixty is a hint. A couple of hundred starts to be an argument.
And the protocol doesn't make you right. It makes you honest, which mostly means it turns confident conclusions into open questions. That feels like going backwards. It isn't. Most of us have spent months trading a rule that came out of an afternoon exactly like the one above, and that costs a lot more than waiting a quarter.
Don't try to reconstruct the past. Open a text file, write tomorrow's date, list the combinations you plan to check, and count them before you look at anything.
Common questions
How many condition combinations should I test at once?
Six to twelve in a session. It isn't a rule: every extra combination raises the bar each result has to clear, and past a dozen you're mostly generating noise you'll have to correct away again.
Isn't testing combinations just overfitting by another name?
Only if you report the winner and hide the search. Testing combinations is how you find conditional edges, which are the realistic kind. Overfitting is what happens when you skip the count, drop the failures, and treat the best row as a conclusion.
What if every combination looks weak?
That's a useful result, not a wasted session. It usually means your edge lives in something you haven't tagged yet, or your baseline is genuinely thin. Either way you know more than the trader who found a beautiful 26-trade row.
When is a surviving combination safe to trade?
When it holds up on trades that didn't exist when you discovered it. Until then it's a candidate, however good the original numbers looked.
If you'd rather not keep the bookkeeping by hand, there's a free interactive demo of the discovery step on the EdgeFlow homepage you can poke at without signing up.