Nine winners out of fourteen. Up eleven R on the month. You already know the feeling, and you know what it makes you want to do next.
Here is the uncomfortable part. A trader with a genuine edge and a trader on a hot streak have exactly the same month. Same clean screenshots, same confidence, same itch to size up. There is no internal signal that tells the two apart. Most of us have done this at least once: doubled risk right at the top of a run, gave it all back over the next three weeks, then blamed our psychology when the real problem was that we never had proof in the first place.
The good news is you don't need a statistics degree to sort this out. You need your journal and about ten minutes. Five questions. Answer them honestly and you land on one of three verdicts, and the third one is the verdict almost nobody is willing to give you.
Quick vocabulary before we start, because I'll use it throughout. R is one unit of risk. If you risk $200 on a trade and make $400, that's +2R. Expectancy is just your average result per trade in R. If you're not sure what separates a setup from an actual edge, start here and come back.
Question 1: how many trades is this actually built on?
Open your journal and count. Not total trades. Count the trades that match the specific thing you think is working.
This trips people up constantly. Someone says "I've got 400 trades logged," and then it turns out the pattern they're excited about is London session, daily above the 20 EMA, entry on the first pullback, and that's 22 trades. Twenty-two is your sample. The other 378 are a different question.
Now the maths, in plain terms. Say those 22 trades produced 13 winners. That's a 59% win rate. What's the honest range that number could represent? Somewhere around 38% to 77%. Not a typo. That's the 95% range, meaning it's wide enough to cover roughly 95 out of every 100 samples like yours. I'll use that same 95% level everywhere in this post so the numbers stay comparable. With a sample that small, a true 40% shooter and a true 75% shooter can both produce your result without anything strange happening. This is what a win rate confidence interval is for: it stops you reading a point estimate as a fact.
Push the same 59% out to 100 trades and the range tightens to roughly 50% to 69%. Better. Still wide enough that if you're trading a 1:1 payoff, break-even sits inside your range and you genuinely cannot rule it out.
There's no magic number that makes a sample "enough," which is why how many trades it takes to validate a setup depends on your payoff structure rather than on a round number someone put in a course. But under 30 trades on one specific pattern, you're not measuring anything yet. You're collecting.
Question 2: how big is the effect, and is it big enough to see?
Big edges announce themselves. Small edges hide inside noise for years.
Take a strategy that wins 40% of the time, makes +2R on winners and loses 1R on losers. Expectancy is (0.40 × 2R) − (0.60 × 1R) = +0.20R per trade. Genuinely good. Over 60 trades you'd expect about +12R.
Here's the part that gets left out. That +12R is an average, not a promise. Imagine running the exact same strategy over 60 trades a hundred separate times. The results scatter, and at the same 95% level we used above, they scatter from roughly −10R to +34R.
Read the left end of that again. A genuinely good, genuinely profitable strategy can finish 60 trades underwater. Nothing broken, nothing to fix. Just variance. The right end lies the other way: +34R off a +0.20R strategy feels like proof of genius when it was mostly a good run.
Now the small edge. A strategy that wins 52% at 1:1 has an expectancy of +0.04R per trade. Over those same 60 trades you'd expect +2.4R, and that same 95% range runs from about −13R to +17R. The signal is worth two and a bit R. The noise is thirty R wide. Completely invisible. To actually see that edge separate from noise you'd need something in the thousands of trades, which for most of us is a career, not a quarter.
So ask: how big does my journal say the effect is? Then ask whether that size is even detectable at my sample size. A +0.03R edge and a small sample is not a finding. It's a rounding error with a story attached.
Question 3: how many different ideas did you test before you found this one?
This is the question that catches the most false edges, and almost nobody asks it.
You went looking. You sliced your journal by session, then by weekday, then by whether the higher timeframe agreed, then by instrument. Say that's three sessions × two structure states × three setups. Eighteen slices. You found one that looks great and got excited about it.
If your trading had zero edge at all, pure coin flips, and you tested eighteen slices at a normal one-in-twenty threshold, you would expect roughly one of them to look "significant" anyway. Not because anything is there. Because you looked eighteen times.
Nobody does this dishonestly. We do it because searching is how you find things. The fix isn't to stop searching, it's to count your searches out loud. Write down how many cuts you tried before this one showed up. Ten slices means the winner needs to be much more convincing than if you'd tested one pre-planned hypothesis. And whatever came out of the search is a candidate, not a conclusion. Which leads straight to the next question.
Question 4: has it survived data it hadn't seen?
This is the one that does the heavy lifting, and it's simple enough to do on a laptop in five minutes.
Sort your trades by date. Take the first two-thirds and use only that chunk to find your pattern. Then apply the rule, unchanged, to the final third that you haven't looked at. If the effect holds up in the part you didn't use to build it, you've got something. If it collapses, what you found was a description of the first two-thirds rather than a property of the market. That split is the entire logic behind in-sample versus out-of-sample testing for traders, and the chronological ordering matters: shuffling the trades randomly leaks the future into the past and quietly flatters everything.
This is the piece we built EdgeFlow around, honestly. Once there are enough tagged trades to cut — around forty — it keeps the newest 30% back and only calls a pattern verified if it also holds there. Under that, splitting would leave both halves too thin to mean anything, so it runs on everything and labels the result unverified rather than pretending. Either way it ranks findings by the cautious lower bound of the range instead of the flattering middle number, so a candidate that looks like +0.8R but could plausibly be +0.05R doesn't get to outrank a steady +0.3R.
The stronger version of this test is forward data, because the market genuinely hasn't seen it either. Lock the rule, then log the next 30 trades that qualify without changing anything. Forward validating a trading edge is slower and more annoying than backtesting, and it's also the only version that can't be fooled by hindsight.
Question 5: can you state it in advance, in one sentence, with numbers?
Try to write your edge as a sentence a stranger could follow tomorrow morning without asking you anything.
Bad: "I do well when the market gives clean structure."
Better: "Long continuation on EURUSD during the London session when the daily closed above the 20 EMA, entry on the first pullback to the prior day's high, stop below the pullback low, target 2R, no trade on days with red-folder news (high-impact scheduled economic releases) before 10:00."
If you can't write the second version, you don't yet have a rule you can test. You have a feeling about your own trading, and feelings recalculate themselves after every result. The written version can be wrong, and that's the point. A rule specific enough to fail is a rule specific enough to validate.
There's a nasty variant of this worth naming. If you find yourself adding a condition after a loss ("well, that one doesn't count, the range was choppy"), the rule is quietly editing itself to fit history. That's not analysis, that's a moving target with good manners.
The three verdicts
Run the five questions and you'll land in one of three places.
It's probably real. Decent sample on the specific pattern, effect big enough to be visible at that sample, you didn't test forty ideas to find it, and it held up on data you hadn't touched. Trade it at normal size, keep logging, and re-check in another 50 trades. "Probably" is doing real work in that sentence. Nobody can promise you more than probably.
It's probably not there. The range around your result comfortably includes break-even, and it fell apart out of sample. Stop feeding it. This verdict feels bad and it's the cheapest one you'll ever get, because the alternative is finding out over the next six months with real money.
You can't tell yet. This is the honest answer for most patterns in most journals, and it's the one the internet refuses to give you. Twenty trades and a good feeling is not enough evidence to declare either way, and forcing a yes or no onto twenty trades is exactly where false confidence is manufactured. "Not yet knowable" isn't a failure. It's an accurate description of where you are.
That third outcome is why EdgeFlow reports edge, no edge, and insufficient data as three separate results instead of two, and why it won't generate a trade plan off a pattern that hasn't verified. A tool that always produces an answer is not being helpful, it's being agreeable.
What "not yet" looks like on Monday morning
Sitting in the not-yet bucket doesn't mean sitting on your hands. It means:
Keep trading the pattern at your normal risk, not double. Tag every trade the same way every time, so the sample you're building is actually comparable. Pick a number in advance, say 40 more qualifying trades, and don't re-run the analysis every Friday hoping the picture improved. Checking constantly is its own version of testing many ideas, and it will hand you a "significant" week eventually.
Then re-run the five questions when you hit the number. Not before.
Running it once, end to end
You think your best pattern is New York reversals off the prior day's high.
Count: 26 trades match. Sample is thin, so nothing here will be conclusive. Effect: 15 winners, average winner +1.8R, average loser −1R, so expectancy comes out around +0.6R per trade. That's a big effect, and big effects are at least detectable at 26 trades, so this stays alive. Searches: you checked four session-and-timeframe combinations before this one stood out, so treat it as a candidate rather than a discovery. Unseen data: the first 17 trades produced most of the edge and the last 9 were roughly flat. That's a warning sign, not a verdict, because 9 trades tells you almost nothing either. Stated in advance: yes, you can write the rule cleanly.
Verdict: not yet. Promising, not proven. Keep the size, keep the tags, revisit at 60.
That's a genuinely useful answer. It's also the answer that no course, no signal group and no chart-of-the-day account will ever give you, because it doesn't sell anything.
If you want your journal to sort trades into those three buckets for you instead of you doing it by hand every couple of months, that's the whole idea behind EdgeFlow's edge analysis. Either way, run the five questions before you size up.