Most articles about journal mistakes are really articles about discipline. Log every trade, be honest, don't skip days. Fine advice. It's also not the problem I see most often.
The problem I see most often is a trader who is disciplined, who does log everything, and whose numbers are still wrong. Not wrong as in badly calculated. The arithmetic checks out. Wrong as in the number is answering a question nobody asked.
That's a measurement failure, not a character failure. And measurement failures have names. Here are six, all of which get worse when your sample is small. Most of us have been caught by at least three.
Quick vocabulary first. R means "one unit of risk." Risk $200 and make $400, that's +2R. Lose the full stop, that's −1R. Using R instead of dollars lets you compare a small trade to a big one. Expectancy is your average R per trade.
Trap 1: The trades that never made it into the journal
This one is called survivorship bias, and it's the quietest of the six because there's nothing in your data to look at. The evidence is missing by definition.
It doesn't look like cheating. It looks like housekeeping. A trade goes nowhere, you scratch it for +$3 and think "that wasn't really a setup, it doesn't belong in the stats." You take something on a Friday afternoon out of boredom, lose 0.4R, and don't log it because it wasn't part of the plan. Over three months that's maybe twenty trades gone, and they're not a random twenty. They're systematically the mediocre and embarrassing ones.
So imagine you took 100 trades and logged 74. The journal says +0.35R average. The 26 you dropped averaged −0.15R. Real expectancy across all 100 is about +0.22R. You've overstated your edge by more than half, and every decision downstream of it inherits that error.
There's a dead simple check, and it costs five minutes. Take your journal's total R, convert it to dollars using your actual risk per trade, and compare it to your broker's realized P&L for the same period. If the two don't line up, trades are missing or mis-recorded. That reconciliation is the most useful habit in trading journal data integrity, and almost nobody does it.
The rule I'd suggest: if money moved, it goes in the journal. Tag it "not my setup" and filter it out later if you want. What you can't do is un-delete a trade six months from now, when you finally want to know what your boredom costs you.
Trap 2: One trade counted as three
Scale-outs break journals in a genuinely subtle way, and I don't think most traders realize it's happening.
Say you risk 1R, take a third off at +1R, a third at +2R, and leave the last third to trail. Plenty of people log that as three rows, because that's how the fills came in and that's how the broker reports it.
Watch what happens. Take a sample of 40 trades. Twenty are winners you scaled out of. Twenty are losers where you got stopped for the full −1R, in one clean fill.
- Logged as fills: 20 winners × 3 rows = 60 winning rows, plus 20 losing rows. That's 80 rows and a 75% win rate.
- Logged as trades: 20 wins, 20 losses. 50% win rate.
Same trades. Same money. The win rate moved 25 points because winners got chopped into pieces and losers didn't.
Here's the part that trips people up: your total R is unaffected. The dollars are the dollars. What breaks is every per-trade statistic sitting on top of it. Win rate, average win, average loss, win/loss ratio, longest losing streak, and, worst of all, your sample size. You think you have 80 observations. You have 40. When you go asking how many trades you need to validate a setup, you're answering it with a denominator that's inflated by a factor of two.
This is one reason EdgeFlow treats a scaled-out position as one trade with one blended R outcome rather than several. It's a modeling decision, it's arguable, and most journals never state which side they've picked. So ask your tool. If you can't find out, count your open-to-flat round trips by hand and see whether the software agrees.
Trap 3: The month that isn't finished yet
Winners and losers don't resolve on the same timetable. Winners often run into a target and close. Losers, especially loosely managed ones, tend to hang around.
So when you cut a sample at the end of a month and only count closed trades, you're not taking a random slice. You're taking the trades that finished, which skews toward the ones that went your way quickly.
Concrete version. You close 38 trades in August at an average of +0.4R, so +15.2R. You also have 6 positions still open, sitting at an average of −0.7R unrealized, because they're the ones that haven't worked yet. Mark them where they are and it's +15.2R − 4.2R = +11R across 44 trades, or +0.25R each. Your "+0.4R month" was really a +0.25R month with bad news pending.
Six open trades out of 44 isn't extreme. In a slower swing style it's normal. Which is why monthly stats are a weaker unit of analysis than most traders assume: the cutoff is arbitrary, and an arbitrary cutoff isn't a neutral one.
Trap 4: 1R stopped meaning the same thing halfway through
R is only useful if it's a stable unit. The moment your risk per trade changes, R from January and R from June are different currencies, and adding them is like adding euros to dollars because both start with a symbol.
The usual version: you traded 80 trades at $50 risk while you were building confidence, then went to $200 risk. Your first 80 average +$15 a trade, so +$1,200. Your next 40 average −$60 a trade, so −$2,400. In dollars you're down $1,200 and it feels like the strategy broke.
In R, though: +0.3R average for 80 trades (+24R), −0.3R for 40 trades (−12R). Net +12R over 120 trades, about +0.1R a trade. The strategy didn't break. You got unlucky at four times the size.
Both numbers are true. The dollar number tells you what happened to your account. The R number tells you what happened to your process. Mixing them up is how people abandon a working method after one badly-timed size increase.
If you change your risk deliberately, write the date in your journal. Then you can split the sample at that line instead of pretending it isn't there.
Trap 5: 200 trades that are really about 30
Sample size assumes the observations are more or less independent. Trades often aren't.
The first flavor is instrument concentration: 200 trades in the journal, 130 on one index future, 70 spread thin across six other symbols. Your "200-trade edge" is a 130-trade edge on one instrument plus noise. If that instrument had a clean trending stretch, that's what you measured.
The second is time clustering, and it hides better. Say 55 of those 200 trades happened in one three-week window when the market trended hard in one direction. Twelve long trades taken in the same afternoon, in the same impulse, are not twelve independent tests of your setup. They're closer to one observation sampled twelve times.
The check: group your trades by instrument, and separately by week, and look at the counts before you look at any performance number. If one bucket holds more than a third of the sample, your headline stat is mostly describing that bucket.
And don't fix it by filtering, at least not casually. Slicing to "just the good instrument, just the London session, just A+ grade" feels like precision and is actually how a 200-trade sample becomes a 22-trade sample with no warning label. That collapse gets its own treatment in sample attrition when filtering your journal.
Trap 6: One trade doing all the work
Last one, and the easiest to check.
120 trades, +18R total, +0.15R expectancy. Looks like a real edge. Now remove the biggest winner, a +14R runner you held through an unusually strong move. What's left is +4R over 119 trades, about +0.03R each, which is basically flat after fees and slippage.
That doesn't automatically mean you have no edge. Fat winners are the whole point of some strategies, and cutting the best trade out of a trend-following sample is unfair to the method. But you need to know that's the shape of your distribution, because it changes what to expect: long stretches of nothing punctuated by rare payoffs, and you'd better have the patience and the account to sit through the nothing.
Two numbers make this visible immediately. Look at your median trade next to your average trade. If the average is +0.15R and the median is −0.2R, most of your trades lose and a few outliers carry the result. Then recompute expectancy without your top trade. There's more on which numbers earn their place in trading journal metrics.
What I'd actually do about it
Sample size isn't one number you clear and then relax about. It's a property of the specific question you're asking, and every trap above quietly shrinks or inflates the sample sitting behind that question.
So before trusting any stat in your journal, ask three things. What counts as one trade here, and does my software agree with me? Which trades could have been in this sample but aren't? And how much of this result comes from one instrument, one week, or one trade?
If you can't answer those, the number isn't wrong. It's just not about what you think it's about.
One more warning, because it's the natural next mistake. Once you know these traps exist, the temptation is to keep re-cutting the data until the numbers behave. Try enough slices on 200 trades and something will look brilliant by pure chance. That's overfitting your own trading journal, and it does more damage than the original problem because it arrives dressed as rigor.
Nobody can promise you a number of trades that makes any of this go away. What you can do is stop your tooling from making it worse: show a likely range instead of one confident-looking figure, rank setups by the cautious lower end of that range rather than the flattering point estimate, split your history chronologically so a rule found on old trades has to survive newer ones, and say plainly when there isn't enough data yet instead of producing a plan anyway. That's the shape we built EdgeFlow around.
If you want to see your own history measured that way, the trading analytics side is where to start, and there's a free interactive demo on the homepage if you'd rather poke at it first.
Frequently asked questions
Should I delete scratch trades from my journal?
No. Tag them and filter them when you have a reason to. Deleting them removes the only record of what your borderline decisions cost you, and it systematically biases your stats upward because scratches are rarely your best trades.
Should I log partial exits as separate trades?
Log the fills if you like, but analyze at the trade level: one entry to flat, one blended R outcome. Counting partials as separate trades inflates win rate and doubles your apparent sample without adding a single new observation.
How many trades is enough?
There's no universal number, and anyone giving you one is guessing. It depends on your win rate, how skewed your winners are, how many variables you're testing, and how much of your sample is genuinely independent. A concentrated 200 can be weaker evidence than a well-spread 80.
My journal and my broker statement don't match. Which is right?
The broker. Start there, find the gap, fix the journal. A journal that doesn't reconcile to realized P&L can't support any conclusion, however good the charts look.