Every conversation about trading journals is about features. Charts, tags, screenshots, AI summaries, calendar heatmaps, replay. Nobody asks the boring question underneath.
Is the arithmetic right?
Does the number in the box equal what happened in your account? If it doesn't, everything built on top is decoration: your expectancy, your win rate by session, your filtered A+ setup analysis.
Broken stats almost never look broken. A corrupted win rate is still a plausible percentage, and a wrong expectancy is still a tidy number with two decimals. Nothing flags it, so you go on sizing your positions off fiction for eight months.
Two definitions first. R is a trade measured in units of what you risked: risk $200, make $400, that's +2R; lose the full $200, that's −1R. Expectancy is your average R per trade across a group of trades. Those two are the numbers most likely to be wrong, and the ones you're most likely to trust.
Seven checks. Run them on whatever you're using. Run them on ours.
1. Does the bottom line reconcile with your broker?
Start here, because it catches almost everything else downstream.
Pick a closed month and add up every trade in your journal. Now open your broker statement and look at realised P&L for the same period. Not the equity change, which includes deposits and open positions.
Say your journal says +$3,140 and the broker says +$2,610. That $530 gap is not rounding. It sits inside every R-multiple too, because the same wrong fills or missing fees produced both numbers. Match to within a few dollars and you're fine; if it doesn't match, the rest of this list tells you where to look.
2. What does your journal do with a scaled exit?
You risk $200. Take half off at +1R, let the rest run to +3R. The honest record is one trade, one decision, weighted result: (0.5 × 1R) + (0.5 × 3R) = +2R. One row.
Some journals, particularly ones importing raw fills, log that as two: a +1R winner and a +3R winner. That quietly rewrites your win rate.
Watch what it does across a sample. Imagine 200 trades, 88 winners, a 44% win rate. You scale out of 60, and like most of us you almost never scale out of a loser, so call all 60 winners. Split in half, those rows become 120, so the journal holds 260 rows, 148 of them winners.
44% becomes 57%.
Your trading didn't change here; the way a database counted one decision did. Your average R per row drops at the same time, because you've added a pile of small +1R partials, so your payoff ratio — average winner divided by average loser — sinks with it.
Find one scaled trade you remember clearly. One row or several? If it's several, every aggregate you have sits on an inflated trade count.
3. Is R calculated from your real average entry?
You wanted 3 lots of EURUSD at 1.0850. You got 1 lot there and 2 more at 1.0856 as price moved against you. Stop at 1.0820, exit at 1.0910. Your journal recorded the entry as 1.0850, so it thinks you risked 30 pips. Your weighted average entry — the average price you really paid across both fills — is 1.0854, so you risked 34.
- What the journal reports: 60 pips on 30 pips of risk = +2.00R
- What actually happened: 56 pips on 34 pips of risk = +1.65R
Same thing in shares. You want 600 at $48.20 and get 200 there, another 400 at $48.36, stop at $47.90. The journal thinks you risked $0.30 a share, so $180. Your real average is $48.31, so you risked $0.41 a share, or $244. Divide by $180 and every R comes out flattering.
Not dramatic on one trade. If that pattern repeats across a few hundred, a measured +0.20R expectancy could really be nearer +0.16R. How big the gap gets depends on how often you get partial fills: fill in one clip most days and it barely touches you, work size into a thin market and it's constant.
Your stop deserves the same check. If the journal computes R from your planned stop rather than where you actually got filled on a gap, every loser is understated, and understated losers next to overstated winners is how a break-even system starts looking like an edge.
4. Are fees and financing inside the R, or sitting in a separate column?
Most journals show gross R next to net dollars, which is where costs go to hide.
Say you risk $200 a trade and pay $5 round-turn — the cost of getting in and back out again — in commission and spread. That's 0.025R per trade, which sounds like nothing. Across 400 trades it's 10R, and if your measured expectancy is +0.10R your total is 40R gross and 30R net. A quarter of your edge, in a column you never look at. Swing traders have a second one: overnight financing on leveraged positions.
Check it directly: open one trade's detail view and see whether the R-multiple moves when commissions are entered. If it doesn't, your R is a pre-cost number — the performance of a hypothetical trader who pays nobody.
We invited you to run these on us, so here is where EdgeFlow sits on its own list. R is a number you supply. On a CSV import we'll work it out from your entry, stop and exit if those columns are there, but that's the price you typed or your broker exported, not a weighted average rebuilt from individual fills. Fees, commission and swap get read in and stored when your file has them, and they are deliberately kept out of the R-multiple — every expectancy figure we show you is a gross number. So check 3 is yours to get right, and on check 4 you should assume our R is pre-cost and do the subtraction yourself. What we do stand behind is the layer above: that the aggregate math on your inputs — expectancy, win rate, the edge statistics — is computed correctly.
Ratio metrics are hit hardest. Profit factor is gross profit divided by gross loss, so costs shift it further than you'd expect, one of several reasons profit factor is a weaker number than it looks.
5. Whose clock is your journal using?
Timestamps break in ways that don't look like breakage. Three clocks are usually in play: broker server time, often UTC+2 or UTC+3; UTC, which most databases store; and your local time, which gets displayed. Break the conversion anywhere in that chain and every time-based conclusion is suspect.
The day boundary is where it bites. A trade taken at 22:40 your time rolls into the next calendar day on a server two hours ahead. Do that regularly and your Tuesday bucket fills up with Monday's US session.
So you filter by day of week, find Tuesday running +0.34R over 41 trades while Friday is negative, and start planning your week around it. Except a chunk of those "Tuesday" trades were late-Monday New York trades, so what you found was a session effect wearing a weekday label.
Daylight saving makes it worse. Europe and the US switch on different dates, so for a few weeks a year the offset moves by an hour and your session tags move with it. Nothing errors. The data just gets a little more wrong every March and October.
Spot check: find your earliest and latest trade of any week and confirm the displayed times match what you remember. If a trade you took at the London open is filed under Asia, stop trusting every session comparison you've run.
6. Are there trades that never closed?
Filter your journal to open positions and look at the bottom of the list. Most of us have a row down there we'd stopped thinking about: still open, months old, deeply underwater.
Some journals count an open trade as $0 P&L and include it in the count, dragging expectancy toward zero. Some exclude it, which is more honest. Some leave it sitting as a winner because you logged a target and never the exit.
Run the numbers. 150 closed trades, 69 winners, a 46% win rate. There are also 6 open rows, all underwater, averaging −1.5R at current price. Mark them to market — value them at today's price as if you closed them right now — and you have 156 trades, 69 winners, 44.2%, and 9R of unrealised loss your expectancy has never seen.
Anything on that list older than your typical hold time is a hole in your statistics.
7. Does the trade count match, and did you log the ones you hated?
The last check is partly about software, partly about you.
Software half: count the rows in your journal for one month against your broker statement. Duplicate imports and dropped rows are both common, and both invisible unless you count.
Human half is bigger. Which trades didn't make it in?
Imagine you skip logging three tilt trades a month: twelve minutes after a stop-out, sized wrong, no setup, closed in shame. Over a year that's 36 trades you never recorded, averaging maybe −2R each.
Your journal says 400 trades at +0.12R, so +48R. Add the missing 36 trades at −2R, so −72R, and the true picture is 436 trades at −0.055R per trade.
The sign flipped. Not because the strategy is bad. Because the sample was curated by your ego. No software catches this one, which is why it pays to be deliberate up front about what you track in a trading journal and to treat "log everything" as a rule rather than an intention.
Correctness deserves to be on the comparison list
The opinion this post exists to argue: accuracy belongs on the list when you choose a journal, right next to price and features. It never is. Comparison pages line up integrations, chart types and monthly cost, and none ask whether the expectancy calculation was ever verified against known inputs.
We hand-verified EdgeFlow's core metrics across two independent testing passes: feed in trades whose correct expectancy, win rate and R-multiples you already worked out on paper, then confirm the output matches. You can demand that from any vendor, us included. The EdgeFlow vs TradeZella comparison covers the feature-level view, but run these seven checks on whichever tool you pick.
Ask any vendor how expectancy is computed. A formula is a good answer; a feature description isn't.
What to do when a check fails
Don't rebuild eighteen months of history in a weekend. You'll get bored, cut corners, and produce fresh errors. Fix the input first, so everything from today forward is clean, then repair backwards only as far as you need. Testing a hypothesis about London-session reversals? Only those trades have to be correct.
Recompute one metric by hand as your acceptance test. Twenty trades, expectancy on paper, checked against the screen. Enough to catch a systematic error, few enough to finish. If you're unsure which numbers earn that effort, the metrics worth checking is a shorter list than most dashboards suggest.
Then be careful how fast you start slicing. Filtering down to a flattering subset produces confident nonsense from accurate inputs, and correct arithmetic on 11 trades is still a small-sample trap, just an honest one.
Getting your numbers right is a low bar, and it still isn't the one the comparison tables measure. Correct data doesn't give you an edge. It tells you whether you have one.
Common questions
How do I know if my trading journal is accurate?
Reconcile one month against your broker statement. If realised P&L and trade count both match, the core is probably sound. If either is off, work through fills, fees, scaled exits and open positions.
Should a scale-out be one trade or two in my journal?
One trade, with a weighted result. Logging each exit as its own row inflates trade count and win rate, because you scale out of winners far more often than losers.
Why do my journal's session stats look wrong?
Usually a timezone mismatch between broker server time, stored UTC and displayed local time, made worse by daylight saving. Check a handful of timestamps against trades you clearly remember.
Start with checks one and two. An hour with a broker statement and a calculator tells you more about whether your trading journal deserves your trust than any feature list.