Trading Journal

Journaling a Prop Evaluation: The Math the Rules Impose

An evaluation is a small-sample test with a hard stop attached. Daily loss limits truncate your distribution and 30 trades cannot separate a real edge from a good week. What to record so the funded account survives what comes next.

E

EdgeFlow

Most people journal an evaluation the same way they journal anything else. Date, pair, entry, exit, result, a line about how it felt. Then they pass, get the funded account, and blow it inside six weeks while staring at a spreadsheet that said they were profitable.

The spreadsheet wasn't lying. It was just measuring something that no longer exists.

An evaluation is not a normal trading period. It's a short test with a hard stop bolted onto it, and both of those things bend your numbers in ways that are easy to miss. If you want your journal to still be useful on the other side of the pass, you have to record the rules alongside the trades.

Two things a challenge does to your data

Let's name them plainly, because the rest of this post depends on both.

One: the sample is tiny. Evaluations don't take many trades. Take 30 as a working number for the rest of this post — if your own number is lower, the argument only gets worse. Either way it's nowhere near enough to tell a real edge from a good stretch.

Two: the loss limits chop off your worst outcomes. A daily loss limit and a max drawdown mean your biggest disasters can't be recorded as trades. Either you stop before they happen, or the account dies and the data disappears with it. Statisticians call this a truncated distribution. Traders usually call it "my stats look great."

Neither of these is a reason to avoid prop firms. They're a reason to distrust the numbers a passed eval produces.

The 30-trade problem, with actual arithmetic

Say your strategy risks 1% per trade, targets 2R, and wins 45% of the time. R just means "one unit of risk" — if you risk $500 and make $1,000, that's +2R. Expectancy, the average result per trade, works out like this:

(0.45 × 2R) − (0.55 × 1R) = +0.35R per trade

Over 30 trades that's an expected +10.5R, which clears most 8–10% profit targets comfortably. Great. That's the strategy actually working.

Now run the same math for someone with no edge at all. To be exactly break-even at 2R you'd need a win rate of about 33.3%:

(0.333 × 2R) − (0.667 × 1R) = 0R per trade

Zero edge. Long-run flat. But 30 trades is not the long run. The spread of outcomes around that zero is roughly 1.4R per trade, and randomness accumulates with the square root of the sample, so across 30 trades the typical swing is about 1.4 × √30 ≈ 7.7R in either direction.

An 8R profit target is barely one standard deviation away from zero. Ballpark, that's something like a one-in-seven shot for a trader with literally no edge — before we even count the ones who get there by taking three lucky 4R runners.

That figure ignores the max drawdown rule, which knocks out a real share of those runs before trade 30 ever arrives — a no-edge trader risking 1% has a decent chance of touching a 5–6% drawdown somewhere along the way and never finishing the test. So treat one-in-seven as an order of magnitude, not a pass rate. The point isn't the exact number. It's that it's nowhere near zero.

It also explains something the forums argue about constantly: why so many people pass on the second or third attempt. Buying more attempts is buying more draws from the same distribution. Eventually one of them comes up.

None of this means you got lucky. It means a pass, on its own, doesn't distinguish between the two cases. If you want the real version of that argument with the sample-size math worked out properly, I wrote it up in how many trades you need to validate a setup, and the companion piece on telling a real edge from a lucky streak covers what evidence would actually settle it.

The daily loss limit is quietly editing your results

This is the part almost nobody journals, and it's the one that bites later.

Imagine a 5% daily loss limit while you risk 1% a trade. In theory that's five losses. In practice, most of us stop after two or three, because the fourth one puts the whole account inside a bad candle's reach. So your worst recorded day is roughly −3R. Not because your strategy has a −3R floor, but because the rule built one.

On a live account with your own money, no such floor exists. The −6R day is real. It happens when you're stubborn, or when a news print gaps through three stops, or when you size up after a good week. Those days are in your true distribution. They were never in your eval distribution.

So the average trade in your eval log is measured against a left tail that got surgically removed. And the traders you're comparing yourself to on Discord are the survivors — anyone whose tail actually showed up got liquidated and isn't posting their stats.

Two separate biases, same direction. Both make the eval look better than the process is.

The practical consequence: your evaluation expectancy is not comparable to your live expectancy, and never will be. They're measurements of different games. If you don't know what expectancy actually is or how it's calculated, start here before going further — the whole argument rests on it.

The third thing: you don't trade the same way

Beyond the math, there's the behavior. During an eval most people:

  • skip the high-impact news they'd normally trade
  • take partials way earlier than their plan says
  • stop trading entirely once they're 1% from target
  • size down after any losing day
  • avoid their B-setups completely

Then the account gets funded, the fear drops, and they go back to trading their actual strategy. The one that was never tested.

I'd argue that's a large part of why funded accounts die early, and it isn't a psychology failure, or at least not only. Follow the logic: the tested process and the deployed process were different processes, and nobody wrote down where they diverged. You can't debug a strategy when the version you measured isn't the version you're running.

What to actually record

Everything you'd normally track still applies — entry, stop, target, R, setup tags, screenshots. The pillar guide on what a trading journal is for covers the base layer. On top of that, an eval needs four extra fields, and they take about ten seconds each.

1. Remaining daily budget at entry. How much of your daily loss limit was still available when you clicked. "Full", "half", "last trade of the day" is granular enough. This one field is what lets you separate free trades from constrained ones later.

2. Rule-forced exit, yes or no. Did you close because the plan said so, or because the drawdown was getting close? Tag it. These trades are not part of your strategy's performance. They're part of the rules' performance.

3. Rule-forced skip. The trades you didn't take because of the eval, logged as missed opportunities with the setup tags you would have used. Painful, and easily the most valuable column you'll have after the pass. It tells you what your strategy looks like unclipped.

4. The behavior delta. One short line: what did you do differently today because you were on a challenge? "Cut at 1.2R instead of running to structure." That's it.

You want these tagged in a structure you can filter, not buried in a notes field. In EdgeFlow this falls out naturally because every trade gets tagged across four layers — technical setup, execution, environment, and management — and "closed early due to drawdown proximity" is a management fact, not a setup fact. Keeping it in the right layer is what stops it from poisoning your read on the setup itself. It's the same reason the Tradezella comparison comes down to a single question: plenty of tools will show you your stats, but the question during an eval isn't what your win rate was, it's whether that number means anything yet.

After you pass, treat the eval as discovery

Here's the one opinion I'll actually plant a flag on.

A passed evaluation is a hypothesis, not a validation. The rules that made it a good test of discipline made it a bad test of expectancy. So don't size up on the strength of it.

What I'd do instead: keep the funded account at eval-level risk for another 40–60 trades, journaled the same way, and treat those as your out-of-sample check. Same setups, same tags, no rule-forced exits this time. Then compare. If the strategy holds up without the training wheels, you have two independent samples pointing the same direction, which is worth vastly more than one sample of 60. That's the whole logic behind forward-validating an edge rather than re-reading the same data until it says yes.

And be honest about the confidence band. With 30 trades at 45% wins, the true win rate could plausibly sit anywhere from the high twenties to the low sixties. That range spans "dead strategy" to "excellent strategy," which is exactly the problem.

So pick an end of that range to plan from. The bottom end is the one that matters when you're deciding how hard to push a fresh funded account, because sizing up is a bet on the top end being true, and you have no evidence for that yet. That's the number EdgeFlow puts in front of you instead of the flattering point estimate — the cautious lower bound of the range, per setup, so the setup that survives scaling is the one you lean on and the one that only looks good at its optimistic end waits for more trades. Sometimes there's no honest lower bound to show yet. That's worth knowing too, and it's the answer more often than anyone wants.

Nobody can promise you the edge is real. The point is to know how much you don't know before the position size makes that question expensive.

One last thing. If you've been logging trades through evaluations for a year and nothing has improved, the problem probably isn't the logging — it's that the log has no question attached to it. That's a separate failure, and I've written about journaling without improving elsewhere.

Pass the eval. Just don't let the pass tell you more than it knows.

There's a free interactive demo on the EdgeFlow homepage that walks through the same logic on a sample journal, if you want to see the shape of it before touching your own data.

Continue reading

Related articles