Confluences

How Many Tags in a Trading Journal? Four Is the Wrong Fix

The standard advice on how many tags a trading journal should have is to cap yourself at about four, because combinations fragment the sample. That diagnoses the problem correctly, then solves it by collecting less information.

E

EdgeFlow

·Updated

Ask how many tags a trading journal should have and you'll get the same answer nearly everywhere. Keep your trading journal tags down. Three, maybe four. Any more and your sample fragments into combinations too small to mean anything, so you end up drawing conclusions from two trades and calling it research.

Quick vocabulary, because I want to be precise about what we're arguing over. A tag is any label you attach to a trade: session, higher-timeframe bias, entry model, whether you moved to break-even. A combination is what you get when you look at several tags at once, like "London + aligned bias + entered after confirmation." Your sample is just how many trades sit inside whatever you're looking at.

Here's my problem with the four-tag rule. The maths behind it is correct. I'm not going to pretend otherwise. But the fix doesn't do what people think it does, and the cost of it lands somewhere you won't notice for about a year.

First, the part they get right

Let's actually run the numbers, because the tag fragmentation argument deserves respect.

Say you tag four things:

  • higher-timeframe bias: aligned or not aligned (2 options)
  • session: Asia, London, New York (3 options)
  • entry: straight off the level, or after a confirmation (2 options)
  • volatility: expanding or contracting (2 options)

Multiply those out. 2 × 3 × 2 × 2 = 24 possible combinations. You have 50 trades logged. That's an average of just over two trades per combination.

And "average" is being generous, because trades don't spread themselves evenly. If you mostly trade London, maybe nine of your combinations have zero trades in them, three have eight or nine trades, and the rest have one or two. So you click through your journal, land on "London + aligned + confirmation + expanding volatility," see a 100% win rate and +2.4R average, and feel something.

R is just trade size measured in units of risk. Risk $150, make $300, that's +2R. Expectancy is your average R per trade across a group. And a 100% win rate over three trades is not an edge, it's a coincidence with good presentation. Anyone who has watched a filter chain collapse in real time knows the feeling, and I've written about how fast that collapse happens when you stack filters.

So yes. Four tags over 50 trades really does fragment. The people warning you about this are not making it up.

Now the part that doesn't hold

Here's what happens when you follow the advice and drop volatility to get back to three tags.

You now have 12 combinations instead of 24. Average trades per combination goes from 2.08 to 4.17. The number doubled. Feels like progress.

You collected exactly zero additional trades.

That's the whole thing, and I want to sit on it for a second because it took me longer than I'd like to admit to see it. Dropping a tag does not add evidence. It merges buckets. Your 50 trades are still 50 trades. All you did was agree to stop distinguishing between two things that were already different.

And they were different. Volatility didn't stop affecting your results because you stopped writing it down. If expanding-volatility mornings genuinely behave differently from dead, contracting ones, that effect is still in your data. It's just been folded into the average, where it now shows up as noise you can't name. Your "London + aligned + confirmation" bucket looks cleaner and is quietly mixing two regimes.

You didn't solve the small-sample problem. You hid it.

The cost you don't feel for a year

Small sample fixes itself. Slowly, annoyingly, but it does. Trade for another eight months and those 50 trades become 250, and combinations that had two trades now have ten or twelve. Time is on your side there.

Missing information never fixes itself.

Picture it concretely. It's February. You capped yourself at four tags like a sensible person, so you never logged whether there was high-impact news within the hour of your entry. Fine. Now it's October, you have 280 trades, and you notice something. Your worst clusters seem to sit around economic releases. You want to check.

You can't. Not properly. You can go back through screenshots for the ones you saved, guess at the rest, and reconstruct something half-accurate that's contaminated by knowing how each trade turned out. Which is the one thing a filter must never be. Grade a trade after you know the result and the grade stops being a condition and starts being an outcome in disguise.

That's the asymmetry the four-tag rule ignores. A statistical problem is temporary. An information problem is permanent. Capping your tags trades the first for the second, and the second is the one you can never undo.

Most of us have run into this from the other direction, too: you finally think of the right question to ask your journal, and the journal has no idea what you're talking about. That's a bad afternoon.

Collecting and querying are two different jobs

The confusion at the heart of this is that people treat "how much I record" and "how much I slice" as the same dial. They aren't.

Recording is cheap and permanent. Two extra dropdowns at trade close costs you eight seconds and buys you an option you can exercise any time in the next five years.

Querying is where the danger lives. Every extra condition you stack on a filter cuts your sample, and if you keep clicking until something looks good, you will find something good. You'd find it in coin flips. Randomness clusters, and a filter bar with enough options is a search engine for flattering subsets. That's a real trap and I've gone through the mechanics of it in testing condition combinations.

But notice where the trap actually is. It's in the asking, not the storing. The four-tag rule tries to fix a bad querying habit by amputating your data collection. That's like fixing overtrading by closing your account.

Better rule, and it costs nothing: collect wide, query narrow.

Tag everything you can capture objectively and quickly. Then, when you sit down to analyse, test one thing at a time against your full baseline. Not "London + aligned + confirmation + expanding" against nothing. Just "aligned bias" against all 280 trades, then "confirmation entries" against all 280 trades. One condition, full sample, readable answer. Figuring out which single condition is carrying your edge is a far better use of a Sunday than hunting for the perfect stack of five.

You'll get to combinations eventually. You should. But you get there with a hypothesis you formed first, not by clicking until the number turns green. The evidence on whether stacking conditions genuinely helps is messier than most traders expect, which I've looked at separately in do confluences actually increase win rate.

Make "not enough sample" an actual answer

Here's the piece that makes wide tagging safe, and honestly it's the whole trick.

Most journals give you two answers: a number, or an empty screen. Three trades produce a number. Two hundred trades produce a number. Same font, same box, same confidence. Nothing on screen tells you which one you're allowed to believe.

You need a third answer. Not enough sample to say.

Not a warning label tucked underneath the result. The result itself. If a combination has six trades, the correct output is not "+1.9R," it's "we don't know, come back later." Because "+1.9R over six trades" is a fact about six specific mornings and nothing more. It describes them perfectly. It predicts nothing.

Watch how little it takes to move. Six trades at +1.9R average is +11.4R in total. Turn one of those winners into a −1R stop, the kind that misses by four ticks, and the same bucket now reads +1.4R. Turn two and you're at +0.9R, below your overall baseline. Nothing about your trading changed. Only which side of the coin landed.

Once that verdict exists, the whole argument for capping tags falls apart. The reason people cap is fear of being fooled by thin combinations. But if thin combinations announce themselves as thin, you're not fooled, you're informed. You keep the detail and lose the false confidence, instead of losing both.

This is the design decision behind how EdgeFlow handles it. Every trade gets tagged across four layers, technical setup, execution, environment and management, so the detail is there whether or not you need it yet. Combinations without the evidence to support a claim come back as exactly that, and the ranking leans on the cautious lower bound rather than the flattering headline.

How many tags should a trading journal actually have?

Honest answer: as many as you can log in under thirty seconds without starting to hate the process. That's the real ceiling, and it's usually somewhere between eight and twelve fields, not four.

I'm not telling you to log 40 fields. There's a real cost here and it isn't statistical, it's human: if closing a trade takes four minutes, you'll stop doing it by March. Half the reason the four-tag advice exists is that people quit journals.

So the constraint you should optimise is seconds per trade, not number of tags. Those are different constraints and they suggest different fixes.

Things worth tagging tend to share three traits. They're objective, so the clock or the chart decides, not your mood. They're recordable at entry, before you know the outcome. And they're impossible to reconstruct later, which is the strongest argument of all. Session times you could rebuild from timestamps. Whether you hesitated for ninety seconds before clicking, you could not.

If you want the longer version of that list, it's in what to track in a trading journal, which covers the layers in more detail than I'll repeat here.

One honest caveat. I can't promise wide tagging makes you profitable. Nobody can. More tags means more ways to fool yourself if you go hunting through them without discipline, and that risk is genuine. What I'm confident about is narrower: the information you don't collect today is information you cannot analyse in a year, and no amount of patience gets it back. That trade is permanent, and it's being recommended to beginners as prudence.

Cap the conclusions you draw. Don't cap the data you keep.

If you want to see how layered tagging works when the tool is willing to say "not enough data yet," the edge building framework walks through it, and there's a free interactive demo on the homepage you can poke at without signing up.

Continue reading

Related articles