Market Environment

Why the Setup That Looked Perfect Last Month Stopped Working

Three explanations compete: the edge was never real, the market environment changed, or your execution drifted. Each demands the opposite response, and most traders reach for the wrong one. How to separate them using data you already have.

E

EdgeFlow

Most of us have a version of this story.

The setup worked. You have the screenshots. A few months where the entries did roughly what you expected, where the curve went up in a way that felt earned rather than lucky, where you finally stopped renegotiating your rules every Sunday night.

Then it stopped.

Usually not with a bang. No account-ending trade. Just a long grinding stretch where the same entries keep failing, and every single loss has a perfectly reasonable explanation once you look at the chart afterwards.

Search for help and you get two answers. One camp says you lost discipline and should journal your feelings. The other says all edges decay, that's just markets. Either can be true. Neither is a diagnosis.

There are three real candidates. The annoying part is that they ask you to do opposite things.

Three suspects, three opposite fixes

One: it was never as good as the number said. Small sample, a couple of enormous winners doing the heavy lifting, or the plain fact that you found the rule and measured the rule on the same pile of trades.

Two: the market environment changed. The setup is fine. The conditions it needs stopped turning up, or turned up wearing a different outfit.

Three: your execution drifted. Same setup, same market, different you. Entries a fraction later. Stops a fraction tighter. Winners cut a fraction sooner.

Now look at what each one wants from you.

If it was never real, you cut size or shelve it until it earns its way back. If the environment changed, you keep every rule exactly as written and get stricter about when you're allowed to press the button. If execution drifted, you don't touch the strategy at all, because the rules are innocent and the problem is you.

Those aren't three flavours of the same fix. Rebuild a setup that was fine and just under-supplied and you've broken something that worked. Grind on discipline for a setup that never had an edge and you lose money slowly with beautiful form. You can execute a dead strategy flawlessly and still go broke.

That's why guessing here is expensive. The wrong answer doesn't just fail to help. It points you the opposite way from the fix.

One set of numbers, carried all the way through

Made up, but the shape is one you'll recognise. Imagine 220 logged trades on a single setup, say a pullback into a level after a break of structure. Everything in R, where 1R is what you risk on a trade.

  • First 165 trades: +0.31R per trade. 41% win rate, average winner +2.2R, average loser −1.0R.
  • Last 55 trades: −0.01R per trade. 31% win rate, winners still around +2.2R, losers still around −1.0R.

Notice what didn't change. The payoff is the same. You're still getting 2.2 out of your winners and losing your full stop when you're wrong. You just stopped hitting.

Fifty-five trades of nothing. Enough to hurt, nowhere near enough to be obvious. Let's run the three tests on it.

Test one: was it ever as big as you recorded?

Do this one first. Almost everybody does it last, because it stings and because blaming the market beats auditing your own spreadsheet. That ordering is expensive, and it's the strongest opinion in this post: check your own number before you go looking for what the market did to you.

Three checks.

Pull the top winners out. At +0.31R across 165 trades you banked roughly 51R. If your three best trades were +9.8R, +7.5R and +6.2R, that's 23.5R from three fills. The other 162 trades produced about 28R, which is +0.17R each. Still positive, still a business. But your recent 55 trades aren't a break from a +0.31R edge. They're an ordinary bad run inside a +0.17R edge, and from the inside those feel identical while meaning very different things.

Count how many ideas you tried. Did you test this setup once, or did you work through eleven filter combinations on a wet Sunday and keep the one that looked best? If it's the second, the headline number was always flattering. Sifting a lot of slices and reporting only the winner is the same mistake backtesters get lectured about, done by hand, on a smaller sample, with no record of how many slices you tried.

Ask whether the rule ever met data it hadn't already seen. If you found the conditions by staring at those 165 trades and then measured them on the same 165 trades, you don't have a result. You have a description of the past. That gap is why people bother with an in-sample and out-of-sample split, and a date filter in a spreadsheet is enough to do it.

Two of those three coming back badly means stop. You've found your answer, and the other two tests are just a comfortable way to avoid it.

Test two: did the environment genuinely change?

This is the one that gets abused as an excuse, which is a shame, because it's the most measurable of the three.

The trick is small. Don't compare recent performance to old performance. Compare recent performance inside one environment to old performance inside that same environment. Hold the condition still and let time move.

Back to the numbers. Split both periods by whatever environment tag you keep. Let's say trending days versus chop.

First 165 trades:

  • Trending: 88 trades, +0.55R each
  • Chop: 77 trades, +0.03R each

Last 55 trades:

  • Trending: 13 trades, +0.50R each
  • Chop: 42 trades, −0.17R each

Read that twice, because it tells a different story than the headline did.

The setup didn't stop working on trending days. It's doing roughly what it always did, +0.50R against +0.55R, and thirteen trades could never distinguish those anyway. What moved is the mix. Trending days were 53% of your old sample. They're 24% of the new one. You didn't lose an edge. You lost the supply of the conditions the edge needed, and you carried on trading regardless.

The fix for that looks nothing like the fix for a broken setup. You keep every rule. You add one gate. You accept a quarter as many trades, which for most of us hurts more than the drawdown did.

Two honest caveats. Thirteen trades confirms nothing; it's consistent with the setup still working, not evidence of it. And the whole comparison collapses if you tagged the environment afterwards. Going back through a year of charts now, labelling days "trending" or "choppy" while you can see how each one resolved, will hand you precisely the answer you were hoping for. Hindsight labelling isn't data.

Which is the practical argument for environment being a field you complete at entry, not a note you add later. In EdgeFlow it's one of four layers tagged on every trade, alongside technical setup, execution and management, so "did the regime change" is a query you run instead of a feeling you have. A spreadsheet column does the same job, as long as it exists before you need it.

Worth remembering too that a condition rarely has one fixed meaning. A sweep — price running the obvious stops before reversing — is strong when the higher timeframe agrees and noise when it doesn't, which is why combinations of conditions usually tell you more than any single tag.

Test three: did you drift?

Environment change and execution drift look identical on an equity curve. Only your own records separate them, which is why this test needs data you might not have been keeping.

Compare the two periods on things that have nothing to do with the market:

  • Distance between planned entry and actual fill. If that crept from 0.2R to 0.6R, you're chasing.
  • Share of trades moved to break-even early. Going from 20% to 55% is a different strategy wearing the old name.
  • Average hold time on winners. Four hours down to ninety minutes means you're harvesting a different distribution.
  • Maximum favourable excursion against what you actually banked. If recent trades ran +1.8R in your favour and you closed them at +0.4R, the market delivered and you handed it back.

That last one is the cleanest signal in the whole diagnosis. When the trades went your way and you didn't collect, the setup isn't the problem. Splitting those two questions properly, is this the idea or is this me, gets its own treatment in the edge or execution problem breakdown.

There's a trap sitting right here. If execution drifted and you respond by rewriting the rules, you've now optimised your strategy around your own sloppiness. That's how traders end up with a setup that only "works" when they're tired.

Sometimes the honest answer is "not enough data"

Run all three tests and you may land on: can't tell yet. That's a legitimate result and you should be willing to stop there.

Fifty-five trades is a narrow window. At a true 41% win rate, the ordinary wobble in a 55-trade sample is around 6.6 percentage points. Your observed 31% sits about one and a half of those below expectation. That is a bad run. It is not, on its own, proof of anything.

Losing streaks make it worse. Seven losers in a row at a 41% win rate is roughly a 1-in-40 event for any given stretch of seven, and there are a lot of stretches of seven inside 165 trades. Expect to meet that streak more than once. It will feel like the edge died every time.

A drawdown by itself is evidence of nothing. The real question isn't "am I losing", it's "is what I'm seeing still plausible if the original model were true?" Uncomfortably often the answer is yes. It's also why a tool that reports an edge as a range rather than one confident number is doing you a favour, and why it should be willing to tell you the sample is too thin to judge. Nobody can promise a clean verdict from 55 trades.

What you can do is refuse to treat an ambiguous result as permission to rebuild everything.

Once you have an answer

Whichever suspect it turns out to be, the change you make now is a hypothesis, not a conclusion. You found it by looking backwards. It has to earn its way forwards.

Write down the change and the date you started applying it. Then judge it only on trades taken after that date, against a target you set in advance. It's dull, and it's the whole difference between adjusting and flailing. Forward validation is what stops you running this same diagnosis again in four months with exactly the same uncertainty.

To do this properly next time, the requirement is unglamorous. Environment recorded at entry. Execution recorded separately from the setup. Enough trades in each bucket for the comparison to mean anything. That's most of what structured edge discovery actually is. Not clever statistics, just fields filled in before you knew you'd need them.

The setup that stopped working usually didn't. Something around it did, and there are only three candidates. Find out which one before you change a single rule.

If your journal can't answer "did the regime change" today, the free demo on our homepage shows what tagging environment as its own layer looks like in practice.

Continue reading

Related articles