I Backtested the Same Candle Event Long and Short. Both Came Back Profitable.
The same event backtested profitably in both directions, and the reason had nothing to do with the pattern itself.
The dashboard doesn’t judge. It just paints the row green or it doesn’t, and on this run, near the top of the results, the same event name showed up twice. Event_Bull_FVG [Short]. Event_Bull_FVG [Long]. Both checked. Both PROFIT.
A bullish fair value gap is, by definition, a setup you buy. Three candles, a gap between the first and third that price hasn’t filled, read by nearly every trader who uses it as a sign of imbalance to the upside. Seeing it show up profitable as a long wasn’t news. Seeing the exact same event, same underlying rows in the dataframe, come back green as a short too, was the kind of thing that makes you stop scrolling.
The first read, and why it’s wrong
The easy explanation is that the pattern just works both ways, some kind of volatility signature that pays off regardless of direction. It’s a tempting story because it lets you keep the pattern. It’s also checkable, and it doesn’t survive the check.
Here’s the math. Under a flat 10-point TP and 10-point SL, if the same set of bars gets traded both long and short, the two outcomes are locked together. Price either tags +10 first or -10 first. If it’s +10 first, the long wins and the short’s stop gets hit at the exact same bar. If it’s -10 first, the reverse. There’s no scenario, on the same entries, where both sides win independently of each other. Which means if Event_Bull_FVG [Long] and Event_Bull_FVG [Short] were trading the same 35 or so occurrences, their win rates would have to add up to something close to 100%.
They didn’t. 60.0% short plus 54.5% long is 114.5%. Not close to 100, over it. That’s not noise rounding, that’s the two numbers describing two different things.
Checking the trade counts instead of the win rates
The tell was sitting in a column I’d been ignoring. Event_Bull_FVG [Short] had 35 completed trades. Event_Bull_FVG [Long] had 44. If both were drawn from the same bars, traded in opposite directions, those counts would match exactly, every single time, because every occurrence produces one resolved outcome on each side. They didn’t match, not for this pattern, and not for Event_Price_Touch_Upper either, where the split was even more obvious: 37 trades short, 99 trades long.
Different counts means different populations of bars. The event wasn’t being traded both ways on the same occurrences. It was firing on two disjoint sets of candles, and something upstream of the pattern itself was deciding which occurrences went into the long bucket and which went into the short bucket.
The row that gave it away
The pattern list had another entry sitting a few rows below Event_Bull_FVG [Long]: Event_Bull_FVG & Event_Trend_Up [Long], a combined pattern requiring both the FVG and an explicit trend-up condition at once. I expected it to narrow the trade count, maybe tighten the win rate, the way adding a second condition normally does.
| Pattern | Trades | WR% | Net PnL | Avg PnL |
|---|---|---|---|---|
Event_Bull_FVG [Long] |
44 | 54.5 | 18.7 | 0.42 |
Event_Bull_FVG & Event_Trend_Up [Long] |
44 | 54.5 | 18.7 | 0.42 |
Identical. Same trade count, same win rate, same net PnL down to the decimal. Adding “and the trend is up” to the condition changed nothing, because every long trade in this dataset already required the trend to be up. The trend filter was enforcing that condition on every long entry before the pattern mask ever got evaluated. The combined pattern wasn’t adding information. It was restating a rule that had already been applied.
That’s what split the bars into two populations. Event_Bull_FVG [Long] only ever fires on occurrences where price was already above the 200-period MA. Event_Bull_FVG [Short] only ever fires where it was below. Two different market regimes, two different sets of candles, wearing the same event name because the event itself, a three-candle imbalance, doesn’t know or care what side of the moving average it happened on.
What the pattern actually turned out to be measuring
Once that’s clear, the “bullish gap that also works as a short” stops being strange. It was never one bidirectional signal. It was two separate, regime-conditioned setups that happened to share a label: a continuation trade when the gap shows up above the 200 MA, and something closer to a fade when it shows up below it, in a context where the broader move is already down and a short-lived upward gap gets absorbed. Both plausible. Neither has much to do with what a fair value gap is supposed to mean on its own.
Which raises the less comfortable question. If the trend filter is already doing the real work of separating these two regimes, and the FVG condition rides along without changing the outcome inside either regime, how much of the 54.5% and 60.0% win rates actually belongs to the candle pattern at all, versus just belonging to the fact that a 200-period moving average is a coarse, well-worn way of saying “this is currently an uptrend” or “this is currently a downtrend,” full stop, no candle shape required. A flat dashboard row doesn’t show you that distinction. It just shows you a name, a direction, and a green checkmark, and lets you assume the name is where the edge came from.
Why this matters more than it looks like it should
This is the same trap as the multiple-comparisons problem from a different angle. It’s not about how many patterns you tested. It’s about whether the pattern you kept is actually measuring what its name claims, or whether a blunter filter sitting upstream already did the separating and the named event is just along for the ride. The dashboard can’t tell you which. It doesn’t have an opinion on semantics, it only has PnL and win rate, and those numbers look identical whether the FVG mattered or whether it was decorative.
The practical fix isn’t complicated once you know to look for it. Any time a pattern shows up profitable in both directions, check the trade counts before you check the win rates. If they match, the two rows are mirror images of the same math and one of them is lying by omission. If they don’t match, go find out which upstream filter split the population, and then test the pattern with that filter removed, on its own, to see if anything real is left. Most of the time, like here, there won’t be much. The moving average was doing the work. The candle just happened to be standing nearby when the checkmark turned green.