What Happens When You Backtest Every Possible Indicator Combination at Once

Running every pairwise combination of your indicator signals doesn't discover an edge, it guarantees a predictable number of accidents that look like one.


Twenty-four event columns. Two hundred seventy-six pairs. Three hundred patterns tested against the same block of XAUUSD M15 candles before you’ve finished your coffee.

That’s not a hypothetical number. That’s what falls out the moment you take a dozen base signals, add a one-bar lag to each, and feed the whole set into itertools.combinations with a max group size of two. The function doing the work here, generate_combinations(), doesn’t know the difference between a real setup and a coincidence. It doesn’t know anything. It just multiplies.

Layer Count
Base events (trend, RSI, BB, pivot, FVG) 12
With one-bar lag added 24
Single-event patterns 24
Two-event combinations (C(24,2)) 276
Total patterns tested 300

Most people who build a miner like this don’t think of it as running three hundred experiments. They think of it as “testing a strategy.” That distinction is where things go wrong, quietly, long before a single live trade gets placed.

The number nobody counts

If you ran one backtest on one set of rules and it cleared your win rate bar, you’d feel reasonably confident you’d found something. One test, one result, low odds it’s an accident.

Now run three hundred tests on the same data. Even if every single pattern had exactly zero real edge, pure variance guarantees some fraction of them will clear your threshold anyway. Not because the market rewarded a genuine insight. Because noise, given enough tries, produces winners on its own. The loop in this script that sorts res_train by net PnL and keeps the top 30 has no way to tell the difference between a pattern that works and a pattern that got lucky on this particular eleven months of gold.

This isn’t a defect specific to this codebase. It’s what happens any time strategy discovery turns into a search problem instead of a hypothesis you’re actually testing. The search doesn’t know what a real edge feels like. It only knows how to rank whatever came out the other side.

What “testing” actually means once you’re searching

Mechanically, here’s what’s happening. pat_masks builds a boolean array for each of the three hundred patterns, every one just a row-wise AND across a handful of event columns like Event_Trend_Up, Event_RSI_Cross50_Down, Event_Bear_FVG. The simulator walks the array bar by bar, checks whether the mask fired on the previous candle, applies the trend filter to decide whether a long, a short, or neither is even allowed, and if a trade opens, tracks it forward up to 120 bars against a flat 10-point TP and a flat 10-point SL. It’s fast, it’s vectorized where it can be, and it is completely blind to the fact that pattern #217 and pattern #218 might share three of their four underlying conditions.

That correlation matters more than the raw count of three hundred does. Event_RSI_OB and Event_RSI_OB_lag1 aren’t independent signals. They’re the same signal one bar apart, still overlapping most of the time the RSI stays extended. Any two-event combination built using both inherits that redundancy instead of adding real informational diversity. So the effective number of genuinely distinct hypotheses being tested here is lower than three hundred. It’s just not low enough to ignore, and worse, you don’t actually know what it is. Somewhere between one and three hundred is not a number you can correct for cleanly, and that ambiguity is exactly what lets people trust a top-30 list more than they should.

Where the split earns its keep, and where it runs out of room

This is precisely the situation the train/test split exists to catch, and it’s still the only validation step in this whole pipeline that actually means something. Everything upstream of it, the win rate filter, the PnL sort, the MIN_COMPLETED_TRADES threshold, is triage performed on data the search has already seen and already optimized against. The test set is the first point where a pattern has to prove itself on bars it never touched during selection.

But a single 70/30 split wasn’t built to absorb the leakage from a three-hundred-pattern sweep. It was built to check one strategy, or a small number of deliberately chosen candidates, against unseen data. Point it at the top 30 survivors of a combinatorial search instead, and you’re asking one holdout window to filter out a wave of noise-driven patterns that were specifically selected, by construction, for looking good on the training set. Some of that top 30 will be real. Some of it will only exist because of how price happened to move through 2024 and into 2025 on this particular instrument, during this particular session window.

The test set kills most of the fakes. It doesn’t kill all of them, and it can’t tell you which of the survivors are genuine edges versus which ones got lucky in two separate, non-overlapping windows in a row. That’s a rarer event than getting lucky once. But with three hundred starting candidates, rare events still show up on the other side.

The fix isn’t fewer combinations

The obvious instinct is to shrink the search space, drop MAX_COMBINED_EVENTS back to one, hand-pick a handful of setups you already have conviction in. That’s not wrong, but it treats the symptom rather than the cause. The real fix is adjusting what you’re willing to call a discovery, given how hard you had to search to find it.

A pattern clearing your win rate bar out of three hundred candidates needs to clear a much higher bar than a pattern you tested in isolation, because you already know some fraction of those three hundred were going to look good regardless of whether the market gave you anything real. Treat the training sort as a shortlist generator, not a results page. Treat the test set as the actual experiment, not a formality on the way to a config file. And treat anything that only survives one round of that gauntlet as unproven, no matter how clean the PnL column looks when you print it.

The real question this kind of miner never answers on its own is how many of those three hundred patterns would have cleared the bar even in a market with no memory at all, no structure, nothing but noise. That number is bigger than most people expect. It’s worth sitting with before you trust what the top of that sorted list is telling you.