How Many Trades Before I Can Trust a Backtest

There's no magic number, but there is a real answer, and it comes from how wide the confidence interval around your win rate still is at the sample size you're looking at.


The question people actually mean

“How many trades before I can trust this” usually isn’t a question about statistics — it’s a question about whether it’s safe to go live. But the honest answer runs through statistics whether you want it to or not, because a win rate computed from a small sample carries a wide margin of error, and that margin doesn’t shrink because you’re impatient to deploy the strategy. A backtest showing a 58% win rate over 25 trades and one showing 58% over 400 trades are reporting the same headline number with two completely different levels of confidence behind it, and treating them the same is how people go live on noise.

Why 30 trades is a myth worth retiring

There’s a persistent number that circulates in trading forums — 30 trades, treated as some kind of statistical threshold for significance. It has a distant, misapplied connection to the central limit theorem, which does involve the number 30 as a rough rule of thumb for when a sample mean starts to approximate a normal distribution. But a win rate isn’t a continuous mean, it’s a proportion, and the relevant statistics for a proportion’s confidence interval depend on the actual observed rate, not on a flat sample-size cutoff. A 55% win rate needs a meaningfully larger sample to pin down precisely than an extreme win rate like 90% or 10% does, because the variance of a binomial proportion is maximized right around 50%, which is exactly the neighborhood most legitimately validated forex strategies live in.

What the confidence interval actually looks like

Using a standard proportion confidence interval calculation, a strategy showing a 56% win rate at different sample sizes looks roughly like this:

Trades Approx. 95% CI on win rate
20 34% – 78%
50 42% – 70%
100 46% – 66%
300 50% – 62%
500 51% – 61%

At 20 trades, a strategy showing 56% could plausibly have a true win rate anywhere from clearly unprofitable to remarkably strong — the backtest has told you almost nothing yet. Somewhere around 300 to 500 trades, the interval narrows enough to actually distinguish a genuinely validated 52–62% strategy from one that’s secretly closer to break-even or secretly overfit. This is why the healthy-looking win rate range that shows up repeatedly in properly validated systems isn’t a coincidence of good strategy design alone — it’s also the range where a reasonably sized sample can actually confirm what it’s showing you, as opposed to extreme win rates that either need enormous samples to trust or are a near-certain sign of curve-fitting regardless of sample size.

Trade count isn’t the same as time elapsed

The other trap is treating elapsed backtest time as a proxy for trade count. A strategy scoped to a specific session window — say, only trading during the London/New York overlap, 13:00–16:00 UTC — might only generate one or two qualifying setups a day even across a full year of data, which is a few hundred trades, sitting right at the edge of a trustworthy sample. A strategy with looser session filters covering the full London session, 08:00–16:00 UTC, might generate several times that in the same calendar period. Two backtests can span the identical two-year date range and produce wildly different levels of statistical confidence purely because of how their start_hour and end_hour fields in the JSON config are set, and it’s easy to see “two years of data” on a report and assume that implies a large sample without actually checking the trade count.

Why this interacts with walk-forward and Monte Carlo, not replaces them

Getting the raw trade count into a trustworthy range is a precondition for the rest of the validation process, not a substitute for it. Four hundred trades sitting inside a single, unbroken backtest window still doesn’t tell you whether the edge would have survived being re-optimized out-of-sample across a walk-forward chain, and it doesn’t tell you anything about the range of drawdowns a different ordering of those same four hundred trades could have produced. What an adequate trade count buys you is the ability to trust the headline win rate and average R-multiple enough to make those next two analyses meaningful. Running walk-forward validation or Monte Carlo resampling on a sample of 40 trades just propagates the same wide uncertainty through more sophisticated-looking tools — the outputs will look more rigorous without actually being more trustworthy.

Win rate isn’t the only number with a confidence interval

It’s easy to fixate on win rate specifically because it’s the number most often quoted, but expectancy — the average result per trade once wins and losses are weighted by their size — carries its own uncertainty, and it’s arguably the more important figure, since a strategy with a lower win rate and a favorable risk-reward ratio can easily out-earn a higher-win-rate strategy with a poor one. Expectancy’s confidence interval is driven by the variance in trade size as well as the variance in outcome, which means a handful of outsized winning or losing trades can swing the estimate more than an equivalent handful would swing a simple win rate. A backtest with a stable win rate across 300 trades but only three or four trades that account for a disproportionate share of total profit is a sign the expectancy figure specifically still needs a larger sample before you’d want to size a live account around it, even if the win rate alone looks settled.

Sitting through the sample-building period without flinching

There’s a psychological cost to all of this that doesn’t show up in any of the tables. Getting from an undersized sample to a trustworthy one, whether through extended backtesting or early live trading at minimal size, takes real calendar time, and during that stretch the strategy will inevitably produce a losing streak that looks, in the moment, indistinguishable from the strategy simply not working. The entire point of having done the confidence-interval math beforehand is to know, going in, that a losing streak of a given length is a statistically unremarkable event at your current sample size rather than evidence the edge was never real. Traders who skip this step and go live on a 40-trade backtest have no such anchor, so the first rough stretch reads as a crisis rather than as expected variance, and the resulting intervention — cutting size, changing parameters, abandoning the system — usually happens well before the sample has grown large enough to actually judge anything.

The uncomfortable implication for a niche pattern

If your strategy is built around a specific, narrow structural pattern — something that only sets up under fairly particular conditions — you may simply not be able to reach a trustworthy sample size within a reasonable backtest window, and that’s a real constraint, not a problem to optimize away. The temptation at that point is to loosen the entry criteria until trade frequency goes up, but that risks diluting the very specificity that made the pattern a real edge in the first place. The more honest path is usually to extend the backtest history further back if reliable tick data exists for it, accept a longer validation timeline before going live, or run the strategy live at minimal size specifically to accumulate real trade count data faster than the backtest alone can provide — treating early live trading as an extension of the sample-gathering process rather than as the finish line the backtest was building toward.