Monte Carlo Simulation for a Trading Strategy

A single backtest equity curve is one path out of thousands the same trades could have taken; Monte Carlo simulation is what tells you how much of that curve was luck.


The equity curve you see is one draw, not the answer

A backtest produces one equity curve, built from one specific sequence of trades in one specific order. That order feels like a fact, because it’s literally the order the trades occurred in historically. But the sequence in which wins and losses arrive is itself a random variable, separate from whether the strategy has a real edge at all. Ten winning trades followed by four losers produces a very different equity curve, on paper, than four losers followed by ten winners — same trades, same win rate, wildly different drawdown experience along the way. A single backtest shows you exactly one of those orderings and nothing about how representative it is of the range of things that could plausibly have happened.

This matters most for anything downstream of the equity curve that depends on path, not just on endpoint — maximum drawdown, the longest losing streak, the capital required to survive to the payoff. Two strategies with an identical final return and an identical 56% win rate can carry very different real-world risk if one of them tends to cluster its losses and the other doesn’t, and a single historical sequence won’t reliably show you which is which.

What the simulation actually does

Monte Carlo simulation takes the population of individual trade outcomes your backtest or walk-forward process already generated — the actual distribution of wins, losses, and their sizes — and reshuffles them into thousands of alternate sequences, typically through resampling with replacement. Each reshuffled sequence produces its own equity curve, and across a few thousand of them you get a distribution rather than a single line: a range of plausible maximum drawdowns, a range of plausible final balances, a range of losing-streak lengths. Instead of asking “what happened,” you’re asking “what’s the reasonable range of things that could have happened, given the actual statistical properties of this strategy’s trades.”

The output that matters most in practice is usually the drawdown percentile table, because that’s the number that determines whether you can survive holding the strategy live long enough for its edge to play out. A backtest might show a maximum historical drawdown of 8%. The Monte Carlo distribution built from the same trades might show that a 15% drawdown sits at the 90th percentile — meaning under a plausible reshuffling of the exact same wins and losses, one in ten paths would have handed you nearly double the drawdown the single historical sequence happened to show. If your position sizing was built around that 8% figure, you were sizing for a specific lucky ordering, not for the strategy’s actual risk profile.

Where the 52–62% win rate band matters here specifically

This is also where an inflated backtest win rate does its worst damage. A strategy showing a 78% win rate off a curve-fit optimization will, when resampled, still show unrealistic-looking equity paths, because the resampling is only as honest as the trade population it’s built from — garbage in, garbage out. But a strategy validated into the more believable 52–62% range gives you a resampled distribution that’s actually informative, because you’re not laundering an already-fictional result through a statistical process and coming out the other side with false confidence. Monte Carlo simulation is a tool for understanding the variance around a real edge. It does nothing to detect or correct an edge that was never real to begin with — that’s the walk-forward process’s job, and it has to happen first.

Building the cost drag into the resampling, not just the average

A common shortcut that quietly undermines this whole process is resampling trade P&L figures that were computed with average, flat assumptions for spread and slippage rather than the actual cost incurred per trade. Spread and slippage aren’t constant — they widen during news, they’re wider going into weekend close, they behave differently during the Asian session’s thinner liquidity than during the London/New York overlap. If your original backtest recorded each trade’s actual realized cost rather than a blanket average, your resampled distribution reflects genuine variance in the cost drag too, not just in the raw price movement. This is a small implementation detail that has an outsized effect on how trustworthy the drawdown percentiles end up being.

Plain resampling has a blind spot around regime

The standard approach — resampling individual trades independently, with replacement — assumes each trade’s outcome is unrelated to the ones around it. That’s a convenient assumption and often close enough to reasonable, but it quietly erases any serial structure in the data, including the structure that shows up when a pattern’s edge is decaying. If your trades were pulled from a period spanning a genuine regime shift — the pattern’s underlying market structure weakening partway through the sample — plain resampling will happily mix early, healthy-edge trades with late, decaying-edge trades into the same simulated sequence, producing paths that never actually correspond to a real market condition. Block bootstrap resampling, which resamples contiguous chunks of consecutive trades rather than individual ones, preserves some of that time-local structure and tends to give a more honest picture when you suspect the sample straddles a regime boundary. It’s slower to set up and rarely the default in off-the-shelf tools, but it’s worth the extra step whenever the trade population being resampled spans more than a year or two, simply because the odds of crossing a genuine regime shift somewhere in that window aren’t small.

Reading the tails without flinching

The temptation, once the simulation is run, is to look at the median outcome and stop there, because the median is usually reassuring. The more useful habit is to sit with the 90th and 95th percentile drawdowns specifically, because those are the paths that will actually test whether you keep running the system. If the 95th percentile drawdown is deep enough that you know, honestly, you’d intervene and shut the strategy off before it got there, you’ve learned something the single backtest curve could never have told you: your position sizing is miscalibrated against your own risk tolerance, even though the strategy itself might be sound.

Percentile Drawdown (illustrative) What it represents
50th 6% Typical path, roughly what the single backtest resembled
75th 10% A meaningfully rougher stretch, still plausible
90th 15% One in ten resampled paths goes here
95th 19% The path you need to be honestly prepared for before going live

Why this changes the intervention conversation, not just the sizing

The deeper value of running this analysis before going live, rather than after a rough live stretch prompts you to run it defensively, is psychological. A trader who has already seen the 90th percentile drawdown on paper, before it happens, experiences that drawdown live as “this is the scenario I already accounted for” rather than “something has gone wrong and I need to intervene.” That distinction is the difference between holding a validated system through a statistically normal rough patch and abandoning it at exactly the point where abandoning it destroys the expected value the whole validation process was built to capture. Monte Carlo simulation doesn’t make the drawdown smaller. It makes the drawdown expected, and expected pain is a fundamentally different experience to sit through than pain that arrives as a surprise.