Deflated Sharpe Ratio for a Retail MT5 Backtest
A Sharpe ratio computed from one backtest doesn't tell you if the edge is real — it tells you the result of one trial, and the deflated version is what corrects for how many trials you actually ran to get it.
The number that hides how it was found
A backtest report hands you a Sharpe ratio like it’s a fact about the strategy. It isn’t. It’s a fact about one specific run, produced after however many parameter combinations you tried before landing on the one you’re looking at. Run an optimizer across a hundred combinations of sl, tp, and session hours, and the best Sharpe ratio out of that hundred is not the same statistical object as the Sharpe ratio of a strategy you designed once and tested once. The more trials you run before picking a winner, the more that winner’s reported performance is inflated purely by the act of selecting the best of many, regardless of whether any of them contain a genuine edge. This is the exact mechanism behind why a 70%+ win rate usually turns out to be overfitting rather than skill — the same selection pressure that inflates a win rate inflates a Sharpe ratio, just in a way that’s harder to eyeball because a Sharpe ratio already looks like a rigorous, risk-adjusted number.
The plain Sharpe ratio also carries a second, quieter problem specific to forex trading: it assumes your returns are normally distributed. A pattern-based strategy with a fixed stop and a wider target produces a return distribution that’s skewed and fat-tailed almost by construction — lots of small losses capped by sl, occasionally interrupted by a large winner. Plain Sharpe ratio doesn’t know or care about that shape. Two strategies with an identical mean and standard deviation of returns can have very different true reliability if one has a long left tail of rare, large losses and the other doesn’t, and the raw Sharpe number treats them as identical.
What the probabilistic Sharpe ratio actually asks
The Probabilistic Sharpe Ratio (PSR), developed by Bailey and López de Prado, reframes the question from “what’s the Sharpe ratio” to “what’s the probability that the true Sharpe ratio actually exceeds some benchmark, given the sample size and the shape of the return distribution you observed.” It takes your observed Sharpe ratio, the number of trades or return periods behind it, and the skewness and kurtosis of those returns, and produces a probability rather than a single point estimate. A strategy showing a modest Sharpe ratio over a large, well-behaved sample can return a high PSR — meaning you can be confident that number reflects something real — while a strategy showing an impressive Sharpe ratio over a small or heavily skewed sample can return a low PSR, telling you the impressive number is well within the range noise alone could have produced.
What the deflated version corrects on top of that
The Deflated Sharpe Ratio (DSR) takes this a step further and folds in the selection bias problem directly. It asks the same probabilistic question as the PSR, but against a benchmark that’s been adjusted upward based on how many independent trials were run to find this result — the more parameter combinations your optimizer tried, and the more those trials’ outcomes varied from each other, the higher the bar the observed Sharpe ratio has to clear before it’s considered statistically meaningful rather than the best result of an essentially random search. This is the rigorous version of an intuition most people already have informally: a backtest that was hand-picked out of two hundred optimizer runs needs to look a lot better than a backtest that was the only version you ever tried, to earn the same amount of trust.
Computing it against your own trade log
You don’t need anything beyond a per-trade return series and a record of how many optimization trials produced it. A rough working version in Python looks like this, using scipy for the distribution functions:
import numpy as np
from scipy.stats import skew, kurtosis, norm
def probabilistic_sharpe_ratio(returns, benchmark_sr=0.0):
n = len(returns)
sr = np.mean(returns) / np.std(returns, ddof=1)
g3 = skew(returns)
g4 = kurtosis(returns, fisher=False) # non-excess kurtosis
numerator = (sr - benchmark_sr) * np.sqrt(n - 1)
denominator = np.sqrt(1 - g3 * sr + ((g4 - 1) / 4) * sr**2)
z = numerator / denominator
return norm.cdf(z)
def deflated_sharpe_ratio(returns, n_trials, sr_variance_across_trials):
# Expected max Sharpe under the null across n_trials random trials
euler_mascheroni = 0.5772156649
expected_max_sr = (
np.sqrt(sr_variance_across_trials)
* ((1 - euler_mascheroni) * norm.ppf(1 - 1 / n_trials)
+ euler_mascheroni * norm.ppf(1 - 1 / (n_trials * np.e)))
)
return probabilistic_sharpe_ratio(returns, benchmark_sr=expected_max_sr)
The part that trips people up isn’t the formula, it’s the two inputs: n_trials and the variance of Sharpe ratios across those trials. If you’re running MT5’s built-in Strategy Tester optimizer, this means logging the Sharpe ratio (or a proxy for it, since MT5’s own criterion isn’t always this exact statistic) for every combination the optimizer actually evaluated, not just the top result — the deflation math needs the full distribution of outcomes across trials, not just the winner, to estimate how much of an advantage “best of N” gives you under pure chance.
Reading the output against the 52–62% heuristic
This is where the DSR formalizes something the site has already leaned on informally. A strategy validated into the 52–62% win rate range, tested over a few hundred trades with modest skew, tends to produce a respectably high DSR even though the headline win rate doesn’t look dramatic — because the sample is large enough and the distribution well-behaved enough for the observed Sharpe ratio to actually mean something. A strategy showing a 78% win rate off forty trades and an aggressive optimizer sweep will very often produce a low DSR, sometimes low enough to be statistically indistinguishable from a random trial that got lucky, despite looking far more impressive on the surface. The DSR doesn’t replace the win-rate-and-sample-size intuition already covered elsewhere on this site — it’s the mathematical version of the same check, useful specifically because it gives you a single number to compare across strategies rather than three separate heuristics you have to weigh by feel.
Where it fits in the validation stack
The honest way to use this is per walk-forward segment, not just once on an aggregate backtest. Computing a DSR for each out-of-sample window in a walk-forward chain, using that window’s own trial count and return distribution, tells you whether the edge holds up statistically stage by stage rather than producing one flattering number from data that spans a regime shift partway through. And practically, having a DSR figure on hand does real work for the psychology of not intervening — when a validated system hits a rough stretch, a number you calculated in advance, saying the strategy’s edge was statistically well-supported before you ever went live, is a steadier thing to hold onto than a gut feeling about whether this losing streak means something’s wrong.