A Sharpe ratio is an estimate with sampling error, and the best of N backtests is the maximum of N noisy estimates. This calculator applies the Bailey & López de Prado corrections for both: the probabilistic Sharpe ratio (PSR) for sample length and fat tails, and the deflated Sharpe ratio (DSR) for selection among trials. Every default is editable and dated October 2026; the reasoning is in the deflated Sharpe ratio guide. Everything runs in your browser; nothing is uploaded.

Your backtest

Results

Deflated Sharpe ratio — P(true SR > luck ceiling)

—

E[max SR] under N zero-skill trials

—
annualized luck ceiling

Probabilistic Sharpe ratio vs SR*

—

Minimum track record length to beat SR*

—
—

Sensitivity to the trial count

Trials NE[max SR] (ann.)Deflated SharpeReading
Assumptions (edit me — defaults dated October 2026)

Units: the formulas run in per-period units. Annualized inputs are divided by √f (f = observations per year); T = years × f. Skewness and kurtosis are taken as measured on the per-period return series. Confidence levels for MinTRL are fixed at 95% (Z = 1.645) and 99% (Z = 2.326), one-sided. Trials are assumed independent — correlated trials have a smaller effective N, so the honest shortcut is to count distinct strategy families at full weight.

How to read this

Probabilistic Sharpe ratio. PSR(SR*) = Φ((SR − SR*)·√(T−1) / √(1 − γ₃·SR + ((γ₄−1)/4)·SR²)), computed in per-period units: the probability that the true Sharpe exceeds the benchmark SR*, given that your estimate came from T observations with the skew and kurtosis you entered. Negative skew and fat tails widen the estimate’s error bars, so the same Sharpe earns a lower PSR.

Expected maximum Sharpe. If you evaluate N variants with zero true skill and keep the best, you have sampled the maximum of N noisy estimates. Extreme-value theory gives E[max SR] ≈ σ_SR·[(1−γ)·Z(1−1/N) + γ·Z(1−1/(N·e))] with γ ≈ 0.5772. The default σ_SR is the standard error of a Sharpe at your sample length (about 1.0 annualized for one year of daily data), which is why 100 trials on one year of daily data produce a best-by-luck Sharpe near 2.5. The growth is logarithmic in N, so the answer is not sensitive to getting N exactly right — but it is very sensitive to pretending N was 1.

Deflated Sharpe ratio. The PSR evaluated with SR* set to that luck ceiling. A DSR of 95%+ means your result is unlikely to be the best of a random search of the size you ran; a DSR near 50% means it is a coin flip; a DSR well below 50% means a zero-skill search would be expected to do as well or better.

Minimum track record length. MinTRL = 1 + (1 − γ₃·SR + ((γ₄−1)/4)·SR²)·(Z_α/(SR − SR*))² periods: how long a track record you need before the PSR against SR* reaches the chosen confidence. Shown in years for 95% and 99%, plus the length needed to beat the luck ceiling itself. If SR ≤ SR*, no track record length suffices.

What the calculator does not fix: survivorship in the data, look-ahead in the pipeline, or unmodelled costs — a leaked backtest deflates beautifully and is still fiction. Read the deflated Sharpe guide for the worked example, the backtest overfitting guide for why selection is the central problem, and the CSCV/PBO explainer for the complementary probability-of-overfitting estimate. To see the maximum-of-N effect rather than compute it, run 100 zero-edge paths in the equity-curve simulator and look at the best one.

Formulas: Bailey & López de Prado, “The Sharpe Ratio Efficient Frontier” (2012) and “The Deflated Sharpe Ratio” (2014). Normal CDF via the Hart/West rational approximation; inverse via Acklam’s algorithm with a Halley refinement step.