UrverkOverfitting Audit

Overfitting audit certificate

LIKELY_OVERFIT
Likely overfit

Your strategy did not clear the gate. At this sample size, and after accounting for the number of variants tried, its measured edge is not reliably distinguishable from selection luck. This is the outcome for most strategies we audit, including our own fund's. Learning it now, for the price of an audit, is far cheaper than learning it with live capital.

What each gate says about your strategy

Does your edge survive selection luck and its own estimation error?Fail

Your reading: 0.3869 (gate: 0.95)

After accounting for how many variants were tried and the noise in a Sharpe measured over this many observations, your advantage is not distinguishable from luck at the 95% confidence floor we require.

How often would this strategy look best purely by an accident of where the data is split?Fail

Your reading: 0.6143 (gate: 0.50)

Across many in-sample and out-of-sample splits, the configuration you submitted often loses its lead. That is a sign the in-sample ranking is driven by the particular split rather than a durable edge.

Did the advantage hold up out of sample?Not evaluated

Not computed. This needs a baseline series to compare against out of sample.

The detection floor on your sample

At 90 observations, the smallest edge this gate could certify at all is about +6.44 of incremental Sharpe. Your observed uplift of +2.890 is below that floor, so this sample cannot separate it from noise however the other gates read.

The floor is set by how many observations you supplied, not by your strategy. More history lowers it. A short sample cannot prove a small edge no matter how clean the numbers look.

The numbers behind the verdict

Observations90 (periods per year: 252)
Candidate Sharpe2.890
Baseline Sharpe0.000 (zero skill null, no baseline supplied)
Uplift over baseline+2.890
Selection luck haircut (E max of 40 trials)+3.391
Deflated uplift-0.500
Incremental DSR (confidence the edge beats baseline plus luck)0.387
Probability true Sharpe is above 00.952
Sharpe standard error±1.741
PBO (probability of backtest overfitting)0.61

These are the raw statistics the gate ran on your data. The verdict above is a mechanical function of them at pre-registered thresholds, not a judgment call.

The honest fine print