UrverkOverfitting Audit

Overfitting audit certificate

SURVIVES_DEFLATION
Survives deflation

Your strategy cleared every gate. Its edge survived selection luck deflation and its own estimation error at our 95% confidence floor. Read this as a statistical finding, not a promise. It means we could not rule your edge out as noise at this sample size, not that it will make money going forward. A live track record, not a backtest, is the only thing that proves forward edge.

What each gate says about your strategy

Does your edge survive selection luck and its own estimation error?Pass

Your reading: 0.9981 (gate: 0.95)

Your advantage over the baseline is large enough that, after accounting for how many variants were tried and the noise in a Sharpe measured over this many observations, we are at least 95% confident the true edge is positive.

How often would this strategy look best purely by an accident of where the data is split?Not evaluated

Not computed. This needs the other strategy variants you tried, supplied as trial_returns. With only the one series there is nothing to rank.

Did the advantage hold up out of sample?Not evaluated

Not computed. This needs a baseline series to compare against out of sample.

The detection floor on your sample

At 1500 observations, the smallest edge this gate could certify at all is about +0.67 of incremental Sharpe. Your observed uplift of +1.187 is above that floor, so it is inside what this sample size can resolve.

The floor is set by how many observations you supplied, not by your strategy. More history lowers it. A short sample cannot prove a small edge no matter how clean the numbers look.

The numbers behind the verdict

Observations1500 (periods per year: 252)
Candidate Sharpe1.187
Baseline Sharpe0.000 (zero skill null, no baseline supplied)
Uplift over baseline+1.187
Selection luck haircut (E max of 1 trials)+0.000
Deflated uplift+1.187
Incremental DSR (confidence the edge beats baseline plus luck)0.998
Probability true Sharpe is above 00.998
Sharpe standard error±0.409

These are the raw statistics the gate ran on your data. The verdict above is a mechanical function of them at pre-registered thresholds, not a judgment call.

The honest fine print