vrp_put_write standalone gate KILL
Did not survive the pre-registered gate. This is a first-class result, not an error: the honest thing to do with a hypothesis that fails is publish the kill.
The equity volatility risk premium, measured over the longest free daily history available: the CBOE S&P 500 PutWrite Index (a fully collateralized, cash-secured at-the-money SPX put write rolled monthly), charged a pre-registered implementation drag and scored against plain SPY total return over the identical sessions. Tests whether harvesting the premium beats simply owning the index it is written on, both statistically and on the tail. Run standalone rather than as a risk-parity sleeve because the long history is the entire point and the fund's own universe cannot reach it.
- 4 of 9 pre-registered gates failed. Failed: Walk-forward out-of-sample uplift, Probability of backtest overfitting (PBO), Incremental DSR, the specificity gate (0.95 floor), Worst month, against the baseline's. A hypothesis is killed if ANY gate fails, and the thresholds were fixed in its pre-registration file before the run, so nothing here was tuned after seeing the result.
- The frictional assumption moves the answer. The pre-registered implementation drag is 1.00% a year. Doubled to 2.00% the candidate scores Sharpe 0.542 against the baseline's 0.610. The published index carries no fund fee, no commission and no bid-ask on any of its option round trips, so a drag is the only thing standing between the index and an investable claim.
The pre-registered gate
5 of 9 gates passed, scored over 1996-08-02 to 2026-08-04 (7543 sessions) against spy_buy_and_hold. Cumulative trial count across every hypothesis this shop has ever tried, at the time of this run: 37.
| Deflated incremental edge | +0.009 PASS (threshold +0.000) |
| Walk-forward out-of-sample uplift | -0.254 FAIL (threshold +0.000) |
| Probability of backtest overfitting (PBO) | 0.90 FAIL (threshold 0.50) |
| Incremental DSR, the specificity gate (0.95 floor) | 0.52 FAIL (threshold 0.95) |
| Worst month, against the baseline's | +0.012 FAIL (threshold +0.000) |
| Expected shortfall (95% CVaR), against the baseline's | -0.008 PASS (threshold +0.000) |
| Drawdown through the 2018 february vol spike, against the baseline's | -0.012 PASS (threshold +0.000) |
| Drawdown through the 2018 q4 selloff, against the baseline's | -0.035 PASS (threshold +0.000) |
| Drawdown through the 2020 q1 covid crash, against the baseline's | -0.047 PASS (threshold +0.000) |
The numbers the gate scored
| Sharpe, candidate vs baseline | +0.619 vs +0.610 (uplift +0.009, standard error +0.186) |
| Annual return, candidate vs baseline | 7.43% vs 10.40% |
| Annualized volatility, candidate vs baseline | 12.97% vs 19.31% |
| Maximum drawdown, candidate vs baseline | 37.60% vs 55.19% |
| Worst month, candidate vs baseline | -17.73% vs -16.52% |
| Walk-forward out-of-sample Sharpe, candidate vs baseline | +0.640 vs +0.893 over 6 windows |
| PBO combinations evaluated | 924 |
Detection floor: what this sample could actually certify
The premise of a long-history test is that length buys resolution. That is checked rather than asserted, by inverting the exact statistic the gate decides on at this sample's own observation count and the candidate's own skew and kurtosis.
| Observations | 7542 |
| Smallest incremental Sharpe this sample could certify | +0.311 |
| Smallest certifiable at the fund's usual window (3000 observations) | +0.498 |
| Observed deflated uplift | +0.009 |
The observed edge is BELOW that floor, so this sample could not have certified it whatever the point estimate said. That floor scales with the square root of the sample, so an edge smaller than it stays uncertifiable here by arithmetic rather than by effort: closing the gap would take a sample orders of magnitude longer than the one that exists.
Check the source
Every number above is read straight from these committed files, nothing recomputed for this page:
config/hypotheses/standalone/vrp-put-write.yaml # the pre-registration, written before this run existed data/backtest/validation/vrp-put-write.json # the verdict, machine readable data/backtest/validation/vrp-put-write.md # the same verdict, written out in full