vx_short_carry standalone gate KILL
Did not survive the pre-registered gate. This is a first-class result, not an error: the honest thing to do with a hypothesis that fails is publish the kill.
The VIX-futures term-structure carry, measured over the full free settlement history of the VX contract: a fully collateralized short of the two nearest listed contracts, weighted by the published daily roll that every short-volatility exchange-traded product tracked, from the product's first session on 2004-03-26. Collateral earns the one-month T-bill rate; the roll's own mechanically determined turnover is charged a pre-registered transaction cost. Baseline is buy-and-hold of the whole US equity market over the identical sessions. Tests whether selling the front of the volatility curve beats owning the equity risk it is insuring, on risk-adjusted terms AND on the tail. Run standalone rather than as a risk-parity sleeve because the long history is the entire point and the fund's own universe is floored at 2011.
- 11 of 11 pre-registered gates failed. Failed: Deflated incremental edge, Walk-forward out-of-sample uplift, Probability of backtest overfitting (PBO), Incremental DSR, the specificity gate (0.95 floor), Worst month, against the baseline's, Expected shortfall (95% CVaR), against the baseline's, Drawdown through the 2008 global financial crisis, against the baseline's, Drawdown through the 2011 us downgrade, against the baseline's, Drawdown through the 2015 august selloff, against the baseline's, Drawdown through the 2018 february vol spike, against the baseline's, Drawdown through the 2020 q1 covid crash, against the baseline's. A hypothesis is killed if ANY gate fails, and the thresholds were fixed in its pre-registration file before the run, so nothing here was tuned after seeing the result.
- The answer moves with the cost assumption. At the pre-registered 15 bps a side the candidate scores Sharpe 0.552 against the baseline's 0.656. At 45 bps it scores 0.450. It is below the baseline at the pre-registered rate already, so the stress only widens a gap that a friendlier cost assumption would not have closed. The stressed rate is a pre-registered diagnostic, not a selection: it was fixed before the run and the verdict is scored at the pre-registered rate either way.
- The source archives disagree on some prices. The two independent published archives behind this series overlap on 7610 dates and disagree on 1. Precedence is fixed by a mechanical rule that reads which archive covers a date, never which price looks right, so on a disputed date it can take the wrong value. None of the disputed prices entered a position the walk actually held, so no disputed value reached the verdict. That is checked against the roll schedule rather than assumed.
The pre-registered gate
0 of 11 gates passed, scored over 2004-03-26 to 2026-06-16 (5591 sessions) against us_market_buy_and_hold. Cumulative trial count across every hypothesis this shop has ever tried, at the time of this run: 41.
| Deflated incremental edge | -0.104 FAIL (threshold +0.000) |
| Walk-forward out-of-sample uplift | -0.539 FAIL (threshold +0.000) |
| Probability of backtest overfitting (PBO) | 0.53 FAIL (threshold 0.50) |
| Incremental DSR, the specificity gate (0.95 floor) | 0.32 FAIL (threshold 0.95) |
| Worst month, against the baseline's | +0.785 FAIL (threshold +0.000) |
| Expected shortfall (95% CVaR), against the baseline's | +0.088 FAIL (threshold +0.000) |
| Drawdown through the 2008 global financial crisis, against the baseline's | +0.399 FAIL (threshold +0.000) |
| Drawdown through the 2011 us downgrade, against the baseline's | +0.539 FAIL (threshold +0.000) |
| Drawdown through the 2015 august selloff, against the baseline's | +0.445 FAIL (threshold +0.000) |
| Drawdown through the 2018 february vol spike, against the baseline's | +0.885 FAIL (threshold +0.000) |
| Drawdown through the 2020 q1 covid crash, against the baseline's | +0.533 FAIL (threshold +0.000) |
The numbers the gate scored
| Sharpe, candidate vs baseline | +0.552 vs +0.656 (uplift -0.104, standard error +0.225) |
| Annual return, candidate vs baseline | 4.63% vs 11.25% |
| Annualized volatility, candidate vs baseline | 70.06% vs 19.08% |
| Maximum drawdown, candidate vs baseline | 99.26% vs 54.57% |
| Worst month, candidate vs baseline | -95.73% vs -17.21% |
| Walk-forward out-of-sample Sharpe, candidate vs baseline | +0.267 vs +0.807 over 6 windows |
| PBO combinations evaluated | 924 |
Detection floor: what this sample could actually certify
The premise of a long-history test is that length buys resolution. That is checked rather than asserted, by inverting the exact statistic the gate decides on at this sample's own observation count and the candidate's own skew and kurtosis.
| Observations | 5590 |
| Smallest incremental Sharpe this sample could certify | +0.392 |
| Smallest certifiable at the fund's usual window (3000 observations) | +0.544 |
| Observed deflated uplift | -0.104 |
The observed edge is BELOW that floor, so this sample could not have certified it whatever the point estimate said. That floor scales with the square root of the sample, so an edge smaller than it stays uncertifiable here by arithmetic rather than by effort: closing the gap would take a sample orders of magnitude longer than the one that exists.
Check the source
Every number above is read straight from these committed files, nothing recomputed for this page:
config/hypotheses/standalone/vx-short-carry.yaml # the pre-registration, written before this run existed data/backtest/validation/vx-short-carry.json # the verdict, machine readable data/backtest/validation/vx-short-carry.md # the same verdict, written out in full