BALLAST.

← Kill log

vx_short_carry standalone gate KILL

Did not survive the pre-registered gate. This is a first-class result, not an error: the honest thing to do with a hypothesis that fails is publish the kill.

The VIX-futures term-structure carry, measured over the full free settlement history of the VX contract: a fully collateralized short of the two nearest listed contracts, weighted by the published daily roll that every short-volatility exchange-traded product tracked, from the product's first session on 2004-03-26. Collateral earns the one-month T-bill rate; the roll's own mechanically determined turnover is charged a pre-registered transaction cost. Baseline is buy-and-hold of the whole US equity market over the identical sessions. Tests whether selling the front of the volatility curve beats owning the equity risk it is insuring, on risk-adjusted terms AND on the tail. Run standalone rather than as a risk-parity sleeve because the long history is the entire point and the fund's own universe is floored at 2011.

Why it failed.
  • 11 of 11 pre-registered gates failed. Failed: Deflated incremental edge, Walk-forward out-of-sample uplift, Probability of backtest overfitting (PBO), Incremental DSR, the specificity gate (0.95 floor), Worst month, against the baseline's, Expected shortfall (95% CVaR), against the baseline's, Drawdown through the 2008 global financial crisis, against the baseline's, Drawdown through the 2011 us downgrade, against the baseline's, Drawdown through the 2015 august selloff, against the baseline's, Drawdown through the 2018 february vol spike, against the baseline's, Drawdown through the 2020 q1 covid crash, against the baseline's. A hypothesis is killed if ANY gate fails, and the thresholds were fixed in its pre-registration file before the run, so nothing here was tuned after seeing the result.
  • The answer moves with the cost assumption. At the pre-registered 15 bps a side the candidate scores Sharpe 0.552 against the baseline's 0.656. At 45 bps it scores 0.450. It is below the baseline at the pre-registered rate already, so the stress only widens a gap that a friendlier cost assumption would not have closed. The stressed rate is a pre-registered diagnostic, not a selection: it was fixed before the run and the verdict is scored at the pre-registered rate either way.
  • The source archives disagree on some prices. The two independent published archives behind this series overlap on 7610 dates and disagree on 1. Precedence is fixed by a mechanical rule that reads which archive covers a date, never which price looks right, so on a disputed date it can take the wrong value. None of the disputed prices entered a position the walk actually held, so no disputed value reached the verdict. That is checked against the roll schedule rather than assumed.

The pre-registered gate

0 of 11 gates passed, scored over 2004-03-26 to 2026-06-16 (5591 sessions) against us_market_buy_and_hold. Cumulative trial count across every hypothesis this shop has ever tried, at the time of this run: 41.

Deflated incremental edge-0.104 FAIL (threshold +0.000)
Walk-forward out-of-sample uplift-0.539 FAIL (threshold +0.000)
Probability of backtest overfitting (PBO)0.53 FAIL (threshold 0.50)
Incremental DSR, the specificity gate (0.95 floor)0.32 FAIL (threshold 0.95)
Worst month, against the baseline's+0.785 FAIL (threshold +0.000)
Expected shortfall (95% CVaR), against the baseline's+0.088 FAIL (threshold +0.000)
Drawdown through the 2008 global financial crisis, against the baseline's+0.399 FAIL (threshold +0.000)
Drawdown through the 2011 us downgrade, against the baseline's+0.539 FAIL (threshold +0.000)
Drawdown through the 2015 august selloff, against the baseline's+0.445 FAIL (threshold +0.000)
Drawdown through the 2018 february vol spike, against the baseline's+0.885 FAIL (threshold +0.000)
Drawdown through the 2020 q1 covid crash, against the baseline's+0.533 FAIL (threshold +0.000)

The numbers the gate scored

Sharpe, candidate vs baseline +0.552 vs +0.656 (uplift -0.104, standard error +0.225)
Annual return, candidate vs baseline 4.63% vs 11.25%
Annualized volatility, candidate vs baseline 70.06% vs 19.08%
Maximum drawdown, candidate vs baseline 99.26% vs 54.57%
Worst month, candidate vs baseline -95.73% vs -17.21%
Walk-forward out-of-sample Sharpe, candidate vs baseline +0.267 vs +0.807 over 6 windows
PBO combinations evaluated 924

Detection floor: what this sample could actually certify

The premise of a long-history test is that length buys resolution. That is checked rather than asserted, by inverting the exact statistic the gate decides on at this sample's own observation count and the candidate's own skew and kurtosis.

Observations5590
Smallest incremental Sharpe this sample could certify +0.392
Smallest certifiable at the fund's usual window (3000 observations) +0.544
Observed deflated uplift -0.104

The observed edge is BELOW that floor, so this sample could not have certified it whatever the point estimate said. That floor scales with the square root of the sample, so an edge smaller than it stays uncertifiable here by arithmetic rather than by effort: closing the gap would take a sample orders of magnitude longer than the one that exists.

Check the source

Every number above is read straight from these committed files, nothing recomputed for this page:

config/hypotheses/standalone/vx-short-carry.yaml   # the pre-registration, written before this run existed
data/backtest/validation/vx-short-carry.json   # the verdict, machine readable
data/backtest/validation/vx-short-carry.md   # the same verdict, written out in full