BALLAST.

← Kill log

auction_concession standalone gate KILL

Did not survive the pre-registered gate. This is a first-class result, not an error: the honest thing to do with a hypothesis that fails is publish the kill.

The Treasury auction supply concession, measured per event over the whole free auction record: long the auctioned sector's constant-maturity zero-coupon Treasury for the five sessions after each nominal coupon auction, bills otherwise, across every coupon tenor Treasury has issued since 1980. Long only, never short, never levered, each sector held at the same weight the baseline holds it. Baseline is that same equal-weight portfolio of the same sectors held on every session. Both curves are scored as excess returns over the same one-month bill series, so the comparison isolates TIMING and nothing else. Tests whether the concession dealers are paid for absorbing announced supply is large enough, and still present enough, to beat simply owning the duration all the time. Run standalone rather than as a risk-parity sleeve because the forty-six-year history is the point and the fund's own universe is floored at 2011 by its tail hedge's inception.

Why it failed.
  • 4 of 13 pre-registered gates failed. Failed: Deflated incremental edge, Walk-forward out-of-sample uplift, Probability of backtest overfitting (PBO), Incremental DSR, the specificity gate (0.95 floor). A hypothesis is killed if ANY gate fails, and the thresholds were fixed in its pre-registration file before the run, so nothing here was tuned after seeing the result.
  • The edge is concentrated before publication. Before 2013-01-01 the candidate ran Sharpe 0.697 against the baseline's 0.599, an advantage of +0.099. From 2013-01-01 onward it runs -0.078 against 0.014, -0.092, and the advantage CHANGED SIGN rather than merely shrinking, so in the modern era the candidate is worse than its baseline and not just less good than it was. The effect was published in the academic literature at that split date. Scored on its own the modern era has 3368 observations and therefore a detection floor of +0.452, which its -0.092 advantage does NOT clear. Nor does the pre-publication era clear its own floor of +0.282 on its +0.099 advantage, so neither half of the sample would have been certifiable alone. The split is a pre-registered diagnostic, not a selection: neither half was scored through the gate and neither changed the verdict.
  • The result is cost-model dependent. At the pre-registered 2 bps a side the candidate scores Sharpe 0.413 against the baseline's 0.452. At 5 bps it scores 0.206, BELOW the baseline, so the strategy dies somewhere between 2 and 5 bps a side. The pre-registered rate is a modern estimate applied to decades when fixed commissions ran and no index fund existed, so the true cost over the decades carrying the edge is plausibly at or beyond the rate that kills it.

The pre-registered gate

9 of 13 gates passed, scored over 1979-10-31 to 2026-06-30 (11631 sessions) against always_invested_treasury_zeros. Cumulative trial count across every hypothesis this shop has ever tried, at the time of this run: 42.

Deflated incremental edge-0.039 FAIL (threshold +0.000)
Walk-forward out-of-sample uplift-0.028 FAIL (threshold +0.000)
Probability of backtest overfitting (PBO)0.84 FAIL (threshold 0.50)
Incremental DSR, the specificity gate (0.95 floor)0.40 FAIL (threshold 0.95)
Worst month, against the baseline's-0.069 PASS (threshold +0.000)
Expected shortfall (95% CVaR), against the baseline's-0.011 PASS (threshold +0.000)
Drawdown through the 1987 black monday, against the baseline's-0.049 PASS (threshold +0.000)
Drawdown through the 1994 bond massacre, against the baseline's-0.142 PASS (threshold +0.000)
Drawdown through the 1998 ltcm, against the baseline's-0.050 PASS (threshold +0.000)
Drawdown through the 2008 global financial crisis, against the baseline's-0.039 PASS (threshold +0.000)
Drawdown through the 2013 taper tantrum, against the baseline's-0.086 PASS (threshold +0.000)
Drawdown through the 2020 treasury market dysfunction, against the baseline's-0.049 PASS (threshold +0.000)
Drawdown through the 2022 rate shock, against the baseline's-0.171 PASS (threshold +0.000)

The numbers the gate scored

Sharpe, candidate vs baseline +0.413 vs +0.452 (uplift -0.039, standard error +0.147)
Annual return, candidate vs baseline 0.65% vs 4.01%
Annualized volatility, candidate vs baseline 1.62% vs 9.87%
Maximum drawdown, candidate vs baseline 9.40% vs 33.97%
Worst month, candidate vs baseline -1.90% vs -8.79%
Walk-forward out-of-sample Sharpe, candidate vs baseline +0.261 vs +0.289 over 6 windows
PBO combinations evaluated 924

Detection floor: what this sample could actually certify

The premise of a long-history test is that length buys resolution. That is checked rather than asserted, by inverting the exact statistic the gate decides on at this sample's own observation count and the candidate's own skew and kurtosis.

Observations11630
Smallest incremental Sharpe this sample could certify +0.241
Smallest certifiable at the fund's usual window (3000 observations) +0.476
Observed deflated uplift -0.039

The observed edge is BELOW that floor, so this sample could not have certified it whatever the point estimate said. That floor scales with the square root of the sample, so an edge smaller than it stays uncertifiable here by arithmetic rather than by effort: closing the gap would take a sample orders of magnitude longer than the one that exists.

Check the source

Every number above is read straight from these committed files, nothing recomputed for this page:

config/hypotheses/standalone/auction-concession.yaml   # the pre-registration, written before this run existed
data/backtest/validation/auction-concession.json   # the verdict, machine readable
data/backtest/validation/auction-concession.md   # the same verdict, written out in full