auction_concession standalone gate KILL
Did not survive the pre-registered gate. This is a first-class result, not an error: the honest thing to do with a hypothesis that fails is publish the kill.
The Treasury auction supply concession, measured per event over the whole free auction record: long the auctioned sector's constant-maturity zero-coupon Treasury for the five sessions after each nominal coupon auction, bills otherwise, across every coupon tenor Treasury has issued since 1980. Long only, never short, never levered, each sector held at the same weight the baseline holds it. Baseline is that same equal-weight portfolio of the same sectors held on every session. Both curves are scored as excess returns over the same one-month bill series, so the comparison isolates TIMING and nothing else. Tests whether the concession dealers are paid for absorbing announced supply is large enough, and still present enough, to beat simply owning the duration all the time. Run standalone rather than as a risk-parity sleeve because the forty-six-year history is the point and the fund's own universe is floored at 2011 by its tail hedge's inception.
- 4 of 13 pre-registered gates failed. Failed: Deflated incremental edge, Walk-forward out-of-sample uplift, Probability of backtest overfitting (PBO), Incremental DSR, the specificity gate (0.95 floor). A hypothesis is killed if ANY gate fails, and the thresholds were fixed in its pre-registration file before the run, so nothing here was tuned after seeing the result.
- The edge is concentrated before publication. Before 2013-01-01 the candidate ran Sharpe 0.697 against the baseline's 0.599, an advantage of +0.099. From 2013-01-01 onward it runs -0.078 against 0.014, -0.092, and the advantage CHANGED SIGN rather than merely shrinking, so in the modern era the candidate is worse than its baseline and not just less good than it was. The effect was published in the academic literature at that split date. Scored on its own the modern era has 3368 observations and therefore a detection floor of +0.452, which its -0.092 advantage does NOT clear. Nor does the pre-publication era clear its own floor of +0.282 on its +0.099 advantage, so neither half of the sample would have been certifiable alone. The split is a pre-registered diagnostic, not a selection: neither half was scored through the gate and neither changed the verdict.
- The result is cost-model dependent. At the pre-registered 2 bps a side the candidate scores Sharpe 0.413 against the baseline's 0.452. At 5 bps it scores 0.206, BELOW the baseline, so the strategy dies somewhere between 2 and 5 bps a side. The pre-registered rate is a modern estimate applied to decades when fixed commissions ran and no index fund existed, so the true cost over the decades carrying the edge is plausibly at or beyond the rate that kills it.
The pre-registered gate
9 of 13 gates passed, scored over 1979-10-31 to 2026-06-30 (11631 sessions) against always_invested_treasury_zeros. Cumulative trial count across every hypothesis this shop has ever tried, at the time of this run: 42.
| Deflated incremental edge | -0.039 FAIL (threshold +0.000) |
| Walk-forward out-of-sample uplift | -0.028 FAIL (threshold +0.000) |
| Probability of backtest overfitting (PBO) | 0.84 FAIL (threshold 0.50) |
| Incremental DSR, the specificity gate (0.95 floor) | 0.40 FAIL (threshold 0.95) |
| Worst month, against the baseline's | -0.069 PASS (threshold +0.000) |
| Expected shortfall (95% CVaR), against the baseline's | -0.011 PASS (threshold +0.000) |
| Drawdown through the 1987 black monday, against the baseline's | -0.049 PASS (threshold +0.000) |
| Drawdown through the 1994 bond massacre, against the baseline's | -0.142 PASS (threshold +0.000) |
| Drawdown through the 1998 ltcm, against the baseline's | -0.050 PASS (threshold +0.000) |
| Drawdown through the 2008 global financial crisis, against the baseline's | -0.039 PASS (threshold +0.000) |
| Drawdown through the 2013 taper tantrum, against the baseline's | -0.086 PASS (threshold +0.000) |
| Drawdown through the 2020 treasury market dysfunction, against the baseline's | -0.049 PASS (threshold +0.000) |
| Drawdown through the 2022 rate shock, against the baseline's | -0.171 PASS (threshold +0.000) |
The numbers the gate scored
| Sharpe, candidate vs baseline | +0.413 vs +0.452 (uplift -0.039, standard error +0.147) |
| Annual return, candidate vs baseline | 0.65% vs 4.01% |
| Annualized volatility, candidate vs baseline | 1.62% vs 9.87% |
| Maximum drawdown, candidate vs baseline | 9.40% vs 33.97% |
| Worst month, candidate vs baseline | -1.90% vs -8.79% |
| Walk-forward out-of-sample Sharpe, candidate vs baseline | +0.261 vs +0.289 over 6 windows |
| PBO combinations evaluated | 924 |
Detection floor: what this sample could actually certify
The premise of a long-history test is that length buys resolution. That is checked rather than asserted, by inverting the exact statistic the gate decides on at this sample's own observation count and the candidate's own skew and kurtosis.
| Observations | 11630 |
| Smallest incremental Sharpe this sample could certify | +0.241 |
| Smallest certifiable at the fund's usual window (3000 observations) | +0.476 |
| Observed deflated uplift | -0.039 |
The observed edge is BELOW that floor, so this sample could not have certified it whatever the point estimate said. That floor scales with the square root of the sample, so an edge smaller than it stays uncertifiable here by arithmetic rather than by effort: closing the gap would take a sample orders of magnitude longer than the one that exists.
Check the source
Every number above is read straight from these committed files, nothing recomputed for this page:
config/hypotheses/standalone/auction-concession.yaml # the pre-registration, written before this run existed data/backtest/validation/auction-concession.json # the verdict, machine readable data/backtest/validation/auction-concession.md # the same verdict, written out in full