The kill log
Every strategy idea we test is pre-registered, thresholds and all, before the result is ever seen, and every verdict lands here, including the failures. Most ideas fail this gate. That is the point: a survivor means something because failure here is public, not edited out.
"Variants tried" counts every distinct strategy configuration ever scored by the deflated Sharpe gate's own multiple testing correction (data/backtest/trial-ledger.jsonl), including an early overlay search that predates the formal registry below, so it can run ahead of the row count in the tables. "Killed" and "forward experiments" count only the hypotheses pre-registered and run through that registry (config/hypotheses/, data/research/, data/backtest/validation/), which the two tables below show in full. A hypothesis earns a citable edge only once it also clears the trust ladder's VERIFIED rung, the same bar this codebase enforces before letting any number reach a headline. As of today, nothing has. The ledger itself is hash-chained and staged for external timestamping, so its history cannot be quietly rewritten; the verification page shows how to check it.
- event_driven_index_add (aggressive): SHIP-OFF → KILL. Provisionally SHIP-OFF, then corrected to KILLED when the deflation gate was tightened: its incremental deflated Sharpe of 0.56 sits below the 0.95 significance floor, so the edge is below the calibrated detection floor and indistinguishable from noise. We caught our own gate shipping sub-threshold noise and killed it.
Every pre-registered hypothesis
| Hypothesis | Profile | Deflated uplift | Walk-forward OOS uplift | PBO | Incremental DSR (0.95 floor) | Verdict | Reason | Detail |
|---|---|---|---|---|---|---|---|---|
| trend_aggressive | aggressive | +0.020 | -0.023 | 0.63 | n/a | KILL | Failed walk-forward out-of-sample uplift and probability of backtest overfitting (PBO). | Full verdict → |
| event_driven_index_add | aggressive | +0.042 | +0.062 | 0.00 | 0.56 | KILLcorrected from SHIP-OFF | Provisionally SHIP-OFF, then corrected to KILLED when the deflation gate was tightened: its incremental deflated Sharpe of 0.56 sits below the 0.95 significance floor, so the edge is below the calibrated detection floor and indistinguishable from noise. We caught our own gate shipping sub-threshold noise and killed it. | Full verdict → |
| activist_13d | aggressive | +0.024 | -0.009 | 0.67 | 0.52 | KILL | Failed walk-forward out-of-sample uplift and probability of backtest overfitting (PBO) and incremental DSR (0.95 floor). | Full verdict → |
| equity_dedup_collapse | aggressive | -0.051 | +0.000 | 0.46 | 0.42 | KILL | Failed deflated incremental edge and walk-forward out-of-sample uplift and incremental DSR (0.95 floor). | Full verdict → |
| equity_dedup_reweight | aggressive | -0.089 | +0.000 | 0.41 | 0.37 | KILL | Failed deflated incremental edge and walk-forward out-of-sample uplift and incremental DSR (0.95 floor). | Full verdict → |
A hypothesis is KILLED if ANY gate fails; the incremental DSR (probability the edge is real, once selection luck is deflated out) is the decisive specificity gate, with a pre-registered 0.95 floor.
Standalone gate: hypotheses the fund's own universe cannot reach
Every hypothesis in the table above is shaped like a fund sleeve and is therefore scored on the fund's own price history, which begins in 2011 because that is when the book's tail hedge began trading. Some questions are only worth asking over a much longer history than that. Those are pre-registered and scored the same way, through the same deflated Sharpe, walk-forward and PBO gate plus a tail bound, but standalone rather than as a sleeve, so they live in their own artifacts. They belong here for the same reason everything else does.
| Hypothesis | Window | Sharpe vs baseline | Gates passed | Detection floor | Verdict | Why we are not trading it | Detail |
|---|---|---|---|---|---|---|---|
| turn_of_month | 1926-07-01 to 2026-06-30 26274 sessions | +1.089 vs +0.634 uplift +0.455 | 10 of 10 | +0.455 vs floor +0.165 clears it, at 26273 observations | SHIP-OFF does not clear program bar ↓ | Does not clear its own research program's selection bar. The edge is concentrated before publication. The result is cost-model dependent. | Full verdict → |
| vrp_put_write | 1996-08-02 to 2026-08-04 7543 sessions | +0.619 vs +0.610 uplift +0.009 | 5 of 9 | +0.009 vs floor +0.311 below it, at 7542 observations | KILL | 4 of 9 pre-registered gates failed. The frictional assumption moves the answer. | Full verdict → |
| vx_short_carry | 2004-03-26 to 2026-06-16 5591 sessions | +0.552 vs +0.656 uplift -0.104 | 0 of 11 | -0.104 vs floor +0.392 below it, at 5590 observations | KILL | 11 of 11 pre-registered gates failed. The answer moves with the cost assumption. The source archives disagree on some prices. | Full verdict → |
| auction_concession | 1979-10-31 to 2026-06-30 11631 sessions | +0.413 vs +0.452 uplift -0.039 | 9 of 13 | -0.039 vs floor +0.241 below it, at 11630 observations | KILL | 4 of 13 pre-registered gates failed. The edge is concentrated before publication. The result is cost-model dependent. | Full verdict → |