Null result

The overlay that was retired

A rule that raised exposure to 1.8x the further price sat below trend, retired in 2026 after it was shown to subtract value at every scale.

The question: The original test gave p = 1.5e-06. How can a rule with that p-value be wrong?

What was found

It's the project's best story about why a p-value isn't a verdict. The original test said: across 57 quarters, the overlay beats buy-and-hold on quarterly Sharpe 64.9% of the time, at p = 1.49e-06. It was reproduced literally and the number is correct. The mistake was in what was concluded from it: Sharpe is NOT additive across windows, so averaging it quarter by quarter and reading it as a statement about the whole period is an aggregation error. Whole-period Sharpe was 0.721 against 0.840 for doing nothing. It won many quarters by a little (median difference +0.048) and lost the high-variance ones.

27.6%49.2%70.9%92.5%114.1%k = 0k = 0.5k = 1k = 1.5k = 0 is buy-and-hold · k = 1 is what ran in productionFull historyExploration (<2023)Post-2017Holdout 2023+
Compound annual growth against how much the tilt is scaled. k = 0 is doing nothing; k = 1 is what ran in production. Monotonically downward in all four windows, holdout included.analysis/scripts/phase8_overlay_leverage_2026-09-07.py

Try it yourself

Drag the tilt's scale across the full 14 years of history. The question isn't whether it works, it's where the maximum is — and the maximum is at zero.

Exposure = 1 + k x (overlay exposure − 1) -> measured annual growth
Compound annual growth, full history (%)

Points of annual growth given up:

The four anchor points (k = 0 / 0.5 / 1 / 1.5) are measured; what lies between them is linear interpolation, not a simulation.

Illustrative example numbers for practice — not real data.

How it was tested, step by step

  1. The rule: measure how many standard deviations price sits from a power-law trend fitted without looking ahead, and the further below it sits, the more exposure is taken, up to 1.8x.
  2. The test that killed it: instead of asking "does it work?", the question became "at what scale does it work best?". Defining exposure = 1 + k x (exposure − 1), k was swept. The growth-optimal k is 0.00 — buy-and-hold — in all four windows, including the exploration window it was originally validated on.
  3. And it isn't only leverage: at MATCHED average exposure, the overlay's shape has a Sharpe of 0.733 against 0.840 for an equivalent constant position.
  4. The mechanism, which is what closes it: the residual is normalized over a rolling 365-day window, so it goes maximally negative EARLY in a decline — when the fall is recent relative to the past year — and re-centres as the bear market drags on. It's "buy while it falls", not "buy the bottom". At maximum exposure, BTC is on average 48.7% below its high and the next 90 days return −1.3%; at neutral, they return +48.7%.
What this does NOT say. Honesty about the verdict itself: the 104 maximum-exposure days cluster into 3-5 bear episodes, so the effective number of independent observations is 3 to 5, not 104. The negative information coefficients (from −0.08 at 5 days to −0.24 at 180) are NOT statistically significant (p between 0.31 and 0.51), because exposure is extremely persistent. The defensible claim isn't "it significantly predicts worse returns": it's "there is no evidence it helps and every point estimate says it hurts". What IS kept is the residual as a descriptive valuation indicator: "price is X sigmas below its long-run trend" remains real information. What fails is turning it into an exposure multiplier.

Tracking

This finding does have a real number that gets published and kept updated.

Data last updated on 2026-09-09.

← All research