The dealer-positioning corner of the internet is loud on calls and silent on outcomes. Runir keeps the opposite discipline: every level we publish is graded against the next session and kept in public, win or lose. That archive is now 478,548 scored walls deep. Here is what it says.
How often the walls actually hold
Cumulative held rate across every scored wall. The shaded band is the 95% Wilson interval: it starts wide and tightens as the sample grows. We render the uncertainty rather than smooth it away.
Scored in public
84.1% held · 16833 of 20018 tested
By wall, and by regime
The same record, split by the level type and by the dealer gamma regime it broke under. Each bar carries its own sample size; the axis starts at the 50% coin-flip so real differences read honestly.
The gamma flip is the tell: it holds 94% of the time when volatility is amplifying versus 85% in a neutral regime (n=6,240). The regime changes the level's meaning, and the archive shows it.
Does the model beat the base rate?
Publishing a hit rate is one thing; a calibrated probability is another. Our breach model is scored against the raw base rate on the same archive, out of sample.
The lift over the base rate is statistically significant on the pre-registered stratum.
And here is the part a headline number can hide: is a "20% chance it breaks" actually a 20%? Each bin of out-of-sample predictions is plotted against what really happened. Dots on the diagonal mean the probabilities are honest.
This is the moat: not a prettier chart, a kept record. See it on a live name at $NVDA →, or read how every number here is computed.