Predicted vs realized, across every calibration ledger
3,679 graded predictions across cc_scanner, fortress_fight, and fortress_rebuild,
scored against actual daily highs and closes. Bars and dots compare what the models claimed
to what the market did.
cc_scanner - survival
81% 94%
predicted vs realized, 2,139 matured picks across all sigma bands.
fortress_fight - touch
39% 28%
raw factor 0.71 over 1417 expired contracts.
fortress_fight - breach
83 breaches
raw breach factor 0.31 (predicted rate vs realized) over 1417 graded contracts.
fortress_rebuild - OTM at expiry
85% 93%
114 of 123 resolved. Touch: 31% predicted, 28% realized.
⚠️
Correction factors computed but NOT armed. 3 expiry week(s) < 6 (one cycle teaches one regime); 22d span < 60d (too short to have seen vol vary); challenger lost to the raw model out of sample (champion/challenger check). Engine behavior is unchanged (factors pinned at 1.0). Raw would have been touch 0.71, breach 0.31.
🔥
Where the models are UNDER-confident: realized touch ran well above predicted on IBIT (75% vs 36%), META (68% vs 40%), DELL (64% vs 42%). Tight strikes on fast movers get challenged more often than the dashboards claim; expect mid-week management on those.
fortress_fight: touch calibration curve
Each dot is a predicted-touch bucket of expired contracts (dot size = count, hover for
exact rates). On the dashed line, prediction equals reality. Dots below the line = the engine over-warned;
above = under-warned. Also on the books: 229 in progress (censored), 69 unexpired early touches (preview only).
Correction-factor arming gates
Raw factors are written every grading pass; the engine
only applies them when all four gates pass.
✓
Graded contracts ≥ 30
1417 / 30
✓
Breach events ≥ 10
83 / 10
✕
Expiry weeks ≥ 6 (one cycle teaches one regime)
3 / 6
✕
History span ≥ 60d (must see vol vary)
22d / 60d
✕
Challenger beats raw out of sample (champion check)
raw
Metric
N
Predicted
Realized
Raw factor
Touch
1417
39.2%
28.0%
0.71
Breach
1417
realized 83 breach event(s)
0.31
IV haircut
correct the input: touch model on iv × h
0.70
Champion/challenger (Brier, lower is better; holdout week 2026-W30, 887 train / 530 holdout): flat 0.1853, haircut 0.1684, raw 0.1434 → winner raw. Ties go to raw; a correction only arms while it wins.
cc_scanner: sigma bands, then ticker by ticker
2,139 matured picks graded against daily highs (5,555 logged since 20260703). Survival = finishing
OTM. Red touch bars: realized ran hotter than predicted.
PredictedRealized
Survival % by sigma band
0.0-0.5σ - n=330
64% → 89%
0.5-1.0σ - n=1027
78% → 92%
1.0-1.5σ - n=463
88% → 98%
>1.5σ - n=319
97% → 100%
Touch % by sigma band
0.0-0.5σ - n=330
57% → 67%
0.5-1.0σ - n=1027
37% → 22%
1.0-1.5σ - n=463
24% → 10%
>1.5σ - n=319
6% → 0%
Predicted vs realized touch by ticker (min 3 matured picks)
Sorted by how hot reality ran vs the model. Red connectors: realized above predicted
(model too relaxed). Gray connectors: realized below predicted (model too scared).
Ticker
0% — predicted • realized — 100%
Gap (real - pred)
IBIT
+39 ▲ n=4
META
+28 ▲ n=109
DELL
+22 ▲ n=111
NVDA
+10 ▲ n=63
AMZN
+10 ▲ n=102
SOFI
+6 ▲ n=94
RIOT
+4 ▲ n=111
GLD
0 n=16
RKLB
-3 n=18
ENPH
-5 n=26
COPX
-5 n=22
NEM
-6 n=9
GOOG
-6 n=309
SPY
-7 n=59
GLXY
-10 n=9
HIMS
-16 n=123
SPCX
-16 n=63
APP
-17 n=35
IREN
-20 n=193
IGV
-20 n=74
COIN
-22 n=63
MARA
-24 n=11
NOW
-26 n=115
INTC
-26 n=133
MU
-27 n=178
MDB
-36 n=86
fortress_rebuild: CC-survival forecasts
Rebuild-local ledger (same calibration machinery as fight): 195 predictions since 2026-07-03,
123 resolved. Buckets group picks by the survival odds the model claimed.
Predicted survivalRealized OTM rate
0-60% - n=4
55% → 100%
60-70% - n=5
65% → 100%
70-80% - n=25
77% → 84%
80-90% - n=47
85% → 89%
90-100% - n=42
94% → 100%
Metric
Value
Predictions logged
195
Resolved (expiry passed)
123
Predicted touch (avg)
31.4%
Realized touch
28.5%
Touch grades from exact daily highs
123 / 123
Diversification complements
Out-of-sample search for tickers that zig when the $1M model book zags — low correlation, positive behaviour on the book's worst days, and enough volatility to sell covered calls against. 2 nights logged (2026-07-24 → 2026-07-25).
No confident add yet — 1 days of history, needs ≥60 nights at 70% persistence and 60% out-of-sample delivery. Candidates are ranked below; a complement moves opposite the model book (corr ≤ 0.30) AND still pays sellable CC premium (vol ≥ 25%).
Ticker
Group
Nights
Persist
Med corr
Crisis day
OOS deliver
Verdict
VG
Midstream (ENERGY)
1
100%
-0.36
+1.60%
young
noise
XOP
Oil E&P (ENERGY)
1
100%
-0.34
+0.69%
young
noise
IHI
MedTech (HEALTH)
1
100%
-0.27
+0.56%
young
noise
TOST
Fintech (TECH)
1
100%
-0.26
+0.90%
young
noise
FLUT
Casinos (CONS)
1
100%
-0.26
+0.57%
young
noise
LDOS
Def Tech (INDL)
1
100%
-0.23
-0.22%
young
noise
EXPE
Travel (CONS)
1
100%
-0.18
+0.69%
young
noise
PARR
Refiners (ENERGY)
1
100%
-0.13
+0.53%
young
noise
“Crisis day” = mean candidate return on the model composite's worst days (average correlation lies in a crash; this is the all-weather test). “OOS deliver” = share of matured claims whose low correlation held in the 30 days AFTER the flag. 0 WATCH.
Generated by calibration_report.py (weekly cron, Sat 08:30 SGT).
Sources: runs/cc_scanner/PORTFOLIO/ledger_calibration.json, runs/fight/calib_factors.json,
runs/rebuild/calibration_stats.json. Graders grade only fully expired windows; duplicate
nightly re-forecasts of the same contract are deduped to one bet. This page mirrors the
graders' published numbers and derives nothing of its own.