Predicted vs realized, across every calibration ledger
2,137 graded predictions across cc_scanner, fortress_fight, and fortress_rebuild,
scored against actual daily highs and closes. Bars and dots compare what the models claimed
to what the market did.
cc_scanner - survival
80% 92%
predicted vs realized, 1,176 matured picks across all sigma bands.
fortress_fight - touch
38% 22%
raw factor 0.57 over 887 expired contracts.
fortress_fight - breach
30 breaches
raw breach factor 0.19 (predicted rate vs realized) over 887 graded contracts.
fortress_rebuild - OTM at expiry
85% 99%
73 of 74 resolved. Touch: 30% predicted, 22% realized.
⚠️
Correction factors computed but NOT armed. 2 expiry week(s) < 6 (one cycle teaches one regime); 15d span < 60d (too short to have seen vol vary). Engine behavior is unchanged (factors pinned at 1.0). Raw would have been touch 0.57, breach 0.19.
🔥
Where the models are UNDER-confident: realized touch ran well above predicted on META (89% vs 36%), NVDA (71% vs 37%), DELL (68% vs 42%), SOFI (56% vs 35%). Tight strikes on fast movers get challenged more often than the dashboards claim; expect mid-week management on those.
fortress_fight: touch calibration curve
Each dot is a predicted-touch bucket of expired contracts (dot size = count, hover for
exact rates). On the dashed line, prediction equals reality. Dots below the line = the engine over-warned;
above = under-warned. Also on the books: 392 in progress (censored), 53 unexpired early touches (preview only).
Correction-factor arming gates
Raw factors are written every grading pass; the engine
only applies them when all four gates pass.
✓
Graded contracts ≥ 30
887 / 30
✓
Breach events ≥ 10
30 / 10
✕
Expiry weeks ≥ 6 (one cycle teaches one regime)
2 / 6
✕
History span ≥ 60d (must see vol vary)
15d / 60d
✓
Challenger beats raw out of sample (champion check)
haircut
Metric
N
Predicted
Realized
Raw factor
Touch
887
37.7%
21.5%
0.57
Breach
887
realized 30 breach event(s)
0.19
IV haircut
correct the input: touch model on iv × h
0.56
Champion/challenger (Brier, lower is better; holdout week 2026-W29, 313 train / 574 holdout): flat 0.1087, haircut 0.1071, raw 0.1398 → winner haircut. Ties go to raw; a correction only arms while it wins.
cc_scanner: sigma bands, then ticker by ticker
1,176 matured picks graded against daily highs (4,745 logged since 20260703). Survival = finishing
OTM. Red touch bars: realized ran hotter than predicted.
PredictedRealized
Survival % by sigma band
0.0-0.5σ - n=183
64% → 87%
0.5-1.0σ - n=613
78% → 89%
1.0-1.5σ - n=254
88% → 97%
>1.5σ - n=126
98% → 100%
Touch % by sigma band
0.0-0.5σ - n=183
56% → 75%
0.5-1.0σ - n=613
36% → 28%
1.0-1.5σ - n=254
23% → 11%
>1.5σ - n=126
5% → 0%
Predicted vs realized touch by ticker (min 3 matured picks)
Sorted by how hot reality ran vs the model. Red connectors: realized above predicted
(model too relaxed). Gray connectors: realized below predicted (model too scared).
Ticker
0% — predicted • realized — 100%
Gap (real - pred)
META
+52 ▲ n=72
NVDA
+34 ▲ n=24
DELL
+26 ▲ n=68
SOFI
+21 ▲ n=66
SPY
+14 ▲ n=28
GOOG
+13 ▲ n=125
AMZN
+8 ▲ n=71
GLD
0 n=9
RKLB
-2 n=9
ENPH
-4 n=8
GLXY
-8 n=4
NOW
-9 n=46
HIMS
-11 n=64
APP
-14 n=29
SPCX
-18 n=50
IGV
-19 n=41
RIOT
-19 n=72
COIN
-23 n=43
IREN
-23 n=106
INTC
-24 n=85
MU
-32 n=92
MDB
-37 n=63
fortress_rebuild: CC-survival forecasts
Rebuild-local ledger (same calibration machinery as fight): 148 predictions since 2026-07-03,
74 resolved. Buckets group picks by the survival odds the model claimed.
Predicted survivalRealized OTM rate
0-60% - n=4
55% → 100%
60-70% - n=4
64% → 100%
70-80% - n=11
77% → 91%
80-90% - n=23
86% → 100%
90-100% - n=32
94% → 100%
Metric
Value
Predictions logged
148
Resolved (expiry passed)
74
Predicted touch (avg)
30.0%
Realized touch
21.6%
Touch grades from exact daily highs
74 / 74
Generated by calibration_report.py (weekly cron, Sat 08:30 SGT).
Sources: runs/cc_scanner/PORTFOLIO/ledger_calibration.json, runs/fight/calib_factors.json,
runs/rebuild/calibration_stats.json. Graders grade only fully expired windows; duplicate
nightly re-forecasts of the same contract are deduped to one bet. This page mirrors the
graders' published numbers and derives nothing of its own.