Predicted vs realized, across every calibration ledger
4,992 graded predictions across cc_scanner, fortress_fight, and fortress_rebuild,
scored against actual daily highs and closes. Bars and dots compare what the models claimed
to what the market did.
cc_scanner - survival
83% 95%
predicted vs realized, 3,144 matured picks across all sigma bands.
fortress_fight - touch
39% 29%
raw factor 0.76 over 1674 expired contracts.
fortress_fight - breach
131 breaches
raw breach factor 0.42 (predicted rate vs realized) over 1674 graded contracts.
fortress_rebuild - OTM at expiry
85% 91%
159 of 174 resolved. Touch: 31% predicted, 28% realized.
⚠️
Correction factors computed but NOT armed. 4 expiry week(s) < 6 (one cycle teaches one regime); 29d span < 60d (too short to have seen vol vary). Engine behavior is unchanged (factors pinned at 1.0). Raw would have been touch 0.76, breach 0.42.
🔥
Where the models are UNDER-confident: realized touch ran well above predicted on DELL (63% vs 43%), AMZN (57% vs 39%). Tight strikes on fast movers get challenged more often than the dashboards claim; expect mid-week management on those.
fortress_fight: touch calibration curve
Each dot is a predicted-touch bucket of expired contracts (dot size = count, hover for
exact rates). On the dashed line, prediction equals reality. Dots below the line = the engine over-warned;
above = under-warned. Also on the books: 168 in progress (censored), 129 unexpired early touches (preview only).
Correction-factor arming gates
Raw factors are written every grading pass; the engine
only applies them when all four gates pass.
✓
Graded contracts ≥ 30
1674 / 30
✓
Breach events ≥ 10
131 / 10
✕
Expiry weeks ≥ 6 (one cycle teaches one regime)
4 / 6
✕
History span ≥ 60d (must see vol vary)
29d / 60d
✓
Challenger beats raw out of sample (champion check)
logit
Metric
N
Predicted
Realized
Raw factor
Touch
1674
38.6%
29.3%
0.76
Breach
1674
realized 131 breach event(s)
0.42
IV haircut
correct the input: touch model on iv × h
0.76
Champion/challenger (Brier, lower is better; holdout week ?, 0 train / 1361 holdout): flat 0.1554, haircut 0.1501, logit 0.1490, raw 0.1494 → winner logit. Ties go to raw; a correction only arms while it wins.
cc_scanner: sigma bands, then ticker by ticker
3,144 matured picks graded against daily highs (6,155 logged since 20260703). Survival = finishing
OTM. Red touch bars: realized ran hotter than predicted.
PredictedRealized
Survival % by sigma band
0.0-0.5σ - n=404
63% → 88%
0.5-1.0σ - n=1281
79% → 92%
1.0-1.5σ - n=773
88% → 98%
>1.5σ - n=686
97% → 100%
Touch % by sigma band
0.0-0.5σ - n=404
58% → 64%
0.5-1.0σ - n=1281
37% → 22%
1.0-1.5σ - n=773
24% → 8%
>1.5σ - n=686
6% → 0%
Predicted vs realized touch by ticker (min 3 matured picks)
Sorted by how hot reality ran vs the model. Red connectors: realized above predicted
(model too relaxed). Gray connectors: realized below predicted (model too scared).
Ticker
0% — predicted • realized — 100%
Gap (real - pred)
DELL
+20 ▲ n=134
AMZN
+18 ▲ n=135
META
+12 ▲ n=140
NVDA
+9 ▲ n=87
RIOT
+9 ▲ n=142
IBIT
+2 ▲ n=8
GLD
0 n=22
SOFI
-1 n=114
CRWV
-2 n=16
RKLB
-4 n=27
COPX
-5 n=26
NEM
-8 n=41
GOOG
-9 n=369
ENPH
-11 n=70
GLXY
-13 n=32
BMNR
-13 n=40
SPY
-15 n=99
HIMS
-15 n=163
SPCX
-15 n=78
IGV
-16 n=104
APP
-17 n=37
COIN
-18 n=177
QCOM
-18 n=60
CLSK
-19 n=35
IREN
-20 n=302
NOW
-22 n=141
MU
-26 n=252
MARA
-26 n=25
INTC
-27 n=160
MDB
-32 n=108
fortress_rebuild: CC-survival forecasts
Rebuild-local ledger (same calibration machinery as fight): 236 predictions since 2026-07-03,
174 resolved. Buckets group picks by the survival odds the model claimed.
Predicted survivalRealized OTM rate
0-60% - n=5
54% → 100%
60-70% - n=9
66% → 100%
70-80% - n=30
76% → 80%
80-90% - n=71
85% → 90%
90-100% - n=59
94% → 97%
Metric
Value
Predictions logged
236
Resolved (expiry passed)
174
Predicted touch (avg)
31.4%
Realized touch
28.2%
Touch grades from exact daily highs
174 / 174
Diversification complements
Out-of-sample search for tickers that zig when the $1M model book zags — low correlation, positive behaviour on the book's worst days, and enough volatility to sell covered calls against. 9 nights logged (2026-07-24 → 2026-08-01).
No confident add yet — 8 days of history, needs ≥60 nights at 70% persistence and 60% out-of-sample delivery. Candidates are ranked below; a complement moves opposite the model book (corr ≤ 0.30) AND still pays sellable CC premium (vol ≥ 25%).
Ticker
Group
Nights
Persist
Med corr
Crisis day
OOS deliver
Verdict
VG
Midstream (ENERGY)
4
100%
-0.33
+2.16%
young
noise
XOP
Oil E&P (ENERGY)
4
100%
-0.32
+1.04%
young
noise
IHI
MedTech (HEALTH)
3
100%
-0.32
+0.91%
young
noise
FLUT
Casinos (CONS)
3
100%
-0.25
+0.30%
young
noise
LDOS
Def Tech (INDL)
4
100%
-0.25
-0.12%
young
noise
TOST
Fintech (TECH)
3
100%
-0.22
+0.68%
young
noise
EXPE
Travel (CONS)
4
100%
-0.18
+0.90%
young
noise
PARR
Refiners (ENERGY)
4
100%
-0.14
+1.66%
young
noise
“Crisis day” = mean candidate return on the model composite's worst days (average correlation lies in a crash; this is the all-weather test). “OOS deliver” = share of matured claims whose low correlation held in the 30 days AFTER the flag. 0 WATCH.
Generated by calibration_report.py (weekly cron, Sat 08:30 SGT).
Sources: runs/cc_scanner/PORTFOLIO/ledger_calibration.json, runs/fight/calib_factors.json,
runs/rebuild/calibration_stats.json. Graders grade only fully expired windows; duplicate
nightly re-forecasts of the same contract are deduped to one bet. This page mirrors the
graders' published numbers and derives nothing of its own.