Fortress toolchain - predicted vs realized

Calibration Report

Scored 2026-08-01 cc_scanner 6,155 picks logged - 3,144 matured fortress_fight 1674 graded - span 29d fortress_rebuild 174 resolved wall 0m 30s (29.5s)

Predicted vs realized, across every calibration ledger

4,992 graded predictions across cc_scanner, fortress_fight, and fortress_rebuild, scored against actual daily highs and closes. Bars and dots compare what the models claimed to what the market did.

cc_scanner - survival
83% 95%
predicted vs realized, 3,144 matured picks across all sigma bands.
fortress_fight - touch
39% 29%
raw factor 0.76 over 1674 expired contracts.
fortress_fight - breach
131 breaches
raw breach factor 0.42 (predicted rate vs realized) over 1674 graded contracts.
fortress_rebuild - OTM at expiry
85% 91%
159 of 174 resolved. Touch: 31% predicted, 28% realized.

fortress_fight: touch calibration curve

Each dot is a predicted-touch bucket of expired contracts (dot size = count, hover for exact rates). On the dashed line, prediction equals reality. Dots below the line = the engine over-warned; above = under-warned. Also on the books: 168 in progress (censored), 129 unexpired early touches (preview only).

perfectly calibrated 0-10% bucket - n=227 - predicted 4.2% - realized 1.3% 10-20% bucket - n=338 - predicted 16.2% - realized 5.6% 20-30% bucket - n=219 - predicted 24.3% - realized 18.3% 30-40% bucket - n=170 - predicted 34.4% - realized 22.4% 40-50% bucket - n=239 - predicted 44.1% - realized 32.2% 50-60% bucket - n=103 - predicted 54.9% - realized 42.7% 60-70% bucket - n=56 - predicted 65.5% - realized 55.4% 70-80% bucket - n=135 - predicted 75.9% - realized 59.3% 80-90% bucket - n=103 - predicted 84.3% - realized 80.6% 90-100% bucket - n=84 - predicted 97.0% - realized 89.3% 0 25 50 75 100 0 25 50 75 100 Predicted touch % Realized touch %

Correction-factor arming gates

Raw factors are written every grading pass; the engine only applies them when all four gates pass.

Graded contracts ≥ 30
1674 / 30
Breach events ≥ 10
131 / 10
Expiry weeks ≥ 6 (one cycle teaches one regime)
4 / 6
History span ≥ 60d (must see vol vary)
29d / 60d
Challenger beats raw out of sample (champion check)
logit
MetricNPredictedRealizedRaw factor
Touch167438.6%29.3%0.76
Breach1674realized 131 breach event(s)0.42
IV haircutcorrect the input: touch model on iv × h0.76
Champion/challenger (Brier, lower is better; holdout week ?, 0 train / 1361 holdout): flat 0.1554, haircut 0.1501, logit 0.1490, raw 0.1494 → winner logit. Ties go to raw; a correction only arms while it wins.

cc_scanner: sigma bands, then ticker by ticker

3,144 matured picks graded against daily highs (6,155 logged since 20260703). Survival = finishing OTM. Red touch bars: realized ran hotter than predicted.

Predicted Realized

Survival % by sigma band

0.0-0.5σ - n=404
63%88%
0.5-1.0σ - n=1281
79%92%
1.0-1.5σ - n=773
88%98%
>1.5σ - n=686
97%100%

Touch % by sigma band

0.0-0.5σ - n=404
58%64%
0.5-1.0σ - n=1281
37%22%
1.0-1.5σ - n=773
24%8%
>1.5σ - n=686
6%0%

Predicted vs realized touch by ticker (min 3 matured picks)

Sorted by how hot reality ran vs the model. Red connectors: realized above predicted (model too relaxed). Gray connectors: realized below predicted (model too scared).

Ticker
0% — predicted • realized — 100%
Gap (real - pred)
DELL
+20 ▲ n=134
AMZN
+18 ▲ n=135
META
+12 ▲ n=140
NVDA
+9 ▲ n=87
RIOT
+9 ▲ n=142
IBIT
+2 ▲ n=8
GLD
0 n=22
SOFI
-1 n=114
CRWV
-2 n=16
RKLB
-4 n=27
COPX
-5 n=26
NEM
-8 n=41
GOOG
-9 n=369
ENPH
-11 n=70
GLXY
-13 n=32
BMNR
-13 n=40
SPY
-15 n=99
HIMS
-15 n=163
SPCX
-15 n=78
IGV
-16 n=104
APP
-17 n=37
COIN
-18 n=177
QCOM
-18 n=60
CLSK
-19 n=35
IREN
-20 n=302
NOW
-22 n=141
MU
-26 n=252
MARA
-26 n=25
INTC
-27 n=160
MDB
-32 n=108

fortress_rebuild: CC-survival forecasts

Rebuild-local ledger (same calibration machinery as fight): 236 predictions since 2026-07-03, 174 resolved. Buckets group picks by the survival odds the model claimed.

Predicted survival Realized OTM rate
0-60% - n=5
54%100%
60-70% - n=9
66%100%
70-80% - n=30
76%80%
80-90% - n=71
85%90%
90-100% - n=59
94%97%
MetricValue
Predictions logged236
Resolved (expiry passed)174
Predicted touch (avg)31.4%
Realized touch28.2%
Touch grades from exact daily highs174 / 174

Diversification complements

Out-of-sample search for tickers that zig when the $1M model book zags — low correlation, positive behaviour on the book's worst days, and enough volatility to sell covered calls against. 9 nights logged (2026-07-24 → 2026-08-01).

No confident add yet — 8 days of history, needs ≥60 nights at 70% persistence and 60% out-of-sample delivery. Candidates are ranked below; a complement moves opposite the model book (corr ≤ 0.30) AND still pays sellable CC premium (vol ≥ 25%).
TickerGroupNightsPersistMed corrCrisis dayOOS deliverVerdict
VGMidstream (ENERGY)4100%-0.33+2.16%youngnoise
XOPOil E&P (ENERGY)4100%-0.32+1.04%youngnoise
IHIMedTech (HEALTH)3100%-0.32+0.91%youngnoise
FLUTCasinos (CONS)3100%-0.25+0.30%youngnoise
LDOSDef Tech (INDL)4100%-0.25-0.12%youngnoise
TOSTFintech (TECH)3100%-0.22+0.68%youngnoise
EXPETravel (CONS)4100%-0.18+0.90%youngnoise
PARRRefiners (ENERGY)4100%-0.14+1.66%youngnoise
“Crisis day” = mean candidate return on the model composite's worst days (average correlation lies in a crash; this is the all-weather test). “OOS deliver” = share of matured claims whose low correlation held in the 30 days AFTER the flag. 0 WATCH.
Generated by calibration_report.py (weekly cron, Sat 08:30 SGT). Sources: runs/cc_scanner/PORTFOLIO/ledger_calibration.json, runs/fight/calib_factors.json, runs/rebuild/calibration_stats.json. Graders grade only fully expired windows; duplicate nightly re-forecasts of the same contract are deduped to one bet. This page mirrors the graders' published numbers and derives nothing of its own.