Fortress toolchain - predicted vs realized

Calibration Report

Scored 2026-09-03 cc_scanner 7,758 picks logged - 6,639 matured fortress_fight 2489 graded - span 62d fortress_rebuild 300 resolved wall 0m 6s (6.5s)

Predicted vs realized, across every calibration ledger

9,428 graded predictions across cc_scanner, fortress_fight, and fortress_rebuild, scored against actual daily highs and closes. Bars and dots compare what the models claimed to what the market did.

cc_scanner - survival
86% 88%
predicted vs realized, 6,639 matured picks across all sigma bands.
fortress_fight - touch
38% 34%
raw factor 0.89 over 2489 expired contracts.
fortress_fight - breach
365 breaches
raw breach factor 0.80 (predicted rate vs realized) over 2489 graded contracts.
fortress_rebuild - OTM at expiry
83% 86%
257 of 300 resolved. Touch: 35% predicted, 35% realized.

fortress_fight: touch calibration curve

Each dot is a predicted-touch bucket of expired contracts (dot size = count, hover for exact rates). On the dashed line, prediction equals reality. Dots below the line = the engine over-warned; above = under-warned. Also on the books: 107 in progress (censored), 15 unexpired early touches (preview only).

perfectly calibrated 0-10% bucket - n=336 - predicted 4.0% - realized 2.4% 10-20% bucket - n=503 - predicted 16.4% - realized 13.1% 20-30% bucket - n=336 - predicted 24.2% - realized 26.5% 30-40% bucket - n=265 - predicted 34.5% - realized 30.9% 40-50% bucket - n=331 - predicted 44.2% - realized 37.8% 50-60% bucket - n=165 - predicted 55.2% - realized 43.0% 60-70% bucket - n=106 - predicted 65.0% - realized 59.4% 70-80% bucket - n=190 - predicted 75.8% - realized 64.7% 80-90% bucket - n=153 - predicted 84.5% - realized 79.1% 90-100% bucket - n=104 - predicted 96.5% - realized 89.4% 0 25 50 75 100 0 25 50 75 100 Predicted touch % Realized touch %

Correction-factor arming gates

Raw factors are written every grading pass; the engine only applies them when all four gates pass.

Graded contracts ≥ 30
2489 / 30
Breach events ≥ 10
365 / 10
Expiry weeks ≥ 6 (one cycle teaches one regime)
9 / 6
History span ≥ 60d (must see vol vary)
62d / 60d
Challenger beats raw out of sample (champion check)
raw
MetricNPredictedRealizedRaw factor
Touch248938.1%33.8%0.89
Breach2489realized 365 breach event(s)0.80
IV haircutcorrect the input: touch model on iv × h0.90
Champion/challenger (rolling-origin walk-forward, 7 folds, 2175 pooled out-of-sample rows): flat Brier 0.1799 / ECE 0.0807; haircut Brier 0.1785 / ECE 0.0889; logit Brier 0.1768 / ECE 0.0752; raw Brier 0.1714 / ECE 0.0440; trend Brier 0.1764 / ECE 0.0867 (experimental, loses to raw) → winner raw. Admission needs Brier no worse than raw AND ECE better; ties go to raw; a correction only arms while it wins. Experimental shapes are scored but never armed.

fortress_fight: week by week

Each row is one expiry week of fully graded contracts. Ratio = realized ÷ predicted touch: below 1.0 the model over-warned that week, above 1.0 it under-warned. A correction that would have looked right on the first rows is judged by the later ones.

1.0 = calibrated2026-W28: n=313, predicted 34.3%, realized 24.0%, ratio 0.70W282026-W29: n=574, predicted 39.6%, realized 20.2%, ratio 0.51W292026-W30: n=530, predicted 41.8%, realized 38.7%, ratio 0.93W302026-W31: n=257, predicted 34.8%, realized 36.6%, ratio 1.05W312026-W32: n=368, predicted 43.4%, realized 50.3%, ratio 1.16W322026-W33: n=212, predicted 36.2%, realized 40.6%, ratio 1.12W332026-W34: n=120, predicted 27.3%, realized 43.3%, ratio 1.59W342026-W35: n=114, predicted 29.6%, realized 24.6%, ratio 0.83W352026-W36: n=1, predicted 28.0%, realized 0.0%, ratio 0.00W36
WeekNPred touchReal touchRatioPred breachReal breach
2026-W2831334.3%24.0%0.7016.5%6.7%
2026-W2957439.6%20.2%0.5119.1%1.6%
2026-W3053041.8%38.7%0.9320.2%10.0%
2026-W3125734.8%36.6%1.0516.8%18.7%
2026-W3236843.4%50.3%1.1620.3%38.9%
2026-W3321236.2%40.6%1.1217.2%19.3%
2026-W3412027.3%43.3%1.5913.2%31.7%
2026-W3511429.6%24.6%0.8314.3%10.5%
2026-W36128.0%0.0%0.0013.9%0.0%

Touch by prior-20-day trend (ex-ante, experimental)

Rows bucket each graded contract by its ticker’s return over the 20 trading days BEFORE the forecast, the one trend fact knowable at pick time. If realized splits by bucket while predicted does not, the model is trend-blind. Walk-forward verdict on the trend challenger: loses to raw out of sample.

Prior 20d trendNPredicted touchRealized touchGap
down >10%135740%36%-4 pts
flat +/-10%106636%32%-4 pts
up >10%5627%5%-21 pts

Scoring-run history

One row per weekly grading pass (cumulative sample at that date). Watch the raw factors converge and the winner column stay on raw.

RunFight NTouch pred / realRaw touch factorRaw breach factorWinnerScanner NScanner surv pred / realRebuild NRebuild touch pred / real
2026-09-03248938% / 34%0.890.80raw663986% / 88%30035% / 35%

cc_scanner: sigma bands, then ticker by ticker

6,639 matured picks graded against daily highs (7,758 logged since 20260703). Survival = finishing OTM. Red touch bars: realized ran hotter than predicted.

Predicted Realized

Survival % by sigma band

0.0-0.5σ - n=536
64%81%
0.5-1.0σ - n=1984
79%84%
1.0-1.5σ - n=1857
88%90%
>1.5σ - n=2262
97%94%

Touch % by sigma band

0.0-0.5σ - n=536
57%68%
0.5-1.0σ - n=1984
37%30%
1.0-1.5σ - n=1857
23%14%
>1.5σ - n=2262
7%8%

Predicted vs realized touch by ticker (min 3 matured picks)

Sorted by how hot reality ran vs the model. Red connectors: realized above predicted (model too relaxed). Gray connectors: realized below predicted (model too scared).

Ticker
0% — predicted • realized — 100%
Gap (real - pred)
ETHA
+74 ▲ n=34
NEM
+54 ▲ n=185
BMNR
+28 ▲ n=253
IBIT
+20 ▲ n=12
DELL
+17 ▲ n=236
AMZN
+16 ▲ n=221
NVDA
+13 ▲ n=177
SPY
+11 ▲ n=196
MDB
+9 ▲ n=239
IGV
+6 ▲ n=196
RIOT
+5 ▲ n=191
META
+5 ▲ n=222
GLD
-1 n=157
NOW
-1 n=237
GOOG
-2 n=654
COPX
-2 n=180
SOFI
-3 n=152
RKLB
-4 n=27
HIMS
-7 n=240
CRWV
-10 n=51
COIN
-15 n=418
QCOM
-16 n=185
ENPH
-16 n=149
APP
-19 n=182
SPCX
-19 n=151
GLXY
-19 n=155
AMD
-20 n=22
MU
-20 n=477
IREN
-22 n=518
CLSK
-22 n=157
MARA
-22 n=132
INTC
-24 n=221
SNDK
-39 n=12

fortress_rebuild: CC-survival forecasts

Rebuild-local ledger (same calibration machinery as fight): 401 predictions since 2026-07-03, 300 resolved. Buckets group picks by the survival odds the model claimed.

Predicted survival Realized OTM rate
0-60% - n=15
54%87%
60-70% - n=20
67%90%
70-80% - n=62
76%77%
80-90% - n=119
85%84%
90-100% - n=84
94%93%
MetricValue
Predictions logged401
Resolved (expiry passed)300
Predicted touch (avg)35.2%
Realized touch35.1%
Touch grades from exact daily highs299 / 299

Diversification complements

Out-of-sample search for tickers that zig when the $1M model book zags — low correlation, positive behaviour on the book's worst days, and enough volatility to sell covered calls against. 40 nights logged (2026-07-24 → 2026-09-03).

No confident add yet — 41 days of history, needs ≥60 nights at 70% persistence and 60% out-of-sample delivery. Candidates are ranked below; a complement moves opposite the model book (corr ≤ 0.30) AND still pays sellable CC premium (vol ≥ 25%).
TickerGroupNightsPersistMed corrCrisis dayOOS deliverVerdict
ABNBInternet (TECH)2196%+0.07+0.27%100%WATCH
BILLFintech (TECH)2196%-0.13+1.29%100%WATCH
EXPETravel (CONS)2196%-0.17+0.90%100%WATCH
HCCCoal (ENERGY)2196%+0.23-1.51%100%WATCH
NVOGLP-1 (HEALTH)2196%-0.20+0.78%100%WATCH
PARRRefiners (ENERGY)2196%-0.12+1.66%67%WATCH
UBERInternet (TECH)2196%+0.00-0.26%100%WATCH
VGMidstream (ENERGY)2196%-0.31+2.16%100%WATCH
VKTXGLP-1 (HEALTH)2196%+0.25-1.04%0%WATCH
XOPOil E&P (ENERGY)2196%-0.29+1.04%100%WATCH
CAVARestaurants (CONS)2095%+0.00-0.08%67%WATCH
CNRCoal (ENERGY)2095%+0.23-1.95%33%WATCH
FLUTCasinos (CONS)2095%-0.20+0.30%67%WATCH
IHIMedTech (HEALTH)2095%-0.29+0.91%100%WATCH
LDOSDef Tech (INDL)2095%-0.05-0.12%100%WATCH
LLYGLP-1 (HEALTH)2095%-0.20+0.56%100%WATCH
NFLXInternet (TECH)2095%-0.10+0.94%100%WATCH
SNOWSoftware (TECH)2091%+0.15+1.06%0%WATCH
TOSTFintech (TECH)2095%-0.01+0.68%100%WATCH
ZSCyber (TECH)2095%-0.01+0.83%100%WATCH
DKNGCasinos (CONS)1995%-0.09+0.35%0%WATCH
OIHOil Svcs (ENERGY)1986%+0.23-0.78%33%WATCH
AMRCoal (ENERGY)1777%+0.26-2.38%33%WATCH
AMGNGLP-1 (HEALTH)1676%-0.12+0.68%100%WATCH
PLTRDef Tech (INDL)1150%+0.20-1.07%0%noise
XYZFintech (TECH)480%+0.20+0.65%youngnoise
BTUCoal (ENERGY)314%+0.22-2.69%33%noise
WFRDOil Svcs (ENERGY)480%+0.22-1.30%youngnoise
FTNTCyber (TECH)1150%+0.23-0.20%100%noise
XHBBuilders (CONS)686%+0.23-0.41%youngnoise
TXSteel (MATL)480%+0.23-0.31%youngnoise
SQMRare Earth (MATL)480%+0.24-1.67%youngnoise
“Crisis day” = mean candidate return on the model composite's worst days (average correlation lies in a crash; this is the all-weather test). “OOS deliver” = share of matured claims whose low correlation held in the 30 days AFTER the flag. 24 WATCH.
Generated by calibration_report.py (weekly cron, Sat 08:30 SGT). Sources: runs/cc_scanner/PORTFOLIO/ledger_calibration.json, runs/fight/calib_factors.json, runs/rebuild/calibration_stats.json. Graders grade only fully expired windows; duplicate nightly re-forecasts of the same contract are deduped to one bet. This page mirrors the graders' published numbers and derives nothing of its own.