Scorecard
Two checks, not one number to trust blindly. First: how reliable has the model been on years of out-of-sample data it never saw during fitting. Second: a live, append-only record of this site's own daily predictions, scored after the fact.
Descriptive statistics, not signals — touching a level is not earning money on it. Every number on this page ships with its sample size (n) and confidence interval. See methodology.
When we say X%, how often does it happen?
Every grid probability is bucketed by what we predicted (10-point-percentage bins) and compared against the realized rate on out-of-sample data — 2018–2026, data the fitting process never touched (fitting period: 2010–2017). If the model is honest, the realized bar should sit right on top of the predicted bar in every bucket.
Futures (index, energy, metals)
MAE: 0.84 pp · n = 164,920,140
predictedrealized (out-of-sample)
0-10%
predicted 3.4% · realized 3.5% · n = 56,042,860
10-20%
predicted 15.1% · realized 15.0% · n = 20,652,607
20-30%
predicted 24.8% · realized 24.8% · n = 16,450,897
30-40%
predicted 34.3% · realized 34.3% · n = 11,760,851
40-50%
predicted 45.2% · realized 45.1% · n = 12,768,405
50-60%
predicted 53.8% · realized 53.7% · n = 7,300,159
60-70%
predicted 65.9% · realized 65.8% · n = 9,908,539
70-80%
predicted 74.0% · realized 73.6% · n = 12,967,003
80-90%
predicted 86.1% · realized 86.2% · n = 16,535,975
90-100%
predicted 90.3% · realized 89.9% · n = 532,844
FX majors (mid-price)
MAE: 0.49 pp · n = 776,797,176
predictedrealized (out-of-sample)
0-10%
predicted 2.4% · realized 2.4% · n = 410,478,455
10-20%
predicted 14.7% · realized 15.2% · n = 85,032,162
20-30%
predicted 24.5% · realized 25.3% · n = 46,212,284
30-40%
predicted 35.0% · realized 35.6% · n = 48,007,249
40-50%
predicted 45.2% · realized 46.3% · n = 27,125,949
50-60%
predicted 55.3% · realized 56.2% · n = 38,443,925
60-70%
predicted 64.0% · realized 64.9% · n = 42,457,281
70-80%
predicted 75.8% · realized 76.4% · n = 36,150,469
80-90%
predicted 83.5% · realized 83.9% · n = 42,889,402
90-100%
no held-out observations in this bucket
Buckets are ranked honestly — including the ones with the largest gap, not just the flattering ones. Weight is nheld (held-out sample size) per grid point, so thinly-sampled corners of the grid don't get equal say with the well-sampled center.
Month by month since 2018
Same check as the reliability tables above, but every single month instead of a pooled bucket: 09:30 ET, +0.5 ATR, on CL, ES, GC, NQ, RTY, YM. Predicted is the number fixed during the fitting period (data through 2017 only) — nothing after that touched the numbers.
overall gap
0.13 pp
|mean predicted − realized| over all 103 months
worst month
15.77 pp
2018-01 — including the months it fit worse
Brier score
0.239
over 13,278 market-sessions
20267 months · avg gap 2.49 pp · n = 900
202512 months · avg gap 5.21 pp · n = 1,544
202412 months · avg gap 4.06 pp · n = 1,554
202312 months · avg gap 5.34 pp · n = 1,542
202212 months · avg gap 3.47 pp · n = 1,548
202112 months · avg gap 2.08 pp · n = 1,548
202012 months · avg gap 4.19 pp · n = 1,550
201912 months · avg gap 4.38 pp · n = 1,548
201812 months · avg gap 5.49 pp · n = 1,544
Live track record
Accumulates from launch day. Every morning's map is frozen with a timestamp before the session opens (09:30 ET) and scored after the close (16:00 ET) — no cherry-picking possible. Frozen columns are never edited; only the outcome is filled in afterward.
No predictions recorded yet — this section fills in automatically starting the first trading day after launch. Here's the schema every row will follow: