Model calibration audit

For every bucket of predicted probability, what did the model actually hit? A well-calibrated model has actual win rate ≈ predicted prob. A negative gap means the model is overconfident in that bucket (dangerous); positive gap means it's underpredicting (safe). Rows turn red when |gap| > 5pp AND N ≥ 10.

Calibration cohort by sport (click to filter):
All sports MLB N=3,582TENNIS_WTA N=909TENNIS_ATP N=873MMA_MIXED_MARTIAL_ARTS N=388NFL_PRESEASON N=145LALIGA N=116NCAAF N=95SERIEA N=87EPL N=86LIGUE1 N=84UCL N=61BUNDESLIGA N=53NBA N=12NCAAB N=11NHL N=9AMERICANFOOTBALL_NFL N=2

Combined "all sports" is rarely meaningful — sports differ in market efficiency, signal availability, and base rates. Use the chips to drill into a single sport. N<50 (red) means the calibration is brittle; N≥200 (green) is trustworthy.

Filter: window=90d

Overall: N = 6,513 · mean predicted 57.5% · actual win rate 55.2% · gap -2.3pp · Brier 0.234 · log-loss 0.658
Calibration by predicted-probability bucket.
Predicted-prob bucket N Mean predicted Actual win rate Gap (actual − predicted) Brier
<50% 1072 42.8% 39.8% -3.0pp 0.233
50-55% 1904 52.3% 51.2% -1.0pp 0.250
55-60% 1287 57.4% 55.7% -1.6pp 0.247
60-65% 953 61.6% 52.7% -9.0pp 0.257
65-70% 535 67.2% 68.0% +0.9pp 0.217
70-75% 186 72.2% 74.2% +2.0pp 0.191
75-80% 331 76.6% 75.5% -1.1pp 0.186
80-90% 152 86.2% 86.8% +0.7pp 0.112
90%+ 93 93.8% 94.6% +0.8pp 0.051