For every bucket of predicted probability, what did the model actually hit? A well-calibrated model has actual win rate ≈ predicted prob. A negative gap means the model is overconfident in that bucket (dangerous); positive gap means it's underpredicting (safe). Rows turn red when |gap| > 5pp AND N ≥ 10.
Combined "all sports" is rarely meaningful — sports differ in market efficiency, signal availability, and base rates. Use the chips to drill into a single sport. N<50 (red) means the calibration is brittle; N≥200 (green) is trustworthy.
Filter: window=90d · sport=mma_mixed_martial_arts
| Predicted-prob bucket | N | Mean predicted | Actual win rate | Gap (actual − predicted) | Brier |
|---|---|---|---|---|---|
| <50% | 110 | 40.7% | 34.5% | -6.1pp | 0.219 |
| 50-55% | 54 | 52.4% | 53.7% | +1.3pp | 0.248 |
| 55-60% | 36 | 57.8% | 44.4% | -13.4pp | 0.258 |
| 60-65% | 44 | 62.1% | 61.4% | -0.7pp | 0.240 |
| 65-70% | 37 | 67.3% | 73.0% | +5.7pp | 0.197 |
| 70-75% | 32 | 71.9% | 84.4% | +12.4pp | 0.151 |
| 75-80% | 26 | 77.7% | 69.2% | -8.4pp | 0.218 |
| 80-90% | 36 | 83.0% | 75.0% | -8.0pp | 0.195 |
| 90%+ | 13 | 93.9% | 92.3% | -1.6pp | 0.076 |