Every hour a job records the v5.3 ensemble forecast in predictions_log; after full time another job settles it against the real score. Only settled forecasts (verified_at IS NOT NULL) within the window count: accuracy = share of predicted home/draw/away outcomes that match reality. The database behind the count: 1,371,656 matches, 27,183 teams, 1,228 leagues since 2015, 26,842 forecasts settled in total. Raw data: SStats API.
Average model confidence is 45.3% against 47.9% actual accuracy — the model slightly underestimates itself on average, but the average hides a skew: low confidences are understated, high ones overstated (see calibration).
| Outcome | Predicted | Hit | Actually happened |
|---|---|---|---|
| Home win | 7,019 | 3,523 | 4,790 |
| Draw | 532 | 149 (28.0%) | 2,722 |
| Away win | 3,644 | 1,700 | 3,683 |
The model calls a draw in only 4.7% of cases vs 24.3% actual. This is structural: per-league DRAW thresholds (default 0.40) cut most draw signals. We publish this openly; shifting thresholds by hand lowers overall accuracy per backtests.
| Stated confidence | Forecasts | Actual accuracy | Stated avg |
|---|---|---|---|
| 0.30–0.40 | 11,042 | 21.8% | 20.9% |
| 0.40–0.50 | 10,356 | 47.7% | 44.3% |
| 0.50–0.60 | 3,507 | 57.3% | 54.0% |
| 0.60–0.70 | 1,028 | 59.0% | 64.1% |
| 0.70+ | 909 | 54.8% | 83.2% |
The working band is 0.50–0.60, where the model beats its own words. Above 0.70 it systematically overstates (83.2% claimed vs 54.8% actual). The old «bet on ≥0.60» advice is withdrawn.
Monthly reports live at September 2026. Full methodology: v5.3. Terms: glossary.