Prediction Model Research
Verified football prediction accuracy data. All statistics are computed from live model predictions β not simulations or backtests.
Model Comparison
Each model independently predicts every match. The ensemble combines all models to reduce overfitting. All metrics computed on out-of-sample match data.
| Model | Predictions | Correct | Accuracy | Brier Score |
|---|---|---|---|---|
| Bayesian Poisson | 15,292 | 8,098 | 53.0% | 0.205 |
| GB-RF (Gradient Boosted) | 15,647 | 8,034 | 51.3% | 0.200 |
| Dixon-Coles | 15,647 | 7,978 | 51.0% | 0.197 |
| Ensemble (Combined) | 15,647 | 7,619 | 48.7% | 0.190 |
Accuracy by League
Ensemble model accuracy across top 12 leagues. Minimum 500 evaluated matches per league for statistical significance.
| League | Matches | Accuracy |
|---|---|---|
| Primeira Liga | 965 | 56.9% |
| SΓΌper Lig | 656 | 54.9% |
| Eredivisie | 893 | 51.6% |
| La Liga | 1,495 | 49.2% |
| Bundesliga | 1,233 | 49.8% |
| Ligue 1 | 804 | 48.0% |
| Serie A | 961 | 47.9% |
| MLS | 1,060 | 47.8% |
| Championship | 1,447 | 46.7% |
| Premier League | 1,139 | 45.2% |
| J.League | 790 | 42.2% |
Key Findings
1. The Bayesian Poisson model is our strongest single predictor
At 53.0% accuracy across 15,292 matches, the Bayesian Poisson model consistently outperforms other single models. Its strength comes from incorporating prior distributions for team attack/defense strength, reducing variance when data is sparse (e.g., early season, newly promoted teams).
2. Lower-division and non-European leagues are harder to predict
J.League (42.2%) and Championship (46.7%) have the lowest accuracy. These leagues have higher variance, less media coverage for injury data, and more unpredictable team performance week-to-week. Top European leagues generally cluster in the 47-51% range.
3. Ensemble reduces variance at the cost of slightly lower raw accuracy
While the ensemble (48.7%) has lower raw accuracy than Bayesian Poisson alone (53.0%), it achieves the best Brier score (0.190 vs 0.205). This means the ensemble's probability estimates are better calibrated β when it says 60%, it means 60%, not 65% or 55%. For value betting, calibration matters more than raw accuracy.
Methodology
- Data source: Match results from public football data APIs, updated every 15 minutes during active match windows.
- Out-of-sample: All accuracy metrics are computed on matches after the prediction was made. No data leakage, no backtesting on training data.
- Home/draw/away: Accuracy is measured on the 3-way match outcome. A prediction is "correct" only if the exact result (H/D/A) matches.
- Brier score: Lower is better. 0.25 = random guessing on 3 outcomes. 0.19 = our ensemble.
- Confidence calibration: Out-of-sample calibration verified. Model probabilities track actual frequencies within Β±3% across all confidence bands.
Read our methodology deep dives: