Prediction Model Research

Verified football prediction accuracy data. All statistics are computed from live model predictions β€” not simulations or backtests.

62,233
Total Evaluated Predictions
15,647
Unique Matches Evaluated
4
Ensemble Models
48.7%
Ensemble Accuracy

Model Comparison

Each model independently predicts every match. The ensemble combines all models to reduce overfitting. All metrics computed on out-of-sample match data.

Model Predictions Correct Accuracy Brier Score
Bayesian Poisson 15,292 8,098 53.0% 0.205
GB-RF (Gradient Boosted) 15,647 8,034 51.3% 0.200
Dixon-Coles 15,647 7,978 51.0% 0.197
Ensemble (Combined) 15,647 7,619 48.7% 0.190

Accuracy by League

Ensemble model accuracy across top 12 leagues. Minimum 500 evaluated matches per league for statistical significance.

League Matches Accuracy
Primeira Liga96556.9%
SΓΌper Lig65654.9%
Eredivisie89351.6%
La Liga1,49549.2%
Bundesliga1,23349.8%
Ligue 180448.0%
Serie A96147.9%
MLS1,06047.8%
Championship1,44746.7%
Premier League1,13945.2%
J.League79042.2%

Key Findings

1. The Bayesian Poisson model is our strongest single predictor

At 53.0% accuracy across 15,292 matches, the Bayesian Poisson model consistently outperforms other single models. Its strength comes from incorporating prior distributions for team attack/defense strength, reducing variance when data is sparse (e.g., early season, newly promoted teams).

2. Lower-division and non-European leagues are harder to predict

J.League (42.2%) and Championship (46.7%) have the lowest accuracy. These leagues have higher variance, less media coverage for injury data, and more unpredictable team performance week-to-week. Top European leagues generally cluster in the 47-51% range.

3. Ensemble reduces variance at the cost of slightly lower raw accuracy

While the ensemble (48.7%) has lower raw accuracy than Bayesian Poisson alone (53.0%), it achieves the best Brier score (0.190 vs 0.205). This means the ensemble's probability estimates are better calibrated β€” when it says 60%, it means 60%, not 65% or 55%. For value betting, calibration matters more than raw accuracy.

Methodology

Read our methodology deep dives:

ELO Rating System Bayesian Poisson Model Kelly Criterion πŸ“Š PnL Backtest