How the match predictions are made
Last updated 20 September 2026
Every fixture page carries a prediction: a percentage for each result, expected goals for both sides and a likeliest score. It comes from one of the simplest models in football, and this page is all of it — including where it is weaker than the bookmakers.
What goes in
Only results from the league season in progress: goals scored and conceded, and nothing else. From them each club gets two numbers, both 1.0 for an average side — an attack strength (how many it scores against the league average) and a defence strength (how many it concedes).
The league average is that season’s goals per team per match, and the home advantage is the league’s own home-to-away goal ratio so far. Combined:
- expected goals, home = league average × home attack × away defence × home advantage;
- expected goals, away = league average × away attack × home defence.
Why four matches do not decide a season
A rate taken from four matches is mostly noise, and taken literally it produces nonsense. In September 2026 this site said Spurs v Aston Villa was a 77% draw with expected goals of 0.1 to 0.2, because Spurs had not scored in four matches and the model read their attack as exactly zero. The bookmakers had Spurs as home favourites that day.
Each rate is now a blend rather than a plain average: four matches of league-average scoring are added to what a club has actually done. A club that has played four matches is therefore half itself and half the league; after twelve it is three-quarters itself; by March the blend barely shows. A side that has not scored in four matches now reads half of average instead of zero.
The size of that prior was not chosen by taste. It was gridded in a walk-forward test over three seasons, scoring each value on matches it had not seen, and reported separately for early-season matches, which is the only regime it really changes. Anything between three and six matches of prior scored the same; four was the best of them early.
From expected goals to percentages
Each side’s goals are treated as a Poisson count, and the two are combined across every scoreline from 0-0 to 8-8. Adding up the cells gives the home win, draw and away win percentages, the likeliest score, both-teams-to-score, over 2.5 goals, and expected points.
Football’s standard refinement for this — the Dixon-Coles correction, which nudges up the low-scoring draws that an independent Poisson misses — is implemented and switched off. It was tested on our data along with time decay, and neither made a difference that survived its own confidence interval. Leaving it on would have been a decoration.
A prediction appears only once both sides have played three matches in the season.
How well it does
Measured walk-forward over the 2023/24, 2024/25 and 2025/26 Premier League seasons — 1,049 matches, each predicted only from what came before it — and scored with the Ranked Probability Score, where lower is better:
| Predictor | RPS | Top pick correct |
|---|---|---|
| League base rate (always the same home/draw/away split) | 0.2330 | 42.6% |
| SquadCheck | 0.2062 | 51.4% |
| Bookmakers’ closing odds, margin removed | 0.1949 | 54.6% |
So the model is well clear of knowing nothing, and still about 0.011 RPS short of the closing line — which is the honest place for a free model built on goals alone. The Model vs Market page keeps the same comparison running on this season’s matches as they are played.
What the prediction does not know
- Injuries and suspensions. They are not an input. Each team’s Power Loss is shown beside the prediction rather than inside it, so you can see what the model has not priced.
- Lineups. The predicted eleven is built separately and does not feed the probabilities.
- Odds. Nothing is calibrated to the market; the market is only ever the benchmark.
- Home and away form separately. A club has one attack and one defence number; only the league-wide home advantage separates the two sides.
- Anything outside the league. Cup runs, European midweeks, weather, a new manager — none of it is in the data.
Treat a percentage as a rough prior on the match, not a verdict. Where the model and the absence list disagree — a heavy favourite missing three starters — the disagreement is the interesting part, and the site shows you both.