Football prediction glossary

Prediction sites use statistical vocabulary without defining it, which makes the numbers harder to judge than they need to be. These are the terms that appear on ScorePredicts pages, defined in the sense they are actually used.

Written by Aleksandar PejčićFounder, developer and modeller
Last updated
17 September 2026
Reading time
4 min read

Probability and confidence

These four describe what a percentage on a prediction is claiming.

Probability #
How likely an outcome is, as a percentage. A probability is a claim about a large number of similar situations, not a prediction about one match: 70% means outcomes like this one happen roughly 70 times in 100.
Trust score #
The percentage shown next to a prediction for the selected outcome. On a Best Pick it is the model's final probability, adjusted for how that market has performed in the past; on a Value Analysis item it is the model's own probability before the market blend. A trust score of 64% is a stated 64% chance - it is not a rating of how good the prediction is, and a low one is not a warning.
Calibration #
Whether stated probabilities match observed reality. A model is calibrated if, of everything it called at 70%, close to 70% actually happened. Calibration is what makes a probability useful rather than decorative, and it can be checked by anyone.
Overconfidence #
A model whose stated probabilities are consistently higher than what it achieves - its 70% calls land at 60%. Overconfidence is the normal failure mode of an uncalibrated model and it inflates every figure the model shows.

Measuring performance

These five are how a prediction record is judged, and why a single accuracy figure is not enough.

Hit rate #
The share of predictions that turned out correct. Simple, and limited: it treats a 51% call and a 94% call as the same event, discarding the part that carried the information.
Brier score #
A scoring rule for probability forecasts. It squares the distance between the stated probability and the outcome, so a confident wrong answer costs more than a hedged one. Lower is better, and saying 50% to everything scores 0.25.
Sample size #
How many graded predictions a figure is calculated from, usually written n. Without it a percentage is not a statistic: forty predictions can produce almost any hit rate through chance alone.
Confidence interval #
The range a measured figure could plausibly take given the sample size. A 95% interval means that if the true value were outside it, results like the ones observed would be unusual. Narrow intervals mean more data, not a better model.
Baseline #
The score achievable without a model, used as the comparison point. For match result the standard baseline is always picking the home team, which lands around 44-47% - so any accuracy figure has to be read against that rather than against zero.

Model and data

These five describe how a prediction is produced.

Feature #
One input value a model reads for a fixture - a team's recent scoring rate, days since its last match, the referee's card average. ScorePredicts' match-result models read 71 of them per fixture; corner and card models use their own sets.
Expected goals (xG) #
The number of goals a team would be expected to score from the chances it created, based on the quality of each attempt rather than whether it went in. It describes performance more stably than the actual score, which is why models read it.
Ensemble #
The full set of specialised models behind a forecast. On ScorePredicts that means one model per market - and per side for goals, corners and cards - chosen by the fixture's competition group, rather than several models voting on the same question.
Isotonic regression #
The calibration method that converts raw model output into probabilities that mean what they say. It fits a curve that only ever rises, so the model's ordering is preserved while the stated percentages are pulled towards what those forecasts actually achieved.
Saturation #
When a model's outputs cluster in a narrow band instead of spreading across the range - most of its forecasts landing between 55% and 65%, for instance. A saturated model can still be useful for ranking fixtures, but its calibration cannot be reported honestly, because there are too few predictions outside the cluster to measure.

Markets and prices

These four describe the comparison between a model and the price a market has set.

Market-implied probability #
The probability a market price corresponds to, before the seller's margin is removed. Prices are the most informed public forecast of a football match, which makes them the natural benchmark for a model rather than a target to beat.
Margin #
The amount by which the probabilities implied by a set of prices add up to more than 100%. It is the price-setter's built-in markup, and it has to be removed before a price can be compared with a model's probability at all.
Consensus fair probability #
A market probability estimated from several independent price sources at once - margins removed, outliers filtered, the median taken. More robust than any single source, because one mispriced source cannot move it.
Probability gap #
How far a model's probability sits above the consensus fair probability for the same outcome, in percentage points. A gap is a statement about a large number of similar situations, not about one match, and a small one is within the noise of both estimates.