From raw match data to a calculated probability
Every forecast on the platform follows the same pipeline. Here is what happens at each stage, what the output actually means, and what it cannot tell you.
The pipeline
Five stages, run for every fixture across every competition we cover.
- 1
Collect and validate
For each fixture we pull recent form, expected goals, squad availability, home and away splits, scheduling load, head-to-head history and how the wider market has priced the match. Every field is validated before anything downstream touches it, because a single bad score corrupts every calculation built on top of it.
- 2
Build the features
Raw data is turned into the measurements the models actually read: form weighted so recent matches count more, attacking and defensive strength adjusted for the quality of opposition faced, and league-level baselines so a goal in one competition is not treated as equivalent to a goal in another.
- 3
Run the ensemble
Several specialised models score the fixture independently - one reads form and strength, another expected goals, another situational context such as rest days and availability. Their outputs are combined rather than averaged blindly, weighted by how much each has actually mattered in comparable matches.
- 4
Calibrate
Raw model output is adjusted so the published numbers mean what they say. This is the step most platforms skip. If we publish 70%, roughly 70% of those situations should resolve that way over a large sample - otherwise the figure is decoration rather than information.
- 5
Rank and publish
Calibrated probabilities are compared against a normalised consensus of publicly available market pricing to identify where the model disagrees most. Each forecast is published with a confidence signal and the context behind it, then logged so its outcome can be recorded.
What the numbers actually mean
This is the part worth reading slowly, because it is the most commonly misunderstood thing on any prediction site.
When the model publishes 70% for an outcome, it means the model calculates a 70% chance. It does not mean the outcome will happen. Across many similar situations, roughly seven in ten resolve that way - and roughly three in ten do not.
A well-calibrated model at 70% should be wrong about 30% of the time. If it were never wrong at that level, the number would simply be incorrect. So a high-probability forecast that fails is not a broken model. It is a probability behaving exactly as a probability should.
No single result tells you anything meaningful about model quality, in either direction. Only large samples do.
The confidence signal
Every forecast carries a signal from Neutral to Strong. Strong ratings are deliberately rare - the scale exists to separate the genuinely clear-cut from the merely likely, not to make everything look convincing.
- Neutral
The model sees no meaningful separation. Read it as a coin-flip with extra steps.
- Solid
A measurable lean, but well within the range where the alternative is unremarkable.
- Good
A clear directional read supported by more than one input.
- High
Strong agreement across the ensemble, with the supporting data pointing the same way.
- Strong
The model's highest rating, issued sparingly. Still not a certainty, and never presented as one.
The six markets we model
Each is modelled separately rather than derived from the match result, because the factors that drive them genuinely differ.
Match result (1X2)
Home win, draw or away win. The foundational market, and the one where team strength and form carry the most weight.
Double chance
Two of the three outcomes combined. Useful where the model sees a clear favourite but rates the draw as live.
Both teams to score
Driven far more by defensive record and attacking consistency than by which side is stronger overall.
Total goals
Over and under lines, built from expected goals for both sides and the pace at which the competition typically plays.
Total corners
Corners accumulate gradually across ninety minutes, which makes them less vulnerable to a single moment than a goals line.
Total cards
Referee tendency matters as much as the fixture itself, so this market reads officiating history alongside team discipline.
Value analysis
Alongside each probability, the model compares its own figure against a normalised consensus of publicly available market pricing for the same outcome. Where the two disagree meaningfully, the fixture is flagged for a closer look.
Normalising matters. Published market pricing includes a built-in margin, so comparing against it raw would make almost everything look like a disagreement. We strip that margin out first, so the comparison is between two estimates of the same thing.
A flagged disagreement is not a recommendation and it is not a prediction of profit. It means our model and the wider market have reached different conclusions about the same match. Sometimes that is because the model has read something the market has not. Sometimes it is because the market knows something the model does not - a late fitness doubt, for instance. It is a starting point for your own judgement, not a substitute for it.
Where the limits are
Stated plainly, because a model presented without its limitations is a marketing document rather than a tool.
Data quality sets the ceiling
Coverage is thinner in smaller competitions. Where the underlying data is sparse, the model is less confident and you should be too.
Late news moves faster than any model
A fitness call an hour before kick-off can change a match materially. We update to kick-off, but no system reacts to information it has not yet been given.
Football is genuinely random
A deflection, a marginal offside or one refereeing decision can overturn the most reasonable analysis. That irreducible randomness is not a flaw to be engineered away.
Past accuracy is not future accuracy
Our published record describes what has already happened. It is not a projection, a target, or a rate you should expect to personally experience.
Read a real forecast
Every fixture page shows the full breakdown - probabilities, the data behind them, and the confidence signal. Three best picks a day are free.