Three claims that must stay separate
- Forecast claim: a model assigns better probabilities than a declared baseline on later fixtures.
- Price claim: a model probability differs from a timestamped market price after a declared margin method.
- Performance claim: a fixed decision and staking rule produces settled results after constraints and costs. Peer-reviewed football forecasting and market research supports keeping forecast and market evidence sample-specific, while football benchmark research supports controlled forecast comparisons.
Evidence for one does not establish the others. Peer-reviewed football forecasting research studies forecasts and fixed-odds efficiency in a defined sample; it does not prove that every modern market can or cannot be beaten.
Evidence ladder
| Stage | Required record | Typical invalid shortcut |
|---|---|---|
| Forecast | Pre-match probability, target, model version | Retrospective class label |
| Baseline | Same fixtures and cutoff | Comparing different coverage |
| Price | Operator, market, decimal odds, retrieval time | Using closing odds for an earlier decision |
| Availability | Status and accepted terms | Assuming every displayed price was available |
| Settlement | Rule version, result and void handling | Deleting voids or losses |
| Return | Stake rule, bankroll path and costs | Quoting selected wins |
| Uncertainty | Sample and interval or resampling | Treating a short run as permanent |
Forecast quality comes first
Football benchmark research illustrates model comparison under a defined dataset and protocol. Use proper probability scores and calibration, not only winner accuracy. Calibration guidance explains why a 60% forecast needs repeated frequency evidence.
Illustrative expected-value calculation
At decimal odds 2.20 and an estimated win probability of 0.48, expected net return per unit under a simplified win-or-lose model is:
0.48 × 1.20 − 0.52 × 1 = 0.056
The illustrative estimate is 0.056 units before errors, limits, voids and other costs. OpenStax's expected-value material supports the weighted-outcome calculation. The crucial 0.48 remains an estimate; a miscalibrated probability can reverse the sign.
Operational reality changes the test
Operator rules govern settlement, voids and feature availability. Betfair's sportsbook rules are one operator example, not a universal rulebook. A credible record uses the rules and price that applied to each accepted decision and retains rejected or unavailable cases in operational reporting.
What would support a strong claim
A prospective, timestamped record across a predeclared eligible set; reproducible model and price snapshots; no test-period rule changes; settled results; coverage and accepted-stake reporting; appropriate uncertainty; and continued monitoring after release. Even then, the conclusion belongs to that period, market and process.
Pre-register the market experiment
Before collecting outcomes, define competitions, markets, forecast cutoff, operators, minimum price availability, margin treatment, selection rule, staking rule, maximum exposure, void treatment and stopping date. Save the specification and do not optimize it against the result period.
| Failure | Required record |
|---|---|
| No model probability | Coverage failure |
| No eligible price | Market availability failure |
| Price rejected or changed | Execution failure |
| Selection void | Rule-based settlement |
| Stake constrained | Accepted versus intended stake |
| Feed outage | Operational exclusion under predeclared rule |
Betfair's sportsbook rules are one first-party example of settlement and feature conditions. Use the applicable current operator rules for each record.
Report probability and return together
A positive financial sample with poor probability scoring may reflect variance or selective prices; a strong probability score may still have no realizable return after price and execution constraints. Publish both layers and their denominators.
Use uncertainty intervals or resampling appropriate to the dependent match and selection structure, and show the bankroll path rather than only the final total. A completed prospective study can support a sample-specific conclusion; it cannot guarantee the next period.
Continue the workflow
Use betting market data in football models to create the timestamped price record required by the second evidence stage.
Continue learning
- Next guide: Elo Rating System for Football
- Related guide: How AI Predicts Football Matches
Assumptions and limitations
The calculation is illustrative and ignores market-specific complications. This page does not say AI can reliably produce profit, does not identify a bet, and does not treat past performance as a guarantee.

