Three layers that should not be confused
- Event data records provider-defined actions such as shots and passes. Opta's definitions illustrate why the vocabulary must be named.
- Derived metrics transform those events, as xG or xT models do.
- Forecasts estimate an outcome that has not happened and require future-data evaluation.
StatsBomb's open-data repository makes one event structure inspectable. Proprietary providers can add broader coverage and tracking context, but access does not remove the need for definitions and version control.
Minimum evidence contract
Every analysis should identify:
- Target question and outcome period.
- Provider, field definition, competition coverage, and retrieval date.
- Inclusion, exclusion, correction, and missing-data rules.
- Information cutoff for every predictor.
- Baseline, evaluation metric, and future test sample.
- Assumptions, uncertainty, and known failure groups.
Football outcome-model research shows that feature and model comparisons are sample-dependent. Calibration guidance explains why a predicted probability must be assessed across future observations, not celebrated after one correct result.
Topic map
Metric guides cover xG, xA, xT, xPts, possession, shots, defending, referees, home advantage, form, head-to-head records, weather and scheduling. Method guides cover Poisson, broader statistical models, dashboards, provider selection and xG research workflows. Each retained URL answers one distinct practical question rather than restating this overview.
A practical route from question to evidence
Use this sequence before opening a spreadsheet or model:
Peer-reviewed football model-evaluation research supports the baseline, time-order, and reporting protocol used below.
- Write the decision in one sentence, including the information cutoff. "Describe last season" and "forecast next Saturday" are different tasks.
- Choose the smallest data layer that can answer it. A shot-quality question may need event data; a formation-spacing question may require tracking data.
- Record the provider definition and version. Familiar labels can hide different event-coding or model rules.
- Keep raw observations separate from derived features and forecasts. That makes corrections and audits possible.
- Compare against a simple baseline on later matches. Added complexity has value only when it improves the chosen evaluation measure out of sample.
- Report where the method fails, not only its overall average.
| Output | Minimum useful context | Common category error |
|---|---|---|
| Match statistic | Definition, denominator, competition, date | Treating description as prediction |
| Derived metric | Inputs, transformation, model version | Comparing providers as if definitions match |
| Forecast probability | Cutoff, training period, test period, calibration | Reading a score as certainty |
| Market comparison | Forecast timestamp, price timestamp, market rules | Calling a raw probability gap profit |
This route also improves retrieval: the metric page can answer a definition question, while the modelling and limitation pages retain the deeper validation questions.
Assumptions and limitations
This pillar does not endorse a provider, universal threshold, or betting strategy. Coverage and products can change. Examples are educational and remain conditional on their definitions, data, model and timestamp.
