Define late before counting
Possible definitions include after a fixed match minute, the final stated interval of regulation time, or all added time. Choose one before reading the results. Under the IFAB duration law, the referee makes allowance for time lost, so actual second-half duration can differ across matches.
Build actual exposure
| Match | Late-window start | Final elapsed time | Minutes at risk | Goal in window |
|---|---|---|---|---|
| A | 80:00 | 95:20 | 15.33 | 1 |
| B | 80:00 | 92:10 | 12.17 | 0 |
| C | 80:00 | 98:00 | 18.00 | 2 |
The table is illustrative. It yields 3 goals across 45.50 minutes at risk, or 3 / 45.50 = 0.0659 goals per minute. That is a descriptive rate for the invented sample, not a population forecast.
Preserve event and clock definitions
The Wyscout open-data paper describes a public event dataset that can support reproducible timing analysis. Verify how the chosen dataset records periods, added time, disallowed events and corrections. Do not silently map 90+5 to minute 95 in one source and minute 90 in another.
Compare like match states
Late windows disproportionately contain matches with a known score, tactical urgency, substitutions, fatigue, sendings-off and competition incentives. Compare or model at least:
- score difference and which team is trailing;
- player count and red-card timing;
- pre-match team strength;
- home and away status;
- competition and season;
- actual remaining exposure.
The peer-reviewed Bayesian in-play model illustrates why a forecast should condition on current match information rather than apply one unconditional percentage to every late state.
Use survival or interval methods carefully
A time-to-event model can represent the changing hazard of a next goal and censor a match at full time. A simpler interval model can predict whether at least one goal occurs in a declared window. Either method needs time-ordered validation and calibration, with no post-window data in the features.
Report uncertainty and coverage
Publish the number of matches, at-risk minutes, events and censored windows for every result, retaining the event provenance documented by the Wyscout open-data paper or the selected provider. Add an interval estimate or full model uncertainty, and show how many eligible windows were lost to missing clocks or event corrections. A precise-looking rate from a narrow state can still be too uncertain to support a useful conditional forecast.
Publication checklist
- State the late-window boundary and time convention.
- Report actual minutes at risk, not only match count.
- Name the event provider and correction policy.
- Separate descriptive rate from conditional forecast.
- Report uncertainty and subgroup sample sizes.
- Validate on later matches and disclose definition sensitivity.
- Join any live price only at its actual timestamp.
Continue with Goal Timing Statistics for broader period and exposure methods.
Continue learning
- Next guide: Red-Card Impact in Football
- Related guide: Second-Half Goal Statistics
Assumptions and limitations
The worked sample is invented. Public event data may not capture the exact whistle, review state or all clock corrections. Results depend on the late-window definition, competitions and match states, and no universal percentage should be transferred without local validation.

