Why the label can mislead
Club names persist while the underlying teams change. A five-match sample can span different managers, divisions, squads and incentives. Counting four wins from five describes those fixtures but does not identify a stable causal matchup.
Relevance checklist
| Check | Keep the match when... |
|---|---|
| Recency | The football context remains comparable |
| Personnel | Core tactical roles are still relevant |
| Competition | Rules and incentives match the target |
| Venue | Home-away conditions are represented correctly |
| Strength | Current team level is not being replaced by old status |
| Mechanism | A tactical claim is stated before the result review |
A testable workflow
Peer-reviewed football model-evaluation research supports the baseline, time-order, and reporting protocol used below.
- Define the target match and data cutoff.
- Build a baseline from current team strength, venue and available information.
- Add a head-to-head feature using a predeclared recency rule.
- Evaluate both versions on later matches.
- Keep the feature only if it improves the chosen proper score robustly.
Peer-reviewed work comparing soccer outcome models shows that features need model-level evaluation rather than intuitive acceptance. TimeSeriesSplit provides one implementation for preserving temporal order.
Small-sample arithmetic
Four wins in five meetings is an observed rate of 4 / 5 = 80%. It is not an 80% forecast for the next match. The denominator is small, matches are not identically distributed, and selection of a memorable rivalry can occur after seeing the pattern.
When qualitative history may still help
Past meetings can reveal a hypothesis, such as repeated difficulty defending one build-up pattern. Verify that the relevant players and tactical structures remain, then seek broader matches where the same mechanism occurred. The evidence is the repeatable mechanism, not the team-name streak alone.
Measure how much history is still relevant
Create an overlap record for each older meeting:
| Relevance field | Example question |
|---|---|
| Squad overlap | How many current expected starters appeared? |
| Manager continuity | Are the tactical decision-makers unchanged? |
| Competition and venue | Is the setting comparable? |
| Time elapsed | How old is the match? |
| Team-strength change | Were either side promoted, relegated or transformed? |
Do not turn this checklist into hidden subjective weights after seeing the result. If a historical feature is retained, specify its decay or inclusion rule first and compare a forecast with and without it.
A safer reader interpretation
Write "Team A won four of the last five listed meetings" rather than "Team A dominates this fixture." Then show the dates and scorelines. The first statement is auditable; the second implies persistence that the small, changing sample does not establish.
For modelling, recent team-strength and venue features usually provide a clearer baseline because they apply to every fixture. Head-to-head information earns inclusion only if it improves later-match evaluation beyond that baseline. If it does not, the record may remain useful as historical context without being promoted to predictive evidence.
Related resources
Use form tables for current windows or statistical model design for feature tests.
Continue learning
- Next guide: Limits of Football Statistics
- Related guide: Football Data Providers
Assumptions and limitations
The five-match example is illustrative. No fixed recency or sample threshold is universally valid. Competition changes, unbalanced venue histories and survivor selection can make historical meetings incomparable.

