My Signals
✦ SportSignals+ just now
Value SmartBetsNEW Props Predictions Live My Bets Alerts

Pre-Season Football Statistics: A Validation Test

Fact-checkedPublished Updated 4 min readGuide 8 of 25

Latest review: Added an auditable friendly-match record, lineup exposure and format controls, multi-season incremental validation, and explicit missing-data fallbacks.

Current

The supporting evidence is within its scheduled review window.

Evidence checked
Review due
In this article (11 sections)

In short

Pre-season matches are heterogeneous observations: minutes, substitutions, opposition strength, training load, venue, and competitive incentives differ. Treat any friendly-derived feature as experimental and retain it only if it improves a predeclared model on later competitive fixtures.

Pre-season friendly match with squad rotation and player substitutions
SportSignals illustration
Key Takeaways
  • Record competition status, match length, venue, substitution rules and whether the fixture was closed-door.
  • Team totals conceal academy-heavy halves and staggered substitutions.
  • Fit a baseline from prior competitive matches, squad continuity and venue.
  • Do not publish only a memorable successful summer.

1. Define eligible matches

Record competition status, match length, venue, substitution rules and whether the fixture was closed-door. Exclude or flag matches whose format cannot be reconciled.

2. Measure player exposure

Team totals conceal academy-heavy halves and staggered substitutions. Store lineups and minutes, then calculate features only for comparable exposure. Missing lineup data should be visible, not silently treated as full strength.

3. Adjust opposition and objective

Use a pre-match opponent rating and distinguish conditioning fixtures from near-competitive lineups where documented. Do not infer intent from the final score alone.

4. Build an incremental test

Fit a baseline from prior competitive matches, squad continuity and venue. Add pre-season result, xG or lineup features, then score both on the first declared block of competitive fixtures. TimeSeriesSplit provides a chronological evaluation pattern.

5. Report all seasons

Do not publish only a memorable successful summer. Football benchmark research shows why dataset and evaluation design control model conclusions.

Field Release requirement
Friendly scope Declared competitions and formats
Lineup completeness Missingness count
Baseline Competitive-data model without friendlies
Test period Fixed before model selection
Decision Keep, revise or remove feature

StatsBomb's open repository illustrates the event and lineup structures a reproducible test needs, although its coverage is not a universal friendly dataset.

Build an auditable pre-season row

For every friendly, store opponent, venue, date, format, player minutes, substitution limits and the stated competition phase. Mark closed-door or shortened matches. Keep penalties after a draw separate from regulation scoring.

Signal More useful form Main caveat
Result Goal and chance process by lineup Objectives can differ sharply
Team xG Split by likely starters and reserves Provider coverage may be incomplete
Player output Per minute with role and opposition Exposure is often tiny
Tactical shape Repeated use with personnel Coaches may be experimenting
Fitness Verified availability and minutes Workload plans are partly private

Incremental validation

At each historical league opener, build the normal pre-season baseline first. Add the friendly features using only matches available before that date. Compare the models through the opening weeks across multiple seasons and teams. Report seasons where the feature hurts as well as helps.

Avoid choosing friendly weights from the same opening fixtures used for evaluation. If the data is sparse or inconsistently collected, an explicit missing flag and conservative fallback are more honest than manufacturing precision. Pre-season information can be contextually useful without supporting a standalone predictive claim.

For readers, show the likely-starter minutes beside any team total. That one addition prevents a reserve-heavy friendly from being read as if it represented the expected league lineup. Also state whether the match used standard duration and substitution rules.

Use form tables for competitive windows or model training and testing for the release design.

Continue learning

Assumptions and limitations

No claim is made that pre-season data always helps or never helps. Friendly coverage is often selective, and transfer activity or late squad changes can invalidate early estimates.

Was this article helpful?
Sources and evidence4 sources, checked 14 Jul 2026
  1. StatsBomb Open Data (StatsBomb)Supports: First-party open football event, lineup, match, and selected 360 data, including documented file structure and licence conditions. Accessed 14 Jul 2026.
  2. TimeSeriesSplit (scikit-learn)Supports: Time-ordered model validation and avoiding training on future observations. Accessed 13 Jul 2026.
  3. Evaluating soccer match prediction models: a deep learning approach and feature optimization for gradient-boosted trees (Machine Learning)Supports: Peer-reviewed football benchmark design, model comparison, feature selection, and evaluation limits. Accessed 14 Jul 2026.
  4. Common pitfalls and recommended practices (scikit-learn)Supports: First-party guidance on leakage, inconsistent preprocessing, randomness, and reproducible evaluation. Accessed 14 Jul 2026.

David Adams

Sports Analyst at SportSignals

David writes every guide in this library, checks it against current operator rules and the named statistical sources, and records what changed in each update. The same byline runs on SportSignals News.

More from Football Statistics for BettingEditorial standards

18+

Gambling involves risk. Never bet more than you can afford to lose. If you feel gambling is affecting your life, free and confidential support is available.