Four layers of evaluation
| Layer | Question | Suitable evidence |
|---|---|---|
| Forecast | Were probabilities useful? | Proper scores, calibration, uncertainty |
| Price | Was the quote favourable under the frozen model? | Timestamped accepted price and EV |
| Execution | Was the decision implemented as specified? | Stake, acceptance, costs, settlement |
| Outcome | What happened? | Complete results and return distribution |
Do not use the realised outcome to rewrite the probability that was available beforehand. Outcome bias research shows that people evaluate the same decision differently after learning its result (Baron and Hershey, 1988).
Two decisions, opposite outcomes
Decision A: p = 0.50 at odds 2.20. EV = (0.50 * 2.20) - 1 = 0.10 units. The bet loses.
Decision B: p = 0.40 at odds 2.20. EV = (0.40 * 2.20) - 1 = -0.12 units. The bet wins.
Under the stated estimates, A was positive EV and B was negative EV. That conclusion still depends on whether 0.50 and 0.40 were defensible, timestamped probabilities. The outcomes alone cannot answer that.
What would support an edge claim
The probability-quality requirements below follow calibration guidance, while the execution fields are the local SportSignals audit contract:
- Forecasts preserved before events.
- One locked model and selection rule.
- Evaluation on later fixtures not used in tuning.
- Calibration and proper scores against relevant baselines.
- Executable accepted prices, costs and all qualifying rows.
- Uncertainty intervals around performance.
- Repeated confirmation after drift checks.
Calibration guidance explains how predicted probabilities should relate to later frequencies. Positive returns without probability-quality evidence may reflect price selection, variance, missing rows or a real advantage; the record must distinguish them.
Why short records mislead
Expected value is an average over repeated comparable decisions, not a schedule for when profit appears. OpenStax's expected-value material also distinguishes expectation from outcome variability. Do not impose one universal number of bets as sufficient; variance depends on prices, edge size, dependence, staking and selection frequency.
A fair review meeting
- Hide outcomes during the first decision-quality pass.
- Check target, inputs, cutoff, model version and probability.
- Recalculate EV and sensitivity from the accepted price.
- Reveal result and settlement only after the process grade.
- Aggregate results by predeclared market and probability band.
- Record changes prospectively rather than editing historical labels.
Extreme short-period performance can also move back toward a longer-run level without any change in ability; regression-to-the-mean research explains why extreme repeated measurements need care.
Run a blinded decision review
Export each forecast, input summary, accepted price, EV range, and rule outcome while hiding the match result and financial return. Ask a reviewer to grade target alignment, data cutoff, probability evidence, execution, and adherence to the declared rule. Reveal outcomes only after the process grades are saved.
Compare process grades with later results in aggregate. If high-grade decisions are not calibrated or repeatedly rely on unavailable prices, the process rubric needs revision. If low-grade decisions happen to win, do not promote their features into the model. This exercise directly counters the result-dependent evaluation documented in outcome-bias research.
Next step
Use Closing Line Value for the next part of this topic.
Continue learning
- Next guide: Overround and Value
- Related guide: Regression to the Mean in Betting Analysis
Assumptions and limitations
The examples use invented probabilities and ignore costs. No single test proves stable skill, and a sound process can still lose money. This page is an evaluation framework, not a promise that enough volume turns an invalid model into an edge.

