Translate labels into tests
| Property | Measurement | Evidence required |
|---|---|---|
| Probability quality | Log loss, Brier score and calibration | Complete later outcomes and pre-event prices |
| Margin | Complete-market booksum | Simultaneous quotes for every outcome |
| Movement | Price path and timestamp | Stable market and source IDs |
| Limits | Executable or accepted stake | Current user-specific placement record |
| Acceptance | Rejection and price-change rate | Every attempted decision |
| Settlement | Rule consistency and disputes | Current rules and settled examples |
An operator may score differently by sport, competition, market and lead time. Do not convert one strong main-market result into a universal category.
What research supports
A 10,699-match historical study found statistically meaningful forecast-quality differences among ten bookmakers and across leagues (Štrumbelj and Robnik-Šikonja, 2010). A later 37-competition comparison found that bookmaker choice and market size affected odds-derived forecasts, and that exchange odds were not always best in smaller markets (Štrumbelj, 2014). These papers support measurement, not a permanent 2026 label for any named firm.
Example evaluation design
- Select one market, league, price cutoff and season before analysis.
- Capture complete prices from every included source.
- Convert prices with the same declared method.
- Score each source on the same set of matches.
- Report coverage and excluded rows.
- Compare margins and executable stakes separately from probability quality.
- Confirm any result on a later period.
An 11-league study found efficiency conclusions changed when mean prices were replaced by the best price available across bookmakers (Angelini and De Angelis, 2019). Price source and shopping policy are therefore part of the method.
Regulatory status is a different question
A source described informally as sharp may still be unavailable or inappropriate in a reader's jurisdiction. In Great Britain, consumers can use the Gambling Commission public register to verify current licence status. Licensing does not establish price accuracy, and forecast accuracy does not establish licensing.
How to write the conclusion
Prefer: “Source A had lower log loss than Source B on regulation-time Premier League 1X2 closing prices in this dated sample, with common coverage reported.”
Avoid: “Source A is the sharpest bookmaker.” The shorter statement hides market, period, method, uncertainty and missing data.
Use one common-coverage scorecard
Build the primary accuracy comparison only from fixtures, markets, selections, and timestamps present at every source. Report source-specific coverage separately. For each common row, preserve the complete market and apply the same probability conversion before calculating log loss, Brier score, and calibration. This controls the source and league differences documented by Štrumbelj and Robnik-Šikonja.
Add uncertainty around source differences and inspect whether rankings reverse by league, lead time, and odds band. If Source A is stronger at major-league closes but weaker at early lower-league prices, publish both findings. Do not average them into one permanent label. Repeat the scorecard on a later period before changing a product benchmark, and keep limit and acceptance evidence outside the forecast-accuracy score.
Next step
Use Pinnacle Sharp Bookmakers for the next part of this topic.
Continue learning
- Next guide: Value Betting vs Matched Betting
- Related guide: Closing Line Value
Assumptions and limitations
This page does not rank current operators and does not recommend opening an account. “Soft” and “sharp” remain useful shorthand only when the underlying measurement is stated. Product access, limits and prices can vary by customer and jurisdiction.

