How We Grade Our Own Projections
Anyone can look accurate in hindsight. The only honest record is one where the forecast is written down before the event and graded by a machine afterwards. That is how the Track record page works.
What gets locked
Several times a day the pipeline recomputes every upcoming game — new lineups, a changed starter, an updated forecast. Each run overwrites the stored projection for games that have not started. Once a game's scheduled first pitch has passed, its projection is no longer touched. What is stored is the win probability for the home team, both teams' expected runs, the starters it assumed, and whether both lineups were confirmed.
What gets measured
- Winners picked: how often the side above 50% won.
- Brier score and log loss: how good the probabilities were, not just the picks.
- Calibration by probability band: did 60% calls win about 60% of the time?
- Run error: the average gap between projected and actual total runs.
- The same figures by month and, separately, for the postseason.
Backtest versus live record
The Methodology page reports a walk-forward backtest over 2022–2026 — more than 12,000 games, all projected using only information available the night before. The live record is smaller but has one extra guarantee: it could not have been influenced by anything we learned later, including changes to the model itself. When the two disagree, trust the live record and give it time — a few hundred games is still a small sample in baseball.