Pitcher strikeouts
It passes the backtest. The live forward test is still running: 0 of 200 settled results over 0 of 7 dates.
Promising: Real skill in the backtest, but one check is still open: calibration spread, or the live forward test. Tier 1: Lead market. Promoted to Validated when the evidence says so.
The backtest
engine v14 · every graded start · walk-forwardSkill is how much lower the mean squared error of the served number is than always guessing the league-wide rate known at the time. Every date is predicted using only earlier dates, with calibration curves refit each date the way the site serves them.
Model better
Compared with the pitcher's own start-by-start average, shrunk toward the league. Difference in mean squared error, model minus player: -0.549 (95% interval -0.741 to -0.372). Negative means the model is more accurate.
| Line | Skill | AUC | Predicted | Happened |
|---|---|---|---|---|
| Over 5.5 | 16.4% | .742 | 35.5% | 34.7% |
| Over 6.5 | 15.2% | .759 | 23.2% | 22.4% |
| Over 7.5 | 13.2% | .776 | 14.2% | 13.0% |
0 of 200 settled results · 0 of 7 dates
Only predictions frozen before first pitch, under the engine now serving, count. The count restarts with every engine version, which is why nothing is Validated yet.
The last engine with a finished live sample (engine v8, 312 results over 12 dates) showed skill of 24.8% (13.9% to 35.4%).
Book more accurate
Log loss, model minus the no-vig closing price: +0.092 (95% interval +0.060 to +0.134) on 117 predictions over 5 dates, engine v8. Positive means the book was more accurate. This comparison changes no status; it answers whether the number is a reason to bet.
The model counts strikeouts well, but at the book's own line it has not beaten a coin flip. Strikeout gaps show Pass until it does, on 100 or more pre-game starts over 7 or more dates, with the interval clear of zero.
Read the research →