dynamicsystemsarchitecture.org

MLB Backtest Analysis

The same four-channel framework used elsewhere on this site, applied to MLB run-margin prediction and walk-forward tested — with the synthetic-data disclosure stated up front, not buried.

Synthetic data — stated plainly, not a footnote. This analysis runs on realistic synthetic game data, not scraped real MLB history. The numbers below are real results of a real walk-forward process, but they characterize how the framework behaves on this synthetic set — not a claim about real-market performance. Recomputed independently from the raw results file before publishing, not just copied from the summary: MAE and improvement figures below were calculated directly from all 240 games in the results file, matching the source analysis closely (1.581 / 1.259 / 1.303 vs. the source's 1.58 / 1.26 / 1.30). One real discrepancy, stated rather than hidden: the source document references 324 total games; the results file actually available contains 240. Treat the fold-level numbers below as verified against 240 games, not 324.

Three models, four channels

The framework splits MLB game features into channels — market, team strength, context, matchup — and compares three prediction approaches on run margin:

Results, independently recomputed

ModelMAE (runs)Improvement over Model 1
Model 1 — Market1.581
Model 2 — Fundamental1.25920.4%
Model 3 — Residual correction1.30317.5%

Walk-forward tested across 12 sequential folds through the 2023–2024 window; Model 2 (fundamentals alone) won the majority of individual folds, not just the aggregate.

What this doesn't support

The source analysis is direct about this, and it's worth repeating rather than softening: prediction error of about ±1.3 runs doesn't clear the roughly 2-run edge needed to overcome standard vigorish betting margins directly. The honest application isn't "bet on this," it's calibrating fair-line estimates and comparing them against posted odds — and even that claim needs real market data, not synthetic data, before it means anything.

What would need to change to make this a real result

Stated directly in the source material rather than glossed over: real MLB history in place of synthetic data (the actual next step, not yet done), real scraped moneyline odds instead of synthetic market proxies, and accounting for what the synthetic data explicitly doesn't model — injury volatility, pitcher rest and workload, weather, and mid-season roster changes.

Comments

We'll only use this to contact you if we reply.