From shadow to live: how we test a model before trusting it
Our MLB model ran 79 graded paper bets against sharp closing prices before its first real bet. Inside the shadow-testing discipline every Modal model goes through.
Anyone can backtest their way to a beautiful curve — pick the right window, tune on the answers, publish the chart. Shadow testing is the harder discipline: the model runs live, on games that haven't happened, at prices it doesn't control, with every prediction timestamped before first pitch. It just doesn't stake money yet.
What our MLB shadow produced
Seventy-nine paper bets graded against sharp closing prices: 59.5% hit rate, +6.3% return. Against soft prices that would be unremarkable; against the sharpest baseline in the sport it's a real signal. Just as valuable was what the shadow changed: it showed the model's largest claimed edges were usually errors, which wrote the discipline gates the live engine now runs — price ceilings, edge caps, and a market blend the raw model didn't have.
Graduation, with a tripwire
The model went live on real money only after the shadow record justified it — and with a pre-agreed bar to stay live: positive closing-line value and a 60%+ win rate after twenty settled bets, or back to the lab, publicly. A model that can't be demoted isn't being tested; it's being marketed.
The same pipeline is grinding for our next verticals right now. Nothing on Modal reaches a paid tier without surviving its shadow first.
Frequently asked questions
What is shadow testing in sports prediction?+
Running a model live — real games, real prices, real timestamped predictions — but with paper stakes. It proves or breaks a model on exactly the conditions it will face, without a euro at risk, and produces a graded record anyone can audit.
Every fixture, fully modelled — the correct-score grid, the derived markets, and the written read.
Start free