Elo for baseball: rating teams in a sport built on randomness
Elo ratings came from chess but fit baseball surprisingly well — if you tune them for a sport where the best teams lose 60 times a year. Here's how ours works.
Elo was invented to rank chess players, where the better player wins the vast majority of games. Baseball is the opposite extreme: the best team loses sixty times a year. Making Elo work here is a tuning problem — and the tuning is where the value lives.
The core idea
After every game, the winner takes rating points from the loser. Beat a strong team, gain a lot; beat a weak one, gain little. The gap between two ratings converts directly into a win probability, and home advantage adds its measured nudge. The rating is self-correcting by construction: overrated teams bleed points until the number matches reality.
Tuning for chaos
The crucial dial is the K-factor — how hard one result moves a rating. Set it high, and a rating chases noise (in a sport where good teams lose 40% of the time, that's fatal). Set it low, and it misses real change like a mid-season collapse. We fit ours walk-forward: every candidate setting is tested on games it hasn't seen, and predictive accuracy picks the winner. No hand-tuning, no vibes.
Elo is the floor, not the ceiling
A pure team rating misses baseball's biggest variable — who's pitching. That's why our Elo is the base layer, fused with the starter's FIP-weighted rating and blended against the market. Team quality sets the floor; the mound sets tonight's number. See it live on the MLB board.
Frequently asked questions
What is an Elo rating in sports?+
A number that rises when you win and falls when you lose, scaled by opponent strength — beating a good team earns more than beating a bad one. The gap between two teams' ratings converts directly into a win probability.
Every fixture, fully modelled — the correct-score grid, the derived markets, and the written read.
Start free