← Journal
track record2 min read

Reading the receipts: a public, graded track record

An analysis you can't check is just a tip. Here's how we grade every prediction against the real result — in the open — and why calibration matters more than a headline accuracy number.

Momus
Published 23 July 2026

Every football prediction service claims a good hit rate. Almost none let you check it. The screenshot of last week's winners is not evidence — it's a highlight reel with the losses cropped out. The fix is boring and non-negotiable: grade every call against the real result, keep the losses in, and publish the whole thing. That's the track record, and it's the receipt behind everything else we do.

Every prediction, scored — wins and losses

When a fixture we've analysed finishes, it's graded automatically against the final score: did the model's most-likely outcome happen? The result lands in the results archive, match by match, with a ✓ or a miss — no quiet deletions. You can scroll the record the same way we can. If a run goes cold, you'll see it.

Accuracy is the wrong headline

Here's the part most "X% accurate" claims get wrong: raw accuracy barely tells you if a model is good. A model that predicts the favourite every time will look accurate in a league full of strong home sides — and be useless, because it's told you nothing the table didn't. The professional test is calibration: over hundreds of matches, do the things it calls 60% actually happen about 60% of the time? We go deeper on why in are AI football predictions accurate?.

That's why the track record leads with a Brier score (a calibration measure — lower is better) alongside the accuracy figure, and shows both for the model and the market on the same fixtures. Which brings us to the real test.

Beating the market, not a coin flip

Anyone can beat a coin flip. The bar that matters is the market — the sharpest baseline there is. So the record puts them head to head on the identical set of settled matches: model accuracy versus market accuracy, model calibration versus market calibration. If the model can't at least hang with the closing line, it isn't adding anything. If it edges it, that's the only claim worth making — and it's checkable.

Why we'd rather show a loss

Publishing losses costs us a good screenshot and buys us the only thing that actually converts a sceptic: trust. A model that hides its misses is asking for faith. A model that grades itself in the open is asking you to check — and that's the entire difference between analysis and a tip. It's the same principle behind telling a real model from a mystery box.

Read the receipts on the track record, then browse them match by match in the results archive.

track recordcalibrationaccuracy
Share on X →
See it on the desk

Every fixture, fully modelled — the correct-score grid, the derived markets, and the written read.

Start free
More from the journal