Market calibration: do favorites win as often as their price says?
How to test whether prediction-market prices are well calibrated, with price bins, a reliability chart, the Brier score and the sample-size check most studies skip.
If prediction-market prices are probabilities, they can be graded like probabilities. A forecaster who says "70%" on a hundred different questions should be right about seventy times. If they are right ninety times, they were underconfident. If they are right fifty times, they were overconfident.
That property is called calibration, and it is one of the most direct ways to look for mispricing in event contracts. This guide shows how to measure it.
What calibration means
A set of prices is well calibrated if, for every price level p, the contracts priced at p resolve YES a fraction p of the time.
Calibration is not the same as accuracy. A market that prices every NFL game at 50¢ could be perfectly calibrated if favorites and underdogs are evenly mixed, while telling you nothing useful. Good markets are both calibrated and sharp, meaning their prices are often far from 50¢ when the outcome is predictable.
Step 1: Collect settled markets
You need a large set of markets that have already settled, each with:
- a price observed at a consistent time before the event, and
- the outcome: 1 if it settled YES, 0 if NO.
Pick the price carefully. The last trade before settlement is a trap: for an event that is effectively decided before the market closes (a game that is over in the fourth quarter, a data release that already happened), the last price will be near 0 or 1 and calibration will look perfect for the wrong reason. Use a snapshot from a fixed time, such as the opening price, or the price an hour or a day before the event starts.
Step 2: Bucket by price
Group the contracts into price bins, usually 10¢ wide. For each bin, compute:
- Count of contracts in the bin.
- Average price, the implied probability.
- YES rate, the fraction that settled YES.
In a spreadsheet, with prices in column B and outcomes in column C:
Count =COUNTIFS($B:$B, ">="&F2, $B:$B, "<"&G2)
Avg price =AVERAGEIFS($B:$B, $B:$B, ">="&F2, $B:$B, "<"&G2)
YES rate =AVERAGEIFS($C:$C, $B:$B, ">="&F2, $B:$B, "<"&G2)
Gap =YES rate − Avg price
where F2 and G2 hold the lower and upper edge of the bin.
Step 3: Add a margin of error
This is the step most calibration charts leave out, and it decides whether you have found anything.
A bin's YES rate is a sample proportion. Its standard error is approximately:
SE = SQRT(p × (1 − p) / n)
where p is the average price and n is the count. A rough 95% interval is the YES rate ± 2 × SE.
Here is a hypothetical table:
| Price bin | Count | Avg price | YES rate | Gap | ±2 SE |
|---|---|---|---|---|---|
| 0.10–0.20 | 180 | 0.15 | 0.11 | −0.04 | ±0.053 |
| 0.40–0.50 | 240 | 0.45 | 0.46 | +0.01 | ±0.064 |
| 0.80–0.90 | 200 | 0.85 | 0.89 | +0.04 | ±0.050 |
Every gap here is smaller than its margin of error. Even though the pattern looks like a favorite–longshot bias, this sample cannot distinguish it from chance. You would need several times more contracts, or a larger gap, before treating it as real.
Step 4: Chart it
A reliability diagram plots average price on the x-axis and YES rate on the y-axis, one point per bin, with a diagonal line for perfect calibration. Points above the line mean contracts at that price won more often than priced (underpriced YES). Points below mean they won less often (overpriced YES). Error bars from step 3 show which deviations matter.
Step 5: Score it
To summarize calibration and sharpness in one number, use the Brier score: the mean squared difference between price and outcome.
Brier = AVERAGE((price − outcome)²)
In a spreadsheet: =SUMPRODUCT((B2:B1000-C2:C1000)^2)/COUNT(B2:B1000).
Lower is better. Always pricing 50¢ scores 0.25. Perfect foresight scores 0. Brier scores are most useful for comparison: the same markets at one hour versus one day before the event, or one category versus another.
Pitfalls
- Correlated outcomes. Twenty contracts on the same game, or on adjacent temperature ranges for the same day, are not twenty independent observations. Count events, not contracts, when judging sample size, or keep one contract per event.
- Mixing categories. Sports, economics and weather markets may be calibrated differently. Pooling them can hide or create patterns.
- Changing conditions. A bias found in one season may be traded away in the next. Check whether the gap is stable over time.
- Liquidity. Prices in thin markets are noisier. Consider filtering by volume, and report what you filtered.
From calibration gap to trading edge
A calibration gap is not automatically profitable. Suppose contracts priced at 85¢ really do settle YES 89% of the time. Buying at the 85¢ mid isn't possible: you pay the ask, plus a fee. If the ask is 86¢ and the fee is 1¢, your cost is 87¢ against an 89% win rate, an expected edge of 2¢ per contract before variance. A wider spread or a higher fee could erase it. Work through the numbers in Expected value and edge, then size any position conservatively using the Kelly criterion.
Run it in Jordan
"Do Kalshi NFL favorites win as often as their price says?" is one of the starting questions in Jordan. Jordan collects the settled game markets for the series onto a sheet, buckets them by price with COUNTIFS and AVERAGEIFS formulas, builds the comparison table and a chart, and shows you every change before it lands. Because the full dataset is on the sheet, you can change the bin width or the price snapshot and watch the table update.
Frequently asked questions
What does it mean for a prediction market to be well calibrated?
It means that across many contracts, the price matches the observed frequency: contracts priced near 30¢ resolve YES about 30% of the time, contracts near 80¢ about 80% of the time, and so on at every price level.
How many markets do I need for a calibration study?
Enough that each price bin's margin of error is smaller than the gap you care about. For a bin priced near 50¢, detecting a 3-point gap with reasonable confidence takes on the order of a thousand independent events. Wider gaps or prices near 0 or 1 need fewer.
What is a good Brier score?
It depends entirely on the questions. Always guessing 50% gives 0.25, and anything lower shows some skill. Compare Brier scores across time horizons or categories on the same kind of questions rather than against a universal benchmark.
What is the favorite-longshot bias?
It is the tendency, documented in many betting markets, for low-probability outcomes to be priced too high and high-probability outcomes slightly too low. Whether it appears in a particular prediction market, and whether it is large enough to trade after costs, has to be measured.