Calibration: what makes a probability trustworthy
Anyone can say '70% chance'. Calibration is the test of whether that number means anything at all.
Last updated: 10 July 2026
Suppose a forecaster tells you there is a 70% chance gold finishes higher this week. What would it take for that to be a good forecast? Not that gold rises — a 70% forecast should be wrong roughly three times in ten, and a model that is never wrong at 70% was lying about the 70%.
The definition
A model is calibrated if, across all the times it said 70%, the thing happened about 70% of the time. Gather every 70% forecast it has ever made; roughly seven in ten should have come true. Do the same at 30%, at 55%, at 90%. If each bucket matches its label, the probabilities are real quantities you can reason with. If not, they are decoration.
This is a property of a long run of forecasts, never of a single one. No individual prediction can be judged right or wrong in probability terms — a 90% forecast that fails is not a mistake, it is the one time in ten. Which is precisely why a forecaster who shows you only their winners is telling you nothing.
Why overconfidence is the usual failure
Most market forecasts are badly calibrated in one particular direction: they are too confident. Things declared 90% certain happen maybe 65% of the time. Sharp, precise-sounding calls sell better than honest hedging, so the incentive runs entirely towards overstatement.
Overconfidence is expensive. If you believe a move is 90% likely when it is really 65%, you size the position for a certainty that isn't there, and the losses arrive far more often than you planned for. The cost of a miscalibrated probability is not embarrassment; it is ruin.
How Coneview keeps itself honest
- •Walk-forward backtesting. The model is trained only on data before a given date, then asked to forecast the period after it. It never sees the answer it is being graded on.
- •A baseline to beat.Directional accuracy alone is meaningless — markets drift upward, so "always say higher" scores well. We compare every forecast against that naive baseline and report the edge: how much better than nothing the model actually is.
- •The Brier score. This grades the probabilities themselves rather than just up-or-down. Saying 95% and being wrong is punished far more harshly than saying 55% and being wrong — exactly as it should be.
- •We publish the failures. On many instruments and horizons the model has no edgeover the baseline. We label those rows "no edge" and show them anyway.
The uncomfortable part
Markets are close to efficient over short horizons. A great deal of the time, the honest calibrated answer is near 50/50, and no model — ours emphatically included — has a directional edge worth acting on. When that happens, Coneview says so, in the interface, in plain language.
This is not modesty. A forecast of 50/50 is genuinely useful information: it tells you the direction is noise, and that the range is the only part of the forecast carrying signal. A product that manufactured a confident direction there would be selling you a coin toss dressed as insight.
See the numbers for yourself on our public accuracy scorecard, or read about how the forecast cone is built.
See it on a real instrument
Every forecast on Coneview ships with the backtest that says how much to trust it.
Open Coneview →