What a score can—and cannot—teach.
A resolved forecast gives you feedback. The number becomes useful when you keep its limits in view.
The result and the forecast are different objects
A highly likely event can fail to occur, and an unlikely event can happen. That is not a loophole in probability; it is part of what probability describes. A single surprising result therefore cannot establish that a forecast was foolish, just as one favorable result cannot establish that its reasoning was sound. You need a way to evaluate stated probabilities without pretending every forecast promised a particular outcome.
A scoring rule provides structured feedback on the probabilities you wrote down. It rewards assigning more probability to the observed outcome, while also accounting for the weight placed on alternatives. The score does not inspect your research notes or judge whether the question was fair. Those parts still require a separate review.
Follow one calculation all the way through
Imagine three paths with probabilities of 25%, 50% and 25%. If the middle path occurs, the observed outcome vector is zero, one, zero. Convert the probabilities to fractions: 0.25, 0.50 and 0.25. Subtract the corresponding observed value from each probability, square the differences and add them together. The sum is 0.0625 plus 0.25 plus 0.0625, or 0.375.
Forkcast divides that multiclass sum by two, producing a normalized Brier score of 0.1875. The normalization places the score between zero and one. Zero is the best possible score: all probability was on the outcome that occurred. One is the largest error: all probability was on a different outcome. Lower is better, but the useful interpretation depends on the question and the comparison.
Confidence makes the feedback sharper
A forecast that assigns 100% to one path receives a perfect score if that path occurs. If another path occurs, the normalized score is one. A more distributed forecast gives up the possibility of a perfect score on that question in exchange for a less severe score when another outcome arrives. The arithmetic makes unsupported certainty expensive when it is wrong.
That does not mean you should flatten every forecast to avoid embarrassment. If one outcome is strongly supported, a high probability may be appropriate. The point is to make the strength of your statement visible. The scoring rule should encourage honest estimates, while your research process determines what estimate the available evidence can actually support.
Do not compare unlike collections carelessly
An average score is affected by which questions you choose. A person forecasting easy, well-defined events can have a lower average than someone tackling genuinely difficult questions, even if the second person brings substantial skill to the harder task. Differences in time horizon, available evidence and outcome structure also affect the comparison.
Within a personal journal, keep the sample in view. How many forecasts have resolved? Did you select only topics where you already felt confident? Did unresolved or awkward questions disappear from your review? A low average across three hand-picked forecasts is descriptive of those three forecasts. It is not a credential, a performance guarantee or a general measure of your ability to understand markets.
Review the process beside the number
After recording an outcome, read the original question and the evidence checkpoint. Did you use the named source? Were the paths mutually exclusive and exhaustive? Did a material ambiguity become visible only at resolution? A score computed from a poorly defined question can look precise while answering very little about your forecasting process.
Then inspect the probabilities in the context of the evidence available before resolution. Avoid reasoning backward from the known result as though it had always been obvious. Identify one process choice worth retaining and one worth changing. Perhaps you found the right primary source early but gave excessive weight to repeated commentary. That specific lesson is more actionable than a broad conclusion that you are good or bad at forecasting.
Build a record, not a verdict
Forkcast scores the last saved probability allocation before you record the outcome. Once resolved, the forecast is locked against further editing. If you want to evaluate earlier versions as well, export them before revising and keep the comparison explicit. A final estimate made close to the result is a different object from an estimate made a month earlier.
Let the journal accumulate a useful history of questions, assumptions and outcomes. Watch for patterns across comparable forecasts instead of turning each result into a referendum on your judgment. The score can direct your attention, but it cannot replace it. Better forecasting is a practice of asking clearer questions, using evidence carefully and giving your future self an honest record to learn from.