Free research tool
Brier Score Calculator: Probability Forecast Accuracy
Score probabilities after the outcomes are known. Enter one probability and one binary result per row to calculate mean squared error and compare it with a reference forecast on the same cases.
Score resolved forecasts
Your scenario
0.065 Brier score
- Resolved forecasts
- 4
- Mean squared probability error (lower is better)
- 0.065
- Reference score on these same outcomes
- 0.25
- Sample event rate (descriptive)
- 50%
- Error of a constant fitted to this sample
- 0.25
- Skill versus your constant reference
- 74%
The binary 0–1 convention: average (probability − outcome)², with probabilities divided by 100. Zero is perfect and one is worst. Positive skill means lower error than your chosen constant reference on the same cases. The sample's event rate is fitted after outcomes are known; its error is a descriptive benchmark, not a forecast available in advance. A small or selectively chosen sample does not establish forecasting skill; Brier score alone does not separate calibration from resolution or measure trading profit.
Continue in Vultax
Compare the score with a documented market study
Read Vultax's published prediction-market calibration study with its population, observation dates and limitations. Then explore the prediction desk to investigate the market context behind a probability.
Worked example
Forecasts of 70%, 30%, 80% and 20%, with outcomes 1, 0, 1 and 0, have errors 0.09, 0.09, 0.04 and 0.04. The average Brier score is 0.065. A constant 50% reference scores 0.25 on those same outcomes, so this sample has 74% Brier skill versus that reference.
Binary Brier score = average [(probability % ÷ 100 − outcome)²]. Brier skill = 1 − your score ÷ reference score, when the reference score is greater than zero.
Lower error is better; a small sample is still small
This tool uses the binary 0–1 convention, where zero is perfect and one is worst. A correct 70% forecast gets less error than a correct 55% forecast, but an incorrect 70% forecast is penalised more. The score evaluates the whole probability rather than just whether it crossed 50%.
Do not select only forecasts that worked. Define the population and observation window, keep the original pre-outcome probabilities, and include all eligible resolved cases. Many correlated markets can contain less independent information than their row count suggests.
Accuracy, calibration and trading profit answer different questions
Brier score combines aspects of calibration and resolution. A lower score alone does not show that a forecaster's 70% calls happen 70% of the time. That requires calibration checks over enough observations.
A forecast can score well and still lose money if the entry price, fees and execution are unfavourable. Vultax's published calibration study documents its sample and method; use the payout calculator to compare a separate trading scenario.
Questions about this calculation
- What is a good Brier score?
- Lower is better, but the useful comparison is with a justified reference on the same outcomes and population. A 50% constant reference scores 0.25 for binary events. An outcome's base rate can make that a weak reference.
- Why is my Brier skill score undefined?
- When the reference is perfect on the entered cases, its error is zero. Dividing by zero cannot produce a meaningful skill ratio, so the tool reports the ratio as undefined rather than inventing a percentage.