We Tested Polymarket Prices Against 6,776 Resolved Markets. Here's Where They Bend.

Original Vultax research: a calibrated fair-value model beat raw Polymarket prices by ~9.6% on out-of-sample Brier score. Where the favorite-longshot effect lives, what the model can honestly flag, and why it still isn't a trading edge.

By Vultax Research3 min readMethodology

Live now — refreshed every 10 minutes

loading…

Brier, Polymarket day-ahead prices, 7 d
Markets in that sample
Brier, Kalshi day-ahead prices, 7 d
Markets in that sample
A faint diagonal line crossed twice by a gently bending gold curve on black
Validation sample

6,776 markets

Resolved Yes/No markets with known outcomes.

Brier improvement

≈ +9.6%

Out-of-sample, 0.0352 → 0.0319 versus raw market prices.

Serving posture

Watch context

notRecommendation=true on every served fair-value row.

The question a fair-value model must answer

Prediction-market prices are good probability estimates on average — that is the standard result, and it held in our data. But 'good on average' leaves room for systematic, repeatable distortion, and the best-documented one is the favorite-longshot effect: low-probability outcomes trade above their true frequency and near-certain outcomes trade slightly below it.

A fair-value model, honestly defined, answers one question: given where this market is priced, where did markets priced like it historically resolve? That is a calibration question, and calibration questions have a standard scoring rule.

How to read the live figures

The strip at the top of this page is rewritten every ten minutes from Polymarket's resolved markets of the last seven days with the CLOB's yes price 24 hours before each market's end, and Kalshi's settled markets with the candlestick close 24 hours before settlement, the same rule on both venues, volume-weighted, markets under a day old skipped. The Vi IQ beneath it scores the six domains the Vultax terminal scores for a pair, computed for this page's subject from published inputs, with a domain that cannot be computed shown as unavailable rather than filled in. The figures in the body of the article are the measurement it was written on and are dated; the strip is what the same measurement says now.

The Brier score is the mean squared error between the day-ahead price and the outcome, so lower is better and 0.25 is a coin flip; the sample sizes beside them say how much weight to give each. The pack behind the page breaks each venue into price buckets so the favourite-longshot bend described above can be seen forming week by week.

Context from elsewhere

The published evidence on calibration has grown quickly. Kalshi's own August 2026 study over more than 2.2 million data points found no systematic bias in aggregate, forecasts improving with time, participation and volume, and markets with about $10,000 of volume or twenty traders already informative; its offered yardstick is the 0.10–0.15 Brier scores of skilled superforecasters. A 2026 academic study of 594 Kalshi macro markets reported an overall Brier of 0.0987, with federal-funds-rate contracts near-perfect (0.0001) and unemployment-rate contracts weakest (0.1302), and the one-hour horizon far better than the thirty-day one (0.0812 against 0.1587). Horizon, more than venue, decides the number.

Volume-weighted calibration inherits a problem this site has written about elsewhere. Columbia researchers flagged about a quarter of Polymarket's volume over three years as wash trading, peaking near 60% in December 2024 and highest in sports markets; a calibration curve weighted by volume gives those prints a vote. The weekly test on this page weights by volume for comparability and states the floor it uses; the bend it finds at the extremes should be read with that in mind.

Method and result

We assembled resolved Yes/No Polymarket markets with known outcomes — 6,776 after filtering for resolvability — and fit a calibration mapping from market price to resolved frequency, evaluating strictly out of sample. The scoring rule is the Brier score: mean squared error between stated probability and outcome, where lower is better.

Raw market prices scored approximately 0.0352. The calibrated mapping scored approximately 0.0319 — an improvement of roughly 9.6%. The improvement concentrates where theory says it should: at the probability extremes, where longshot enthusiasm and near-certainty discounting are strongest.

  • Sample: 6,776 resolved Yes/No markets.
  • Score: out-of-sample Brier, 0.0352 → 0.0319.
  • Improvement: ≈ 9.6% versus raw prices.
  • Effect concentrated near the 0 and 1 extremes.

Why this is not a trading edge

The distance between a calibration result and a tradeable edge is wide, and it is worth walking through it explicitly. First, the improvement is an aggregate: it says the population of prices bends in a known way, not that any specific market is wrong today. Second, the largest dislocations sit at the extremes, where bid-ask spreads are widest relative to the gap — a two-cent theoretical dislocation inside a three-cent spread is not an opportunity, it is noise with good posture.

Third, capture requires timing and settlement: the market can stay 'dislocated' until resolution, and the honest benchmark for any live-edge claim is close-line value — did flagged prices systematically beat the closing line? That measurement is future work on our roadmap, and until it exists, the product treats every fair-value flag as watch context. Every served row carries notRecommendation=true, and this article is bound by the same rule.

How to use the fair-value watch honestly

Used as designed, the fair-value watch is an attention allocator. A market flagged outside its calibrated band is a market worth understanding: check the liquidity, check the event room, check who is trading it and whether venues agree. Sometimes the price is stale; sometimes the flag is thin-category noise; sometimes the market simply knows something the base rates do not.

That triage — flag, then investigate structure — is the entire honest use of the model. Anyone promising more from a calibration curve is selling past their evidence.

Questions this page answers

How accurate are prediction markets?
Well calibrated on liquid contracts near resolution: the venues' own studies and academic work put Brier scores in the 0.02–0.10 range close to settlement, against 0.25 for a coin flip. A month out the same markets are much less informative, and prices at the extremes bend away from the outcome frequencies.
What is a Brier score?
The mean squared difference between a probability and the outcome, 0 for a perfect forecaster and 0.25 for one who always says 50%. The live strip shows the day-ahead Brier score on both venues for the last seven days.
What is the favourite-longshot bias?
The tendency for low-priced contracts to win less often than their price implies and high-priced ones more often. The article measures it on 6,776 resolved markets; the weekly pack checks whether it persists.
Is a mispriced market a trading edge?
Not on its own. The bend is small, fees and spreads sit on top of it, and it is measured after the fact across thousands of markets. The section above on why this is not an edge is the durable part of this page.

Revision notes

This page is kept current. Each entry records what changed and when; the live figures above refresh on their own.

  • Added the August 2026 Kalshi calibration study and the 2026 macro-markets paper, and the wash-trading caveat on volume-weighted curves; the pack now runs a day-ahead calibration on both venues every week.
  • Live figures and a Vi IQ for this subject now refresh every ten minutes on this page from outside sources and Vultax's own tables; a context section, the questions below and these revision notes were added.

Sources and evidence

Vultax articles are market intelligence and education, not investment advice. This is a description of a backtested calibration result with explicit limitations; it is not a claim of live tradeable edge, and no market flagged by the model is a recommendation.

See the live data behind this article.

Order books, whale flow, Vi IQ scores and arbitrage routes from the same venues, in the terminal. Free, no signup.

New studies, by email.

Measured, sourced, and sent once per study. No digests, no promotions.

One email when a study publishes. Unsubscribe anytime. Privacy