We Tested Polymarket Prices Against 6,776 Resolved Markets. Here's Where They Bend.
Original Vultax research: a calibrated fair-value model beat raw Polymarket prices by ~9.6% on out-of-sample Brier score. Where the favorite-longshot effect lives, what the model can honestly flag, and why it still isn't a trading edge.
Live now — refreshed every 10 minutes
loading…
- Brier, Polymarket day-ahead prices, 7 d
- –
- Markets in that sample
- –
- Brier, Kalshi day-ahead prices, 7 d
- –
- Markets in that sample
- –

- Validation sample
6,776 markets
Resolved Yes/No markets with known outcomes.
- Brier improvement
≈ +9.6%
Out-of-sample, 0.0352 → 0.0319 versus raw market prices.
- Serving posture
Watch context
notRecommendation=true on every served fair-value row.
The question a fair-value model must answer
Prediction-market prices are good probability estimates on average — that is the standard result, and it held in our data. But 'good on average' leaves room for systematic, repeatable distortion, and the best-documented one is the favorite-longshot effect: low-probability outcomes trade above their true frequency and near-certain outcomes trade slightly below it.
A fair-value model, honestly defined, answers one question: given where this market is priced, where did markets priced like it historically resolve? That is a calibration question, and calibration questions have a standard scoring rule.
How to read the live figures
The strip at the top of this page is rewritten every ten minutes from Polymarket's resolved markets of the last seven days with the CLOB's yes price 24 hours before each market's end, and Kalshi's settled markets with the candlestick close 24 hours before settlement, the same rule on both venues, volume-weighted, markets under a day old skipped. The Vi IQ beneath it scores the six domains the Vultax terminal scores for a pair, computed for this page's subject from published inputs, with a domain that cannot be computed shown as unavailable rather than filled in. The figures in the body of the article are the measurement it was written on and are dated; the strip is what the same measurement says now.
The Brier score is the mean squared error between the day-ahead price and the outcome, so lower is better and 0.25 is a coin flip; the sample sizes beside them say how much weight to give each. The pack behind the page breaks each venue into price buckets so the favourite-longshot bend described above can be seen forming week by week.
Context from elsewhere
The published evidence on calibration has grown quickly. Kalshi's own August 2026 study over more than 2.2 million data points found no systematic bias in aggregate, forecasts improving with time, participation and volume, and markets with about $10,000 of volume or twenty traders already informative; its offered yardstick is the 0.10–0.15 Brier scores of skilled superforecasters. A 2026 academic study of 594 Kalshi macro markets reported an overall Brier of 0.0987, with federal-funds-rate contracts near-perfect (0.0001) and unemployment-rate contracts weakest (0.1302), and the one-hour horizon far better than the thirty-day one (0.0812 against 0.1587). Horizon, more than venue, decides the number.
Volume-weighted calibration inherits a problem this site has written about elsewhere. Columbia researchers flagged about a quarter of Polymarket's volume over three years as wash trading, peaking near 60% in December 2024 and highest in sports markets; a calibration curve weighted by volume gives those prints a vote. The weekly test on this page weights by volume for comparability and states the floor it uses; the bend it finds at the extremes should be read with that in mind.
Method and result
We assembled resolved Yes/No Polymarket markets with known outcomes — 6,776 after filtering for resolvability — and fit a calibration mapping from market price to resolved frequency, evaluating strictly out of sample. The scoring rule is the Brier score: mean squared error between stated probability and outcome, where lower is better.
Raw market prices scored approximately 0.0352. The calibrated mapping scored approximately 0.0319 — an improvement of roughly 9.6%. The improvement concentrates where theory says it should: at the probability extremes, where longshot enthusiasm and near-certainty discounting are strongest.
- Sample: 6,776 resolved Yes/No markets.
- Score: out-of-sample Brier, 0.0352 → 0.0319.
- Improvement: ≈ 9.6% versus raw prices.
- Effect concentrated near the 0 and 1 extremes.
Why this is not a trading edge
The distance between a calibration result and a tradeable edge is wide, and it is worth walking through it explicitly. First, the improvement is an aggregate: it says the population of prices bends in a known way, not that any specific market is wrong today. Second, the largest dislocations sit at the extremes, where bid-ask spreads are widest relative to the gap — a two-cent theoretical dislocation inside a three-cent spread is not an opportunity, it is noise with good posture.
Third, capture requires timing and settlement: the market can stay 'dislocated' until resolution, and the honest benchmark for any live-edge claim is close-line value — did flagged prices systematically beat the closing line? That measurement is future work on our roadmap, and until it exists, the product treats every fair-value flag as watch context. Every served row carries notRecommendation=true, and this article is bound by the same rule.
How to use the fair-value watch honestly
Used as designed, the fair-value watch is an attention allocator. A market flagged outside its calibrated band is a market worth understanding: check the liquidity, check the event room, check who is trading it and whether venues agree. Sometimes the price is stale; sometimes the flag is thin-category noise; sometimes the market simply knows something the base rates do not.
That triage — flag, then investigate structure — is the entire honest use of the model. Anyone promising more from a calibration curve is selling past their evidence.
Questions this page answers
- How accurate are prediction markets?
- Well calibrated on liquid contracts near resolution: the venues' own studies and academic work put Brier scores in the 0.02–0.10 range close to settlement, against 0.25 for a coin flip. A month out the same markets are much less informative, and prices at the extremes bend away from the outcome frequencies.
- What is a Brier score?
- The mean squared difference between a probability and the outcome, 0 for a perfect forecaster and 0.25 for one who always says 50%. The live strip shows the day-ahead Brier score on both venues for the last seven days.
- What is the favourite-longshot bias?
- The tendency for low-priced contracts to win less often than their price implies and high-priced ones more often. The article measures it on 6,776 resolved markets; the weekly pack checks whether it persists.
- Is a mispriced market a trading edge?
- Not on its own. The bend is small, fees and spreads sit on top of it, and it is measured after the fact across thousands of markets. The section above on why this is not an edge is the durable part of this page.
Revision notes
This page is kept current. Each entry records what changed and when; the live figures above refresh on their own.
- Added the August 2026 Kalshi calibration study and the 2026 macro-markets paper, and the wash-trading caveat on volume-weighted curves; the pack now runs a day-ahead calibration on both venues every week.
- Live figures and a Vi IQ for this subject now refresh every ten minutes on this page from outside sources and Vultax's own tables; a context section, the questions below and these revision notes were added.
Sources and evidence
- Vultax prediction desk (preview)
Where fair-value watch context is served.
- Brier score (background)
The proper scoring rule used for validation.
- Favorite-longshot bias (background)
The market-structure effect the calibration captures.
- Kalshi Research — Prediction market calibration: how accurate is Kalshi? (20 Aug 2026)
The venue's own study over more than 2.2 million data points: forecasts improve with time, participation and volume; markets with about $10,000 or twenty traders were already informative; the comparison offered is superforecaster Brier scores of 0.10–0.15.
- Krause — Calibration and Forecast Accuracy of Macroeconomic Prediction Markets: Evidence from Kalshi, 2025–2026 (SSRN)
594 Kalshi macro markets, 1,860 observations, nine series: overall Brier 0.0987; federal-funds-rate markets near-perfectly calibrated (0.0001); unemployment-rate markets the weakest (0.1302); one-hour horizon 0.0812 against 0.1587 at thirty days, as summarised in the abstract.
- Gizmodo — Study finds around a quarter of Polymarket trades are fake (Columbia, Nov 2025)
About 25% of Polymarket volume over three years flagged as wash trading by counterparty structure, peaking near 60% in December 2024; roughly one wallet in seven flagged; many flagged wallets made no profit.
- CoinDesk — Polymarket's trading volume may be 25% fake, Columbia study finds
Sports markets carried the highest flagged share (about 45% of historical volume); election markets about 17%.
Vultax articles are market intelligence and education, not investment advice. This is a description of a backtested calibration result with explicit limitations; it is not a claim of live tradeable edge, and no market flagged by the model is a recommendation.
See the live data behind this article.
Order books, whale flow, Vi IQ scores and arbitrage routes from the same venues, in the terminal. Free, no signup.
New studies, by email.
Measured, sourced, and sent once per study. No digests, no promotions.
Continue reading
Best Polymarket Tools in 2026: 22 Tools Compared, 4 Publish Their Latency
22 Polymarket tools scored on stated latency, P&L method, history depth, liquidity context and price, from each vendor's own pages on 5 September 2026. Only four state how fast their alerts are, and no leaderboard says whether its P&L is net of fees.
Esports Out-Traded the Fed on Polymarket Today: 21.4% vs 20.5% of Volume, and Five Video Games in Our Top Five
On 6 September 2026 esports was the largest category among Polymarket's hundred busiest markets, ahead of the September Fed decision, and the five biggest markets in Vultax's own fills were all League of Legends and Counter-Strike playoffs. Who trades them, hour by hour.
The Fed Is a Coin Flip on Polymarket: 49.5% Hike, 49.5% Hold, and Kalshi Agrees to Half a Cent
Ten days before the 16 September FOMC decision, Polymarket prices a 25 bp hike and no change at 49.5% each on $99 million of volume; Kalshi's contracts closed at 49 and 49. Thirty days of both venues, the gap between them, and the 5,120 wallets trading it.