How Accurate Are Prediction Markets? We Ran the Numbers
The famous "94% accurate" stat measures almost nothing. We survey what calibration studies actually show about prediction markets — where they beat polls, where they break — and commit to a public audit through the 2026 midterms.

Liquid prediction markets are genuinely good probability forecasters — and almost every viral claim about how good measures the wrong thing. The honest summary of two decades of published evidence: large markets are well calibrated near resolution, meaning events they price at 70% happen roughly 70% of the time; they beat polls most clearly months out, not on election eve; and their accuracy decays in predictable places — long horizons, thin order books, and longshots. The “94% accurate” figure that circulates in headlines captures almost none of that. Here is where that number actually comes from, what the academic record supports, and the public calibration audit EventMarkets is launching through the 2026 midterms — so the next accuracy claim you read can be checked against data nobody’s marketing team touched.
What should “accurate” mean for a probability?
A prediction market never says an event will happen. It quotes a price, and the price is a probability: a contract at 70¢ claims the event happens 70% of the time — which is equally a claim that it fails 30% of the time. If 70¢ contracts never lost, the market would be broken, because they should have been trading at 99¢.
A 70-cent contract that loses three times out of ten isn’t failing — it’s doing exactly what it said it would.
So grading any single market “right” or “wrong” is a category error. The real test is calibration: gather every contract that traded at a given price at some fixed point before resolution and check the frequencies. If about 70% of the 70¢ contracts resolved yes, about 40% of the 40¢ contracts, and so on down the line, the market is calibrated. Plot that across all price buckets and you have a calibration curve — the honest report card.
Calibration alone isn’t sufficient, because a forecaster who answers 50% to everything is perfectly calibrated and perfectly useless. The standard fix is the Brier score: the average squared gap between the forecast probability and the outcome (1 or 0). Zero is perfect. Permanently shrugging “50-50” scores 0.25. Anything meaningfully below 0.25 is real signal. Keep both benchmarks in mind — they do a lot of work below.
Where does the “94% accurate” number come from?
Not from Polymarket, for a start. The figure traces to an independent analysis by data scientist Alex McCullough, who built a public Dune dashboard of resolved Polymarket markets and found the market-favored outcome came true roughly 90% of the time a month before resolution — and about 94% of the time four hours before, as reported by CoinDesk in March 2025. The number is real. As a measure of forecasting skill, it is close to empty.
Two reasons. First, it’s a hit rate, not a calibration test. Four hours before resolution, most markets have already converged toward 5¢ or 95¢; scoring those “correct” is grading questions that were effectively settled. Second, a hit rate throws away the probabilities that make markets useful. It counts a 60¢ favorite that loses as a miss, even though the market explicitly said it would lose 40% of the time — and it counts a 99¢ near-certainty that fails as no worse than a 51¢ coin flip that does, when the first is a catastrophically larger error.
The platforms are starting to publish better numbers. In August 2026, Kalshi released a study by its own researchers covering 2.3 million bets from its 2021 launch through mid-2026, scored properly in Brier terms: about 0.02 just before markets close and roughly 0.09 three months out, per Semafor’s report, which also notes the accuracy was strongest in hard-to-manipulate categories like elections and economic data and weaker in sports. Against the 0.25 coin-flip benchmark, those are strong results — and the three-month figure is a genuinely useful disclosure. The caveat writes itself: it is Kalshi grading Kalshi’s homework. As of August 2026, Polymarket has published no equivalent audit of its own markets; every accuracy number attached to it is third-party.
What does the academic record actually show?
The peer-reviewed evidence is older, slower, and more useful than any platform stat. The load-bearing studies:
| Study | Data | Headline finding |
|---|---|---|
| Berg, Nelson & Rietz (2008) | Iowa Electronic Markets vs. 964 polls, 1988–2004 | Market closer to the outcome 74% of the time; clearest edge beyond 100 days out |
| Erikson & Wlezien (2008) | Same Iowa markets, reanalyzed | Properly projected polls beat the market |
| Rothschild (2009) | 2008 state races: Intrade vs. debiased polls | Debiased market forecasts win early and in uncertain races |
| Page & Clemen (2013) | Real-money markets across event types | Well calibrated near expiration; biased at long horizons |
| Clinton & Huang (working paper) | 2,500+ markets, $2.4B traded, 2024 election | Better than chance, but prices diverged across platforms and herded |
The founding result is the Iowa Electronic Markets — a small, real-money academic exchange running since 1988. Berg, Nelson and Rietz compared its presidential vote-share prices against 964 national polls across five elections and found the market closer to the final result 74% of the time, with the advantage concentrated at long range: more than 100 days out, the market beat the polls in every election studied.
The sharpest rebuttal came from Erikson and Wlezien, who argued raw polls are the wrong opponent. Polls measure today’s opinion, not November’s; project them forward properly and the projections beat the market. Rothschild’s follow-up, comparing debiased Intrade prices against debiased poll aggregates across 2008’s state races, tilted back toward markets — especially early, and in races that weren’t foregone conclusions. A fair reading of all three: the market’s edge is real at long horizons and shrinks toward a tie against sophisticated poll aggregation as election day approaches.
On calibration itself, Page and Clemen found prices well calibrated when expiration is close but significantly biased for events months away — and biased in a specific direction: the favorite-longshot bias, documented in betting markets for decades, in which rare events are overpriced and near-certainties underpriced. Practical translation: a 5¢ contract on a longshot months from resolution usually overstates the true chance.
Where do markets beat polls — and where don’t they?
Markets win on speed and breadth. They update in minutes, not field-and-weight cycles, and they aggregate anything traders know — early votes, court rulings, weather — not just survey answers. Brexit night 2016 is the canonical case, and it cuts both ways: when UK polls closed, Betfair’s exchange implied roughly a 90% chance of Remain, and Remain lost. Yet Auld and Linton’s study of that night found something humbling: both the betting market and currency traders were strikingly slow to digest the incoming county results, taking hours to converge on what the counts were showing. (The study initially credited the betting market with about an hour’s head start over the pros; a later corrigendum traced that to a timestamp error — the corrected data shows both moved on broadly the same clock.) Markets were wrong at the top of the night and slow to get right. Speed is real, but it is not clairvoyance.
The 2024 US election is the modern talking point: Polymarket and Kalshi carried Donald Trump as a clear favorite while major polling averages sat near even, and he won. But one election is one draw from the distribution — a 60% favorite winning tells you almost nothing about calibration. And the detailed autopsy is sobering: Clinton and Huang examined more than 2,500 political markets and $2.4 billion in trading across the final five weeks and found identical contracts priced differently on different exchanges, daily moves that reversed the next day (herding, not news), and hit rates that varied sharply by venue — highest on small, position-capped PredictIt, lower on the big-money platforms. Record volume bought speed. It did not buy efficiency.
How do prediction markets fail?
Thin markets. Every calibration result above comes from liquid markets. A market with $2,000 of volume and a 15-point spread doesn’t inherit that record — its “price” is the last handshake between two strangers, and a modest order can move it ten points. The favorite-longshot bias is at its worst here, and long-horizon thin markets combine both failure modes at once.
Whales. Prices can be leaned on. Rothschild and Sethi’s transaction-level study of Intrade’s 2012 presidential market found a single trader who placed about a third of all bets on Mitt Romney over the final two weeks, lost at least $4 million, and held Romney’s price near a 30% floor while other markets drifted away. The market corrected eventually — at the whale’s expense — but “eventually” lasted two weeks in the most-watched political market of its era.
Resolution risk. A contract pays only if the resolution mechanism agrees the event happened. In July 2025, a Polymarket market on whether Ukraine’s president would wear a suit — with more than $160 million riding on it — resolved “No” even after Reuters and the BBC described his NATO summit outfit as a suit. The call went through the UMA oracle, and traders argued publicly that the oracle’s majority-vote incentives, not the facts, decided the outcome. However you read that dispute, the lesson stands: no probability is better than the rulebook that settles it.
The EventMarkets calibration audit
Surveying other people’s numbers only goes so far, and the best current disclosures are self-graded. So, as of August 2026, we are committing to an independent public audit through the midterms, with the rules published before any results exist:
- Universe: every Kalshi and Polymarket market tied to the November 3, 2026 midterms — Senate, House, and governors’ races — plus Fed-decision and CPI markets as a non-political control group.
- Snapshots: closing midpoint prices at fixed horizons — 90, 30, 7, and 1 days before resolution, plus 4 hours out, so the “94%” framing gets tested on its own terms.
- Scoring: calibration curves by price decile and Brier scores by horizon, platform, and category, with sample sizes shown for every cell.
- No survivorship games: the market list is frozen at each snapshot; nothing is dropped after the fact, and disputed resolutions are reported separately rather than quietly excluded.
- Publication: full results and the raw snapshot data, after the midterm markets resolve — whatever they say about the markets, or about this article.
The results will run here once the last midterm market settles.
Quick answers
Are prediction markets more accurate than polls? Months before an election, the published evidence says yes — the Iowa markets beat polls at horizons past 100 days in every election studied. Close to election day, sophisticated poll projections tie or win. Both “always” and “never” are wrong.
Is Polymarket really 94% accurate? That number comes from an independent dashboard, not Polymarket, and it measures the hit rate four hours before resolution — when most markets are effectively settled. It is not a calibration measure, and Polymarket has published no audit of its own as of August 2026.
A market said 80% and the event didn’t happen. Was it wrong? Not on that evidence alone — an honest 80% forecast should fail one time in five. Only a pattern across many markets, with 80¢ contracts winning well below 80% of the time, demonstrates miscalibration.
What is a good Brier score? Zero is perfect; 0.25 is what a permanent “50-50” answer scores. Kalshi’s self-published figure is roughly 0.09 three months out and 0.02 near close. Our audit will test comparable numbers independently.
What makes a specific market untrustworthy? Low volume, wide bid-ask spreads, distant resolution dates, cheap longshot prices, and ambiguous resolution rules. Any one is a caution flag; several together mean the price is noise.
When will EventMarkets publish its audit? After the November 3, 2026 midterm markets resolve, using the methodology frozen above — calibration curves, Brier scores, and the raw data.