Are Prediction Markets Accurate? What the Evidence Actually Shows

The Short Answer
Are prediction markets accurate? In aggregate, mostly yes: across large samples they are reasonably well calibrated, and the best long-run study found market prices closer to the final election result than individual polls 74 percent of the time. But accuracy is conditional. It degrades on long-dated contracts, thin markets, rare events and the closing minutes of sports contracts. And the ugliest episodes of 2025 and 2026 turned on resolution rules rather than on forecasting.
People search "are prediction markets accurate" for one of two reasons: they just read that markets called an election the pollsters missed, or they just watched a market price a real event at 1 percent hours before it happened. Both things happened, and neither settles the question, because a probability is not a prediction.
Key Takeaways
- Accuracy here is only measurable in aggregate. A market that says 20 percent should be wrong one time in five.
- Page and Clemen found favourite/longshot bias growing with time to expiry in 2013. A 2026 preprint on 353 million Kalshi and Polymarket trades found a separate, persistent compression toward 50 percent in political markets. Both point at underconfidence, but they are not measuring the same conditioner.
- The loudest disputes of the past two years came out of resolution rules rather than forecasting errors, and the CFTC-regulated exchange was not spared.
What Accuracy Even Means for a Probability
The most common mistake here is treating a price as a prediction. If a market says 80 percent and the thing does not happen, that is not a failure, it is the expected outcome one time in five. Kalshi and Polymarket both priced Robert Prevost at roughly 1 to 2 percent before he became Pope Leo XIV in May 2025. That proves miscalibration only if events those markets price at 1 percent happen far more often than 1 percent of the time, across hundreds of markets.
So accuracy gets measured over large samples, using two standard tools. A calibration curve sorts every price into buckets and asks how often the events in each bucket actually happened, then plots that against the price; a perfectly calibrated venue traces the 45-degree diagonal. A Brier score is the mean squared error between the stated probability and the 0 or 1 outcome, so lower is better, and answering 50 percent to everything scores 0.25. Almost none of the venue-level research below reports Brier scores; it works in calibration curves. Worth knowing if you go hunting for a headline Brier number per platform. We have not found one published against a stated sample, so treat any figure quoted without that sample as decoration.
Are Prediction Markets Accurate in Aggregate? The Calibration Evidence
Start at the peer-reviewed end of the literature, which as of August 2026 is thinner than the citation traffic suggests. Page and Clemen, in the Economic Journal (2013), found prediction markets "reasonably well calibrated when time to expiration is relatively short, but prices are significantly biased for events farther in the future". The bias runs in a favourite/longshot direction: high-likelihood events underpriced, low-likelihood events overpriced.
Thirteen years later, a 2026 preprint (not peer reviewed) covering 353 million trades across roughly 429,000 binary contracts on Kalshi and Polymarket reported "persistent underconfidence in political markets, where prices compress toward 50%", on both platforms, with the compression strongest on large Kalshi trades. Note the difference from Page and Clemen: the conditioner there is time to expiry, here it is trade size. Both point at underconfidence without measuring the same thing. The preprint's own summary is the part worth keeping: "calibration is therefore conditional: a price's meaning depends on what, when and how much is traded."
Sports contracts, which by 2026 made up the bulk of Kalshi's traded volume, add a wrinkle. That mix swings hard with the sports calendar and the election cycle, so check the venue's current figures before repeating it. A 2026 preprint (not peer reviewed) on 23 million Kalshi moneyline trades, meaning straight win-or-lose contracts on one team, found the calibration curve sitting on the ideal diagonal through the middle of a game and then departing sharply near expiry, turning step-like in the final ten minutes. The paper reads that shape as consistent with insurance demand from traders holding losing positions. If that reading is right, the practical consequence is that a nearly dead contract trading at two or three cents in the last minutes is dearer than its real chance, so the cheap side of a late sports market is the one to avoid. The same work found cross-game parlays, single contracts that pay only if every leg wins, priced above the product of their own legs even when each leg was well calibrated, which is one more reason the sports betting comparison is worth reading.
Prediction Markets vs Polls: Accuracy, and the Caveat Nobody Quotes
The canonical result comes from Berg, Nelson and Rietz (International Journal of Forecasting, 2008), who compared Iowa Electronic Markets prices with 964 polls across the 1988 to 2004 presidential elections. The market landed closer to the eventual two-party vote split 74 percent of the time, with mean absolute error of 1.82 percentage points against 3.37 for polls.
Now the caveat that never survives the retelling. That comparison is against raw, unadjusted individual polls, chosen deliberately because most settings lack the history needed to model adjustments. Beating one raw poll is a much weaker claim than beating a modern poll average, and the comparison a reader actually wants has not been run at that scale: we have found no large-sample study testing markets against present-day poll aggregates. That absence is itself an answer. Treat rough parity as the honest default, and file "markets beat the averages" under untested.
The counter-evidence is real too. A 2019 study by Dana, Atanasov, Tetlock and Mellers, published in Judgment and Decision Making, paired market orders in an IARPA forecasting tournament with each trader's own self-reported probability. Aggregating those self-reports was at least as accurate as the prices, and combining the two beat prices alone. The self-report edge on its own was not statistically significant, so polls did not win. The combination result was, which means the market failed to squeeze the information out of its own traders.
The Documented Failures, and What Caused Them
| Case | When | What markets showed | Underlying cause |
|---|---|---|---|
| Brexit referendum | Jun 2016 | Remain a heavy favourite | Wealth-weighted odds |
| Papal conclave | May 2025 | Leo XIV at 1 to 2 percent | Rare, opaque process |
| Polymarket flu and measles forecasts | 2025 to 2026 | Beaten by FluSight ensemble and baselines | Thin volume, impossible outcomes priced |
| Ukraine mineral deal | Mar 2025 | Resolved Yes with no deal | Oracle vote capture |
| Kalshi Khamenei contract | Feb to Mar 2026 | Settled at last traded price, not $1 | Disputed contract carveout |
Only the first three are forecasting failures. The disease work (a 2026 preprint, not peer reviewed) is the most damning of them: on flu, Polymarket was "dominated by the FluSight ensemble", the CDC's expert-curated multi-team forecast, and the best combination of the two gave the market zero weight; on measles it lost to simple statistical baselines. The diagnosed causes were low trading volume and probability mass sitting on arithmetically impossible outcomes. That study looked at Polymarket only, so it says nothing either way about Kalshi. Conclave markets, as Dartmouth economist Eric Zitzewitz put it, "only get a data point every decade or two".
Brexit is the row people quote most and the one to quote most carefully. Reported Remain probabilities vary so much between bookmakers, exchanges and the hour they were sampled that we will not put a number on any of them, and the spread between those sources is itself part of the lesson. What is well documented is the mechanism. Ladbrokes stake data reported by the Financial Times had the average Remain bet running roughly six times the size of the average Leave bet, while Leave drew more bets by count. That is bookmaker money rather than exchange order flow, but the lesson carries: prices are wealth-weighted rather than person-weighted.
The last two rows are not forecasting failures at all. In both, the price was defensible and the settlement was the problem. A roughly $7 million Polymarket contract on a Ukraine minerals deal resolved Yes with no deal signed, after what Polymarket called an "unprecedented" governance attack on its oracle. Polymarket settles disputed markets through a third-party system in which governance token holders vote on the outcome, rather than through the exchange itself, so anyone who accumulates enough voting weight can force a settlement the facts do not support. Users were not refunded.
Kalshi, the CFTC-regulated venue, settled a contract on Ayatollah Khamenei at its last traded price rather than at $1, invoking a death carveout in the contract terms. The $54 million figure attached to that market in coverage is cumulative trading volume, not the size of the disputed payout. Kalshi says the carveout was in the published rules from the start, and that it refunded trading fees and covered net losses so that no trader finished down. A class action disputes the first half of that, alleging the carveout was not in the rules summary users actually saw, and the case was unresolved as of August 2026. Whichever way it lands, the practical lesson is to read a venue's full contract terms rather than the summary card. Resolution risk is not a crypto problem.
Are Polymarket Odds Accurate? The 2024 Test Case
2024 is where "markets beat polls" became folk wisdom. A 2025 preprint (not peer reviewed) put Polymarket's mean Trump win probability above 55 percent from mid-October, peaking near 67, while the polling series it compared against hovered in the 40 to 50 percent band and never firmly favoured either candidate. Those two series are not the same quantity, and the paper does not resolve that, which limits how much the comparison can carry. Its authors flagged other limits: one market against aggregated polls, an unrepresentative trader base, and Michigan and Wisconsin results where polling "may be potentially superior".
The counterweight gets less coverage than it deserves. Clinton and Huang, in a working paper (not peer reviewed), examined more than 2,500 political markets across the Iowa Electronic Markets, Kalshi, PredictIt and Polymarket in the campaign's final five weeks, over $2 billion in transactions, and found "little evidence of efficiency": identical contracts priced differently across exchanges, arbitrage peaking in the last fortnight. The same paper reports the share of markets on each venue that predicted outcomes better than chance, at 93 percent on PredictIt, 78 percent on Kalshi and 67 percent on Polymarket. Better than chance is a low bar and is not the same test as calibration, so read the ordering rather than the levels. The ordering is awkward for the received story, because Polymarket, usually cited as 2024's success story, came last. Our Polymarket vs Kalshi comparison covers the rest.
One widely repeated claim about 2024 does not survive contact with the data. An on-chain transaction analysis of the French trader known as Theo, whose huge October position on Trump was reported everywhere as having bent Polymarket's odds, found capital flowing into both sides of the market at once, "consistent with heterogeneous-beliefs trading rather than one-sided manipulation". The same work found that naive aggregation overstated that market's October turnover by roughly two and a half times once share minting and burning were stripped out. Minting creates paired Yes and No shares at once, which inflates reported volume without anyone taking a directional position. That correction is specific to one market in one month, so it is not a discount you can apply to Polymarket's reported volume in general.
How to Judge Whether a Specific Market Is Trustworthy
The aggregate findings above only become useful when you point them at the market open in your other tab. Six checks, in the order they are worth doing:
- Resting depth, not headline volume. Open the order book and read how much size is sitting within a cent or two of the current price on each side. Reported volume includes minting and churn; resting depth is what your order actually meets, and it is what tells you whether the price reflects real opinion or two people and a market maker. Thin books are the named cause in the disease and conclave failures above.
- Time to resolution. The peer-reviewed finding is that bias grows with distance from expiry, so the same question resolving next week deserves more trust than one resolving next year. No study publishes a clean cutoff, so treat it as a gradient rather than a threshold: the further out the date, the more you should assume the price is pulled toward 50 and the less you should read into a move of a few cents.
- Domain. Recurring, publicly scored events calibrate well, because the crowd has a track record to learn from: elections, macro releases, scheduled sport. Rare closed-door processes have none. A conclave, a leadership succession or a one-off court ruling is the category where the crowd has nothing to be calibrated against.
- The resolution rules, read like a contract. Three clauses do most of the damage: the named source of truth, any carveout that changes what counts as Yes, and the conditions under which the venue can settle early. The Khamenei market is the worked example, where a carveout rather than the event decided the payout. Read the full contract terms, not the one-line summary on the market page.
- Cross-platform divergence. Look the same question up on the other venue. A gap of a cent or two is ordinary friction. A wide gap is usually a rules mismatch, a different resolution source or a different deadline, so read both rulebooks before you treat it as an arbitrage opportunity.
- Product type. Two shapes price worst in the research: multi-leg parlays, which sit above the product of their own legs, and near-dead longshots in the closing minutes of a sports contract.
Checking calibration yourself beats taking anyone's word for it, ours included, and the method is short enough to describe in full. Both major venues publish market data through public APIs, and it is the settled markets you want: pull the closing price and the actual outcome for a few hundred resolved contracts in one category, sort them into ten buckets by price, then compare each bucket's average price with the share of contracts in it that resolved Yes. Well calibrated comes out close to a straight line. Compression toward 50 shows up as the low buckets resolving Yes less often than their prices implied and the high buckets more often. Our Kalshi API guide and Polymarket API guide cover the endpoints. Keeping those histories current across venues is the problem we work on at Predictefy, along with the prediction market data tools around it. None of this is financial, legal or tax advice, and calibration in aggregate says nothing about whether you finish ahead after fees.
Frequently Asked Questions
Are prediction markets accurate?
In aggregate, reasonably. Large studies find prices well calibrated through the middle of a contract's life: outcomes priced at 70 percent happen roughly 70 percent of the time. Calibration weakens with distance from resolution, on thin order books and on rare one-off events, and no single outcome can prove a market right or wrong, because a 20 percent event is supposed to happen one time in five.
How accurate are prediction markets?
Accurate enough to be useful, not accurate enough to quote as fact. The benchmark study of 964 polls across five presidential elections found Iowa Electronic Markets prices closer to the final vote 74 percent of the time, with mean absolute error of 1.82 percentage points against 3.37 for polls. Newer work on Kalshi and Polymarket trades finds accuracy depends on how much is traded, how far off resolution is and what the contract is about.
Are prediction markets more accurate than polls?
Often, but the famous comparisons are against single raw polls rather than modern poll averages, which the original authors stated explicitly. No large-sample study has run markets against present-day poll aggregates, so rough parity is the honest default rather than market superiority. A 2019 study also found that asking traders for their own probability estimates and averaging them was at least as accurate as the market price.
Are Polymarket odds accurate?
They were directionally right in 2024, sitting above 55 percent for Trump from mid-October while polling averages showed something near a coin flip. But a working paper covering more than 2,500 political markets in that campaign's final five weeks found little evidence of efficiency, including identical contracts priced differently across exchanges, and Polymarket ranked last of the four venues studied on the share of markets that beat chance.
What does prediction market accuracy research show?
Two decades of it converge on one point: calibration is conditional. Page and Clemen found favourite/longshot bias growing with time to expiry in the Economic Journal in 2013, and a 2026 preprint (not peer reviewed) covering 353 million trades found political prices compressing toward 50 percent. Polymarket's disease markets lost to the CDC FluSight ensemble and to simple baselines where volume was low. Most of the recent work is unrefereed preprints, so weigh it accordingly.
Conclusion
Accurate enough that ignoring them is silly, conditional enough that quoting a single price as fact is sillier. What you are reading is a wealth-weighted clearing price, net of fees and the cost of capital locked up until resolution. Its reliability depends on resting depth, time to resolution and how carefully the rules were written, so check those three before you check the price. Applying this to November? Our guide to 2026 midterm election odds is the place to start, with the caveat that nobody can publish accuracy results for an election that has not happened yet.