Do weather models beat Kalshi's temperature markets? Two weeks of forward testing say no
Seven strategies, two venues, 504 settled Kalshi markets and 660 on Polymarket. Ensemble divergence loses at every threshold and loses worse the more confident the models are. The one idea that looked real died in forward testing.
Short answer: no. Not with ensemble divergence, not by fading it, not by city, not by time of day, not by entering at market open, not on Polymarket, and not by quoting instead of crossing. Seven strategies, two venues, 504 settled Kalshi markets and 660 on Polymarket, judged on live forward tests rather than backtests alone. Every one of them loses.
This page exists because the opposite is widely claimed and rarely tested.
The setup
The idea is reasonable, which is why so many people repeat it. An ensemble forecast, 82 members across GFS and ECMWF, gives a probability distribution for the day's high temperature. When that implied probability diverges from the market price by enough, trade the side the models favour.
The lab ran it from 16 to 30 July 2026: 239,388 quote snapshots, 504 settled Kalshi markets, 660 settled Polymarket markets. Nothing moved real money. Entries were taken at the price the book was showing and charged the fee Kalshi would actually charge.
It loses, and it loses worse the more confident the models are
Deciding at 10am ET, held to settlement:
| Divergence | Trades | Win rate | Net per trade |
|---|---|---|---|
| 5pp or more | 329 | 33% | −3.5¢ |
| 10pp or more | 254 | 31% | −5.6¢ |
| 15pp or more | 198 | 32% | −4.8¢ |
| 20pp or more | 146 | 28% | −5.6¢ |
| 25pp or more | 111 | 28% | −5.1¢ |
Read the first column against the last. The more strongly the models disagreed with the market, the worse the trade did. That is the opposite of what an edge looks like.
The mechanism: the ensemble is honest about long shots and overconfident about favourites
| Ensemble said | Markets | Actually happened | |
|---|---|---|---|
| 0–9% | 244 | 9.8% | honest |
| 10–19% | 88 | 17.0% | honest |
| 20–29% | 62 | 27.4% | honest |
| 30–39% | 50 | 32.0% | honest |
| 40–49% | 32 | 18.8% | overconfident |
| 50–59% | 16 | 31.2% | overconfident |
| 60% or more | 12 | roughly 0–20% | badly overconfident |
The models are well calibrated right up until they become confident, and then they fall apart. Every losing trade is the lab buying one of those confident calls.
The market prices that asymmetry correctly. That is what an efficient market looks like from the inside: not that the forecast is useless, but that everything useful in it is already in the price, including its known failure mode.
The other six
Fade it instead. If the models are overconfident, bet against them. Roughly break even before fees and negative after. Winning 63% of fades still loses, because fading an overconfident favourite means buying an expensive one.
Segment by city or direction. No durable pocket. New York cleared at +3.9¢ over 20 trades, which is a hot streak rather than a finding. A promising Chicago-good, Miami-bad split did not survive the full sample.
Time of day. At honest sample sizes, 77 to 94 trades at 10am and 2pm, the market's favourite wins exactly as often as its price implies. A tempting +7.9¢ at 6pm was 34 trades at a 100% win rate, which is noise wearing a result's clothing.
Enter at market open instead of same day. This is the one that looked real, and it is the most useful thing in the study. See below.
Polymarket. Not softer. Same rule, 998 settled trades, −$118, every tier between −10¢ and −13¢. Higher win rates than Kalshi and bigger losses per loss.
Market making. A rebate-dependent bleed. Even modelled as a completely free maker it loses, and about −0.4¢ per fill survives even at the tightest inventory discipline. That residual is adverse selection: on a market this slow, the person lifting your quote usually knows something you do not. Whether a liquidity rebate covers it is not a question paper can answer, and chasing a rebate is not the same thing as finding an edge.
The mirage, and why the forward test is the whole method
The lead-time idea looked like the real thing. Enter at market open rather than same day, at a 10pp divergence:
| Sample | Net per trade |
|---|---|
| Backtest, 55 trades | +4.6¢ |
| Backtest, 72 trades | +2.2¢ |
| Backtest, 216 trades, full | −2.8¢ |
| Live forward test, 123 trades | −11.5¢ |
The edge shrank as the sample grew, crossed zero, and then the forward test came in at −11.5¢ and stayed negative at every tier. Nothing was wrong with the arithmetic. The first 55 trades were simply a small sample that happened to look good, and every additional trade dragged it back toward the truth.
The lab keeps a standing rule out of this: when a backtest and a forward test disagree, the forward test wins. Anyone claiming a weather edge from a backtest alone has skipped the step that kills it.
Two facts worth having even if you never trade this
Kalshi and Polymarket resolve on different weather stations. LaGuardia against Central Park for New York, O'Hare against Midway for Chicago. The two venues can disagree and both settle correctly. Anyone treating a price gap between them as free money is buying resolution risk and calling it arbitrage.
The advice to trade these 3 to 5 days out cannot be followed. Kalshi lists daily temperature markets roughly 1 to 1.5 days ahead. There is no 3 day market to trade. The entire tradeable window is about a day and a half, and its best point, market open, is efficient.
What would have to be true for an edge to exist
Something has to be wrong with the price, and the ensemble is not it, because the market already knows where the ensemble breaks.
The one mechanism that reliably produces an edge in a prediction market is being faster than everyone else, and daily weather markets move too slowly for speed to be worth anything. That thesis belongs in markets that settle in minutes, not in a contract that resolves tomorrow afternoon.
Experiment closed 30 July 2026. The loggers are off and the numbers above are final.
The lab's full board publishes every Sunday. The Seats lists what each strategy does.