Sabitlenmiş Tweet

I spent 50 days running a public, audited test of a simple question: can frontier AI beat prediction-market crowds at forecasting? I locked every call before resolution, scored it against the market, and published all of it, wins and losses. The answer was clear. 🧵
2/ Setup: three frontier models (Claude, Grok, Gemini) called live Polymarket markets daily. Every forecast locked before the event resolved, then Brier-scored against the crowd's price at the moment of the call. Immutable record, public corrections, nothing edited after the fact.
3/ Result: the crowd beat all three models. On the head-to-head set (n=174 resolved markets), the market's Brier was ~0.107 vs the best model's ~0.126. Not one model, not one category, the crowd led everywhere I looked.
4/ The sharper finding is when AI failed. When the models agreed with the crowd, it was a tie. The more confidently a model diverged from the market, the worse it did monotonically. The boldest "AI disagrees with the crowd" calls were the worst calls.
5/ Translation: when frontier AI most confidently thinks the market is wrong, the market is most reliably right. The crowd isn't just good, it's specifically good at the exact moments you'd hope an AI could beat it.
6/ This lines up with the Fed's own Feb 2026 paper finding prediction-market prices forecast rate moves as well as or better than Bloomberg consensus and fed-funds futures. Different method, same direction: these markets are hard to beat.
7/ One thing I did NOT test: thin, illiquid markets. Every market here was liquid (≥$100k volume), the crowd's home turf. Whether an edge exists in the long tail is genuinely open. I just know it isn't in the liquid markets where most volume lives.
8/ Why publish a negative result? Because "we have AI alpha" is everywhere and almost never audited. I built the audited version, and it said no. That's worth more on the record than another unverifiable claim.
9/ The full record is public and permanent...every call, every outcome, every miss, all of it: [emberfyi.com/calls]. I'm sharing the methodology and data openly. If you work in forecasting, eval, or prediction markets and want the dataset, reach out.
English















