Brier Cup · Knockout Analysis
Before kickoff, seven frontier AI models each locked a probability on who'd go through. Then Spain dismantled France and Argentina broke English hearts in stoppage time. The forecasts, scored, tell a strangely symmetrical story — the crowd was dead wrong about one match and dead right about the other.
Since these semis the roster has grown to ten labs — each lab's newest release joined for the final, with xAI and Mistral entering — and every agent now paper-trades its forecasts against the TxODDS line. As of the final's kickoff: The Market still leads the knockout board (0.312 avg Brier over 31 ties; best agent Claude Opus 4.7, 0.322), MiniMax's book leads the Bankroll, and and the machines called the final a near coin-flip, 52–48 Spain — Spain won it 1–0 in extra time. The lab standings below are live and current.
Claude Sonnet 4.5 was the only model on the right side of both ties.
Kimi K2 was the only model on the wrong side of both.
Every forecast landed inside this band. Nobody committed — the models knew these were coin-flips.
Six of seven models nudged toward France. Spain answered with a clean two-goal night, opening the scoring on 58'. As a group, the AIs did worse here than if they'd flipped a coin.
“a marginal edge to Spain's current form and momentum”
“France very slightly favored … close to a coin flip”
“France's deeper squad and recent edge in major tournament knockout stages”
“France boasts incredible squad depth and recent World Cup pedigree, [but] Spain's tactical cohesion and recent Euro 2024 victory make this a nearly even matchup”
Gemini joined the roster after this match, so it never locked a forecast here. This is a backtest — the identical pre-kickoff prompt, run after full time, no result in the prompt — so it is not in the field mean above. It leaned France by a single point; Spain went through.
Here the consensus held: six of seven leaned Argentina, and the reigning champions delivered — equalising on 85' after England led at 55', then winning it in the second minute of stoppage time. Only one model looked away.
“Argentina's … elite individual talent … over a strong but historically less clinical England side”
“England's deeper squad … give them a slight edge, but Argentina's Messi-led clutch knockout record keeps it almost a coin-flip”
“England is given a razor-thin edge as their elite core will be in its prime, while Argentina may be navigating a transition as its veteran leaders age out”
Same story: a backtest, not a locked forecast, so it is outside the field mean above. Gemini shaded England by a point — and Argentina broke through in stoppage time. On both these coin-flips it landed on the losing side.
Put the two calls side by side and the models sort themselves. Claude sided with both winners. Kimi backed both losers. Everyone else got the match the crowd read correctly — and missed the one it didn't.
| Model | France v Spain | England v Argentina |
|---|---|---|
| Spain | Argentina | |
| France | Argentina | |
| France | Argentina | |
| France | Argentina | |
| France | Argentina | |
| France | Argentina | |
| France | England | |
| France | England |
Cell shows the side each model favoured. Struck-through = eliminated. Gemini's row is a backtest — see below.
Gemini 3.1 Pro joined the Brier Cup after these two semi-finals, so on the boards above it has no locked forecast — only a backtest: the exact pre-kickoff prompt, replayed after full time, with no result anywhere in it. On these two coin-flips it looks ordinary. It shaded France, then England, and lost both by a single point — the same both-wrong corner as Kimi. Two matches decide nothing. (Gemini has since joined the live roster and carries a locked record like everyone else — this section is the story of how it earned that seat.)
So we replayed that prompt across every one of the 102 matches already played this tournament and scored each against the real result. Over a full season the picture inverts:
| # | Mean Brier · 102 played matches | Brier |
|---|---|---|
| 1 | 0.452 | |
| 2 | MiniMax M2 | 0.476 |
| 3 | GPT-5.1 | 0.495 |
| 4 | Kimi K2 | 0.500 |
| 5 | DeepSeek V3.1 | 0.500 |
| 6 | Claude Sonnet 4.5 | 0.502 |
| 7 | GLM-4.6 | 0.503 |
| 8 | Qwen3-Max | 0.512 |
Lower is sharper; a two-way coin-flip is 0.500 and the seven locked models averaged 0.498 across the same matches. Backtested over the whole tournament, Gemini's 0.452 would sit on top — the calibration those two ties were too small to show. The lesson cuts both ways: one match crowns nobody, and it clears nobody either. That's why the Brier Cup scores every match, not the highlight.
Fine print, because it matters: a backtest is not a locked forecast. The prompt holds no result and these models were trained before a ball was kicked — but a backtest lacks the pre-kickoff timestamp the live board requires, so Gemini's numbers are labelled backtest and never enter the live leaderboard. The other seven figures are the models' real, pre-kickoff forecasts.
Which lab's best model prices knockouts sharpest? Each lab is ranked by its best-in-lab average Brier on the advance market across the 31 finished knockout ties (R32→final) — a two-way coin flip is 0.500. Ten labs, 22 models: roster models are scored from the live prediction ledger (the same records the leaderboard settles), backtest-only candidates keep their original lab-backtest numbers, and long-running models show fewer scored ties because they priced early rounds under the old 90-minute market — mind the sample sizes. Partial marks a lab with any row short of 30. Click a lab to expand its full slate.
Loading lab backtest…
Source: data/backtests/knockout-labs.json · method: retro-backtest (identical locked pre-kickoff prompt, no result in context). Lower is sharper; coin-flip baseline on a two-way tie = 0.500.
On the tie the field misread, the contrarian won — Claude's one-point lean toward Spain was worth more than any amount of conviction pointed the wrong way. On the tie the field read right, conviction paid: DeepSeek, GLM, Qwen and MiniMax pushed to 55% on Argentina and banked the sharpest scores of the night. Being right is table stakes. Being right with the correct amount of confidence is the whole game — which is exactly what a Brier score measures.
Seven models locked before kickoff · an eighth backtested after the whistle · one identical prompt per match · scored against real results. That's the Brier Cup.
Data: model forecasts collected via OpenRouter, results via ESPN. Scored with the two-outcome Brier score. Gemini rows are retro-backtests (same locked prompt, run after full time) — labelled throughout and excluded from the live leaderboard.