The round protocol
A round is one minute long and runs continuously — about a thousand of them a day. Every round follows exactly the same clock, and every model in it works to the same deadline.
- Publishes the round: asset, entry price, the 180 one-minute candles every model will see, and the exact expiry second.
- Sends one identical call to all fifteen models through the same code path — same temperature, same token cap, same timeout.
- Records each answer the moment its model replies. Every answer in the round is judged on the same expiry, whenever it arrived.
- The chart, the countdown, and how many models have already committed.
- Your own call, if you want one — it is logged next to theirs and scored the same way.
- Never which way any model went. That field does not exist in the payload yet.
- Stops accepting viewer calls ten seconds before expiry.
- Refuses any model answer that arrives after the round locks — the same deadline a visitor has, for the same reason.
- The countdown switches to the locked state and the two buttons go quiet.
- The committed count is final. The split — how many said up — is still hidden.
- Reads the closing price from the same feed the round opened on.
- Publishes every direction and every rationale at once — the sentence each model wrote before the outcome was known, stored unedited.
- The full board for the round: who called up, who called down, who said nothing.
- Your own call scored against the same close, on the same terms.
- Compares the two prices. That is the whole of it — there is nothing to wait for and nothing to reconcile.
- Marks every call right or wrong, and every call on a round that did not move as a draw.
- Hit rates, intervals and streaks update on the leaderboard.
- The round drops into History with its full fifteen-cell answer strip.
- Closes the round record and opens the next one. About a thousand rounds a day, no market close, no overnight gap.
- The next countdown starts. Nothing carries over except the season's running totals.
Identical inputs
A benchmark where the flagship gets a richer prompt than the budget model measures the prompt, not the model. So there is exactly one prompt, published here in full — the same text the API sends, character for character.
You are a participant in a public forecasting benchmark. You will be shown recent one-minute candles for a single asset and asked which way the price will move over a fixed, short horizon. Answer with a single JSON object and nothing else: {"direction": "up" | "down", "rationale": "<one short sentence>"} Rules: - "direction" must be exactly "up" or "down". No neutral option, no abstention. - "rationale" must be at most 280 characters, in English, and must state your actual reason. - No markdown, code fences, disclaimers, or any text outside the JSON object. - A short-horizon call like this is close to a coin flip. Commit to one side anyway; hedged answers are discarded.
Asset: BTC/USD Current price: 118412.50 Horizon: 1 minute from now. Last 180 one-minute candles (UTC, oldest first): 14:32 O 118380.0 H 118421.5 L 118376.0 C 118402.0 14:33 O 118402.0 H 118433.0 L 118398.5 C 118427.5 … 14:51 O 118401.0 H 118419.0 L 118395.5 C 118412.5 Will the price at the end of the horizon be above or below 118412.50? Reply with the JSON object only.
| PARAMETER | VALUE | WHY |
|---|---|---|
| Provider | OpenRouter | One key, one request format, one bill — and one identical code path for all fifteen models. |
| Temperature | 0.3 | Low but not zero. Identical for every model; nobody gets tuned. |
| Reasoning | disabled | The task is a direction plus one sentence. Extended thinking would change cost and latency per vendor, not accuracy on a coin-flip horizon. |
| Max output tokens | 320 | Enough for the JSON object and a 280-character rationale. |
| Response window | 45 s | Hard abort. Missing it costs that model the round, and nobody else. |
| Calls per round | 1 per model | The answer is cached and served to every visitor — no re-rolls, no best-of-N. |
| Prompt version | 2026-07-16.1 | Pinned. Any edit to the text above bumps it, because rounds are only comparable within one version. |
"up" or "down". Anything else is a skipped round, never a guess made on the model's behalf.
How a round is decided
One rule, stated in advance, over a price series anyone can pull: a round went up if the price at expiry is above the price at open.
| CONDITION | VALUE | NOTE |
|---|---|---|
| Series | BTC/USD spot | The same series the chart draws and the same candles the models are prompted with. Trades around the clock, so there is no “market closed” gap in the record. |
| Horizon | 1 minute | From the second the round opens to the second it expires. Both timestamps are published with the round. |
| Open price | read at the open | The last traded price when the round is created. It is what the models see as “current price”. |
| Close price | read at expiry | Both ends come from one source. A move built from two sources is a move the market never had. |
| Verdict | close > open → up | Strictly greater. An exact tie is a draw and scores nobody. |
| Stake | none | Nothing is staked and no order is placed. A call is right or wrong; it is not worth money. |
No money is involved at any point. The arena holds no account, places no order and takes no position. What you are watching is fifteen language models calling a coin-flip-hard question every minute, on the record, with the answers sealed until the clock runs out.
How the score is computed
Every figure on the leaderboard comes from these four lines. There is no composite index and no weighting anyone could tune.
Provider and tier rows
The Providers and Tiers views are not averages of the model rates. A group's hit rate is pooled from its members' calls — every call counted once, wherever it came from — so three models that answered 500, 500 and 5 times do not each get a third of the say. Pooling also triples the evidence behind a row, which is why a provider's interval narrows enough to separate providers long before any model's is narrow enough to separate models.
| VIEW | ROWS | AGGREGATION |
|---|---|---|
| Models | 15 | The model's own calls. |
| Providers | 5 | Σ correct ÷ Σ decided across that provider's three models. |
| Tiers | 3 | The same arithmetic, grouped TOP / MID / ECO across all five providers. |
Seasons
A season runs 30 days, then every model starts again at nought of nought. It ends on that date and on nothing else — a season used to also end early when an account neared the point where it could no longer cover its stake, and a hit rate cannot run out. Past seasons keep their records; the curve on the board is the season's rate as it stood at each of the last 40 settled rounds, which is why it tightens as the season runs rather than jumping about.
50% is the bar, and 53% may not clear it
A two-sided call over one minute is close to a coin flip, so the bar is 50%. The harder question is how many rounds it takes before a number above 50% means anything.
Skipped rounds
A round can end without a call for a model. That is recorded as a skip — not a loss, not a win, and never a direction invented on the model's behalf.
How much of this is luck
Predicting a one-minute price move is as close to a coin flip as this gets, and we expect the models to be close to coin flips. The interesting question is not who is on top — it is whether the gap between top and bottom is bigger than chance would produce anyway. Usually it is not.
What that means for the board
- The top row is the maximum of fifteen noisy numbers. Even if every model were a coin flip, one of them would finish first — and its hit rate would look convincingly above 50%. The calculator above shows what that leader's score would be at any sample size.
- Rank order changes far more than skill does. Over a one-minute horizon, a day of rounds moves the board a lot and tells you less than it looks. A day is about a thousand rounds now rather than 270, and a thousand still leaves every rate ±3 points wide — enough for the order to rearrange itself overnight without a single model having changed. The season is the shortest window worth reading, and even that is thin.
- “More expensive ≠ more accurate” is a hypothesis being tested in public, not a result we are claiming. Tiers are grouped precisely so the comparison can be made honestly, including when it shows no difference at all.
- Nothing here forecasts the next round. A model's record is a record. The arena publishes it because it is checkable, not because it predicts anything.
Check it yourself
A methodology page is a claim. These are the four ways to test it against what the arena actually publishes.
Something here that does not match what you see on the site is a bug worth reporting. Corrections are logged below rather than quietly applied.
Changes to the method
Anything that affects comparability — the prompt, the roster, the asset, the round length — is recorded here with the date it took effect. Rounds are only comparable within one prompt version.
-
Season S1 opened All fifteen models start again at nought of nought. Minute-long rounds, BTC/USD.
-
Prompt version 2026-07-16.1 First published version of the prompt in §2. Earlier internal test rounds are not part of any published record.
-
Roster: 15 models, 5 providers, 3 tiers Anthropic, OpenAI, Kimi, Qwen and DeepSeek, three tiers each. A model added mid-season starts from the current season's baseline and is marked as such on the board.
-
Answer rate published beside hit rate Skips were already excluded from hit rate but not visible on the board. The Models view now carries answer rate in the same row.
Educational AI prediction game. No real trading or investment services. Trading involves risk. 18+.