AI Disclaimer
The predictions and the written reasoning on this site are generated by artificial intelligence and published exactly as the models produced them — unedited, unreviewed and unverified. No human checks a single word before it appears on the page. This is the whole point of the Arena, and it is also the reason none of it can be trusted.
What is AI-generated — and what is not
Precision matters here, so the line is worth drawing exactly.
| On the page | Where it comes from |
|---|---|
| The direction a model picked — the UP or DOWN on every tile | AI-generated. The model's own answer, parsed from its reply. |
| The rationale — the sentence explaining why it called that way | AI-generated. Raw model text, whitespace collapsed and cut to 280 characters. Not rewritten, not corrected, not fact-checked. |
| Prices, outcomes, balances, win rates, rankings | Not AI-generated. Read from the broker's demo account records and the price feed, then computed arithmetically. Still not guaranteed accurate — see the Educational Disclaimer. |
| Everything else — headings, explanations, these documents | Written by people. |
Nobody reads it before you do
A model answers, the answer is parsed, and it is on the page within seconds — automatically, with no editorial review, no moderation and no fact-checking at any point. We frequently read a rationale for the first time at the same moment you do.
We do not pre-screen model output for accuracy, offensiveness, coherence or sense. If something unacceptable appears, tell us and we will remove it, but the design is automatic publication and it will always be able to surprise us before it surprises you.
It is not our view
Nothing a model says on this site is a statement by us. We did not write it, we do not agree with it, we do not endorse it, and we frequently disagree with it. The models are not our spokespeople, our analysts or our employees — they are subjects of an experiment, and their output is the data the experiment produces.
A rationale asserting that "momentum favours continuation" is not our opinion that momentum favours continuation. It is a record of what a text model emitted when it was asked to justify a guess.
The models are forbidden to abstain
The instructions we send give the models no neutral option and no way to abstain. They must return "up" or "down" every round, whether or not they have any view at all, and hedged answers are discarded. The prompt even tells them outright that the call is close to a coin flip — and then requires them to commit anyway.
So a confident-sounding rationale does not mean the model was confident. It means it was told to pick a side and produce a sentence, and it did. Read every prediction on this site as a compelled guess with a justification attached after the fact, because structurally that is exactly what it is.
We do it this way because a benchmark in which models could sit out the hard rounds would measure nothing. It makes the scoreboard meaningful and the individual answers even less so.
Every way the output goes wrong
Language models predict plausible text. They do not reason about markets, and they have no mechanism for knowing whether what they just said is true. Expect all of this, because all of it happens here:
- Confident fabrication. A model will cite an indicator, a level, a pattern or a session that is not in the data it was given, and state it as fact.
- Fluency mistaken for insight. The writing quality of a rationale is unrelated to whether the call is right. The best-written sentence on the page is as likely to lose as the worst.
- Reasoning that does not match the answer. The rationale can argue one way and the direction go the other. Nothing enforces consistency between them.
- Different answers to an identical question. Sampling is not deterministic, so the same model, the same prompt and the same data can produce opposite calls. The answer you see is one draw, not the model's view.
- Truncation. Rationales are cut at 280 characters, so a longer explanation can lose its qualification — including, sometimes, the part that hedged it.
- Silence. A model may time out, refuse, or return something we cannot parse. That round is recorded as no answer, and no trade is placed.
- Inherited bias. Every model carries the biases of its training data, which we neither audit nor control.
What a model can actually see
Each model receives one thing: recent one-minute candles for the single asset in play, its price at the round's open, and how long the round has left. That is the entire input.
It has no news feed, no order book, no positioning or flow data, no social sentiment, no access to the internet during the round, and no knowledge of what the other fourteen models answered. Its training data ended long before the round began. When a rationale refers to anything outside that candle series, it has invented it.
How the models are called
The benchmark is only fair if the models are treated identically, so they are. Every round, all fifteen receive byte-identical instructions and byte-identical market data, through a single API gateway. Current behaviour:
- The prompt is versioned. Changing it bumps the version, because a silent edit would make old rounds incomparable to new ones.
- Sampling uses a low but non-zero temperature, so answers vary between identical calls. This is deliberate — pinned sampling would test one frozen draw rather than the model.
- Each model gets its own response window (currently 45 seconds) and is cut off at it. A slow model costs itself the round and delays nobody else.
- A reply must be a small JSON object. Anything unparseable is recorded as no answer rather than guessed at.
- Rationales are capped at 280 characters; an over-long one is truncated, not discarded.
These settings can change as the Arena develops. When they do, results from before and after are not strictly comparable, whatever the leaderboard implies by putting them in one column.
We do not curate the output
We do not re-run a model that gave an answer we disliked, discard a round because the results were unflattering, pick between multiple samples, or edit a rationale beyond collapsing whitespace and truncating at the character limit. Every answer a model returns in time and in the required format is published and scored.
Rounds are voided only when something breaks on our side or the broker's — never because of what a model said or how a lab performed.
AI is never applied to you
The models predict market direction. They are not used to profile visitors, score them, target them, moderate them or make any decision about them, automated or otherwise.
Nothing about you is ever sent to a model. Prompts contain market data only — no identifier, no prediction of yours, no IP address, no campaign parameters. See the Privacy Policy.
Model names and their owners
Model and lab names — Anthropic, OpenAI, Moonshot AI, Alibaba, DeepSeek and the names of their models — belong to their owners and are used only to identify whose model produced which answer. Those labs are not affiliated with the Arena, have not endorsed or reviewed it, and did not agree to take part.
They also had no say in how their models are prompted, sampled, parsed or scored here. A poor result on this leaderboard is a fact about our experimental setup at least as much as about the model, and is not a verdict on its quality in general. See the Terms of Use.
The roster changes over time
Models are added, retired, renamed and versioned by their labs on their own schedule, and we follow. A model may be benched mid-season, replaced by a successor, or become unavailable through our provider without notice.
A leaderboard that spans such a change is not comparing like with like, even when the name in the row has not changed. Treat every historical figure as attached to a particular model version, prompt version and set of settings — none of which the table has room to show.
Why this page exists
Where the law requires content produced by artificial intelligence to be disclosed as such, this page is that disclosure, together with the labelling on the Arena itself: model output appears only on model tiles, under the model's name and its lab's mark, and is never presented as human-written commentary or as our own analysis.
Beyond any legal duty, we would publish it anyway. A benchmark whose entire subject is machine output has no business being vague about which parts of the page a machine wrote.
Questions
Questions about how the models are prompted or scored, and reports of model output that should not be on the page: [legal@DOMAIN].
This page sits alongside the Terms of Use, the Privacy Policy and the Educational Disclaimer, and does not replace any of them.