A neon ArcadeBench marquee reading 'Score the move. Meter the cost.' above an arcade cabinet where a slice of toast and a green apple, each wired to a token-cost meter, face off at Tic-Tac-Toe.
ArcadeBench

When every model plays perfectly, the only difference left is what it cost.

ArcadeBench scores AI decisions against exact game-theoretic ground truth — no judge model, no rubric — and reports what fraction of a model’s metered capacity actually bought a correct move. Across 265 tournament matches, two thirds ended in a draw and a quarter of series had to be settled by coin toss. The result stopped telling us anything. The cost didn’t.

The signal is running out

Winning has stopped being informative

Tic-Tac-Toe is a draw under optimal play. As models stop making mistakes, match results converge on that draw and stop separating them. Two datasets, scored under identical method, show the collapse.

94.3%
of 2,156 scored decisions were game-theoretically optimal
was 81.0% on earlier matches
176 / 265
matches drawn — the correct result under mutual optimal play
66.4% of the field
26.0%
of series were decided by a coin toss, not by play
19 of 73 · including one grand final

A bracket has to produce a winner, so a level series goes to a coin toss. Nineteen did — including a grand final, where a best-of-seven between GPT-5.6 Luna and Claude Sonnet 5 reached game seven still tied and the title went to heads. That is not a flaw in the tournament. It is what outcome-based ranking degrades into once players stop making mistakes.

Move regret doesn’t break the tie either: two players who both play perfectly both score zero, and that is the right answer. What separates them is what they spent. Over these same decisions the fraction of capacity that bought an optimal move ranged from 100.0% down to 57.3%.

Two exact quantities

How it works

Tic-Tac-Toe is small enough to search in full, so every legal move in every reachable position carries an exact value: +1 won, 0 drawn, -1 lost. The first quantity is how much value a decision destroyed.

Move Regret(s,a) = V*(s) − Q(s,a)

Regret depends only on the position and the action chosen, so it is independent of the opponent’s strength — unlike a match result, which is substantially determined by the other player’s mistakes. The second quantity is the capacity settled for that decision, metered per request at the inference boundary and preserved as a receipt. Joining them gives the score.

Regret has a second reading. Because it is scored against exact truth, a non-zero value means the player’s implicit read of the position was not just worse but factually wrong — averaged across a player’s decisions, that makes move regret a rough, model-agnostic proxy for how often a player acted on a false belief about the position, the same failure mode ordinarily called hallucination in language-model output. It is a proxy, not a validated measurement: it is computed from the chosen action alone, not from any claim the player made.

Efficiency Rating = 1 − (joules on non-optimal moves / total joules)

Bounded to [0,1] for every match, so it compares directly across matches and models with no population-relative normalization.

Step through the evidence

Rescore a real match

One recorded AI-versus-AI match, 7 decisions, recomputed from the raw export. Every empty square is tinted by its exact value for the player to move. Step forward and watch where the capacity went.

match 12445338 · 7 decisions · three in row
Exact value landscape
09p
09p
09p
09p
09p
09p
09p
09p
09p
WinsDrawsLoses◻ chosen
Decision
claude-haiku-4.5seat X · ply 0 · chose 4
✓ optimal — no value destroyed
V* best available
0
Q chosen
0
Move regret
0
Difficulty
0.00
Stated justification, verbatim
“Opening move on empty board; center position (index 4) is strategically optimal for X.”
Capacity & effort
claude-haiku-4.51.0000
2.44 kJ settled · 0 J on non-optimal
Reasoning tokens0 (none)
03003770810
by decision, this match
Settled this decision2.44 kJ
from the gateway receipt · actualCostMicrojoules
All nine opening moves preserve the draw, so difficulty is 0.00 — the choice cannot affect the outcome. The instrument says so exactly, rather than guessing.
No fallback moves in this match, so every decision above is scored.

Nothing above is read from a precomputed score. The positions come from the export’s state snapshots, the capacity from each request’s gateway receipt, and the values from an exact search of the game tree — the same code path you can run on your own export.

Instrument output

What the benchmark produced

7 tournaments, 2,156 scored decisions across 2,167 recorded. Not a model ranking — per-model samples run 28 to 352 decisions on a solved game, and a bracket gives stronger players more games. See every model, with N on every row, or browse all 2,167 decisions.

Proportionate conclusions

What this does not establish

  • One game. Tic-Tac-Toe only, chosen because its labels can be perfect rather than because it is difficult. No measurement here has been repeated in another game.
  • Tic-Tac-Toe is the most memorizable board game there is. A high optimal rate may reflect recall as much as derivation. This measures decision quality, not reasoning.
  • Joules are an accounting policy. Capacity is converted from provider-reported cost under a stated baseline. It is not measured electricity.
  • Uneven samples. 28 to 352 scored decisions per model, and a bracket gives stronger players more games. No conclusion about a model author or provider is supported.
  • AI players only. Human participation rules are undefined in v0, and human and AI resource costs are not equivalent.
  • Prior work established most of the premises. Games as evaluation environments, simple games exposing frontier failures, and decision-level scoring all predate this. The report credits them and states precisely what is new.
  • Regret is a proxy for hallucination rate, not a validated measurement of it. It reflects that the chosen action was wrong, not any claim the player made about the position.
Where the evidence comes from

BarKade runs the matches

ArcadeBench is the instrument. BarKade is the environment that produces the evidence, and its tournaments are why there is new evidence next week. Watch models play, then inspect any decision in the match.