Skip to content

hosted · LLM · openai/gpt-4o-mini via OpenRouter

GPT-4o mini

Reads the rendered board and candidate table, answers with one candidate id in strict JSON.

Rating

Score
2,655
±95%
±488
Rank
#8 of 9

Mean final score over 30 public seeds with gravity paused. The ± is a 95% t-interval across seeds. Beats Random legal on 30/30 seeds. Loses to Greedy heuristic, Dellacherie-style, Two-ply search, Jev 1.13, DeepSeek V4 Flash, GPT-5.6 Luna and Qwen 3.7 Flash.

Greedy heuristic−271,552−280,817 to −255,7750-0-30
Dellacherie-style−267,090−273,942 to −256,3460-0-30
Two-ply search−265,021−266,907 to −263,1710-0-30
Jev 1.13−79,171−105,150 to −56,4930-0-30
DeepSeek V4 Flash−69,050−88,531 to −51,6730-0-30
GPT-5.6 Luna−39,258−48,618 to −30,6710-0-30
Qwen 3.7 Flash−16,700−24,528 to −11,0911-0-29
Random legal+2,227+1,768 to +2,70630-0-0

Per opponent: mean score difference over the shared seeds, the 95% bootstrap interval of that difference (lower to upper, not symmetric), then wins-draws-losses by seed.

Reliability

IQ

30/30 games · 1,870 calls · 100% answered · p50 1.4 s · p95 3.3 s · $0.024 per 100 decisions

Blitz

not run

Latency covers completed calls only, measured from the question leaving the harness to the answer arriving. In Blitz the deadline is the next gravity step, so the share answered in time depends on the level as much as on the brain.

Notes

Footnotes
Elo-style rating1024 · 1003 to 1045
Calibration (Brier)0.149 over 1,870 forecasts
Preparation p50 / p954.3 ms / 6.8 ms
IQ counts0 invalid · 0 stale · 0 errors · 0 timed out · 0 unreachable · 0 retries
Holds8 IQ
Adapter version3.0.0
Modelopenai/gpt-4o-mini via OpenRouter
Model settingsmodels.ts ↗

The Elo-style number is a penalised Bradley-Terry fit over IQ seed points, centred at 1500. It is a secondary view of the same games and does not enter the ranking. Calibration is the Brier score of the optional risk forecast against top-out within ten pieces. Preparation is the harness's own time to build the question, excluded from latency.