GPT-5.6 Luna
Reads the rendered board and candidate table, answers with one candidate id in strict JSON.
Rating
- 41,912
- ±10,133
- #6 of 9
Mean final score over 30 public seeds with gravity paused. The ± is a 95% t-interval across seeds. Beats Qwen 3.7 Flash on 22/30 seeds, GPT-4o mini on 30/30 seeds and Random legal on 30/30 seeds. Loses to Greedy heuristic, Dellacherie-style, Two-ply search, Jev 1.13 and DeepSeek V4 Flash.
Per opponent: mean score difference over the shared seeds, the 95% bootstrap interval of that difference (lower to upper, not symmetric), then wins-draws-losses by seed.
Reliability
30/30 games · 6,449 calls · 100% answered · p50 1.2 s · p95 1.7 s · $0.040 per 100 decisions
not run
Latency covers completed calls only, measured from the question leaving the harness to the answer arriving. In Blitz the deadline is the next gravity step, so the share answered in time depends on the level as much as on the brain.
Replays
IQ · gravity paused
Notes
Footnotes
The Elo-style number is a penalised Bradley-Terry fit over IQ seed points, centred at 1500. It is a secondary view of the same games and does not enter the ranking. Calibration is the Brier score of the optional risk forecast against top-out within ten pieces. Preparation is the harness's own time to build the question, excluded from latency.