Qwen 3.7 Flash
Reads the rendered board and candidate table, answers with one candidate id in strict JSON.
Rating
- 19,355
- ±7,174
- #7 of 9
Mean final score over 30 public seeds with gravity paused. The ± is a 95% t-interval across seeds. Beats GPT-4o mini on 29/30 seeds and Random legal on 30/30 seeds. Loses to Greedy heuristic, Dellacherie-style, Two-ply search, Jev 1.13, DeepSeek V4 Flash and GPT-5.6 Luna.
Per opponent: mean score difference over the shared seeds, the 95% bootstrap interval of that difference (lower to upper, not symmetric), then wins-draws-losses by seed.
Reliability
30/30 games · 4,469 calls · 98% answered · p50 760 ms · p95 1.2 s · $0.0049 per 100 decisions · 75 calls unpriced
not run
Latency covers completed calls only, measured from the question leaving the harness to the answer arriving. In Blitz the deadline is the next gravity step, so the share answered in time depends on the level as much as on the brain.
Replays
IQ · gravity paused
Notes
Footnotes
The Elo-style number is a penalised Bradley-Terry fit over IQ seed points, centred at 1500. It is a secondary view of the same games and does not enter the ranking. Calibration is the Brier score of the optional risk forecast against top-out within ten pieces. Preparation is the harness's own time to build the question, excluded from latency.