Skip to content

hosted · LLM · deepseek/deepseek-v4-flash via OpenRouter

DeepSeek V4 Flash

Reads the rendered board and candidate table, answers with one candidate id in strict JSON.

Rating

Score
71,705
±95%
±19,580
Rank
#5 of 9

Mean final score over 30 public seeds with gravity paused. The ± is a 95% t-interval across seeds. Beats GPT-5.6 Luna on 19/30 seeds, Qwen 3.7 Flash on 27/30 seeds, GPT-4o mini on 30/30 seeds and Random legal on 30/30 seeds. No evidence of a difference from Jev 1.13. Loses to Greedy heuristic, Dellacherie-style and Two-ply search.

Greedy heuristic−202,502−225,666 to −175,5321-0-29
Dellacherie-style−198,040−216,243 to −178,0120-0-30
Two-ply search−195,971−213,151 to −176,8840-0-30
Jev 1.13−10,121−44,431 to +20,91316-0-14
GPT-5.6 Luna+29,792+13,133 to +49,41119-0-11
Qwen 3.7 Flash+52,350+34,741 to +72,11827-0-3
GPT-4o mini+69,050+51,673 to +88,53130-0-0
Random legal+71,277+53,759 to +90,67230-0-0

Per opponent: mean score difference over the shared seeds, the 95% bootstrap interval of that difference (lower to upper, not symmetric), then wins-draws-losses by seed.

Reliability

IQ

30/30 games · 8,137 calls · 99.8% answered · 15 timed out · p50 1.3 s · p95 9.9 s · $0.0069 per 100 decisions · 15 calls unpriced

Blitz

not run

Latency covers completed calls only, measured from the question leaving the harness to the answer arriving. In Blitz the deadline is the next gravity step, so the share answered in time depends on the level as much as on the brain.

Notes

Footnotes
Elo-style rating1526 · 1468 to 1583
Calibration (Brier)0.040 over 8,113 forecasts
Preparation p50 / p954.0 ms / 5.3 ms
IQ counts0 invalid · 0 stale · 0 errors · 15 timed out · 0 unreachable · 0 retries
Holds0 IQ
Adapter version3.0.0
Modeldeepseek/deepseek-v4-flash via OpenRouter
Model settingsmodels.ts ↗

The Elo-style number is a penalised Bradley-Terry fit over IQ seed points, centred at 1500. It is a secondary view of the same games and does not enter the ranking. Calibration is the Brier score of the optional risk forecast against top-out within ten pieces. Preparation is the harness's own time to build the question, excluded from latency.