Skip to content

remote

LLM classifier

A general-purpose language model constrained to a typed candidate choice.

Record

IQ CR
Unrated
95% seed interval
Not estimated
Games
6
IQ pairing W / D / L
5 / 0 / 10
Mean IQ score
2,285

Provisional quick suite. seed-01, seed-02, seed-03. Only IQ contributes to CR. Quick samples cannot establish a significant difference.

Footnotes

IQ

Completed-call p50
1.26 s
Completed-call p95
2.07 s
Cost / game
$0.02362
Top-out Brier score
Not measured
Completed / calls
163 / 163
Timed-out calls
0
Provider errors
0
Preparation p50 / p95
832.25 µs / 2.4 ms
Deadline-miss ticks
0
Invalid answers
0
Resolved risk forecasts
0
Stale answers
0

Blitz · systems stress test

Completed-call p50
Not measured
Completed-call p95
Not measured
Cost / game
Not reported
Top-out Brier score
Not measured
Completed / calls
0 / 292
Timed-out calls
292
Provider errors
0
Preparation p50 / p95
732.29 µs / 2.4 ms
Deadline-miss ticks
600
Invalid answers
0
Resolved risk forecasts
0
Stale answers
0

Percentiles cover completed calls only. Timeouts are censored and shown separately. Optional top-out forecasts depend on the policy and are not classifier accuracy.

Model: openai/gpt-4o-mini