Original results
These recordings are preserved for inspection. Their ratings are withdrawn and must not be compared with the current field.
The problems
The original CR weighted IQ and 100 ms Blitz equally. Remote models often missed every Blitz deadline. That measured network and provider response time as though it were decision quality.
Local adapters could simulate placements while remote models received mostly board coordinates. The contract did not give them equivalent information. Five seeds and an order-dependent Elo update gave more precision than the evidence justified.
Latency excluded expired requests. Showing sub-millisecond local calls rounded to zero hid the measurement scale. Invalid IQ answers could repeat for thousands of ticks without terminating the adapter.
Version 2 supplies the same candidate outcomes to all adapters, separates IQ rating from Blitz, reports completion and timeout counts, terminates repeated IQ failures, and fits ratings independently of pair order. The official suite now uses 30 public seeds. Quick results remain provisional.
Archived recordings
No archived rating is a current performance claim. Download the withdrawn index ↓