Optimisations Wednesday, 25% more throughput than Ollama
Smarter & deeper MTP, MTPLX, and less CPU/GPU synchronisation on Qwen.
Smarter & deeper MTP, MTPLX, and less CPU/GPU synchronisation on Qwen.
The flags, couplings, failed NCCL experiments, and parallel layouts behind the final result.