APAC HPC-AI Competition
3rd of 49 university teams. DeepSeek-R1 inference on two 8×H200 nodes, and the open experiment record that became the ICPP 2026 paper.
The problem
Forty-nine university teams, two tracks, and a year of access to hardware most students never touch: NSCC Singapore’s ASPIRE-2A+, NCI Australia and Firmus. I led the AI track as a first-year, on a team of undergraduates and final-year engineers.
- AI track. Maximise inference throughput for DeepSeek-R1 (671B parameters, 37B active) with SGLang across two 8×H200 nodes.
- HPC track. Minimise NWChem run time across multi-node CPU clusters. Nathan Culshaw led this side.
What I found
- Fewer GPUs, arranged deliberately, beat more GPUs arranged by default. A tuned single node (PP2 · TP4 · DP4 attention) hit 17,417 tok/s, 81% over the 16-GPU baseline and 2.05× the per-GPU efficiency. Keeping collectives on NVLink and replicating within a node beats sharding across InfiniBand.
- The boring stuff mattered most. CUDA version alignment alone was worth +35%. Forty-plus hand-tuned NCCL configurations never beat the defaults.
- Getting the model running was the easy part. Most of the time went into finding where SGLang actually bottlenecked on H200s: concurrency, memory management and the configuration surface, not the model.
Outcome
- 3rd of 49. Presented on stage at SupercomputingAsia 2026 in Osaka.
- The paper. The work became “Less is More: Optimising SGLang Distributed DeepSeek-R1 Inference on a Two-Node H200 Cluster”, accepted at the BID ’26 workshop at ICPP 2026. First author with Nathan Culshaw.
- Upstream. Credited SGLang contributor, so fast inference on H200 clusters is something other people can reproduce rather than something I got working once.
The record
Every configuration tried is in the open repository: launch scripts, Slurm logs and the workbooks used to tabulate results. Every number in the paper can be reproduced from it.
The team
- Josh Riantoputra · Team captain
- Luca Lowndes · AI lead
- Nathan Culshaw · HPC lead
- Isaac Barnes · Engineer
- Giacomo Bonomi · Engineer
Notes from the competition
The working notebooks behind this result, written on the clusters as we went.
- SGLang and DeepSeek · installing, serving and benchmarking R1
- A field guide to Aspire 2A+ · the NSCC H100 cluster
- SGLang optimisations that moved throughput · the tuning that placed
- A field guide to Massive 3 · Monash’s M3, the daily driver
Stack
DeepSeek-R1, SGLang, CUDA, NCCL, Slurm, Singularity, NVIDIA H100/H200