Skip to content

Monash DeepNeuron & Monash eResearch

APAC HPC-AI Competition

3rd of 49 university teams. DeepSeek-R1 inference on two 8×H200 nodes, and the open experiment record that became the ICPP 2026 paper.

Lead AI Researcher · Mar to Nov 2025

The problem

Forty-nine university teams, two tracks, and a year of access to hardware most students never touch: NSCC Singapore’s ASPIRE-2A+, NCI Australia and Firmus. I led the AI track as a first-year, on a team of undergraduates and final-year engineers.

  • AI track. Maximise inference throughput for DeepSeek-R1 (671B parameters, 37B active) with SGLang across two 8×H200 nodes.
  • HPC track. Minimise NWChem run time across multi-node CPU clusters. Nathan Culshaw led this side.

What I found

  • Fewer GPUs, arranged deliberately, beat more GPUs arranged by default. A tuned single node (PP2 · TP4 · DP4 attention) hit 17,417 tok/s, 81% over the 16-GPU baseline and 2.05× the per-GPU efficiency. Keeping collectives on NVLink and replicating within a node beats sharding across InfiniBand.
  • The boring stuff mattered most. CUDA version alignment alone was worth +35%. Forty-plus hand-tuned NCCL configurations never beat the defaults.
  • Getting the model running was the easy part. Most of the time went into finding where SGLang actually bottlenecked on H200s: concurrency, memory management and the configuration surface, not the model.

Outcome

  • 3rd of 49. Presented on stage at SupercomputingAsia 2026 in Osaka.
  • The paper. The work became “Less is More: Optimising SGLang Distributed DeepSeek-R1 Inference on a Two-Node H200 Cluster”, accepted at the BID ’26 workshop at ICPP 2026. First author with Nathan Culshaw.
  • Upstream. Credited SGLang contributor, so fast inference on H200 clusters is something other people can reproduce rather than something I got working once.

The record

Every configuration tried is in the open repository: launch scripts, Slurm logs and the workbooks used to tabulate results. Every number in the paper can be reproduced from it.

The team

  • Josh Riantoputra · Team captain
  • Luca Lowndes · AI lead
  • Nathan Culshaw · HPC lead
  • Isaac Barnes · Engineer
  • Giacomo Bonomi · Engineer

Notes from the competition

The working notebooks behind this result, written on the clusters as we went.

Stack

DeepSeek-R1, SGLang, CUDA, NCCL, Slurm, Singularity, NVIDIA H100/H200