M3 (Massive 3)

Notes from Monash University's M3 supercomputer. Slurm-scheduled, partition-diverse, and the cluster I run most day-to-day experiments on. This is everything I had to work out the hard way.

The cluster at a glance

Always do CUDA-sensitive work (pip installs, sgl-kernel builds, model loads) inside a GPU allocation. Login nodes have no GPU, and anything that probes for CUDA devices during install will segfault.

Partitions

PartitionGPU TypesPer-nodeExample srun
gpuA40 / A100 / T4 / L40Svariessrun --pty -p gpu -t 00:30:00 --gres=gpu:L40S:1 --cpus-per-task=16 --mem=64G bash
m3hH1004× H100srun --pty --partition=m3h --qos=m3h -t 00:30:00 --gres=gpu:4 --cpus-per-task=72 --mem=985G bash
m3gV100variessrun --pty --partition=m3g -t 00:30:00 --gres=gpu:V100:1 --cpus-per-task=18 --mem=64G bash

To request more than 4 GPUs, add --qos=shortq, without it your job hangs in the queue behind the default QoS limits.

Scratch and caches

Home quota is small. Scratch is big and fast. Redirect every cache that matters before anything Python-ish runs:

export SCRATCH="$HOME/vf38_scratch/<username>"
mkdir -p "$SCRATCH"/{huggingface,uv,pycache,vllm,sglang}

export HF_HOME="$SCRATCH/.caches/huggingface"
export UV_CACHE_DIR="$SCRATCH/.caches/uv"
export PYTHONPYCACHEPREFIX="$SCRATCH/.caches/pycache"
export VLLM_CACHE_DIR="$SCRATCH/.caches/vllm"
export SGLANG_CACHE_DIR="$SCRATCH/.caches/sglang"

Drop that block into ~/.bashrc or the top of every job script. If you don't, your home quota dies the first time you pull a 600 GB checkpoint.

Persistent sessions with smux

Slurm's srun dies when the ssh connection dies. For longer interactive work, smux keeps a Slurm allocation alive across disconnects:

smux new-session --partition=m3h --qos=m3h --time=00:15:00 \
  --gres=gpu:4 --cpus=72 --mem=985G
smux attach-session       # reconnect later
smux list-sessions        # list your active allocations

Use this for anything interactive that might last longer than a coffee break.

CUDA toolchain

M3 ships many CUDA versions as lmod modules. List them once:

ls -a /apps/cuda
module avail cuda

Load the version that matches your PyTorch wheel:

module load cuda/12.6

For bleeding-edge kernels (DeepGEMM on Hopper, for example) you may need 12.9, if it's not in /apps/cuda, install it per-user inside a GPU session; see the Aspire notes for the runfile recipe.

Python environment

Miniconda + uv is the combination that causes the fewest headaches:

conda create -p "$SCRATCH/py312" python=3.12 -y
conda activate "$SCRATCH/py312"
pip install uv
uv cache dir   # confirm it points at $SCRATCH

Or, if you want the project-local venv pattern:

uv venv
source .venv/bin/activate

Match PyTorch to the loaded CUDA:

uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126

SGLang install

Inside a GPU session so the build-time CUDA probes succeed:

export TORCH_CUDA_ARCH_LIST="9.0"    # H100 on m3h
uv pip install "sglang[all]>=0.5.0"
uv pip install --force-reinstall "sgl-kernel==0.2.4" sentencepiece

Quick sanity check:

python -m sglang.bench_offline_throughput --help

Sanity checks on a GPU node

Run these as the first thing inside every allocation, catches 80% of "why isn't my job working" questions before they happen:

hostname
nvidia-smi -L
nvidia-smi
echo "CUDA_VISIBLE_DEVICES=$CUDA_VISIBLE_DEVICES"

python3 - << 'EOF'
import torch
print("torch:", torch.__version__)
print("cuda:", torch.version.cuda)
print("available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
for i in range(torch.cuda.device_count()):
    print(i, torch.cuda.get_device_name(i))
EOF

Slurm batch template

For a 1-node, 4-GPU H100 job on m3h:

#!/bin/bash
#SBATCH --job-name=sglang
#SBATCH --partition=m3h
#SBATCH --qos=m3h
#SBATCH --time=00:30:00
#SBATCH --gres=gpu:4
#SBATCH --cpus-per-task=72
#SBATCH --mem=985G
#SBATCH --output=run/%x.%j.out

source "$SCRATCH/py312/bin/activate"
module load cuda/12.6

nvidia-smi -L
time python3 -m sglang.bench_offline_throughput \
  --model-path deepseek-ai/DeepSeek-R1 \
  --dataset-path "$SCRATCH/ShareGPT_V3_unfiltered_cleaned_split.json" \
  --num-prompts 2000 --load-format dummy --seed 2025 --dtype bfloat16 \
  --tp 4 --trust-remote-code \
  --warmups compile-deep-gemm

Submit with sbatch, monitor with squeue -u $USER, tail the log once it starts.

Things I wish I'd known

See also