Skip to content
All writing

4 min

A field guide to Massive 3

Scratch storage, Slurm, caches, and module notes for Monash University's M3 cluster.

View MarkdownOpen in ClaudeOpen in ChatGPT

Notes from Monash University's M3 supercomputer. Slurm-scheduled, partition-diverse, and the cluster I run most day-to-day experiments on. This is everything I had to work out the hard way.

The cluster at a glance#

  • Login host: m3.massive.org.au
  • Scheduler: Slurm (sbatch, squeue, scancel)
  • Scratch: $HOME/vf38_scratch/<username>, the only sensible place for big files and caches
  • GPU partitions: mixed, see table below
  • Persistent sessions: smux (Monash's tmux wrapper over Slurm)

Always do CUDA-sensitive work (pip installs, sgl-kernel builds, model loads) inside a GPU allocation. Login nodes have no GPU, and anything that probes for CUDA devices during install will segfault.

Partitions#

PartitionGPU TypesPer-nodeExample srun
gpuA40 / A100 / T4 / L40Svariessrun --pty -p gpu -t 00:30:00 --gres=gpu:L40S:1 --cpus-per-task=16 --mem=64G bash
m3hH1004× H100srun --pty --partition=m3h --qos=m3h -t 00:30:00 --gres=gpu:4 --cpus-per-task=72 --mem=985G bash
m3gV100variessrun --pty --partition=m3g -t 00:30:00 --gres=gpu:V100:1 --cpus-per-task=18 --mem=64G bash

To request more than 4 GPUs, add --qos=shortq, without it your job hangs in the queue behind the default QoS limits.

Scratch and caches#

Home quota is small. Scratch is big and fast. Redirect every cache that matters before anything Python-ish runs:

export SCRATCH="$HOME/vf38_scratch/<username>"
mkdir -p "$SCRATCH"/{huggingface,uv,pycache,vllm,sglang}

export HF_HOME="$SCRATCH/.caches/huggingface"
export UV_CACHE_DIR="$SCRATCH/.caches/uv"
export PYTHONPYCACHEPREFIX="$SCRATCH/.caches/pycache"
export VLLM_CACHE_DIR="$SCRATCH/.caches/vllm"
export SGLANG_CACHE_DIR="$SCRATCH/.caches/sglang"

Drop that block into ~/.bashrc or the top of every job script. If you don't, your home quota dies the first time you pull a 600 GB checkpoint.

Persistent sessions with smux#

Slurm's srun dies when the ssh connection dies. For longer interactive work, smux keeps a Slurm allocation alive across disconnects:

smux new-session --partition=m3h --qos=m3h --time=00:15:00 \
  --gres=gpu:4 --cpus=72 --mem=985G
smux attach-session       # reconnect later
smux list-sessions        # list your active allocations

Use this for anything interactive that might last longer than a coffee break.

CUDA toolchain#

M3 ships many CUDA versions as lmod modules. List them once:

ls -a /apps/cuda
module avail cuda

Load the version that matches your PyTorch wheel:

module load cuda/12.6

For bleeding-edge kernels (DeepGEMM on Hopper, for example) you may need 12.9, if it's not in /apps/cuda, install it per-user inside a GPU session; see the Aspire notes for the runfile recipe.

Python environment#

Miniconda + uv is the combination that causes the fewest headaches:

conda create -p "$SCRATCH/py312" python=3.12 -y
conda activate "$SCRATCH/py312"
pip install uv
uv cache dir   # confirm it points at $SCRATCH

Or, if you want the project-local venv pattern:

uv venv
source .venv/bin/activate

Match PyTorch to the loaded CUDA:

uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126

SGLang install#

Inside a GPU session so the build-time CUDA probes succeed:

export TORCH_CUDA_ARCH_LIST="9.0"    # H100 on m3h
uv pip install "sglang[all]>=0.5.0"
uv pip install --force-reinstall "sgl-kernel==0.2.4" sentencepiece

Quick sanity check:

python -m sglang.bench_offline_throughput --help

Sanity checks on a GPU node#

Run these as the first thing inside every allocation, catches 80% of "why isn't my job working" questions before they happen:

hostname
nvidia-smi -L
nvidia-smi
echo "CUDA_VISIBLE_DEVICES=$CUDA_VISIBLE_DEVICES"

python3 - << 'EOF'
import torch
print("torch:", torch.__version__)
print("cuda:", torch.version.cuda)
print("available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
for i in range(torch.cuda.device_count()):
    print(i, torch.cuda.get_device_name(i))
EOF

Slurm batch template#

For a 1-node, 4-GPU H100 job on m3h:

#!/bin/bash
#SBATCH --job-name=sglang
#SBATCH --partition=m3h
#SBATCH --qos=m3h
#SBATCH --time=00:30:00
#SBATCH --gres=gpu:4
#SBATCH --cpus-per-task=72
#SBATCH --mem=985G
#SBATCH --output=run/%x.%j.out

source "$SCRATCH/py312/bin/activate"
module load cuda/12.6

nvidia-smi -L
time python3 -m sglang.bench_offline_throughput \
  --model-path deepseek-ai/DeepSeek-R1 \
  --dataset-path "$SCRATCH/ShareGPT_V3_unfiltered_cleaned_split.json" \
  --num-prompts 2000 --load-format dummy --seed 2025 --dtype bfloat16 \
  --tp 4 --trust-remote-code \
  --warmups compile-deep-gemm

Submit with sbatch, monitor with squeue -u $USER, tail the log once it starts.

Things I wish I'd known#

  • $SCRATCH is not automatic, the variable isn't set for you. Export it yourself in every script.
  • --qos=m3h is mandatory on m3h, leave it off and the job sits in the queue forever without a clear error.
  • smux sessions count against your active-job quota, don't leave five of them running.
  • Module order matters, load cuda/12.6 before activating your conda env, otherwise LD_LIBRARY_PATH gets clobbered.
  • /apps is read-only, everything you install goes in $SCRATCH.
  • H100s on m3h are contended at end of semester, submit early, especially for multi-GPU jobs.

See also#

Start here

Serving DeepSeek-R1 at scale

Four notebooks from the APAC HPC-AI Competition that became the ICPP 2026 paper: install, cluster, tuning, and the daily driver.

  1. 01SGLang and DeepSeek5 min
  2. 02A field guide to Aspire 2A+4 min
  3. 03SGLang optimisations that moved throughput6 min
  4. 04A field guide to Massive 34 min