# A field guide to Massive 3

> Scratch storage, Slurm, caches, and module notes for Monash University's M3 cluster.

- Author: Luca Lowndes
- Published: 2025-12-04
- Topics: HPC, Slurm, Monash
- Canonical: https://luca.lowndes.net/blog/massive-3

Notes from Monash University's **M3** supercomputer. Slurm-scheduled, partition-diverse, and the cluster I run most day-to-day experiments on. This is everything I had to work out the hard way.


## The cluster at a glance

- **Login host**: `m3.massive.org.au`
- **Scheduler**: Slurm (`sbatch`, `squeue`, `scancel`)
- **Scratch**: `$HOME/vf38_scratch/<username>`, the only sensible place for big files and caches
- **GPU partitions**: mixed, see table below
- **Persistent sessions**: `smux` (Monash's tmux wrapper over Slurm)

Always do CUDA-sensitive work (pip installs, `sgl-kernel` builds, model loads) inside a GPU allocation. Login nodes have no GPU, and anything that probes for CUDA devices during install will segfault.


## Partitions

| Partition | GPU Types | Per-node | Example srun |
| :-- | :-- | :-- | :-- |
| `gpu` | A40 / A100 / T4 / L40S | varies | `srun --pty -p gpu -t 00:30:00 --gres=gpu:L40S:1 --cpus-per-task=16 --mem=64G bash` |
| `m3h` | H100 | 4× H100 | `srun --pty --partition=m3h --qos=m3h -t 00:30:00 --gres=gpu:4 --cpus-per-task=72 --mem=985G bash` |
| `m3g` | V100 | varies | `srun --pty --partition=m3g -t 00:30:00 --gres=gpu:V100:1 --cpus-per-task=18 --mem=64G bash` |

To request more than 4 GPUs, add `--qos=shortq`, without it your job hangs in the queue behind the default QoS limits.


## Scratch and caches

Home quota is small. Scratch is big and fast. Redirect every cache that matters before anything Python-ish runs:

```bash
export SCRATCH="$HOME/vf38_scratch/<username>"
mkdir -p "$SCRATCH"/{huggingface,uv,pycache,vllm,sglang}

export HF_HOME="$SCRATCH/.caches/huggingface"
export UV_CACHE_DIR="$SCRATCH/.caches/uv"
export PYTHONPYCACHEPREFIX="$SCRATCH/.caches/pycache"
export VLLM_CACHE_DIR="$SCRATCH/.caches/vllm"
export SGLANG_CACHE_DIR="$SCRATCH/.caches/sglang"
```

Drop that block into `~/.bashrc` or the top of every job script. If you don't, your home quota dies the first time you pull a 600 GB checkpoint.


## Persistent sessions with smux

Slurm's `srun` dies when the ssh connection dies. For longer interactive work, `smux` keeps a Slurm allocation alive across disconnects:

```bash
smux new-session --partition=m3h --qos=m3h --time=00:15:00 \
  --gres=gpu:4 --cpus=72 --mem=985G
smux attach-session       # reconnect later
smux list-sessions        # list your active allocations
```

Use this for anything interactive that might last longer than a coffee break.


## CUDA toolchain

M3 ships many CUDA versions as lmod modules. List them once:

```bash
ls -a /apps/cuda
module avail cuda
```

Load the version that matches your PyTorch wheel:

```bash
module load cuda/12.6
```

For bleeding-edge kernels (DeepGEMM on Hopper, for example) you may need 12.9, if it's not in `/apps/cuda`, install it per-user inside a GPU session; see the Aspire notes for the runfile recipe.


## Python environment

Miniconda + `uv` is the combination that causes the fewest headaches:

```bash
conda create -p "$SCRATCH/py312" python=3.12 -y
conda activate "$SCRATCH/py312"
pip install uv
uv cache dir   # confirm it points at $SCRATCH
```

Or, if you want the project-local venv pattern:

```bash
uv venv
source .venv/bin/activate
```

Match PyTorch to the loaded CUDA:

```bash
uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
```


## SGLang install

Inside a GPU session so the build-time CUDA probes succeed:

```bash
export TORCH_CUDA_ARCH_LIST="9.0"    # H100 on m3h
uv pip install "sglang[all]>=0.5.0"
uv pip install --force-reinstall "sgl-kernel==0.2.4" sentencepiece
```

Quick sanity check:

```bash
python -m sglang.bench_offline_throughput --help
```


## Sanity checks on a GPU node

Run these as the first thing inside every allocation, catches 80% of "why isn't my job working" questions before they happen:

```bash
hostname
nvidia-smi -L
nvidia-smi
echo "CUDA_VISIBLE_DEVICES=$CUDA_VISIBLE_DEVICES"

python3 - << 'EOF'
import torch
print("torch:", torch.__version__)
print("cuda:", torch.version.cuda)
print("available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
for i in range(torch.cuda.device_count()):
    print(i, torch.cuda.get_device_name(i))
EOF
```


## Slurm batch template

For a 1-node, 4-GPU H100 job on `m3h`:

```bash
#!/bin/bash
#SBATCH --job-name=sglang
#SBATCH --partition=m3h
#SBATCH --qos=m3h
#SBATCH --time=00:30:00
#SBATCH --gres=gpu:4
#SBATCH --cpus-per-task=72
#SBATCH --mem=985G
#SBATCH --output=run/%x.%j.out

source "$SCRATCH/py312/bin/activate"
module load cuda/12.6

nvidia-smi -L
time python3 -m sglang.bench_offline_throughput \
  --model-path deepseek-ai/DeepSeek-R1 \
  --dataset-path "$SCRATCH/ShareGPT_V3_unfiltered_cleaned_split.json" \
  --num-prompts 2000 --load-format dummy --seed 2025 --dtype bfloat16 \
  --tp 4 --trust-remote-code \
  --warmups compile-deep-gemm
```

Submit with `sbatch`, monitor with `squeue -u $USER`, tail the log once it starts.


## Things I wish I'd known

- **`$SCRATCH` is not automatic**, the variable isn't set for you. Export it yourself in every script.
- **`--qos=m3h` is mandatory on `m3h`**, leave it off and the job sits in the queue forever without a clear error.
- **`smux` sessions count against your active-job quota**, don't leave five of them running.
- **Module order matters**, load `cuda/12.6` *before* activating your conda env, otherwise `LD_LIBRARY_PATH` gets clobbered.
- **`/apps` is read-only**, everything you install goes in `$SCRATCH`.
- **H100s on `m3h` are contended at end of semester**, submit early, especially for multi-GPU jobs.


## See also

- [Aspire 2A+](/blog/aspire-2a-plus), NSCC Singapore equivalent, PBS Pro rather than Slurm.
- [SGLang & DeepSeek](/blog/sglang-and-deepseek), install and benchmark flags.
- [SGLang Optimisations](/blog/sglang-optimisations), the flags that actually move throughput.
