Founding and lead engineer at EvryStay, where I built the product as the only engineer, helped take it through a A$4.5M raise at a A$17M post-money valuation, and now lead a team of five.
First author and speaker at ICPP 2026 on serving DeepSeek-R1 (671B) with SGLang, work that grew out of placing 3rd in the APAC HPC-AI Competition. Credited SGLang contributor. Youngest ever UniHack winner. Dean’s Award for Academic Excellence in first year. Deferred my double degree at Monash to do this properly. All by 20.
I believe in keeping code simple and data structures lightweight. Whether I’m optimising kernels, shipping product, or training for my next marathon, I’m driven by efficient progress.
8th APAC HPC-AI Competition
3rd of 49 university teams across the Asia-Pacific region, presented on stage at SupercomputingAsia 2026 in Osaka
Lead AI Researcher: DeepSeek-R1 671B (37B active) on H100 (NSCC ASPIRE-2A+) and two 8× NVIDIA H200 nodes
Fewer GPUs arranged deliberately beat more GPUs arranged by default: a tuned single node (PP2 · TP4 · DP4 attention) hit 17,417 tok/s, 81% over the 16-GPU baseline and 2.05× the per-GPU efficiency
CUDA version alignment alone was worth +35%. 40+ hand-tuned NCCL configurations never beat the defaults
HPC Track: optimised NWChem Fortran chemistry simulations across multi-node CPU clusters
Team: Josh Riantoputra (Captain), Nathan Culshaw (HPC Lead), Isaac Barnes, Giacomo Bonomi; mentor Simon Michnowicz
The work became the Firmus and HPC-AI Advisory Council partnership on SGLang best practices, and the ICPP 2026 paper
UniHack 2025
1st of 850+ participants and 135+ projects, youngest winner on record
Growth Garden, AI-powered personal development app
Architected the conversational AI system in Langflow with DataStax, part of the frontend, and the winning pitch
Less is More: Optimising SGLang Distributed DeepSeek-R1 Inference on a Two-Node H200 Cluster
Accepted at Benchmarking in the Data Center (BID ’26), a workshop of the 55th International Conference on Parallel Processing
Shows that keeping collectives on NVLink and replicating within a node beats sharding across InfiniBand: 17,417 tok/s, 81% over the 16-GPU baseline
Presenting Less is More at BID ’26, ICPP 2026
In a session alongside speakers from NVIDIA and the HPC-AI Advisory Council
Accelerating the Future: Overcoming Bottlenecks in GPU-Based AI and HPC
Panellist with Jeff Adie (NVIDIA), Pengzhi Zhu (Director, HPC-AI Lab, HPC-AI Advisory Council) and Sandeep (IIT Palakkad)
Moderated by Dr Samar Aseeri and Dr Jens Domke
3rd of 49 university teams, APAC HPC-AI Competition/ Presented at SupercomputingAsia 2026, Osaka
1st of 850+ participants, UniHack 2025/ Youngest winner on record
Dean’s Award for Academic Excellence/ Monash University
Dean’s Honours List/ Monash University
Highest Grade in Cohort, FIT1045/ Monash University
Dux of Mathematical Methods/ Elwood College
EvryStay
Joined as the first and only engineer. Designed and built Eve, an AI concierge that handles guest conversations end to end for short-term rental hosts, property managers and hotels, plus the property management system integrations and AI service architecture underneath it
The platform I built was the product investors backed: EvryStay raised A$4.5M at a A$17M post-money valuation, opening and closing the round inside a week
Hired and now lead five full-time engineers in person. I own the architecture, the AI service boundary, code review and the final technical call, while still shipping daily: 300+ merged PRs in my first seven months
Built Dummy, our Slack-native company agent that turns a feature idea into branches, a full-stack preview and draft PRs, so feedback happens on something concrete and the idea to shipped feature cycle is much shorter
Firmus Technologies & HPC-AI Advisory Council
Contributed optimisations to SGLang, the inference framework behind hundreds of thousands of GPUs worldwide, credited in the v0.5.5 release notes
Wrote deployment guidance for running SGLang on H200 clusters as part of a joint effort between Firmus, the HPC-AI Advisory Council, Monash eResearch and Monash DeepNeuron
Monash DeepNeuron & Monash eResearch
Maximised inference throughput for DeepSeek-R1 (671B parameters, 37B active) across two 8× H200 nodes, on infrastructure from NSCC Singapore, NCI Australia and Firmus
Fewer GPUs arranged deliberately beat more GPUs arranged by default: a tuned single node (PP2 · TP4 · DP4 attention) hit 17,417 tok/s at 2.05× the per-GPU efficiency of 16 GPUs
The boring stuff mattered most: CUDA version alignment alone was worth +35%, and 40+ hand-tuned NCCL configs never beat the defaults
3rd of 49 university teams, presented on stage at SupercomputingAsia 2026 in Osaka. The work became the Firmus partnership and the ICPP 2026 paper
Monash DeepNeuron
Implemented Neural Cellular Automata methods from recent papers in PyTorch to model self-organising systems such as embryonic gastrulation, plus on-device and cost-function optimisations using WebGL
Ran hands-on workshops in Victorian high schools on how deep learning models work and how to use generative AI well
Jemma
Melbourne, VIC, best role to date
Freelance AI Engineer
Built the Scag Mechanic and IM Advisor assistants for model comparison, troubleshooting and service recommendations
Structured proprietary product data, shipped embeddable chat widgets for web and mobile, and put brand persona and legal guardrails around both
DummyView →
EvryStay’s company agent, and it lives in Slack. Built so the non-technical side of the company can see their ideas before engineers spend time on them: describe a feature in a thread and Dummy builds a working version, with real branches across our repos, a full-stack preview at a permanent URL, desktop and mobile QA screenshots, and draft PRs. An independent review agent checks the code, sensitive changes need an allowlisted human to approve in Slack, and nothing it does can merge, deploy, or touch production. Built on the open source QM harness.
Swarm Inference LabView →
Contributor to a self-configuring cluster runtime for heterogeneous LLM inference over a LAN: immutable content-addressed model snapshots, artifact leases, and an Apple Silicon multi-worker stage ring
APAC-AI-HPC-2025View →
The full open experiment record behind the ICPP paper: launch scripts, Slurm logs and workbooks. Every number is reproducible
PiVoiceView →
Local speech-to-speech voice interface for terminal coding agents, running entirely on Apple Silicon
Growth GardenView →
AI-powered personal development app. Winner of UniHack 2025, Australia’s largest student hackathon, 1st of 850+ participants
Neural Cellular AutomataView →
Developing NCA models to generate complex emergent patterns and explore self-organising biological systems using PyTorch
BookMarkerView →
Privacy-focused Chrome extension that remembers your exact scroll position on any webpage
High-powered performance: DeepNeuron computes its way to third place
Monash University students triumph at leading competition UNIHACK 2025
HPC-AI Advisory Council Announces Results of 8th APAC HPC-AI Competition
Monash students secure third place at APAC HPC-AI Competition
Monash students triumph at UNIHACK 2025 with Growth Garden
Australia’s Biggest Student Hackathon Crowns Growth Garden as Champion
Street Corn Cultivates Victory at UNIHACK 2025
Monash University
Bachelor of Engineering (Honours), Mechatronics / Bachelor of Computer Science, Algorithms & Software
High Distinction WAM · Dean’s Award for Academic Excellence · Top of cohort, FIT1045 · Deferred 2026 to lead engineering at EvryStay
Elwood College
VCE, Dux of Mathematical Methods
AIE
Certificate III in Information Technology, completed alongside VCE
Harvard University
CS50x, Introduction to Computer Science, completed at 16
Machine Learning Specialization/ Stanford / Andrew Ng
Deep Learning Specialization/ DeepLearning.AI