Founding engineer · AI systems · HPC research

I build ambitious
AI products and the
systems behind them.

I’m Luca Lowndes, founding and lead engineer at EvryStay. I built the product as its only engineer, helped take it through a A$4.5M raise, and now lead a team of five while continuing to ship.

I’m also a first author and speaker at ICPP 2026 for work on serving DeepSeek-R1 671B with SGLang, a credited SGLang contributor, and the youngest UniHack winner on record.

Founding engineer ICPP 2026 first author SGLang contributor UniHack winner APAC HPC-AI 3rd place Founding engineer ICPP 2026 first author SGLang contributor UniHack winner APAC HPC-AI 3rd place
01

Experience

From zero to product,
then product to team.

I work closest to the hard boundary between research, product, and production: turning an idea into something people can depend on.

Mar 2025 — Nov 2025

Melbourne, VIC

Monash DeepNeuron & Monash eResearch

Lead AI Researcher, APAC HPC-AI Competition Team

Optimised distributed inference for DeepSeek-R1 across two 8×H200 nodes using infrastructure from NSCC Singapore, NCI Australia, and Firmus.

The finding

A deliberately configured single node beat a default 16-GPU setup: PP2 · TP4 · DP4 attention reached 17,417 tok/s and 2.05× the per-GPU efficiency.

The lesson

The “boring” systems work mattered most. CUDA version alignment alone delivered +35%, while more than 40 hand-tuned NCCL configurations never beat the defaults.

Outcome Presented at SupercomputingAsia 2026 in Osaka; the work became the ICPP paper.

Feb 2025 — Present

Melbourne, VIC

Monash DeepNeuron

AI & HPC Technical Team, Outreach

Researching Neural Cellular Automata for self-organising biological systems and running practical AI workshops in Victorian high schools.

Jan 2025 — Nov 2025

Hybrid

Scag Australia · International Mowers

Freelance AI Engineer

Built customer-facing AI agents for model comparison, troubleshooting, and service recommendations, with structured proprietary data, embeddable web/mobile chat, brand persona, and legal guardrails.

02

Research

Less is more.

Our results challenge the instinct to distribute the biggest model across every available GPU. Topology-aware parallelism was both faster and more efficient.

ICPP 2026 · BID ’26 First author

Less is More: Optimising SGLang Distributed DeepSeek-R1 Inference on a Two-Node H200 Cluster

Luca Lowndes · Nathan Culshaw

Accepted at Benchmarking in the Data Center, a workshop of the 55th International Conference on Parallel Processing. To appear in the ACM Digital Library.

+81% throughput over
16-GPU baseline
Best result
17,417 tok/s
Hardware
8× NVIDIA H200
Model
DeepSeek-R1 671B
Serving stack
SGLang
Core insight

Keep collectives on NVLink and replicate within a node instead of sharding critical communication across InfiniBand.

01

Talk · Singapore · Sep 2026

Presenting “Less is More” at BID ’26

Speaking alongside representatives from NVIDIA and the HPC-AI Advisory Council.

02

Panel · ICPP 2026

Accelerating the Future

Panellist on overcoming bottlenecks in GPU-based AI and HPC, with NVIDIA and the HPC-AI Advisory Council.

03

Talk · Osaka · Jan 2026

SupercomputingAsia 2026

Presented the APAC HPC-AI Competition work after placing third among 49 university teams.

03

Selected projects

Useful things,
built all the way through.

Agents that write production-shaped code, models that grow biological structure, and experiment records designed to survive scrutiny.

2026 Company agent

Dummy

Describe a feature in Slack. Dummy builds a working version.

It creates real branches across three repos, a seeded full-stack preview at a permanent URL, desktop and mobile QA screenshots, and draft pull requests.

Slack thread
Plan + build
Review agent
Preview + PRs
2025 Hackathon winner

Growth Garden

Habit tracking as a garden you grow by getting things done.

On a team of six, I architected the contextual AI assistant in Langflow, built the DataStax model and live sync, and led the winning pitch.

1st of 135+ projects
2025 — Present Research

Neural Cellular Automata

Learning complex growth from simple local rules.

Built the 3D temporal training path for time-resolved mouse embryo data to explore gastrulation-like development, including reduced voxel handling, dynamic frame batches, differentiable IoU, loss instrumentation, PyTorch models, and browser inference work.

Visit neuralca.org
2025 Open experiment record

APAC-AI-HPC-2025

The reproducible record behind every number in the ICPP paper.

Launch scripts, Slurm logs, configuration notes, and workbooks for every experiment across two 8×H200 nodes.

github.com/LucaLow/APAC-AI-HPC-2025 Open repository
04

Recognition & education

A fast start,
not a finished story.

I deferred university in 2026 to build EvryStay properly, while continuing research, open-source work, and conference speaking.

Awards

04
03

APAC HPC-AI Competition

3rd of 49 university teams · Presented at SupercomputingAsia, Osaka

01

UniHack 2025

1st of 850+ participants · Youngest winner on record

HD

Engineering Dean’s Award

Academic Excellence · Top of the FIT1045 cohort, Monash

D

Dux of Mathematical Methods

Elwood College

Education

02

Monash University

BEng (Hons) Mechatronics / BCS Algorithms & Software

  • High Distinction WAM
  • Deferred in 2026 for EvryStay

Academy of Interactive Entertainment

Certificate III in Information Technology

  • Completed alongside VCE
05

Technical toolkit

Deep enough for the kernel.
Broad enough for the product.

01

Inference & HPC

SGLang, CUDA, NCCL, tensor parallelism, pipeline parallelism, data parallelism, Slurm, Singularity, NVIDIA H100/H200

02

Machine learning

PyTorch, TensorFlow, neural cellular automata, distributed model serving, evaluation, loss design

03

Product engineering

TypeScript, Next.js, Python, Postgres, LLM agents, tool use, multi-repo systems, production architecture

04

Systems & delivery

C, C++, Git, CI/CD, code review, technical leadership, hiring, architecture ownership

Let’s talk

Building something
worth obsessing over?

Research, product, distributed AI systems, or the point where all three collide.

Luca@lowndes.net