Skip to content
All writing

2 min

Sixteen chips just lost to eight

My paper "Less is More" has been accepted for presentation at the BID Workshop at ICPP 2026.

View MarkdownOpen in ClaudeOpen in ChatGPT

My paper "Less is More: Optimising SGLang Distributed DeepSeek-R1 Inference on a Two-Node H200 Cluster" has been accepted for presentation at the BID Workshop at ICPP 2026, the International Conference on Parallel Processing.

I took a 671-billion-parameter AI model and made it run 81% faster on half the hardware. The industry's answer to every AI problem is to buy more chips. Measured properly, the chips were spending their time talking to each other instead of thinking. Understanding beat spending.

(Technical crowd: TP=4, DP=4, PP=2 with DP-attention, 17,417 tokens per second.)

The biggest cost inside any AI product is inference. Every guest message our concierge answers at EvryStay costs compute, and the difference between an AI company with real margins and one burning runway is exactly the discipline in this paper. That's why EvryStay treats performance engineering as product strategy, not plumbing. It's why our pre-seed opened and closed inside a week. It's why a product built from a beachside office in Mornington is instant for guests and priced so hotels actually say yes. And it's why I'd rather build here, with real ownership of every system I touch, than be one of ten thousand engineers somewhere else.

Co-authored with Nathan Culshaw. The experiments ran on Firmus Technologies' H200 cluster. If you're in hospitality and want to see what obsessive engineering feels like as a guest experience, my DMs are open.

Originally posted on LinkedIn