An AI model company buys wholesale electricity by the megawatt & resells it as cognitive work.
“The base cost of compute tends to be around 10 or 13 or $15 million per megawatt. In the case of Anthropic, the revenue has gone as high as $50 million per megawatt. And what that now enables them to do is, hey, if I spend 10 bucks on inference capacity, I actually generate 50 bucks of revenue. And then I can turn around and incrementally spend all of that profit on training.”
— Dylan Patel, SemiAnalysis, on the Dwarkesh Podcast1
Gross profit measured as a function of electricity proves model companies can be profitable on a contribution basis.
Anthropic’s gross margin was −94% in 2024 : $1.94 of compute for every $1 of revenue.2
By 2025 the corner turned, & Anthropic swung from −94% to a 40-50% gross margin.
In 2026 revenue passed cost : $50m per megawatt against a $10-15m cost. Anthropic booked its first profitable quarter, $10.9b of revenue & $559m of operating profit.
That 5% operating margin sits well below the 70 to 80% gross margin the megawatt math implies. Training runs & headcount consume the difference.
Gross profit per megawatt is not just about intelligence, but also efficiency. A model that serves the same intelligence at a fraction of the compute generates more profit per megawatt, even at a lower price.
GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index : identical to Claude Opus 4.8.3 But it does it on 18 billion active parameters, at a 90 to 97% reduction in cost. Fewer active parameters & less attention compute means more tokens per megawatt.
The frontier itself keeps climbing in waves. Three jumps of 3 points or more carry half of the 25-point gain from 37 to 63. The efficient models chase a target that resets every quarter.
Which is what the profit per megawatt buys. Patel’s last clause is the hinge : “turn around and incrementally spend all of that profit on training.” The margin funds the factory.
Nvidia paid $6b to acquire Poolside & invested another $1b on that premise. Model building becomes an industrial process : thousands of experiments across a search space, not artisanal hand-tuning.
“The model is the output. The ability to keep building better models, faster & more efficiently each time is the actual innovation. The Model Factory is the compounding asset.”
— Jason Warner, Poolside4
Laguna S 2.1 went from kickoff to release in 52 days.
Inference margin funds the factory that makes the next model cheaper to build & more efficient to run.
-
Dylan Patel on the Dwarkesh Podcast, “Anthropic & OpenAI will have most of the world’s compute by 2028” (Aug 2026). Apple Podcasts ↩︎
-
The Information, “Anthropic’s Gross Margin Flags Long-Term AI Profit Questions.” The Information ↩︎
-
Artificial Analysis Intelligence Index. Artificial Analysis ↩︎