An AI model company buys wholesale electricity by the megawatt & resells it as cognitive work.

“The base cost of compute tends to be around 10 or 13 or $15 million per megawatt. In the case of Anthropic, the revenue has gone as high as $50 million per megawatt. And what that now enables them to do is, hey, if I spend 10 bucks on inference capacity, I actually generate 50 bucks of revenue. And then I can turn around and incrementally spend all of that profit on training.”

— Dylan Patel, SemiAnalysis, on the Dwarkesh Podcast1

Gross profit measured as a function of electricity proves model companies can be profitable on a contribution basis.

Anthropic’s gross margin was −94% in 2024 : $1.94 of compute for every $1 of revenue.2

Revenue per megawatt crossing above the cost of compute, from a −94% gross margin in 2024 to $50m per megawatt in 2026

By 2025 the corner turned, & Anthropic swung from −94% to a 40-50% gross margin.

In 2026 revenue passed cost : $50m per megawatt against a $10-15m cost. Anthropic booked its first profitable quarter, $10.9b of revenue & $559m of operating profit.

That 5% operating margin sits well below the 70 to 80% gross margin the megawatt math implies. Training runs & headcount consume the difference.

Anthropic's gross margin swung from −94% to profitable in two years

Gross profit per megawatt is not just about intelligence, but also efficiency. A model that serves the same intelligence at a fraction of the compute generates more profit per megawatt, even at a lower price.

GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index : identical to Claude Opus 4.8.3 But it does it on 18 billion active parameters, at a 90 to 97% reduction in cost. Fewer active parameters & less attention compute means more tokens per megawatt.

Waterfall of the Artificial Analysis Intelligence Index frontier climbing from 37 to 63 points in waves, with three jumps of at least 3 points in orange and smaller steps in steel

The frontier itself keeps climbing in waves. Three jumps of 3 points or more carry half of the 25-point gain from 37 to 63. The efficient models chase a target that resets every quarter.

Which is what the profit per megawatt buys. Patel’s last clause is the hinge : “turn around and incrementally spend all of that profit on training.” The margin funds the factory.

Nvidia paid $6b to acquire Poolside & invested another $1b on that premise. Model building becomes an industrial process : thousands of experiments across a search space, not artisanal hand-tuning.

“The model is the output. The ability to keep building better models, faster & more efficiently each time is the actual innovation. The Model Factory is the compounding asset.”

— Jason Warner, Poolside4

Laguna S 2.1 went from kickoff to release in 52 days.

Inference margin funds the factory that makes the next model cheaper to build & more efficient to run.


  1. Dylan Patel on the Dwarkesh Podcast, “Anthropic & OpenAI will have most of the world’s compute by 2028” (Aug 2026). Apple Podcasts ↩︎

  2. The Information, “Anthropic’s Gross Margin Flags Long-Term AI Profit Questions.” The Information ↩︎

  3. Artificial Analysis Intelligence Index. Artificial Analysis ↩︎

  4. Jason Warner, Poolside. X ↩︎