Wednesday Sep 23

Apple Sells AI Without A Token Meter

23SEP
1 PLUG, 1T PARAMSMAC STUDIONO METER

Johny Srouji's pitch for the new Macs: buy the box, run the model, pay nobody per token. Four Mac Studios ran a trillion-parameter model from one wall outlet. Nvidia declined to comment.

The M5 Ultra Mac Studio starts at $5,499. The 256GB memory version with 16TB of storage costs $18,299. Apple's argument is arithmetic: one upfront invoice versus a cloud bill that never stops.

Apple holds 4.6 percent of enterprise desktops. Windows holds 91.3 percent, per IDC's Linn Huang. Microsoft hosts a Windows event next month, and Satya Nadella has already been talking about unmetered intelligence.

The target is the cloud AI business model itself. Every API price cut this week is measured in tokens. Apple wants the unit of account to be hardware instead.

full brief & sources

⚡ Why this matters

  • Two labs cut token prices on the same day Apple said tokens should not have a price. That is a fight over the unit of account, not over chips.
  • Local inference on a desk changes who signs the contract: IT hardware budgets instead of cloud commits. Different buyer, different sales motion.
  • If a trillion-parameter model runs from a wall outlet, the data center's moat is latency and scale, not capability.

🔍 What happened

  • Apple hardware chief Johny Srouji told Reuters on September 22: "There's no cost per token. You're just using the machine again and again."
  • Apple demonstrated four Mac Studios running a trillion-parameter model as one cluster, powered from a single wall outlet.
  • The M5 Ultra Mac Studio starts at $5,499. A configuration with 256GB of memory and 16TB of storage costs $18,299.
  • IDC's Linn Huang puts Apple at 4.6 percent of enterprise desktops versus 91.3 percent for Windows. Microsoft holds a Windows event next month.
  • Satya Nadella has used the phrase unmetered intelligence for Microsoft's own direction. Nvidia declined to comment on Apple's claims.

💬 Smart takes

  • Johny Srouji, Apple: "There's no cost per token. You're just using the machine again and again." Twelve words that reprice the whole category.
  • Linn Huang, IDC: the enterprise desktop is still 91 percent Windows. Apple's AI pitch has to beat procurement habits before it beats Nvidia.
  • Skeptic: a trillion-parameter model on four Macs runs one user at a time. The cloud sells concurrency. Apple is selling a very fast single seat.

🧭 Where this goes

  1. LikelyMicrosoft answers at next month's Windows event with local-model hardware claims of its own.
  2. LikelyMac Studio clusters become the default for law firms and studios that cannot send data to a cloud.
  3. PossibleAnthropic or OpenAI license a distilled model to run natively on Apple silicon.
  4. Wild CardApple publishes a cost-per-token comparison against the cloud labs and starts a pricing fight it usually avoids.

🥄 The Spoon Take

Srouji is not selling a computer. He is selling the end of the meter. The cloud labs spent Tuesday cutting per-token prices, which concedes the point: the meter is the problem. Apple's bet is that a CFO would rather buy an $18,000 box once than sign a bill that scales with success. For a lot of workloads, the CFO is right.

🤔 Pushback

Local inference serves one team at a time. Most enterprise AI demand is bursty and concurrent, which is exactly what the cloud is good at.