The Inference Shift
27MAY
Ben Thompson, Stratechery analyst, splits AI compute into two distinct kinds of inference. "Answer inference" has a human in the loop and speed matters. "Agentic inference" runs without a human, so latency is tolerable. The agentic side will be the bigger market - and it does NOT need Nvidia's speed premium.
Two kinds of inference. Two kinds of chip. Two kinds of datacenter. The market split everyone missed.
When a human is waiting, speed matters and Nvidia's HBM advantage is worth the premium. When an agent is doing overnight work, latency is fine. Slower DRAM, slower chips, slower locations all become viable. Thompson reads this as good news for Chinese fabs (no speed crown) and orbital compute (light-second delays acceptable). Bad news for Nvidia's pricing power on the agentic half.
For PMs: stop assuming one compute roadmap fits all your workloads. For execs procuring AI: split your inference RFPs into answer-tier and agentic-tier. For investors: the "Nvidia at any multiple" trade just got narrower.