AMD Buys Chips With Models Baked In
9AUG
Some AI chips will stop being general purpose. AMD is buying Taalas, a Toronto startup that burns a model's weights permanently into transistors. No memory reads, no reprogramming, ten times less power.
Every token a GPU makes requires pulling all model weights out of memory. That read is the speed ceiling. Taalas removes it by turning the weights themselves into hardware.
The demo hit 17,000 tokens per second on Llama 3.1 8B, at one-tenth the power draw of an Nvidia H200. Founder Ljubisa Bajic is a former AMD and Nvidia architect who co-founded Tenstorrent.
The catch is obvious. A chip with a model fused into it cannot run the next model. That is a bet on which weights are worth freezing.