Sunday Aug 9

AMD Buys Chips With Models Baked In

9AUG
17K TOKENS/SECBAKED INNO UPDATES

Some AI chips will stop being general purpose. AMD is buying Taalas, a Toronto startup that burns a model's weights permanently into transistors. No memory reads, no reprogramming, ten times less power.

Every token a GPU makes requires pulling all model weights out of memory. That read is the speed ceiling. Taalas removes it by turning the weights themselves into hardware.

The demo hit 17,000 tokens per second on Llama 3.1 8B, at one-tenth the power draw of an Nvidia H200. Founder Ljubisa Bajic is a former AMD and Nvidia architect who co-founded Tenstorrent.

The catch is obvious. A chip with a model fused into it cannot run the next model. That is a bet on which weights are worth freezing.

full brief & sources

⚡ Why this matters

  • First time a major GPU vendor has bought model-in-silicon technology.
  • Inference demand now exceeds training demand, and inference is where fixed silicon wins.
  • Model lifespan becomes a hardware procurement question, not just a research one.

🔍 What happened

  • AMD agreed to acquire Taalas, founded in Toronto in 2023. Terms undisclosed.
  • Taalas hard-wires trained model weights directly into transistors, removing DRAM reads from the inference path.
  • Claimed throughput is 17,000 tokens per second on Llama 3.1 8B at one-tenth an H200's power draw.
  • Founders are Ljubisa Bajic, Drago Ignatovic and Lejla Bajic. Bajic co-founded Tenstorrent and previously worked at AMD and Nvidia.
  • Taalas raised $219 million in total, including $169 million in February 2026, with over $170 million still unspent at signing.
  • The deal is expected to close in the fourth quarter of 2026, pending regulatory approval.

💬 Smart takes

  • The Register: the acquisition targets inference performance by etching models into silicon rather than scaling general-purpose accelerators.
  • Hiroki Miyano, AI newsletter writer: Anthropic's custom-silicon news and AMD's Taalas deal landing the same week is probably not a coincidence.
  • Skeptic: models still change every few months, so a chip that cannot be reprogrammed is a depreciating asset on a very short clock.

🧭 Where this goes

  1. LikelyNvidia answers with a fixed-function inference part or an acquisition of its own inside 12 months.
  2. Likelyfixed-silicon inference lands first on small, stable open-weight models, not frontier ones.
  3. Possiblemodel providers start publishing long-term-support versions so hardware partners can commit.
  4. Wild Carda frontier lab freezes one model as a hardware standard and sells it as a permanent low-cost tier.

🥄 The Spoon Take

The industry has been buying general-purpose compute because nobody knew which model would matter. Fixing weights into silicon is a bet that some models will stop changing. That bet is worth more than the chip.

🤔 Pushback

Etched silicon only pays off if a model stays useful for years, and nothing in the last three has.