Thursday Aug 6

SK Hynix And Sandisk Standardize AI Flash

6AUG
512 GB PER STACKHBMHBFSSD

AI chips lean on HBM, the fast but scarce memory stacked next to GPUs. SK hynix and Sandisk published the first open spec for High Bandwidth Flash, a cheaper middle tier above SSDs.

The spec went through the Open Compute Project at the FMS summit: up to 512 GB per stack from 8-high and 16-high NAND dies, with bandwidth grades from 0.4 to 3.0 terabytes per second.

AI inference is memory-bound and HBM supply is every roadmap's choke point. HBF trades speed for capacity: model weights and KV caches sit on flash, not scarce DRAM.

The consortium includes Google and Tenstorrent, and any vendor can build against the spec. For NAND makers it is the first on-ramp into AI money that went to HBM and DRAM.

full brief & sources

⚡ Why this matters

  • AI inference is memory-bound; HBM cost and supply now gate every accelerator roadmap.
  • An open capacity tier changes inference economics: long context and big models get cheaper to serve.
  • It is the NAND industry's route into AI capex that has bypassed it so far.

🔍 What happened

  • SK hynix and Sandisk published the first High Bandwidth Flash specification through the Open Compute Project at FMS 2026 in Santa Clara, Aug 4-6.
  • HBF stacks NAND flash dies the way HBM stacks DRAM: up to 512 GB per stack in 8-high and 16-high configurations.
  • Three bandwidth grades, from roughly 0.4 TB/s to 3.0 TB/s per stack.
  • Positioning: HBM-like bandwidth with NAND-like capacity, a new tier between HBM and SSDs in the memory hierarchy.
  • The consortium behind the spec includes Sandisk, SK hynix, Google, and Tenstorrent.

💬 Smart takes

  • The consortium: an open OCP spec means any GPU or accelerator vendor can design against HBF without betting on a single supplier.
  • The inference view: KV caches and model weights are the target workloads, where capacity per dollar beats raw latency.
  • Skeptic: NAND is still orders of magnitude slower than DRAM on latency, a spec is not a product, and Samsung and Micron are absent from the table.

🧭 Where this goes

  1. Likelyfirst HBF silicon samples from SK hynix and Sandisk within 12 months, aimed at 2027-2028 inference systems.
  2. Likelyaccelerator vendors add HBF controllers alongside HBM interfaces; Tenstorrent moves first since it is already in the consortium.
  3. LikelySamsung, Micron, and Kioxia respond, either joining the OCP spec or pushing rival capacity-tier designs.
  4. Possiblecloud providers spec HBF into inference fleets for long-context serving before 2028.
  5. Wild CardHBF becomes the default weight-storage tier and caps HBM growth, shifting AI memory economics away from DRAM vendors.

🥄 The Spoon Take

This is plumbing news, which is exactly why it matters. Inference cost is set by memory, not FLOPs, and HBM is the bottleneck everyone prices around. An open, second-sourced capacity tier is the kind of boring standard that quietly resets the cost curve, if the parts actually ship.

🤔 Pushback

Flash endurance and latency may confine HBF to niche read-heavy workloads, and without Samsung and Micron the standard could stay a two-vendor spec.