Saturday Aug 8

ByteDance Trains A 10 Trillion Model

8AUG
8T SHIPPED10T BUILDING

The Financial Times reports ByteDance is pre-training a model of up to 10 trillion parameters. Bigger than Anthropic's Mythos 5. Founder Zhang Yiming told the team not to distill from rivals.

That ceiling is under consideration, not a shipped result. The training run alone takes three to six months.

The no-copying instruction is the part worth reading twice. A US official recently accused Moonshot of lifting weights from a rival lab while building Kimi K3. Skipping that shortcut is slow and expensive.

Size is a spending signal, not a capability one. A 2,000-person team is running it. Judge the result when something ships.

full brief & sources

⚡ Why this matters

  • It is the largest disclosed training run from a Chinese lab and a direct answer to distillation accusations.
  • Refusing distillation is a costly choice that says something about how ByteDance wants to be seen.
  • It resets the scale conversation at a point when many assumed scaling had given way to efficiency.

🔍 What happened

  • The Financial Times reports ByteDance is pre-training a model of up to 10 trillion parameters.
  • The work is run by ByteDance's roughly 2,000-person Seed team.
  • That is about three times the size of Moonshot's Kimi K3 and above estimates for Anthropic's Mythos 5 at 8 trillion.
  • Founder Zhang Yiming directed the team to avoid model distillation entirely.
  • The direction follows a US official's allegation that Moonshot distilled Anthropic's Fable model.
  • The model is in pre-training, a phase that typically runs three to six months before fine-tuning.

💬 Smart takes

  • Financial Times: reported the training run and the scale target via people familiar with the work.
  • XenoSpectrum: argued raw parameter counts cannot measure the actual gap with Anthropic.
  • Skeptic: ten trillion is described as an upper bound under consideration, which is not a commitment.
  • Skeptic: parameter count stopped predicting benchmark performance somewhere around 2024.

🧭 Where this goes

  1. Likelyno weights or benchmarks appear before late 2026 given the pre-training timeline.
  2. Likelyother Chinese labs disclose their own scale targets in response.
  3. Possiblethe shipped model lands well below 10 trillion after efficiency work during training.
  4. Wild CardByteDance publishes training provenance documentation to prove the no-distillation claim.

🥄 The Spoon Take

The number is not the story. The instruction is. Distillation is cheap, fast, and increasingly treated as theft, and ByteDance just chose the expensive path in public. That reads less like a research decision than a positioning one, aimed at regulators and partners who have started asking where a model's capabilities actually came from.

🤔 Pushback

Nobody outside ByteDance can verify a no-distillation claim, which makes it a statement of intent rather than a fact.