Tuesday Aug 11

Meta Fits A Model On One GPU

11AUG
30B PARAMS1 GPU

Meta open-sourced a model that runs on a gaming PC. Muse Glimmer packs 30 billion parameters into 24 gigabytes of memory using 4-bit compression. It matches bigger closed models on coding and math tests.

No cloud, no API key, no per-token bill. Muse Glimmer runs fully offline on a single consumer GPU.

It scores 94.7 on the AIME math benchmark and 51.2 on SWE-Bench Pro coding. A speculative-decoding trick called DFlash triples the output speed on an RTX 5090. Apache 2.0 means anyone can build on it for free.

This is the local-agent argument getting real: your laptop, not a data center. Enterprises worried about sending data to the cloud now have a credible offline option.

full brief & sources

⚡ Why this matters

  • Open-weight models are catching up to closed frontier models fast.
  • Running locally kills the data-privacy objection enterprises raise about cloud AI.
  • It's a real alternative to paying per-token for agentic coding work.

🔍 What happened

  • Meta Superintelligence Labs released Muse Glimmer on Aug 10 under Apache 2.0.
  • 30B parameters, distilled from Meta's larger Muse model, with a built-in vision encoder.
  • 4-bit quantized versions fit 24GB and 32GB consumer GPU memory.
  • Scores category-best on MCP Atlas, SWE-Bench Pro, AIME 2026, and Charxiv Reasoning.
  • DFlash speculative decoding lifts an RTX 5090 from 74.9 to 233.4 tokens per second.
  • Ships with a 2B vision encoder feeding a 28B text decoder.

💬 Smart takes

  • Meta: positions this as proof open models can match closed ones on agentic tasks.
  • Skeptic: benchmark-best claims from the model's own maker deserve independent verification before belief.

🧭 Where this goes

  1. Likelyother labs respond with their own compact, GPU-local agent models within weeks.
  2. Possibleenterprises pilot Muse Glimmer for on-premise coding agents where data can't leave the building.
  3. Wild Carda compact open model like this ends up embedded directly in a laptop OS.

🥄 The Spoon Take

The frontier used to mean the biggest model money could rent by the hour. Now it also means a 30B model that fits on the GPU you already own. That's a second front opening in the model wars: not just smartest, but smallest-that's-still-good-enough.

🤔 Pushback

Self-reported benchmark scores from the lab that built the model aren't the same as independent evaluation.