Monday Aug 17

Qwen 3.8 Thinks For 21 Minutes

17AUG
XHIGH MODEPELICAN

Alibaba's new open model runs on a laptop and codes, sees, and calls tools. Simon Willison calls it the best local model he's tested. One catch: the default setting overthinks everything.

Qwen 3.8 27B is a 17GB file. It handles vision, tool use, and a 262,000-token context. Willison ran it on a MacBook and an NVIDIA Spark.

The default reasoning mode burns tokens on everything. One pelican drawing took 21 minutes and 22,000 thinking tokens. Turn reasoning down and the same job takes two minutes.

It also drove a real coding agent and nailed image bounding boxes. Willison's verdict: a miracle file, held back only by speed.

full brief & sources

⚡ Why this matters

  • A 17GB open file now covers vision, coding, tool use, and long context on consumer hardware.
  • Local models this capable change the math on privacy-sensitive and offline AI work.
  • Reasoning-effort defaults matter: the same model is brilliant or comically slow depending on one setting.

🔍 What happened

  • Alibaba's Qwen lab released Qwen 3.8 27B on Friday - Apache 2 licensed, vision-capable, 262K context.
  • Simon Willison ran the 17GB build on an M5 MacBook Pro and an NVIDIA DGX Spark.
  • At the default xhigh reasoning setting, one pelican SVG took 21 minutes and 22,276 reasoning tokens.
  • With reasoning off, the same prompt finished in just over two minutes.
  • The model drove the Pi coding agent through real repo questions and built a working bounding-box tool.
  • A llama.cpp multi-token-prediction build ran about 72% faster than the LM Studio default.

💬 Smart takes

  • Simon Willison: "The fact that a 17GB file can do all of this stuff on my home machines is a miracle."
  • Willison on the default: "This is a hilarious default. It's absolutely not a good way to run the model."
  • Skeptic: at 15-30 tokens per second, hosted APIs still answer several times faster - speed, not smarts, keeps local models off daily-driver duty.

🧭 Where this goes

  1. Likelycommunity MTP and MLX optimizations cut local inference times sharply within weeks.
  2. Likelyindependent benchmarks confirm the 27B beats Qwen's own closed 3.7-Plus on several tasks.
  3. Possiblereasoning-effort dials become a standard control across all open model releases.
  4. Wild Carda laptop-class open model becomes the default coding agent driver for small teams within a year.

🥄 The Spoon Take

The story isn't one model - it's the floor rising. A free 17GB file now does what expensive hosted setups did a year ago. When capability is solved at laptop scale, speed becomes the last moat for the API labs.

🤔 Pushback

Willison's 15-30 tokens per second is painful in practice - most users will drift back to fast hosted models within a week.