Qwen 3.8 Thinks For 21 Minutes
17AUG
Alibaba's new open model runs on a laptop and codes, sees, and calls tools. Simon Willison calls it the best local model he's tested. One catch: the default setting overthinks everything.
Qwen 3.8 27B is a 17GB file. It handles vision, tool use, and a 262,000-token context. Willison ran it on a MacBook and an NVIDIA Spark.
The default reasoning mode burns tokens on everything. One pelican drawing took 21 minutes and 22,000 thinking tokens. Turn reasoning down and the same job takes two minutes.
It also drove a real coding agent and nailed image bounding boxes. Willison's verdict: a miracle file, held back only by speed.