Saturday Aug 22
+14.6 PTSTAPPED ITMISSED IT

Agents are bad at using real software. Alibaba launched Qwen-UI-Agent, a model built to read screens and click through phone and desktop apps. It beat Claude Opus 4.8 by 14.6 points on mobile.

Most agent demos run on APIs. Real work runs on messy screens with buttons that move. Qwen-UI-Agent is trained for the second one.

Numbers: 82.1% on the MobileWorld benchmark, 92.2% on real devices, 81.5% on ScreenSpot-Pro for pointing at the right pixel. State of the art on all five tests it ran.

A Chinese lab now leads the one agent skill enterprises actually need. Most business software has no clean API. Whoever drives the screen well drives the workflow.

full brief & sources

⚡ Why this matters

  • Most enterprise software has no clean API. Whoever drives the screen well drives the workflow.
  • A Chinese lab now leads the agent skill that matters most for real deployments.
  • Benchmarks run on physical devices, not simulators, are a harder and more honest test.

🔍 What happened

  • Alibaba released Qwen-UI-Agent on August 20, covering phones, desktops, web and deep search.
  • It scored 82.1% on MobileWorld, beating GPT-5.6 Sol by 12.0 points and Claude Opus 4.8 by 14.6.
  • On MobileWorld-Real, a 400-task benchmark run on physical hardware across 100+ apps, it hit 92.2%.
  • It posted 81.5% on ScreenSpot-Pro, which measures pointing at the correct pixel.
  • Alibaba built a device farm of 100+ phones and 150+ apps for training and evaluation.
  • State of the art on all five benchmarks reported in the technical report.

💬 Smart takes

  • Qwen technical report: frames the core problem as the simulation-to-real gap, and built a physical device farm specifically to close it.
  • Developers Digest: reads the release as a bet on owning the GUI agent runtime, not just a model score.
  • Skeptic: self-reported numbers on a benchmark the lab built itself. MobileWorld-Real is Alibaba's own and nobody has reproduced 92.2% independently.

🧭 Where this goes

  1. LikelyWestern labs respond with their own real-device evaluation suites within a quarter.
  2. Likelyautomation vendors start quoting GUI agent scores in sales decks.
  3. Possibleprocurement teams in the US and EU refuse the model on origin grounds regardless of the numbers.
  4. Possiblethe gap closes fast once others train on real devices instead of simulators.
  5. Wild Cardan independent reproduction comes in far below 92.2% and the lead evaporates.

🥄 The Spoon Take

Screen driving is the unglamorous part of agents and it is where most deployments die. Your billing system, your ad platform, your internal admin tool: none of them have a clean API. A model that clicks reliably is worth more to a product team than one more reasoning benchmark.

🤔 Pushback

Every number here comes from Alibaba, including the benchmark. Nobody outside the lab has reproduced it yet.

Monday Aug 17
XHIGH MODEPELICAN

Alibaba's new open model runs on a laptop and codes, sees, and calls tools. Simon Willison calls it the best local model he's tested. One catch: the default setting overthinks everything.

Qwen 3.8 27B is a 17GB file. It handles vision, tool use, and a 262,000-token context. Willison ran it on a MacBook and an NVIDIA Spark.

The default reasoning mode burns tokens on everything. One pelican drawing took 21 minutes and 22,000 thinking tokens. Turn reasoning down and the same job takes two minutes.

It also drove a real coding agent and nailed image bounding boxes. Willison's verdict: a miracle file, held back only by speed.

full brief & sources

⚡ Why this matters

  • A 17GB open file now covers vision, coding, tool use, and long context on consumer hardware.
  • Local models this capable change the math on privacy-sensitive and offline AI work.
  • Reasoning-effort defaults matter: the same model is brilliant or comically slow depending on one setting.

🔍 What happened

  • Alibaba's Qwen lab released Qwen 3.8 27B on Friday - Apache 2 licensed, vision-capable, 262K context.
  • Simon Willison ran the 17GB build on an M5 MacBook Pro and an NVIDIA DGX Spark.
  • At the default xhigh reasoning setting, one pelican SVG took 21 minutes and 22,276 reasoning tokens.
  • With reasoning off, the same prompt finished in just over two minutes.
  • The model drove the Pi coding agent through real repo questions and built a working bounding-box tool.
  • A llama.cpp multi-token-prediction build ran about 72% faster than the LM Studio default.

💬 Smart takes

  • Simon Willison: "The fact that a 17GB file can do all of this stuff on my home machines is a miracle."
  • Willison on the default: "This is a hilarious default. It's absolutely not a good way to run the model."
  • Skeptic: at 15-30 tokens per second, hosted APIs still answer several times faster - speed, not smarts, keeps local models off daily-driver duty.

🧭 Where this goes

  1. Likelycommunity MTP and MLX optimizations cut local inference times sharply within weeks.
  2. Likelyindependent benchmarks confirm the 27B beats Qwen's own closed 3.7-Plus on several tasks.
  3. Possiblereasoning-effort dials become a standard control across all open model releases.
  4. Wild Carda laptop-class open model becomes the default coding agent driver for small teams within a year.

🥄 The Spoon Take

The story isn't one model - it's the floor rising. A free 17GB file now does what expensive hosted setups did a year ago. When capability is solved at laptop scale, speed becomes the last moat for the API labs.

🤔 Pushback

Willison's 15-30 tokens per second is painful in practice - most users will drift back to fast hosted models within a week.

Monday Aug 3
#2 GLOBALALIBABACLAUDE

Alibaba unveiled Qwen3.8-Max, its largest model ever, on Monday. The 2.4 trillion parameter model ranks second globally on image benchmarks. It still trails Claude on text, and full release lands next week.

A mixture of experts design keeps costs down. Only ninety five billion of the total parameters activate per request. That's how Alibaba keeps inference cheap at frontier scale.

Reuters frames this as a fierce race among Chinese firms building cheaper models. Its parameter count sits close to Moonshot's Kimi K3, which has two point eight trillion.

Alibaba hasn't published a full benchmark table yet. So today's numbers are still just the company's own claims. Independent testing will decide if that vision ranking actually holds.

full brief & sources

⚡ Why this matters

  • Chinese labs keep closing the gap with US frontier models, fast.
  • Parameter count and open weights are becoming Alibaba's key recruiting pitch to developers.
  • Cost matters as much as capability now that mixture-of-experts design cuts inference bills.

🔍 What happened

  • Alibaba unveiled Qwen3.8-Max on Monday, August 3.
  • The model has 2.4 trillion parameters, close to Moonshot's 2.8 trillion parameter Kimi K3.
  • Only 95 billion parameters activate per request under its mixture-of-experts design.
  • It ranks second globally on Arena.AI's image and video leaderboard, behind a Claude Fable 5 variant.
  • On text tasks it still trails Claude Fable 5 and three Anthropic Opus variants.
  • Full release through Alibaba Cloud's Model Studio is set for next week.

💬 Smart takes

  • Alibaba: the model completed a full software-engineering project in 16 days during internal testing.
  • Reuters: Chinese tech companies are "locked in a fierce and fast-moving battle" to build powerful models cheaply.
  • Skeptic: parameter count is a marketing number. Qwen3.8-Max still trails Claude on the benchmark that matters most, text reasoning.

🧭 Where this goes

  1. LikelyAlibaba leans on the vision leaderboard ranking as its main marketing hook once the model ships next week.
  2. LikelyUS labs keep their parameter counts secret, making direct comparisons harder to verify.
  3. Possibleindependent benchmarks show a smaller gap, or a bigger one, than Alibaba's own numbers suggest.
  4. Wild CardQwen3.8-Max's cost advantage pulls meaningful US enterprise workloads away from Anthropic and OpenAI within months.

🥄 The Spoon Take

Alibaba keeps playing the same card: bigger parameter count, lower price, open weights. It's working on developers even if text benchmarks still favor Claude. Watch the vision leaderboard ranking, not the headline parameter count. That's where Qwen3.8-Max actually earned second place.

🤔 Pushback

Alibaba hasn't published a benchmark table yet, so every number here is still Alibaba's own claim.