Monday Sep 14

Devin Now Runs On A Chinese Base Model

14SEP
64% CHEAPERON KIMI K3VS FABLE 5.1

Cognition's SWE-2 is post-trained on Moonshot's 2.8-trillion-parameter Kimi K3. It scores 50.0% on FrontierCode against Fable 5.1's 50.9%, at 64% lower cost. Chinese open weights reach the frontier.

SWE-2 shipped September 10 inside Devin Desktop and CLI. Cognition calls it the first RL run at multi-trillion-parameter scale. Its own RL adds 5 to 6 points over the Kimi base.

Three effort levels trained in one run. Medium takes 58% fewer turns and costs 81% less than SWE-1.7. Mean steps per task dropped from 127 to 53.

The fine print: Terminal-Bench 4 is 27.3% versus 55.8% for Fable 5.1. FrontierCode is Cognition's own benchmark. No API, no per-token price, no model card yet.

full brief & sources

⚡ Why this matters

  • A US coding-agent company at a $48B valuation now ships its flagship on Chinese open weights. That is the supply chain, not a side experiment.
  • Near-frontier coding at roughly a third of the price changes the build-vs-buy math for anyone paying per task.
  • It lands the same week Amodei asks for a crackdown on distillation from frontier models. Open weights are the loophole nobody has to distill.

🔍 What happened

  • Cognition released SWE-2 on September 10. It is post-trained with reinforcement learning from Kimi K3, Moonshot's 2.8-trillion-parameter open-weight model.
  • Cognition's table: 50.0% on FrontierCode 1.1 Main versus 50.9% for Fable 5.1 and 53.3% for GPT-6 Astra. 73.0% on DeepSWE 1.1. 92.8% on Terminal-Bench 2.1, top of the table.
  • Cost claim: 64% cheaper than Fable 5.1 at the FrontierCode point, about a quarter of Astra's cost. Anchor: Fable 5.1 Medium at $3.28 per task. SWE-2's own per-task price is not published.
  • Three reasoning effort levels, medium, high and max, trained in a single RL run with a linear cost penalty per level. Medium averages 53 steps per task against 127 for SWE-1.7.
  • Terminal-Bench 4: 27.3% for SWE-2 against 55.8% for Fable 5.1 and 57.9% for Astra. Cognition prints the row but leaves it out of the headline.
  • Availability is Devin Desktop and CLI today, Devin Web and Fusion rolling out. No standalone API, context window or model card. Cognition raised $2B+ at $48B on September 8.

💬 Smart takes

  • Cognition: SWE-2 is 'within one point of Fable 5.1 while being 64% cheaper' and 'our closest model yet to the frontier.'
  • Nitish Garg, CellCog CEO: on par with the frontier holds on three benchmarks and not on the fourth. Every rival number is Cognition's own run in the rival's harness.
  • Skeptic: the benchmark is Cognition's, the harness is Cognition's, the price is relative. Wait for an outside run.

🧭 Where this goes

  1. LikelyCognition ships an SWE-2 API with a per-token price within a quarter, and the cost claim gets tested.
  2. Likelyat least one other US agent company announces a Kimi K3 or DeepSeek V4 base by October.
  3. PossibleWashington adds open-weight Chinese bases to the distillation and export-control conversation.
  4. PossibleMoonshot restricts the license on the next Kimi release once it sees who is building on it.
  5. Wild CardAnthropic or OpenAI drops a coding-only model priced against SWE-2 rather than against each other.

🥄 The Spoon Take

Two years ago the story was Chinese labs distilling American models. This week an American company post-trains a Chinese open model and gets within a point of Fable on its own benchmark. Cognition's real product is the RL recipe and the harness. The base is a commodity, and the cheapest good one is Chinese and open. The question is not which lab. It is which base plus whose harness.

🤔 Pushback

Terminal-Bench 4 at half the frontier score says the model still breaks on the hardest long-horizon work. And a Devin-only model with no token price is a plan feature, not a market price.