Cognition's SWE-2 is post-trained on Moonshot's 2.8-trillion-parameter Kimi K3. It scores 50.0% on FrontierCode against Fable 5.1's 50.9%, at 64% lower cost. Chinese open weights reach the frontier.
SWE-2 shipped September 10 inside Devin Desktop and CLI. Cognition calls it the first RL run at multi-trillion-parameter scale. Its own RL adds 5 to 6 points over the Kimi base.
Three effort levels trained in one run. Medium takes 58% fewer turns and costs 81% less than SWE-1.7. Mean steps per task dropped from 127 to 53.
The fine print: Terminal-Bench 4 is 27.3% versus 55.8% for Fable 5.1. FrontierCode is Cognition's own benchmark. No API, no per-token price, no model card yet.