Cognition's new SWE-2 lands within a point of the top coding models — and says it runs 64% cheaper
The maker of Devin says its new SWE-2 gets within a single point of Anthropic's Fable 5.1 on its own coding benchmark, at about 64% lower cost — and it's built on top of a Chinese open model, Moonshot's Kimi K3.

Cognition — the company behind the Devin coding agent — has launched SWE-2, which it calls its "most advanced coding model yet." The pitch is the defining one of 2026: near-frontier performance at a fraction of the cost.
Close to the frontier, not past it
On Cognition's own FrontierCode 1.1 benchmark — which tests whether an AI-written pull request would actually be merged by a human maintainer — SWE-2 scores 50.0%. For context, Cognition puts Anthropic's Fable 5.1 at 50.9% and OpenAI's GPT-6 Astra at 53.3% on the same test.
So SWE-2 does not beat the frontier models — it lands just short. The claim is what it costs to get there: Cognition says SWE-2 runs about 64% cheaper than Fable 5.1 at that score; coverage has summarised the savings as up to 70% across the benchmarks Cognition tested — a secondary characterisation, not a figure in Cognition's own post. It's available now across Devin's Desktop, CLI, Web and Fusion products.
That's the same shape as DeepSeek's V4.1 Flash launch this morning: the 2026 race is less about who tops the leaderboard than who gets close enough at a low enough price that the leaderboard stops being the point.
Built on a Chinese open model
The detail worth pausing on: SWE-2 is built on top of Kimi K3, the 2.8-trillion-parameter open model from China's Moonshot AI. A US startup taking an open Chinese base model and post-training it to near-frontier coding is a neat illustration of how the open-weights supply chain now crosses borders — and of how much of the frontier is downstream of a handful of large open bases.
Cognition also makes two technical claims: that it "scaled RL to the multi-trillion-parameter regime for the first time," and that it used a cost-penalised reward function, tuned to the slope of its base model's Pareto curve, to optimise the whole cost–performance curve across effort levels rather than chasing a single benchmark number. It reports SWE-2 uses 58% fewer turns and costs 81% less than its own previous SWE-1.7 on average.
The caveat that matters here
Every number above is Cognition's own, self-reported, on its own benchmark. FrontierCode is Cognition's test; the cost comparisons are Cognition's; and "first to scale RL to the multi-trillion-parameter regime" is a claim, not an independently verified fact. Independent evaluation — and real-world use inside Devin — will tell the fuller story. But if the numbers hold even roughly, SWE-2 is another data point in the year's clearest trend: the gap between the frontier and the cheap seats keeps narrowing.
Ask Relay — he reads every question himself and replies personally by email.
