AI ONLINE22 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Models & Releases

Cognition's new SWE-2 lands within a point of the top coding models — and says it runs 64% cheaper

The maker of Devin says its new SWE-2 gets within a single point of Anthropic's Fable 5.1 on its own coding benchmark, at about 64% lower cost — and it's built on top of a Chinese open model, Moonshot's Kimi K3.

Priya AnandBy Priya AnandBusiness Editor
10 September 2026
Listen to this postread by Relay

Cognition — the company behind the Devin coding agent — has launched SWE-2, which it calls its "most advanced coding model yet." The pitch is the defining one of 2026: near-frontier performance at a fraction of the cost.

Close to the frontier, not past it

On Cognition's own FrontierCode 1.1 benchmark — which tests whether an AI-written pull request would actually be merged by a human maintainer — SWE-2 scores 50.0%. For context, Cognition puts Anthropic's Fable 5.1 at 50.9% and OpenAI's GPT-6 Astra at 53.3% on the same test.

So SWE-2 does not beat the frontier models — it lands just short. The claim is what it costs to get there: Cognition says SWE-2 runs about 64% cheaper than Fable 5.1 at that score; coverage has summarised the savings as up to 70% across the benchmarks Cognition tested — a secondary characterisation, not a figure in Cognition's own post. It's available now across Devin's Desktop, CLI, Web and Fusion products.

That's the same shape as DeepSeek's V4.1 Flash launch this morning: the 2026 race is less about who tops the leaderboard than who gets close enough at a low enough price that the leaderboard stops being the point.

Built on a Chinese open model

The detail worth pausing on: SWE-2 is built on top of Kimi K3, the 2.8-trillion-parameter open model from China's Moonshot AI. A US startup taking an open Chinese base model and post-training it to near-frontier coding is a neat illustration of how the open-weights supply chain now crosses borders — and of how much of the frontier is downstream of a handful of large open bases.

Cognition also makes two technical claims: that it "scaled RL to the multi-trillion-parameter regime for the first time," and that it used a cost-penalised reward function, tuned to the slope of its base model's Pareto curve, to optimise the whole cost–performance curve across effort levels rather than chasing a single benchmark number. It reports SWE-2 uses 58% fewer turns and costs 81% less than its own previous SWE-1.7 on average.

The caveat that matters here

Every number above is Cognition's own, self-reported, on its own benchmark. FrontierCode is Cognition's test; the cost comparisons are Cognition's; and "first to scale RL to the multi-trillion-parameter regime" is a claim, not an independently verified fact. Independent evaluation — and real-world use inside Devin — will tell the fuller story. But if the numbers hold even roughly, SWE-2 is another data point in the year's clearest trend: the gap between the frontier and the cheap seats keeps narrowing.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Priya Anand — Business Editor. Priya tracks the money and the market: raises, deals, pricing, and the economics shaping where AI goes next. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →