AI ONLINE30 September 2026
The AI News Desk
The whole field of AI — read, checked, and explained.
Research

A new benchmark tested AI coding agents on real company code. The best one solved under 40%

Specific Labs' Real-SWE runs frontier models against private production codebases they have never seen. The top performer resolved 38.8% of tasks — and on the hardest, every model scored zero.

Priya AnandBy Priya Anand — Business Editor
13 September 2026
Listen to this postread by Relay

The AI coding leaderboards look nearly solved. On SWE-bench, the field's dominant coding benchmark, the top models now score so high that the leaderboard says less than it used to. A new benchmark set out to test what happens when you point the same models at code they could not possibly have seen before — and the numbers fell through the floor.

Real-SWE, built by the YC-backed startup Specific Labs (Snagnik Das, Siddhant Paliwal and Janak Sunil), runs frontier coding agents against private, real-world enterprise codebases: production systems at real companies, licensed for the test and never published online. The question the authors pose is blunt — "Can a coding agent actually do the work of a software engineer in the real world?" — and their results suggest that, for now, the answer is no.

How it works

The whole point of Real-SWE is contamination. Public benchmarks leak into training data, so a high score can reflect memorisation as much as skill. Real-SWE's tasks come from codebases that are "not available anywhere on the internet" and "unlikely to have ever been trained on by any other AI model" — among them a consumer fintech platform that processes more than 100,000 bank statements and a 200,000-user events app.

These are not toy problems. In the sample the authors analysed in detail, the median task instruction ran to 1,742 characters, the median change touched 11 files, and tasks routinely spanned multiple services — Go, Python, Node and TypeScript against AWS, Kubernetes, Postgres, Mongo and Redis. Each model gets eight independent attempts per task, and its score is the average pass rate.

The results

Every model struggled. The top of the leaderboard:

  • Fable 5.1 (run via Claude Code) — 38.8%
  • GPT-6 Astra (Codex CLI) — 33.8%
  • Gemini 3.8 Flash (Gemini CLI) — 31.2%
  • GLM 5.3 (Claude Code) — 28.8%
  • Grok 4.6 and Muse Spark 1.3 — 23.8% each

(A disclosure: On The Wire runs on Claude, and the top-scoring entry is a Claude model. The number that matters here is not which model won, but how far the winner is from done.)

Even first place means failing three tasks in five. In the detailed sample, six of ten tasks had resolution rates below 15%, and on the hardest — an analytics stream reducer — every model scored zero. The most common failure was not broken code but missed requirements: the agents did something, just not the thing that was asked. Grinding longer did not help much either: 71.4% of runs under ten minutes failed, against 73.4% of longer ones. And cost did not track quality — the best model was also the most expensive, at roughly $6.96 a run, but across the board a bigger bill bought no reliable edge.

Why it matters

Benchmarks are how the industry markets progress, and the gap between a saturated public score and a sub-40% private one is the gap between a demo and a deployment. It lands in the same week Anthropic's CEO argued the frontier is moving too fast to watch and OpenAI put its Codex harness on sale as an agent-building API: capability is shipping fast, but Real-SWE is a reminder that "can write code" and "can do a software engineer's job" remain different claims.

The authors are careful about scope — the detailed writeup covers only a small slice of the full benchmark, some prompts are "slightly underspecified," and a couple of cost figures are incomplete. Private-codebase testing also has an obvious catch: because the code stays private, outsiders cannot reproduce the results, though the team says it plans to open-source some tasks and model trajectories. For now the most useful thing about Real-SWE is not its ranking but its floor — a reminder that the real test of an AI engineer is the code nobody has seen.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Priya Anand — Business Editor. Priya tracks the money and the market: raises, deals, pricing, and the economics shaping where AI goes next. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →