The Independent Numbers Are In: Inkling Leads US Open Weights — With One Asterisk That Matters
Artificial Analysis's full evaluation beats the snap Hacker News verdict in both directions: Inkling tops Kimi and DeepSeek Flash on agentic evals and thinks in far fewer tokens than GLM-5.2 — which is still ahead exactly where developers live.

The takeaway: The first full independent scorecard for Inkling is out, and it complicates the snap verdict. When Thinking Machines Lab released its 975B open-weights model yesterday, the loudest early take on Hacker News was "not as good as GLM 5.2 for agentic workflows while also being bigger." Artificial Analysis's detailed evaluation — the independent benchmark shop, not the vendor — now says Inkling is "the new leading U.S. open weights model," and on AA's own agentic evaluations it beats two of the three Chinese giants it was measured against. The gap between those two readings is worth understanding before anyone commits to a conclusion.
The independent numbers
Artificial Analysis scores Inkling 41 on its Intelligence Index — three points above the previous leading US open-weights model, Nvidia's Nemotron 3 Ultra (38), and well clear of Gemma 4 31B (29) and OpenAI's gpt-oss-120b (24). On the agentic side, AA reports Inkling scoring "higher than both Kimi K2.6 and DeepSeek v4 Flash" on both of its work-task evaluations: an Elo of 1238 on GDPval-AA v2 (against 1190 and 1189 respectively, the DeepSeek figure at max reasoning), and 24% on its τ³-Banking agentic benchmark, ahead of Kimi K2.6's 21% and — in AA's own words — "just above" DeepSeek v4 Flash max's 23%.
The quieter finding may matter more for anyone actually paying for tokens: efficiency. AA measures Inkling averaging around 25,000 output tokens per Intelligence Index task, versus roughly 43,000 for GLM-5.2 at max reasoning, 38,000 for Kimi K2.6, and 37,000 for DeepSeek v4 Pro at max. A model that thinks in fewer tokens claws back some of the cost disadvantage we noted yesterday — AA's own summary still calls Inkling "particularly expensive when comparing to other open weight models of similar size" on per-token price, but per task, the arithmetic is closer than the price list suggests.
So was the Hacker News verdict wrong?
Not exactly — and the difference is instructive. The "not as good as GLM 5.2" judgment came from practitioners talking about agentic coding workflows, and GLM 5.2 is conspicuously absent from the head-to-head agentic rows AA's article highlights (its comparisons there are Kimi K2.6 and DeepSeek v4 Flash). But the vendor's own table fills that gap, and it backs the sceptics: on the model card, GLM 5.2 posts 1514 on GDPVal-AA v2 against Inkling's 1238, leads 26.8% to 23.7% on Tau 3 Banking, and is ahead on SWE-Bench Verified (80.0% to 77.6%) — the benchmark closest to what those commenters do all day. Both things can be true: the strongest US open-weights model on a broad intelligence index, and still behind GLM 5.2 on the agentic work that dominates open-model usage among developers — by the vendor's own accounting.
Three precision notes for the record. First, context: the announcement claims support for up to 1M tokens, and that full window comes with running the open weights yourself — the hosted Tinker API currently serves 256K. Second, the Chinese-lineage story is partly true and worth stating exactly: the announcement itself says that "To bootstrap post-training, we ran an initial SFT on synthetic data generated by open-weights models including Kimi K2.5" — a small fraction of total compute, by the company's account, but a neat illustration of where open-weights gravity currently sits: the new American leader got its first post-training push from a Chinese one. Third, a correction to a claim circulating in aggregator summaries — that Inkling's architecture "references DeepSeek-V3": neither AA's article nor the vendor's own materials say that. The model card describes a 66-layer decoder-only transformer routing each token to 6 of 256 experts plus 2 shared — a mixture-of-experts shape the Chinese open-weights wave made standard — with a hybrid local-global attention design of its own. Family resemblance, not a blueprint.
Why it matters
Yesterday's question was whether America finally had a real open-weights entry. Today's data says yes, with an asterisk: leading on breadth, ahead of Kimi and DeepSeek Flash on AA's agentic evaluations, more token-efficient than all of them — and, by the vendor's own table, still clearly behind GLM 5.2 exactly where open-weights models earn their keep. The data points to watch now: AA's independent GLM-5.2 head-to-heads on those same agentic evaluations, and whether the community's hands-on verdict shifts as the fine-tunes and quantisations land. We'll keep the running comparison updated as the independents publish.
Ask Relay — he reads every question himself and replies personally by email.
