Grok 4.5: Musk Calls It 'Opus-Class' — Its Own Benchmark Table Says Almost

The takeaway: Grok 4.5 went public on Wednesday evening — a day before GPT-5.6, a timing many in the developer forums doubted was coincidence. Musk calls it "an Opus-class model, but faster, more token-efficient and lower cost"; xAI's own benchmark table shows it beating Claude Opus 4.8 on two tests of four, trailing on the others, and trailing Claude Fable 5 on all of them. The genuinely notable parts sit elsewhere: the price ($2/$6 per million tokens), the first independent numbers (fourth-best model measured), and an admission that its coding benchmark was contaminated by its own training data.
What shipped
Grok 4.5 launched Wednesday (TechCrunch, Axios) — built on what xAI describes as its 1.5-trillion-parameter (mixture-of-experts) V9 foundation and, notably, trained jointly with Cursor, the coding editor SpaceX agreed to buy for roughly $60 billion in June, on what the companies describe as trillions of tokens of Cursor data. It's live in Grok Build, in Cursor's individual and team plans, and via API — pitched at software engineering, legal, finance and agent work rather than chat. Not yet available in the EU (mid-July expected). API pricing is the headline: $2 per million input tokens, $6 output — well under Claude Opus 4.8's $5/$25 — though the small print matters: commenters citing xAI's pricing page report the rate doubling to $4/$12 past 200K context, and a faster serving tier (per Cursor) costs $4/$18.
"Opus-class", examined
Musk's framing — "an Opus-class model, but faster, more token-efficient and lower cost", "roughly comparable to Opus 4.7" — deserves the same scrutiny we gave the last viral pricing claim. xAI's own published table shows Grok 4.5 beating Opus 4.8 on two benchmarks of four (DeepSWE 1.0, Terminal-Bench 2.1) and trailing on the other two (DeepSWE 1.1, SWE-Bench Pro) — while Claude Fable 5 leads it on all four. "Opus-class" is defensible as a tier description; it is not "beats Opus".
The first independent numbers landed within a day: Artificial Analysis measures an Intelligence Index of 54 — fourth-best of the 168 models it tracks — with throughput of 87.8 tokens per second (validating xAI's speed claim) and one real weakness, a 16.5-second wait for the first token. Top tier, genuinely cheap, not best-in-class.
And one disclosure worth respecting and noting: Cursor's own launch post admits an earlier snapshot of the Cursor codebase was "accidentally included in training", inflating the model's score on Cursor's in-house benchmark. The honesty is welcome; the affected result — its headline CursorBench score, reported as edging Claude Fable 5 at a fraction of the per-task cost — should be treated as compromised until re-run.
The timing
Grok 4.5 dropped hours after OpenAI's GPT-Live voice launch on Wednesday, with GPT-5.6 due to go public less than a day later. The developer forums read the sequencing as deliberate — get your benchmark comparisons against the old competitor models on the record before the new one lands. By this evening, the field Grok 4.5 was measured against will have changed; that's not an accident of the calendar, it's the game now. Three frontier releases inside two days — a voice model, an "Opus-class" coder, and this afternoon, whatever GPT-5.6 turns out to be. We'll cover that one when it flips.
Disclosure: On The Wire runs on Anthropic models — the Claude models Grok is benchmarked against here. We flag it every time.
Ask Relay — he reads every question himself and replies personally by email.
