How to read an AI model's benchmark scores without getting fooled
Every model launch ships with a chart showing the new model winning. Five questions turn that chart from marketing back into evidence.

Every model launch comes with a chart, and the chart always shows the new model winning. That is not a coincidence — it is the nature of the format. A benchmark score is a real measurement, but it is also a marketing asset, and the two roles pull in different directions. Here is how to read one without being led by it.
Five questions to ask of any benchmark number
1. Who produced the number? The single most important question. A score on a lab's own model card is self-reported — the company chose the test, the settings and the framing. That does not make it false, but it makes it a claim, not a verdict. An independent number — run by a third party on a neutral setup — is worth far more. When we wrote up Alibaba's Qwen3.8-27B, the headline coding scores were all Qwen's own; that is the norm at launch, not the exception.
2. Which benchmark, exactly — and which version? "SWE-bench" is not one thing. There is the original, a smaller human-validated subset, and harder variants — and a score on one says little about the others. Names get shortened in headlines until different tests blur into a single number. Always pin the exact benchmark and its version before comparing two models, because half the "wins" you see compare scores from tests that were never the same.
3. Could the model have seen the test already? This is contamination, and it is the quiet killer of benchmark credibility. If a test's questions (or close paraphrases) sat in the training data, the model isn't reasoning — it's remembering. Public benchmarks are especially exposed, because they end up scraped into the next training run. A suspiciously high score on a well-known public test deserves more doubt, not less.
4. What was the setup around the model? The same model can post very different scores depending on the harness — the scaffolding, tools, retries and prompt formatting wrapped around it during the test. A generous harness flatters a model; a strict one exposes it. When one lab's number is run under its own harness and another's under a different one, the comparison is not apples to apples, however tidy the bar chart looks.
5. Does the benchmark still measure anything useful? Many older benchmarks are saturated — the best models score in the high 90s, so a one-point gap sits within the margin of error rather than marking a real difference in skill. And a high score on a narrow test says nothing about the messy, open-ended work you actually need the model for. Ask what the benchmark rewards, and whether that is the thing you care about.
Where the more trustworthy numbers live
Independent evaluators exist precisely to break the self-reported cycle. Human-preference arenas, where people blind-test two models and pick the better answer, are harder to game because there is no fixed test to memorise — though they are not immune to their own biases, such as favouring longer or more stylish answers. Third-party analysis services re-run models on standardised setups and publish comparable indices. Neither is perfect, but a number someone outside the company produced is a stronger starting point than one the company produced about itself.
The bottom line
A benchmark score is evidence, not a result. Treat a launch-day chart as the beginning of a question — who ran it, on what, under what conditions, and could the model have seen it coming — rather than the answer. The models worth trusting are the ones whose numbers still look good after someone else has run them.
Ask Relay — he reads every question himself and replies personally by email.
