Thomson Reuters built its own frontier model for $40M
The data giant trained 'Thomson' on Westlaw, Reuters and its other archives, says its own tests put it on par with frontier models, and released a small open-weight version — a bet that if you own the data, you can skip the inference tax.

Thomson Reuters has trained its own frontier-class AI model, called Thomson, and says it did it for about $40 million — a fraction of what the leading labs spend — by starting from an open-source base and adding something no lab can buy: its own archives.
The company announced Thomson on 24 August, calling it its first proprietary large language model. Rather than train from scratch, Thomson Reuters took a strong open-source foundation and specialised it on decades of its own content — Westlaw, Practical Law, Checkpoint and Reuters — data it owns outright and that most model-builders can only license, if at all.
And it says it has used less than 10% of its content so far. Thomson is the first pass, not the finished article.
The claim, and the caveat
Thomson Reuters says its own evaluations put Thomson "on par with the latest frontier models across a range of tasks." That is a self-reported result — the model has not been independently benchmarked, and "a range of tasks" is doing quiet work in that sentence. It is a claim to weigh, not a scoreboard to trust, and the same caution applies to every vendor's launch-day numbers.
What is more concrete is the economics. The company frames Thomson as trained and run "at a fraction of the cost of comparable frontier models" and "without the heavy inference costs" of calling someone else's frontier API on every query. For a business running millions of legal and tax lookups, the inference bill is close to the whole game — and it is exactly the cost curve the frontier labs have been racing down from the other direction.
"Thomson shows there is another path. Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient and entirely under your control." — Joel Hron, Chief Technology Officer, Thomson Reuters
Where it lands
Thomson goes to work first inside CoCounsel Legal, powering a Tabular Analysis feature. And in a move that is becoming the norm rather than the exception, Thomson Reuters is releasing a small, open-weight version on Hugging Face for academic and non-commercial use — the same Hugging Face now weighing a $13bn sale that has become the default shop window for open models, from Meta's Muse-Glimmer on down.
Why it matters
The interesting thing here is not the model; it is the strategy. The dominant story of the last two years has been that intelligence is something you buy — a metered API from one of a handful of labs. Thomson's bet is that if you already own a deep, proprietary dataset, you can build a specialised model that is good enough for your domain, cheaper to run, and entirely yours — no rate limits, no data leaving the building, no dependence on a supplier who is also, increasingly, a competitor.
That will not be true for everyone. Most companies do not have Westlaw. But every incumbent sitting on a data moat — a bank, an insurer, a health system, a publisher — is now watching to see whether "specialise an open base on your own archive" is a real alternative to renting frontier intelligence by the token. Thomson Reuters just spent $40 million to argue that it is.
Ask Relay — he reads every question himself and replies personally by email.
