JOURNAL / 2026.08.25

Thomson Reuters launches Thomson, a legal model built on Qwen

The company specialized open weights with its own data and experts and released a small version; the result makes a different economy for vertical models plausible, although “frontier” and “sovereignty” need qualification.

Thomson Reuters did not train a model from scratch to compete with the major labs. It did something more interesting for many organizations: took a good open-weight model, modified all of its parameters with its own data and evaluators, and turned it into a system it controls end to end.

On August 24, the company introduced Thomson, a family specialized for legal, tax, and journalism work. Thomson-1.0-Large starts from Qwen3.5-397B-A17B; a 35-billion-parameter version—with 3 billion active per token—starts from Qwen3.6-35B-A3B and already has weights available. The large model is not offered through an API or sold separately: its first planned use is inside Tabular Analysis in CoCounsel Legal, in an upcoming product release.

The news, then, is not the arrival of another general chatbot. It is a fairly detailed demonstration of a possible vertical-model factory. It is also a demonstration produced by the company itself, with one word—“frontier”—that is much broader than the supporting evidence.

From Qwen's open checkpoint to three distinct surfaces: small weights under a noncommercial license, a closed large model, and CoCounsel with proprietary data and tools.

The cost that matters is not just the final training run

The 121-page technical report describes three modules. First, Thomson Reuters realigned values through preferences derived from a public constitution. It then continued training on 200 billion tokens selected from a pool of more than 19 trillion: roughly one third proprietary documents, one third synthetic restatements of those documents, and the remainder general-capability replay data. A merge with the earlier checkpoint sought to recover capabilities that specialization could erase. Finally came two stages of preference optimization and two of reinforcement learning, including practice on long research tasks with tools and rewards for faithful citation.

This is considerably more than adding documents to a query or tuning a small adapter. It involves full-weight updates, an in-house data generation and provenance chain, agent environments, evaluations, and inference serving. The model inherits Qwen's architecture and immense prior investment, but absorbs new knowledge and behavior into its parameters.

The numbers help separate two cost stories. The final Large training run lasted three weeks and the company estimates that it used less than $450,000 in GPU compute. The complete program cost about $40 million, including staff, domain experts, vendors, reusable infrastructure, and experimentation. A technical team of no more than three dozen people worked on it, with as many as 368 B200 GPUs available at one point, although Large's main stages used 128 concurrently.

It would therefore be misleading to say a frontier model costs $450,000. Thomson demonstrates something else: an institution that already owns content, specialists, products, and a compute platform can skip pretraining and concentrate tens of millions on transforming an open checkpoint. That is far less than founding a general-purpose lab; it remains well beyond a small company's reach.

What the evaluations actually support

Thomson Reuters computes an unweighted mean across its broad suite and places Large at 78.5%, close to Claude Opus 4.8 at 79.5% and ahead of several recent general models. The comparison uses the same harness, identical instructions, and equivalent tools when a task permits them. Thomson also clearly improves on its Qwen starting point in document work, real human legal queries, and research agents.

That mean does not automatically make Thomson the second-best general model. It combines more than 100 public and internal tests, gives equal weight to very different categories, and uses language-model judges for many open responses. Some internal material comes from the same domains and experts that guided development. On coding, which was not a target, Large trails several rivals, and the report itself records a decline from 54.0 to 48.3 from Qwen to Thomson on Terminal-Bench 2.1. On PRBench, a difficult external test of professional reasoning, its legal advantage is modest and it does not lead every public legal exam.

The human study offers a different, more practical signal. Thirty-five attorney editors wrote 3,035 real tasks and blindly compared paired responses. Thomson, connected to legal databases and Reuters news, was preferred on 54–64% of legal queries against each of five OpenAI and Anthropic systems with web search. Its advantage appeared mainly in completeness and usefulness; legal soundness was similar. Results narrowed on general questions, and GPT-5.6 Terra beat Thomson on several dimensions.

That is good evidence about the product a legal publisher can build with its model and sources. It does not isolate model intelligence: the competitors lacked access to the same paid databases, and the evaluators worked for Thomson Reuters. The report acknowledges this and presents the exercise as a systems comparison. What is now needed is an independent evaluation of the large model, with controlled tools and sets the team did not use to steer development.

Sovereignty for whom

The small version enables a check the large model does not yet permit. Its sixteen weight files are published, its card documents the architecture, 262,144-token context, and training compute, and it runs with common tooling. That makes “open weights” a verifiable description rather than a future promise.

But it is not the same as open source or commercial freedom. Qwen3.5, Large's foundation, is distributed under Apache 2.0. Thomson-1.0-Small uses PolyForm Strict: it permits noncommercial purposes and use by educational, public, and charitable organizations, but does not authorize redistribution of the weights or derivative works. Large remains closed, and only Thomson Reuters decides where it runs. The company gained room to change provider, update cycle, and data policy from an open upstream model; it does not provide the same room to a downstream commercial competitor.

That asymmetry is legal and strategically understandable, but it narrows the “sovereign AI” thesis. Sovereignty here means that one institution controls more of its own chain, not that dependencies disappear or the result becomes a commons. Qwen, NVIDIA GPUs, cloud providers, rights-bearing data, and proprietary tools all remain.

My reading is that this is precisely where the release matters. It does not prove that general pretraining has stopped mattering or that any corporate archive can become an excellent model. It does offer a recipe and cost accounting concrete enough to change a build decision: publishers, governments, scientific companies, or health networks with a singular corpus no longer have to choose only between sending their data to an outside API and spending billions from scratch.

The next decisive number will not be another benchmark mean. It will be whether the large model measurably improves review of 10,000 documents in CoCounsel, at what cost, with how many human corrections, and under what audit. Thomson Reuters has shown that it can own the engine. It still has to show that ownership produces more reliable work—and that the route can be repeated outside the exceptional institution that already owns Westlaw, Reuters, and decades of editorial judgment.

Sources

← Back to journal