Loading prices...
All news
Flat vector illustration of a glowing amber house outline with deep roots reaching into the ground on the left, next to a faint, rootless outline of an identical house on the right, symbolizing the difference between owning an AI model and renting one

Thomson Reuters spent $40M to own its AI instead of renting from OpenAI

17:15 · 24.08.2026
Source: The Decoder
1

Thomson Reuters built its own AI language model for legal work instead of renting one from OpenAI or Anthropic, spending roughly $40 million on staff and computing power over more than two years, The Decoder reported.

The model, called Thomson, sits on top of Alibaba's open Qwen, most recently Qwen3.5-397B. Working with Imperial College, Thomson Reuters first retrained the Chinese model for safety, ethics, and political neutrality, an intermediate version it calls Snowdon, after the mountain in Wales. Pre-training on the company's own content came next, followed by post-training with domain experts and agentic reinforcement learning inside Thomson Reuters' own tool environments. Less than 10% of the available content has gone into training so far.

The $40 million figure dwarfs the $450,000 number that has circulated more widely, which covers only the final training run of the current version. Even $40 million understates the real investment: decades of content from Westlaw, Practical Law, Checkpoint, and Reuters, plus the working hours of hundreds of domain experts who built the underlying data.

It's renting a house versus buying a house. You are building equity in something that you own for the long-term, and that compounds over time.

Joel Hron, CTO, Thomson Reuters

The company's own benchmarks show a mixed picture. On Stanford LegalBench, Thomson scores 0.823, trailing Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark it sits right behind Opus 4.8. It leads on instruction following and PrBench Legal, but falls off sharply on reasoning and especially coding, and the comparison itself is skewed: Thomson runs with test-time scaling, while GPT-5.5 runs without a reasoning mode.

In the company's in-house Deep Research benchmark, Thomson scores 0.53 on factual accuracy with web access alone, versus 0.65 for GPT-5.4. Give Thomson access to Thomson Reuters' own content, and it edges past GPT-5.4, 0.83 to 0.82. Evaluation lead Andrew Bean credits a big uplift that comes from being able to train on and practice with the company's own tools, something outside providers can't replicate. GPT-5.4 improves almost as sharply with the same content, though, which means data access is doing nearly as much work as the specialized training itself.

Fine-tuning a frontier model was the alternative, and Thomson Reuters rejected it for three reasons. Research chief Jonathan Schwartz says standard fine-tuning techniques tend to degrade general capability, and fine-tuning still leaves a company locked into its provider for inference costs and the model roadmap. Fine-tuning reshapes a general model into a specialist, but the second reason Thomson Reuters gives is data: the performance jump comes from training inside proprietary tools like Westlaw, access the company isn't willing to hand to an outside lab. The third is compounding value: every expert review inside the company's workflow becomes training data it keeps, rather than value that evaporates once it leaves for a third-party provider.

  • Training cost: about $40 million over more than two years, versus the widely cited $450,000 that covers only the final run
  • Base model: Alibaba's Qwen3.5-397B, retrained for safety and neutrality as an intermediate version called Snowdon
  • Stanford LegalBench score: 0.823, behind Gemini 3.1 Pro and GPT-5.5
  • Deep Research accuracy with company data: 0.83, edging past GPT-5.4's 0.82
  • Training content used so far: less than 10% of what's available

Thomson's first job is narrower than a general-purpose assistant. It's replacing the Tabular Analysis feature inside CoCounsel Legal, where a smaller, cheaper model fits the economics of high-volume document review. The product stays multi-model, administrators can switch between models, and Thomson isn't meant to orchestrate tasks so much as handle subtasks like citation checking. Customer data doesn't go into training, the company says.

A smaller version is headed to Hugging Face as an open-weight model under a non-commercial license, with a technical report and developer portal to follow, and Hron says early, non-binding talks are underway with law firms about direct licensing. Whether the model stays ahead depends less on the current benchmarks than on what the company called the real advantage: knowing which intelligence is worth owning outright.

Nothing here should be taken as financial advice — just information to consider.

Published: 17:15 · 24.08.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!