cd /news/artificial-intelligence/thomson-reuters-built-its-own-ai-mod… · home topics artificial-intelligence article
[ARTICLE · art-82542] src=thomsonreuters.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Thomson Reuters built its own AI model that now ranks among the best

Thomson Reuters announced early benchmarking results for its proprietary AI model, Thomson, which performs competitively with leading frontier models such as Claude Opus 4.8 and outperforms GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro across legal and general benchmarks. The model, built on an open-source foundation and trained on Thomson Reuters' authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, is set to launch later this summer. Thomson Reuters acquired Safe Sign Technologies in 2024 to develop the model, which uses less than 10% of the company's content in training so far.

read5 min views1 publishedJul 31, 2026
Thomson Reuters built its own AI model that now ranks among the best
Image: source

#

Jul 31, 2026 |

AI and product innovation Thomson Reuters Built Its Own AI Model That Now Ranks Among the World’s Best

The most capable AI models no longer come only from frontier AI labs. One now comes from Thomson Reuters.

Today, we are sharing early benchmarking results for Thomson, a first of its kind AI model. Across a range of benchmarks assessing legal and general capabilities, Thomson performed competitively with the strongest frontier models on the market, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro.

Why? Because it knows the work.

Launching later this summer, Thomson is the newest layer of the Thomson Reuters AI strategy, and a demonstration of what becomes possible when authoritative content, expert judgment, professional tools, and model development come together.

In 2024, Thomson Reuters acquired Safe Sign Technologies, an AI research company. At the time, the market was betting that access to increasingly powerful general-purpose models would be enough.

We made a different bet. We believed the future of professional AI would require more than general-purpose intelligence. It would require models built specifically for the domains, standards, and consequences of professional work. We believed that the distinct advantages of Thomson Reuters decades of world-class content and expertise could be best expressed in a model that we ourselves crafted.

Thomson is the result of that bet. And it is why Thomson Reuters will continue to set the standard for Fiduciary-Grade AI™.

Meet Thomson

Thomson starts from a strong open-source foundation, so it performs general-purpose work just as effectively as the frontier models. It then goes further: trained using state of the art mid-training and post-training techniques on decades of authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, content professionals have staked their reputations on for generations.

That training was shaped by hundreds of subject matter experts who evaluated outputs, identified failure modes, and validated that the model reasons the way legal professionals actually work. The same professional standard governs how Thomson is deployed. Customer data is never used to train the model.

The result is a model that thinks and reasons like a lawyer while outperforming models multiple times larger on the work that matters.

**Thomson Matches the Best. And Beats the Rest. **

We evaluated Thomson against the leading general-purpose models on the market for general professional work and categories spanning:

| Legal | Coding | | Tax | Math | | Accounting | Multilingualism | | Journalism | Agentic tasks | | Safety | Long context | | Reasoning | Following instruction |

Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate. Thomson Reuters has achieved that performance by combining exceptional AI talent with authoritative proprietary content and deep domain expertise. And with less than 10% of Thomson Reuters content used in its training so far, there remains significant opportunity to expand its capabilities.

Instruction Following is a composite average of the IFEval and FollowBench benchmarks. Reasoning is a composite average of the GPQA Diamond, HLE, and MMLU-Pro benchmarks. Coding is a composite average of the SWE-Bench Pro and Terminal-Bench 2.1. Long Context is a composite average of the Infinity Bench as well as some internal benchmarks developed by Thomson Reuters.

These evaluations show Thomson’s competitiveness with industry recognized benchmarks. Our internal evaluation and training cover a wide range of scenarios, including carefully designed agentic use cases optimized for real-world professional work, tens of thousands of real-world queries written by experts, end-to-end deep research training with human-calibrated judges, and data-centric mid-training on a large amount of our content. To increase safety and robustness, we conduct training and evaluations consistent with Thomson Reuters values, and stress-test the models through human and automated red-teaming. Thomson is still early in its development. To date, less than 10% of Thomson Reuters content has been used in its training, leaving significant opportunity to expand its domain knowledge and capabilities through additional training, rigorous evaluation, and expert validation.

Thomson also showcases its strength when a native integration with Thomson Reuters content is added. When compared with leading frontier models given unrestricted access to the web, Thomson’s access to proprietary data sources such as Westlaw, Practical Law, and Reuters news ensures both superior completeness and factuality (i.e. the ability to back up claims through accurate citations to trusted sources).

This evaluation covers 53 legal research queries written by our internal Subject Matter Experts to represent real world questions. LLM’s are connected via an in-house agentic harness to Westlaw/Practical Law for TR content and Brave search engine for web search. Completeness and Factuality are scored using LLM’s as a judge. Completeness is scored based on SME-written rubrics that list every element that would be required for a good answer to the question. Factuality is based on extracting the claims made in each report and checking whether the cited sources provide evidence for each claim. These metrics were developed and calibrated against SME scoring.

The First Deployment. Not the Last.

Thomson’s first integration will launch in August inside Tabular Analysis in CoCounsel Legal, and that choice was deliberate. Tabular Analysis performs high-volume, structured document review against a clear, measurable accuracy standard. It is exactly the kind of work where a purpose-built model has a demonstrable advantage over a general-purpose alternative, and where that advantage is immediately visible to the professionals relying on the output. Thomson will become the default model powering Tabular Analysis, and over the next year we will continue to integrate it across the Thomson Reuters product portfolio in legal and tax.

Content. Expertise. Tools. Now the Model.

Thomson Reuters has always brought together authoritative content, deep domain expertise, and the tools professionals rely on every day. Thomson adds the fourth element: a model purpose-built to power it all, trained on content competitors cannot access and validated to the standard professional’s demand. It delivers frontier-level performance at a fraction of the size and operating cost of many general-purpose models. A model only Thomson Reuters could build.

This is our commitment to Fiduciary-Grade AI™ in action: AI designed for professionals with duties of care and accountability, where almost right is not good enough.

General-purpose AI is built for everyone. Thomson is built for the professionals who cannot afford to be wrong.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @thomson reuters 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/thomson-reuters-buil…] indexed:0 read:5min 2026-07-31 ·