Thomson Reuters’ AI Model Competes with Global Leaders: A Closer Look at Benchmark Achievements and Industry Impact

Thomson Reuters has made strides in the realm of artificial intelligence with the announcement that its homegrown large language model, named Thomson, now matches some of the leading general-purpose AI models globally. The company recently published its first benchmarking data, claiming competitive, if not superior, performance against models such as Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5. This key development originates from the company’s acquisition of Safe Sign Technologies in 2024, marking a significant step in Thomson Reuters’ AI journey. For a more detailed exploration, the announcement made on their innovation blog provides deeper insights.

Thomson, relying on an open-source foundational model, has been further refined through both mid-training and post-training enhancements utilizing Thomson Reuters’ extensive proprietary content from platforms like Westlaw and Practical Law. Notably, only 10% of their content has been utilized so far, indicating potential for further growth and improvement. The model’s first deployment is scheduled for August, where it will be integrated into Tabular Analysis within the CoCounsel Legal platform, showcasing its capabilities in high-volume structured documentary analysis.

Offering specifics, Thomson’s benchmark performances indicate that it has excelled in legal domains, specifically in the PrBench Legal Hard benchmark, where it outperformed its competitors. However, its performance showed variability across other domains. Although the model ranked first in certain tests, it did not top a majority of the categories, and factors like the use of test-time scaling and different operational modes across models add layers of complexity to the benchmarking results.

In another revealing evaluation, Thomson, when combined with Thomson Reuters’ vast proprietary content, demonstrated superior performance in legal research queries against frontier models that had unrestricted web access. This was anticipated, given the strength of the Westlaw and Practical Law data sources when compared to general web data. Yet, it leaves an open question as to how these frontier models would perform if given similar access to Thomson Reuters’ content.

Moreover, with no independent external verification of these results, Thomson Reuters acknowledges the necessity for third-party evaluation to validate their claims. While some improvements for OpenAI’s model might be anticipated if tested in a more optimal mode, Thomson’s performance to date positions it impressively among much larger competitors.

Thomson underscores a strategic direction for the company, suggesting not only a commitment to developing competitive AI solutions internally but also emphasizing their vision of “Fiduciary-Grade AI.” Despite the apparent competitive edge, questions remain about the scalability and operational efficacy compared to more established frontier labs. As tools evolve and integration capabilities improve, the dynamic between proprietary models and general-purpose AI will continue to intrigue and challenge legal professionals’ approach to technological advancements. For further analysis and upcoming developments, the full coverage on LawNext details the ongoing saga of Thomson’s AI journey here.