The Short Answer
Thomson Reuters built a legal and tax model by post-training Alibaba's open-weight Qwen 3.5 on Westlaw, Practical Law, Checkpoint and Reuters content — roughly $40M over two years, with a final training run of about $450K. The gap between those figures is the lesson: the compute was under 2% of the cost. On ibl.ai you own all the code and the data, which is where the other 98% of that investment actually lives.
The widely repeated version of this story is that a domain-leading model now costs $450,000. That is not what happened, and organizations budgeting from it will be badly wrong.
What did Thomson Reuters actually build?
Thomson Reuters built a proprietary model for legal and tax reasoning by post-training an existing open-weight foundation rather than pre-training its own.
The base is Qwen 3.5, Alibaba's open-weight foundation model. On top of it the company applied mid-training and post-training techniques concentrated entirely on professional legal work, using content from Westlaw, Practical Law, Checkpoint and Reuters.
One detail signals how much room is left: the company says it has used less than 10% of its content library for training so far.
The reported investment is approximately $40 million over more than two years in staff and computing power, with the final training run of the current version costing about $450,000. First deployment is in Tabular Analysis within CoCounsel Legal.
Why is the gap between $40M and $450K the real story?
The gap matters because it shows the training run is the cheapest component of a domain model, and it is the only component most organizations price.
$450,000 is roughly 1.1% of the total. The remaining ~$39.5 million went to two years of staff, data curation, evaluation design, and safety work.
That ratio is the honest budget shape for anyone considering the same path. The GPU invoice is not the barrier.
The barrier is having decades of proprietary content, the people who understand it well enough to structure it for training, and an evaluation methodology credible enough that domain experts accept the result.
Read the other way, it is encouraging: the expensive parts are things an established institution may already own. What it cannot shortcut is the curation.
What does starting from an open-weight base actually require?
Starting from an open-weight base requires an explicit step most coverage omits: deciding what the model inherited before putting it in front of customers.
Thomson Reuters worked with Imperial College to first retrain the base model for safety, ethics, and political neutrality — before the legal post-training.
That is the under-discussed part of this announcement, and the most transferable. Adopting an open-weight foundation is not only a capability decision; it is a provenance decision.
The base model arrives with training data you did not choose and behaviours you did not specify, and in a regulated product somebody has to own that inheritance.
It is also a reminder that "open weights" and "unaccountable" are not the same thing. You can inspect, evaluate, and correct an open-weight model precisely because you hold it. You cannot do any of that to an API.
Does the model outperform frontier models, or match them?
The reporting says Thomson performs comparably to frontier models on its domain tasks, with external testing by legal academics — not that it beats them across the board.
This distinction is worth preserving because several summaries have upgraded "comparable" to "outperforms," and the inflated version is both unnecessary and easy to falsify.
Comparable frontier-level performance on your own domain, from a base you did not pay to pre-train, at a fraction of frontier cost, is already a significant result. It is the claim that survives scrutiny, and it is the one a buyer can act on.
Overstating it invites the obvious counter — that a general frontier model still wins on general tasks — which was never the point of a domain model in the first place.
What should an organization take from this?
An organization should take the strategy, not the budget, and the strategy has four parts.
The model layer is commoditizing; your data is not. Open-weight foundations have reached a quality threshold where they serve as viable bases for enterprise post-training. That reframes the question from "can we match the frontier" to "what do we have that nobody else can train on."
Curation is the actual project. Budget for the 98%, not the 1%. If your proprietary content is not structured, labelled, and rights-cleared, that work is the program.
Post-training is not the only lever, and usually not the first. Retrieval over your own systems captures a large share of the same advantage with none of the training cost, which is where most institutions should start. Revolut's own foundation-model paper made the point in passing: strong downstream results came from a simple model over good embeddings of their data. We worked through that trade-off in Revolut built its own foundation model — most banks can't.
Whatever you train, you have to run it somewhere you control. A domain model trained on privileged material, then served from a vendor's endpoint, reintroduces the exposure the exercise was meant to remove.
On ibl.ai you own all the code and the data, run it model-agnostic across any LLM — including an open-weight base you have post-trained yourself — and pay with no per-seat pricing. 1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on. For the buyer's view of the product this model ships inside, see our CoCounsel alternative comparison.
Frequently asked questions
Can a mid-size firm replicate this for $450,000?
No. That figure is the final training run only, against a reported total of roughly $40 million over two years. A firm without a curated proprietary corpus and the staff to structure it should not budget from the $450K number.
Why start from a Chinese open-weight model?
Because it was the strongest available open foundation, and holding the weights allows the inheritance to be corrected — which is what the safety and neutrality retraining with Imperial College was for. That option does not exist with a closed API.
Is fine-tuning the right first step for most organizations?
Usually not. Retrieval over your own systems captures much of the domain advantage without a training program, and it is reversible. Post-training makes sense once retrieval is saturated and the corpus is genuinely large and clean.
The bottom line
The headline number that matters is not $450,000 and not $40 million. It is the ratio between them.
Domain models are within reach of institutions that hold irreplaceable data — and the cost of getting there is almost entirely the unglamorous work of curating that data, not the compute. Price the 98%.