Abstract
The competitive advantage in enterprise AI is no longer the model — it is the proof. As open and Chinese-developed models reach performance parity with proprietary flagships at a fraction of the cost, regulated businesses face a strategic inflection point: capability is abundant, but trust remains scarce and hard-won. This paper argues that for banks, insurers, and health systems operating under demanding regulatory scrutiny, the central question has shifted from which model should we use to who can prove a model is safe, accurate, and fit to deploy.
Drawing on 2026 industry data — including open models now accounting for a reported majority of AI usage among U.S. companies on major model marketplaces — the paper maps the five-stage lifecycle every regulated model must traverse (Select, Qualify, Adapt, Operate, Retire) and identifies where governance most commonly fails: late, reactive, and impossible to reconstruct under audit pressure. It makes the case that independent assurance must enter at the point of selection, not after deployment, transforming compliance from a trailing checkpoint into a front-loaded standard.
The paper also examines the two honest paths to deploying open models — as a service or in-house — along with their respective trade-offs on compliance, data sovereignty, and operational burden. It concludes with a call to reorient AI strategy around repeatable qualification discipline rather than model selection: the model is rented and will be replaced within a year; the assurance infrastructure, evidence trail, and institutional trust built around it are what compound in value. In a regulatory environment demanding documented proof — not reassurance — that discipline is both the differentiator and, increasingly, the price of entry.
The most capable freely available AI models in the world increasingly come from open labs, many of them Chinese. They are powerful and remarkably cheap. For a bank, insurer, or health system, the interesting question is no longer whether these models are good enough. It is whether anyone can prove a given model is safe, accurate, and fit to use. This article lays out how that shift is progressing and what it means for anyone deploying AI in regulated industries where the rules are unforgiving.
What just changed
A growing number of labs have now released models that have reached competitive performance, dramatically expanding the field of options. Many release genuinely strong models to the public, free to download and run, and refresh them every few weeks. The performance gap with big proprietary systems has narrowed to the point where, on many everyday business tasks, it is hard to tell them apart.
The clearest sign of the shift is not a benchmark, but the behavior. On the largest neutral marketplace where companies route their AI traffic, open models went from a rounding error at the start of 2025 to a reported majority of the usage attributed to U.S. companies by the middle of 2026. Businesses are voting with their workloads, and they are moving toward models that are open and inexpensive.
For most industries this is simply good news: more capability for less money. For regulated industries it is more complicated, and to my mind more interesting. The same qualities that make these models cheap they are open, they come from many places, they change constantly also make them somewhat harder to govern. A bank cannot adopt a model just because it is popular. An insurer cannot explain to a regulator that it chose a system because it topped a leaderboard last month. The center of gravity in enterprise AI has moved from a simple question, which model should we use, to a far harder one: Who can prove this model is safe, accurate, and right for the job. The hard part of AI is no longer finding a capable model. It is proving the one you chose is fit to deploy and being able to show your work.
This matters if you run a regulated business
Strip away the noise and the appeal of these models comes down to four things that speak directly to the pressures a regulated enterprise already feels. The first is cost. Open models can run a fraction of the price of the proprietary flagships while holding their own on the tasks that fill a working day drafting, summarizing, classifying, answering routine questions. When you are processing millions of documents or claims under a fixed-price contract, that difference does not trim the budget but reshapes the economics of the work.
The second is control over your own data. Because the models are yours to download, you can run them inside your own secure environment. Sensitive records patient data, customer information, claims files never have to leave your system or travel to someone else’s servers. For a compliance officer, that single fact removes the objection that sinks most AI proposals before they start.
The third is independence. Relying entirely on one AI vendor leaves you exposed to its price changes, its product decisions, and its politics. When you own the model, a sudden price hike, a discontinued product, or a new trade restriction becomes someone else’s problem, not a threat to your operations. The fourth is ownership. An open model can be trained further on your own history your actuarial tables, your legal precedents, your past decisions to become a specialist that understands your business better than any general-purpose system. The result is not a rented capability. It is an asset you own.
None of this makes the decision easy. It simply moves the difficulty to where it belongs: not “is this model good,” but “can we trust it, and can we prove it.” And the moment a regulated business seriously considers one of these models, it runs into a wall.
A regulated model moves through five stages, and each one carries its own risk:
Select: Choosing which model to bring in. Risk: picking on benchmark hype or price alone, missing licensing/jurisdiction problems that surface later.
Qualify: Proving it's fit before it touches production. Risk: shallow testing that automated scores pass but real-world edge cases fail.
Adapt: Fine-tuning to the business or wiring in guardrails and retrieval. Risk: the tuning introduces new bias or leakage nobody checks for.
Operate: Running it against live workloads. Risk: drift — the model's behavior degrades as data and conditions change, silently.
Retire/replace: Moving to a newer model when the current one is surpassed. Risk: without a clear transition process, the organization either holds on to outdated technology or replaces it blindly.
The model itself only occupies the middle three stages, and only for a matter of months. The assurance wrapped around it spans the entire lifecycle and outlives any single model. Every model you deploy will complete this cycle and most likely exit it within a year or two. The question is whether the discipline that carried it through survives to carry the next one, because that discipline is the thing you keep.
Most organizations bolt on governance at the Operate stage, when a model is already live and an auditor is already asking questions. By then, documentation is reconstructed after the fact, logs are in engineer-formats no auditor accepts, and model versions were tracked inconsistently. This is a real, documented failure pattern in 2026 (the Forbes "audit-trail gap" article describes exactly this: deployments scoped for business value with documentation as an afterthought). Fixing it retroactively is slow, costly, and sometimes impossible.
Bringing in the right external discipline changes where on the timeline assurance enters. Instead of assurance as a late checkpoint, it becomes the entry criterion at Select and Qualify before a single regulated decision runs through the model.
- The licensing/jurisdiction problem is caught at Select, not discovered in an audit.
- Bias, leakage, and domain-fitness are proven at Qualify in days, not the months an internal team would take standing this up from scratch.
- The evidence trail is generated as the model is adopted, not reconstructed under audit pressure later.
- When the model is inevitably replaced, the qualification process is already repeatable, so the next model clears in days too.
The frameworks and approaches across the AI value chain/model lifecycle exist to move that work to the front of the lifecycle turning assurance from a scramble at the end into a standard at the start.
The wall every regulated buyer hits
There are only two honest ways to use one of these open models, and neither is a free lunch.
You can use the model as a service, calling it over the internet the way you would any cloud tool. That library does not depreciate when the next model ships. It gets more valuable. It is the easiest path and the cheapest to start. But for most regulated work it is a dead end: the hosted versions of these models generally cannot meet the standards that govern healthcare, payments, and financial data, and your information may be processed under foreign law. For sensitive workloads, that is usually where the conversation ends.
Or you can bring the model inside download it and run it on infrastructure you control. This solves the data problem cleanly: nothing leaves your environment. The trade is that you now own the operation of it, and you are running something that may be outdated within months, because the field moves that fast.
What matters to you | Use it as a service | Bring it in-house |
|---|---|---|
Compliance | Usually cannot meet healthcare, payment, or financial-data standards. | Your data stays inside your walls. The safe default for regulated work. |
Reliability of access | Access can be cut off with little warning, for reasons outside your control. | Once it’s yours, it keeps working regardless of politics or trade rules. |
Effort | Someone else runs it. | You run and maintain it, and plan for the model to be replaced over time. |
Cost | Lowest to begin, you pay as you go. | Pays off at scale and when the capability is core to what you sell. |
Whichever path you take, the risk does not disappear but moves – and that is the real lesson. The choice of model is not where success or failure is decided. It is decided by everything wrapped around the model, and by the evidence you can put in front of a regulator, a board, or a customer.
The asset is the assurance, not the model
Here is the conviction underneath everything above: whatever model you choose today will be outclassed within a year. Chasing the best model is chasing a moving target. What lasts is everything else the guardrails, the checks, the record of what the system did and why, the human oversight, and above all the independent evidence that the model behaves as it should. That is the part that compounds in value. The model is rented. The trust is built.
Industry testing has made this concrete. In recent evaluations, a free open model was brought up to the quality of a top proprietary system on demanding tasks at roughly a fraction of the cost without changing the model at all. The gains came entirely from the engineering around it. Most failures in real deployments, it turns out, come not from the model but from weak scaffolding and the absence of rigorous testing.
If that is right, then the strategic priority for a regulated business invert. Do not organize an AI program around picking a winner among models; organize it around the discipline that lets you adopt any model quickly, run it safely, and prove it works. Build the ability to test a new model in days rather than months. Keep the model at arm’s length from the moment of decision, so that in a loan denial or a claim rejection it is rules and evidence that decide, and the model only drafts and explains. Insist on evidence before deployment, not reassurance after. Do these things and the identity of this quarter’s best model stops being a strategic question at all. Capability is now the easy part. Proof is the hard part and the valuable one. That is where I would put the investment.
The frameworks to do this exist. They enable independent qualification at the point of selection. Repeatable testing that clears a new model in days rather than months. Routing of routine work to the most cost-effective model that clears the quality bar, with expensive systems reserved for genuinely hard problems. Specialist models trained on a client's own history i.e., policies, precedents, past decisions and handed back as owned intellectual property rather than a rented service. Assurance that extends across images, audio, and video, not text alone. Each engagement builds a reusable body of evidence about how models actually behave, which is what keeps qualification costs falling and proof available on demand.
EXL works with regulated businesses to put this discipline in place from model selection and independent qualification through to operation and replacement. If the challenge described here is one you are facing, it is worth a conversation.
The timing favors this view
Regulation is moving in exactly this direction. Europe’s AI rules are now in force, with the most consequential obligations landing through 2026. In the United States, in the absence of a single federal law, dozens of states have written their own well over a thousand AI bills in the past year alone and industry regulators are adding their own expectations, from insurance supervisors demanding clear audit trails for automated underwriting to established rules in healthcare and payments. The common thread across all of it is the same word: proof. Regulators no longer want assurances that a system is responsible. They want documented, independent evidence.
That is why the market for testing and assuring AI rather than merely building it is one of the fastest-growing corners of the industry, worth several billion dollars today and expanding at double-digit rates. The direction of travel is unmistakable: the discipline I have described is shifting from a competitive edge to a condition of doing business.
The bottom line. The new generation of open models has made AI capability abundant and cheap. It has not made it safe by default, and for a regulated business the difference is everything. Stop treating model selection as the main event. Adopt the best of what is available, run it on your own terms, and put your energy into the assurance around it because when a regulator, a board, or a customer asks the only question that now matters, the answer that counts is not which model you chose. It is where is the proof.