The model you chose last year is not the model you would choose today
Most organisations picked their AI model provider once and never revisited it. The commercial, technical and operational arguments for running more than one all point the same way, and the operational one is where this usually fails.
Most organisations picked their AI model provider once, early, and have not revisited the decision. One contract, one integration, one set of controls to learn. It was the right call at the time.
It was also a call made about a market that no longer exists.
The position I would argue for is straightforward. Run more than one model, and treat the choice of model as a routing decision made per use case and reviewed on a cadence.
Not because variety is a virtue in itself. Because the alternative is a standing bet that one supplier will remain the best answer for every job you put in front of it, at every price point, for as long as you run it.
There are three arguments for this, and they are usually made separately by three people who do not talk to each other often enough. The commercial one, the technical one and the operational one. They point the same way.
The commercial argument
Most AI work in an enterprise is not hard. Classifying an inbound email, extracting five fields from a form, routing a request, summarising a short document. That work sits alongside a much smaller volume of genuinely difficult reasoning, and it is the difficult reasoning that sets your unit price when everything runs through the same endpoint.
Gartner forecast in May 2026 that worldwide AI spending would reach $2.59 trillion in 2026, up 47 per cent year on year, with vendors and hyperscalers driving most of it. Whatever you make of the number, the direction is the useful part. Supply-side investment at that scale means a release cadence you cannot plan around and prices that will move in both directions.
There is a negotiating point underneath this that finance teams grasp faster than technology teams do. An organisation with one integrated provider renegotiates from a weak position, because the alternative to agreeing is a rebuild. An organisation that can move some of its workload elsewhere in a fortnight is having a different conversation.
The technical argument
Not every task wants the same model, and the gap is about quality as much as cost.
Gartner predicted in April 2025 that by 2027 organisations would use small, task-specific models at more than three times the volume of general-purpose large language models, noting that accuracy on general-purpose models falls away on tasks needing specific business domain context. The specialised model is often not merely cheaper. It is better at the narrow thing.
A single-provider architecture has no way to express that difference. Every task gets the same model whether the job deserves it or not.
Then there is availability. If a core process runs through one provider’s endpoint, that provider’s outage is your outage, their deprecation notice is your migration project and their rate limit is your queue. Fallback is not exotic engineering. It is the same thinking applied to a payments gateway or a messaging provider, and nobody argues about it there.
The operational argument, which is where this usually fails
This is the part worth spending the money on.
Model-agnostic gets written into the target architecture, an abstraction layer gets built, and then nobody ever changes the model. The capability exists on paper and dies in practice, because switching was treated as a technical property of the system rather than as something a person has to propose, test and approve.
Running more than one model is an operating discipline. It needs three things that are usually missing.
- An evaluation set per use case. Thirty to fifty real examples with agreed correct outputs, owned by the business function that consumes the output rather than by the build team. Without it, whether the new model is better becomes a matter of opinion, and opinion loses to inertia every time.
- Cost and quality attributed per use case. Not a monthly invoice from a provider, but a view of what each workflow costs to run and how well it performs. You cannot route by cost if you do not know which use case is spending the money.
- A change route. Who proposes a model change, who runs the evaluation, who approves it and who tells the users. Where that route does not exist, the answer to every proposed change is no by default.
Changing the model is a change to the service
The adoption risk here gets budgeted at zero, and it should not be.
People calibrate on behaviour. A team that has spent three months learning how the drafting assistant phrases things, when it declines and how it formats output, has built a working model of the tool in their heads. Change the model underneath them on a Tuesday without saying so and that mental model breaks. Output that is no worse on any measure still costs you trust, and usage follows trust.
So a model change is a release. It needs a note to users, a named owner and a route for someone to say the new one is worse.
Treat it as an invisible infrastructure swap and you pay for it in adoption, which is the metric that eventually decides whether any of this was worth doing.
When one model is still the right answer
Standardising on one provider genuinely holds when the footprint is narrow and stable. Two or three contained use cases, modest volume, a small team, and more value in simplicity than in flexibility. Adding a routing layer to that is engineering you do not need.
It stops holding at a recognisable point. When AI sits inside a core process rather than beside it. When the bill becomes a line someone asks about at board level. When a provider outage would stop a business function rather than inconvenience a team.
Most organisations cross that line without noticing, and review the provider decision afterwards, under pressure.
What to do this week
Three things, none of which need a budget.
- List every AI use case you run and the model behind it. If nobody can produce that list in an afternoon, that is the finding.
- Take one low-risk use case and build a thirty-example evaluation set with the team that owns the output. Run it against your current model and one alternative. You now have evidence rather than a preference.
- Write down who would have to approve a model change and how long it would take. If the answer is unclear, you do not have a multi-model capability. You have an abstraction layer.
The goal is not to run six models. It is to be able to change the one you run without it becoming a project.
