Ask a compliance team to list the AI in their firm and you get a list of projects: the chatbot, the document tool, the pilot running in operations. It’s a reasonable list, and in most firms I’ve looked at it accounts for a minority of the AI actually in use.
The rest arrived as features. The case management system added summarisation. The email security tool started scoring intent. The recruitment platform began ranking candidates. Nobody ran a project, nobody completed an assessment, and no one signed anything. A supplier shipped a release, a toggle defaulted to on, and a model started forming views about your customers and your staff.
Your AI register lists the systems you decided to build. Your exposure is mostly in the ones you were sold.
The governance process is attached to the wrong event
Most AI governance triggers on a decision — a project starts, a budget is approved, an assessment gets filled in. That works for AI you set out to acquire. It catches nothing when the capability arrives inside a product procured three years ago for an entirely different reason, and the vendor’s roadmap does the deciding.
So the first useful move isn’t a new framework, it’s a change of question. Instead of “what AI have we deployed?”, ask “which of our suppliers has shipped an AI feature in the last eighteen months, and which of those are switched on for us?” You can’t answer that from the AI register. You answer it from the supplier list, one line at a time, and the list is always longer than anyone expects.
Bought AI fails differently from built AI
The standard complaint about vendor AI is that you can’t see inside it. True, and not really the problem — you can’t see inside a model you built either, which is why we test behaviour rather than read weights. Three other differences matter more.
It changes underneath you. A vendor swaps the model behind a feature — a newer version, a cheaper provider, a reworked prompt — and the behaviour shifts. Nothing changed on your side, and nothing appears in your change management. The validation you did in March now describes a system that no longer exists, and the only signal is that the outputs feel slightly different, which nobody is watching for.
The chain is longer than the contract. Your supplier has a model provider, and that provider has its own release cycle, deprecation schedule and region choices. Your contractual visibility usually stops at the first hop, while the thing most likely to change your outcomes sits two or three hops down.
The failure lands on your side of the table. When a vendor’s screening model quietly starts flagging a different population, it’s your customer who gets declined and your firm that has to explain it.
Be honest about what a questionnaire tells you
Due diligence questionnaires are the standard control here, and they’re worth roughly what they cost. A completed DDQ tells you the supplier employs someone who can write “human oversight” and “bias testing” in the right boxes. It is evidence that a policy exists, not evidence of behaviour, and it’s a self-assessment from the party with the least incentive to volunteer bad news.
That isn’t an argument for abandoning them. It’s an argument against treating a returned questionnaire as the end of the assessment — which is what usually happens, because the file is now complete and the file was the goal.
What to ask for instead
Four things, in the order I’d fight for them.
Notice of material change to the model. This is the highest-value clause in the whole agreement and the one almost nobody asks for. Define “material” against your use, not the vendor’s: a change that could alter outcomes for your customers, whether or not the supplier considers it a new feature.
The right to test, with somewhere to do it. A sandbox and permission to run your own cases through the feature, before go-live and after any change. This converts every other assurance from a promise into something you can check.
Disclosure of the chain. Which model provider, in which jurisdiction, whether your data is retained, and whether anything you send is used for training. If a supplier can’t answer this quickly, they don’t know, which is itself the answer.
A pinned version, or a change window. Some suppliers will hold a version for you. Most won’t, and then what you want is advance notice and a window in which to test before it reaches production.
You will not get all four from a large platform vendor where you have no negotiating power. What you can always do is establish which ones you didn’t get and size the risk accordingly, rather than filing the DDQ and calling it managed.
Test what you bought the way you’d test what you built
An eval set doesn’t care who wrote the software. If you have a set of real cases with an agreed view of what a good outcome looks like — the approach I set out in you can’t unit-test an agent — you can run it against a vendor feature on a schedule and watch the score. It’s a modest amount of work, and it’s the only mechanism I know of that tells you the day a supplier changed something without saying so.
Accountability doesn’t transfer with the invoice
Buying a system does not buy you out of responsibility for what it does. Supervisors have taken that line on outsourcing for years, and the AI-specific rules are more explicit. The EU AI Act runs obligations along the value chain and treats a deployer who puts their own name on a high-risk system, or substantially modifies it, as a provider with a provider’s duties. In EU financial services, DORA has since January 2025 set out what a contract for third-party ICT services must contain, including rights to monitor and to exit.
Different instruments, same message: “that’s handled by our supplier” describes where the work happens, not where the accountability sits. It’s the same continuous-ownership habit I described when writing about AI risk standards, extended across the supplier boundary.
Start with one page, not a programme: every supplier, the AI features they now ship, whether each is switched on, what decisions it touches, and whether you would find out if it changed. That last column is the point. For most firms it reads “no” all the way down, and it will tell you more about your exposure than the register you already maintain.
If your team is working out how to supervise AI that arrived through suppliers rather than through a project, that’s the kind of thing we help teams get right. Talk to us if it’s useful, or see how we run it in-house.
Sources: EU AI Act, Article 25 — Responsibilities along the AI value chain, DORA (Regulation (EU) 2022/2554).