Almost every AI system I’ve helped put into a regulated workflow ends up with the same line in its control document: a human reviews the output before anything happens. It’s the sentence that gets the thing approved: cheap to write, obviously sensible, and it moves the risk conversation along.
Then the system goes live and nobody looks at that control again. Six months on, the reviewer is approving ninety-something per cent of what the model hands them, a few seconds each, and the control document still says exactly what it said on day one. On paper, nothing has changed. In practice, the control stopped existing somewhere around week three.
A control you never measure isn’t a control. It’s a sentence.
Approval is not review
The failure isn’t laziness, and it’s genuinely hard to see. A reviewer who agrees with the model every time might not be rubber-stamping — the model might simply be right. High agreement is the expected outcome of a good system and the expected outcome of a collapsed control, and from the outside the two look identical.
So the question worth asking is never “is there a human in the loop?” It’s narrower and less comfortable: on the day the model is wrong, will this person notice, and can they act on it? If not, you don’t have oversight — you have someone absorbing accountability for a decision they didn’t really make, which is worse than no control at all, because it stops anyone looking further.
I’ve written before that irreversible actions should have a human on them. This is the half I underrated: putting a human on the action is the easy part. Keeping them able to disagree is the work.
Why reviewers stop disagreeing
Three forces push in the same direction, and none of them are anyone’s fault.
Accuracy is the problem. A model that’s wrong 40% of the time keeps a reviewer sharp. A model that’s wrong 3% of the time teaches them, over hundreds of cases, that checking carefully is wasted effort — because it almost always is. The better your system gets, the faster its human check erodes. Parasuraman and Riley named this decades ago as the “misuse” of automation — it behaves the same whether the automation is an autopilot or an alert-triage model.
The queue is the problem. If case volume doubles and review headcount doesn’t, time-per-case halves and the review quietly becomes a glance. Nobody decided that; it’s arithmetic. I’ve watched teams celebrate a productivity number that was, read another way, a precise measure of how much less scrutiny each case now gets.
Disagreement costs something. This is the one people don’t say out loud. If overriding the model means writing a justification, fielding a question about why your override rate is high, and being the person who slowed things down — while agreeing costs nothing and is invisible — you have built a system with a thumb on the scale.
The order you show things in
One design detail changes reviewer behaviour more than almost anything else, and costs nothing: whether they see the model’s conclusion before or after they’ve looked at the evidence.
Lead with the answer — “recommend: close, low risk” at the top of the screen — and everything below it gets read as confirmation. The reviewer isn’t forming a judgement; they’re deciding whether it’s worth disagreeing with one. Show the underlying material first, let them form a view, then reveal what the model concluded, and you get a genuine second opinion. Same person, same case, same model, different control.
It’s slower, which is why teams don’t do it — so reserve it for decisions where being wrong actually costs something.
What makes a reviewer real
Four things, in rough order of how often they’re missing.
Give them the evidence, not the conclusion. A reviewer who can only see the model’s output has nothing to review it against. They need the inputs it worked from — enough to reconstruct the decision, not just read it.
Budget the time honestly. Decide how long a real review of one case takes, multiply by volume, and check whether the people you have can do it. If they can’t, that isn’t a staffing problem — it’s a control that is mathematically impossible to perform. Sample properly instead, and say that’s what you’re doing.
Make disagreement cheap. An override should be one click with an optional note, not a form. And if agreeing is invisible while disagreeing gets scrutinised, fix that asymmetry first; it does more damage than any weakness in the model.
Test the reviewer, not just the model. Occasionally route a case where you know the model is wrong, and see whether it gets caught. This feels adversarial and isn’t — it’s the only direct measurement of whether the control works. Everything else is inference. Tell people you do it, and why.
The number that tells you the truth
Track the override rate — the share of cases where the reviewer changes or rejects the output — over time, per reviewer, and split by how consequential the case is.
You’re not looking for a target. You’re looking for movement. A rate drifting towards zero is the signature of a control decaying, and better to notice that yourself than have it noticed for you. Treat it as a vital sign, not a KPI — the moment anyone is measured on hitting a number, the number stops telling you anything.
This gives you something an eval set can’t. As I’ve argued elsewhere, evals tell you how often the model is right. The override rate tells you whether anyone would find out if it stopped being right.
The regulator has already noticed
If this reads as a nice-to-have, it isn’t. The EU AI Act’s human oversight article requires that the people assigned to oversee a high-risk system are enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”, and that they can “decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output”. That is a legal requirement that the reviewer be genuinely able to say no — written into the text, not inferred from it.
Worth sitting with, whether or not the Act applies to you. The question is moving from “who signs off?” to “what evidence do you have that the sign-off means something?” That second one is hard to answer from a control document, and easy if you’ve been watching the override rate all along.
If your team has AI running inside a real workflow and you want to know whether the human checks are doing anything, that’s the kind of thing we work through on a team’s own cases. Talk to us if it’s useful, or see how we run it in-house.
Sources: Article 14, EU AI Act · Parasuraman & Riley, “Humans and Automation: Use, Misuse, Disuse, Abuse”, Human Factors (1997)