Ask a room of risk and compliance people how they plan to control an AI system, and within a minute somebody says “we’ll keep a human in the loop.” The room relaxes. The phrase does a lot of work — it sounds like oversight, it sounds like accountability, and it drops neatly into a policy document.
Then I go and look at the loop. What I usually find is one person at the end of a queue, approving a few hundred AI outputs a day, with no practical way to check any of them.
A human in the loop isn’t a control. It’s a job you’ve given someone — and the control is whether they can actually do it.
The approval step that approves everything
Here’s the number I ask for first: what proportion of the AI’s output does the reviewer change or reject? When the answer is “almost none,” one of two things is true. The system is genuinely excellent, or the review isn’t really happening.
From the outside those two look identical. Same green dashboard, same clean audit trail, same signed-off record for every item. That trail is the trap: it is evidence of clicks, not of thought. It will satisfy an auditor who asks whether there was human oversight, and not survive one who asks what the human saw.
I’ve built these queues, and the shape of the mistake is always the same. The reviewer is shown the model’s answer and a confirm button. Not the inputs. Not what the model retrieved, or ignored. Just a conclusion, and a request to agree with it.
Why capable people rubber-stamp
This isn’t laziness, and it isn’t a training problem. It’s what the design makes rational.
Automation bias is real and it has a name. People systematically under-scrutinise a machine’s recommendation, accepting suggestions they’d have questioned from a colleague. It’s established enough that the EU AI Act writes it into law: the human oversight duties for high-risk systems require the person to stay aware of the tendency to rely or over-rely on the output.
The effort is wildly asymmetric. Agreeing costs one click. Disagreeing costs twenty minutes of digging, a written justification and an awkward conversation with whoever owns the queue. Nobody is ever asked to explain an approval. Under volume, any sensible person takes the cheap path — and the design chose it for them.
The answer arrives before the judgement. Show someone a conclusion first and you’ve made them a checker rather than a decider. There is a real difference between “here’s the case, what do you think?” and “here’s the answer, do you agree?” The second gets a yes far more often, from the same person on the same file.
The throughput target has already spent the time. If the business case promised the AI would cut handling time sharply, that saving was booked before anyone costed the review. The reviewer inherits a quota that only works if they don’t look properly — and nobody tells them that’s the deal.
What a reviewer actually needs
Four things, and none of them are a longer policy.
The evidence, not just the answer. The reviewer needs what the system was given, what it consulted and what it did — the same trace I’ve argued is the real artefact when an agent goes wrong. A conclusion on its own can only be judged for plausibility, and plausibility is what these systems are best at producing.
Time that someone has actually budgeted. Work out the honest minutes-per-case, multiply by volume, and put that number in the plan before anyone books the saving.
A cheap, unpunished way to say no. If rejecting means a form, an escalation and a delay to the queue, you’ve priced rejection as a punishment. The override path should be as easy as the approval path, and nobody should have to defend the fact that they looked.
A case they’re able to judge at all. Some decisions simply cannot be checked from the output. Whether a screening alert is a false positive — checkable, given the underlying record. Whether the model weighed the right four hundred documents — not checkable by anyone in ninety seconds. When the answer is “not checkable,” human review is the wrong control, and you need a different one: a narrower scope, a hard limit in code, or sampling with real depth.
The number that tells you it’s working
Track the override rate — how often the human changes or rejects the machine — per reviewer, over time. It’s the most informative number in the whole arrangement, and almost nobody watches it.
A rate of zero is not a triumph. It’s an alarm. And the pattern to watch for isn’t a sudden break, it’s a slow fade: a rate that drifts down month after month as reviewers learn the system is usually right and quietly stop reading. Controls like this don’t fail loudly. They decay.
Pair it with one habit. Take a small random sample of approved items each month and have someone re-do them properly, from the evidence, without seeing the original decision. That blind re-check is the only thing that distinguishes “the system is excellent” from “nobody is looking.”
Review fewer things, properly
The instinct in a regulated firm is to review everything, and it’s the wrong one. Reviewing 100% of cases at nine seconds each is worse than reviewing 10% with time to think: the first gives you a false record of oversight, the second gives you real evidence.
Sort by consequence, not by volume. The cases that deserve a person are the irreversible ones, the large ones, the ones involving a customer who can’t absorb a mistake — not simply the ones that arrived today.
Let the boring cases through. If a category of decision is low-stakes and easily reversed, spend the review budget elsewhere and accept that some will be wrong. That is a choice you can defend.
Spend what you save on the hard end. The cases where the model is uncertain, the ones sitting near a threshold, the ones that don’t resemble anything it was trained on — exactly the cases a nine-second review is guaranteed to miss.
If your policy says “human in the loop,” make it answer three questions: who reviews, what can they see, and what happens when they say no. If it can’t, the policy isn’t describing a control. It’s describing a hope, with a signature attached.
If your team is putting AI into decisions that carry real consequences and wants oversight that would survive being examined, that’s the kind of thing we help teams get right. Talk to us if it’s useful, or see how we run it in-house.