Human in the Loop Is Not Enough: Who Is Actually Accountable for the AI Decision?
Sep 05, 2026
Last updated: September 5, 2026
For more than a decade, sub-postmasters across the United Kingdom rang a helpline to report the same thing: the accounting system said money was missing from their branches, but it was not. Call center staff were told to tell them they were the only ones having trouble with the software. They were not. Their contracts obliged them to accept the system's figures and cover the shortfall themselves. Some remortgaged their homes. Some went to prison.
At every step there was a person: an auditor, an investigator, a prosecutor, a manager on the end of the phone. Each had a human being telling them the numbers were wrong, and a computer saying otherwise, and each believed the computer.
Sir Wyn Williams, who chaired the public inquiry, reported in July 2025 that roughly 736 sub-postmasters were wrongly prosecuted between 2000 and 2013, and thirteen took their own lives. The Post Office, he found, "maintained the fiction that its data was always accurate."
It would be comforting to file that as a scandal and move on. It is not. It is what people do when they are handed an answer.

In 2023, researchers at three German university hospitals asked 27 radiologists to assign BI-RADS scores to a set of mammograms, with an AI system offering its own suggestion for each image. When the AI was right, the radiologists were accurate about 80 percent of the time. When it was wrong, accuracy collapsed: 19.8 percent among the least experienced readers and 45.5 percent among those with more than 15 years of mammography experience.
Nobody bullied those radiologists. They were experts looking straight at the evidence, under no obligation to agree, and one wrong suggestion halved their accuracy.
The reason is priming or anchoring. Once you have been shown an answer, you are no longer looking for one; you are checking, and attention drifts toward whatever fits. The research literature calls it automation bias and anchoring bias; a 2026 study in European Radiology found each in about a third of cases where the AI suggestion had been manipulated. It is a property of human cognition, not a failure of diligence. You do not fix it by trying harder.
Every one of those radiologists was a human in the loop. So was everyone in the Post Office chain.
The moral crumple zone
In a NASA-supported study of glass-cockpit pilots, Kathleen Mosier and Linda Skitka found something stranger than over-reliance: pilots reported seeing warning indicators that never appeared. The automation did not bypass their judgment; it rewrote their memory. Parasuraman and Manzey's 2010 review concluded that automation bias cannot be eliminated by training alone and affects experts as reliably as novices. So when a board is told the AI has human oversight, the honest translation is usually: we have found someone to sign it.
Madeleine Clare Elish named the arrangement in 2019: the moral crumple zone. The human becomes the component that absorbs the moral and legal impact when the system fails, as a crumple zone absorbs a crash. The system is protected. The nearest person is not.
Ben Green's survey of 41 government policies requiring human oversight of algorithms found that these policies provide false assurance, allowing agencies and vendors to shirk accountability.
The Post Office chain is a textbook case. So is the Netherlands, where risk profiling flagged childcare benefit claims as fraudulent and roughly 26,000 parents were wrongly accused between 2005 and 2019. The parliamentary inquiry Ongekend onrecht was blunt: officials had the power to assess proportionality on a case-by-case basis and did not use it. The cabinet resigned.
Those humans held real discretion and behaved as though they held none.
The law already agrees, and almost nobody has read it
I have written elsewhere about why human judgment still matters as AI takes over more of the work. This post is about something narrower and more uncomfortable: whether the specific person at the checkpoint is capable of exercising that judgment, and what makes them capable. The European legislature has already answered part of that in statute, and most organizations rolling out human-in-the-loop have not read it.
Article 14 of the EU AI Act does not say a human must review the output. It says the overseer must understand the system's capacities and limitations, monitor it for anomalies, interpret its output, and remain aware of what the Act calls "automation bias." They must be able, in the phrase worth pinning to a wall, to "decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output." Article 26(2) tells deployers to give that job to people with "the necessary competence, training and authority, as well as the necessary support." That is a capable person, not a process step.
The timing has moved. The AI Omnibus deferred these obligations to December 2027 for standalone high-risk systems and to August 2028 for embedded ones. That is not breathing room.
Building people who can overrule a machine takes longer than writing the policy that says they may.
The four capabilities a checkpoint actually needs
Automation bias is a self-leadership problem before it is a governance problem. Four things separate a checkpoint that holds from one that only looks like it.
Agency. The person can say no, and saying no is survivable. If the reviewer is measured on throughput, you have designed compliance and called it oversight. Agency is the practiced habit of choosing to stop outsourcing your decisions, matched by real authority.
Judgment. The reviewer needs to form a view before seeing the machine's answer. Once the output is on screen, you are no longer deciding; you are rating. Decision-making under pressure is the whole job here.
Accountability. In that same cockpit research, pilots who felt accountable cross-checked the automation and made fewer errors. Not accountability in the abstract sense that somebody will be blamed later, but a named person who knows the decision is theirs. Diffuse accountability is how a crumple zone forms.
Learning. When someone overrides the system, or fails to, something has to change. Most organizations run single-loop learning: retrain the model, move the threshold, add a warning. Double-loop learning, in Argyris and Schön's sense, asks the underlying question: should this be delegated to a machine at all, and what belief led us to think it should?
Diagnose before you prescribe
My first career was in physiotherapy, and it left me with a rule I have never shaken: prescription without diagnosis is malpractice.
A checkpoint asks a person to sign off on a prescription they never diagnosed. That is the design flaw. Stop asking whether there is a human in the loop and ask four questions about that person. Can they say no and survive it? Can they form a view before the machine gives them one? Does the decision carry their name? When they get it wrong, does anything upstream change?
If the answer to any of those is no, you do not have oversight. You have a signature.
The organizations that come through this will not have the best models. They will have taken human potential in the age of AI seriously enough to build people who can stand behind a decision, including the decision to overrule it.
Through my work as a leadership speaker and with Executive Leadership Teams, I support developing these four capabilities for humans in the loop. Contact me if you want to discuss this.
FAQ
What does "human in the loop" actually mean?
Human in the loop describes any AI system where a person reviews, approves, or can intervene in the output before it takes effect. In practice, it ranges from a reviewer with genuine authority to reject an output to someone clicking approve in a queue. The phrase says nothing about which of those you have.
Is human in the loop enough to make AI safe?
No. Decades of research show that people in monitoring roles reliably defer to the machine, experts included. Parasuraman and Manzey's 2010 review found that automation bias cannot be eliminated by training alone and affects experts as much as novices. A human at the checkpoint reduces risk only if that person has the authority, the independent judgment, and the named accountability to disagree.
Who is accountable when an AI system makes a bad decision?
In practice, it lands on the nearest human operator. Madeleine Clare Elish calls this the moral crumple zone: the person absorbs the blame while the system that produced the error is left intact. Design against that. Accountability belongs with a named person who had the information, the competence, and the authority to decide otherwise, and with the leaders who chose to deploy the system.
What makes a human capable of overseeing an AI decision?
Four things. Agency, meaning they can refuse the output and that refusal is survivable. Judgment, meaning they can form an independent view before seeing the system's answer. Accountability, meaning the decision carries their name. Learning, meaning overrides change something upstream instead of disappearing. The EU AI Act frames the same requirement as competence, training, authority, and support.
How do you train people to exercise judgment over AI output?
Training alone will not do it, and the research is clear on that. You have to change the conditions of the role as well as the capability of the person. Require an independent assessment before the AI output is visible. Name one accountable decision-maker, not a committee. Treat override rates as a metric people are expected to have rather than penalized for. Review overrides in a forum that can question the deployment decision itself. Then train self-awareness and self-regulation, because a reviewer who cannot notice their own deference under time pressure will defer.
Turn Insight Into ActionReading about leadership can change what you know.Self-leadership changes what you do.Develop your leadership effectiveness → Explore Self-leadership Discover your leadership strengths & gaps → Choose an Executive Coach Unlock the potential of your team → Workshops & Masterclasses Ignite your event or conference → Book Andrew Bryant as your Leadership Speaker "Prescription without diagnosis is malpractice" |