An AI customer service agent receives a request for a refund. It has access to the customer's purchase history, previous conversations and the company's refund policy. It checks the facts against the policy and issues the refund. Fine.
But what if the same AI decides to close a customer's account for suspected fraud? Rejects a vulnerable customer's complaint? Denies someone access to a financial product? The technology may well be capable of making all of these decisions. The more important question is whether it should make them alone.
Why this question is becoming more urgent
Until recently, most customer-facing AI recommended or assisted. AI agents can now retrieve information, make decisions, access other systems, carry out actions and communicate the outcome without a person intervening at each step.
That shift is what makes "human in the loop" a live operational question rather than an abstract governance principle. Once AI can act as well as advise, the question stops being "should AI help our employees?" and becomes “which customer-facing actions are organisations actually prepared to let AI take on their own?”
Three models of oversight
"Human in the loop" can describe several different arrangements, each suited to different decisions. The three main models, echoed in the EU's Ethics Guidelines for Trustworthy AI, are:
Model | What the AI does | What the human does | Better suited to |
Human in the loop | Proposes an action | Approves or rejects before it happens | High-impact or ambiguous decisions |
Human on the loop | Acts within defined boundaries | Monitors outcomes and can intervene | Medium-risk, largely predictable workflows |
Human out of the loop | Acts independently | No real-time involvement | Low-risk, easily reversible tasks |
None of these is inherently correct. A password reset, a routine delivery update or a straightforward FAQ answer works well out of the loop; requiring sign-off for each would add friction without reducing risk. The harder question is how an organisation decides which model applies to which decision.
The organising principle: what happens if the AI gets it wrong?
The most useful starting question is not whether AI is "safe enough" in general, but what the consequence would be if a specific decision were wrong. That question does most of the work, but on its own it is too blunt. It does not tell you whether monitoring is enough or whether a person needs to approve the action before it happens. Three further factors matter alongside it: how easily the decision can be reversed, how confident the AI actually is in this specific case, and how much ambiguity the situation involves.
Four questions for AI autonomy
1. Consequence: What happens if this decision is wrong: financial loss, harm to the customer, loss of access, reputational damage?
2. Reversibility: Can the decision be corrected, and can the customer be restored to their previous position, or does the action cause lasting harm once taken?
3. Confidence: How certain is the AI in this specific case, rather than in general? A model with strong average accuracy can still be uncertain about an individual, unusual case.
4. Ambiguity: Does the situation fit cleanly within predefined rules, or does it involve circumstances a policy was never written to anticipate?
Together, these questions provide a practical way to decide how much autonomy a decision should have. If human oversight is required, it must also be capable of genuinely influencing the outcome.
The four factors point towards one of three broad postures. Low-consequence, reversible and predictable cases can run autonomously within defined boundaries. Moderate-consequence cases or those involving genuine uncertainty call for AI to act within policy, with monitoring and a clear escalation route. High-consequence decisions involving irreversibility or ambiguity are more likely to require human review before the action is taken rather than afterwards.
Why nominal oversight fails
Putting a person at the end of an AI workflow does not automatically create meaningful oversight, and that follow-up question, whether a human can genuinely intervene, is where good intentions often break down in practice.
An employee asked to approve 500 AI-generated recommendations a day, in seconds each, is technically a human in the loop. In practice, automation bias can encourage people to trust the system, volume makes scrutiny difficult, reviewers may lack context, and targets focused on speed can discourage careful judgement. The ICO makes a similar point. Rubber-stamping does not constitute meaningful human involvement, and reviewers need genuine authority and competence to challenge the system.
Meaningful oversight requires information, authority, time and competence. Without them, the human is not supervising the AI. They are signing off on it. This is one of the specific failure modes that well-designed AI guardrails are meant to catch, rather than something a policy document alone can prevent.
Where humans genuinely add value
None of this means people are simply a safer default. Humans can be inconsistent, slow to decide, and just as prone to bias as a model trained on historical data; a rushed or poorly briefed reviewer will sometimes make a worse call than the system they are checking.
Where human judgement earns its place is with problems that resist reduction to a rule. Consider a customer repeatedly requesting compensation after a disrupted journey. AI can calculate what they are contractually owed. A person may recognise context the data does not capture, such as distress, a family emergency, or a technically correct outcome that would be disproportionately harsh. Humans can also challenge the underlying business rule itself, rather than simply apply it, and they allow clear accountability to remain with a person or organisation responsible for the decision. That is different from the routine division of labour between people and AI in everyday CX work. What matters here is what happens when the decision itself is consequential, ambiguous or contested. Who should have the final say?
Designing autonomy at the level of the action
Rather than treating a whole AI system as either autonomous or human-controlled, the more workable approach is to set the level of autonomy per action, using the four questions above. FAQs and password resets can run fully autonomously. Refunds under a defined threshold might run within policy with monitoring, while larger ones require approval. Account closures, eligibility decisions and actions involving suspected fraud or potentially vulnerable customers are more likely to warrant meaningful human review before action is taken.
The exact thresholds will differ by organisation and risk appetite, but the principle carries across all of them. Autonomy is a property of the decision, not a single setting for the system as a whole. The same AI agent might resolve a delivery query without any human involvement, yet be required to escalate before it closes an account. This is as much a matter of AI governance in a customer experience context as a technical one, and in some contexts a regulatory one too. In the UK, solely automated decisions that have a legal or similarly significant effect on an individual are subject to specific restrictions and safeguards under data protection law, including the right to obtain human intervention and to contest the decision.
There is also a temptation to treat human escalation as the default answer to AI risk. But routing everything to a person has its own cost. People are slower and more expensive than a model, and can be less consistent on repetitive, data-heavy tasks. If every AI decision needs sign-off, an organisation has not built autonomy, it has built a more complicated workflow with an AI system embedded in it. What matters is not maximum human involvement, but involvement matched to consequence, reversibility, confidence and ambiguity.
The goal is not more human involvement, it is better-designed boundaries
Human in the loop is sometimes presented as a simple fix for AI risk: if something matters, put a person in the process. The reality is less tidy. A person in the process is not the same as a person able to meaningfully oversee it, and too much oversight can be as costly as too little.
The real work is deciding, decision by decision, where AI can be trusted to act, where it needs monitoring, and where a genuinely capable human needs the final word before anything happens. As AI takes on more responsibility for customer decisions, the organisations that benefit most will be those that are clearest about where that AI can act, where it needs supervision, and where a human should retain the final say.

