‘There is a human-in-the-loop’ has become one of the most reassuring sentences in AI.
It’s also one of the least useful.
All this does is tell us that a person appears somewhere in the process. It doesn’t tell us whether they understand the output, have enough information to challenge it or if they can change the outcome.
A person reviewing an AI output doesn’t automatically make a system safe. They need the time, authority, information and competence to intervene. The workflow needs to make intervention possible. The organisation needs to know what should be recorded, monitored and escalated when they do. You need to make sure it's the right human, in the right loop.
Without that, human oversight is just the appearance of control.
Start with the decision
I would start in the same place I start with most questions about AI: what problem are we trying to solve?
The right oversight model depends on what the system is doing and, more importantly, what its output can influence. A tool drafting routine internal text needs different controls from one supporting a decision about a person’s care.
The intended purpose matters. So do the potential harm, the level of autonomy, the quality of the evidence available and whether an outcome can be reversed.
Before deciding where the human needs to be in the loop, I would ask:
- What decision or action can the AI influence?
- Who is accountable for that decision?
- What does the person need to know to challenge the output properly?
- What can they actually do if they believe it is wrong?
- What evidence will show that the control is working?
That last question is important. It is very easy to design an oversight process that works on paper. It is much harder to know whether people can use it consistently and effectively under the conditions in which they work.
Then you can balance the severity of the potential harm against the likelihood of it occurring, and decide what level of ‘accuracy’ the process actually requires. Where the task involves subjective interpretation, ‘wrong’ is a much softer definition. The focus is less on prescribing a single correct outcome and more on defining the guardrails.
Oversight begins before deployment
HIQA’s July 2026 guidance is useful because it treats responsible AI as a lifecycle.
Its good-practice examples for planning and procurement cover evidence, risk, rights, data, testing, regulatory requirements, interoperability and the potential burden on staff.
All of that is human oversight.
By the time an AI output appears in front of a member of staff, many of the most important decisions have already been made. People have decided whether the tool is suitable for the service, which populations it can be used with, where it should sit in the workflow and which safeguards are required.
Those decisions need input from procurement, technical, governance and clinical or care teams. Not everyone needs the same level of technical knowledge, but they do need enough shared understanding to challenge the claims being made about the system.
If each group sees only its own part of the problem, gaps will survive into deployment.
Make challenge possible
Once the tool is in use, the person overseeing it needs clear boundaries.
They should understand what the system is intended to do, what it is not intended to do and what a reliable output looks like. They need to know its limitations, when independent verification is required and where responsibility sits when the output influences a decision.
They also need the expertise to recognise more than something that looks superficially good.
This is where an unstructured approach to AI use creates risk. People are best placed to challenge AI when it is being used within the work they already understand. If someone does not know what a good outcome looks like in that context, they are not equipped to validate the accuracy or quality of what the system has produced.
The workflow then needs to give them a meaningful way to act.
Can they override the output? Can they pause the process? Is the escalation route clear? Does the interface provide the information they need to question the recommendation? Do they have enough time to do it properly?
Putting a confirmation button beneath an AI recommendation is not the same as giving someone control over it.
Training cannot compensate for a process that makes the wrong action easier. Equally, a well-designed process will fail if people have been taught only how to operate the tool and not how to recognise when it is going wrong.
Recognising when AI is going wrong is a skill in itself. Traditional review relied, at least in part, on intuitive warning signs: grammatical errors, poor structure or unpolished design. Those signals told us an output warranted closer scrutiny. Generative AI can produce something fluent, structured and convincing in seconds. The appearance of accuracy is now cheap to produce.
Its errors also take unfamiliar forms. An LLM may follow the wording of an instruction while missing its intent, introduce a plausible but unsupported detail, or drift from an agreed process without producing anything that looks obviously wrong. These are not necessarily the kinds of mistakes a workforce is accustomed to finding. People need to be taught how AI failure presents rather than being reminded that it can fail.
Monitor how people use the system
Technical monitoring typically looks at accuracy, drift, latency and failure rates. Those measures are necessary, but they do not tell us whether human oversight is working.
We also need to understand how people interact with the system.
Are outputs being overridden? Are the same concerns appearing across different teams? Do particular groups experience different outcomes? Are staff bypassing the approved workflow because it creates too much friction? Are people becoming less likely to question recommendations over time?
None of those signals gives us an answer on its own.
A high override rate could mean the system is performing poorly, or that it is being used outside its intended purpose. A very low override rate could mean it is working well, or that people have become too willing to accept its recommendations.
The point is to connect technical performance with operational reality. That is how we identify automation bias, problems with the user experience, gaps in capability or a mismatch between the process that was designed and the work people are actually doing.
Incidents should change the system
HIQA’s guidance says AI-related incidents should be recorded and actively reviewed to support learning and quality improvement. It also calls for mechanisms to respond to and learn from feedback, concerns and complaints.
That learning needs to result in a change.
If an incident shows that staff misunderstood a limitation, the answer might include clearer guidance or different training. But it may also require a workflow change, an interface change or an additional control.
If the answer to every incident is more training, there is a risk that the organisation is treating a system problem as a user problem.
The same applies when monitoring identifies a new risk. Relevant users need to know what has changed, why it matters and how the way they work with the system needs to change as a result.
This is why AI literacy cannot be a one-off course. The capability of the workforce needs to develop as the system, the evidence and the way the tool is being used develop.
Plan how to stop
The least discussed part of human oversight is withdrawal.
What happens if performance deteriorates? What if the supplier changes the model, a critical dependency fails or the service decides the tool is no longer appropriate?
A safe deployment needs a named owner, clear thresholds for suspension and a route back to the previous process. The data involved must remain protected and continuity of care must be maintained.
Staff also need to retain enough capability to deliver the service without becoming entirely dependent on the tool.
If an organisation cannot explain how it would stop using an AI system safely, it has not finished planning how to deploy it.
In fact, it’s highly likely that if organisations do need to pull back on AI deployments in ~6 months' time, they may find themselves without the resource or tacit knowledge to be able to manually complete the same work.
Human oversight is not one checkpoint and it is not one role. It is the combination of governance, workflow, technical monitoring, domain expertise, workforce capability, organisational learning and exit planning across the full lifecycle.
Humans are part of that system. They are not the system.
Unlock the power of your data & AI
Speak with us to learn how you can embed org-wide data & AI fluency today.



.png)
.png)


.png)
.png)



.png)
.jpg)
.png)
.png)
.png)
.png)