A number on its own tells you nothing

Teams often lead with an accuracy figure when describing an AI system. Eighty-seven percent accurate sounds like a strong result. It says nothing about who is affected when the system is wrong, how often those errors cluster in one group of people, or what happens next when an error occurs.

Naming the type of harm first

AI harm is not one thing. It can be individual, affecting one person's outcome. It can be a group harm, where an error rate is higher for one demographic than another. It can be organisational, through reputational or regulatory exposure. It can be societal, where a pattern of automated decisions shifts access to opportunity at scale. Naming which type of harm is in play changes what evidence you need to collect and who needs to sign off.

Fairness is not the same as accuracy

A model can be accurate overall while performing worse for a specific group. Testing has to look at accuracy by group, not only in aggregate. This is why bias testing and fairness testing sit alongside accuracy testing as separate, necessary steps rather than nice-to-haves.

Human oversight is a control, not a formality

Automation bias, the tendency to trust a system's output without genuinely reviewing it, is a real risk to human oversight controls. A human reviewer who rubber-stamps every recommendation is not providing meaningful oversight, even if a person is technically in the loop.

A simple test before launch

Before any AI system goes live, it is worth mapping who could be harmed, what type of harm is realistic, how probable and how severe it is, and what the mitigation hierarchy looks like: avoid, reduce, transfer, accept or monitor. That mapping should happen before the accuracy number becomes the headline.

We walk through this harm and stakeholder mapping process in detail in the AI Governance, Risk & Compliance Practitioner programme.