In 2021, the Dutch government resigned after a parliamentary inquiry into the country’s childcare benefits affair found that thousands of parents had been wrongly accused of fraud and treated unfairly by the tax authorities. Many were forced to repay large sums of money, sometimes over relatively minor administrative errors, with serious consequences for affected families. Automated risk-profiling tools had been used as part of the process, with characteristics including nationality influencing how applications were assessed, something Amnesty International documented in its review of the Dutch childcare benefits scandal. This was not simply a case of technology going wrong. The problem was much wider, involving the government’s approach to suspected fraud, the way the system was used and the failure of the safeguards around it.
The resignation of the government was significant because political responsibility was ultimately accepted for what had happened. In a corporate environment, responsibility can be more difficult to pin down. Management may blame the technology provider, the provider may argue that the system was not used as intended, those responsible for testing may be criticised for not identifying the problem, while the board may say it relied on assurances from management and external experts. When AI gets it wrong, who is actually accountable?
AI is already doing work that would previously have required human judgement, and in some organisations it is also influencing decisions that would once have been left to people. When a system performs well most of the time, people will naturally begin to trust it and, over time, may stop questioning its output as closely as they once did. Yet a system being right most of the time does not make the occasions when it is wrong unimportant, particularly where its decisions can affect people’s jobs, finances or access to services.
Board Oversight of AI
A core part of the board’s role is approving the strategic direction of the company, monitoring financial performance, approving major transactions and overseeing risk. Businesses are now investing significant amounts of money in AI systems to automate processes and improve efficiency, with some of these systems influencing decisions affecting employees, customers and the business itself. It becomes difficult to treat such systems as purely technical or operational matters when the consequences of their use may create significant risks for the organisation.
For significant AI systems, it is not enough for a board simply to accept management’s assurance that the system has been properly developed and tested. Directors should understand what problem the system is supposed to solve, what data it relies on, how extensively it has been tested, what safeguards are in place and how its performance will be measured after deployment. What happens when it gets something wrong? Could it unfairly disadvantage certain groups? Has there been an independent security or technical assessment? The fact that considerable time and money have already gone into developing or acquiring the system should not make those questions harder to ask.
None of this requires directors to understand the algorithm behind the system or know how to write the code. They need to understand what the system is supposed to do, the risks that come with its use and what could happen if it fails. Where the technical issues go beyond the board’s expertise, external advice can help. Boards already rely on external auditors for independent assurance over financial statements, and for an AI system carrying significant risk, there may be a case for independent technical or security assurance.
If the board is not satisfied with the answers it receives, it can delay approval or decide not to proceed with the system at all. And approval does not mean the system should be left alone afterwards. A tool that performs well during testing may behave differently once it is used with real people, changing data and actual business processes. An internal drafting tool, for example, does not require the same level of scrutiny as a system used in recruitment, credit or fraud decisions.
Management’s Responsibility
Management is much closer to the development, testing and everyday use of these systems and will often know first when something is not working as expected. They must test properly and disclose problems honestly, including problems that may delay a project in which considerable time and money have already been invested. If serious problems emerge, management should be prepared to change or suspend the system. In some cases, it may have to abandon it altogether.
Amazon faced something similar with an experimental recruitment tool it developed to help review job applicants. The system learned from resumes submitted over several years, but because the historical data reflected a technology workforce that had been heavily male dominated, it began learning patterns that disadvantaged women. Reuters reported that the tool penalised resumes containing terms such as “women’s” and that Amazon eventually abandoned the project despite the resources already spent developing it.
When an AI system later causes harm, however, it does not automatically follow that the board has failed in its oversight. Systems can fail even where reasonable questions have been asked, proper testing has been carried out, and management has disclosed the problems it knows about. The board may have done everything reasonably expected of it and the system still failed. As I have argued before in discussing whether board approval really means directors were properly informed, the information directors had at the time, and how they responded to it, becomes important in deciding whether there was a failure of oversight. Different people will have had different responsibilities for the system, from its development and testing to its use and oversight. The accountability question is whether they did what was reasonably expected of them.
If the board simply accepted assurances from management, consultants or developers, especially where there were already warning signs, it becomes much harder to say that it exercised proper oversight. That does not mean responsibility will always rest with the board. A vendor may have overstated what the system could do, failed to disclose known weaknesses or failed to meet agreed testing or security requirements.
Management still has work to do after deployment. A system may have been properly approved and tested and still end up being used for purposes it was never designed for. Employees may start relying on its recommendations too heavily, or the quality of its output may change as the underlying data changes. And where the system is being used for decisions that directly affect people, even the relatively small number of cases in which it gets things wrong may have serious consequences.
Human Oversight
The Dutch childcare affair is worth returning to here because humans were involved in the process. It was not simply a case of an automated system producing an output that nobody ever looked at. Human officials were involved, but that did not necessarily provide an effective safeguard. A person expected to review an output cannot do much with that responsibility if they do not understand why someone has been flagged, do not have enough information to question the result or have become so accustomed to the system being right that they assume it must be right again.
That kind of human review becomes even more difficult with so-called black-box systems, where it may be possible to see the information going in and the recommendation or decision coming out without being able to explain clearly how the result was reached. If a human reviewer cannot understand how the result was reached, it is difficult to see what meaningful challenge really looks like. For the board, the issue is whether the organisation understands the system well enough to know its limitations, recognise unusual outcomes and respond when something appears to have gone wrong.
AI systems will continue to find their way into ordinary business decisions. Some of those systems will fail even after serious attempts have been made to test and control them. When that happens, the fact that the AI failed tells only part of the story.
AI can make accountability harder to trace because several people may be involved in developing, approving and using the system. A vendor may have built it, management may operate it and the board may never see the underlying code. None of that means responsibility disappears. When something goes wrong, attention still has to turn to the people around the system: what they knew, what they were expected to oversee and how they responded when problems emerged.
