The Four Changes Boards Have Not Made
In March an engineer at Meta put a colleague’s technical question to an internal AI agent. The agent posted its answer straight to the forum, without the engineer authorising it. The answer was wrong. Someone acted on it, widened a set of access controls, and for roughly two hours company and user data sat readable by people with no clearance to see it. Meta classified it Sev-1 and found no evidence anyone had exploited it.
Models are wrong all the time. The instructive part here is that the agent could publish an unauthorised answer into a channel people act on, and nobody had ever decided it could. The reasoning error only mattered because that gap was sitting there waiting for it.
The question that follows an incident of that kind is the one a board finds hardest to answer. Whose job here was it to have seen it coming?
I wrote in May that the action which used to require a human at the moment of decision now requires a human at the moment of design. That is close to common ground now, and boards have begun to act on it: reviewing Fortune 100 companies, EY found the share disclosing that a board committee had been charged with AI oversight rose from about 11% in 2024 to roughly 40% a year later.
Almost all of that movement is agenda-level. The institution itself is largely unchanged, and directors can see the gap in their own practice: in the 2026 “What Directors Think” survey from Diligent Institute and Corporate Board Member, 66% said they use AI for board work, while only 22% reported having governance processes in place to guide that use. The people appointed to oversee adoption are ahead of the structures meant to hold it, on their own devices, in their own boardroom.
Four things about the board itself have to change, and a board that makes three of them has made none.
Who gets recruited
The reflex is to appoint one AI-literate director and treat the competence question as settled. That is the wrong size of response, and it reaches for the wrong competence.
A director who can describe how a model is trained is not thereby able to say whether a deployment puts authority in the right place. What a board needs is judgement about where authority sits, what triggers an escalation, and where a human is genuinely required rather than merely listed on a diagram as present. That is a design competence, different in kind from a subject-matter seat.
One technologist can leave a board feeling covered while it remains unable, as a body, to test how a system is built. The feeling of coverage is the risk. It is the governance analogue of an audit committee that takes a clean control report every quarter and never asks whether the control was tested.
In practice this is a line on the skills matrix, named alongside cyber and audit. It changes the nominating interview: ask a candidate how they would satisfy themselves that a deployed agent cannot exceed its intended authority, and where they would expect the escalation triggers to sit. The answer separates a director who reasons about control points from one who reaches for the vendor’s assurances.
What directors are evaluated on
Boards refresh their composition constantly and evaluate against this axis almost never.
A director used to be assessed, formally or in the nominating committee’s private judgement, on the ability to probe a decision and read management. That assumed the decision was the unit under review and could be interrogated afterwards. Continuous execution weakens the assumption: decisions are made at machine speed by systems no director will inspect one output at a time.
The Bank of England’s deputy governor for financial stability, Sarah Breeden, put the constraint plainly at the ECB’s Sintra forum in June. Existing frameworks, she said, were not built to contemplate autonomous agents, and relying on a human in the loop for all agent actions is unlikely to be realistic. She was addressing supervisory frameworks for market stability rather than board practice, but the constraint carries. If reviewing every output does not scale, the evaluable capacity is the ability to oversee the system producing them.
This has to reach the instrument. A director self-assessment can carry one direct question: over the past year, did this director engage the design of at least one deployed system, how it is authorised, where it escalates, how it is retired, or only its outputs and its business case? A re-nomination memo can be required to answer in specifics. A director who cannot engage the architecture gets called “light on tech”, a phrase that lets everyone relax. That director is light on the core oversight task, the way a director who could not read a balance sheet would once have been.
Re-nomination is where this becomes real or lapses. A board can rewrite its skills matrix, run one honest evaluation cycle against the new axis, and then re-nominate the same slate out of collegiality. The nominating and governance committee is where enforcement has to happen, and it is not a committee accustomed to enforcing much.
What the charter obliges
Assigning oversight and chartering it are different acts, and the EY number above counts what boards disclose, not what they have obliged a committee to do.
Start from the duty a Delaware board already carries. Under the Caremark line, through Marchand, directors must implement and monitor board-level reporting for mission-critical risks, and compliance monitoring inside management is not a substitute for oversight at board level. Writing on the CLS Blue Sky Blog in March, Pierluigi Matera maps that onto the present case: “Courts are not asked to evaluate the quality of an algorithm or to second-guess model architecture. They are asked whether directors made a sustained, good-faith effort to understand the role the system played within the firm’s process for monitoring risk, whether validation and periodic review mechanisms existed, and whether escalation pathways were meaningfully designed.” His conclusion is that “AI does not heighten the standard. It sharpens the clarity of process”, and that “The inquiry is procedural rather than technological.”
No Delaware court has yet sustained a Caremark claim based on AI oversight. The duty is settled; its application here is argued by analogy and not yet decided. That is worth saying directly, because it is the honest description of where the law stands.
What it points to is a charter that runs the whole life of a system rather than stopping at launch approval. Gartner expects that by 2027 some 40% of enterprises will demote or decommission autonomous agents over governance gaps found only after a production incident. Decommissioning belongs in the charter. On the mission-critical logic the duty reaches from how an agent is authorised to how it is retired, and a charter that goes quiet after approval leaves the largest part of the risk unowned.
What management is paid to protect
Boards have spent a decade learning to put risk into pay. Clawbacks, risk-adjusted metrics, deferral, the whole apparatus that followed the last financial crisis. Almost none of it prices the integrity of the decision architecture.
Where incentive plans reward deployment velocity and automation-driven cost reduction, they reward shipping the agent and booking the saving, and they price nothing for whether the controls around it hold. If a board pays management to move faster than the controls the board has just made itself responsible for, it should expect exactly that. A compensation committee that has never asked whether its plan rewards decision-architecture integrity has priced that integrity at zero by default.
The lever is familiar. The portion of an incentive tied to AI-driven efficiency, the cost an agent took out or the speed it added, can be made contingent on an attestation that the controls around that agent were tested and held through the period. With that link in place the plan stops paying management to outrun the board’s own obligation.
The four are one
Each of these is hollow without the others, which is why partial adoption feels like progress and delivers none.
Recruit for design competence but never re-nominate on it and the board drifts back to the comfortable slate. Rewrite the evaluation but leave the charter silent and directors are judged against a task no committee is bound to perform. Charter the full lifecycle but pay management to sprint past the controls and the obligation the charter creates is one the board is funding people to ignore.
Return to the question an incident leaves behind. It has an answer only on a board that made all four changes: a committee chartered to oversee a system’s design across its whole life, staffed by directors recruited and re-nominated for the capacity to do it, whose management is paid to protect the architecture the company depends on. On a board missing any one of them the honest answer is still nobody, and the next incident will ask again.
That redesign is the board’s own, and it is the board’s to make before an incident makes the case for it.
The models underlying this argument, on how decision authority is structured and assessed, are set out on the Frameworks page. On the agents nobody has inventoried, see Shadow Agents; on oversight that records a human presence without providing one, The Timestamp Problem.