An AI Board Reporting Protocol for Decisions That Matter

By Mark | leadership | 7 min read

An AI report can be accurate and useless to a board. A list of pilots may describe activity without revealing which decisions the organization has delegated, what could go wrong, and when directors will act.…

An AI report can be accurate and useless to a board. A list of pilots may describe activity without revealing which decisions the organization has delegated, what could go wrong, and when directors will act. Conversely, approval of every model change can slow useful work while drawing directors into operational judgments.

The question is not how often directors should discuss AI. It is what information they need to oversee consequential use, which decisions management should make, and what events should change the level of scrutiny. A reporting protocol should make those boundaries explicit before an incident or a contested investment forces an improvised decision.

This is an approach to business judgment and oversight, not a statement of legal duties or legal advice. The right boundaries will differ with the organization, the decisions affected, its operating capacity, and applicable obligations assessed with counsel.

Start With Decisions, Not Tools

Management should maintain an inventory of material AI uses, including capabilities embedded in purchased software. The useful unit of analysis is not a vendor or a model name. It is a decision or workflow: who uses the output, what action may follow, and what happens if the output is wrong. A customer service assistant that drafts a response for review is different from a system that sends a binding communication without review, even if both use the same underlying model.

Each inventory entry should identify an accountable business owner, the decision affected, the population exposed, the source of the data, the degree of automation, the vendor dependencies, and the means of detecting and correcting failure. Management can then group uses by potential consequence rather than treating every application as equally important.

A simple three level scheme may suffice. Lower consequence uses affect internal drafting or analysis and remain subject to ordinary review. Material uses affect customers, employees, financial reporting, or a significant operating process, but have defined human controls and manageable reversibility. Critical uses could cause serious harm, a difficult to reverse decision, or a major interruption if they fail. Classification should consider the severity and reach of a plausible error, the ease of correction, and how much independent judgment remains with a person. It should not depend solely on whether the product is marketed as AI.

These are management classifications, not regulatory labels. A small organization may need only a short register. The important discipline is explaining why each use receives the attention it does.

Give Management Authority and Preserve Board Judgment

Management owns selection, testing, deployment, monitoring, and remediation. Business leaders should decide whether a use meets its intended purpose. Technology and risk teams should test reliability, data handling, access, and monitoring appropriate to the consequence of failure. Counsel should advise on the obligations that actually apply. An identified executive should resolve disagreements among these functions and own the resulting decision.

The board should oversee the allocation of authority and the effectiveness of that process. It can approve the organization's risk boundaries, receive reporting on the most consequential uses, challenge whether material benefits justify residual risks, and decide whether a proposed exposure falls outside an agreed boundary. Directors should not choose model settings, certify every output, or substitute a quarterly vote for continuous management control.

For a critical use, management might require documented review before deployment, named approval authority, a test of plausible failure scenarios, a way to stop or limit the use, and scheduled reassessment. For a lower consequence use, ordinary procurement and manager review may be proportionate. The board needs to see that this difference exists and is applied consistently, not to sign every approval.

Delegating everything to a technical team can leave no owner for customer or employee consequences. Escalating every choice to directors produces delay. An authority map should say which executive may accept a risk, which committee receives notice, and which questions require full board consideration.

Build a Board Packet Around Exceptions and Evidence

A board packet should show changes since the prior report, the highest consequence uses, incidents and near misses, control exceptions, and requests for direction. Each material entry should say what management knows, what remains uncertain, who owns the next action, and when directors will receive an update.

Three kinds of evidence deserve particular attention. First, provenance: where relevant inputs and underlying services come from, what rights and limitations attach to data and vendor use, and which assumptions depend on a third party. A complete history of every training example may not be available. The board should instead understand material gaps and how management has limited reliance on unverified claims.

Second, performance: the benchmark before launch, operational measures, and differences across important settings or affected groups. A favorable average can conceal a serious failure. Management should describe sample limitations.

Third, change: whether inputs, behavior, business context, or a vendor's service have shifted enough to call the original approval into question. Drift is not proof of failure. It is a prompt to investigate whether a previously acceptable control still works. Reports should separate a measured deterioration from a suspected one and should explain what the organization will do in either case.

A Hypothetical Reporting Decision

Consider a hypothetical company using a vendor model to suggest which customer complaints require urgent review. Employees make the final response, but the model determines which cases they see first. Management classifies the use as material because delays may disadvantage customers, while a trained employee can still correct the priority before action.

Before deployment, the business owner records the decision flow and an acceptable delay range. The technology team tests the model against a relevant set of past complaints. The procurement and legal teams review what customer information reaches the vendor and how the vendor may use it. Management sets a weekly check for missed urgent cases and retains a manual queue for use if the system is paused.

The next board packet reports the number of complaints handled, how many urgent cases were identified late, whether manual review found an unexpected pattern, and whether a vendor update changed model behavior. It identifies the management owner and next review date. These are hypothetical fields, not actual results.

Suppose the review finds a repeated increase in late identification among complaints submitted in a particular language. Management first checks data quality and the review process. If the cause is unclear and the delay could materially affect customers, the accountable executive temporarily routes that group through the manual queue and notifies the designated board committee. Directors do not debug the model. They ask whether the temporary control is workable, whether affected customers require a response, whether the exposure extends beyond this workflow, and what evidence will support resumption.

This example also shows a tradeoff. Immediate suspension of all automated triage might protect against one error but create a backlog that delays every complaint. Continuing unchanged might preserve speed at the cost of uneven service. A targeted pause with a manual route may be more proportionate, provided management has the capacity to run it and can measure its effect.

Set Triggers Before Pressure Arrives

The board and management should agree on categories of triggers, while leaving operational numbers to be calibrated against real capacity and consequences. Routine variation within a documented range may remain with the business owner. A sustained change outside that range, an unexplained outcome affecting a sensitive decision, loss of a key data right, or a significant vendor change may require executive review and committee notice. An event that could cause serious harm, a widespread inability to correct decisions, or a failure of the fallback process may justify an immediate pause and prompt board notification.

These are illustrative thresholds, not universal rules. A numerical trigger without a sensible denominator, baseline, or response owner creates false precision. Equally, a trigger stated only as “material concern” can leave people debating vocabulary while exposure continues. The protocol should pair observable signals with an owner, a time for reassessment, an interim control, and a record of who decided to continue, narrow, or stop use.

A pause should have a path out. Management should document the defect or uncertainty, the corrective action, the evidence needed for a limited restart, and who authorizes it. The board should receive a concise account of decisions that crossed its agreed boundary. This produces a record of judgment rather than a stack of assurances.

No packet eliminates uncertainty. Effective oversight makes limitations visible, concentrates effort where consequences are greatest, and revises boundaries when experience warrants it.

Executive Imperatives

Executive Imperative: Ask management for a decision based inventory of material AI uses. For each, require the decision affected, accountable owner, level of automation, plausible consequence, available fallback, and reason for its assigned risk level. Keep the board view focused on consequential uses and exceptions.

Executive Imperative: Agree on written decision rights. Specify who approves deployment and changes, who can accept residual risk, when a committee receives notice, and what moves to the full board. Reserve operational testing and remediation for management while retaining board scrutiny of exposures outside agreed boundaries.

Executive Imperative: Require evidence rather than confidence statements. For material uses, review the limits of data and vendor provenance, the basis for performance claims, signals of change after deployment, unresolved control gaps, and the next date for reassessment.

Executive Imperative: Rehearse one plausible pause and restart. Choose a consequential workflow, test the notice path and fallback capacity, and record what would justify narrowing, suspending, and resuming it. Update the protocol if the exercise reveals that an owner, threshold, or interim control is unclear.