Redefining Technology

Energy & UtilitiesLeadership Insights & Strategy

AI leadership playbooks for energy and utilities: the decision guides a leadership team runs on

AI leadership playbooks are the standing decision guides a utility leadership team uses so that recurring AI decisions — how to fund one, whether to build or buy, what to do when a model is wrong during a live network event, which vendor, which workforce change, when to stop — are settled by a written route rather than re-argued from scratch.

Illustration of a utility executive team around a boardroom table reviewing an AI decision pack, with an offshore wind farm visible through the window
Energy & Utilities · Leadership Insights & Strategy

Key takeaways

  1. An unexercised playbook is a document, not a capability. A utility already knows this about black-start and storm response; the same rule applies to AI decisions, and the test is whether leadership has ever rehearsed one against an invented scenario before a real event forced it.
  2. Six decisions recur often enough in a utility to deserve a standing guide: funding, build versus buy, model incident, vendor selection, workforce change and stopping a use case. Everything else can be handled case by case; these six cannot, because they arrive under time pressure and with money or reliability attached.
  3. A playbook is not a policy. A policy states what is not allowed; a playbook states how a specific decision gets made — its trigger, its decision rights, the inputs the sponsor must bring, the gate criteria and the evidence it leaves behind.
  4. Decision rights fail on the accountable column, not the responsible one. Every playbook needs exactly one named accountable executive; committees that decide collectively produce decisions nobody can reconstruct eighteen months later when a regulator asks.
  5. The stop playbook is the one most often missing and the one that most changes behaviour. Where a stop condition is agreed before a use case is funded, retiring it costs a sponsor nothing; where it is not, use cases are defended long past the evidence and the next proposal is funded against that memory.

Abbreviations used on this page

RACI
Responsible, accountable, consulted, informed — the decision-rights notation used in every playbook on this page
SCADA
Supervisory control and data acquisition
EMS
Energy management system (the control-centre application suite)
ADMS
Advanced distribution management system
OMS
Outage management system
CMMS
Computerised maintenance management system (the work-order system of record)
AMI
Advanced metering infrastructure (smart meters and their head end)
DER
Distributed energy resources — rooftop solar, batteries, heat pumps, EV chargers
SAIDI
System average interruption duration index — the headline reliability KPI
ETR
Estimated time of restoration (the figure given to a customer during an outage)
TOTEX
Total expenditure — the combined capex and opex category regulators increasingly assess as one
RIIO
Revenue = Incentives + Innovation + Outputs — Ofgem's price-control framework for GB network companies

Free · 8 questions · ~3 minutes

Score your playbook library

Eight questions, one at a time, about three minutes. Answer them and we build your personalised playbook report — your stage on the ladder, your score on each of the four dimensions, and the specific gap standing between you and the next stage — and send it to your inbox. Your result doubles as the coverage map for your first drafting session.

0 of 8 answered

Question 1 of 8Playbook coverage

How many of the six recurring AI decisions — funding, build or buy, model incident, vendor, workforce change and stopping a use case — have a written decision guide?

Coverage is the cheapest thing to fix and the strongest predictor of how much senior time an AI programme consumes per decision.

How the score maps to a stage
  • 04 — Stage 1, No playbook. No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
  • 59 — Stage 2, Drafted. Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
  • 1014 — Stage 3, Adopted. Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
  • 1520 — Stage 4, Exercised. Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
  • 2124 — Stage 5, Institutional. Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.

What AI leadership playbooks are — and what they are not

A definition, the five parts every playbook needs, and the difference between a decision route and an AI policy.

AI leadership playbooks are standing decision guides: one page per recurring decision, stating what fires it, who decides, what the sponsor must bring, what the gate tests and what evidence the decision leaves behind. They exist so that a utility leadership team stops re-deriving the same positions — on funding, on build versus buy, on what to do when a model is wrong at 02:00 — every time a new use case reaches the table.

The distinction that matters most is between a playbook and a policy. A policy is a boundary: it states what is not allowed, who must be consulted, which data may leave the estate. It is necessary, and it decides nothing. A playbook is a route: it takes a specific, recurring question and describes how it gets answered, by whom, against what criteria, in what order. Utilities that write the policy first typically spend a year producing approvals without producing decisions, because the policy generates gates without generating routes through them.

The form is not novel and it is not ours. The NIST AI Risk Management Framework (opens in a new tab) ships with a companion Playbook of suggested actions (opens in a new tab) organised around govern, map, measure and manage, and ISO/IEC 42001 (opens in a new tab) defines an AI management system built on documented decisions, assigned roles and continual improvement. Both describe the shape of the artefact. What neither can supply is the part that makes a playbook work in a network business: the specific trigger in your estate, the specific executive who holds the decision, and the gate criteria expressed in SAIDI minutes, ETR accuracy or TOTEX rather than in general risk language.

How an AI decision reaches a conclusion, with and without a playbook

The two routes a recurring AI question can take through a utility. The top lane is not disorganised — it is undocumented, which is a different and more expensive problem. The middle lane is the playbook route; the bottom lane is the loop that keeps it accurate. Most leadership teams are in the top lane.

  • Human in the loop
  • Where value leaks
  • Data & feeds
  • System-of-record action

The process, in words

  • Without a playbook, a question arrives from the board, a regulator or a sponsor, gets a paper written in whatever format the author prefers, and reopens the same first principles every time — is this opex or capex, who signs it, what happens if it does not work. The decision is settled by whoever is in the room that week and nothing durable is filed, so the next question starts from the same place.
  • On the playbook route, a defined trigger fires rather than someone requesting a slot. The sponsor assembles the named input pack and nothing beyond it, decision rights route the paper to exactly one accountable executive with a written consulted list, the gate scores it against criteria agreed before the case existed, and the outcome — go, hold or stop — is filed with the score, any dissent and the playbook version in force.
  • The learning loop is what stops the route becoming fiction. Scheduled exercises and real events both feed back into a dated, owned revision of the playbook, and the next decision of that type runs against the new version. Without this loop a library ages faster than the estate changes: names leave, systems upgrade, and the written route quietly stops describing reality.
Step-by-step insights
The trigger is the part most libraries omit
Almost every drafted playbook describes a process and forgets to say what starts it. That omission is why unused libraries stay unused: with no trigger, using the playbook is a voluntary act by a busy executive, and voluntary acts lose to time pressure. A trigger is a specific, observable event — a use case asking for money beyond discovery, an AI clause appearing in a renewal, a model output contradicting the control room during an abnormal state, two consecutive review cycles below threshold. Written that way, the playbook starts itself, and the question 'should we run this?' never has to be asked.
Why the input pack must be closed, not open
Sponsors will bring whatever they think helps, which under uncertainty means everything. A closed input pack — these six items, nothing else scored — does two useful things at once. It caps preparation cost, so a small use case is not priced out of the gate by paperwork; and it makes incompleteness visible, so a chair can return a paper in two lines rather than debating an eighty-slide deck. The discipline compounds: after one paper is returned for a missing stop condition, every subsequent paper has one.
One accountable name, and what the consulted list is for
The accountable column carries a single name because reconstruction depends on it: eighteen months later, 'the steering committee agreed' is not an answer a regulator or an inquiry can work with. The consulted list does a different job — it is the pre-agreed set of people whose absence invalidates the decision, typically the control-room authority for anything touching real-time operations, the safety function for anything touching field work, and employee representatives for anything that changes a role's task set. Naming them in advance prevents the most common gate failure, which is a technically sound decision taken without the one person who could have said why it would not work on the ground.
The gate scores criteria that existed before the case did
A gate that invents its criteria while reading the paper is a debate, not a gate. The criteria belong in the playbook, agreed when nobody's proposal is on the table and therefore nobody's interests are engaged — which is the only time an organisation can set an honest threshold. This is also why the stop condition must be written before the first score: a threshold agreed in advance is a fact, and the same threshold proposed after a use case exists is an accusation.
The decision record is the by-product that becomes the asset
Filing the paper, the score, any dissent and the playbook version costs a few minutes per decision and is the difference between a library that can be audited and one that merely exists. Its value shows up in three places a utility will recognise: a price-control or rate-case submission that needs to show how investment decisions were governed, an incident review that needs to show what was known when, and a leadership change where a successor can read twenty decisions instead of interviewing twenty people.
The loop is why exercises outrank approvals
Approval tests whether a document is acceptable; an exercise tests whether it is true. The bottom lane exists because the estate keeps moving — an ADMS release relocates a fallback, a reorganisation vacates an accountable name, a vendor contract changes who holds the model — and none of those changes announces itself to the library. A scheduled exercise, plus a replay after every real event, is the only mechanism that reliably finds the drift before an incident does.

The five stages in detail

For each stage: what it looks like inside a real utility, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps leadership teams there, and what leaving costs.

The ladder runs from No playbook to Institutional, and its shape is worth noticing before the detail: the value released is close to flat across the first two rungs and inflects at Adopted, when live decisions actually start passing through a gate. A library that exists and is never used releases almost exactly as much value as no library at all, which is why the count of approved documents is the least informative metric an AI programme can report.

Decision throughput released along the playbook ladder

Value stays flat while playbooks are documents — most leadership teams sit on that flat section — and inflects at Adopted, when a real decision first passes through a gate. The second inflection is Exercised: rehearsal is what converts a route that works on paper into one that holds under a live network event. Drawn from the ladder on this page, not from a measured dataset.

Leadership decision throughput by stage

  • Stage 1 · No playbook — 24% of operators. No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
  • Stage 2 · Drafted — 31% of operators. Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
  • Stage 3 · Adopted — 27% of operators. Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
  • Stage 4 · Exercised — 13% of operators. Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
  • Stage 5 · Institutional — 5% of operators. Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the IEA's assessment that sector-wide AI adoption is not a given.

Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions inside a network or generation business, the diagnostic signals are checks you can run against your own governance records this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

No playbook

24% of operators sit here

No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.

Stage 1 is not the absence of AI activity — most utilities at this stage have several models running and at least one that works. It is the absence of any repeatable route by which a decision about that activity gets made. Each question arrives fresh, gets a bespoke paper, and is settled by the balance of seniority in the room on the day. The decision may well be correct; it is simply not reproducible, and nothing about it makes the next one cheaper.

The tell is the re-argument. Ask how the last three AI funding decisions were made and you will hear three different processes described by three different sponsors, each convinced theirs was the normal one. The same underlying questions — is this opex or capex, who signs it off, what happens if it does not work — are re-opened every time, because none of the answers was written anywhere a successor could find. Senior time is spent re-deriving positions the organisation has already held.

This stage is cheap to leave and expensive to occupy, and the cost is asymmetric. A good decision made this way earns no institutional credit, because it cannot be pointed at; a bad one is catastrophic, because there is no record showing what was considered. In an industry where a regulator, a system operator or a public inquiry may reasonably ask how a decision was reached, an undocumented route is a liability that grows quietly with every use case added.

In practice

The vegetation model funded three times

A distribution business built an imagery model that ranks spans by vegetation-encroachment risk. It was funded as a discovery piece by the asset director, re-argued four months later as an innovation project when discovery money ran out, and argued a third time as part of the next capital plan — three papers, three formats, three different sets of success criteria. Nothing was wrong with any of them. The cost was six months of executive attention spent on a decision the organisation had already effectively taken twice.

What it looks like

  • AI questions reach the executive committee as one-off papers, each in its author's own format
  • There is no written list of which AI decisions leadership actually owns
  • The funding argument is re-run from scratch for every use case
  • Nobody can produce the reasoning behind an AI decision taken six months ago

Diagnostic signals you can check this week

  • Ask three executives who approves a model that writes into the ADMS. Three different answers means you are here
  • Ask for the paper behind the last AI funding decision. If it takes more than a day to find, there is no decision record
  • Look for the words 'trigger' and 'gate' in any AI governance document. Their absence is the definition of stage 1
  • Ask what would happen at 03:00 if a live model were badly wrong. If the answer starts with a name rather than a route, there is no playbook

Anti-pattern · Writing an AI policy instead

The instinctive fix is a policy — acceptable use, prohibited use, a statement of principles, a sign-off from legal. It is genuinely useful and it solves a different problem. A policy tells people what they may not do; it does not tell an executive committee how to decide whether to fund a DER forecasting platform, or what to do at 02:00 when a model has mis-ranked a feeder during a storm. Utilities that write the policy first typically wait another year before anything decides anything, because the policy generates approvals without generating routes.

What holds you here

There is no written route for any recurring AI decision, so each one consumes senior time re-deriving positions the organisation has already held.

Highest-leverage next move

Write down the six decisions that recur, name one accountable executive for each, and draft the two that hurt most — usually funding and model incident — on a single page each.

Cost of leaving

Effort
4–8 weeks
Team
One executive sponsor, a programme lead, and half a day each from the four executives who actually hold the decisions
Risk
Low — nothing operational changes, but the choice of which six decisions to write sets the library's credibility
To next stage
4–8 weeks

If this is you, the next step is

A two-week engagement: interview the decision holders, name the triggers, draft the first two playbooks.

Map your six recurring AI decisions

Stage 2

Drafted

31% of operators sit here

Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.

Stage 2 is the most common resting place for a utility AI programme and the easiest to mistake for progress. The artefacts are real: a decision framework, a risk taxonomy, sometimes an impressively detailed RACI matrix produced by a consultancy. The library exists on a page. What has not happened is any decision passing through it, which means every assumption in it is untested — including the assumption that the people named are the people who actually decide.

The structural cause is authorship. The playbooks were written by whoever was asked to write them, usually a function with an interest in coherence rather than one holding the decision. So the funding playbook is written by transformation rather than by the finance director who will apply it; the model-incident playbook is written by the data team rather than by the head of control who will be woken. Each document is internally consistent and externally unowned, and the first live decision routes around it because the person deciding never agreed to be bound.

The other tell is drift against reality. A drafted playbook ages badly: it names a head of digital who has left, a governance forum that merged, an ADMS release that shipped. Nobody notices, because nothing is exercising it. By the time a real incident tests the text, roughly a third of it is wrong — and the credibility damage from a playbook that fails in use is worse than having none at all, because it teaches the organisation that the library is decorative.

In practice

The framework that survived the reorganisation on paper only

A vertically integrated utility approved an AI decision framework in the spring: five decision types, an escalation ladder, a named owner per type. In the autumn the digital and data functions merged, two of the five named owners changed and the risk committee's terms of reference were rewritten. The framework document was not touched. When a supplier proposed an AI-bearing asset-health module in December, procurement followed the standard capital route and nobody opened the framework, because the person it named no longer worked there.

What it looks like

  • A set of AI decision guides exists, typically produced by a transformation or risk function
  • The documents were approved at a governance forum and circulated
  • No live decision has yet been taken through a playbook's gate
  • The playbooks describe an organisation chart that has already changed

Diagnostic signals you can check this week

  • Ask for the last three decisions taken through a playbook's gate. If the answer is none, you are at stage 2 whatever the documents say
  • Check every name in the library against the current organisation chart — count how many are wrong
  • Ask who wrote the funding playbook and who applies it. Different people, with no shared session, is the signature
  • Look for a version number and a date on each playbook. Their absence means nobody expects the text to change

Anti-pattern · Adding more playbooks to fix an unused library

When a library is not used, the reflex is to extend it: more decision types, more detail, an appendix of templates. It is the wrong lever. An unused library is not too thin, it is unowned — and doubling its length lowers the odds that any executive reads it under time pressure. The correct move is to take the single playbook that hurts most, sit the person who actually holds that decision down with it, cut it to one page in their words, and route the next real decision through it, even clumsily. One used playbook beats six approved ones.

What holds you here

The playbooks were written by a function that does not hold the decision, so the first live decision routes around them.

Highest-leverage next move

Take the highest-pain playbook, rewrite it with its accountable executive in their own words on one page, and put the next real decision of that type through its gate.

Cost of leaving

Effort
6–10 weeks
Team
Each playbook's accountable executive for two sessions, plus a lead who is allowed to cut text
Risk
Medium — the first live decision through a gate will expose the parts that were written for tidiness rather than use
To next stage
6–10 weeks

If this is you, the next step is

We facilitate the first gate meeting, then rewrite the playbook from what actually happened in it.

Route one real decision through a gate

Stage 3

Adopted

27% of operators sit here

Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.

Stage 3 is the first stage at which the library changes behaviour rather than describing it. The mechanism is mundane and it is the whole trick: a standing slot in the calendar, a defined input pack, and a chair willing to return an incomplete paper. Once a sponsor has had a paper returned for missing the stop condition, the next paper has a stop condition. The playbook stops being a document and becomes a queue discipline.

The character of the executive conversation changes here too. At stage 2 the meeting spends its first twenty minutes agreeing what is being decided and by what standard; at stage 3 that is settled before anyone sits down, so the discussion is about the specific evidence in front of it. Utilities notice this as a throughput effect — more AI decisions taken per quarter, with less senior time each — well before they notice it as a governance effect.

The constraint that emerges is rehearsal, and it emerges asymmetrically. Playbooks with a scheduled trigger — funding rounds, vendor renewals, workforce planning — get exercised naturally by the calendar, so they improve. Playbooks whose trigger is an event nobody schedules — a model badly wrong during a storm, a use case that must be stopped — are never exercised at all, and the first time they run is the first time anyone reads them, at 02:00, under pressure, with the network in an abnormal state.

In practice

The paper that was returned, and the quarter that followed

A network operator's AI gate met monthly. In its third month a well-regarded sponsor brought a DER-forecasting business case with no stop condition and no named consulted list; the chair returned it with two lines of feedback. It was resubmitted a month later, complete, and approved. The visible effect was one month lost. The unseen effect was that all four papers submitted the following quarter arrived complete, and the gate's average decision time fell from two meetings to one.

What it looks like

  • AI funding and vendor decisions arrive at a standing gate, not at whichever forum has space
  • Sponsors submit the playbook's input pack because papers without it are returned
  • Each decision produces a record naming the approver, the criteria and the date
  • The accountable executive for each playbook is a single named person, not a committee

Diagnostic signals you can check this week

  • Look at the last four AI papers. Do they share a structure, or does each reflect its author?
  • Ask whether a paper has ever been returned for an incomplete input pack. Never is a warning sign, not a compliment
  • Check whether the model-incident playbook has ever run. Scheduled playbooks improve; event-triggered ones rot
  • Ask a sponsor what happens if their use case misses its threshold. A confident, specific answer means the gate is real

Anti-pattern · Letting the gate become a reporting meeting

Once a gate is established it attracts attendance, and attendance turns decisions into updates. The agenda fills with progress reports, the decision items slip to the end, and within two quarters the forum is a status meeting with a decision-shaped item nobody has time to argue about. The defence is structural: keep the gate to decisions only, cap it at ninety minutes, and give progress reporting a different meeting with a different chair. A gate that cannot say no in under an hour has stopped being a gate.

What holds you here

Only the calendar-triggered playbooks get used, so the event-triggered ones — incident and stop — are still untested when they first run for real.

Highest-leverage next move

Schedule a desktop exercise of the model-incident playbook against an invented storm scenario, with the people who would actually be woken in the room.

Cost of leaving

Effort
3–6 months
Team
A gate chair with authority to return papers, a secretary who maintains the decision record, and the accountable executives
Risk
Medium — the first refusal is a political event, and how it is handled sets whether the gate has teeth
To next stage
3–6 months

If this is you, the next step is

Terms of reference, input pack templates and the first three gate meetings run with you.

Stand up the gate and its input pack

Stage 4

Exercised

13% of operators sit here

Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.

Stage 4 imports a discipline the industry already has and rarely applies to AI. A utility does not assume its black-start procedure works because it is written; it exercises it, discovers that a phone number is wrong and a substation key is in the wrong cabinet, and updates the procedure. The same reasoning applies exactly to an AI model-incident playbook, and the same class of finding comes out: the fallback that has never been switched, the escalation path to a role that no longer exists, the assumption that the vendor answers out of hours.

The economics of rehearsal are what make it worth executive time. An exercise costs two hours from six people and finds three defects; the same defects found during a real storm cost restoration minutes, customer trust and, in some jurisdictions, a reportable event. Utilities that already run emergency exercises can attach an AI inject to the existing programme rather than creating a new one — a scenario in which the outage-prediction model is confidently wrong during the drill is cheap to add and disproportionately informative.

What separates a real exercise from theatre is whether the text changes. An exercise that ends with everyone agreeing the playbook worked has usually been run without injects, or without anyone empowered to say the escalation was wrong. The output artefact is not a satisfied room; it is a dated diff. Programmes that keep the diff visible — this playbook has changed four times, here is why — build a kind of institutional confidence that no amount of approval can manufacture.

In practice

The drill that found the fallback nobody had switched

A DSO ran a two-hour desktop exercise on its outage-prediction model: the inject was a wind event where the model badly under-predicted faults on one 33 kV group. The control-room lead reached for the documented fallback — reverting the ADMS field to the rule-based estimate — and discovered the revert required a change ticket that took four hours in normal service. Nobody had ever switched it. The playbook changed that week: the revert became a control-room-authority action with a standing pre-approval, and the exercise was repeated to confirm it.

What it looks like

  • The model-incident and stop playbooks are drilled at least twice a year, with injects
  • Exercises are attended by the people who would really be woken, not delegates
  • Every exercise produces a dated change to the playbook or an explicit decision not to change it
  • Rehearsal findings are reported alongside operational exercise results, not separately

Diagnostic signals you can check this week

  • Ask for the date of the last AI exercise and the change it produced. A date with no change means the exercise had no injects
  • Check whether the people in the exercise were the decision holders or their delegates
  • Ask the control room whether they can revert an AI-fed field without a change ticket, and whether they have ever done it
  • Count the versions of the model-incident playbook. A version 1.0 that is two years old has never been tested

Anti-pattern · Rehearsing only the approval, never the refusal

Exercises gravitate to the pleasant scenario: a good business case that passes the gate, a model that degrades gently and is caught by monitoring. The decisions that actually damage utilities are the refusals — declining a case a powerful sponsor wants, stopping a live use case that a site has become fond of, telling a regulator that a model contributed to a customer-facing error. Rehearse those. If nobody in the exercise has been made uncomfortable, the exercise has not tested the part of the playbook that will fail.

What holds you here

Exercises happen but their findings live in the exercise report rather than in the playbook, so the library improves more slowly than the estate changes.

Highest-leverage next move

Make the dated playbook diff the exercise's required output, and put the change list in front of the same forum that approved the original text.

Cost of leaving

Effort
6–12 months
Team
An exercise designer, the emergency-planning function, and two hours per quarter from each accountable executive
Risk
Medium — a well-designed exercise will expose that a documented authority does not exist in practice, which is politically uncomfortable and the entire point
To next stage
9–18 months

If this is you, the next step is

We write the scenario, run the two-hour exercise and hand you the change list.

Design and run your first AI inject

Stage 5

Institutional

5% of operators sit here

Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.

Stage 5 is defined by succession, not by sophistication. The test is what happens when the executive who sponsored the AI programme leaves: at stage 3 the library survives as documents and quietly stops being used; at stage 5 the new incumbent finds a versioned route, a decision history and a rehearsal calendar already in the diary, and their first quarter is spent on decisions rather than on rebuilding the machinery for taking them.

The evidence discipline is what makes this durable, and it is where the external world starts pulling in the same direction. The reference frameworks are converging on exactly this shape: the NIST AI Risk Management Framework ships a companion playbook of suggested actions organised around govern, map, measure and manage, and ISO/IEC 42001 defines an AI management system with documented decisions, roles and continual improvement. A utility that can produce the paper, the gate score, the dissent and the playbook version for a decision taken eighteen months ago is answering a question that regulators, auditors and insurers are all beginning to ask in similar words.

Sustaining stage 5 is a maintenance problem rather than a design problem, and it is the stage most likely to regress quietly. Estates change: an ADMS upgrade moves where a fallback lives, a reorganisation vacates an accountable name, a new vendor arrangement changes who holds the model. None of these fires an alarm. The defence is a review date on every playbook and an owner whose objectives include it — unglamorous, and the reason some libraries are still accurate five years after they were written.

In practice

The successor who inherited a route

A transmission business changed chief information officers. In her first month the incoming CIO asked how AI decisions were taken and was given eleven pages: six playbooks, each versioned and owned, a decision register listing twenty-three AI decisions with their approvers and criteria, and a rehearsal calendar with the next model-incident exercise already booked. Her first substantive act was to change one gate criterion. At stage 3 the same appointment would have consumed a quarter agreeing how decisions get made.

What it looks like

  • Every playbook has a version, an owner, a review date and a change history
  • Decision records are retrievable in minutes and reference the playbook version used
  • The library is cited in regulatory submissions and external assurance work
  • A new executive inherits the route rather than reinventing it in their first quarter

Diagnostic signals you can check this week

  • Ask how long it takes to produce the reasoning behind a named AI decision from eighteen months ago. Minutes is stage 5; a week is not
  • Check whether any playbook has an owner whose objectives reference it
  • Look for the library in a regulatory submission or an external assurance report
  • Ask whether the last executive change caused any playbook to be rewritten from scratch. It should not have

Anti-pattern · Treating the library as a compliance artefact

Once the playbooks are cited externally, the temptation is to optimise them for the reader rather than the user — longer, more defensive, cross-referenced to standards, and progressively less usable at 02:00. The library then bifurcates: an official version for assurance and an informal version the control room actually follows, which is the worst of both, because the evidence trail now describes a process nobody uses. Keep one version. If it is too long to act on under pressure, it is too long to be true.

What holds you here

Sustaining the library is a maintenance discipline — the constraint becomes drift between the written route and an estate that keeps changing.

Highest-leverage next move

Put a review date and a named owner on every playbook, and reconstruct one past decision from evidence alone each quarter as a standing check.

Cost of leaving

Effort
Continuous
Team
A named owner per playbook, a standing review cycle, and the emergency-planning function for the exercise calendar
Risk
Concentrated — low frequency, high consequence; the failure mode is silent drift between the written route and the real estate

If this is you, the next step is

We pick one past AI decision and try to reconstruct it from your evidence alone, then report what is missing.

Audit a decision the way a regulator would

Where utility leadership teams actually sit today

The distribution across the ladder, why the Drafted rung is so crowded, and what the sector's own research says about the barriers.

Most utility leadership teams sit at Drafted. The distribution below is weighted heavily toward libraries that exist on paper and have never settled a live decision — a majority of teams can produce an AI governance document, and a small minority can produce a dated record of the last exercise that changed one. The gap between those two states is the whole subject of this page.

Illustrative distribution of utility leadership teams across the ladder

Illustrative, not measured: the shares are the model behind this page's ladder, synthesised from the adoption and barrier research named beneath the chart, and they are labelled as such wherever they appear. Drafted is both the mode and the plateau — the drop from Drafted to Adopted is the largest single transition loss on the ladder.

Share of leadership teams

  • 24% — 1 · No playbook
  • 31% — 2 · Drafted (the plateau)
  • 27% — 3 · Adopted
  • 13% — 4 · Exercised
  • 5% — 5 · Institutional

Source: Illustrative distribution, synthesised from IEA, Eurelectric and EPRI adoption research

The barriers the sector's own research names are decision barriers as much as technical ones. The IEA's assessment of AI for energy optimisation (opens in a new tab) lists unfavourable regulation, lack of access to data, interoperability concerns, critical gaps in skills and general resistance to change as the obstacles to sector-wide deployment — and four of those five are settled at a leadership table rather than in an engineering one. A utility that cannot decide quickly and defensibly about data access or regulatory treatment does not have a modelling problem.

Adoption of such AI applications at a sector-wide level, however, is not a given.

The regulatory context sharpens the point. Network businesses answer to price-control and rate-case machinery — Ofgem's regulatory programmes (opens in a new tab) in Great Britain, FERC (opens in a new tab) and state commissions in the United States — which ask, in effect, how an investment decision was governed and why the consumer should fund it. Reliability obligations run in parallel: NERC's reliability standards (opens in a new tab) apply to the systems AI increasingly informs, and the EU AI Act's regulatory framework (opens in a new tab) treats AI managing critical infrastructure as high-risk, with documentation and human-oversight duties attached. Each of those regimes rewards exactly the artefact a used playbook produces as a by-product: a reconstructable decision record. Eurelectric (opens in a new tab) and McKinsey's electric power and natural gas insights (opens in a new tab) both track the same widening gap between sector ambition and governed delivery.

The playbook library: six decisions that recur

The six standing decision guides a utility leadership team needs — what fires each one, what the sponsor must bring, what the gate tests and what evidence it leaves behind.

Six AI decisions recur often enough in a utility to deserve a written route: funding, build versus buy, model incident, vendor, workforce change and stopping a use case. Everything else can reasonably be handled case by case. These six cannot, because each arrives under time pressure, with money, reliability or people attached, and each will be asked about later by someone who was not in the room — a regulator, an auditor, a board committee or a successor.

PlaybookTrigger — what fires itInputs the sponsor must bringDecision gate — what it testsEvidence it leaves behind
FundingA use case asks for money beyond discovery, or a use case must be placed in the next price-control or rate-case submissionOperational problem in the estate's own units; opex/capex treatment and TOTEX position; the holdout design; the stop conditionWhether the value is stated in SAIDI minutes, ETR accuracy, TOTEX or customer cost rather than in model accuracy — and whether the funding route matches the regulatory cycleBusiness case, gate score, funding route chosen, stop condition, approver and date
Build versus buyA vendor product covers a substantial share of a proposed use case, or an internal team proposes to build something a market product exists forMarket scan; the share of the decision that is genuinely estate-specific; data that only you hold; the five-year run cost of each optionWhether the decision differentiates your network and whether the data required is uniquely yours — not whether the team would enjoy building itComparison against the two axes, the option chosen, the exit position if the vendor route is taken
Model incidentA model output materially contradicts reality during a live network event, or reaches an operational decision while wrongNothing — this playbook must run with no preparation, from the control room's own informationSeverity class, who is woken, whether the fallback is switched, what must be preserved before anything is retrained, and who must be told outside the businessIncident record, model version frozen, decision log for the affected window, external notifications made
VendorAn AI-bearing product enters procurement, or an AI clause appears in a renewal of an existing EMS, ADMS, CMMS or AMI contractThe evidence pack: model documentation, data provenance, retrain cadence, failure modes, exit and portability terms, incident response commitmentsWhether the supplier can be held to an evidence standard you could show a regulator, and whether you can leave without losing the decisionCompleted evidence pack, contractual AI schedule, named supplier incident contact and out-of-hours route
Workforce changeA use case changes a role's task set, competency requirement or headcount plan — at any stage, including discoveryWhich roles, which tasks, what changes in the working day, the consultation position and its timingWhether employee representatives were consulted before the gate rather than informed after it, and whether the change is described in task terms rather than headcount termsConsultation record with dates, the role-by-role task description, any position not accepted and why
StopA live use case misses its stated evidence threshold in two consecutive review cycles, or its stop condition firesThe threshold as originally written; the two cycles of measurement; the reusable assets; the exit route for any dependent processWhich of the four exits applies — retire, absorb, hand back or park — and what leadership says publicly about itClosure record, assets banked for reuse, the public statement made, and the change to the next funding paper's threshold
The playbook library. Each row is one page in practice. 'Trigger' is deliberately an observable event rather than a request, so the playbook starts itself; 'evidence' is the by-product that later answers a regulator, an auditor or a successor.

Two rules keep the library usable. The first is the one-page rule: if a playbook cannot be acted on from a single page, it will not be acted on at 02:00, and length is almost always a sign that the document is being written for an assurance reader rather than for the person who must use it. The second is that the trigger is written as an observable event, never as a request — because a playbook whose trigger is 'when leadership decides to use it' is a voluntary act by a busy executive, and voluntary acts lose to time pressure every time.

  • Funding — the argument is about the regulatory clock, not the money

    In a regulated network business the decision is rarely whether the value exists; it is whether the spend is opex or capex, how it sits within TOTEX, and whether the case can reach the next price-control or rate-case window. A use case that misses that window by a quarter can wait a year or more for a funding route, and the sponsor will usually blame the technology.

  • Build versus buy — two axes, and neither of them is engineering appetite

    The only two questions that predict regret are whether the decision genuinely differentiates your network and whether the data it needs is uniquely yours. Load forecasting for a specific constrained group with your own AMI and DER history sits differently from meter-to-cash document handling, and the playbook exists to make that difference explicit before anyone falls in love with an architecture.

  • Model incident — the only playbook that must run with zero preparation

    Every other playbook is used by someone who has had time to prepare. This one is used by a control-room lead during an abnormal system state, possibly at night, and it must therefore be short, unambiguous about authority, and rehearsed. It is also the playbook most commonly missing, because its trigger never appears in anyone's calendar.

  • Vendor — buy the evidence, not just the model

    The failure mode is a capable product with no documentation you could show a regulator and no exit that leaves you with the decision. The playbook's job is to make the evidence pack a procurement requirement rather than a post-contract request, and to establish an out-of-hours supplier route before the first incident rather than during it.

  • Workforce change — a decision-rights question before an engagement one

    The playbook covers timing and rights only: consultation is an input to the gate, not a communication after it, and the paper records the representatives' position whether or not it was accepted. How that conversation is actually held across an estate — the site-by-site engagement — is the subject of the AI roadshow page, and the two are designed to be used together.

  • Stop — the playbook that changes behaviour before it is ever used

    Agreeing a stop condition before funding removes the political cost of being wrong later, which is why sponsors who have one propose bolder use cases. Where no stop route exists, use cases are defended long past the evidence, and the next proposal is funded against the memory of the last one that would not die.

Decision rights: who actually decides, and who must be asked

One accountable name per playbook, a consulted list agreed in advance, and a routing rule that keeps small decisions away from the executive committee.

Decision rights fail on the accountable column, almost never on the responsible one. Utilities are good at saying who will do the work and poor at saying who owns the outcome, so AI decisions get attributed to forums — a steering committee, a digital board, a risk panel — which is precisely the attribution that cannot be reconstructed eighteen months later when a regulator, an insurer or an internal auditor asks how a decision was reached and on what basis.

PlaybookAccountable — one nameResponsible — does the workConsulted — absence invalidates the decisionInformed — told after
FundingChief financial officer, or the regulated-business finance directorUse-case sponsor with the transformation or data leadRegulatory affairs (TOTEX and price-control treatment); the operational director whose KPI is claimedExecutive committee; internal audit
Build versus buyChief information officer or head of digitalEnterprise architecture with the use-case sponsorThe system owner for the EMS, ADMS, CMMS or AMI touched; information security; procurementFinance; the affected operational directorate
Model incidentHead of control, or the on-duty system operations manager out of hoursThe model owner and the control-room shift leadSafety; cyber-security duty officer; regulatory affairs when a reportable threshold may be crossedExecutive committee at next working start; the supplier if theirs
VendorChief procurement officer, jointly with the CIO for anything touching operational systemsProcurement lead with the model ownerInformation security; data protection; legal; the operational system ownerFinance; the AI gate; the supplier-management function
Workforce changeThe operational director whose people are affectedHR business partner with the use-case sponsorEmployee representatives; safety; the site or depot leadership affectedExecutive committee; the AI gate
StopThe same executive who was accountable for funding itThe use-case owner with the evaluation leadThe operational owner relying on the output; finance; anyone whose process consumes itExecutive committee; the gate, as a threshold lesson
Decision rights for the six playbooks, in RACI form. The accountable column carries exactly one role. The consulted column is the set whose absence invalidates the decision — not a distribution list — and it is agreed when the playbook is written, not when the paper arrives.

The single most useful line in that table is the last one: the executive accountable for stopping a use case is the same one who was accountable for funding it. Splitting those two creates the pathology every utility recognises — the sponsor advocates, an independent reviewer recommends closure, the sponsor defends, and the decision becomes about the two people rather than the evidence. Keeping both with one name makes stopping an act of stewardship rather than a defeat, and it is the cheapest structural fix on this page.

The consulted column needs the same discipline. It is not a distribution list; it is the short set of people whose absence makes the decision invalid, and it should be provocative enough that someone occasionally objects to being on it. In a network business it almost always includes the control-room authority for anything that reaches real-time operations, the safety function for anything that changes field work, and regulatory affairs for anything whose treatment affects a price-control or rate-case position.

Routing rule: which decisions belong at which level

Plot the decision on two axes — how reversible it is, and how far it reaches into operational systems. Three of the four quadrants should never reach an executive committee, and the routing rule is what keeps the gate available for the quadrant that should.

Delegate to the sponsor

  • Reversible, advisory only
  • Discovery work, analyst-facing models, back-office drafting
  • Route: sponsor decides, logged in the register, no gate

Delegate with a standing rule

  • Reversible, but reaches operational systems
  • Fields a control engineer can ignore or revert within a shift
  • Route: system owner decides against a written standing rule; gate informed

Executive gate

  • Hard to unwind, advisory only
  • Multi-year vendor commitments, workforce changes, public positions
  • Route: the playbook's accountable executive, at the gate

Gate plus rehearsal

  • Hard to unwind and reaches live operations
  • Automated switching support, protection-adjacent advice, ETR published to customers
  • Route: the gate, and no go-live until the incident playbook has been exercised on it
Reversibility — top: Easily reversed within a shift, bottom: Hard to unwind — contract, headcount, public commitment
Reach into operational systems — left: Advisory only — a person still decides, right: Writes into EMS, ADMS, OMS or field work

Workforce decisions sit in the bottom-left quadrant and are the ones most often mis-routed, because the technology decision and the people decision are taken months apart by different executives. The playbook's contribution is narrow and specific: consultation is an input to the gate rather than a communication after it, and the paper records the employee representatives' position whether or not it was accepted. How that conversation is actually held across control centres, depots, generating stations and contact centres is a separate discipline with its own page — the AI roadshow for energy and utility leaders covers the route, the stop types and the job question. This page stops at the decision right and the timing.

The money playbooks: funding, build versus buy and vendors

Opex pilot or capex platform, where the regulatory cycle actually binds, the two axes that decide build versus buy, and the evidence pack a supplier must clear.

The funding playbook exists because in a utility the hard part is almost never whether the value is real — it is which pocket the money comes from and when the window opens. A use case with an unarguable operational case can still stall for a year because it was framed as capex four months after the capital plan closed, or as opex in a business whose opex envelope is fixed until the next price-control or rate-case period. The playbook's job is to force that question at the start, when it is cheap to answer.

The funding playbook, in order

  1. State the problem in the estate's own units

    Not 'improve fault prediction' but 'reduce unplanned SAIDI minutes on the 11 kV network in these three groups' or 'raise ETR accuracy on wind-driven faults above the threshold customers are told'. A case written in model terms cannot be funded by anyone whose budget is measured in operational terms, and it cannot later be stopped cleanly either, because nothing measurable was promised.

  2. Decide the funding class before the amount

    Opex pilot, capitalised platform, innovation allowance or price-control submission — each has a different approval path, a different consumer-recovery position and a different time to money. Choosing the class first often changes the shape of the proposal: a decision that must reach the next submission window is scoped differently from one that can be run inside a discretionary opex envelope this quarter.

  3. Place it against the regulatory calendar

    Mark the submission or rate-case dates and work backwards. Ofgem's price-control and regulatory programmes (opens in a new tab) and the equivalent state and federal processes in the United States, coordinated through FERC (opens in a new tab), set windows that no amount of internal urgency will move. A quarter's slip in an internal decision can mean a year's slip in funding.

  4. Design the holdout before the build

    Name the feeders, the region, the depot or the customer segment that stays on the current process, and agree it with the operational owner in the same paper. Without it, seasonality, a network reconfiguration or an unusually mild winter will claim the credit or take the blame, and the gate will have no way to tell.

  5. Write the stop condition into the funding paper

    A specific, measurable threshold and the number of review cycles it may be missed for before the stop playbook fires. Written before the first score, it is a fact everyone agreed to; proposed afterwards, the same sentence reads as an attack on the sponsor. This single line is the highest-leverage sentence in the library.

Funding classWhat it suitsTime to moneyRegulatory treatmentWhat the gate should argue about
Discovery opexProving a data path exists and the problem is measurable at allWeeksOrdinary operating cost; rarely contestedWhether a decision would actually change if the answer were known
Opex pilotOne region, one asset class, one decision, with a holdoutOne to two quartersOperating cost within the existing envelope; TOTEX position notedWhether the holdout is real and whether the operational owner has signed up to the KPI
Capitalised platformShared data, serving and monitoring layers reused by several use casesTwo to four quartersCapitalised under the business's software policy; recovered over the asset lifeWhether use cases two and three genuinely exist, or are being assumed to justify the layer
Innovation allowanceGenuinely uncertain work with a shareable learning outputDepends entirely on the scheme's windowRegulator-defined; usually requires published learningWhether the business will actually adopt the result if it works, or treat it as a research artefact
Price control or rate caseMulti-year capability that must be recovered from consumersOne to several yearsAssessed against consumer benefit and delivery governanceWhether the decision record will stand up to the regulator asking how the investment was governed
Funding classes for a utility AI use case. The right-hand column is the question the gate should actually argue about — the amount is usually the least contested part of the paper.

The last row is where the playbook library pays for itself. A regulator assessing a multi-year AI capability is assessing governance as much as engineering: who decided, against what criteria, with what evidence, and what happens if it does not work. A business that can answer with a decision register is answering a different, easier question from one assembling the story retrospectively — and the assembly is always more expensive than the record would have been.

Build versus buy, on the only two axes that predict regret

Plot the decision, not the technology. Engineering appetite, vendor relationships and the availability of a demo are all irrelevant to this placement, and all three are what actually drive the choice when there is no playbook.

Buy, and integrate well

  • Differentiating decision, commodity data
  • Typically planning and simulation tooling
  • Fight for integration and exit terms, not for features

Build — or partner with your data

  • Differentiating decision on data only you hold
  • Constraint forecasting, feeder-level DER behaviour, asset health on your own fleet history
  • The only quadrant where building is usually right

Buy without hesitation

  • Undifferentiated decision, commodity data
  • Document handling, meter-to-cash exceptions, contact-centre drafting
  • Building here is the most expensive habit in the sector

Buy the engine, keep the features

  • Undifferentiated decision, but your data is the input
  • Vegetation imagery, inspection triage, load disaggregation
  • Buy the model, retain feature definitions and the right to leave with them
Does the decision differentiate your network? — top: Specific to your topology, licence area or asset mix, bottom: Everyone in the sector decides this the same way
Is the data uniquely yours? — left: Commodity or vendor-held data, right: Your AMI, SCADA, GIS and asset history only

The bottom-right quadrant is where most utility regret is manufactured. The product is genuinely good, the data going into it is yours, and the contract quietly makes the feature definitions the supplier's — so three years later the decision cannot move without rebuilding the inputs from scratch. The vendor playbook's evidence pack exists mostly to prevent that outcome, and the single clause that matters most is the one describing what you leave with.

RequirementWhy it exists in a utilityWhat a good answer looks likeWhere it is checked
Model documentationA regulator or inquiry may ask what the model does and on what it was trainedA written description of intended use, limits, known failure modes and evaluation data — not a marketing sheetGate paper; retained in the decision record
Data provenance and rightsYour AMI, SCADA and customer data may not lawfully train a general productExplicit statement of what is used, retained, and whether it improves the vendor's other customers' modelsData protection and legal review before the gate
Retrain cadence and change noticeA silent model update can change operational behaviour with no change controlA stated cadence, advance notice of material changes, and the right to hold a versionContractual AI schedule; verified at the first release
Failure modes and degradation behaviourThe control room needs to know what wrong looks like before it happensNamed failure modes with observable symptoms, and what the product does when inputs are staleModel-incident exercise, before go-live
Exit and portabilityDecisions outlive suppliers; the feature definitions are the assetYou leave with the feature definitions, the historical outputs and the labelled data you suppliedContract; rehearsed as a desktop exit scenario
Incident response and out-of-hours routeNetwork events do not respect business hoursA named contact, a response time, and a route that works at 02:00 on a SundayTested during the incident exercise, not during an incident
Standards alignmentAssurance and insurance questions increasingly cite named frameworksEvidence against ISO/IEC 42001 or the NIST AI RMF functions, with scope stated honestlyAssurance review; cited in regulatory submissions
The vendor evidence pack. Every row is a procurement requirement rather than a post-contract request, and every row has a place where it is checked rather than merely received.

Two of those rows are rehearsed rather than filed, and that is deliberate. A documented out-of-hours supplier route that has never been dialled is a phone number, not a capability — the same class of assumption as a fallback nobody has switched. Testing both during a scheduled exercise costs an hour and routinely finds that the number reaches a service desk with no authority to act on an operational system. EPRI's artificial intelligence work (opens in a new tab) and the shared evaluation effort behind the Open Power AI Consortium (opens in a new tab) exist in part because every utility was otherwise assembling this evidence alone, for the same suppliers.

The model-incident playbook: when a model is wrong during a live event

What leadership does in the first hour when an AI output contradicts reality on the network — severity, authority, containment and what must be preserved.

When a model is wrong during a live network event, the first decision is not diagnostic — it is whether to keep using it. That decision belongs to the control room and must be exercisable without a change ticket, an approval or a phone call to a data team. Every other question, including what went wrong, can wait; the network cannot. This is the playbook most utilities do not have, because its trigger never appears in anyone's calendar and nothing else forces it to be written.

The first hour

  1. Declare, and let anyone declare

    An AI incident is a wrong or unavailable model output that has reached, or could reach, an operational decision. Anyone may declare one — a control engineer, a field supervisor, a contact-centre team leader who notices ETRs going out that cannot be right. Declaration is deliberately cheap; the cost of a false declaration is one phone call, and the cost of a delayed one is measured in restoration minutes.

  2. Contain by switching the fallback

    Revert to the previous decision source — the rule-based estimate, the static schedule, the manual assessment — as a control-room authority action requiring no external approval. This is the step that fails most often in exercises, not because the fallback does not exist but because switching it turns out to require a change ticket nobody has ever raised under pressure.

  3. Establish the blast radius

    Which decisions, in which window, used the output: which feeders were switched, which crews dispatched, which ETRs published, which customers told. This is the question that determines everything downstream — the severity class, whether an external notification is due, and how much of the day has to be reviewed — and it is answerable in minutes only if the decision log exists.

  4. Classify and wake the right people

    Severity drives who is woken, not seniority and not who is closest. A classification table agreed in advance means the shift lead does not have to negotiate at 02:00 about whether this warrants calling a director, which is exactly the negotiation that delays containment.

  5. Preserve before you fix

    Freeze the model version, the inputs it saw and the decision log for the affected window before anything is retrained, patched or rolled forward. Utilities lose more evidence to a well-intentioned overnight fix than to any other cause, and the evidence is what a later review, a regulator or an insurer will actually ask for.

  6. Notify outside the business on the class's rule

    Whether the system operator, the regulator or customers must be told is a pre-agreed function of the severity class, not a judgement call taken while tired. Reliability and reporting obligations — including those under NERC's reliability standards (opens in a new tab) and equivalent national regimes — apply to the operational outcome regardless of whether a model was involved.

ClassWhat it looks like on the networkWho is wokenExternal notificationRestart condition
1 · CriticalModel output contributed to a switching, dispatch or protection-adjacent decision during an abnormal system stateOn-duty system operations manager, head of control, executive on call, cyber duty officerSystem operator and regulator per the reporting regime; customer communication if ETRs were affectedFull replay of the affected window, a written cause, and a re-exercised fallback
2 · MajorWrong output reached customer-facing commitments or crew dispatch at scale — ETRs, appointment windows, priority customer handlingHead of control, model owner, customer operations leadCustomer communication; regulator if a service obligation was missedCause identified, holdout comparison re-run, gate informed
3 · ContainedOutput was wrong but the fallback was switched before any operational decision consumed itModel owner and shift lead; logged for the next gateNone, unless a pattern across incidents emergesMonitoring adjusted; incident logged as an exercise input
4 · DegradedInputs stale or the model unavailable; the decision proceeded on the fallback with no error introducedModel owner during working hoursNoneFreshness alerting reviewed; no gate action
Severity classes for an AI model incident in a network business. The point of the table is that nobody negotiates the class at 02:00 — the observable symptom determines it, and the class determines everything else.
Likelihood: highImpact: high

The fallback exists but has never been switched

Documented, tested once during commissioning, and requiring a change ticket in normal service. Under storm conditions the control room finds it cannot revert without raising a ticket that takes hours, so the wrong output stays in the decision path because switching it off is harder than working around it.

PreventionMake the revert a standing control-room authority with a pre-approved change, and switch it once per exercise cycle on a quiet shift.

Likelihood: highImpact: medium

The incident is classified as an IT incident

It goes to the service desk, gets a priority based on system availability, and is worked to an IT restoration target. Meanwhile the operational consequence — crews dispatched on wrong priorities, ETRs published that cannot hold — has no owner, because the operational incident route was never told.

PreventionDefine an AI incident by its operational consequence, not by which system failed, and give it a route into the operational incident process from the first minute.

Likelihood: mediumImpact: high

The model is retrained before the evidence is frozen

A well-meaning engineer retrains or rolls forward overnight to stop the recurrence. The fix works and the evidence is gone: the version that produced the wrong output, the inputs it saw and the reasoning are unreconstructable, which is fatal if a regulator or an insurer later asks.

PreventionPut 'preserve before you fix' ahead of remediation in the playbook, and give the model owner an explicit instruction that freezing outranks fixing.

Likelihood: mediumImpact: medium

The playbook names a role that no longer exists

A reorganisation merged two functions, the named escalation is now vacant, and the shift lead spends fifteen minutes working out who to call. Nothing announced the drift, because event-triggered playbooks are only read when they are needed.

PreventionReview every name after any structural change, and make the exercise a check on the call list as well as on the decisions.

Exercising and evidencing the library

Four exercise formats and the rule that the output is a dated change; then the evidence layers, the eight metrics and the reconstruction test that settles whether the library is real.

A playbook library is real when it has been exercised and when a decision taken against it can be reconstructed from evidence alone — rehearsal proves the route works, records prove it was followed. An unexercised playbook is a document, not a capability, and a utility already believes this about everything else it does. Nobody assumes a black-start procedure works because it is written and approved; it is exercised, it fails in three specific places, and the procedure changes. AI decisions deserve exactly the same treatment and rarely get it, because the library sits with governance functions that do not have an exercise habit while the exercise habit sits with emergency planning, who have not been asked.

The cheapest way to start is therefore not to build an AI exercise programme but to attach an AI inject to the one that already runs. A storm exercise in which the outage-prediction model is confidently wrong about one 33 kV group costs almost nothing to add, uses people who are already in the room, and tests the fallback, the escalation and the notification rule in a single afternoon.

FormatWhat it actually testsWho must be in the roomCadenceEvidence it produces
Desktop walk-throughWhether the text is comprehensible and the names are currentThe accountable executive and the playbook ownerTwice a year per playbookA dated change list, or an explicit decision not to change
Injected drillWhether authority, containment and escalation exist in practice under time pressureThe real decision holders — control-room lead, model owner, on-call executive — not delegatesTwice a year for the incident playbook; annually for stopChange list plus the timings: how long to contain, to classify, to notify
Live shadowWhether the fallback can genuinely be switched in normal serviceControl room and the system owner, on a quiet shiftOnce per release of the affected operational systemA confirmed revert time and any change control that had to be pre-approved
Post-event replayWhether the playbook as written matched what people actually didEveryone who was involved, plus one person who was notAfter every real incident, refusal or stopThe diff between the written route and the actual one, and which of the two should change
Four exercise formats for a playbook library. Cadence assumes an estate of ordinary complexity; the point is that each format tests something the others cannot, and the desktop walk-through alone will not find an authority that does not exist.

The last row is the most valuable and the most often skipped. After a real event, people write an incident report and file it; almost nobody compares the written route with what actually happened. Where the two differ, one of them is wrong — and it is worth being genuinely open about which. Sometimes the playbook was right and was not followed, which is a rehearsal problem; more often the playbook was written for an organisation that no longer exists, and the people improvising were correct.

Exercise-readiness checklist

Eight conditions that separate a rehearsal from a meeting about a rehearsal. Tick as you go — this list works without JavaScript.

0 of 8 ticked

Nothing ticked — start with one inject, not a programme

Zero is the honest answer for most leadership teams and it is not a crisis. Do not design an exercise programme; attach one AI inject to the next emergency exercise that is already scheduled. Two hours, people who are already in the room, and it will produce three findings you can act on immediately.

The library is real if you can reconstruct a decision from evidence alone. That is the whole test, and it is worth running as a standing quarterly exercise: pick an AI decision taken more than a year ago, and try to produce the paper, the criteria it was scored against, who was accountable, who was consulted, what dissent was recorded and which version of the playbook was in force. Everything in this section exists to make that reconstruction take minutes rather than a week.

The evidence stack behind a playbook library

Four layers, each annotated with the ladder stage that first requires it. Nothing here is a product — every layer is defined by what it must guarantee, and in most utilities the first two are a well-disciplined register in an existing system rather than anything new.

  1. Decision register

    Stage 3+

    • Decision recordOne per gate outcome: paper, score, outcome, date
    • Playbook version in forceWhich text the decision was taken against
    • Dissent capturedRecorded positions not accepted, and by whom
  2. Use-case and model register

    Stage 3+

    • Accountable executiveOne name per live use case, kept current
    • System of record touchedEMS, ADMS, OMS, CMMS, AMI — and the fallback for each
    • Risk class and thresholdRouting class from the matrix, plus the stop condition
  3. Incident and exercise log

    Stage 4+

    • Incident recordClass, blast radius, frozen version, notifications made
    • Exercise recordScenario, injects, attendees, timings observed
    • Playbook diffThe dated change each exercise or event produced
  4. Board and regulatory pack

    Stage 5+

    • Quarterly decision summaryDecisions taken, refused and stopped, with reasons
    • Submission extractThe governance evidence a price control or rate case asks for
    • Reconstruction pathThe documented route from a decision reference to its full record

Pipeline described

  1. Decision register (stage 3+) — Decision record: One per gate outcome: paper, score, outcome, date; Playbook version in force: Which text the decision was taken against; Dissent captured: Recorded positions not accepted, and by whom
  2. Use-case and model register (stage 3+) — Accountable executive: One name per live use case, kept current; System of record touched: EMS, ADMS, OMS, CMMS, AMI — and the fallback for each; Risk class and threshold: Routing class from the matrix, plus the stop condition
  3. Incident and exercise log (stage 4+) — Incident record: Class, blast radius, frozen version, notifications made; Exercise record: Scenario, injects, attendees, timings observed; Playbook diff: The dated change each exercise or event produced
  4. Board and regulatory pack (stage 5+) — Quarterly decision summary: Decisions taken, refused and stopped, with reasons; Submission extract: The governance evidence a price control or rate case asks for; Reconstruction path: The documented route from a decision reference to its full record
Step-by-step insights
Decision register — the version field is the one people forget
Most registers capture what was decided and by whom, and omit which version of the playbook was in force. That field is what makes the record interpretable later: a decision that looks strange today may have been exactly right against the criteria of the time, and without the version nobody can tell the difference between a bad decision and a changed standard. It costs one column and it is the difference between an archive and an audit trail.
Use-case register — the fallback belongs here, not in the runbook
Recording the system of record each use case touches and the fallback for each is what lets an incident be scoped in minutes rather than hours. It also produces an uncomfortable but useful count: how many live use cases write into an operational system without a documented, exercised fallback. In most utilities that number is not zero, and nobody has ever seen it in one place before the register exists.
Incident and exercise log — the timings are the leading indicator
Capturing what happened matters less than capturing how long it took: minutes to contain, to classify, to notify. Those three numbers are the closest thing a playbook library has to a health metric, they are comparable across exercises and real events, and a rising containment time is an early sign that the estate has drifted away from the written route — usually because a system upgrade moved where the fallback lives.
Board and regulatory pack — assemble it quarterly, not on demand
A pack assembled quarterly from a live register costs an hour; the same pack assembled retrospectively under regulatory or inquiry pressure costs weeks and is systematically less convincing, because the effort shows. The quarterly summary of decisions taken, refused and stopped is also the single most effective internal advertisement for the library: sponsors who can see that two cases were refused and one stopped treat the gate as real.

The metrics below are the instrumentation for the four dimensions the assessment scores. All of them are readable from the registers above rather than from a survey, which matters: a self-reported view of playbook maturity is reliably one stage optimistic, because approved documents are memorable and absent exercises are not.

MetricHow it is readSourceCadenceHonest from
Playbook coveragePlaybooks written ÷ the six recurring decisionsThe library itselfQuarterlyStage 2
Gate adherenceAI decisions taken through a gate ÷ all AI decisions takenDecision register vs finance and procurement recordsQuarterlyStage 3
Time to decisionTrigger event → gate outcome, elapsed daysDecision registerPer decisionStage 3
Input-pack completenessPapers accepted first time ÷ papers submittedGate secretary's logPer meetingStage 3
Exercise cadence adherenceExercises run ÷ exercises scheduled, by playbookExercise logHalf-yearlyStage 4
Playbook change rateDated changes produced per exercisePlaybook version historyPer exerciseStage 4
Containment timeDeclaration → fallback switched, in minutesIncident and exercise logPer incident or drillStage 4
Reconstruction timeMinutes to produce the full reasoning for a named past decisionStanding quarterly test against the registerQuarterlyStage 5
Instrumentation for a playbook library. 'Honest from' is the ladder stage at which the metric first measures something real — reading a gate-adherence figure at stage 2 measures nothing, because there is no gate.

Two of these deserve a place in a board pack rather than a governance report. Gate adherence answers the question a non-executive actually has — are AI decisions being taken the way we said they would be, or is the gate a formality some decisions route around? And reconstruction time answers the one an auditor or regulator will ask, in a unit anyone can understand. Reported together they say more about a utility's AI governance than any count of models in production. Eurelectric's published work (opens in a new tab) and the broader management literature on AI decision-making (opens in a new tab) both point the same way: the constraint on value is the decision process, not the model.

A 90-day plan: two playbooks around the vegetation-risk model

The No-playbook-to-Exercised move made concrete on one decision — the imagery model that reallocates a distribution business's vegetation-cutting programme. Contains no model development.

A library takes about 90 days to start when it is scoped to one decision, and several years when it is scoped to a function. To make that concrete, the plan below runs the transition on a specific and very common distribution problem: an imagery model that ranks spans by vegetation-encroachment risk and is being used to reallocate the cutting programme. It touches money, reliability, contractor workforce and a live operational consequence if it is wrong, which makes it the ideal first subject — and the model itself already exists, so the quarter contains no model development at all.

From no playbook to one exercised decision, in one quarter

One decision, one licence area, one accountable executive. If any phase needs longer than its window, narrow the scope — one region rather than the estate, one asset class rather than the programme — rather than extending the plan.

  1. Days 1–15

    Name the decision and reconstruct the last three

    Write down exactly which decision is being governed: how the vegetation-cutting programme is reallocated between spans on the strength of a model's risk ranking. Name one accountable executive — normally the asset or network director whose SAIDI it moves. Then attempt to reconstruct how the last three comparable spend-reallocation decisions were made, from records alone. The reconstruction is the diagnostic: what you cannot find is what the playbook must produce.

    One named decision, one accountable executive, an honest evidence gap list

  2. Days 16–45

    Draft the funding and model-incident playbooks

    One page each, written with the executive who holds the decision rather than for them. The funding page states the problem in SAIDI minutes and cutting cost, fixes the opex-or-capex class against the price-control calendar, names the holdout — the districts that stay on the cyclical programme — and, critically, writes the stop condition before the model is scored. The incident page states what happens if the model has under-ranked a span that then faults during a wind event: who is told, what reverts to the cyclical schedule, and what is preserved.

    Two one-page playbooks with named triggers and one accountable name each

  3. Days 46–70

    Exercise both, including the refusal

    Two scenarios, two hours each, with the real decision holders. First, a business case that fails the gate: the sponsor is credible, the number is attractive, the holdout is not real — rehearse saying no and see whether the route survives the seniority in the room. Second, the storm inject: the model under-ranked a span that faulted, a customer is off supply, and a local journalist has asked whether the cutting programme was changed by an algorithm. Change the text of both playbooks the same week.

    Two dated playbook versions and a change list with owners

  4. Days 71–90

    Run one real decision through the gate and file it

    Take the actual reallocation decision through the gate as written. File the paper, the score against the pre-agreed criteria, any dissent, and the playbook version in force. Then publish two things internally: the decision record and the fact that the playbook changed twice because of the exercises. The second publication is what makes the library credible to the next sponsor.

    One reconstructable decision record and a library people believe in

The order matters

  1. Write the stop condition before the first score

    A threshold agreed before anyone knows the answer is a fact everybody signed; the identical sentence proposed after the first disappointing quarter is an attack on a colleague. This single line determines whether the stop playbook is ever usable, and it costs nothing while the case is still hypothetical.

  2. Rehearse the refusal, not just the approval

    Exercises drift toward pleasant scenarios, and pleasant scenarios test nothing. The decision that actually damages a utility is declining a case a powerful sponsor wants, or stopping a live use case a region has grown fond of. If nobody in the room was uncomfortable, the exercise did not reach the part of the playbook that will fail.

  3. One decision through the gate beats six playbooks on a shelf

    Resist the pull to complete the library before using any of it. A single real decision routed through one imperfect playbook teaches the organisation more — and produces more accurate text — than six approved documents nobody has tested. Write the other four after the first two have survived contact.

The stop playbook: retiring a use case without political cost

Four exits, who announces the closure, what happens to the evidence, and why a stop condition written before funding changes how sponsors behave.

A use case is retired without political cost when the condition for retiring it was agreed before it was funded. That is the whole mechanism. Where the threshold exists in the original paper, closure is the system working as designed and the sponsor is executing a commitment; where it does not, closure is a judgement about a colleague's work, and everybody involved knows it — which is why so many utility AI use cases neither succeed nor stop, but persist in a maintained, unloved, quietly expensive state.

The second mechanism is who says it. The closure is announced by the executive who was accountable for funding it, not by a reviewer, an audit function or a new arrival. This is why the decision-rights table keeps funding and stopping under one name: it converts the act from an external judgement into stewardship, and it removes the adversarial dynamic that otherwise makes every stop decision a contest between two people rather than a reading of the evidence.

ExitWhen it is the right oneWhat leadership saysWhat happens to the evidence
RetireThe stated threshold was missed across two review cycles and there is no credible route to it"We set a threshold, we measured against it honestly, and it did not clear. Here is what we learned and where it goes next."Data path, feature definitions and the holdout design are banked for the next use case; the threshold lesson goes to the gate
AbsorbThe value is real but does not justify a standing use case — it belongs inside an existing process or product"This works, and it is no longer a project. It is now part of how the outage process runs."Ownership moves to the process owner; the model joins that system's change control and monitoring
Hand backA vendor or partner can now do this as a product and the differentiation has gone"The market caught up. Keeping this in-house costs us more than it returns."Feature definitions and labelled data are retained under the vendor playbook's exit terms
ParkThe blocker is external and dated — a data-access ruling, a system upgrade, a regulatory determination"This is not wrong, it is early. Here is the specific event that will restart it and who is watching for it."The case, the data path and the named restart trigger are archived together with a review date
Four exits from a live AI use case. Most utilities have only the first and treat it as failure, which is why so few use cases are ever closed — the other three are cheaper, more honest and much more common in practice.
  • Write the threshold in the funding paper, in operational units

    "Reduces unplanned SAIDI minutes on these groups by a stated amount against the holdout, measured over two review cycles." A threshold in model-accuracy terms cannot be honestly stopped against, because everyone can argue that a slightly better model is a quarter away.

  • Name the number of cycles before you know the answer

    Two consecutive misses is a defensible default in a business with seasonal effects; one is noise and three is a year. Fixing the count in advance prevents the most common failure, which is an indefinite extension granted one quarter at a time by people who are individually being reasonable.

  • Bank the assets explicitly, and say so publicly

    The data path, the feature definitions, the holdout design and the integration work usually outlive the model. Closing a use case while naming what carries forward is what makes a stop feel like a return rather than a write-off, and it is factually true far more often than it is said.

  • Send the threshold lesson back to the gate

    Every stop is evidence about how the organisation sets thresholds. If three use cases have now missed thresholds set the same way, the gate criteria are wrong rather than the sponsors. That feedback is the difference between a library that learns and one that merely records.

The sector's own framing supports treating adoption as a decision that can go either way. The IEA is explicit that sector-wide AI adoption is not a given (opens in a new tab), and its wider digitalisation work (opens in a new tab) describes an energy system where the constraint is organisational as often as technical. A leadership team with a working stop route is better placed on both counts: it can commit harder to the use cases that clear their thresholds precisely because it has a credible way to release the ones that do not.

What governed AI decisions look like in public

Three publicly reported programmes, read against the playbook ladder. None is an Atomic Loops engagement — each links to the organisation's own published material.

Utility decision routes are rarely published, but the structures built to hold them are — which makes those structures the best available public evidence for this page's argument. In each case below the durable artefact is institutional rather than technical: a shared evaluation body, a stated method for turning grid planning into a repeatable capability, and an annually republished set of digital commitments. Each maps onto a different rung of the ladder.

Three programmes read against the ladder

Outcomes as reported by the organisations themselves. Verify figures against the linked source before reusing them; we have not independently audited them. Images are illustrative energy-sector scenes from our library, not photographs of these organisations or their sites.

Illustration of analysts at control-room consoles reviewing power system waveform and network analytics on large curved displaysEPRI — Open Power AI ConsortiumSector research institute · utilities, technology providers and researchers25
Challenge
Every utility evaluating the same AI products was assembling the same evidence alone — model documentation, benchmarks, use-case prioritisation — at full cost, and with no shared standard for what an acceptable answer looked like.
Approach
EPRI founded and leads the Open Power AI Consortium as a standing collaboration between utilities, technology providers and researchers, publishing shared datasets, models and benchmarking tools, and running demonstration work such as the AI for Power Challenge to translate prioritised use cases into tested projects.
Reported outcome
EPRI reports the consortium engaging over 300 organisations, including more than 120 utilities and energy companies and over 150 technology providers, with the stated aim of accelerating adoption and reducing duplication across the sector.
What it shows about the curveThis is Institutional at sector scale: the evidence standard behind the vendor playbook is far cheaper to hold collectively than alone, and a shared benchmark is exactly the artefact a gate needs when a supplier's own material is the only alternative.

EPRI — Open Power AI Consortium (opens in a new tab)

Illustration of two engineers in high-visibility clothing reviewing monitoring data from a glass-walled control booth above an instrumented utility tunnelDuke EnergyUS investor-owned utility · 8.2 million electric customers34
Challenge
Grid planning decisions depended on simulations that took weeks to run, against a capital programme the company has publicly described as around USD 145 billion over the following decade — so the decision cycle, not the analysis, was the binding constraint.
Approach
Duke Energy has publicly described building its Intelligent Grid Services — a suite of custom applications for anticipating demand and identifying where the grid needs updating — with AWS, and separately deploying AI to scan digital channels for scams targeting its customers, each announced as a bounded capability with a named executive owner rather than as a general AI initiative.
Reported outcome
Duke Energy's chief information officer is quoted in the company's own announcement stating the aim to "run those same simulations in 15 minutes or less", against a baseline the release describes as taking weeks.
What it shows about the curveBounded, individually announced capabilities with a named executive and a stated before-and-after are what a gate produces. The alternative — an undifferentiated AI programme — cannot be funded, measured or stopped as separate decisions.

Duke Energy — smart grid solutions with AWS (opens in a new tab)

Illustration of two utility staff in high-visibility clothing examining a large network schematic display showing a transmission tower and fault indicatorsNESO (National Energy System Operator, GB)GB electricity system operator · national control centre35
Challenge
A system operator's digital commitments are scrutinised by a regulator, by industry participants and by the public, so decisions about data, AI and the systems they touch cannot rest on internal memory.
Approach
NESO publishes a Digitalisation Strategy and Action Plan and republishes it as commitments change — the June 2025 edition describing "a flexible-led approach, which outlines how we're transforming our people, processes, data and technology" and naming "embracing early AI adoption" among its principles, alongside its separately published work on AI.
Reported outcome
NESO reports its digital and AI commitments in public documents aligned to its business plan and to Ofgem's regulatory requirements, so a stated intention has a dated, citable version rather than an internal one.
What it shows about the curvePublishing the commitments is the Institutional move applied to the evidence dimension: an externally dated document cannot quietly drift, and a successor inherits a position rather than a set of recollections.

NESO — June 2025 Digitalisation Strategy and Action Plan (opens in a new tab)

Read together the three describe the same pattern from different directions: decisions become durable when the structure holding them is external to any individual. A shared evaluation body outlives a procurement team, a bounded capability with a stated before-and-after outlives its sponsor, and a published commitment outlives an executive's memory of what was agreed — see also EPRI's press release on launching the consortium (opens in a new tab), Duke Energy's AI scam-detection announcement (opens in a new tab) and NESO's published work on AI (opens in a new tab).

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

AI leadership playbook
A one-page standing decision guide for a recurring AI decision, stating its trigger, decision rights, required inputs, gate criteria and the evidence it leaves behind. Distinct from an AI policy, which sets boundaries rather than describing a route.
Trigger
The observable event that starts a playbook — a use case asking for money beyond discovery, an AI clause in a renewal, a model output contradicting the control room. Written as an event rather than a request, so using the playbook is not a voluntary act by a busy executive.
Decision gate
The standing forum and moment at which a playbook's criteria are applied to a specific case, producing go, hold or stop. Its defining property is that the criteria existed before the case did.
Input pack
The closed set of items a sponsor must bring to a gate — no more and no fewer. Closing the set caps preparation cost for small use cases and lets a chair return an incomplete paper in two lines.
Decision rights (RACI)
The allocation of responsible, accountable, consulted and informed roles for a decision. In this library the accountable column always carries exactly one name, because a decision attributed to a committee cannot be reconstructed later.
Decision record
The filed by-product of a gate: the paper, the score against the criteria, any dissent, the approver, the date and the playbook version in force. The artefact a regulator, auditor or successor actually needs.
Exercise (inject)
A rehearsal of a playbook against an invented scenario, with a deliberate complication — the inject — introduced to test authority and escalation under pressure. Its required output is a dated change to the playbook, not a satisfied meeting.
Stop condition
A specific, measurable threshold written into the funding paper before the first score, with the number of review cycles it may be missed for. It is what makes retiring a use case an act of stewardship rather than a judgement on a colleague.
Fallback source
The previous decision source a model replaced — a rule-based estimate, a cyclical schedule, a manual assessment — and the switch that reverts to it. In a control room the switch must be a standing authority, not a change ticket.
Model incident
A wrong or unavailable model output that has reached, or could reach, an operational decision. Defined by its operational consequence rather than by which system failed, so it routes into the operational incident process rather than to a service desk.
Blast radius
The set of decisions, in a defined window, that consumed a model's output during an incident — feeders switched, crews dispatched, ETRs published, customers told. Determines severity class, notification duty and how much of the day must be reviewed.
TOTEX
Total expenditure — the combined capital and operating cost category regulators increasingly assess as one, with a capitalisation rate applied. It is why the funding playbook forces the opex-or-capex question before the amount is discussed.

Frequently asked questions

The questions utility leadership teams ask most often when building a playbook library.

What is an AI leadership playbook?

It is a one-page standing decision guide for a recurring AI decision. It states the trigger that fires it, who is accountable and who must be consulted, the inputs the sponsor must bring, the criteria the gate applies and the evidence the decision leaves behind. The point is that the same question — how to fund a use case, what to do when a model is wrong — is settled by a written route instead of being re-argued from first principles by whoever is in the room that week.

How is a playbook different from an AI policy?

A policy is a boundary and a playbook is a route. The policy states what is not allowed, which data may leave the estate and who must approve exceptions; it decides nothing on its own. The playbook takes one recurring decision and describes how it actually gets made, by whom, against what criteria. Utilities that write only the policy typically spend a year generating approvals without generating decisions, because there is no route through the gates the policy created.

How many playbooks does a utility actually need?

Six: funding, build versus buy, model incident, vendor, workforce change and stopping a use case. Those are the decisions that recur often enough, arrive under enough time pressure and carry enough money, reliability or people risk to justify a standing route. Everything else can be handled case by case. Adding more is the most common way to make a library unusable, because length lowers the odds that any executive reads it under pressure.

Who should be accountable for AI decisions in a utility?

One named executive per playbook, never a committee. Funding usually sits with the finance director of the regulated business, build versus buy with the CIO or head of digital, model incident with the head of control or the on-duty system operations manager, vendor with procurement jointly with the CIO, workforce change with the operational director whose people are affected, and stopping a use case with the same executive who was accountable for funding it. That last pairing is the cheapest structural fix available.

What does an unexercised playbook actually cost?

It costs the difference between finding a defect in a two-hour drill and finding it during a storm. Exercises reliably surface three classes of problem: a fallback that cannot be switched without a change ticket, an escalation path to a role that no longer exists, and a supplier out-of-hours number that reaches a service desk with no authority to act. Each is trivial to fix in advance and expensive during a live network event, when the cost is measured in restoration minutes and customer trust.

How do we fund an AI use case in a regulated network business?

Decide the funding class before the amount. Discovery opex proves a data path exists; an opex pilot with a real holdout proves value in one region; a capitalised platform is justified only when use cases two and three genuinely exist; an innovation allowance suits uncertain work with a shareable learning output; and a price-control or rate-case submission suits multi-year capability. Then place the decision against the regulatory calendar and work backwards — a quarter's internal slip can mean a year's slip in funding.

When should a utility build AI rather than buy it?

Only when the decision genuinely differentiates your network and the data it needs is uniquely yours — constraint forecasting on your own topology, feeder-level DER behaviour, asset health on your own fleet history. Undifferentiated decisions on commodity data, such as document handling or meter-to-cash exceptions, should be bought without hesitation. The awkward quadrant is a product built on data only you hold: buy the engine, but retain the feature definitions and the contractual right to leave with them.

What happens when a model is wrong during a live network event?

Contain first, diagnose later. Anyone may declare an AI incident; the control room reverts to the fallback source as a standing authority with no change ticket; the blast radius is established from the decision log — which feeders, which crews, which ETRs, which customers; the severity class determines who is woken and what must be reported externally; and the model version, its inputs and the decision log are frozen before anything is retrained. Utilities lose more evidence to a well-intentioned overnight fix than to any other cause.

How do you stop an AI use case without it looking like failure?

Agree the stop condition before you fund it, and let the executive who funded it announce the closure. A threshold written into the original paper makes closure the system working as designed; the same sentence proposed after a disappointing quarter reads as an attack on a colleague. Name the exit honestly — retire, absorb into an existing process, hand back to a vendor, or park against a dated external blocker — and state publicly which assets carry forward, because the data path and feature definitions usually do.

How often should playbooks be exercised?

Twice a year as a desktop walk-through for every playbook, twice a year as an injected drill for the model-incident playbook, annually for the stop playbook, and a live-shadow test of the fallback with each release of the operational system it touches. Add a post-event replay after every real incident, refusal or stop. The cheapest way to start is to attach an AI inject to an emergency exercise that is already scheduled rather than building a separate programme.

How do these playbooks help with a regulator or a price control?

They produce the evidence those processes ask for as a by-product. A price-control or rate-case assessment of a multi-year AI capability is an assessment of governance as much as engineering: who decided, against what criteria, with what evidence, and what happens if it does not work. A decision register answers that in minutes. Frameworks are converging on the same shape — the NIST AI Risk Management Framework's companion playbook and ISO/IEC 42001 both assume documented decisions, assigned roles and continual improvement.

Where do these playbooks sit relative to an AI roadshow?

They are complements with different jobs. A roadshow is a periodic change instrument: leadership physically taking a strategy round control centres, sites and depots to build site-level ownership and capture use cases. The playbooks are the standing decision guides used between roadshows, when a captured idea asks for money, a supplier proposes an AI module or a live model is wrong at 02:00. A roadshow without a playbook library generates ideas that have nowhere to be decided.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, utilities and manufacturing operators — forecasting, asset health, outage analytics and decision support running against live SCADA, ADMS, OMS and CMMS data, integrated into the systems control rooms and depots already run on rather than delivered as dashboards.

  • · Deployments across generation, transmission, distribution and energy retail
  • · Decision-gate and model-incident playbooks written with utility executive teams
  • · Integration-first delivery: system-of-record write-back, monitoring, rollback
  • · 25 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyEnergy and AI — executive summary (opens in a new tab)
  3. International Energy AgencyAI for energy optimisation and innovation (opens in a new tab)
  4. International Energy AgencyDigitalisation and the energy system (opens in a new tab)
  5. EurelectricEurelectric (opens in a new tab)
  6. EurelectricPublications (opens in a new tab)
  7. North American Electric Reliability CorporationReliability standards (opens in a new tab)
  8. North American Electric Reliability CorporationReliability assessment and performance analysis (opens in a new tab)
  9. NERCNorth American Electric Reliability Corporation (opens in a new tab)
  10. OfgemOfgem (opens in a new tab)
  11. OfgemPolicy and regulatory programmes (opens in a new tab)
  12. FERCFederal Energy Regulatory Commission (opens in a new tab)
  13. NISTAI Risk Management Framework (opens in a new tab)
  14. NIST Trustworthy & Responsible AI Resource CenterAI RMF Playbook (opens in a new tab)
  15. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  16. European CommissionRegulatory framework for AI (opens in a new tab)
  17. Harvard Business ReviewAI and machine learning (opens in a new tab)
  18. McKinsey & CompanyElectric Power & Natural Gas insights (opens in a new tab)
  19. EPRIOpen Power AI Consortium (opens in a new tab)
  20. EPRIEPRI launches the Open Power AI Consortium (opens in a new tab)
  21. EPRIArtificial intelligence (opens in a new tab)
  22. Duke EnergySmart grid solutions developed with AWS (opens in a new tab)
  23. Duke EnergyAI to protect customers and combat scams (opens in a new tab)
  24. NESOJune 2025 Digitalisation Strategy and Action Plan (opens in a new tab)
  25. NESOThe future of the ESO and artificial intelligence (opens in a new tab)

Write the six playbooks — then exercise the two that matter

We interview the executives who actually hold each decision, draft the missing playbooks on a page each in their own words, run the first gate meeting and design the first model-incident inject against a scenario drawn from your own network. You keep the library, the scenarios and the decision-record template either way.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.