Energy & UtilitiesLeadership Insights & Strategy
AI leadership playbooks for energy and utilities: the decision guides a leadership team runs on
AI leadership playbooks are the standing decision guides a utility leadership team uses so that recurring AI decisions — how to fund one, whether to build or buy, what to do when a model is wrong during a live network event, which vendor, which workforce change, when to stop — are settled by a written route rather than re-argued from scratch.

Key takeaways
- An unexercised playbook is a document, not a capability. A utility already knows this about black-start and storm response; the same rule applies to AI decisions, and the test is whether leadership has ever rehearsed one against an invented scenario before a real event forced it.
- Six decisions recur often enough in a utility to deserve a standing guide: funding, build versus buy, model incident, vendor selection, workforce change and stopping a use case. Everything else can be handled case by case; these six cannot, because they arrive under time pressure and with money or reliability attached.
- A playbook is not a policy. A policy states what is not allowed; a playbook states how a specific decision gets made — its trigger, its decision rights, the inputs the sponsor must bring, the gate criteria and the evidence it leaves behind.
- Decision rights fail on the accountable column, not the responsible one. Every playbook needs exactly one named accountable executive; committees that decide collectively produce decisions nobody can reconstruct eighteen months later when a regulator asks.
- The stop playbook is the one most often missing and the one that most changes behaviour. Where a stop condition is agreed before a use case is funded, retiring it costs a sponsor nothing; where it is not, use cases are defended long past the evidence and the next proposal is funded against that memory.
Abbreviations used on this page
- RACI
- Responsible, accountable, consulted, informed — the decision-rights notation used in every playbook on this page
- SCADA
- Supervisory control and data acquisition
- EMS
- Energy management system (the control-centre application suite)
- ADMS
- Advanced distribution management system
- OMS
- Outage management system
- CMMS
- Computerised maintenance management system (the work-order system of record)
- AMI
- Advanced metering infrastructure (smart meters and their head end)
- DER
- Distributed energy resources — rooftop solar, batteries, heat pumps, EV chargers
- SAIDI
- System average interruption duration index — the headline reliability KPI
- ETR
- Estimated time of restoration (the figure given to a customer during an outage)
- TOTEX
- Total expenditure — the combined capex and opex category regulators increasingly assess as one
- RIIO
- Revenue = Incentives + Innovation + Outputs — Ofgem's price-control framework for GB network companies
Free · 8 questions · ~3 minutes
Score your playbook library
Eight questions, one at a time, about three minutes. Answer them and we build your personalised playbook report — your stage on the ladder, your score on each of the four dimensions, and the specific gap standing between you and the next stage — and send it to your inbox. Your result doubles as the coverage map for your first drafting session.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised playbook report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, which of the six playbooks you are missing, and the drafting and rehearsal sequence for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · No playbook
No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
Your next moveWrite down the six decisions that recur, name one accountable executive for each, and draft the two that hurt most — usually funding and model incident — on a single page each.
Stage 2 · Drafted
Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
Your next moveTake the highest-pain playbook, rewrite it with its accountable executive in their own words on one page, and put the next real decision of that type through its gate.
Stage 3 · Adopted
Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
Your next moveSchedule a desktop exercise of the model-incident playbook against an invented storm scenario, with the people who would actually be woken in the room.
Stage 4 · Exercised
Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
Your next moveMake the dated playbook diff the exercise's required output, and put the change list in front of the same forum that approved the original text.
Stage 5 · Institutional
Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.
Your next movePut a review date and a named owner on every playbook, and reconstruct one past decision from evidence alone each quarter as a standing check.
0 / 24
Playbook coverage
— / 6
Decision rights clarity
— / 6
Exercise cadence
— / 6
Evidence & learning loop
— / 6
Your score maps to a stage on the ladder. The dimension breakdown matters more than the total: a library with excellent coverage and no exercise cadence is a set of documents, and it will fail the first time an event tests it. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the ladder. The dimension breakdown matters more than the total: a library with excellent coverage and no exercise cadence is a set of documents, and it will fail the first time an event tests it.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the missing playbooks drafted with your decision holders?
We interview the executives who actually hold each decision, draft the missing playbooks in their words on a page each, and run the first gate and the first exercise with you. You keep the library and the exercise scenarios either way.
How the score maps to a stage
- 0–4 — Stage 1, No playbook. No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
- 5–9 — Stage 2, Drafted. Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
- 10–14 — Stage 3, Adopted. Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
- 15–20 — Stage 4, Exercised. Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
- 21–24 — Stage 5, Institutional. Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.
What AI leadership playbooks are — and what they are not
A definition, the five parts every playbook needs, and the difference between a decision route and an AI policy.
AI leadership playbooks are standing decision guides: one page per recurring decision, stating what fires it, who decides, what the sponsor must bring, what the gate tests and what evidence the decision leaves behind. They exist so that a utility leadership team stops re-deriving the same positions — on funding, on build versus buy, on what to do when a model is wrong at 02:00 — every time a new use case reaches the table.
The distinction that matters most is between a playbook and a policy. A policy is a boundary: it states what is not allowed, who must be consulted, which data may leave the estate. It is necessary, and it decides nothing. A playbook is a route: it takes a specific, recurring question and describes how it gets answered, by whom, against what criteria, in what order. Utilities that write the policy first typically spend a year producing approvals without producing decisions, because the policy generates gates without generating routes through them.
The form is not novel and it is not ours. The NIST AI Risk Management Framework (opens in a new tab) ships with a companion Playbook of suggested actions (opens in a new tab) organised around govern, map, measure and manage, and ISO/IEC 42001 (opens in a new tab) defines an AI management system built on documented decisions, assigned roles and continual improvement. Both describe the shape of the artefact. What neither can supply is the part that makes a playbook work in a network business: the specific trigger in your estate, the specific executive who holds the decision, and the gate criteria expressed in SAIDI minutes, ETR accuracy or TOTEX rather than in general risk language.
How an AI decision reaches a conclusion, with and without a playbook
The two routes a recurring AI question can take through a utility. The top lane is not disorganised — it is undocumented, which is a different and more expensive problem. The middle lane is the playbook route; the bottom lane is the loop that keeps it accurate. Most leadership teams are in the top lane.
- Human in the loop
- Where value leaks
- Data & feeds
- System-of-record action
The process, in words
- Without a playbook, a question arrives from the board, a regulator or a sponsor, gets a paper written in whatever format the author prefers, and reopens the same first principles every time — is this opex or capex, who signs it, what happens if it does not work. The decision is settled by whoever is in the room that week and nothing durable is filed, so the next question starts from the same place.
- On the playbook route, a defined trigger fires rather than someone requesting a slot. The sponsor assembles the named input pack and nothing beyond it, decision rights route the paper to exactly one accountable executive with a written consulted list, the gate scores it against criteria agreed before the case existed, and the outcome — go, hold or stop — is filed with the score, any dissent and the playbook version in force.
- The learning loop is what stops the route becoming fiction. Scheduled exercises and real events both feed back into a dated, owned revision of the playbook, and the next decision of that type runs against the new version. Without this loop a library ages faster than the estate changes: names leave, systems upgrade, and the written route quietly stops describing reality.
Step-by-step insights
- The trigger is the part most libraries omit
- Almost every drafted playbook describes a process and forgets to say what starts it. That omission is why unused libraries stay unused: with no trigger, using the playbook is a voluntary act by a busy executive, and voluntary acts lose to time pressure. A trigger is a specific, observable event — a use case asking for money beyond discovery, an AI clause appearing in a renewal, a model output contradicting the control room during an abnormal state, two consecutive review cycles below threshold. Written that way, the playbook starts itself, and the question 'should we run this?' never has to be asked.
- Why the input pack must be closed, not open
- Sponsors will bring whatever they think helps, which under uncertainty means everything. A closed input pack — these six items, nothing else scored — does two useful things at once. It caps preparation cost, so a small use case is not priced out of the gate by paperwork; and it makes incompleteness visible, so a chair can return a paper in two lines rather than debating an eighty-slide deck. The discipline compounds: after one paper is returned for a missing stop condition, every subsequent paper has one.
- One accountable name, and what the consulted list is for
- The accountable column carries a single name because reconstruction depends on it: eighteen months later, 'the steering committee agreed' is not an answer a regulator or an inquiry can work with. The consulted list does a different job — it is the pre-agreed set of people whose absence invalidates the decision, typically the control-room authority for anything touching real-time operations, the safety function for anything touching field work, and employee representatives for anything that changes a role's task set. Naming them in advance prevents the most common gate failure, which is a technically sound decision taken without the one person who could have said why it would not work on the ground.
- The gate scores criteria that existed before the case did
- A gate that invents its criteria while reading the paper is a debate, not a gate. The criteria belong in the playbook, agreed when nobody's proposal is on the table and therefore nobody's interests are engaged — which is the only time an organisation can set an honest threshold. This is also why the stop condition must be written before the first score: a threshold agreed in advance is a fact, and the same threshold proposed after a use case exists is an accusation.
- The decision record is the by-product that becomes the asset
- Filing the paper, the score, any dissent and the playbook version costs a few minutes per decision and is the difference between a library that can be audited and one that merely exists. Its value shows up in three places a utility will recognise: a price-control or rate-case submission that needs to show how investment decisions were governed, an incident review that needs to show what was known when, and a leadership change where a successor can read twenty decisions instead of interviewing twenty people.
- The loop is why exercises outrank approvals
- Approval tests whether a document is acceptable; an exercise tests whether it is true. The bottom lane exists because the estate keeps moving — an ADMS release relocates a fallback, a reorganisation vacates an accountable name, a vendor contract changes who holds the model — and none of those changes announces itself to the library. A scheduled exercise, plus a replay after every real event, is the only mechanism that reliably finds the drift before an incident does.
The five stages in detail
For each stage: what it looks like inside a real utility, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps leadership teams there, and what leaving costs.
The ladder runs from No playbook to Institutional, and its shape is worth noticing before the detail: the value released is close to flat across the first two rungs and inflects at Adopted, when live decisions actually start passing through a gate. A library that exists and is never used releases almost exactly as much value as no library at all, which is why the count of approved documents is the least informative metric an AI programme can report.
Decision throughput released along the playbook ladder
Value stays flat while playbooks are documents — most leadership teams sit on that flat section — and inflects at Adopted, when a real decision first passes through a gate. The second inflection is Exercised: rehearsal is what converts a route that works on paper into one that holds under a live network event. Drawn from the ladder on this page, not from a measured dataset.
Leadership decision throughput by stage
- Stage 1 · No playbook — 24% of operators. No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
- Stage 2 · Drafted — 31% of operators. Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
- Stage 3 · Adopted — 27% of operators. Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
- Stage 4 · Exercised — 13% of operators. Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
- Stage 5 · Institutional — 5% of operators. Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the IEA's assessment that sector-wide AI adoption is not a given.
Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions inside a network or generation business, the diagnostic signals are checks you can run against your own governance records this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
No playbook
24% of operators sit here
No playbook is the stage where every AI decision is argued from first principles by whoever is in the room, with no record of how the last one was settled.
Stage 1 is not the absence of AI activity — most utilities at this stage have several models running and at least one that works. It is the absence of any repeatable route by which a decision about that activity gets made. Each question arrives fresh, gets a bespoke paper, and is settled by the balance of seniority in the room on the day. The decision may well be correct; it is simply not reproducible, and nothing about it makes the next one cheaper.
The tell is the re-argument. Ask how the last three AI funding decisions were made and you will hear three different processes described by three different sponsors, each convinced theirs was the normal one. The same underlying questions — is this opex or capex, who signs it off, what happens if it does not work — are re-opened every time, because none of the answers was written anywhere a successor could find. Senior time is spent re-deriving positions the organisation has already held.
This stage is cheap to leave and expensive to occupy, and the cost is asymmetric. A good decision made this way earns no institutional credit, because it cannot be pointed at; a bad one is catastrophic, because there is no record showing what was considered. In an industry where a regulator, a system operator or a public inquiry may reasonably ask how a decision was reached, an undocumented route is a liability that grows quietly with every use case added.
In practice
The vegetation model funded three times
A distribution business built an imagery model that ranks spans by vegetation-encroachment risk. It was funded as a discovery piece by the asset director, re-argued four months later as an innovation project when discovery money ran out, and argued a third time as part of the next capital plan — three papers, three formats, three different sets of success criteria. Nothing was wrong with any of them. The cost was six months of executive attention spent on a decision the organisation had already effectively taken twice.
What it looks like
- AI questions reach the executive committee as one-off papers, each in its author's own format
- There is no written list of which AI decisions leadership actually owns
- The funding argument is re-run from scratch for every use case
- Nobody can produce the reasoning behind an AI decision taken six months ago
Diagnostic signals you can check this week
- Ask three executives who approves a model that writes into the ADMS. Three different answers means you are here
- Ask for the paper behind the last AI funding decision. If it takes more than a day to find, there is no decision record
- Look for the words 'trigger' and 'gate' in any AI governance document. Their absence is the definition of stage 1
- Ask what would happen at 03:00 if a live model were badly wrong. If the answer starts with a name rather than a route, there is no playbook
Anti-pattern · Writing an AI policy instead
The instinctive fix is a policy — acceptable use, prohibited use, a statement of principles, a sign-off from legal. It is genuinely useful and it solves a different problem. A policy tells people what they may not do; it does not tell an executive committee how to decide whether to fund a DER forecasting platform, or what to do at 02:00 when a model has mis-ranked a feeder during a storm. Utilities that write the policy first typically wait another year before anything decides anything, because the policy generates approvals without generating routes.
What holds you here
There is no written route for any recurring AI decision, so each one consumes senior time re-deriving positions the organisation has already held.
Highest-leverage next move
Write down the six decisions that recur, name one accountable executive for each, and draft the two that hurt most — usually funding and model incident — on a single page each.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- One executive sponsor, a programme lead, and half a day each from the four executives who actually hold the decisions
- Risk
- Low — nothing operational changes, but the choice of which six decisions to write sets the library's credibility
- To next stage
- 4–8 weeks
If this is you, the next step is
A two-week engagement: interview the decision holders, name the triggers, draft the first two playbooks.
Stage 2
Drafted
31% of operators sit here
Drafted is the stage where playbooks exist as documents — usually written by one team, approved once, and never yet used to settle a live decision.
Stage 2 is the most common resting place for a utility AI programme and the easiest to mistake for progress. The artefacts are real: a decision framework, a risk taxonomy, sometimes an impressively detailed RACI matrix produced by a consultancy. The library exists on a page. What has not happened is any decision passing through it, which means every assumption in it is untested — including the assumption that the people named are the people who actually decide.
The structural cause is authorship. The playbooks were written by whoever was asked to write them, usually a function with an interest in coherence rather than one holding the decision. So the funding playbook is written by transformation rather than by the finance director who will apply it; the model-incident playbook is written by the data team rather than by the head of control who will be woken. Each document is internally consistent and externally unowned, and the first live decision routes around it because the person deciding never agreed to be bound.
The other tell is drift against reality. A drafted playbook ages badly: it names a head of digital who has left, a governance forum that merged, an ADMS release that shipped. Nobody notices, because nothing is exercising it. By the time a real incident tests the text, roughly a third of it is wrong — and the credibility damage from a playbook that fails in use is worse than having none at all, because it teaches the organisation that the library is decorative.
In practice
The framework that survived the reorganisation on paper only
A vertically integrated utility approved an AI decision framework in the spring: five decision types, an escalation ladder, a named owner per type. In the autumn the digital and data functions merged, two of the five named owners changed and the risk committee's terms of reference were rewritten. The framework document was not touched. When a supplier proposed an AI-bearing asset-health module in December, procurement followed the standard capital route and nobody opened the framework, because the person it named no longer worked there.
What it looks like
- A set of AI decision guides exists, typically produced by a transformation or risk function
- The documents were approved at a governance forum and circulated
- No live decision has yet been taken through a playbook's gate
- The playbooks describe an organisation chart that has already changed
Diagnostic signals you can check this week
- Ask for the last three decisions taken through a playbook's gate. If the answer is none, you are at stage 2 whatever the documents say
- Check every name in the library against the current organisation chart — count how many are wrong
- Ask who wrote the funding playbook and who applies it. Different people, with no shared session, is the signature
- Look for a version number and a date on each playbook. Their absence means nobody expects the text to change
Anti-pattern · Adding more playbooks to fix an unused library
When a library is not used, the reflex is to extend it: more decision types, more detail, an appendix of templates. It is the wrong lever. An unused library is not too thin, it is unowned — and doubling its length lowers the odds that any executive reads it under time pressure. The correct move is to take the single playbook that hurts most, sit the person who actually holds that decision down with it, cut it to one page in their words, and route the next real decision through it, even clumsily. One used playbook beats six approved ones.
What holds you here
The playbooks were written by a function that does not hold the decision, so the first live decision routes around them.
Highest-leverage next move
Take the highest-pain playbook, rewrite it with its accountable executive in their own words on one page, and put the next real decision of that type through its gate.
Cost of leaving
- Effort
- 6–10 weeks
- Team
- Each playbook's accountable executive for two sessions, plus a lead who is allowed to cut text
- Risk
- Medium — the first live decision through a gate will expose the parts that were written for tidiness rather than use
- To next stage
- 6–10 weeks
If this is you, the next step is
We facilitate the first gate meeting, then rewrite the playbook from what actually happened in it.
Stage 3
Adopted
27% of operators sit here
Adopted is the stage where live decisions genuinely go through the playbooks — the gate is in the calendar and papers arrive in its format.
Stage 3 is the first stage at which the library changes behaviour rather than describing it. The mechanism is mundane and it is the whole trick: a standing slot in the calendar, a defined input pack, and a chair willing to return an incomplete paper. Once a sponsor has had a paper returned for missing the stop condition, the next paper has a stop condition. The playbook stops being a document and becomes a queue discipline.
The character of the executive conversation changes here too. At stage 2 the meeting spends its first twenty minutes agreeing what is being decided and by what standard; at stage 3 that is settled before anyone sits down, so the discussion is about the specific evidence in front of it. Utilities notice this as a throughput effect — more AI decisions taken per quarter, with less senior time each — well before they notice it as a governance effect.
The constraint that emerges is rehearsal, and it emerges asymmetrically. Playbooks with a scheduled trigger — funding rounds, vendor renewals, workforce planning — get exercised naturally by the calendar, so they improve. Playbooks whose trigger is an event nobody schedules — a model badly wrong during a storm, a use case that must be stopped — are never exercised at all, and the first time they run is the first time anyone reads them, at 02:00, under pressure, with the network in an abnormal state.
In practice
The paper that was returned, and the quarter that followed
A network operator's AI gate met monthly. In its third month a well-regarded sponsor brought a DER-forecasting business case with no stop condition and no named consulted list; the chair returned it with two lines of feedback. It was resubmitted a month later, complete, and approved. The visible effect was one month lost. The unseen effect was that all four papers submitted the following quarter arrived complete, and the gate's average decision time fell from two meetings to one.
What it looks like
- AI funding and vendor decisions arrive at a standing gate, not at whichever forum has space
- Sponsors submit the playbook's input pack because papers without it are returned
- Each decision produces a record naming the approver, the criteria and the date
- The accountable executive for each playbook is a single named person, not a committee
Diagnostic signals you can check this week
- Look at the last four AI papers. Do they share a structure, or does each reflect its author?
- Ask whether a paper has ever been returned for an incomplete input pack. Never is a warning sign, not a compliment
- Check whether the model-incident playbook has ever run. Scheduled playbooks improve; event-triggered ones rot
- Ask a sponsor what happens if their use case misses its threshold. A confident, specific answer means the gate is real
Anti-pattern · Letting the gate become a reporting meeting
Once a gate is established it attracts attendance, and attendance turns decisions into updates. The agenda fills with progress reports, the decision items slip to the end, and within two quarters the forum is a status meeting with a decision-shaped item nobody has time to argue about. The defence is structural: keep the gate to decisions only, cap it at ninety minutes, and give progress reporting a different meeting with a different chair. A gate that cannot say no in under an hour has stopped being a gate.
What holds you here
Only the calendar-triggered playbooks get used, so the event-triggered ones — incident and stop — are still untested when they first run for real.
Highest-leverage next move
Schedule a desktop exercise of the model-incident playbook against an invented storm scenario, with the people who would actually be woken in the room.
Cost of leaving
- Effort
- 3–6 months
- Team
- A gate chair with authority to return papers, a secretary who maintains the decision record, and the accountable executives
- Risk
- Medium — the first refusal is a political event, and how it is handled sets whether the gate has teeth
- To next stage
- 3–6 months
If this is you, the next step is
Terms of reference, input pack templates and the first three gate meetings run with you.
Stage 4
Exercised
13% of operators sit here
Exercised is the stage where playbooks are rehearsed against invented scenarios before a real event tests them, and the rehearsal changes the text.
Stage 4 imports a discipline the industry already has and rarely applies to AI. A utility does not assume its black-start procedure works because it is written; it exercises it, discovers that a phone number is wrong and a substation key is in the wrong cabinet, and updates the procedure. The same reasoning applies exactly to an AI model-incident playbook, and the same class of finding comes out: the fallback that has never been switched, the escalation path to a role that no longer exists, the assumption that the vendor answers out of hours.
The economics of rehearsal are what make it worth executive time. An exercise costs two hours from six people and finds three defects; the same defects found during a real storm cost restoration minutes, customer trust and, in some jurisdictions, a reportable event. Utilities that already run emergency exercises can attach an AI inject to the existing programme rather than creating a new one — a scenario in which the outage-prediction model is confidently wrong during the drill is cheap to add and disproportionately informative.
What separates a real exercise from theatre is whether the text changes. An exercise that ends with everyone agreeing the playbook worked has usually been run without injects, or without anyone empowered to say the escalation was wrong. The output artefact is not a satisfied room; it is a dated diff. Programmes that keep the diff visible — this playbook has changed four times, here is why — build a kind of institutional confidence that no amount of approval can manufacture.
In practice
The drill that found the fallback nobody had switched
A DSO ran a two-hour desktop exercise on its outage-prediction model: the inject was a wind event where the model badly under-predicted faults on one 33 kV group. The control-room lead reached for the documented fallback — reverting the ADMS field to the rule-based estimate — and discovered the revert required a change ticket that took four hours in normal service. Nobody had ever switched it. The playbook changed that week: the revert became a control-room-authority action with a standing pre-approval, and the exercise was repeated to confirm it.
What it looks like
- The model-incident and stop playbooks are drilled at least twice a year, with injects
- Exercises are attended by the people who would really be woken, not delegates
- Every exercise produces a dated change to the playbook or an explicit decision not to change it
- Rehearsal findings are reported alongside operational exercise results, not separately
Diagnostic signals you can check this week
- Ask for the date of the last AI exercise and the change it produced. A date with no change means the exercise had no injects
- Check whether the people in the exercise were the decision holders or their delegates
- Ask the control room whether they can revert an AI-fed field without a change ticket, and whether they have ever done it
- Count the versions of the model-incident playbook. A version 1.0 that is two years old has never been tested
Anti-pattern · Rehearsing only the approval, never the refusal
Exercises gravitate to the pleasant scenario: a good business case that passes the gate, a model that degrades gently and is caught by monitoring. The decisions that actually damage utilities are the refusals — declining a case a powerful sponsor wants, stopping a live use case that a site has become fond of, telling a regulator that a model contributed to a customer-facing error. Rehearse those. If nobody in the exercise has been made uncomfortable, the exercise has not tested the part of the playbook that will fail.
What holds you here
Exercises happen but their findings live in the exercise report rather than in the playbook, so the library improves more slowly than the estate changes.
Highest-leverage next move
Make the dated playbook diff the exercise's required output, and put the change list in front of the same forum that approved the original text.
Cost of leaving
- Effort
- 6–12 months
- Team
- An exercise designer, the emergency-planning function, and two hours per quarter from each accountable executive
- Risk
- Medium — a well-designed exercise will expose that a documented authority does not exist in practice, which is politically uncomfortable and the entire point
- To next stage
- 9–18 months
If this is you, the next step is
We write the scenario, run the two-hour exercise and hand you the change list.
Stage 5
Institutional
5% of operators sit here
Institutional is the stage where the library survives its authors — versioned, owned, exercised on a calendar, and cited in board papers and regulatory submissions.
Stage 5 is defined by succession, not by sophistication. The test is what happens when the executive who sponsored the AI programme leaves: at stage 3 the library survives as documents and quietly stops being used; at stage 5 the new incumbent finds a versioned route, a decision history and a rehearsal calendar already in the diary, and their first quarter is spent on decisions rather than on rebuilding the machinery for taking them.
The evidence discipline is what makes this durable, and it is where the external world starts pulling in the same direction. The reference frameworks are converging on exactly this shape: the NIST AI Risk Management Framework ships a companion playbook of suggested actions organised around govern, map, measure and manage, and ISO/IEC 42001 defines an AI management system with documented decisions, roles and continual improvement. A utility that can produce the paper, the gate score, the dissent and the playbook version for a decision taken eighteen months ago is answering a question that regulators, auditors and insurers are all beginning to ask in similar words.
Sustaining stage 5 is a maintenance problem rather than a design problem, and it is the stage most likely to regress quietly. Estates change: an ADMS upgrade moves where a fallback lives, a reorganisation vacates an accountable name, a new vendor arrangement changes who holds the model. None of these fires an alarm. The defence is a review date on every playbook and an owner whose objectives include it — unglamorous, and the reason some libraries are still accurate five years after they were written.
In practice
The successor who inherited a route
A transmission business changed chief information officers. In her first month the incoming CIO asked how AI decisions were taken and was given eleven pages: six playbooks, each versioned and owned, a decision register listing twenty-three AI decisions with their approvers and criteria, and a rehearsal calendar with the next model-incident exercise already booked. Her first substantive act was to change one gate criterion. At stage 3 the same appointment would have consumed a quarter agreeing how decisions get made.
What it looks like
- Every playbook has a version, an owner, a review date and a change history
- Decision records are retrievable in minutes and reference the playbook version used
- The library is cited in regulatory submissions and external assurance work
- A new executive inherits the route rather than reinventing it in their first quarter
Diagnostic signals you can check this week
- Ask how long it takes to produce the reasoning behind a named AI decision from eighteen months ago. Minutes is stage 5; a week is not
- Check whether any playbook has an owner whose objectives reference it
- Look for the library in a regulatory submission or an external assurance report
- Ask whether the last executive change caused any playbook to be rewritten from scratch. It should not have
Anti-pattern · Treating the library as a compliance artefact
Once the playbooks are cited externally, the temptation is to optimise them for the reader rather than the user — longer, more defensive, cross-referenced to standards, and progressively less usable at 02:00. The library then bifurcates: an official version for assurance and an informal version the control room actually follows, which is the worst of both, because the evidence trail now describes a process nobody uses. Keep one version. If it is too long to act on under pressure, it is too long to be true.
What holds you here
Sustaining the library is a maintenance discipline — the constraint becomes drift between the written route and an estate that keeps changing.
Highest-leverage next move
Put a review date and a named owner on every playbook, and reconstruct one past decision from evidence alone each quarter as a standing check.
Cost of leaving
- Effort
- Continuous
- Team
- A named owner per playbook, a standing review cycle, and the emergency-planning function for the exercise calendar
- Risk
- Concentrated — low frequency, high consequence; the failure mode is silent drift between the written route and the real estate
If this is you, the next step is
We pick one past AI decision and try to reconstruct it from your evidence alone, then report what is missing.
Where utility leadership teams actually sit today
The distribution across the ladder, why the Drafted rung is so crowded, and what the sector's own research says about the barriers.
Most utility leadership teams sit at Drafted. The distribution below is weighted heavily toward libraries that exist on paper and have never settled a live decision — a majority of teams can produce an AI governance document, and a small minority can produce a dated record of the last exercise that changed one. The gap between those two states is the whole subject of this page.
Illustrative distribution of utility leadership teams across the ladder
Illustrative, not measured: the shares are the model behind this page's ladder, synthesised from the adoption and barrier research named beneath the chart, and they are labelled as such wherever they appear. Drafted is both the mode and the plateau — the drop from Drafted to Adopted is the largest single transition loss on the ladder.
Share of leadership teams
- 24% — 1 · No playbook
- 31% — 2 · Drafted (the plateau)
- 27% — 3 · Adopted
- 13% — 4 · Exercised
- 5% — 5 · Institutional
Source: Illustrative distribution, synthesised from IEA, Eurelectric and EPRI adoption research
The barriers the sector's own research names are decision barriers as much as technical ones. The IEA's assessment of AI for energy optimisation (opens in a new tab) lists unfavourable regulation, lack of access to data, interoperability concerns, critical gaps in skills and general resistance to change as the obstacles to sector-wide deployment — and four of those five are settled at a leadership table rather than in an engineering one. A utility that cannot decide quickly and defensibly about data access or regulatory treatment does not have a modelling problem.
Adoption of such AI applications at a sector-wide level, however, is not a given.
The regulatory context sharpens the point. Network businesses answer to price-control and rate-case machinery — Ofgem's regulatory programmes (opens in a new tab) in Great Britain, FERC (opens in a new tab) and state commissions in the United States — which ask, in effect, how an investment decision was governed and why the consumer should fund it. Reliability obligations run in parallel: NERC's reliability standards (opens in a new tab) apply to the systems AI increasingly informs, and the EU AI Act's regulatory framework (opens in a new tab) treats AI managing critical infrastructure as high-risk, with documentation and human-oversight duties attached. Each of those regimes rewards exactly the artefact a used playbook produces as a by-product: a reconstructable decision record. Eurelectric (opens in a new tab) and McKinsey's electric power and natural gas insights (opens in a new tab) both track the same widening gap between sector ambition and governed delivery.
The playbook library: six decisions that recur
The six standing decision guides a utility leadership team needs — what fires each one, what the sponsor must bring, what the gate tests and what evidence it leaves behind.
Six AI decisions recur often enough in a utility to deserve a written route: funding, build versus buy, model incident, vendor, workforce change and stopping a use case. Everything else can reasonably be handled case by case. These six cannot, because each arrives under time pressure, with money, reliability or people attached, and each will be asked about later by someone who was not in the room — a regulator, an auditor, a board committee or a successor.
| Playbook | Trigger — what fires it | Inputs the sponsor must bring | Decision gate — what it tests | Evidence it leaves behind |
|---|---|---|---|---|
| Funding | A use case asks for money beyond discovery, or a use case must be placed in the next price-control or rate-case submission | Operational problem in the estate's own units; opex/capex treatment and TOTEX position; the holdout design; the stop condition | Whether the value is stated in SAIDI minutes, ETR accuracy, TOTEX or customer cost rather than in model accuracy — and whether the funding route matches the regulatory cycle | Business case, gate score, funding route chosen, stop condition, approver and date |
| Build versus buy | A vendor product covers a substantial share of a proposed use case, or an internal team proposes to build something a market product exists for | Market scan; the share of the decision that is genuinely estate-specific; data that only you hold; the five-year run cost of each option | Whether the decision differentiates your network and whether the data required is uniquely yours — not whether the team would enjoy building it | Comparison against the two axes, the option chosen, the exit position if the vendor route is taken |
| Model incident | A model output materially contradicts reality during a live network event, or reaches an operational decision while wrong | Nothing — this playbook must run with no preparation, from the control room's own information | Severity class, who is woken, whether the fallback is switched, what must be preserved before anything is retrained, and who must be told outside the business | Incident record, model version frozen, decision log for the affected window, external notifications made |
| Vendor | An AI-bearing product enters procurement, or an AI clause appears in a renewal of an existing EMS, ADMS, CMMS or AMI contract | The evidence pack: model documentation, data provenance, retrain cadence, failure modes, exit and portability terms, incident response commitments | Whether the supplier can be held to an evidence standard you could show a regulator, and whether you can leave without losing the decision | Completed evidence pack, contractual AI schedule, named supplier incident contact and out-of-hours route |
| Workforce change | A use case changes a role's task set, competency requirement or headcount plan — at any stage, including discovery | Which roles, which tasks, what changes in the working day, the consultation position and its timing | Whether employee representatives were consulted before the gate rather than informed after it, and whether the change is described in task terms rather than headcount terms | Consultation record with dates, the role-by-role task description, any position not accepted and why |
| Stop | A live use case misses its stated evidence threshold in two consecutive review cycles, or its stop condition fires | The threshold as originally written; the two cycles of measurement; the reusable assets; the exit route for any dependent process | Which of the four exits applies — retire, absorb, hand back or park — and what leadership says publicly about it | Closure record, assets banked for reuse, the public statement made, and the change to the next funding paper's threshold |
Two rules keep the library usable. The first is the one-page rule: if a playbook cannot be acted on from a single page, it will not be acted on at 02:00, and length is almost always a sign that the document is being written for an assurance reader rather than for the person who must use it. The second is that the trigger is written as an observable event, never as a request — because a playbook whose trigger is 'when leadership decides to use it' is a voluntary act by a busy executive, and voluntary acts lose to time pressure every time.
Funding — the argument is about the regulatory clock, not the money
In a regulated network business the decision is rarely whether the value exists; it is whether the spend is opex or capex, how it sits within TOTEX, and whether the case can reach the next price-control or rate-case window. A use case that misses that window by a quarter can wait a year or more for a funding route, and the sponsor will usually blame the technology.
Build versus buy — two axes, and neither of them is engineering appetite
The only two questions that predict regret are whether the decision genuinely differentiates your network and whether the data it needs is uniquely yours. Load forecasting for a specific constrained group with your own AMI and DER history sits differently from meter-to-cash document handling, and the playbook exists to make that difference explicit before anyone falls in love with an architecture.
Model incident — the only playbook that must run with zero preparation
Every other playbook is used by someone who has had time to prepare. This one is used by a control-room lead during an abnormal system state, possibly at night, and it must therefore be short, unambiguous about authority, and rehearsed. It is also the playbook most commonly missing, because its trigger never appears in anyone's calendar.
Vendor — buy the evidence, not just the model
The failure mode is a capable product with no documentation you could show a regulator and no exit that leaves you with the decision. The playbook's job is to make the evidence pack a procurement requirement rather than a post-contract request, and to establish an out-of-hours supplier route before the first incident rather than during it.
Workforce change — a decision-rights question before an engagement one
The playbook covers timing and rights only: consultation is an input to the gate, not a communication after it, and the paper records the representatives' position whether or not it was accepted. How that conversation is actually held across an estate — the site-by-site engagement — is the subject of the AI roadshow page, and the two are designed to be used together.
Stop — the playbook that changes behaviour before it is ever used
Agreeing a stop condition before funding removes the political cost of being wrong later, which is why sponsors who have one propose bolder use cases. Where no stop route exists, use cases are defended long past the evidence, and the next proposal is funded against the memory of the last one that would not die.
Decision rights: who actually decides, and who must be asked
One accountable name per playbook, a consulted list agreed in advance, and a routing rule that keeps small decisions away from the executive committee.
Decision rights fail on the accountable column, almost never on the responsible one. Utilities are good at saying who will do the work and poor at saying who owns the outcome, so AI decisions get attributed to forums — a steering committee, a digital board, a risk panel — which is precisely the attribution that cannot be reconstructed eighteen months later when a regulator, an insurer or an internal auditor asks how a decision was reached and on what basis.
| Playbook | Accountable — one name | Responsible — does the work | Consulted — absence invalidates the decision | Informed — told after |
|---|---|---|---|---|
| Funding | Chief financial officer, or the regulated-business finance director | Use-case sponsor with the transformation or data lead | Regulatory affairs (TOTEX and price-control treatment); the operational director whose KPI is claimed | Executive committee; internal audit |
| Build versus buy | Chief information officer or head of digital | Enterprise architecture with the use-case sponsor | The system owner for the EMS, ADMS, CMMS or AMI touched; information security; procurement | Finance; the affected operational directorate |
| Model incident | Head of control, or the on-duty system operations manager out of hours | The model owner and the control-room shift lead | Safety; cyber-security duty officer; regulatory affairs when a reportable threshold may be crossed | Executive committee at next working start; the supplier if theirs |
| Vendor | Chief procurement officer, jointly with the CIO for anything touching operational systems | Procurement lead with the model owner | Information security; data protection; legal; the operational system owner | Finance; the AI gate; the supplier-management function |
| Workforce change | The operational director whose people are affected | HR business partner with the use-case sponsor | Employee representatives; safety; the site or depot leadership affected | Executive committee; the AI gate |
| Stop | The same executive who was accountable for funding it | The use-case owner with the evaluation lead | The operational owner relying on the output; finance; anyone whose process consumes it | Executive committee; the gate, as a threshold lesson |
The single most useful line in that table is the last one: the executive accountable for stopping a use case is the same one who was accountable for funding it. Splitting those two creates the pathology every utility recognises — the sponsor advocates, an independent reviewer recommends closure, the sponsor defends, and the decision becomes about the two people rather than the evidence. Keeping both with one name makes stopping an act of stewardship rather than a defeat, and it is the cheapest structural fix on this page.
The consulted column needs the same discipline. It is not a distribution list; it is the short set of people whose absence makes the decision invalid, and it should be provocative enough that someone occasionally objects to being on it. In a network business it almost always includes the control-room authority for anything that reaches real-time operations, the safety function for anything that changes field work, and regulatory affairs for anything whose treatment affects a price-control or rate-case position.
Routing rule: which decisions belong at which level
Plot the decision on two axes — how reversible it is, and how far it reaches into operational systems. Three of the four quadrants should never reach an executive committee, and the routing rule is what keeps the gate available for the quadrant that should.
Delegate to the sponsor
- Reversible, advisory only
- Discovery work, analyst-facing models, back-office drafting
- Route: sponsor decides, logged in the register, no gate
Delegate with a standing rule
- Reversible, but reaches operational systems
- Fields a control engineer can ignore or revert within a shift
- Route: system owner decides against a written standing rule; gate informed
Executive gate
- Hard to unwind, advisory only
- Multi-year vendor commitments, workforce changes, public positions
- Route: the playbook's accountable executive, at the gate
Gate plus rehearsal
- Hard to unwind and reaches live operations
- Automated switching support, protection-adjacent advice, ETR published to customers
- Route: the gate, and no go-live until the incident playbook has been exercised on it
Workforce decisions sit in the bottom-left quadrant and are the ones most often mis-routed, because the technology decision and the people decision are taken months apart by different executives. The playbook's contribution is narrow and specific: consultation is an input to the gate rather than a communication after it, and the paper records the employee representatives' position whether or not it was accepted. How that conversation is actually held across control centres, depots, generating stations and contact centres is a separate discipline with its own page — the AI roadshow for energy and utility leaders covers the route, the stop types and the job question. This page stops at the decision right and the timing.
The money playbooks: funding, build versus buy and vendors
Opex pilot or capex platform, where the regulatory cycle actually binds, the two axes that decide build versus buy, and the evidence pack a supplier must clear.
The funding playbook exists because in a utility the hard part is almost never whether the value is real — it is which pocket the money comes from and when the window opens. A use case with an unarguable operational case can still stall for a year because it was framed as capex four months after the capital plan closed, or as opex in a business whose opex envelope is fixed until the next price-control or rate-case period. The playbook's job is to force that question at the start, when it is cheap to answer.
The funding playbook, in order
State the problem in the estate's own units
Not 'improve fault prediction' but 'reduce unplanned SAIDI minutes on the 11 kV network in these three groups' or 'raise ETR accuracy on wind-driven faults above the threshold customers are told'. A case written in model terms cannot be funded by anyone whose budget is measured in operational terms, and it cannot later be stopped cleanly either, because nothing measurable was promised.
Decide the funding class before the amount
Opex pilot, capitalised platform, innovation allowance or price-control submission — each has a different approval path, a different consumer-recovery position and a different time to money. Choosing the class first often changes the shape of the proposal: a decision that must reach the next submission window is scoped differently from one that can be run inside a discretionary opex envelope this quarter.
Place it against the regulatory calendar
Mark the submission or rate-case dates and work backwards. Ofgem's price-control and regulatory programmes (opens in a new tab) and the equivalent state and federal processes in the United States, coordinated through FERC (opens in a new tab), set windows that no amount of internal urgency will move. A quarter's slip in an internal decision can mean a year's slip in funding.
Design the holdout before the build
Name the feeders, the region, the depot or the customer segment that stays on the current process, and agree it with the operational owner in the same paper. Without it, seasonality, a network reconfiguration or an unusually mild winter will claim the credit or take the blame, and the gate will have no way to tell.
Write the stop condition into the funding paper
A specific, measurable threshold and the number of review cycles it may be missed for before the stop playbook fires. Written before the first score, it is a fact everyone agreed to; proposed afterwards, the same sentence reads as an attack on the sponsor. This single line is the highest-leverage sentence in the library.
| Funding class | What it suits | Time to money | Regulatory treatment | What the gate should argue about |
|---|---|---|---|---|
| Discovery opex | Proving a data path exists and the problem is measurable at all | Weeks | Ordinary operating cost; rarely contested | Whether a decision would actually change if the answer were known |
| Opex pilot | One region, one asset class, one decision, with a holdout | One to two quarters | Operating cost within the existing envelope; TOTEX position noted | Whether the holdout is real and whether the operational owner has signed up to the KPI |
| Capitalised platform | Shared data, serving and monitoring layers reused by several use cases | Two to four quarters | Capitalised under the business's software policy; recovered over the asset life | Whether use cases two and three genuinely exist, or are being assumed to justify the layer |
| Innovation allowance | Genuinely uncertain work with a shareable learning output | Depends entirely on the scheme's window | Regulator-defined; usually requires published learning | Whether the business will actually adopt the result if it works, or treat it as a research artefact |
| Price control or rate case | Multi-year capability that must be recovered from consumers | One to several years | Assessed against consumer benefit and delivery governance | Whether the decision record will stand up to the regulator asking how the investment was governed |
The last row is where the playbook library pays for itself. A regulator assessing a multi-year AI capability is assessing governance as much as engineering: who decided, against what criteria, with what evidence, and what happens if it does not work. A business that can answer with a decision register is answering a different, easier question from one assembling the story retrospectively — and the assembly is always more expensive than the record would have been.
Build versus buy, on the only two axes that predict regret
Plot the decision, not the technology. Engineering appetite, vendor relationships and the availability of a demo are all irrelevant to this placement, and all three are what actually drive the choice when there is no playbook.
Buy, and integrate well
- Differentiating decision, commodity data
- Typically planning and simulation tooling
- Fight for integration and exit terms, not for features
Build — or partner with your data
- Differentiating decision on data only you hold
- Constraint forecasting, feeder-level DER behaviour, asset health on your own fleet history
- The only quadrant where building is usually right
Buy without hesitation
- Undifferentiated decision, commodity data
- Document handling, meter-to-cash exceptions, contact-centre drafting
- Building here is the most expensive habit in the sector
Buy the engine, keep the features
- Undifferentiated decision, but your data is the input
- Vegetation imagery, inspection triage, load disaggregation
- Buy the model, retain feature definitions and the right to leave with them
The bottom-right quadrant is where most utility regret is manufactured. The product is genuinely good, the data going into it is yours, and the contract quietly makes the feature definitions the supplier's — so three years later the decision cannot move without rebuilding the inputs from scratch. The vendor playbook's evidence pack exists mostly to prevent that outcome, and the single clause that matters most is the one describing what you leave with.
| Requirement | Why it exists in a utility | What a good answer looks like | Where it is checked |
|---|---|---|---|
| Model documentation | A regulator or inquiry may ask what the model does and on what it was trained | A written description of intended use, limits, known failure modes and evaluation data — not a marketing sheet | Gate paper; retained in the decision record |
| Data provenance and rights | Your AMI, SCADA and customer data may not lawfully train a general product | Explicit statement of what is used, retained, and whether it improves the vendor's other customers' models | Data protection and legal review before the gate |
| Retrain cadence and change notice | A silent model update can change operational behaviour with no change control | A stated cadence, advance notice of material changes, and the right to hold a version | Contractual AI schedule; verified at the first release |
| Failure modes and degradation behaviour | The control room needs to know what wrong looks like before it happens | Named failure modes with observable symptoms, and what the product does when inputs are stale | Model-incident exercise, before go-live |
| Exit and portability | Decisions outlive suppliers; the feature definitions are the asset | You leave with the feature definitions, the historical outputs and the labelled data you supplied | Contract; rehearsed as a desktop exit scenario |
| Incident response and out-of-hours route | Network events do not respect business hours | A named contact, a response time, and a route that works at 02:00 on a Sunday | Tested during the incident exercise, not during an incident |
| Standards alignment | Assurance and insurance questions increasingly cite named frameworks | Evidence against ISO/IEC 42001 or the NIST AI RMF functions, with scope stated honestly | Assurance review; cited in regulatory submissions |
Two of those rows are rehearsed rather than filed, and that is deliberate. A documented out-of-hours supplier route that has never been dialled is a phone number, not a capability — the same class of assumption as a fallback nobody has switched. Testing both during a scheduled exercise costs an hour and routinely finds that the number reaches a service desk with no authority to act on an operational system. EPRI's artificial intelligence work (opens in a new tab) and the shared evaluation effort behind the Open Power AI Consortium (opens in a new tab) exist in part because every utility was otherwise assembling this evidence alone, for the same suppliers.


