Redefining Technology

Energy & UtilitiesAI Adoption & Maturity Curve

AI adoption risks in energy and utilities — and how to mitigate them at every stage

AI adoption risks in energy and utilities are the ways an AI programme can damage the operation it is meant to improve — bad decisions coupled to the grid, silent model drift, a widened IT/OT attack surface, and regulatory exposure. They change class at every stage of the maturity curve, so mitigation must be rebuilt at each stage, not inherited.

Generated scene: utility control room with AI risk and mitigation overlays across grid operations displays
Energy & Utilities · AI Adoption & Maturity Curve

Key takeaways

  1. AI adoption risk in a utility does not shrink as the programme matures — it changes class. Stage 1–2 failures waste money invisibly; stage 3–4 failures reach the control room; stage 5 failures reach the network. Each stage needs its own controls, and inheriting the last stage's controls is itself a failure mode.
  2. The cheapest control on the whole curve is a model inventory. Most utilities already run more models than they can list — vendor 'AI-enabled' features inside the ADMS and APM, engineers' scoring spreadsheets — and a risk you cannot enumerate cannot be mitigated.
  3. The riskiest single transition is stage 2 to 3, when model output first couples to grid operations. Three controls make it survivable: an acceptance test on your own network data, a drilled fallback to the previous method, and drift alerting keyed to network events such as DER growth and reconfiguration.
  4. Regulators do not prohibit AI in the decision path — NERC CIP, Ofgem licence conditions and the EU AI Act all converge on the same demand: evidence. A decision log that can reconstruct any AI-influenced decision months later is the one artefact that serves all three at once.
  5. Mitigation posture follows two questions, not one: how severe is a wrong output on the grid, and how quickly can the action be unwound? Severe-and-slow decisions stay advisory forever; contained-and-reversible ones are where autonomy is earned first.

Abbreviations used on this page

SCADA
Supervisory control and data acquisition
EMS
Energy management system (transmission control)
ADMS
Advanced distribution management system
OMS
Outage management system
AMI
Advanced metering infrastructure (smart meters and head-end)
APM
Asset performance management (transformer and line health)
DER
Distributed energy resources (rooftop solar, batteries, EVs)
DERMS
Distributed energy resource management system
OT
Operational technology — the control-system side of the estate
NERC CIP
NERC Critical Infrastructure Protection standards
SAIDI
System average interruption duration index
AI RMF
NIST AI Risk Management Framework

Free · 8 questions · ~3 minutes

Score your risk controls against the curve

Eight questions, one at a time, about three minutes — each scoring a control that exists or does not, never an ambition. Answer them and we build your personalised risk report: your stage on the curve, your score on each of the four control dimensions, and the specific unmitigated risk most likely to bite next. It lands in your inbox.

0 of 8 answered

Question 1 of 8Model & data risk controls

How complete is your inventory of models influencing operational decisions — including vendor features inside the ADMS, OMS or APM, and engineers' scoring spreadsheets?

Every other control operates per model. An incomplete inventory means unmitigated risk you cannot even locate.

How the score maps to a stage
  • 05 — Stage 1, Unmapped. AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
  • 611 — Stage 2, Pilot-exposed. Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
  • 1216 — Stage 3, Grid-coupled. Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
  • 1721 — Stage 4, Portfolio-managed. Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
  • 2224 — Stage 5, Bounded autonomy. Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.

What AI adoption risks are in energy and utilities — and why they change class

A definition, the five risk families, and the property that organises this whole page: as a programme matures, its risks do not shrink — they change class.

AI adoption risks in energy and utilities are the ways an AI programme can harm the operation it serves: a wrong output coupled to a grid decision, a model silently drifting away from the network it describes, operational data leaking across the IT/OT boundary, a regulator asking for a reconstruction nobody can produce, and a control room whose trust in the system is miscalibrated in either direction. They are distinct from ordinary project risk — a failed AI project wastes money, whereas a badly governed successful one can mis-stage storm crews, defer the wrong transformer maintenance or hold a voltage envelope that no longer matches the network.

The organising property is that these risks change class as the programme matures. At stages 1–2 the dominant exposures are invisible and cheap: unregistered models, ungoverned extracts, inherited vendor claims. From stage 3 the exposure is operational: output couples to the EMS, ADMS and OMS, and drift, missing fallbacks and alarm fatigue become the live risks. By stages 4–5 the exposure is systemic and regulatory: common-mode failure across a portfolio, concentration in shared feeds and vendors, envelope validity, and the reconstruction demands of NERC CIP (opens in a new tab), Ofgem (opens in a new tab) licence obligations and the EU AI Act (opens in a new tab), which classes AI safety components in critical infrastructure — energy explicitly included — as high-risk systems. The controls that retire stage-2 risks do nothing against stage-4 ones, which is why mitigation must be rebuilt at each stage rather than inherited.

  • Model and data risk

    Wrong, drifted or unvalidated model output, and the poor-quality or leaked data behind it. In a utility this family lives in specific places: SCADA and historian extracts feeding pilots, AMI interval data in vendor tenancies, health scores inside the APM, forecasts inside the EMS.

  • Operational and safety risk

    The coupling of model output to grid-facing actions — switching recommendations, crew pre-staging, maintenance deferrals, voltage control. Governed by two questions: how severe is a wrong output, and how fast can the action be unwound?

  • Security risk

    Every model, feed and vendor tenancy is new attack surface across the IT/OT boundary. The IEA's Energy and AI analysis records that cyberattacks on energy utilities tripled in four years and grew more sophisticated because of AI — the same technology sits on both sides of this ledger.

  • Regulatory and compliance risk

    Not prohibition but evidence: reliability standards, licence conditions and the EU AI Act's high-risk regime all demand that AI-influenced decisions be explainable, monitored and reconstructable. The risk is being unable to show your workings.

  • Organisational risk

    Control-room trust miscalibration, operator deskilling as advisory systems take load, and automation complacency once things work — the slowest-moving family, and the one that converts a technical failure into an operational incident.

Consequence of an uncontrolled failure, along the curve

The curve every risk on this page is indexed against. As output couples to the grid (stage 3) and then to autonomous action (stage 5), the consequence of the same wrong output rises steeply — which is why controls must be rebuilt at each stage. The stages themselves are named by the risk class that dominates them.

Consequence if a control fails by stage

  • Stage 1 · Unmapped — 24% of operators. AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
  • Stage 2 · Pilot-exposed — 37% of operators. Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
  • Stage 3 · Grid-coupled — 25% of operators. Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
  • Stage 4 · Portfolio-managed — 11% of operators. Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
  • Stage 5 · Bounded autonomy — 3% of operators. Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.

Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with IEA Energy and AI grid-application analysis.

How a wrong output travels at each stage

The same failure — a model producing a wrong number — lands in three different worlds depending on stage. At stages 1–2 it dies in a spreadsheet, invisibly; at stages 3–4 it reaches the control room through governed, logged steps; at stage 5 it reaches the network unless the envelope catches it. The controls at each stage exist to match the blast radius.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • At stages 1–2, operational data leaves the estate as ad hoc extracts, feeds unregistered models, and returns as decks and health indices that quietly inform real deferrals. Nothing is logged, so a wrong output costs little today and cannot be reconstructed later — the exposure is invisibility itself.
  • At stages 3–4, registered flows from the historian and AMI feed acceptance-tested, drift-watched models whose output lands in advisory fields inside the OMS or ADMS. The control room approves or overrides each recommendation with the previous method one switch away, and every step is logged.
  • At stage 5, the accumulated override log justifies a versioned operating envelope inside which enumerated, reversible actions — volt/VAR optimisation, DER dispatch — execute automatically. Anything outside bounds escalates to a person, and the escalation rate is charted as the leading indicator of envelope decay.
Step-by-step insights
The unlogged edge is the most dangerous line on this diagram
The dashed edge from 'deck or health index' to 'engineer may act' is where stage-1 risk actually lives. The failure is not that the engineer acts on the model — it is that no record exists that a model was in the loop at all. When the deferred transformer fails eighteen months later, the investigation finds a decision with no decision path. Every control later on the curve — registers, logs, envelopes — exists to make that edge impossible to draw.
Why the boundary node changes everything
The difference between lane 0's 'exported SCADA history' and lane 1's 'historian & AMI feeds' is not the data — it is the register. A classified, registered flow can be technically enforced, monitored for freshness, and produced in an audit. An emailed CSV can do none of those things, and its riskiest property is that it works fine, normalising a path that widens with every pilot. Utilities that build the sanctioned path early never accumulate the shadow paths that take years to find again.
The advisory field is a risk instrument, not a half-measure
Writing model output into an advisory field the operator already sees — rather than automating, and rather than a separate dashboard — does double duty. It bounds the blast radius, because a person with network context checks every action. And it generates the override log: every accept and reject, with circumstances, is evidence about where the model can and cannot be trusted. That log is the only honest basis for the stage-5 envelope. Skip the advisory years and the envelope is guesswork wearing a signature.
The envelope is a protection-setting, culturally
Utilities already run the exact governance an AI operating envelope needs — for protection settings. Proposed with studies, reviewed by someone who did not propose them, versioned, revisited when the network changes. Treating the AI envelope as a member of that family, on the same review calendar, imports thirty years of discipline for free. Treating it as software configuration — tunable in a settings screen — is how automation quietly outgrows its evidence.
Escalation rate: the one number that watches the watchers
Every stage-5 mechanism can decay silently except one: the escalation rate. When the share of automated decisions falling outside bounds trends upward, something real has changed — new DER connections, a reconfiguration, drifted inputs — and the envelope's claim about the world is aging. Charting it beside SAIDI on the operational dashboard turns envelope decay from a latent audit finding into a routine operational signal that triggers review before an excursion forces one.

The five stages of AI risk in a utility, in detail

Each stage named by the risk class that dominates it. For each: what it looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.

The five stages below are risk postures, not capability levels — each is named for the risk class that dominates it, and the honest question at each is not 'what can we build?' but 'what can currently go wrong, and what control retires it?'. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own estate this week, and the anti-pattern is the mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Unmapped

24% of operators sit here

AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.

Stage 1 is not the absence of AI — it is the absence of a map. Most utilities at this stage are already consuming model output daily without calling it that: the APM suite scores transformer health with a vendor's proprietary model, the OMS vendor has quietly enabled a predictive module in the last upgrade, and a protection engineer maintains a dissolved-gas-analysis spreadsheet whose thresholds came from a machine-learned fit. None of it is on any register, so none of it has an owner, a validation record or a fallback.

The distinguishing risk is that exposure cannot be enumerated. Ask the leadership team how many models influence operational decisions and the answer is a guess that is usually out by a factor of three. That matters because every later control — validation, drift monitoring, audit evidence — operates per model. A risk framework applied to the four models you know about does nothing for the nine you do not.

This stage is cheap to leave and dangerous to stay in, because unmapped risk compounds silently. The maintenance deferral informed by an unvalidated health score does not fail loudly; it shows up eighteen months later as a transformer failure with no record that a model was ever in the loop — which is precisely the reconstruction a regulator will ask for.

In practice

The transformer spreadsheet nobody registered

During an asset-management review, a distribution utility found that maintenance deferrals on 40-year-old transformers were being informed by a health-index spreadsheet one engineer had built from dissolved-gas-analysis exports, with weightings fitted years earlier and never revalidated. It had quietly become the deciding input on capital deferrals worth millions. Nothing was wrong with the engineer's work — what was wrong was that no one else knew the decision path existed.

What it looks like

  • No inventory of models influencing operational decisions
  • Vendor 'AI-enabled' features active inside the ADMS, APM or OMS, unregistered
  • Engineers run private scoring spreadsheets that inform real deferrals
  • AI risk appears on no risk register, so it has no owner

Diagnostic signals you can check this week

  • Ask for the list of models influencing operational decisions — time how long it takes to produce, and who has to be asked
  • Check the last ADMS, OMS and APM vendor release notes for AI features enabled by default
  • Ask three engineers whether any spreadsheet of theirs feeds a real decision; count the surprises
  • Look for AI anywhere on the corporate risk register — absence is the stage-1 signature

Anti-pattern · Writing the AI policy before the AI inventory

The instinctive first move is a policy document — principles, ethics language, an approval workflow for future AI. It feels like control and governs nothing, because it applies to the AI the organisation plans to adopt while saying nothing about the AI already running unregistered in the estate. Inventory first, policy second: a one-page register of what exists, who relies on it and what it decides governs more real risk than any principles document, and the policy written afterwards will describe reality rather than intention.

What holds you here

You cannot mitigate what you cannot list — every control operates per model, and the model count is unknown.

Highest-leverage next move

Build the model inventory: every model, vendor feature and scoring spreadsheet that influences an operational decision, each with a named owner.

Cost of leaving

Effort
4–8 weeks
Team
One OT engineer and one analyst, part-time, with vendor-contract access
Risk
Low — discovery work, nothing in production changes
To next stage
1–2 months

If this is you, the next step is

A short engagement: we sweep the estate — vendor features, spreadsheets, trials — and hand you the register.

Run a shadow-AI discovery

Stage 2

Pilot-exposed

37% of operators sit here

Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.

Stage 2 is where risk becomes visible but controls remain paper. The pilots are real: a wind or load forecasting trial, an asset-health proof of value, an outage-prediction bake-off. To run them, operational data starts moving — historian extracts emailed to a vendor's data scientist, AMI interval data loaded into a cloud tenancy under a contract clause nobody technical has read. The controls that exist are legal ones, and legal controls do not stop a feed, classify an extract or notice that a export contained customer data.

The second stage-2 risk is claim inheritance. Vendor models arrive with accuracy figures earned on another operator's network — a different climate, a different DER penetration, a different mix of overhead and underground circuits. A storm-outage model trained on a coastal utility's history transfers badly to an inland one, and the difference does not surface in a demo. Utilities that skip acceptance testing on their own network data discover the gap after the model is coupled to operations, which is the most expensive possible moment.

What makes this stage treacherous is that its failures are quiet and its successes are loud. A leaked extract or an over-fitted demo costs nothing visible this quarter, while a good-looking pilot earns headlines and budget. The programme accelerates toward stage 3 carrying unexamined data paths and untested claims — which is exactly the cargo you do not want at the moment output first couples to the grid.

In practice

The forecasting pilot that emailed the historian

A generation forecasting pilot at a mid-size utility ran for five months on weekly historian exports a graduate engineer emailed to the vendor as CSV attachments. The pilot's accuracy was genuinely good. The exports, it later turned out, included tags from units subject to critical-infrastructure information rules, and no record existed of which extracts had been sent, when, or what the vendor had retained. The pilot succeeded; the data path it normalised became the finding in the next security review.

What it looks like

  • Pilots consume real SCADA, historian or AMI extracts
  • Data-sharing terms with vendors are contractual, not technical
  • Vendor accuracy claims come from someone else's network
  • Success criteria are model metrics, not operational deltas

Diagnostic signals you can check this week

  • Trace one pilot's data path end to end — every hop, every retention point; note where the map runs out
  • Ask who technically (not contractually) could stop data leaving the estate today
  • Ask the vendor for accuracy figures on a network with your DER penetration and circuit mix — watch for the pivot to the global figure
  • Check whether any pilot has a defined operational holdout for later attribution; at stage 2 the honest answer is usually no

Anti-pattern · Banning cloud AI instead of building the boundary

When a security team discovers stage-2 data paths, the reflex is prohibition: no operational data leaves the estate, no cloud AI services, pilots frozen. The estate then reliably routes around the ban — engineers anonymise poorly and email anyway, vendors run 'demos' on data brought to workshops — and the organisation ends up with the same exposure minus the visibility. The durable fix is a sanctioned path: a classified extract process, an approved tenancy, a data-flow register. Make the safe route cheaper than the workaround and the workaround disappears.

What holds you here

Data paths and vendor claims are both ungoverned, and the pilot's momentum is pushing both toward the grid.

Highest-leverage next move

Stand up the two stage-2 controls: a classified extract process across the IT/OT boundary, and an acceptance test that revalidates every vendor claim on your own network data against a holdout.

Cost of leaving

Effort
3–6 months
Team
One OT security engineer, one data engineer, procurement support
Risk
Medium — the boundary work must not stall the pilots it governs, or it will be bypassed
To next stage
3–6 months

If this is you, the next step is

We map every pilot's data path across the IT/OT boundary and hand you the classified-extract process.

Get the pilot boundary reviewed

Stage 3

Grid-coupled

25% of operators sit here

Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.

Stage 3 is the moment the risk class changes physically: a wrong output stops being a wasted analysis and starts being a wrong pre-staging decision, a deferred maintenance visit, a mis-set switching plan. The programme's risk profile now moves with the network. A distribution network is not a stationary system — DER connections grow monthly, reconfigurations change flow patterns, an AMI rollout changes the very telemetry the model was trained on — and every one of those changes is a drift event that arrives on the network's schedule, not the data science team's.

The controls that matter here are operational, not analytical. Drift monitoring keyed to network events, not just statistical thresholds. Freshness alerting on every feed the model consumes, because a historian interface that silently stops updating is more dangerous than one that visibly fails. And above all a fallback: the previous method — the deterministic storm matrix, the seasonal load profile, the manual health index — kept alive, one switch away, and drilled. An undrilled fallback is a document, and documents do not operate networks.

The subtle stage-3 risk is human: trust miscalibration in the control room. Operators either over-trust the new field (accepting recommendations without the scrutiny they would apply to a colleague) or under-trust it (quietly ignoring it, so the programme's value silently evaporates while its risk remains). Both are measurable — override rates, time-to-accept — and both respond to the same treatment: showing operators the model's failure cases and its bounds, not just its wins.

In practice

The storm model that drifted with the rooftops

A distribution utility coupled an outage-prediction model to its OMS to pre-stage crews ahead of storms. It worked well for two seasons. Over the following eighteen months, rooftop solar penetration in two districts roughly doubled, changing fault signatures and feeder loading patterns the model had learned. Pre-staging accuracy decayed slowly enough that each miss looked like weather luck. Only when a crew spent a storm night positioned sixty kilometres from the actual damage did anyone check the model's error trend against the DER connection register — where the drift had been visible for a year.

What it looks like

  • Model output lands in the EMS, ADMS or OMS, at least in advisory fields
  • Control-room and field decisions are influenced daily
  • Drift and freshness monitoring exist for the coupled models
  • A fallback to the previous method exists — though it may never have been drilled

Diagnostic signals you can check this week

  • Pick one coupled model and ask when its fallback was last exercised — a date and a log entry, not an assurance
  • Check whether drift alerts reference network events (DER growth, reconfiguration, AMI changes) or only statistical thresholds
  • Pull the override rate on advisory recommendations; both near-zero and near-total are alarms
  • Ask what happens to the model's inputs when the historian interface is patched — who is told, and how

Anti-pattern · Improving accuracy while the fallback rots

Once output is coupled, teams instinctively invest where they are strongest: model accuracy. Meanwhile the fallback — the old storm matrix, the manual dispatch process — quietly decays: the person who ran it retires, the spreadsheet stops being updated, the switch is never tested. The estate ends up more dependent on the model precisely as its safety net disappears. Accuracy work is worth doing only after the fallback is drilled on a calendar, because the day you need the old method is by definition the day the model has failed.

What holds you here

Controls are per-model and hand-built, so every new coupled model re-raises the same risks with none of the previous answers.

Highest-leverage next move

Standardise the stage-3 control set — event-keyed drift alerting, freshness SLAs, drilled fallbacks, decision logging — so the next coupled model inherits controls instead of re-inventing them.

Cost of leaving

Effort
6–12 months
Team
One integration engineer, one ML engineer, a named control-room owner, OT security review
Risk
Medium-high — first grid-coupled writes need rollback paths and change-board approval
To next stage
9–18 months

If this is you, the next step is

We design and run the reversion exercise on a quiet day — and leave you the drill calendar.

Drill your fallback with us

Stage 4

Portfolio-managed

11% of operators sit here

Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.

Stage 4 is where individual-model risk is largely tamed and a new class quietly replaces it: systemic risk. A portfolio of eight models serving load forecasting, asset health, outage prediction and vegetation management does not carry eight independent risks — it carries shared ones. Most of the portfolio consumes the same weather feed, the same historian, the same feature pipeline. A single upstream failure now degrades five models at once, in correlated ways, during exactly the weather event when all five matter most. The failure mode nobody designed is the one the portfolio inherited from its own efficiency.

Concentration risk arrives the same way. Standardising on one vendor's platform, one cloud tenancy or one forecasting supplier is operationally sensible and creates a single point whose commercial failure, security compromise or model regression propagates across the estate. The mitigation is not to abandon standardisation but to know precisely what depends on what — a dependency register for models, feeds and vendors — and to have exercised the loss of each critical dependency the way transmission planners exercise the loss of a line: as a contingency with a rehearsed response.

The other stage-4 discipline is portfolio-level evidence. By now the regulator conversation has changed: it is no longer 'do you use AI?' but 'show us how you manage it'. NERC CIP audits, Ofgem's licence-condition conversations and — for European operators — EU AI Act conformity all want the same artefacts: the inventory, the review cadence, the incident record, the reconstruction capability. At stage 4 those artefacts either fall out of the framework automatically, or the framework is theatre.

In practice

The morning the weather feed went down

A utility running six operational models discovered its concentration the morning its commercial weather provider had an outage. Load forecasting degraded to seasonal profiles as designed. What nobody had mapped was that vegetation risk scoring, storm pre-staging and two asset-health models consumed derived features from the same feed — all four degraded simultaneously, and the control room's fallback procedures assumed models failed one at a time. Nothing broke on the network that day. The dependency map got built the same week.

What it looks like

  • A model risk function reviews every operational model on a cadence
  • Controls are inherited from a standard set, not rebuilt per model
  • Dependencies between models and shared feeds are mapped
  • Portfolio-level metrics exist: coverage, drift status, drill currency

Diagnostic signals you can check this week

  • Ask for the dependency map: which models share which feeds, features and vendors — and when the loss of the biggest shared dependency was last exercised
  • Check whether the model risk review is a real gate: find one model change it delayed or refused
  • Count models per critical upstream feed; more than three on one feed with no exercised contingency is unmanaged concentration
  • Ask how long it takes to produce the full portfolio evidence pack for an audit — days is stage 4, weeks is stage 3 with paperwork

Anti-pattern · Scaling the framework by adding paperwork

As the portfolio grows, the risk function's instinct is to add process weight: longer review templates, more sign-offs, quarterly attestation spreadsheets. Teams respond rationally — they route around it, classifying new work as 'analytics' rather than 'models' to avoid the gate, and the inventory silently diverges from reality again, which is stage 1 wearing a stage-4 badge. The framework scales through automation, not attestation: controls inherited by default from the platform, evidence generated as a by-product of operation, and a review gate that is fast precisely because the artefacts already exist.

What holds you here

Every decision still routes through a person, so the portfolio's value is capped by control-room attention — but crossing to autonomy safely demands evidence thresholds most frameworks have not yet defined.

Highest-leverage next move

Define, per candidate decision, the operating envelope and evidence bar under which autonomous execution would be acceptable — before any automation is switched on.

Cost of leaving

Effort
12–24 months
Team
Model risk owner, platform engineers, OT security, internal audit partnership
Risk
Medium — the framework must earn adoption or it will be evaded, recreating unmapped risk
To next stage
18+ months

If this is you, the next step is

We run a loss-of-feed contingency exercise across your portfolio and report what actually degrades.

Stress-test your dependency map

Stage 5

Bounded autonomy

3% of operators sit here

Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.

Stage 5 in a utility is narrower than the phrase 'autonomous grid' suggests, and correctly so. It is an enumerated set of contained, reversible actions — volt/VAR optimisation within band, DER curtailment inside connection agreements, battery dispatch within market and thermal limits — executing inside envelopes an engineer signed, with everything else escalating to a person. Switching plans, protection settings and safety-adjacent decisions stay human forever, not because models cannot rank options but because the consequence-reversibility calculus says advisory is where they belong.

The risk that defines this stage is envelope validity. An envelope is a claim about the world — these bounds are safe given this network — and the network keeps changing underneath it. New DER connections, a reconfiguration, a new interconnector: each quietly erodes the assumptions the envelope encodes. The leading indicator is the escalation rate. When the share of decisions falling outside bounds trends up, the world has moved and the envelope needs review before an incident forces one. Utilities already run this discipline for protection settings; the transfer of that habit to AI envelopes is the whole trick.

The second defining risk is complacency — months of correct automated operation recalibrate humans. Operators stop shadowing the system, the fallback drill slips, the reversion switch goes untested through two reorganisations. The mitigations are deliberately artificial: scheduled hours operating on the fallback method, injected exercise scenarios, and treating a missed drill with the seriousness of a missed relay test. Stage 5 is not a destination where risk work ends; it is a permanent operating discipline, and the utilities that hold it treat the envelope, the drill calendar and the decision log as living operational artefacts.

In practice

The envelope that aged out under new connections

A network operator ran automated voltage optimisation on a group of feeders inside an envelope validated against the connection register at go-live. Over a year, behind-the-meter batteries and two community solar sites shifted the feeders' behaviour; the escalation rate crept from one action in fifty to one in nine. Because escalations were charted as an operational KPI, the trend triggered an envelope review — recomputed bounds, a re-signed policy version — months before any customer saw a voltage excursion. The near-miss that never happened is what stage-5 discipline looks like from inside.

What it looks like

  • Enumerated decisions execute automatically inside explicit bounds
  • The operating envelope is versioned and reviewed like a protection setting
  • Escalation rate is monitored as a leading indicator of envelope decay
  • Kill switches and reversion paths are drilled, with dated log entries

Diagnostic signals you can check this week

  • Ask for the envelope's version history and who signed the last change — a settings screen with no history is the anti-signature
  • Check the escalation-rate trend and whether anyone owns reviewing it
  • Find the last dated kill-switch drill; more than six months old means the switch is theoretical
  • Ask an auditor's question: reconstruct one automated action from last quarter — inputs, model version, envelope version, outcome — and time the answer

Anti-pattern · Letting the envelope drift to match the model

When an automated system keeps escalating, the path of least resistance is to widen the bounds — each widening individually defensible, none re-deriving the envelope from network studies. Two years later the envelope reflects what the model wants to do rather than what the network can tolerate, and no one can say which version was last independently validated. Envelope changes need the discipline of protection-setting changes: proposed with evidence, reviewed by someone who did not propose them, versioned, and periodically re-derived from scratch against the current network.

What holds you here

Sustaining autonomy is a standing evidence-and-drill discipline — the constraint is envelope governance and regulatory reconstruction, not engineering.

Highest-leverage next move

Put envelope reviews, escalation-rate monitoring and kill-switch drills on the same operational calendar as protection-setting reviews — permanently.

Cost of leaving

Effort
Continuous
Team
Platform team, control-room ownership, a standing risk forum with OT and compliance
Risk
Concentrated — low-frequency, high-consequence, regulatory in nature

If this is you, the next step is

We reconstruct three real automated actions end to end and stress-test the envelope, trail and reversion.

Audit an autonomous decision path

Where energy and utilities operators actually sit today

The distribution across the risk curve, why the crowd at stage 2 matters, and what the external research says about the threat side of the ledger.

Most energy and utilities operators sit at stages 1 and 2 — running pilots on real operational data with controls that are contractual rather than technical, above an estate carrying more unregistered models than anyone has listed. The crowding matters because the sector's momentum is real: system operators are publishing digitalisation strategies, vendors are shipping AI features by default inside the ADMS and APM layer, and the physics of a decarbonising grid — more DERs, more variability, tighter margins — pulls forecasting and optimisation models toward the control room. The risk work queues up exactly where the crowd is thinnest: in the controls that make coupling survivable.

Distribution of energy and utilities operators across the risk stages

Illustrative distribution, synthesised from IEA, EPRI and Eurelectric adoption research — charted to show shape, not to report a survey. Stage 2 is the mode: pilots on real data, controls on paper. The drop from stage 2 to stage 3 is where risk work, not model work, is the gate.

Share of operators (illustrative)

  • 24% — 1 · Unmapped
  • 37% — 2 · Pilot-exposed (the crowd)
  • 25% — 3 · Grid-coupled
  • 11% — 4 · Portfolio-managed
  • 3% — 5 · Bounded autonomy

Source: Illustrative; synthesised from IEA, EPRI and Eurelectric research

Cyberattacks on energy utilities have tripled in the past four years and have become more sophisticated because of AI. At the same time, AI is becoming a critical tool to defend against them.

The external research is unusually consistent about both sides of the ledger. The IEA's Energy and AI report (opens in a new tab) catalogues the upside — AI-based fault detection reducing outage durations by 30–50%, better renewables forecasting cutting curtailment — while recording the tripling of attacks on the same page. EPRI's AI research programme (opens in a new tab) and Eurelectric (opens in a new tab) track the sector's adoption and its governance gap, and cross-industry work such as McKinsey's electric power insights (opens in a new tab) finds the familiar pattern of experimentation running ahead of impact. The reading that matters for this page: the sector's problem is not enthusiasm and not evidence of value — it is that controls lag coupling, and the gap between those two lines is exactly where incidents live.

The risk register: nine risks, stage-indexed, with the control that retires each

The centrepiece. Each risk enters the curve early, peaks later, and bites in a specific system — and each has a named control and a named piece of evidence that proves the control exists.

The register below is the working core of this page: the nine risks that recur across utility AI programmes, indexed by the stage where each enters, the stage where it peaks, and the operational system where it bites. Two disciplines make a register like this useful rather than decorative. First, every risk carries a control that is buildable — not 'increase awareness' but 'acceptance test on own network data against a holdout'. Second, every control carries an evidence artefact — the thing an auditor, a regulator or your own risk committee can hold. A control without evidence is an intention; utilities run on evidence.

RiskEntersPeaksWhere it bitesMitigating controlEvidence it is retired
Shadow AI — unregistered models and vendor features13APM, ADMS, OMS vendor modules; engineers' spreadsheetsModel inventory with owners, tied to procurement and change controlRegister reviewed on a cadence; vendor AI-disclosure clause in contracts
Operational data leaking across the IT/OT boundary22Historian and AMI extracts to vendor tenanciesSanctioned classified-extract process; registered, technically enforced flowsData-flow register; boundary rule that can actually stop a transfer
Inherited vendor claims — accuracy earned on someone else's network23Forecasting, asset health, outage predictionAcceptance test on own network data against a holdout, before couplingSigned acceptance report with your DER mix and circuit types in it
Silent model drift after network change34EMS load forecast, OMS storm models, APM scoresDrift alerting keyed to network events — DER growth, reconfiguration, AMI changesAlert log showing fires and responses; retrain cadence adhered to
Missing or rotten fallback33Control room, storm response, dispatchPrevious method maintained one switch away, drilled on a calendarDated drill log entries — not an assurance that reversion 'would work'
Control-room trust miscalibration and alarm fatigue34Advisory fields in EMS/ADMS/OMSAlert budgets; operators shown failure cases and bounds, not just winsOverride-rate trend charted and reviewed — neither ~0% nor ~100%
Common-mode failure across the portfolio45Shared weather feeds, features, historian, one cloud tenancyDependency register plus exercised loss-of-feed contingenciesContingency drill report: what degraded, what the control room did
Undocumented AI-influenced decisions35Everywhere output meets a decisionDecision log linking inputs, model version, output and human actionA timed reconstruction: any decision, months later, in hours
Automation beyond the envelope's validity55Volt/VAR, DER dispatch, battery schedulingVersioned envelope reviewed like a protection setting; escalation-rate monitoring; drilled kill switchEnvelope version history with signatures; escalation trend on the ops dashboard
The stage-indexed AI risk register for energy and utilities. 'Enters' is the stage where the risk first exists; 'peaks' is where it does the most damage — note how often those differ, which is why controls built at entry are cheap and controls built at peak are incident responses.

Read the 'enters' and 'peaks' columns together and the register's main lesson appears: almost every risk is cheapest to retire one stage before it peaks. The inventory that takes four weeks at stage 1 takes a forensic quarter at stage 3, because by then the unregistered models are load-bearing. The acceptance test that costs a fortnight at stage 2 costs an incident at stage 3. This is also why 'we will add governance when we scale' is the sector's most expensive sentence — the register's risks do not wait for scale, they compound under it. Who holds the pen on each control — which decisions need an accountable owner rather than a committee — is its own discipline, covered properly in the AI leadership playbooks for energy utilities; this page stays on the controls themselves.

Mitigation posture: grid consequence × reversibility

Where each decision type belongs, before any model quality argument. Plot the wrong-output consequence against how fast the action can be unwound; the quadrant sets the control intensity. Nothing in the top-left ever automates, however good the model gets.

Advisory forever

  • Switching plans, protection settings, safety-adjacent calls
  • AI ranks options; a person decides, always
  • Control: full decision log plus human accountability

Human-approved

  • Storm crew pre-staging, dispatch recommendations, load transfers
  • Every action approved in the control room, logged
  • Control: drilled fallback plus override-rate monitoring

Batch with review

  • Maintenance scheduling, connection studies, capital deferral scoring
  • Output reviewed in planning cycles, not in real time
  • Control: acceptance testing plus periodic revalidation

Earn autonomy here

  • Volt/VAR in band, DER curtailment within agreements, meter-data cleansing
  • Contained, reversible — the honest stage-5 candidates
  • Control: versioned envelope, escalation monitoring, kill switch
Consequence of a wrong output — top: Severe on the grid, bottom: Contained
Reversibility of the action — left: Hard to unwind, right: Instantly reversible

The matrix is the page's second spine because it answers the question the register raises: how much control is enough? The answer is positional, not universal. A meter-data cleansing model needs an envelope and a log; a switching adviser needs a human forever, and the correct response to 'the model is now accurate enough to automate switching' is that accuracy was never the constraint — consequence and reversibility were. Utilities that write this matrix down early spend their governance budget where the grid actually needs it, instead of spreading uniform process weight across decisions that differ by three orders of magnitude in consequence.

What managed adoption looks like in public

Three publicly reported programmes, read against the risk curve. None is an Atomic Loops engagement — each links to the organisation's own published material, and the card images are generated industry scenes, not operator photography.

The clearest public evidence for controls-first adoption is in what system operators and utilities chose to publish alongside their AI capability. In each case below, the organisation's own material pairs the capability with its governing frame — a digitalisation strategy, a partnership structure, a shared validation effort — which is precisely the pairing the risk register above formalises.

Three programmes read against the risk curve

Outcomes as reported by the organisations themselves — verify against the linked source before reusing; we have not independently audited them. Card images are generated industry scenes and do not depict the organisations' facilities.

Generated scene: national electricity system control centre with renewables forecasting displaysNESO (National Energy System Operator)GB electricity system operator · control room & balancing23
Challenge
Bringing AI into the environment with the least tolerance for a wrong output — the national control room — where forecasting errors move balancing costs and, ultimately, system security.
Approach
NESO's published AI communications describe building an AI Centre of Excellence to pool expertise and strengthen the security and reliability of the network, with machine-learning forecasting supporting — not replacing — control-room decision-making, and the governing frame published openly in its Digitalisation Strategy and Action Plan (June 2025).
Reported outcome
As reported by NESO, AI and machine learning now support forecasting and control-room insight within a published digitalisation strategy, with the capability and its governance developed and communicated together.
What it shows about the curveThe stage-2→3 transition done in the open: the control frame ships with the capability, and advisory mode in the control room is the deliberate posture, not a stepping stone skipped for speed.

NESO — the ESO and artificial intelligence (opens in a new tab)

Generated scene: distribution grid operations with smart-grid analytics overlaysDuke EnergyUS investor-owned utility · six-state service area23
Challenge
Scaling grid analytics and smart-grid capability across a very large distribution estate requires cloud-scale tooling — which puts operational grid data and grid-adjacent decisions into a partner ecosystem, exactly the boundary stage 2 must govern.
Approach
Duke Energy publicly announced a collaboration with AWS to develop smart grid solutions supporting its clean-energy transition — structuring the cloud partnership as a named, governed programme rather than accumulating ad hoc vendor arrangements team by team.
Reported outcome
As reported in Duke Energy's own release, the collaboration develops smart-grid solutions to better serve customers and support the utility's clean-energy transition, with the partnership's scope stated publicly.
What it shows about the curvePartner-based capability still leaves the operator holding grid accountability. A named programme with published scope is itself a control — it replaces the invisible, team-by-team data paths that define pilot-exposed estates.

Duke Energy newsroom — AWS collaboration (opens in a new tab)

Generated scene: grid operations centre with algorithmic accountability overlays — illustrative, not an EPRI facilityEPRI — Open Power AI ConsortiumIndustry research consortium · utilities and technology partners23
Challenge
Every utility validating AI models alone repeats the same model-risk work — acceptance testing, benchmarking, domain evaluation — at every operator, which guarantees the sector's controls lag its adoption.
Approach
EPRI convened the Open Power AI Consortium to develop and benchmark domain-specific AI models and evaluation approaches for the power sector, pooling utilities and technology partners so validation effort is shared rather than duplicated.
Reported outcome
As reported by EPRI, the consortium operates openly with published aims around domain-specific models and benchmarks for the electricity sector.
What it shows about the curveMitigation can be pooled. Shared, sector-specific benchmarks retire the inherited-vendor-claim risk more cheaply than any single operator's acceptance testing — the register's third row, industrialised.

Open Power AI Consortium (opens in a new tab)

One reading disciplines all three: nothing above is an argument that these organisations have eliminated AI risk — it is that each has made its risk posture public and inspectable, which is the property this page's register exists to produce. A published frame can be audited, criticised and improved; an unpublished one can only be discovered, usually by an incident.

The mitigation architecture, layer by layer

The five control layers a de-risked utility AI estate actually needs, annotated with the stage at which each becomes mandatory — and mapped to the frameworks auditors will ask about.

A de-risked AI estate in a utility is five control layers, each becoming mandatory at a specific stage of the curve. The architecture below is deliberately vendor-neutral — every layer is defined by what it must guarantee, not by what product provides it — and it is cumulative: stage 5 runs all five layers, and a programme attempting autonomy without the assurance layer is running a stage-5 exposure on stage-3 evidence. The external frameworks map cleanly onto it: the NIST AI Risk Management Framework (opens in a new tab) (govern, map, measure, manage) spans layers one to five, ISO/IEC 42001 (opens in a new tab) certifies the management system the assurance layer implements, and the evidence artefacts double as material for NERC CIP (opens in a new tab) audits and FERC (opens in a new tab)-jurisdictional reporting in North America, or Ofgem licence-condition conversations in Great Britain.

Control layers required by stage

Each layer is annotated with the stage that first requires it. Build order matters: the boundary and inventory layers are cheap at stage 1–2 and forensic at stage 4.

  1. OT boundary & data controls

    Stage 1+

    • Data-flow registerEvery path from SCADA, historian and AMI to any model or vendor, classified
    • Sanctioned extract processThe safe route made cheaper than the workaround
    • Technical enforcementFlows that can actually be stopped, not just contractually forbidden
  2. Model inventory & validation

    Stage 2+

    • Model registerEvery model, vendor feature and scoring spreadsheet, each with a named owner
    • Acceptance testingVendor claims revalidated on your network data against a holdout
    • Procurement hooksAI-disclosure clauses so new vendor features enter the register at purchase
  3. Operational safeguards

    Stage 3+

    • Advisory write-back with approvalOutput lands in the field the operator reads; a person approves, everything logged
    • Drilled fallbackThe previous method one switch away, exercised on a calendar
    • Event-keyed drift & freshness alertingAlerts tied to DER growth, reconfigurations and AMI changes, paging a named owner
  4. Portfolio risk management

    Stage 4+

    • Dependency registerWhich models share which feeds, features, vendors — with exercised loss contingencies
    • Model risk review gateA cadence with teeth: it can delay or refuse a change
    • Portfolio health metricsCoverage, drift status and drill currency, in one view
  5. Assurance, envelopes & audit

    Stage 3+

    • Decision logInputs, model version, output, human action — reconstructable months later
    • Versioned operating envelopeBounds reviewed like protection settings; kill switch drilled (stage 5)
    • Framework mappingEvidence generated as a by-product, mapped to NIST AI RMF, ISO/IEC 42001 and CIP

Pipeline described

  1. OT boundary & data controls (stage 1+) — Data-flow register: Every path from SCADA, historian and AMI to any model or vendor, classified; Sanctioned extract process: The safe route made cheaper than the workaround; Technical enforcement: Flows that can actually be stopped, not just contractually forbidden
  2. Model inventory & validation (stage 2+) — Model register: Every model, vendor feature and scoring spreadsheet, each with a named owner; Acceptance testing: Vendor claims revalidated on your network data against a holdout; Procurement hooks: AI-disclosure clauses so new vendor features enter the register at purchase
  3. Operational safeguards (stage 3+) — Advisory write-back with approval: Output lands in the field the operator reads; a person approves, everything logged; Drilled fallback: The previous method one switch away, exercised on a calendar; Event-keyed drift & freshness alerting: Alerts tied to DER growth, reconfigurations and AMI changes, paging a named owner
  4. Portfolio risk management (stage 4+) — Dependency register: Which models share which feeds, features, vendors — with exercised loss contingencies; Model risk review gate: A cadence with teeth: it can delay or refuse a change; Portfolio health metrics: Coverage, drift status and drill currency, in one view
  5. Assurance, envelopes & audit (stage 3+) — Decision log: Inputs, model version, output, human action — reconstructable months later; Versioned operating envelope: Bounds reviewed like protection settings; kill switch drilled (stage 5); Framework mapping: Evidence generated as a by-product, mapped to NIST AI RMF, ISO/IEC 42001 and CIP
Step-by-step insights
OT boundary — the layer that cannot be retrofitted cheaply
Every other layer can be added late at a price; the boundary layer is the one whose absence leaves permanent residue. Extracts emailed during two years of pilots cannot be un-sent, and reconstructing which data left the estate — for a security review or a CIP conversation — is archaeology. This is why the data-flow register belongs at stage 1, before there is much to register: the register's value is that it was always there, so its gaps are findings rather than its existence being one.
Inventory & validation — where procurement becomes a risk control
The model register decays unless something feeds it automatically, and the feed is procurement. An AI-disclosure clause — the vendor must declare model-driven features and their update mechanism — turns every purchase into a register entry, closing the loop that otherwise reopens with every ADMS upgrade. Acceptance testing belongs in the same contractual moment: a fortnight of validation on your own feeder data, with your DER penetration, before signature, converts the vendor's global accuracy claim into a local, evidenced one.
Operational safeguards — the three controls that make coupling survivable
Advisory write-back, drilled fallback, event-keyed alerting: the stage-3 triad. The write-back bounds consequence while generating the override log; the fallback bounds the cost of being wrong; the alerting bounds the time you can be wrong without knowing. Note what is absent — nothing in the triad improves the model. That is deliberate. At the coupling moment, control investment outranks accuracy investment, because an accurate model with no fallback is a better-dressed single point of failure.
Portfolio risk — contingency thinking utilities already own
Transmission planners exercise the loss of any single element as a matter of course; the portfolio layer applies the same doctrine to the AI estate. Enumerate the shared dependencies — the weather feed, the feature pipeline, the cloud tenancy, the one vendor — and exercise the loss of each: what degrades, in what order, and what does the control room do? The exercise report is simultaneously an operational rehearsal and the concentration-risk evidence a risk committee can act on. Utilities rarely need convincing here; they need only the reframing that a weather feed is now a grid dependency.
Assurance — evidence as a by-product, or theatre
The test of the assurance layer is whether its artefacts are produced by operation or for inspection. A decision log written by the delivery pipeline, an envelope whose versions accumulate through real reviews, drill entries dated by the drills themselves — these cost nearly nothing at audit time because they already exist. The alternative — evidence assembled in the six weeks before the auditor arrives — is more expensive, less convincing, and reliably one incident out of date. Map the artefacts to NIST AI RMF and ISO/IEC 42001 once, and every subsequent audit is an export.

The build order is the architecture's real content. Boundary and inventory first, because they are cheap early and forensic late; safeguards at the coupling moment, not after it; portfolio and assurance layers grown from artefacts the earlier layers already produce. Programmes that invert the order — governance frameworks first, inventory last — spend a year producing documents that govern an estate they have not yet mapped.

A 90-day plan: putting controls around a storm-outage model

Risk mitigation made concrete on one common energy problem — an outage-prediction model already influencing crew pre-staging through the OMS, currently running with no register entry, no drilled fallback and no decision log.

The fastest way to de-risk an estate is to fully control one live exposure, because every control built for it becomes the template for the next model. The plan below takes the most common uncontrolled exposure we see at distribution utilities: a storm outage-prediction model — often a vendor module inside the OMS — whose output already influences where crews are pre-staged ahead of weather, running without a register entry, an acceptance test on local data, a drilled fallback or a decision log. Ninety days, no model development at all, and at the end the model is either evidenced or retired; both outcomes are wins.

Ninety days from uncontrolled exposure to evidenced control

One model, one owner, four phases. If any phase needs more than its window, narrow the scope — fewer districts, one storm season of history — rather than extending the plan.

  1. Days 1–15

    Register the exposure and baseline it

    Write the model's register entry: what it consumes (weather feed, AMI events, historian tags), what it influences (pre-staging, call-out sizing), and who owns it — name the distribution operations manager, since crew hours and restoration times are their numbers. Pull two storm seasons of OMS history and compute the baseline: pre-staging accuracy, wasted truck rolls, restoration times by district.

    A register entry, a named owner, a measured baseline

  2. Days 16–45

    Acceptance-test it on your own network

    Re-validate the vendor's claims against your own history: districts with high rooftop-solar penetration scored separately, overhead versus underground circuits separately. Define the operational holdout — comparable districts kept on the conventional storm matrix — for later attribution. Map the model's data path across the IT/OT boundary and register every flow; rule explicitly on whether anything touches CIP scope.

    A signed acceptance report and a registered data path

  3. Days 46–70

    Build and drill the operational safeguards

    Keep the conventional storm matrix maintained and one switch away; drill the reversion once, deliberately, on a quiet-weather day, and log the date. Stand up freshness alerts on the weather feed and AMI inputs, and drift alerts keyed to the DER connection register and any planned reconfigurations. Agree the alert budget with the control room so the new alarms displace noise rather than adding to it.

    A dated fallback drill, live event-keyed alerting

  4. Days 71–90

    Turn on evidence and attribute the value

    Switch on the decision log: every pre-staging recommendation, the forecast inputs behind it, and what the duty manager actually did. Rehearse an audit: reconstruct one storm event end to end and time it. Then report the quarter in operational units — pre-staging accuracy and wasted truck rolls against the holdout districts — and let that number, not the model's error metric, decide whether the model keeps its place.

    Reconstruction in hours; a holdout-attributed value number

The order matters

  1. Register before controls

    The register entry comes first because every later artefact — acceptance report, drill log, decision log — attaches to it. Controls built before the register exists become orphan documents that the next reorganisation loses.

  2. Controls before accuracy

    Nothing in the ninety days improves the model, deliberately. If the acceptance test shows the vendor's model underperforms the conventional storm matrix on your network, the correct output of the quarter is retirement — a cheaper and more honest outcome than tuning a model that should not be trusted yet.

  3. Evidence as a by-product, from day one

    The drill log, the acceptance report and the decision log are all generated by doing the work, not written up afterwards. That is what makes the ninety days repeatable: the second model inherits templates, and the audit pack assembles itself.

When the mitigations themselves fail

The controls on this page are not fire-and-forget. Four patterns account for most control decay — each with the cheap preventive that keeps it from happening.

Mitigations fail the same way models do: silently, while everyone believes they are working. A control that has decayed is worse than a control that never existed, because it appears on the register, satisfies the review, and buys confidence the estate no longer deserves. The four patterns below account for most of the decay we see — and every one is caught by a cheap, boring, calendar-driven preventive.

Likelihood: highImpact: high

The fallback that exists only in the audit binder

The conventional method is documented, referenced in every review — and has not been executed since the engineer who ran it retired. The switch has never been tested against the current OMS version. On the storm night it is needed, the fallback fails alongside the model, which is the exact correlated failure the control existed to prevent.

PreventionDrill the reversion on a calendar — twice a year minimum — and treat a missed drill like a missed relay test.

Likelihood: highImpact: medium

Drift monitoring tuned to statistics, blind to the network

The alerts watch distribution shifts in the model's inputs and outputs, but nothing connects them to the DER connection register, the reconfiguration plan or the AMI rollout schedule. The network changes on its own calendar, the statistical thresholds lag by months, and the drift is discovered by an operational miss instead of an alert.

PreventionKey drift alerts to the network-change events you already track — connections, reconfigurations, meter changes — not just to thresholds.

Likelihood: mediumImpact: medium

The register that dies in the reorganisation

The model inventory was accurate the day it was built. Then the innovation team was restructured, its owner changed roles, and eighteen months later the register describes an estate that no longer exists — while new vendor features shipped unregistered. The organisation is back at stage 1, still holding stage-3 paperwork.

PreventionTie register maintenance to processes that survive reorganisations: procurement (AI-disclosure clauses) and OT change control, not a named team's goodwill.

Likelihood: lowImpact: high

The envelope inherited by a decision it never covered

Autonomy thresholds derived from one decision type's override history get applied to a neighbouring decision type that was never in that history — DER curtailment bounds reused for battery dispatch, say. The first bad automated action triggers a blanket automation freeze, and the programme regresses two stages from a single incident.

PreventionEach decision type re-earns autonomy from its own override log — envelope inheritance is prohibited by policy, in writing.

The afternoon audit: eight checks that reveal your real posture

Every check below is observable in an afternoon from systems and logs — no workshops, no self-report. Tick what you can evidence today, not what is planned.

Your real risk posture is readable in an afternoon, because every control that matters leaves an artefact: a register with a review date, a drill log entry, an alert that names an owner, a reconstruction that can be timed. The checklist below is that afternoon, in order. The discipline is to tick only what you can point at — a dated log entry, a signed report — because 'we could produce that if asked' is precisely the sentence this page exists to retire.

ControlArtefact to ask forWhere it livesPass condition
Model inventoryThe register itself, with its last review dateRisk system or governed documentCovers vendor features and spreadsheets; reviewed within the last quarter
Data-flow governanceThe flow register for one named modelSecurity / OT documentationEvery hop listed; at least one flow technically stoppable, not just contractually
Acceptance testingThe acceptance report for the last vendor model adoptedProject recordsTested on own network data against a holdout, before coupling
Fallback readinessThe most recent reversion drill entryOperational logDated within six months, against the current system version
Drift & freshness alertingThe alert configuration and its firing historyMonitoring stackKeyed to network events; alerts page a named owner; fires have responses
Decision loggingA timed reconstruction of one AI-influenced decisionDecision logInputs, model version, output and human action recovered in hours
Envelope governance (if automating)The envelope's version historyPolicy repositorySigned versions; last review dated; escalation trend charted
Exercised failure responseThe last exercise report containing a model-failure injectEmergency-preparedness recordsA model failure has been rehearsed, not just discussed
How to verify each control without a meeting: the artefact to ask for, where it lives, and the pass condition. An assurance is not an artefact.

The eight-control checklist

Tick only what you can evidence with an artefact today. This list works without JavaScript, and the count maps to the stage bands used across this page.

0 of 8 ticked

Nothing evidenced yet — start with the inventory, this month

Zero ticks almost always means unmapped, not uncontrolled-but-known. The inventory is four to eight weeks of part-time work and it converts every other risk on this page from invisible to listed. Do nothing else first.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Shadow AI
Models influencing real decisions without appearing on any register — vendor 'AI-enabled' features switched on inside the ADMS, OMS or APM, and engineers' machine-learned scoring spreadsheets. The defining risk of stage 1.
Model inventory
The maintained register of every model, vendor feature and scoring tool that influences an operational decision, each with a named owner. The cheapest control on the curve and the prerequisite for every other one.
IT/OT boundary
The separation between enterprise IT and the operational technology that monitors and controls the grid. Every model or vendor consuming SCADA, historian or AMI data creates a crossing that must be registered, classified and technically enforceable.
Acceptance test
Revalidation of a vendor's model claims on the buyer's own network data — its DER penetration, circuit mix and weather — against a holdout, before the model couples to operations. Retires inherited-claim risk.
Drift
Degradation of a model as the live network moves away from the one it was trained on — in utilities, driven by DER growth, reconfigurations, AMI changes and new interconnection, arriving on the network's calendar rather than the model team's.
Event-keyed drift alerting
Drift monitoring tied to the network-change events the utility already tracks — connections, reconfigurations, meter rollouts — rather than to statistical thresholds alone, which lag the network by months.
Fallback source
The previous method — the deterministic storm matrix, the seasonal profile, the manual health index — maintained one switch away from the model, and drilled. An undrilled fallback is a document, not a control.
Decision log
The record linking a model's inputs, version and output to the human action taken, reconstructable months later. The single artefact that serves NERC CIP evidence, Ofgem conversations and EU AI Act obligations at once.
Operating envelope
The versioned, signed set of bounds inside which an enumerated decision may execute without human approval. Governed like a protection setting: proposed with evidence, independently reviewed, periodically re-derived against the current network.
Escalation rate
The share of automated decisions falling outside the envelope and routing to a person. Charted as a leading indicator: a rising trend means the network has moved and the envelope's validity is decaying.
Common-mode failure
Simultaneous, correlated degradation of multiple models through a shared dependency — one weather feed, one feature pipeline, one cloud tenancy — typically during the event when all of them matter most. The defining stage-4 risk.
Automation complacency
The recalibration of human vigilance after months of correct automated operation — operators stop shadowing the system and fallback skills decay. Mitigated artificially: scheduled fallback-mode hours and injected exercise scenarios.

Frequently asked questions

The questions utility leaders and engineers ask most often when they start treating AI adoption as a risk discipline.

What are the biggest AI adoption risks for energy and utilities companies?

Five families cover almost everything: model and data risk (drift, unvalidated vendor claims, leaked SCADA and AMI extracts), operational risk (wrong output coupled to grid-facing actions), security risk (new attack surface across the IT/OT boundary — the IEA reports attacks on utilities tripled in four years), regulatory risk (being unable to evidence an AI-influenced decision), and organisational risk (control-room trust miscalibration and automation complacency). Which family dominates depends on your maturity stage, which is why this page indexes every risk to the curve.

Do AI risks decrease as our programme matures?

No — they change class, which is this page's organising claim. Stage 1–2 risks are cheap and invisible: unregistered models, ungoverned extracts. Stage 3–4 risks are operational: drift, missing fallbacks, common-mode failure across a portfolio. Stage 5 risks are governance risks: envelope validity and reconstruction. Total exposure often rises with maturity because consequence rises faster than probability falls. The practical implication is that controls must be rebuilt at each stage; a utility carrying stage-2 controls into stage-4 exposure is the standard pre-incident configuration.

What is shadow AI in a utility, and how do we find it?

Shadow AI is any model influencing real decisions without being registered: vendor 'AI-enabled' features switched on by an ADMS, OMS or APM upgrade, machine-learned modules inside planning tools, and engineers' private scoring spreadsheets that quietly inform deferrals. You find it by sweeping three places: vendor release notes and contracts for model-driven features, the estate's spreadsheets and scripts that feed recurring decisions, and interviews with engineers — asking not 'do you use AI?' but 'what informs this decision?'. Most utilities find three times more than they expected.

Does NERC CIP prohibit AI in operational systems?

No. NERC CIP governs the security of critical cyber assets — categorisation, access control, change management, incident response — and AI systems touching those assets fall inside that scope rather than outside the rules. The practical consequences are two: every data path from OT to a model or vendor needs an explicit CIP-scope ruling, and AI-influenced operational decisions need the same evidential discipline as other control-system changes. A decision log and a registered data-flow map are what turn a CIP conversation about AI from a problem into an export.

How does the EU AI Act treat AI in energy grids?

The EU AI Act classes AI systems used as safety components in the management and operation of critical infrastructure — energy explicitly included — as high-risk. High-risk classification does not prohibit deployment; it mandates a risk-management system, data governance, logging, human oversight, accuracy and robustness requirements, and conformity assessment. For a utility, the striking fact is how closely those obligations match the controls on this page's register: an operator with a model inventory, acceptance tests, decision logs and versioned envelopes has already built most of the compliance posture.

How do we validate a vendor's AI claims before signing?

Run an acceptance test on your own network data, before contract signature, against a holdout. Vendor accuracy figures are earned on someone else's network — different climate, DER penetration and circuit mix — and transfer is the exception, not the rule. Score districts with high rooftop-solar penetration separately, overhead and underground circuits separately, and compare against your incumbent method, not against zero. Write the AI-disclosure and revalidation terms into the contract so the test repeats when the vendor ships model updates. A fortnight of validation is the cheapest insurance on the curve.

What is the single cheapest mitigation to start with?

The model inventory. It is four to eight weeks of part-time discovery work, it changes no production system, and it converts every other risk on this page from invisible to listed — because acceptance testing, drift monitoring, fallbacks and decision logs all operate per model, and you cannot apply them to models you have not found. Close behind it is the fallback drill for any model already coupled to operations: one afternoon on a quiet day, and it is the control whose absence most reliably turns a model failure into an operational incident.

How do we monitor model drift on a distribution network?

Key the monitoring to network events, not just statistics. Distribution networks drift on a known calendar: DER connections (tracked in the connection register), reconfigurations (tracked in switching records), AMI rollouts and meter changes, new large loads. Wire those event streams into the alerting so a threshold review triggers when the network changes, not months later when the statistical drift finally crosses a line. Retain the statistical thresholds as the backstop. A model watched only by statistics finds out about the network from its own errors — which means an operational miss found it first.

Who should own AI risk in a utility?

Three roles, not one. Each model needs a named operational owner — the person whose numbers move if it is wrong, paged when it degrades. The portfolio needs a model-risk owner with a route into the existing corporate risk register and committee — not a new parallel structure. And the technical controls — boundary, alerting, logs — need an engineering owner, usually where OT security already sits. The pattern to avoid is a standing AI ethics committee with no named individuals: committees diffuse exactly the accountability that makes controls real. Decision rights beyond ownership are covered in our leadership playbooks page.

Can we use cloud AI services on SCADA or AMI data?

Yes, with a governed boundary — most of the sector's serious programmes, including publicly announced utility-cloud partnerships, do exactly this. The requirements are: a registered, classified data path for every flow (with an explicit CIP or critical-infrastructure ruling where it applies), technical enforcement so a flow can actually be stopped, contractual retention and use terms verified rather than assumed, and one-way patterns where the workload allows. What is not defensible is the stage-2 default: extracts moving under a contract clause nobody technical has read, through paths nobody could enumerate.

What evidence will a regulator ask for after an AI-influenced decision?

A reconstruction: what data went in, which model version ran, what it output, and what the human did with it — months after the fact. NERC CIP audits, Ofgem licence conversations and EU AI Act obligations all converge on that demand, differing in vocabulary rather than substance. The artefact that answers all three is a decision log produced as a by-product of operation, plus the register entry, the acceptance report and — where automation exists — the envelope's version history. If assembling that pack takes weeks, the evidence layer is theatre; the pass condition is hours.

When is it safe to let AI act autonomously on the grid?

When four conditions hold at once: the action is contained and instantly reversible (volt/VAR in band, DER curtailment inside connection agreements — never switching or protection); an override log from a long advisory period supports the proposed bounds; the envelope is versioned, independently reviewed and re-derived when the network changes; and the kill switch plus reversion path are drilled on a calendar. Severity and reversibility, not model accuracy, are the gate — a more accurate model never moves a decision out of the advisory-forever quadrant.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — load and generation forecasting, asset-health scoring, outage prediction and decision support running against live operational data, integrated into the EMS, ADMS and OMS layer with the fallbacks, monitoring and audit trails that make them safe to leave running.

  • · Production deployments against live SCADA, AMI and historian data
  • · Risk and maturity assessments run jointly with utility engineering teams
  • · Controls-first delivery: fallback drills, drift alerting, decision logs
  • · 16 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyEnergy and AI — executive summary (opens in a new tab)
  3. NISTAI Risk Management Framework (opens in a new tab)
  4. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  5. OfgemOfgem — energy regulation (opens in a new tab)
  6. NERCNERC — reliability and security (opens in a new tab)
  7. NERCNERC Reliability Standards (opens in a new tab)
  8. FERCFederal Energy Regulatory Commission (opens in a new tab)
  9. European CommissionRegulatory framework for AI (EU AI Act) (opens in a new tab)
  10. EurelectricEurelectric — the European electricity industry (opens in a new tab)
  11. McKinsey & CompanyElectric power and natural gas insights (opens in a new tab)
  12. NESOThe ESO and artificial intelligence (opens in a new tab)
  13. NESOJune 2025 Digitalisation Strategy and Action Plan (opens in a new tab)
  14. Duke EnergyDuke Energy collaborates with AWS on smart grid solutions (opens in a new tab)
  15. EPRIArtificial intelligence thought leadership (opens in a new tab)
  16. EPRIOpen Power AI Consortium (opens in a new tab)

Know your exposure before an incident maps it for you

We run the risk assessment with your OT, security and operations leads, sweep the estate for unregistered models, benchmark your controls against comparable operators, and leave you a stage-indexed register with the first three controls costed. You keep everything either way.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.