Energy & UtilitiesAI Adoption & Maturity Curve
AI adoption risks in energy and utilities — and how to mitigate them at every stage
AI adoption risks in energy and utilities are the ways an AI programme can damage the operation it is meant to improve — bad decisions coupled to the grid, silent model drift, a widened IT/OT attack surface, and regulatory exposure. They change class at every stage of the maturity curve, so mitigation must be rebuilt at each stage, not inherited.

Key takeaways
- AI adoption risk in a utility does not shrink as the programme matures — it changes class. Stage 1–2 failures waste money invisibly; stage 3–4 failures reach the control room; stage 5 failures reach the network. Each stage needs its own controls, and inheriting the last stage's controls is itself a failure mode.
- The cheapest control on the whole curve is a model inventory. Most utilities already run more models than they can list — vendor 'AI-enabled' features inside the ADMS and APM, engineers' scoring spreadsheets — and a risk you cannot enumerate cannot be mitigated.
- The riskiest single transition is stage 2 to 3, when model output first couples to grid operations. Three controls make it survivable: an acceptance test on your own network data, a drilled fallback to the previous method, and drift alerting keyed to network events such as DER growth and reconfiguration.
- Regulators do not prohibit AI in the decision path — NERC CIP, Ofgem licence conditions and the EU AI Act all converge on the same demand: evidence. A decision log that can reconstruct any AI-influenced decision months later is the one artefact that serves all three at once.
- Mitigation posture follows two questions, not one: how severe is a wrong output on the grid, and how quickly can the action be unwound? Severe-and-slow decisions stay advisory forever; contained-and-reversible ones are where autonomy is earned first.
Abbreviations used on this page
- SCADA
- Supervisory control and data acquisition
- EMS
- Energy management system (transmission control)
- ADMS
- Advanced distribution management system
- OMS
- Outage management system
- AMI
- Advanced metering infrastructure (smart meters and head-end)
- APM
- Asset performance management (transformer and line health)
- DER
- Distributed energy resources (rooftop solar, batteries, EVs)
- DERMS
- Distributed energy resource management system
- OT
- Operational technology — the control-system side of the estate
- NERC CIP
- NERC Critical Infrastructure Protection standards
- SAIDI
- System average interruption duration index
- AI RMF
- NIST AI Risk Management Framework
Free · 8 questions · ~3 minutes
Score your risk controls against the curve
Eight questions, one at a time, about three minutes — each scoring a control that exists or does not, never an ambition. Answer them and we build your personalised risk report: your stage on the curve, your score on each of the four control dimensions, and the specific unmitigated risk most likely to bite next. It lands in your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised risk report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the unmitigated risks typical of your stage, and the 90-day control plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Unmapped
AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
Your next moveBuild the model inventory: every model, vendor feature and scoring spreadsheet that influences an operational decision, each with a named owner.
Stage 2 · Pilot-exposed
Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
Your next moveStand up the two stage-2 controls: a classified extract process across the IT/OT boundary, and an acceptance test that revalidates every vendor claim on your own network data against a holdout.
Stage 3 · Grid-coupled
Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
Your next moveStandardise the stage-3 control set — event-keyed drift alerting, freshness SLAs, drilled fallbacks, decision logging — so the next coupled model inherits controls instead of re-inventing them.
Stage 4 · Portfolio-managed
Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
Your next moveDefine, per candidate decision, the operating envelope and evidence bar under which autonomous execution would be acceptable — before any automation is switched on.
Stage 5 · Bounded autonomy
Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.
Your next movePut envelope reviews, escalation-rate monitoring and kill-switch drills on the same operational calendar as protection-setting reviews — permanently.
0 / 24
Model & data risk controls
— / 6
Grid-facing safeguards
— / 6
Security & compliance posture
— / 6
Risk governance & assurance
— / 6
Your score maps to a stage on the risk curve. Read the dimension breakdown before the total: the lowest dimension is where your next incident is most likely to originate, and it is where the next control belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the risk curve. Read the dimension breakdown before the total: the lowest dimension is where your next incident is most likely to originate, and it is where the next control belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the register built with your engineers?
We will walk your OT, security and operations leads through the dimension scores, sweep the estate for unregistered models, and leave you a stage-indexed risk register with the first three controls costed. No obligation, and you keep the register either way.
How the score maps to a stage
- 0–5 — Stage 1, Unmapped. AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
- 6–11 — Stage 2, Pilot-exposed. Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
- 12–16 — Stage 3, Grid-coupled. Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
- 17–21 — Stage 4, Portfolio-managed. Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
- 22–24 — Stage 5, Bounded autonomy. Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.
What AI adoption risks are in energy and utilities — and why they change class
A definition, the five risk families, and the property that organises this whole page: as a programme matures, its risks do not shrink — they change class.
AI adoption risks in energy and utilities are the ways an AI programme can harm the operation it serves: a wrong output coupled to a grid decision, a model silently drifting away from the network it describes, operational data leaking across the IT/OT boundary, a regulator asking for a reconstruction nobody can produce, and a control room whose trust in the system is miscalibrated in either direction. They are distinct from ordinary project risk — a failed AI project wastes money, whereas a badly governed successful one can mis-stage storm crews, defer the wrong transformer maintenance or hold a voltage envelope that no longer matches the network.
The organising property is that these risks change class as the programme matures. At stages 1–2 the dominant exposures are invisible and cheap: unregistered models, ungoverned extracts, inherited vendor claims. From stage 3 the exposure is operational: output couples to the EMS, ADMS and OMS, and drift, missing fallbacks and alarm fatigue become the live risks. By stages 4–5 the exposure is systemic and regulatory: common-mode failure across a portfolio, concentration in shared feeds and vendors, envelope validity, and the reconstruction demands of NERC CIP (opens in a new tab), Ofgem (opens in a new tab) licence obligations and the EU AI Act (opens in a new tab), which classes AI safety components in critical infrastructure — energy explicitly included — as high-risk systems. The controls that retire stage-2 risks do nothing against stage-4 ones, which is why mitigation must be rebuilt at each stage rather than inherited.
Model and data risk
Wrong, drifted or unvalidated model output, and the poor-quality or leaked data behind it. In a utility this family lives in specific places: SCADA and historian extracts feeding pilots, AMI interval data in vendor tenancies, health scores inside the APM, forecasts inside the EMS.
Operational and safety risk
The coupling of model output to grid-facing actions — switching recommendations, crew pre-staging, maintenance deferrals, voltage control. Governed by two questions: how severe is a wrong output, and how fast can the action be unwound?
Security risk
Every model, feed and vendor tenancy is new attack surface across the IT/OT boundary. The IEA's Energy and AI analysis records that cyberattacks on energy utilities tripled in four years and grew more sophisticated because of AI — the same technology sits on both sides of this ledger.
Regulatory and compliance risk
Not prohibition but evidence: reliability standards, licence conditions and the EU AI Act's high-risk regime all demand that AI-influenced decisions be explainable, monitored and reconstructable. The risk is being unable to show your workings.
Organisational risk
Control-room trust miscalibration, operator deskilling as advisory systems take load, and automation complacency once things work — the slowest-moving family, and the one that converts a technical failure into an operational incident.
Consequence of an uncontrolled failure, along the curve
The curve every risk on this page is indexed against. As output couples to the grid (stage 3) and then to autonomous action (stage 5), the consequence of the same wrong output rises steeply — which is why controls must be rebuilt at each stage. The stages themselves are named by the risk class that dominates them.
Consequence if a control fails by stage
- Stage 1 · Unmapped — 24% of operators. AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
- Stage 2 · Pilot-exposed — 37% of operators. Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
- Stage 3 · Grid-coupled — 25% of operators. Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
- Stage 4 · Portfolio-managed — 11% of operators. Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
- Stage 5 · Bounded autonomy — 3% of operators. Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.
Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with IEA Energy and AI grid-application analysis.
How a wrong output travels at each stage
The same failure — a model producing a wrong number — lands in three different worlds depending on stage. At stages 1–2 it dies in a spreadsheet, invisibly; at stages 3–4 it reaches the control room through governed, logged steps; at stage 5 it reaches the network unless the envelope catches it. The controls at each stage exist to match the blast radius.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- At stages 1–2, operational data leaves the estate as ad hoc extracts, feeds unregistered models, and returns as decks and health indices that quietly inform real deferrals. Nothing is logged, so a wrong output costs little today and cannot be reconstructed later — the exposure is invisibility itself.
- At stages 3–4, registered flows from the historian and AMI feed acceptance-tested, drift-watched models whose output lands in advisory fields inside the OMS or ADMS. The control room approves or overrides each recommendation with the previous method one switch away, and every step is logged.
- At stage 5, the accumulated override log justifies a versioned operating envelope inside which enumerated, reversible actions — volt/VAR optimisation, DER dispatch — execute automatically. Anything outside bounds escalates to a person, and the escalation rate is charted as the leading indicator of envelope decay.
Step-by-step insights
- The unlogged edge is the most dangerous line on this diagram
- The dashed edge from 'deck or health index' to 'engineer may act' is where stage-1 risk actually lives. The failure is not that the engineer acts on the model — it is that no record exists that a model was in the loop at all. When the deferred transformer fails eighteen months later, the investigation finds a decision with no decision path. Every control later on the curve — registers, logs, envelopes — exists to make that edge impossible to draw.
- Why the boundary node changes everything
- The difference between lane 0's 'exported SCADA history' and lane 1's 'historian & AMI feeds' is not the data — it is the register. A classified, registered flow can be technically enforced, monitored for freshness, and produced in an audit. An emailed CSV can do none of those things, and its riskiest property is that it works fine, normalising a path that widens with every pilot. Utilities that build the sanctioned path early never accumulate the shadow paths that take years to find again.
- The advisory field is a risk instrument, not a half-measure
- Writing model output into an advisory field the operator already sees — rather than automating, and rather than a separate dashboard — does double duty. It bounds the blast radius, because a person with network context checks every action. And it generates the override log: every accept and reject, with circumstances, is evidence about where the model can and cannot be trusted. That log is the only honest basis for the stage-5 envelope. Skip the advisory years and the envelope is guesswork wearing a signature.
- The envelope is a protection-setting, culturally
- Utilities already run the exact governance an AI operating envelope needs — for protection settings. Proposed with studies, reviewed by someone who did not propose them, versioned, revisited when the network changes. Treating the AI envelope as a member of that family, on the same review calendar, imports thirty years of discipline for free. Treating it as software configuration — tunable in a settings screen — is how automation quietly outgrows its evidence.
- Escalation rate: the one number that watches the watchers
- Every stage-5 mechanism can decay silently except one: the escalation rate. When the share of automated decisions falling outside bounds trends upward, something real has changed — new DER connections, a reconfiguration, drifted inputs — and the envelope's claim about the world is aging. Charting it beside SAIDI on the operational dashboard turns envelope decay from a latent audit finding into a routine operational signal that triggers review before an excursion forces one.
The five stages of AI risk in a utility, in detail
Each stage named by the risk class that dominates it. For each: what it looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.
The five stages below are risk postures, not capability levels — each is named for the risk class that dominates it, and the honest question at each is not 'what can we build?' but 'what can currently go wrong, and what control retires it?'. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own estate this week, and the anti-pattern is the mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Unmapped
24% of operators sit here
AI is already running somewhere in the estate, but nobody can list where — the dominant risk is exposure you cannot enumerate.
Stage 1 is not the absence of AI — it is the absence of a map. Most utilities at this stage are already consuming model output daily without calling it that: the APM suite scores transformer health with a vendor's proprietary model, the OMS vendor has quietly enabled a predictive module in the last upgrade, and a protection engineer maintains a dissolved-gas-analysis spreadsheet whose thresholds came from a machine-learned fit. None of it is on any register, so none of it has an owner, a validation record or a fallback.
The distinguishing risk is that exposure cannot be enumerated. Ask the leadership team how many models influence operational decisions and the answer is a guess that is usually out by a factor of three. That matters because every later control — validation, drift monitoring, audit evidence — operates per model. A risk framework applied to the four models you know about does nothing for the nine you do not.
This stage is cheap to leave and dangerous to stay in, because unmapped risk compounds silently. The maintenance deferral informed by an unvalidated health score does not fail loudly; it shows up eighteen months later as a transformer failure with no record that a model was ever in the loop — which is precisely the reconstruction a regulator will ask for.
In practice
The transformer spreadsheet nobody registered
During an asset-management review, a distribution utility found that maintenance deferrals on 40-year-old transformers were being informed by a health-index spreadsheet one engineer had built from dissolved-gas-analysis exports, with weightings fitted years earlier and never revalidated. It had quietly become the deciding input on capital deferrals worth millions. Nothing was wrong with the engineer's work — what was wrong was that no one else knew the decision path existed.
What it looks like
- No inventory of models influencing operational decisions
- Vendor 'AI-enabled' features active inside the ADMS, APM or OMS, unregistered
- Engineers run private scoring spreadsheets that inform real deferrals
- AI risk appears on no risk register, so it has no owner
Diagnostic signals you can check this week
- Ask for the list of models influencing operational decisions — time how long it takes to produce, and who has to be asked
- Check the last ADMS, OMS and APM vendor release notes for AI features enabled by default
- Ask three engineers whether any spreadsheet of theirs feeds a real decision; count the surprises
- Look for AI anywhere on the corporate risk register — absence is the stage-1 signature
Anti-pattern · Writing the AI policy before the AI inventory
The instinctive first move is a policy document — principles, ethics language, an approval workflow for future AI. It feels like control and governs nothing, because it applies to the AI the organisation plans to adopt while saying nothing about the AI already running unregistered in the estate. Inventory first, policy second: a one-page register of what exists, who relies on it and what it decides governs more real risk than any principles document, and the policy written afterwards will describe reality rather than intention.
What holds you here
You cannot mitigate what you cannot list — every control operates per model, and the model count is unknown.
Highest-leverage next move
Build the model inventory: every model, vendor feature and scoring spreadsheet that influences an operational decision, each with a named owner.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- One OT engineer and one analyst, part-time, with vendor-contract access
- Risk
- Low — discovery work, nothing in production changes
- To next stage
- 1–2 months
If this is you, the next step is
A short engagement: we sweep the estate — vendor features, spreadsheets, trials — and hand you the register.
Stage 2
Pilot-exposed
37% of operators sit here
Named pilots run on real operational data with paper controls — the dominant risks are data leaving the OT boundary and vendor claims nobody has tested.
Stage 2 is where risk becomes visible but controls remain paper. The pilots are real: a wind or load forecasting trial, an asset-health proof of value, an outage-prediction bake-off. To run them, operational data starts moving — historian extracts emailed to a vendor's data scientist, AMI interval data loaded into a cloud tenancy under a contract clause nobody technical has read. The controls that exist are legal ones, and legal controls do not stop a feed, classify an extract or notice that a export contained customer data.
The second stage-2 risk is claim inheritance. Vendor models arrive with accuracy figures earned on another operator's network — a different climate, a different DER penetration, a different mix of overhead and underground circuits. A storm-outage model trained on a coastal utility's history transfers badly to an inland one, and the difference does not surface in a demo. Utilities that skip acceptance testing on their own network data discover the gap after the model is coupled to operations, which is the most expensive possible moment.
What makes this stage treacherous is that its failures are quiet and its successes are loud. A leaked extract or an over-fitted demo costs nothing visible this quarter, while a good-looking pilot earns headlines and budget. The programme accelerates toward stage 3 carrying unexamined data paths and untested claims — which is exactly the cargo you do not want at the moment output first couples to the grid.
In practice
The forecasting pilot that emailed the historian
A generation forecasting pilot at a mid-size utility ran for five months on weekly historian exports a graduate engineer emailed to the vendor as CSV attachments. The pilot's accuracy was genuinely good. The exports, it later turned out, included tags from units subject to critical-infrastructure information rules, and no record existed of which extracts had been sent, when, or what the vendor had retained. The pilot succeeded; the data path it normalised became the finding in the next security review.
What it looks like
- Pilots consume real SCADA, historian or AMI extracts
- Data-sharing terms with vendors are contractual, not technical
- Vendor accuracy claims come from someone else's network
- Success criteria are model metrics, not operational deltas
Diagnostic signals you can check this week
- Trace one pilot's data path end to end — every hop, every retention point; note where the map runs out
- Ask who technically (not contractually) could stop data leaving the estate today
- Ask the vendor for accuracy figures on a network with your DER penetration and circuit mix — watch for the pivot to the global figure
- Check whether any pilot has a defined operational holdout for later attribution; at stage 2 the honest answer is usually no
Anti-pattern · Banning cloud AI instead of building the boundary
When a security team discovers stage-2 data paths, the reflex is prohibition: no operational data leaves the estate, no cloud AI services, pilots frozen. The estate then reliably routes around the ban — engineers anonymise poorly and email anyway, vendors run 'demos' on data brought to workshops — and the organisation ends up with the same exposure minus the visibility. The durable fix is a sanctioned path: a classified extract process, an approved tenancy, a data-flow register. Make the safe route cheaper than the workaround and the workaround disappears.
What holds you here
Data paths and vendor claims are both ungoverned, and the pilot's momentum is pushing both toward the grid.
Highest-leverage next move
Stand up the two stage-2 controls: a classified extract process across the IT/OT boundary, and an acceptance test that revalidates every vendor claim on your own network data against a holdout.
Cost of leaving
- Effort
- 3–6 months
- Team
- One OT security engineer, one data engineer, procurement support
- Risk
- Medium — the boundary work must not stall the pilots it governs, or it will be bypassed
- To next stage
- 3–6 months
If this is you, the next step is
We map every pilot's data path across the IT/OT boundary and hand you the classified-extract process.
Stage 3
Grid-coupled
25% of operators sit here
Model output first reaches operational systems and people act on it — the dominant risks become drift, missing fallbacks and control-room trust miscalibration.
Stage 3 is the moment the risk class changes physically: a wrong output stops being a wasted analysis and starts being a wrong pre-staging decision, a deferred maintenance visit, a mis-set switching plan. The programme's risk profile now moves with the network. A distribution network is not a stationary system — DER connections grow monthly, reconfigurations change flow patterns, an AMI rollout changes the very telemetry the model was trained on — and every one of those changes is a drift event that arrives on the network's schedule, not the data science team's.
The controls that matter here are operational, not analytical. Drift monitoring keyed to network events, not just statistical thresholds. Freshness alerting on every feed the model consumes, because a historian interface that silently stops updating is more dangerous than one that visibly fails. And above all a fallback: the previous method — the deterministic storm matrix, the seasonal load profile, the manual health index — kept alive, one switch away, and drilled. An undrilled fallback is a document, and documents do not operate networks.
The subtle stage-3 risk is human: trust miscalibration in the control room. Operators either over-trust the new field (accepting recommendations without the scrutiny they would apply to a colleague) or under-trust it (quietly ignoring it, so the programme's value silently evaporates while its risk remains). Both are measurable — override rates, time-to-accept — and both respond to the same treatment: showing operators the model's failure cases and its bounds, not just its wins.
In practice
The storm model that drifted with the rooftops
A distribution utility coupled an outage-prediction model to its OMS to pre-stage crews ahead of storms. It worked well for two seasons. Over the following eighteen months, rooftop solar penetration in two districts roughly doubled, changing fault signatures and feeder loading patterns the model had learned. Pre-staging accuracy decayed slowly enough that each miss looked like weather luck. Only when a crew spent a storm night positioned sixty kilometres from the actual damage did anyone check the model's error trend against the DER connection register — where the drift had been visible for a year.
What it looks like
- Model output lands in the EMS, ADMS or OMS, at least in advisory fields
- Control-room and field decisions are influenced daily
- Drift and freshness monitoring exist for the coupled models
- A fallback to the previous method exists — though it may never have been drilled
Diagnostic signals you can check this week
- Pick one coupled model and ask when its fallback was last exercised — a date and a log entry, not an assurance
- Check whether drift alerts reference network events (DER growth, reconfiguration, AMI changes) or only statistical thresholds
- Pull the override rate on advisory recommendations; both near-zero and near-total are alarms
- Ask what happens to the model's inputs when the historian interface is patched — who is told, and how
Anti-pattern · Improving accuracy while the fallback rots
Once output is coupled, teams instinctively invest where they are strongest: model accuracy. Meanwhile the fallback — the old storm matrix, the manual dispatch process — quietly decays: the person who ran it retires, the spreadsheet stops being updated, the switch is never tested. The estate ends up more dependent on the model precisely as its safety net disappears. Accuracy work is worth doing only after the fallback is drilled on a calendar, because the day you need the old method is by definition the day the model has failed.
What holds you here
Controls are per-model and hand-built, so every new coupled model re-raises the same risks with none of the previous answers.
Highest-leverage next move
Standardise the stage-3 control set — event-keyed drift alerting, freshness SLAs, drilled fallbacks, decision logging — so the next coupled model inherits controls instead of re-inventing them.
Cost of leaving
- Effort
- 6–12 months
- Team
- One integration engineer, one ML engineer, a named control-room owner, OT security review
- Risk
- Medium-high — first grid-coupled writes need rollback paths and change-board approval
- To next stage
- 9–18 months
If this is you, the next step is
We design and run the reversion exercise on a quiet day — and leave you the drill calendar.
Stage 4
Portfolio-managed
11% of operators sit here
Many models run under a common risk framework — the dominant risks become systemic: shared dependencies, common-mode failure and concentration in single vendors or feeds.
Stage 4 is where individual-model risk is largely tamed and a new class quietly replaces it: systemic risk. A portfolio of eight models serving load forecasting, asset health, outage prediction and vegetation management does not carry eight independent risks — it carries shared ones. Most of the portfolio consumes the same weather feed, the same historian, the same feature pipeline. A single upstream failure now degrades five models at once, in correlated ways, during exactly the weather event when all five matter most. The failure mode nobody designed is the one the portfolio inherited from its own efficiency.
Concentration risk arrives the same way. Standardising on one vendor's platform, one cloud tenancy or one forecasting supplier is operationally sensible and creates a single point whose commercial failure, security compromise or model regression propagates across the estate. The mitigation is not to abandon standardisation but to know precisely what depends on what — a dependency register for models, feeds and vendors — and to have exercised the loss of each critical dependency the way transmission planners exercise the loss of a line: as a contingency with a rehearsed response.
The other stage-4 discipline is portfolio-level evidence. By now the regulator conversation has changed: it is no longer 'do you use AI?' but 'show us how you manage it'. NERC CIP audits, Ofgem's licence-condition conversations and — for European operators — EU AI Act conformity all want the same artefacts: the inventory, the review cadence, the incident record, the reconstruction capability. At stage 4 those artefacts either fall out of the framework automatically, or the framework is theatre.
In practice
The morning the weather feed went down
A utility running six operational models discovered its concentration the morning its commercial weather provider had an outage. Load forecasting degraded to seasonal profiles as designed. What nobody had mapped was that vegetation risk scoring, storm pre-staging and two asset-health models consumed derived features from the same feed — all four degraded simultaneously, and the control room's fallback procedures assumed models failed one at a time. Nothing broke on the network that day. The dependency map got built the same week.
What it looks like
- A model risk function reviews every operational model on a cadence
- Controls are inherited from a standard set, not rebuilt per model
- Dependencies between models and shared feeds are mapped
- Portfolio-level metrics exist: coverage, drift status, drill currency
Diagnostic signals you can check this week
- Ask for the dependency map: which models share which feeds, features and vendors — and when the loss of the biggest shared dependency was last exercised
- Check whether the model risk review is a real gate: find one model change it delayed or refused
- Count models per critical upstream feed; more than three on one feed with no exercised contingency is unmanaged concentration
- Ask how long it takes to produce the full portfolio evidence pack for an audit — days is stage 4, weeks is stage 3 with paperwork
Anti-pattern · Scaling the framework by adding paperwork
As the portfolio grows, the risk function's instinct is to add process weight: longer review templates, more sign-offs, quarterly attestation spreadsheets. Teams respond rationally — they route around it, classifying new work as 'analytics' rather than 'models' to avoid the gate, and the inventory silently diverges from reality again, which is stage 1 wearing a stage-4 badge. The framework scales through automation, not attestation: controls inherited by default from the platform, evidence generated as a by-product of operation, and a review gate that is fast precisely because the artefacts already exist.
What holds you here
Every decision still routes through a person, so the portfolio's value is capped by control-room attention — but crossing to autonomy safely demands evidence thresholds most frameworks have not yet defined.
Highest-leverage next move
Define, per candidate decision, the operating envelope and evidence bar under which autonomous execution would be acceptable — before any automation is switched on.
Cost of leaving
- Effort
- 12–24 months
- Team
- Model risk owner, platform engineers, OT security, internal audit partnership
- Risk
- Medium — the framework must earn adoption or it will be evaded, recreating unmapped risk
- To next stage
- 18+ months
If this is you, the next step is
We run a loss-of-feed contingency exercise across your portfolio and report what actually degrades.
Stage 5
Bounded autonomy
3% of operators sit here
Defined actions execute without approval inside versioned envelopes — the dominant risks become envelope validity, automation complacency and regulatory evidence.
Stage 5 in a utility is narrower than the phrase 'autonomous grid' suggests, and correctly so. It is an enumerated set of contained, reversible actions — volt/VAR optimisation within band, DER curtailment inside connection agreements, battery dispatch within market and thermal limits — executing inside envelopes an engineer signed, with everything else escalating to a person. Switching plans, protection settings and safety-adjacent decisions stay human forever, not because models cannot rank options but because the consequence-reversibility calculus says advisory is where they belong.
The risk that defines this stage is envelope validity. An envelope is a claim about the world — these bounds are safe given this network — and the network keeps changing underneath it. New DER connections, a reconfiguration, a new interconnector: each quietly erodes the assumptions the envelope encodes. The leading indicator is the escalation rate. When the share of decisions falling outside bounds trends up, the world has moved and the envelope needs review before an incident forces one. Utilities already run this discipline for protection settings; the transfer of that habit to AI envelopes is the whole trick.
The second defining risk is complacency — months of correct automated operation recalibrate humans. Operators stop shadowing the system, the fallback drill slips, the reversion switch goes untested through two reorganisations. The mitigations are deliberately artificial: scheduled hours operating on the fallback method, injected exercise scenarios, and treating a missed drill with the seriousness of a missed relay test. Stage 5 is not a destination where risk work ends; it is a permanent operating discipline, and the utilities that hold it treat the envelope, the drill calendar and the decision log as living operational artefacts.
In practice
The envelope that aged out under new connections
A network operator ran automated voltage optimisation on a group of feeders inside an envelope validated against the connection register at go-live. Over a year, behind-the-meter batteries and two community solar sites shifted the feeders' behaviour; the escalation rate crept from one action in fifty to one in nine. Because escalations were charted as an operational KPI, the trend triggered an envelope review — recomputed bounds, a re-signed policy version — months before any customer saw a voltage excursion. The near-miss that never happened is what stage-5 discipline looks like from inside.
What it looks like
- Enumerated decisions execute automatically inside explicit bounds
- The operating envelope is versioned and reviewed like a protection setting
- Escalation rate is monitored as a leading indicator of envelope decay
- Kill switches and reversion paths are drilled, with dated log entries
Diagnostic signals you can check this week
- Ask for the envelope's version history and who signed the last change — a settings screen with no history is the anti-signature
- Check the escalation-rate trend and whether anyone owns reviewing it
- Find the last dated kill-switch drill; more than six months old means the switch is theoretical
- Ask an auditor's question: reconstruct one automated action from last quarter — inputs, model version, envelope version, outcome — and time the answer
Anti-pattern · Letting the envelope drift to match the model
When an automated system keeps escalating, the path of least resistance is to widen the bounds — each widening individually defensible, none re-deriving the envelope from network studies. Two years later the envelope reflects what the model wants to do rather than what the network can tolerate, and no one can say which version was last independently validated. Envelope changes need the discipline of protection-setting changes: proposed with evidence, reviewed by someone who did not propose them, versioned, and periodically re-derived from scratch against the current network.
What holds you here
Sustaining autonomy is a standing evidence-and-drill discipline — the constraint is envelope governance and regulatory reconstruction, not engineering.
Highest-leverage next move
Put envelope reviews, escalation-rate monitoring and kill-switch drills on the same operational calendar as protection-setting reviews — permanently.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team, control-room ownership, a standing risk forum with OT and compliance
- Risk
- Concentrated — low-frequency, high-consequence, regulatory in nature
If this is you, the next step is
We reconstruct three real automated actions end to end and stress-test the envelope, trail and reversion.
Where energy and utilities operators actually sit today
The distribution across the risk curve, why the crowd at stage 2 matters, and what the external research says about the threat side of the ledger.
Most energy and utilities operators sit at stages 1 and 2 — running pilots on real operational data with controls that are contractual rather than technical, above an estate carrying more unregistered models than anyone has listed. The crowding matters because the sector's momentum is real: system operators are publishing digitalisation strategies, vendors are shipping AI features by default inside the ADMS and APM layer, and the physics of a decarbonising grid — more DERs, more variability, tighter margins — pulls forecasting and optimisation models toward the control room. The risk work queues up exactly where the crowd is thinnest: in the controls that make coupling survivable.
Distribution of energy and utilities operators across the risk stages
Illustrative distribution, synthesised from IEA, EPRI and Eurelectric adoption research — charted to show shape, not to report a survey. Stage 2 is the mode: pilots on real data, controls on paper. The drop from stage 2 to stage 3 is where risk work, not model work, is the gate.
Share of operators (illustrative)
- 24% — 1 · Unmapped
- 37% — 2 · Pilot-exposed (the crowd)
- 25% — 3 · Grid-coupled
- 11% — 4 · Portfolio-managed
- 3% — 5 · Bounded autonomy
Source: Illustrative; synthesised from IEA, EPRI and Eurelectric research
Cyberattacks on energy utilities have tripled in the past four years and have become more sophisticated because of AI. At the same time, AI is becoming a critical tool to defend against them.
The external research is unusually consistent about both sides of the ledger. The IEA's Energy and AI report (opens in a new tab) catalogues the upside — AI-based fault detection reducing outage durations by 30–50%, better renewables forecasting cutting curtailment — while recording the tripling of attacks on the same page. EPRI's AI research programme (opens in a new tab) and Eurelectric (opens in a new tab) track the sector's adoption and its governance gap, and cross-industry work such as McKinsey's electric power insights (opens in a new tab) finds the familiar pattern of experimentation running ahead of impact. The reading that matters for this page: the sector's problem is not enthusiasm and not evidence of value — it is that controls lag coupling, and the gap between those two lines is exactly where incidents live.
The risk register: nine risks, stage-indexed, with the control that retires each
The centrepiece. Each risk enters the curve early, peaks later, and bites in a specific system — and each has a named control and a named piece of evidence that proves the control exists.
The register below is the working core of this page: the nine risks that recur across utility AI programmes, indexed by the stage where each enters, the stage where it peaks, and the operational system where it bites. Two disciplines make a register like this useful rather than decorative. First, every risk carries a control that is buildable — not 'increase awareness' but 'acceptance test on own network data against a holdout'. Second, every control carries an evidence artefact — the thing an auditor, a regulator or your own risk committee can hold. A control without evidence is an intention; utilities run on evidence.
| Risk | Enters | Peaks | Where it bites | Mitigating control | Evidence it is retired |
|---|---|---|---|---|---|
| Shadow AI — unregistered models and vendor features | 1 | 3 | APM, ADMS, OMS vendor modules; engineers' spreadsheets | Model inventory with owners, tied to procurement and change control | Register reviewed on a cadence; vendor AI-disclosure clause in contracts |
| Operational data leaking across the IT/OT boundary | 2 | 2 | Historian and AMI extracts to vendor tenancies | Sanctioned classified-extract process; registered, technically enforced flows | Data-flow register; boundary rule that can actually stop a transfer |
| Inherited vendor claims — accuracy earned on someone else's network | 2 | 3 | Forecasting, asset health, outage prediction | Acceptance test on own network data against a holdout, before coupling | Signed acceptance report with your DER mix and circuit types in it |
| Silent model drift after network change | 3 | 4 | EMS load forecast, OMS storm models, APM scores | Drift alerting keyed to network events — DER growth, reconfiguration, AMI changes | Alert log showing fires and responses; retrain cadence adhered to |
| Missing or rotten fallback | 3 | 3 | Control room, storm response, dispatch | Previous method maintained one switch away, drilled on a calendar | Dated drill log entries — not an assurance that reversion 'would work' |
| Control-room trust miscalibration and alarm fatigue | 3 | 4 | Advisory fields in EMS/ADMS/OMS | Alert budgets; operators shown failure cases and bounds, not just wins | Override-rate trend charted and reviewed — neither ~0% nor ~100% |
| Common-mode failure across the portfolio | 4 | 5 | Shared weather feeds, features, historian, one cloud tenancy | Dependency register plus exercised loss-of-feed contingencies | Contingency drill report: what degraded, what the control room did |
| Undocumented AI-influenced decisions | 3 | 5 | Everywhere output meets a decision | Decision log linking inputs, model version, output and human action | A timed reconstruction: any decision, months later, in hours |
| Automation beyond the envelope's validity | 5 | 5 | Volt/VAR, DER dispatch, battery scheduling | Versioned envelope reviewed like a protection setting; escalation-rate monitoring; drilled kill switch | Envelope version history with signatures; escalation trend on the ops dashboard |
Read the 'enters' and 'peaks' columns together and the register's main lesson appears: almost every risk is cheapest to retire one stage before it peaks. The inventory that takes four weeks at stage 1 takes a forensic quarter at stage 3, because by then the unregistered models are load-bearing. The acceptance test that costs a fortnight at stage 2 costs an incident at stage 3. This is also why 'we will add governance when we scale' is the sector's most expensive sentence — the register's risks do not wait for scale, they compound under it. Who holds the pen on each control — which decisions need an accountable owner rather than a committee — is its own discipline, covered properly in the AI leadership playbooks for energy utilities; this page stays on the controls themselves.
Mitigation posture: grid consequence × reversibility
Where each decision type belongs, before any model quality argument. Plot the wrong-output consequence against how fast the action can be unwound; the quadrant sets the control intensity. Nothing in the top-left ever automates, however good the model gets.
Advisory forever
- Switching plans, protection settings, safety-adjacent calls
- AI ranks options; a person decides, always
- Control: full decision log plus human accountability
Human-approved
- Storm crew pre-staging, dispatch recommendations, load transfers
- Every action approved in the control room, logged
- Control: drilled fallback plus override-rate monitoring
Batch with review
- Maintenance scheduling, connection studies, capital deferral scoring
- Output reviewed in planning cycles, not in real time
- Control: acceptance testing plus periodic revalidation
Earn autonomy here
- Volt/VAR in band, DER curtailment within agreements, meter-data cleansing
- Contained, reversible — the honest stage-5 candidates
- Control: versioned envelope, escalation monitoring, kill switch
The matrix is the page's second spine because it answers the question the register raises: how much control is enough? The answer is positional, not universal. A meter-data cleansing model needs an envelope and a log; a switching adviser needs a human forever, and the correct response to 'the model is now accurate enough to automate switching' is that accuracy was never the constraint — consequence and reversibility were. Utilities that write this matrix down early spend their governance budget where the grid actually needs it, instead of spreading uniform process weight across decisions that differ by three orders of magnitude in consequence.
What managed adoption looks like in public
Three publicly reported programmes, read against the risk curve. None is an Atomic Loops engagement — each links to the organisation's own published material, and the card images are generated industry scenes, not operator photography.
The clearest public evidence for controls-first adoption is in what system operators and utilities chose to publish alongside their AI capability. In each case below, the organisation's own material pairs the capability with its governing frame — a digitalisation strategy, a partnership structure, a shared validation effort — which is precisely the pairing the risk register above formalises.
Three programmes read against the risk curve
Outcomes as reported by the organisations themselves — verify against the linked source before reusing; we have not independently audited them. Card images are generated industry scenes and do not depict the organisations' facilities.
NESO (National Energy System Operator)GB electricity system operator · control room & balancing23
- Challenge
- Bringing AI into the environment with the least tolerance for a wrong output — the national control room — where forecasting errors move balancing costs and, ultimately, system security.
- Approach
- NESO's published AI communications describe building an AI Centre of Excellence to pool expertise and strengthen the security and reliability of the network, with machine-learning forecasting supporting — not replacing — control-room decision-making, and the governing frame published openly in its Digitalisation Strategy and Action Plan (June 2025).
- Reported outcome
- As reported by NESO, AI and machine learning now support forecasting and control-room insight within a published digitalisation strategy, with the capability and its governance developed and communicated together.
- What it shows about the curveThe stage-2→3 transition done in the open: the control frame ships with the capability, and advisory mode in the control room is the deliberate posture, not a stepping stone skipped for speed.
NESO — the ESO and artificial intelligence (opens in a new tab)
Duke EnergyUS investor-owned utility · six-state service area23
- Challenge
- Scaling grid analytics and smart-grid capability across a very large distribution estate requires cloud-scale tooling — which puts operational grid data and grid-adjacent decisions into a partner ecosystem, exactly the boundary stage 2 must govern.
- Approach
- Duke Energy publicly announced a collaboration with AWS to develop smart grid solutions supporting its clean-energy transition — structuring the cloud partnership as a named, governed programme rather than accumulating ad hoc vendor arrangements team by team.
- Reported outcome
- As reported in Duke Energy's own release, the collaboration develops smart-grid solutions to better serve customers and support the utility's clean-energy transition, with the partnership's scope stated publicly.
- What it shows about the curvePartner-based capability still leaves the operator holding grid accountability. A named programme with published scope is itself a control — it replaces the invisible, team-by-team data paths that define pilot-exposed estates.
Duke Energy newsroom — AWS collaboration (opens in a new tab)
EPRI — Open Power AI ConsortiumIndustry research consortium · utilities and technology partners23
- Challenge
- Every utility validating AI models alone repeats the same model-risk work — acceptance testing, benchmarking, domain evaluation — at every operator, which guarantees the sector's controls lag its adoption.
- Approach
- EPRI convened the Open Power AI Consortium to develop and benchmark domain-specific AI models and evaluation approaches for the power sector, pooling utilities and technology partners so validation effort is shared rather than duplicated.
- Reported outcome
- As reported by EPRI, the consortium operates openly with published aims around domain-specific models and benchmarks for the electricity sector.
- What it shows about the curveMitigation can be pooled. Shared, sector-specific benchmarks retire the inherited-vendor-claim risk more cheaply than any single operator's acceptance testing — the register's third row, industrialised.
One reading disciplines all three: nothing above is an argument that these organisations have eliminated AI risk — it is that each has made its risk posture public and inspectable, which is the property this page's register exists to produce. A published frame can be audited, criticised and improved; an unpublished one can only be discovered, usually by an incident.