Energy & UtilitiesReadiness & Transformation Roadmap
The grid roadmap for AI integration: how utilities move models from reports into the control loop
A grid roadmap for AI integration is the staged path by which a utility moves AI from offline analytics to supervised closed-loop operation. Progress is measured in integration depth — what models may read, where their output lands, and what they may write back to the grid — and gated by telemetry, operator trust and the OT security boundary.

Key takeaways
- A grid roadmap for AI integration measures progress in integration depth, not model quality: what the model may read (stages 1–2), where its output lands (stage 3), and what it may write back to the grid (stages 4–5). Most programmes report model progress while standing still on all three.
- The grid already runs closed loops — protection, automatic generation control, FLISR self-healing — and it accepts them because they are engineered, bounded and reversible. AI enters the loop the same way: by adjusting the setpoints and policies of engineered schemes inside a validated operating envelope, never by direct device control.
- Most utility AI is stuck at stage 2: models reading live telemetry and rendering onto side screens nobody consults during a storm. The move to stage 3 is a workflow and write-path problem — landing output inside the EMS, ADMS or OMS with accepts and overrides logged — not an accuracy problem.
- The operating envelope is the central artefact of the far stages: versioned, power-flow-validated bounds with rate limits, validity conditions and auto-revert triggers. The accept/override log collected at stage 3 is the dataset that calibrates it — skip the logging and the envelope is guesswork.
- The binding constraints on closed-loop AI are the OT security boundary (NERC CIP and its equivalents) and the currency of the network model — both buildable today. Closed-loop readiness is almost entirely present-tense integration discipline, which is why the roadmap starts at the historian, not at the model.
Abbreviations used on this page
- SCADA
- Supervisory control and data acquisition
- EMS
- Energy management system (transmission control room)
- ADMS
- Advanced distribution management system
- OMS
- Outage management system
- DERMS
- Distributed energy resource management system
- AMI
- Advanced metering infrastructure (smart meters and the head-end system)
- PMU
- Phasor measurement unit — a synchrophasor sensor sampling the waveform many times a second
- DER
- Distributed energy resources — rooftop solar, batteries, EV chargers, flexible load
- VVO
- Volt/VAR optimisation — coordinated control of voltage regulators and capacitor banks
- FLISR
- Fault location, isolation and service restoration — the self-healing switching scheme
- OT
- Operational technology — the control systems and networks that act on the physical grid
- CIP
- Critical Infrastructure Protection — the NERC cyber-security standards for the bulk power system
Free · 8 questions · ~3 minutes
Score your integration depth
Eight questions, one at a time, about three minutes. Answer them and we build your personalised integration report — your stage on the ladder, your score on each of the four dimensions, and the specific boundary currently holding your models out of the loop — and send it to your inbox. Your result doubles as the integration baseline for your next roadmap review.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised integration report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the boundary holding your models out of the loop, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Detached
Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
Your next moveStand up one governed, refreshing feed — historian, AMI and as-operated topology for one substation area — and rebuild one existing analysis on top of it.
Stage 2 · Observing
Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
Your next movePick one decision and land the model's output inside the system that already carries it — an EMS display column, an ADMS advisory, an OMS estimate — with accepts and overrides logged from day one.
Stage 3 · Advising
Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
Your next moveChoose the one decision with contained consequences and high tempo — Volt/VAR setpoints are the usual answer — and start engineering its operating envelope from the override log while running shadow mode.
Stage 4 · Acting
Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
Your next moveExtract the shared machinery — envelope validation service, scheme interfaces, action audit log, reversion pattern — so the next loop is an envelope and a model, not a programme.
Stage 5 · Coordinating
Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.
Your next moveKey every envelope and model review to network-change events, review interacting loops jointly, and report AI actions through the same reliability lens as any other control action.
0 / 24
Telemetry & network model
— / 6
Control-room integration
— / 6
Actuation & safeguards
— / 6
OT security & model governance
— / 6
Your score maps to a stage on the integration ladder. The dimension breakdown matters more than the total: the lowest dimension is the boundary your models are actually stuck at, and in utilities it is far more often the write path and its assurance than the modelling. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the integration ladder. The dimension breakdown matters more than the total: the lowest dimension is the boundary your models are actually stuck at, and in utilities it is far more often the write path and its assurance than the modelling.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the integration boundary mapped on your actual estate?
We walk your operations and engineering leads through the dimension scores, mark your EMS, ADMS, OMS and DERMS estate against the boundary map on this page, and leave you with a sequenced plan for the next stage — including the assured OT path if actuation is in reach. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–4 — Stage 1, Detached. Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
- 5–9 — Stage 2, Observing. Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
- 10–14 — Stage 3, Advising. Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
- 15–19 — Stage 4, Acting. Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
- 20–24 — Stage 5, Coordinating. Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.
What a grid roadmap for AI integration actually is
A definition, the axis the five stages sit on, and the three paths — read, advise, act — a model output can take toward the grid.
A grid roadmap for AI integration is the staged sequence by which a utility deepens the connection between its AI models and the grid itself: first a read path (live telemetry and an as-operated network model), then an advisory path (output landing inside the EMS, ADMS, OMS and work queues where people decide), and finally an actuation path (setpoints and schedules flowing into engineered control schemes inside validated operating envelopes). Its unit of progress is integration depth — what the model may read, where its output lands, what it may write — not model accuracy, which improves on a different and much easier axis.
The framing matters because the grid is not an ordinary software estate. It already runs closed loops, and has for a century: protection clears faults in milliseconds, automatic generation control (AGC) balances the system continuously, and FLISR schemes reroute distribution power without a human in the loop. The grid's operating culture accepts these because they are engineered, bounded, deterministic where it matters, and reversible. That is the standard AI must meet — and it recasts the roadmap's destination. The far stage is not 'AI runs the grid'; it is learned models adjusting the aims of engineered schemes inside bounds a human wrote down and validated. Meanwhile the pressure to get there is real and externally documented: the IEA finds that AI-based fault detection can cut outage durations by 30–50% (opens in a new tab), and that scaled AI-led interventions could save around 300 TWh of electricity globally — value that only exists on the far side of the integration work this page maps.
Two things this roadmap deliberately is not. It is not a funding and submission calendar — how the stages below are paid for, and how they cut against price-control and rate-case windows, is the subject of the AI transformation timeline for utilities. And it is not a generic maturity model with the word 'grid' inserted: every stage boundary on this page is a physical or institutional boundary of the electricity system — the historian's edge, the control-room display, the OT security perimeter, the engineered control scheme. A roadmap that could be re-badged for a retailer or a bank by find-and-replace is measuring the wrong thing.
Decisions AI can safely touch, against integration depth
The curve is flat through stages 1–2 because a model that cannot reach an operational system touches no decisions at all, however accurate it becomes. It inflects at stage 3, when output lands where operators already work, and steepens at stage 4, when contained high-tempo decisions close the loop. The far end flattens honestly: some grid decisions belong to deterministic engineering permanently.
Grid decisions safely reachable by stage
- Stage 1 · Detached — 26% of operators. Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
- Stage 2 · Observing — 36% of operators. Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
- Stage 3 · Advising — 24% of operators. Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
- Stage 4 · Acting — 11% of operators. Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
- Stage 5 · Coordinating — 3% of operators. Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.
Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with IEA, Energy and AI.
Read, advise, act: the three paths a model output can take toward the grid
The three integration paths, top to bottom in order of depth. The read path ends at a side screen — where most utility AI stops. The advisory path lands output inside the operational systems and logs every accept and override. The actuation path runs through envelope validation into an engineered scheme, with supervision and auto-revert around it — a model never touches a device directly.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- On the read path, historian and AMI extracts feed an offline model whose output lands in a report or on a side screen outside every operational system. Whether anyone acts on it is voluntary, and voluntary attention disappears during storms — which is when the model was supposed to matter. Stages 1 and 2 live entirely on this path; so does most utility AI today.
- On the advisory path, live telemetry joined to an as-operated network model feeds a served, monitored model whose output is written into the EMS, ADMS or OMS display — or the work queue — that the operator already reads. The operator decides, and every accept and override is logged with context. That log is simultaneously the adoption metric, the trust record, and the calibration dataset the actuation path will later need.
- On the actuation path, a model's proposed setpoint or schedule passes an envelope and power-flow check before anything moves: hard bounds, rate limits and validity conditions, all versioned. In-bounds proposals flow into an engineered scheme — Volt/VAR optimisation, DERMS dispatch — which actuates with its own interlocks intact. Out-of-bounds proposals escalate to the operator. Supervision watches the escalation rate as a leading indicator, every action is logged for reconstruction, and any breach auto-reverts to the previous control source.
Step-by-step insights
- The extract ceiling — why the read path caps everything above it
- A hand-assembled extract is stale the moment it lands, carries no lineage, and embeds one analyst's private join decisions between historian tags, AMI meters and GIS assets. Nothing built on it can be refreshed, assured or operationalised, which is why the first integration investment is a governed feed rather than a better model. The grid-specific trap is the topology join: telemetry joined to the as-built record rather than the as-operated network produces analysis of a grid that is not the one currently switched — and the error is invisible until a recommendation lands on a feeder that was reconfigured last month.
- The side screen — the most expensive furniture in the utility
- The side screen exists because it needs no change approval: rendering a dashboard next to the control room touches no operational system. That is exactly why it fails. A control engineer's attention during an event is consumed by the EMS, the OMS and the telephone; a second application on a second login receives none of it. The pattern is not a utility quirk — but the utility version has a sharp edge, because grid events are rare, high-stakes and short, so the model's entire annual value can be concentrated in the forty-eight hours its screen goes unwatched. Consultation rates during events are the number to measure, and almost nobody measures them.
- The in-system write — small engineering, structural change
- Landing model output inside the EMS, ADMS or OMS display usually amounts to one field, one column or one ranked queue — engineering measured in weeks. What changes is structural: the recommendation is now on the operator's default glance path, inside the system their procedures reference, and the act of accepting or overriding it can be captured as data. The write also crosses the OT boundary for the first time, which is why it inherits an assurance and change-approval calendar. Teams that budget for the integration but not the assurance discover the schedule belongs to the latter.
- The accept/override log — the roadmap's most undervalued dataset
- Every accepted recommendation is a labelled example of the model being trusted; every override, with its reason, is a labelled example of the boundary of that trust. Accumulated across seasons, the log answers the question stage 4 will ask: inside what conditions has this model's advice been reliably safe to follow? Envelope bounds derived from thousands of logged judgements are defensible to operations leadership, to assurance reviewers and — increasingly — to regulators. Bounds set without that log are set by committee instinct, and committee instinct either strangles the loop with conservatism or, worse, does not.
- The operating envelope — a control document, not a config screen
- A real envelope has four load-bearing parts: hard bounds validated against power-flow studies of the actual network (not generic limits); rate limits so the loop cannot thrash tap changers and switches through excessive operations; validity conditions stating when the model may act at all — topology confirmed as-operated, telemetry fresh, no active switching in the area; and auto-revert triggers that return control to the previous scheme the instant a condition fails. It is versioned like code, owned by a named person in operations, and re-validated when the network changes. The settings screen holds its enforcement values; the envelope is the document, the studies and the sign-off behind them.
- Why actuation always runs through an engineered scheme
- The model never commands a breaker, a tap changer or an inverter directly — it adjusts the aim of a scheme that was engineered to act: a VVO controller, a DERMS dispatch engine, a FLISR policy. The layering is what makes the loop assurable. The scheme keeps its deterministic interlocks and its protection coordination, so the worst case of a bad model proposal is a suboptimal-but-safe instruction the scheme was already designed to bound; protection and under-frequency schemes remain untouched, deterministic and certifiable. This is the same architecture the grid has always used to admit new control intelligence, and AI earns no exemption from it.
The five integration stages in detail
For each stage: what it looks like inside a network business, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what the move to the next stage costs.
Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a real estate — which systems a model touches, what is logged, what is drilled — and the diagnostic signals are checks you can run against your own EMS, ADMS and work-management systems this week. The anti-pattern is the specific mistake most often made trying to leave that stage, and the investment lines are stated in team terms, because in a regulated network business the schedule is set by assurance paths and people, not by compute.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Detached
26% of operators sit here
Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
Detached is not a failure state — it is where every utility correctly starts, and much genuinely useful work lives here permanently: load research, loss studies, hosting-capacity screening, post-event analysis. The problem is not that detached analytics exist. It is that programmes stay detached while believing they are integrating, because the models keep improving and improvement feels like progress. Integration depth, the thing this roadmap measures, has not moved at all: the model still cannot see the grid, and the grid still cannot hear the model.
The signature of the stage is the extract ritual. An engineer requests a historian export, someone in the metering team pulls AMI interval data for the study feeders, a GIS analyst produces a shapefile, and a week of every study is spent joining the three on asset identifiers that never quite reconcile. Each study repeats the ritual, because nothing was kept: no governed feed, no maintained mapping, no shared definition of which topology snapshot the join was made against. Ten studies cost ten times what one costs, and the eleventh will too.
What makes Detached worth leaving quickly is that its outputs age badly in a way the grid now punishes. A hosting-capacity study computed on last year's topology and last year's DER register was tolerable when connections arrived slowly. With distributed solar, storage and EV load growing at the pace the sector's own bodies document, a detached analysis is materially wrong within months — and decisions made on it inherit the error. The first integration investment is therefore not a model at all: it is one governed, refreshing feed from the historian and the AMI head-end, joined to a topology that updates.
In practice
The hosting-capacity study that was wrong by summer
A distribution utility commissioned a feeder-by-feeder hosting-capacity analysis in the winter. The work was competent: power-flow models, seasonal load shapes, DER growth scenarios. It was also an extract — GIS as-built topology, a December AMI pull, the DER register as of the study date. By July, two feeders in the study had been reconfigured, a 5 MW solar cluster had connected, and the planning team was quoting connection applicants numbers the network could no longer honour. The utility's fix was not a better study; it was a monthly refresh pipeline — the first piece of stage-2 plumbing — so the analysis tracked the grid instead of memorialising it.
What it looks like
- Models are built on historian, AMI or GIS exports assembled by hand
- No model consumes a live feed, and none could without a new connection request
- Findings are delivered as reports, slide decks and one-off studies
- The network model used for analysis is the as-built GIS view, not the grid as switched today
Diagnostic signals you can check this week
- Ask how the last analysis got its data. If the answer names people rather than pipelines, you are Detached
- Ask which topology snapshot the analysis was computed against, and whether anyone can say how old it was
- Check whether any model output is refreshed on a schedule without a human re-running it
- Ask what a model would have to do to receive live SCADA points. If the answer is 'raise a request and wait', no read path exists
Anti-pattern · Buying the platform before proving one feed
The instinctive exit from Detached is a data-platform programme: a lakehouse, an integration bus, an eighteen-month roadmap. It fails here for the same reason it fails everywhere, with a grid-specific twist — the hard part of utility data is not volume but joinability, and joinability problems only surface when a real use case tries to join real feeds. Stand up one governed feed — historian points and AMI intervals for one substation's feeders, joined to an as-operated topology — and let its problems teach you what the platform must do. The platform bought first solves the problems a vendor imagined instead.
What holds you here
No live read path exists — every analysis re-assembles the grid from exports, so nothing can be kept current and nothing downstream can ever be operational.
Highest-leverage next move
Stand up one governed, refreshing feed — historian, AMI and as-operated topology for one substation area — and rebuild one existing analysis on top of it.
Cost of leaving
- Effort
- 4–8 months
- Team
- One data engineer, one distribution planning engineer, part-time historian/SCADA support
- Risk
- Low — everything is additive; nothing operational depends on the work yet
- To next stage
- 4–8 months
If this is you, the next step is
A three-week engagement: pick the substation, map the joins, size the pipeline.
Stage 2
Observing
36% of operators sit here
Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
Observing is the most crowded stage on the ladder and the easiest to mistake for the destination. The engineering achievement is real: live feeds, a maintained topology join, models that track the network hour by hour. Load forecasts update themselves, transformer thermal models follow the weather, an anomaly detector watches feeder currents. Everything works — and none of it is wired to a decision. The output renders on a screen beside the control room, or a portal the planning team can open, and whether anyone acts on it is a matter of individual habit.
The stage has a structural failure mode that the sector's own experience documents relentlessly: the side screen fails exactly when it matters. On a quiet Tuesday, a control engineer has attention to spare and may consult the forecast portal. During a storm restoration — the hours the model was built for — the engineer is managing switching, crews and an alarm flood inside the EMS and OMS, and a second application on a second screen receives no attention at all. The model's consultation rate is highest when its value is lowest, and this inversion is invisible unless someone measures it, which at stage 2 nobody does.
The honest accounting of Observing is that it produces evidence, not value. That is worth having: a model that has tracked the live network for two quarters, with its accuracy measured against what actually happened, is exactly the artefact that justifies the integration work of stage 3. The mistake is polishing the model instead of spending that evidence. A forecast that improves from good to very good on a side screen changes nothing; the same forecast landed inside the ADMS display the operator already reads changes decisions the first week. Accuracy is the stage-2 team's comfort zone, and the write path is the exit.
In practice
The overload forecast nobody opened in the heatwave
A network operator built a transformer overload forecaster: AMI-derived loading, weather-driven thermal modelling, live topology. It ran on a web dashboard, and through spring the planning team used it appreciatively. In the July heatwave — its design case — the control room made every load-transfer decision from the EMS loading displays and their own experience, exactly as their procedures directed. Dashboard analytics later showed zero control-room sessions during the event. The model had performed well; forty-eight hours of accurate overload warnings went unseen. The next phase of the programme delivered the same forecast as a loading-forecast column inside the EMS display the engineers already used, with no model changes at all.
What it looks like
- At least one model runs continuously against streaming or frequently refreshed telemetry
- The model sees the grid as switched today, not the as-built record
- Output is a dashboard, portal or side screen — nothing lands inside the EMS, ADMS or OMS
- Telemetry gaps and feed failures are discovered by the model team, sometimes weeks late
Diagnostic signals you can check this week
- List every screen a model renders on, then ask which of them an operator has open during a storm. Usually none
- Ask whether accepts and overrides are recorded anywhere. At stage 2 there is nothing to accept — that absence is the diagnostic
- Check who is alerted when a telemetry feed stops. If it is the data team rather than a monitored SLA, reads are informal
- Ask what changed in any operational procedure because of the model. The stage-2 answer is nothing
Anti-pattern · Improving accuracy to earn adoption
When the side screen is ignored, the reflex is to make the model better, on the theory that operators will trust a sharper number. They will not, because they are not weighing its sharpness — they are working inside systems and procedures the model is simply absent from. A moderately accurate forecast inside the ADMS display outperforms an excellent one behind a separate login, because the default glance captures it. Spend the next two quarters on the write path and the operator workflow, and revisit accuracy when an accuracy point can be priced in avoided load transfers or excursion minutes.
What holds you here
Model output lives outside every operational system, so acting on it is voluntary — and voluntary attention disappears precisely during the events the models were built for.
Highest-leverage next move
Pick one decision and land the model's output inside the system that already carries it — an EMS display column, an ADMS advisory, an OMS estimate — with accepts and overrides logged from day one.
Cost of leaving
- Effort
- 6–12 months
- Team
- One integration engineer, one ML engineer, a named control-room or operations owner, EMS/ADMS vendor liaison
- Risk
- Medium — the first write into an operational display crosses the OT boundary and needs the change and assurance path walked properly
- To next stage
- 6–12 months
If this is you, the next step is
We map the integration, assurance and change-approval route before it becomes the schedule.
Stage 3
Advising
24% of operators sit here
Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
Advising is where AI first becomes part of how the grid is operated rather than commentary on it. The engineering that gets a programme here is mostly unglamorous: an interface into the EMS or ADMS that renders a model output as a field the operator already scans, a work-management integration that turns a condition score into a ranked inspection queue, an OMS hook that replaces a static restoration estimate with a predicted one. None of it is machine learning. All of it is what makes machine learning matter.
The discipline that distinguishes a real stage 3 from a cosmetic one is the accept/override log. Every recommendation shown, every one accepted, every one overridden, with enough context to reconstruct why — feeder, conditions, operator, what was done instead. This log is usually pitched as an adoption metric, and it is, but its deeper role is forward-looking: it is the calibration dataset for the operating envelope that stage 4 will need. When the time comes to let a model write setpoints, the bounds inside which it may do so are derived from thousands of logged human judgements about when its advice was safe to follow. A programme that skipped the logging has to set those bounds by committee guesswork, and committees guess conservatively enough to strangle the loop.
Advising is also where trust is actually built, and it is built asymmetrically. Operators extend trust slowly and withdraw it instantly: one visibly wrong recommendation during a major event can end consultation for a year. The mitigation is scope honesty. A restoration-time model trained on routine faults should say so on its face and widen its uncertainty during major storms — or be suppressed in those modes entirely. Teams that let a blue-sky model advise through a hurricane are spending trust they have not earned, and the sector's operational culture, correctly, does not refund it.
In practice
The inspection queue that replaced the annual list
A network operator's vegetation and asset-inspection programme ran on an annual cycle list. The AI team had condition and risk models at stage 2 for a year — scores on a portal, admired and unused. The integration project wrote the scores into the work-management system as a ranked queue with a reason code per item, and inspection planners kept their override authority: any item could be re-ranked with a logged reason. Within two seasons the override log itself became the most valuable artefact — it surfaced two systematic model blind spots (coastal corrosion, a data gap on one legacy asset class) that accuracy metrics had never caught, and the reviewed acceptance rate became the number the programme reported upward.
What it looks like
- At least one model writes into an operational display, queue or field the operator already reads
- Accepts and overrides are captured with context, and someone reviews them
- Telemetry freshness and model health are monitored, and a named person is paged
- Storm-mode behaviour is measured: the team knows the acceptance rate on bad days, not just good ones
Diagnostic signals you can check this week
- Open the operational system and look for the model's output. If finding it requires a second application, you are still Observing
- Ask for last month's acceptance rate, and whether anyone can split it by operating conditions
- Ask what a specific override taught the team, with an example. A stage-3 team has one ready
- Check that a paged, named owner exists for model health — and that the name survives a team reorganisation
Anti-pattern · Skipping the advisory year and jumping to actuation
Once the write path exists, closing the loop looks like one more integration ticket — the scheme interface is right there. Resist it. The advisory period is not caution theatre; it is the only mechanism that produces the override log the envelope needs, and the only period in which operators can watch the model be wrong safely. A loop closed on three months of clean-weather advisory data carries bounds calibrated on nothing, and its first bad action in anger will not just be reverted — it will take the advisory layer down with it. Run the advisory stage through at least one full seasonal cycle, storms included.
What holds you here
Every decision still routes through a person, so the value ceiling is operator attention — and the envelope evidence a closed loop needs only accumulates as fast as the override log grows.
Highest-leverage next move
Choose the one decision with contained consequences and high tempo — Volt/VAR setpoints are the usual answer — and start engineering its operating envelope from the override log while running shadow mode.
Cost of leaving
- Effort
- 12–18 months, dominated by envelope engineering and OT assurance rather than modelling
- Team
- Integration engineer, ML engineer, protection/automation engineer, OT security assessor, named control-room owner
- Risk
- Medium to high — the write path into a control scheme crosses the CIP or equivalent boundary and inherits its assurance calendar
- To next stage
- 12–18 months
If this is you, the next step is
We review the write path, the override log and storm-mode behaviour against what a closed loop will need.
Stage 4
Acting
11% of operators sit here
Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
Acting is narrower than the word 'autonomy' suggests, and the narrowness is the design. No serious utility lets a learned model command breakers or write to protection. What stage 4 actually looks like is a model proposing setpoints or schedules to a control scheme that was already engineered to act — a Volt/VAR optimisation adjusting regulator and capacitor settings, a DERMS issuing dispatch and curtailment within registered limits, a battery schedule optimiser. The scheme retains its own interlocks and its deterministic behaviour; the model adjusts what the scheme aims for, inside bounds a human wrote down. The grid has run closed loops for a century — protection, automatic generation control (AGC), and more recently FLISR self-healing — and it accepts them precisely because they are bounded, deterministic where it matters, and reversible. AI joins the loop on the same terms or not at all.
The central artefact is the operating envelope, and it deserves more engineering respect than it usually gets. A real envelope is not a min/max pair in a configuration screen. It is a versioned document and its enforcement code: hard bounds validated against power-flow studies of the actual network, rate limits so the loop cannot thrash equipment through excessive tap and switching operations, validity conditions that state when the model may act at all — topology as expected, telemetry fresh, no active switching in the area — and auto-revert triggers that return control to the previous scheme the moment any condition fails. Every element traces to evidence: the bounds to network studies, the validity conditions to the failure modes observed in the advisory year, the escalation thresholds to the override log.
The operational job changes accordingly. Supervising a closed loop is not watching its every action — that would just be stage 3 with worse ergonomics — it is watching its envelope statistics. The escalation rate (how often proposals fall outside bounds) is the leading indicator: rising escalations mean the network has moved away from the envelope's assumptions, through DER growth, reconfiguration or seasonal change, and the envelope needs review before an incident forces one. This is also where the security obligations bite hardest: a component that writes into a control scheme sits inside the OT boundary, and in North America the NERC CIP standards govern how it is built, patched, accessed and monitored. That is not an obstacle erected against AI — it is the same regime every control component lives under, and the roadmap's earlier stages exist partly so the assured path is built once, properly, before the loop needs it.
In practice
The envelope that caught the network changing
A distribution utility ran AI-recommended Volt/VAR schedules on a pilot feeder group through an existing VVO controller. The envelope held voltage inside statutory limits with margin, capped daily tap operations, and required as-operated topology confirmation before each schedule push. For two quarters the loop ran quietly, with escalations under two per cent. Then escalations tripled in a month. Investigation found no model fault: a wave of new rooftop solar connections had shifted the feeders' daytime voltage profile, and proposals that had been routine now pressed the upper bound. The envelope review brought the bounds and the model's training window up to date together — an afternoon of governance that, without the escalation signal, would have been an overvoltage investigation instead.
What it looks like
- At least one closed loop is live: model setpoints or schedules flow into an engineered scheme such as VVO or DERMS dispatch
- A written, versioned operating envelope exists, validated against power flow, with rate limits and validity conditions
- Out-of-envelope proposals escalate to an operator; envelope breaches auto-revert and page
- Reversion to the previous control source is drilled on a calendar, not discovered in an incident
Diagnostic signals you can check this week
- Ask to see the envelope document and its version history. A settings screenshot is not an envelope
- Ask when reversion to the previous control source was last drilled, and whether it was a calendar drill or an incident
- Ask for the escalation-rate trend over six months, and what the last rise turned out to mean
- Ask how a single automated setpoint from three months ago would be reconstructed — model version, envelope version, inputs, action
Anti-pattern · Widening the envelope because the loop has been quiet
After six clean months the pressure arrives from both directions: the model team wants headroom, operations has relaxed, and the envelope looks like conservatism tax. Widening it because nothing has gone wrong is inverted reasoning — nothing has gone wrong inside bounds that were validated; outside them is unvalidated by definition. Envelopes widen the way they were set: a power-flow study of the proposed bounds, a review of escalated proposals showing the widened region would have been safe, a new version with sign-off. Quiet operation is evidence the envelope works, not evidence it is unnecessary — the distinction funds the whole stage.
What holds you here
Each loop's envelope, assurance case and scheme integration is built as a one-off, so the second loop costs nearly what the first did and the portfolio stalls at one or two.
Highest-leverage next move
Extract the shared machinery — envelope validation service, scheme interfaces, action audit log, reversion pattern — so the next loop is an envelope and a model, not a programme.
Cost of leaving
- Effort
- 18–30 months to a portfolio of loops; each additional loop far cheaper on the assured path
- Team
- Platform and integration engineers, protection/automation engineer, OT security assessor, envelope owner in operations, standing review forum
- Risk
- Concentrated — low frequency, high consequence; the governance and evidence obligations are the schedule, not the modelling
- To next stage
- 18–30 months
If this is you, the next step is
We review bounds, validity conditions, reversion and audit trail against a real scenario on your network.
Stage 5
Coordinating
3% of operators sit here
Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.
Coordinating is the stage vendor slides call the self-driving grid, and the honest version is smaller and more valuable than that. What actually exists at stage 5 is a portfolio: a handful of supervised closed loops in the contained, high-tempo corners of the network, a wider ring of advisory services, and — the part that earns the stage its name — shared machinery underneath them. One governed feed layer, one envelope-validation service, one action audit log, one reversion pattern, one assured OT path that every new component reuses. The marginal cost of the next loop collapses, which is the economic point, and every loop behaves identically under failure, which is the operational one.
The genuinely new engineering problem at this stage is interaction. Two loops that are each safe alone are not automatically safe together: a Volt/VAR loop and a DER dispatch loop both influence feeder voltage, and their envelopes must be validated jointly, not merely side by side — the same discipline protection engineers have always applied to interacting schemes, extended to learned controllers. This is also the honest boundary of the stage. System-level questions — can learned models one day assist constraint management across a whole transmission network? — are live research at bodies like EPRI, whose Open Power AI Consortium exists precisely because no single operator can validate such systems alone. A credible roadmap treats those as research to track, not line items to schedule.
What keeps a utility at stage 5 is governance rhythm rather than technology. Envelopes are reviewed on a calendar and re-validated when the network changes, because the network now changes constantly: every DER connection wave shifts the assumptions a bound was set under. Model risk management runs as a standing discipline with the shape that frameworks such as the NIST AI Risk Management Framework and ISO/IEC 42001 formalise — inventory, ownership, review cadence, evidence retention — applied with control-engineering seriousness. And the reporting closes the loop institutionally: AI-driven actions appear in reliability and regulatory reporting with the same reconstructability as any switching operation, which is what makes the capability durable across leadership changes, audits and control periods rather than dependent on the team that built it.
In practice
The joint envelope review that found the interaction
A utility running an AI-informed VVO loop brought a DER flexibility-dispatch loop to the same feeder areas. Each had a validated envelope; the commissioning gate required a joint study. The study found a plausible evening sequence — dispatch instructing battery export while the VVO held a high-voltage schedule set that morning — that could press upper voltage limits at two weakly monitored points. The fix was a shared constraint in the envelope-validation service: dispatch proposals in those areas are checked against the active VVO schedule before release. The sequence has since occurred twice in operation; both times the joint check reshaped the dispatch and no operator needed to intervene. The interaction would never have appeared in either loop's individual assurance case.
What it looks like
- Multiple loops — VVO, DER dispatch, storm pre-positioning, ratings — run from shared feeds, envelope validation and audit machinery
- Envelope and model reviews are keyed to network-change events: DER connection waves, reinforcements, reconfigurations
- Loops are coordinated: their envelopes are checked jointly so two safe controllers cannot combine into an unsafe pair
- AI actions are reconstructable years later and reported through the same reliability lens as any control action
Diagnostic signals you can check this week
- Ask how many loops share one envelope-validation service and one audit log, versus carrying their own
- Ask when envelopes were last reviewed jointly rather than one at a time
- Pick an automated action from two years ago and time how long full reconstruction takes — model, envelope, inputs, outcome
- Ask how the last major network change — a reinforcement, a DER wave — propagated into envelope and model reviews, and how long it took
Anti-pattern · Chasing loop count as the maturity metric
Once the machinery makes new loops cheap, the tempting scoreboard is how many are live, and the portfolio drifts toward decisions that were never good closed-loop candidates — high-consequence switching, storm restoration sequencing — because they are what remains. The matrix on this page exists for exactly this moment: some grid decisions earn a loop, some earn advisory support forever, and some belong to deterministic engineering permanently. A mature stage 5 holds the line and reports value per decision, not loops per year; the discipline of not closing a loop is the same discipline that made the closed ones safe.
What holds you here
Sustaining the stage is a standing governance cost — joint validation, review calendars, evidence retention — and it decays quietly the first year attention moves elsewhere.
Highest-leverage next move
Key every envelope and model review to network-change events, review interacting loops jointly, and report AI actions through the same reliability lens as any other control action.
Cost of leaving
- Effort
- Continuous — the work becomes governance rhythm, joint validation and evidence retention
- Team
- Platform team, envelope governance forum, protection/automation engineering, OT security, regulatory reporting partner
- Risk
- Concentrated and systemic — interaction effects and governance decay, not individual model failure
If this is you, the next step is
We stress-test interacting envelopes and the reconstruction trail across your live loops.
Where utilities actually sit on the integration ladder
The distribution across the five stages, the side-screen plateau at stage 2, and the external evidence for what the far side is worth.
Most utilities sit at stage 2: models genuinely connected to live grid data, with output that never enters an operational system. The distribution below is illustrative — synthesised from the sector research linked beneath it rather than measured from a single survey — but its shape is the consistent finding of every serious study of utility AI: broad experimentation, a crowded read-only middle, and a small minority whose models can act on the grid under supervision.
Distribution of utilities across the five integration stages
Stage 2 is the mode and the plateau: the side screen needs no change approval, so programmes accumulate there. The sharpest drop is into stage 4 — the envelope, assurance and scheme-integration work that most programmes have never scoped. Distribution is illustrative, charted from the synthesis described in the source.
Share of utilities
- 26% — 1 · Detached
- 36% — 2 · Observing (the side-screen plateau)
- 24% — 3 · Advising
- 11% — 4 · Acting
- 3% — 5 · Coordinating
Source: Illustrative distribution, synthesised from IEA Energy and AI and EPRI sector research
The context makes the plateau expensive. The grid the models must integrate with is being rebuilt around them: the IEA estimates that meeting national climate targets means adding or refurbishing 80 million km of grids by 2040 (opens in a new tab) — equal to the entire existing global grid — while Eurelectric's Grids for Speed (opens in a new tab) finds 30% of European grids already over forty years old and distribution investment needing to reach €67 billion a year to 2050. On the supply side, IRENA recorded 585 GW of renewable capacity added in 2024 alone (opens in a new tab) — 92.5% of all new capacity — most of it variable and much of it connecting at distribution voltage, and US electricity demand is growing again (opens in a new tab) after two flat decades. Every one of those trends multiplies the decisions the grid must make per hour; none of them adds operators to make them. That gap — more decisions, same people — is what the advisory and actuation stages exist to close, and research consortia such as EPRI's Open Power AI Consortium (opens in a new tab) exist because the sector knows it must close the gap collectively, with validation no single operator can produce alone.
The integration boundary map: what AI reads and writes, system by system
The grid control stack as an integration surface — what a model may consume from each system, what it may ever write back, at which stage, and the safeguard that makes the write safe.
AI integrates with the grid through a specific, enumerable set of systems, and each has its own read surface, write surface and safeguard. The map below is the page's working tool: find each system you run, read what a model may take from it and what — if anything — a model may ever write back, and note the stage at which each write becomes legitimate. Two features deserve emphasis before the table. First, the write column is short by design: most systems on the stack are read-rich and write-forbidden, and a roadmap that proposes writes everywhere has not understood the stack. Second, one row has no write stage at all — protection — and its permanence is a feature of the roadmap, not a limitation of it.
| System | AI reads (from stage) | AI writes (from stage) | The safeguard that makes the write safe |
|---|---|---|---|
| Historian | Point history, event logs (stage 1) | Never — the historian is the record | Not applicable; derived datasets live outside it |
| SCADA / EMS | Live points, alarms, state estimates (stage 2) | Display fields and advisories (stage 3); ratings and schedule inputs (stage 4) | Assured OT path (CIP), display-only separation, envelope validation for anything a scheme consumes |
| ADMS / DMS | As-operated topology, loading, switching state (stage 2) | Advisory fields and ranked suggestions (stage 3); VVO setpoint schedules (stage 4) | Envelope with power-flow-validated bounds, rate limits, auto-revert to static schedules |
| OMS | Outage events, restoration progress (stage 2) | Predicted restoration times and damage estimates, flagged as predictions (stage 3) | Uncertainty display, storm-mode scope honesty, operator override |
| DERMS | DER registrations, availability, response history (stage 2) | Dispatch and curtailment schedules within registered limits (stage 4) | Registered device limits, joint envelope check against VVO, auto-revert |
| AMI head-end | Interval loads, voltages, outage flags (stage 1–2) | Never for control — read-only by design | Not applicable; AMI-derived signals act only via ADMS/DERMS schemes |
| GIS & network model | Assets, connectivity, as-built record (stage 1) | Suggested corrections routed to data stewards (stage 3) | Human-reviewed correction queue — models never edit the model of record directly |
| Work & asset management | Asset condition, work history, inspection results (stage 1) | Ranked inspection and maintenance queues with reason codes (stage 3) | Planner override authority, logged re-ranking, periodic blind-spot review |
| Protection & UFLS relays | Event records and disturbance data, for analysis only (stage 2) | Never — settings change through protection engineering | Deterministic domain: AI may inform offline setting studies that engineers then apply through the certified process |
Everything in the write column lives inside the OT security perimeter, and the perimeter has a name in each jurisdiction: in North America it is the NERC Critical Infrastructure Protection standards (opens in a new tab), and equivalent network-and-information-security regimes apply elsewhere. The practical consequence for the roadmap is that the first write is slow and every subsequent write is cheap — the zoning, patching, access and monitoring pattern is negotiated once, documented, and reused. Governance frameworks are converging on the same layered picture from the model side: the NIST AI Risk Management Framework (opens in a new tab) and the ISO/IEC 42001 management-system standard both formalise the inventory-ownership-review-evidence cycle that the envelope discipline on this page implements in control-engineering terms. Regulators are pushing the stack in the same direction from a third side: the US federal regulator FERC (opens in a new tab) now requires transmission providers to use ambient-adjusted line ratings under Order 881 — a mandated move from static assumptions to model-informed operations that lands squarely in this table's EMS row, and a preview of how model-informed values will keep arriving inside the control stack whether a utility has a roadmap or not.
Which grid decisions earn the closed loop — and which never should
Two axes — decision tempo and consequence of a wrong action — sort the grid's decisions into four quadrants, and only one of them is where the loop closes first.
The decisions that earn closed-loop AI first are the ones that are both fast and forgiving: high-tempo enough that human approval is the bottleneck, and contained enough that a wrong action is a correctable inefficiency rather than a safety event. Plotting tempo against consequence sorts the grid's decision landscape cleanly, and the sort does real work in a roadmap review — most disagreements about 'how far should AI go' dissolve once the specific decision under discussion is placed on the matrix rather than argued in the abstract.
The closed-loop candidacy matrix
Place each candidate decision by its tempo and the consequence of a wrong action. The top-left quadrant is where loops close first; the top-right stays deterministic permanently — and holding that line is what makes the rest of the roadmap trustworthy.
Close the loop here first
- Volt/VAR setpoint schedules within statutory limits
- DER dispatch and curtailment inside registered device limits
- Battery charge/discharge scheduling against forecasts
- Wrong is inefficient, visible and auto-revertible
Engineered schemes only — permanently
- Protection operation and settings, under-frequency load shedding
- AGC and primary frequency response
- AI may inform offline studies; certified engineering applies them
- Determinism here is what makes loops elsewhere acceptable
Advisory is enough
- Maintenance and inspection prioritisation
- Hosting-capacity screening and connection studies
- Reinforcement and planning analysis
- Value flows fully without any actuation path
Human decides, AI assembles
- Switching plans and outage restoration sequencing
- Storm crew pre-positioning and load transfers
- AI drafts, checks and ranks; the operator commits
- The accept/override log here is permanent, by design
Tempo decides where the human adds nothing
A Volt/VAR schedule adjusting through the day, or DER dispatch tracking a five-minute market, outruns any approval workflow — the human in that loop is latency, not judgement. Conversely, a switching plan built over twenty minutes has room for a person, and the person carries context no model sees: the crew on the ground, the customer on life support, the substation with the known defect.
Consequence decides where the human is the point
Restoration sequencing and load transfers are decisions the organisation must be able to defend afterwards, name by name. Advisory AI makes them faster and better-informed; closing the loop on them would trade accountability for a latency gain nobody asked for. The top-right quadrant goes further: protection and frequency response are certified deterministic domains, and their untouchability is a public commitment, not a technical shortfall.
The matrix is a sequence, not just a sorting
Programmes that respect it build trust in a defensible order: advisory value in the bottom half funds the work, the top-left loop proves the envelope discipline, and the top-right line — never crossed — is what lets operations leadership, boards and regulators extend permission for the rest. Programmes that violate the order, usually by proposing automation in the bottom-right first because the savings look large, spend years rebuilding the trust one proposal cost.

_case_study.webp)