Redefining Technology

Energy & UtilitiesReadiness & Transformation Roadmap

The grid roadmap for AI integration: how utilities move models from reports into the control loop

A grid roadmap for AI integration is the staged path by which a utility moves AI from offline analytics to supervised closed-loop operation. Progress is measured in integration depth — what models may read, where their output lands, and what they may write back to the grid — and gated by telemetry, operator trust and the OT security boundary.

Generated scene: a grid control room with engineers in high-visibility vests at a lit network planning table, in front of a curved video wall showing grid topology, load curves and analytics panels
Energy & Utilities · Readiness & Transformation Roadmap

Key takeaways

  1. A grid roadmap for AI integration measures progress in integration depth, not model quality: what the model may read (stages 1–2), where its output lands (stage 3), and what it may write back to the grid (stages 4–5). Most programmes report model progress while standing still on all three.
  2. The grid already runs closed loops — protection, automatic generation control, FLISR self-healing — and it accepts them because they are engineered, bounded and reversible. AI enters the loop the same way: by adjusting the setpoints and policies of engineered schemes inside a validated operating envelope, never by direct device control.
  3. Most utility AI is stuck at stage 2: models reading live telemetry and rendering onto side screens nobody consults during a storm. The move to stage 3 is a workflow and write-path problem — landing output inside the EMS, ADMS or OMS with accepts and overrides logged — not an accuracy problem.
  4. The operating envelope is the central artefact of the far stages: versioned, power-flow-validated bounds with rate limits, validity conditions and auto-revert triggers. The accept/override log collected at stage 3 is the dataset that calibrates it — skip the logging and the envelope is guesswork.
  5. The binding constraints on closed-loop AI are the OT security boundary (NERC CIP and its equivalents) and the currency of the network model — both buildable today. Closed-loop readiness is almost entirely present-tense integration discipline, which is why the roadmap starts at the historian, not at the model.

Abbreviations used on this page

SCADA
Supervisory control and data acquisition
EMS
Energy management system (transmission control room)
ADMS
Advanced distribution management system
OMS
Outage management system
DERMS
Distributed energy resource management system
AMI
Advanced metering infrastructure (smart meters and the head-end system)
PMU
Phasor measurement unit — a synchrophasor sensor sampling the waveform many times a second
DER
Distributed energy resources — rooftop solar, batteries, EV chargers, flexible load
VVO
Volt/VAR optimisation — coordinated control of voltage regulators and capacitor banks
FLISR
Fault location, isolation and service restoration — the self-healing switching scheme
OT
Operational technology — the control systems and networks that act on the physical grid
CIP
Critical Infrastructure Protection — the NERC cyber-security standards for the bulk power system

Free · 8 questions · ~3 minutes

Score your integration depth

Eight questions, one at a time, about three minutes. Answer them and we build your personalised integration report — your stage on the ladder, your score on each of the four dimensions, and the specific boundary currently holding your models out of the loop — and send it to your inbox. Your result doubles as the integration baseline for your next roadmap review.

0 of 8 answered

Question 1 of 8Telemetry & network model

How would a model obtain the current state of one feeder — loading, voltage, connected DER, topology?

Integration depth is capped by the read path. A model that cannot see the grid live can never do more than write reports about its past.

How the score maps to a stage
  • 04 — Stage 1, Detached. Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
  • 59 — Stage 2, Observing. Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
  • 1014 — Stage 3, Advising. Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
  • 1519 — Stage 4, Acting. Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
  • 2024 — Stage 5, Coordinating. Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.

What a grid roadmap for AI integration actually is

A definition, the axis the five stages sit on, and the three paths — read, advise, act — a model output can take toward the grid.

A grid roadmap for AI integration is the staged sequence by which a utility deepens the connection between its AI models and the grid itself: first a read path (live telemetry and an as-operated network model), then an advisory path (output landing inside the EMS, ADMS, OMS and work queues where people decide), and finally an actuation path (setpoints and schedules flowing into engineered control schemes inside validated operating envelopes). Its unit of progress is integration depth — what the model may read, where its output lands, what it may write — not model accuracy, which improves on a different and much easier axis.

The framing matters because the grid is not an ordinary software estate. It already runs closed loops, and has for a century: protection clears faults in milliseconds, automatic generation control (AGC) balances the system continuously, and FLISR schemes reroute distribution power without a human in the loop. The grid's operating culture accepts these because they are engineered, bounded, deterministic where it matters, and reversible. That is the standard AI must meet — and it recasts the roadmap's destination. The far stage is not 'AI runs the grid'; it is learned models adjusting the aims of engineered schemes inside bounds a human wrote down and validated. Meanwhile the pressure to get there is real and externally documented: the IEA finds that AI-based fault detection can cut outage durations by 30–50% (opens in a new tab), and that scaled AI-led interventions could save around 300 TWh of electricity globally — value that only exists on the far side of the integration work this page maps.

Two things this roadmap deliberately is not. It is not a funding and submission calendar — how the stages below are paid for, and how they cut against price-control and rate-case windows, is the subject of the AI transformation timeline for utilities. And it is not a generic maturity model with the word 'grid' inserted: every stage boundary on this page is a physical or institutional boundary of the electricity system — the historian's edge, the control-room display, the OT security perimeter, the engineered control scheme. A roadmap that could be re-badged for a retailer or a bank by find-and-replace is measuring the wrong thing.

Decisions AI can safely touch, against integration depth

The curve is flat through stages 1–2 because a model that cannot reach an operational system touches no decisions at all, however accurate it becomes. It inflects at stage 3, when output lands where operators already work, and steepens at stage 4, when contained high-tempo decisions close the loop. The far end flattens honestly: some grid decisions belong to deterministic engineering permanently.

Grid decisions safely reachable by stage

  • Stage 1 · Detached — 26% of operators. Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.
  • Stage 2 · Observing — 36% of operators. Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.
  • Stage 3 · Advising — 24% of operators. Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.
  • Stage 4 · Acting — 11% of operators. Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.
  • Stage 5 · Coordinating — 3% of operators. Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.

Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with IEA, Energy and AI.

Read, advise, act: the three paths a model output can take toward the grid

The three integration paths, top to bottom in order of depth. The read path ends at a side screen — where most utility AI stops. The advisory path lands output inside the operational systems and logs every accept and override. The actuation path runs through envelope validation into an engineered scheme, with supervision and auto-revert around it — a model never touches a device directly.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • On the read path, historian and AMI extracts feed an offline model whose output lands in a report or on a side screen outside every operational system. Whether anyone acts on it is voluntary, and voluntary attention disappears during storms — which is when the model was supposed to matter. Stages 1 and 2 live entirely on this path; so does most utility AI today.
  • On the advisory path, live telemetry joined to an as-operated network model feeds a served, monitored model whose output is written into the EMS, ADMS or OMS display — or the work queue — that the operator already reads. The operator decides, and every accept and override is logged with context. That log is simultaneously the adoption metric, the trust record, and the calibration dataset the actuation path will later need.
  • On the actuation path, a model's proposed setpoint or schedule passes an envelope and power-flow check before anything moves: hard bounds, rate limits and validity conditions, all versioned. In-bounds proposals flow into an engineered scheme — Volt/VAR optimisation, DERMS dispatch — which actuates with its own interlocks intact. Out-of-bounds proposals escalate to the operator. Supervision watches the escalation rate as a leading indicator, every action is logged for reconstruction, and any breach auto-reverts to the previous control source.
Step-by-step insights
The extract ceiling — why the read path caps everything above it
A hand-assembled extract is stale the moment it lands, carries no lineage, and embeds one analyst's private join decisions between historian tags, AMI meters and GIS assets. Nothing built on it can be refreshed, assured or operationalised, which is why the first integration investment is a governed feed rather than a better model. The grid-specific trap is the topology join: telemetry joined to the as-built record rather than the as-operated network produces analysis of a grid that is not the one currently switched — and the error is invisible until a recommendation lands on a feeder that was reconfigured last month.
The side screen — the most expensive furniture in the utility
The side screen exists because it needs no change approval: rendering a dashboard next to the control room touches no operational system. That is exactly why it fails. A control engineer's attention during an event is consumed by the EMS, the OMS and the telephone; a second application on a second login receives none of it. The pattern is not a utility quirk — but the utility version has a sharp edge, because grid events are rare, high-stakes and short, so the model's entire annual value can be concentrated in the forty-eight hours its screen goes unwatched. Consultation rates during events are the number to measure, and almost nobody measures them.
The in-system write — small engineering, structural change
Landing model output inside the EMS, ADMS or OMS display usually amounts to one field, one column or one ranked queue — engineering measured in weeks. What changes is structural: the recommendation is now on the operator's default glance path, inside the system their procedures reference, and the act of accepting or overriding it can be captured as data. The write also crosses the OT boundary for the first time, which is why it inherits an assurance and change-approval calendar. Teams that budget for the integration but not the assurance discover the schedule belongs to the latter.
The accept/override log — the roadmap's most undervalued dataset
Every accepted recommendation is a labelled example of the model being trusted; every override, with its reason, is a labelled example of the boundary of that trust. Accumulated across seasons, the log answers the question stage 4 will ask: inside what conditions has this model's advice been reliably safe to follow? Envelope bounds derived from thousands of logged judgements are defensible to operations leadership, to assurance reviewers and — increasingly — to regulators. Bounds set without that log are set by committee instinct, and committee instinct either strangles the loop with conservatism or, worse, does not.
The operating envelope — a control document, not a config screen
A real envelope has four load-bearing parts: hard bounds validated against power-flow studies of the actual network (not generic limits); rate limits so the loop cannot thrash tap changers and switches through excessive operations; validity conditions stating when the model may act at all — topology confirmed as-operated, telemetry fresh, no active switching in the area; and auto-revert triggers that return control to the previous scheme the instant a condition fails. It is versioned like code, owned by a named person in operations, and re-validated when the network changes. The settings screen holds its enforcement values; the envelope is the document, the studies and the sign-off behind them.
Why actuation always runs through an engineered scheme
The model never commands a breaker, a tap changer or an inverter directly — it adjusts the aim of a scheme that was engineered to act: a VVO controller, a DERMS dispatch engine, a FLISR policy. The layering is what makes the loop assurable. The scheme keeps its deterministic interlocks and its protection coordination, so the worst case of a bad model proposal is a suboptimal-but-safe instruction the scheme was already designed to bound; protection and under-frequency schemes remain untouched, deterministic and certifiable. This is the same architecture the grid has always used to admit new control intelligence, and AI earns no exemption from it.

The five integration stages in detail

For each stage: what it looks like inside a network business, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what the move to the next stage costs.

Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a real estate — which systems a model touches, what is logged, what is drilled — and the diagnostic signals are checks you can run against your own EMS, ADMS and work-management systems this week. The anti-pattern is the specific mistake most often made trying to leave that stage, and the investment lines are stated in team terms, because in a regulated network business the schedule is set by assurance paths and people, not by compute.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Detached

26% of operators sit here

Detached means AI and the grid do not share a wire: models run on historian exports and metering extracts, and their output ends in documents rather than in any operational system.

Detached is not a failure state — it is where every utility correctly starts, and much genuinely useful work lives here permanently: load research, loss studies, hosting-capacity screening, post-event analysis. The problem is not that detached analytics exist. It is that programmes stay detached while believing they are integrating, because the models keep improving and improvement feels like progress. Integration depth, the thing this roadmap measures, has not moved at all: the model still cannot see the grid, and the grid still cannot hear the model.

The signature of the stage is the extract ritual. An engineer requests a historian export, someone in the metering team pulls AMI interval data for the study feeders, a GIS analyst produces a shapefile, and a week of every study is spent joining the three on asset identifiers that never quite reconcile. Each study repeats the ritual, because nothing was kept: no governed feed, no maintained mapping, no shared definition of which topology snapshot the join was made against. Ten studies cost ten times what one costs, and the eleventh will too.

What makes Detached worth leaving quickly is that its outputs age badly in a way the grid now punishes. A hosting-capacity study computed on last year's topology and last year's DER register was tolerable when connections arrived slowly. With distributed solar, storage and EV load growing at the pace the sector's own bodies document, a detached analysis is materially wrong within months — and decisions made on it inherit the error. The first integration investment is therefore not a model at all: it is one governed, refreshing feed from the historian and the AMI head-end, joined to a topology that updates.

In practice

The hosting-capacity study that was wrong by summer

A distribution utility commissioned a feeder-by-feeder hosting-capacity analysis in the winter. The work was competent: power-flow models, seasonal load shapes, DER growth scenarios. It was also an extract — GIS as-built topology, a December AMI pull, the DER register as of the study date. By July, two feeders in the study had been reconfigured, a 5 MW solar cluster had connected, and the planning team was quoting connection applicants numbers the network could no longer honour. The utility's fix was not a better study; it was a monthly refresh pipeline — the first piece of stage-2 plumbing — so the analysis tracked the grid instead of memorialising it.

What it looks like

  • Models are built on historian, AMI or GIS exports assembled by hand
  • No model consumes a live feed, and none could without a new connection request
  • Findings are delivered as reports, slide decks and one-off studies
  • The network model used for analysis is the as-built GIS view, not the grid as switched today

Diagnostic signals you can check this week

  • Ask how the last analysis got its data. If the answer names people rather than pipelines, you are Detached
  • Ask which topology snapshot the analysis was computed against, and whether anyone can say how old it was
  • Check whether any model output is refreshed on a schedule without a human re-running it
  • Ask what a model would have to do to receive live SCADA points. If the answer is 'raise a request and wait', no read path exists

Anti-pattern · Buying the platform before proving one feed

The instinctive exit from Detached is a data-platform programme: a lakehouse, an integration bus, an eighteen-month roadmap. It fails here for the same reason it fails everywhere, with a grid-specific twist — the hard part of utility data is not volume but joinability, and joinability problems only surface when a real use case tries to join real feeds. Stand up one governed feed — historian points and AMI intervals for one substation's feeders, joined to an as-operated topology — and let its problems teach you what the platform must do. The platform bought first solves the problems a vendor imagined instead.

What holds you here

No live read path exists — every analysis re-assembles the grid from exports, so nothing can be kept current and nothing downstream can ever be operational.

Highest-leverage next move

Stand up one governed, refreshing feed — historian, AMI and as-operated topology for one substation area — and rebuild one existing analysis on top of it.

Cost of leaving

Effort
4–8 months
Team
One data engineer, one distribution planning engineer, part-time historian/SCADA support
Risk
Low — everything is additive; nothing operational depends on the work yet
To next stage
4–8 months

If this is you, the next step is

A three-week engagement: pick the substation, map the joins, size the pipeline.

Scope your first governed feed

Stage 2

Observing

36% of operators sit here

Observing means models consume live grid telemetry — SCADA, AMI, weather — joined to an as-operated network model, but their output renders on side screens and portals outside every operational system.

Observing is the most crowded stage on the ladder and the easiest to mistake for the destination. The engineering achievement is real: live feeds, a maintained topology join, models that track the network hour by hour. Load forecasts update themselves, transformer thermal models follow the weather, an anomaly detector watches feeder currents. Everything works — and none of it is wired to a decision. The output renders on a screen beside the control room, or a portal the planning team can open, and whether anyone acts on it is a matter of individual habit.

The stage has a structural failure mode that the sector's own experience documents relentlessly: the side screen fails exactly when it matters. On a quiet Tuesday, a control engineer has attention to spare and may consult the forecast portal. During a storm restoration — the hours the model was built for — the engineer is managing switching, crews and an alarm flood inside the EMS and OMS, and a second application on a second screen receives no attention at all. The model's consultation rate is highest when its value is lowest, and this inversion is invisible unless someone measures it, which at stage 2 nobody does.

The honest accounting of Observing is that it produces evidence, not value. That is worth having: a model that has tracked the live network for two quarters, with its accuracy measured against what actually happened, is exactly the artefact that justifies the integration work of stage 3. The mistake is polishing the model instead of spending that evidence. A forecast that improves from good to very good on a side screen changes nothing; the same forecast landed inside the ADMS display the operator already reads changes decisions the first week. Accuracy is the stage-2 team's comfort zone, and the write path is the exit.

In practice

The overload forecast nobody opened in the heatwave

A network operator built a transformer overload forecaster: AMI-derived loading, weather-driven thermal modelling, live topology. It ran on a web dashboard, and through spring the planning team used it appreciatively. In the July heatwave — its design case — the control room made every load-transfer decision from the EMS loading displays and their own experience, exactly as their procedures directed. Dashboard analytics later showed zero control-room sessions during the event. The model had performed well; forty-eight hours of accurate overload warnings went unseen. The next phase of the programme delivered the same forecast as a loading-forecast column inside the EMS display the engineers already used, with no model changes at all.

What it looks like

  • At least one model runs continuously against streaming or frequently refreshed telemetry
  • The model sees the grid as switched today, not the as-built record
  • Output is a dashboard, portal or side screen — nothing lands inside the EMS, ADMS or OMS
  • Telemetry gaps and feed failures are discovered by the model team, sometimes weeks late

Diagnostic signals you can check this week

  • List every screen a model renders on, then ask which of them an operator has open during a storm. Usually none
  • Ask whether accepts and overrides are recorded anywhere. At stage 2 there is nothing to accept — that absence is the diagnostic
  • Check who is alerted when a telemetry feed stops. If it is the data team rather than a monitored SLA, reads are informal
  • Ask what changed in any operational procedure because of the model. The stage-2 answer is nothing

Anti-pattern · Improving accuracy to earn adoption

When the side screen is ignored, the reflex is to make the model better, on the theory that operators will trust a sharper number. They will not, because they are not weighing its sharpness — they are working inside systems and procedures the model is simply absent from. A moderately accurate forecast inside the ADMS display outperforms an excellent one behind a separate login, because the default glance captures it. Spend the next two quarters on the write path and the operator workflow, and revisit accuracy when an accuracy point can be priced in avoided load transfers or excursion minutes.

What holds you here

Model output lives outside every operational system, so acting on it is voluntary — and voluntary attention disappears precisely during the events the models were built for.

Highest-leverage next move

Pick one decision and land the model's output inside the system that already carries it — an EMS display column, an ADMS advisory, an OMS estimate — with accepts and overrides logged from day one.

Cost of leaving

Effort
6–12 months
Team
One integration engineer, one ML engineer, a named control-room or operations owner, EMS/ADMS vendor liaison
Risk
Medium — the first write into an operational display crosses the OT boundary and needs the change and assurance path walked properly
To next stage
6–12 months

If this is you, the next step is

We map the integration, assurance and change-approval route before it becomes the schedule.

Plan the write path into your EMS or ADMS

Stage 3

Advising

24% of operators sit here

Advising means model output lands inside the operational systems — EMS and ADMS displays, OMS estimates, work-management queues — where a person decides, and every accept and override is logged.

Advising is where AI first becomes part of how the grid is operated rather than commentary on it. The engineering that gets a programme here is mostly unglamorous: an interface into the EMS or ADMS that renders a model output as a field the operator already scans, a work-management integration that turns a condition score into a ranked inspection queue, an OMS hook that replaces a static restoration estimate with a predicted one. None of it is machine learning. All of it is what makes machine learning matter.

The discipline that distinguishes a real stage 3 from a cosmetic one is the accept/override log. Every recommendation shown, every one accepted, every one overridden, with enough context to reconstruct why — feeder, conditions, operator, what was done instead. This log is usually pitched as an adoption metric, and it is, but its deeper role is forward-looking: it is the calibration dataset for the operating envelope that stage 4 will need. When the time comes to let a model write setpoints, the bounds inside which it may do so are derived from thousands of logged human judgements about when its advice was safe to follow. A programme that skipped the logging has to set those bounds by committee guesswork, and committees guess conservatively enough to strangle the loop.

Advising is also where trust is actually built, and it is built asymmetrically. Operators extend trust slowly and withdraw it instantly: one visibly wrong recommendation during a major event can end consultation for a year. The mitigation is scope honesty. A restoration-time model trained on routine faults should say so on its face and widen its uncertainty during major storms — or be suppressed in those modes entirely. Teams that let a blue-sky model advise through a hurricane are spending trust they have not earned, and the sector's operational culture, correctly, does not refund it.

In practice

The inspection queue that replaced the annual list

A network operator's vegetation and asset-inspection programme ran on an annual cycle list. The AI team had condition and risk models at stage 2 for a year — scores on a portal, admired and unused. The integration project wrote the scores into the work-management system as a ranked queue with a reason code per item, and inspection planners kept their override authority: any item could be re-ranked with a logged reason. Within two seasons the override log itself became the most valuable artefact — it surfaced two systematic model blind spots (coastal corrosion, a data gap on one legacy asset class) that accuracy metrics had never caught, and the reviewed acceptance rate became the number the programme reported upward.

What it looks like

  • At least one model writes into an operational display, queue or field the operator already reads
  • Accepts and overrides are captured with context, and someone reviews them
  • Telemetry freshness and model health are monitored, and a named person is paged
  • Storm-mode behaviour is measured: the team knows the acceptance rate on bad days, not just good ones

Diagnostic signals you can check this week

  • Open the operational system and look for the model's output. If finding it requires a second application, you are still Observing
  • Ask for last month's acceptance rate, and whether anyone can split it by operating conditions
  • Ask what a specific override taught the team, with an example. A stage-3 team has one ready
  • Check that a paged, named owner exists for model health — and that the name survives a team reorganisation

Anti-pattern · Skipping the advisory year and jumping to actuation

Once the write path exists, closing the loop looks like one more integration ticket — the scheme interface is right there. Resist it. The advisory period is not caution theatre; it is the only mechanism that produces the override log the envelope needs, and the only period in which operators can watch the model be wrong safely. A loop closed on three months of clean-weather advisory data carries bounds calibrated on nothing, and its first bad action in anger will not just be reverted — it will take the advisory layer down with it. Run the advisory stage through at least one full seasonal cycle, storms included.

What holds you here

Every decision still routes through a person, so the value ceiling is operator attention — and the envelope evidence a closed loop needs only accumulates as fast as the override log grows.

Highest-leverage next move

Choose the one decision with contained consequences and high tempo — Volt/VAR setpoints are the usual answer — and start engineering its operating envelope from the override log while running shadow mode.

Cost of leaving

Effort
12–18 months, dominated by envelope engineering and OT assurance rather than modelling
Team
Integration engineer, ML engineer, protection/automation engineer, OT security assessor, named control-room owner
Risk
Medium to high — the write path into a control scheme crosses the CIP or equivalent boundary and inherits its assurance calendar
To next stage
12–18 months

If this is you, the next step is

We review the write path, the override log and storm-mode behaviour against what a closed loop will need.

Audit your advisory layer

Stage 4

Acting

11% of operators sit here

Acting means a model's output changes the state of the grid without a per-action human approval — always through an engineered control scheme, always inside a versioned, power-flow-validated operating envelope, always supervised and reversible.

Acting is narrower than the word 'autonomy' suggests, and the narrowness is the design. No serious utility lets a learned model command breakers or write to protection. What stage 4 actually looks like is a model proposing setpoints or schedules to a control scheme that was already engineered to act — a Volt/VAR optimisation adjusting regulator and capacitor settings, a DERMS issuing dispatch and curtailment within registered limits, a battery schedule optimiser. The scheme retains its own interlocks and its deterministic behaviour; the model adjusts what the scheme aims for, inside bounds a human wrote down. The grid has run closed loops for a century — protection, automatic generation control (AGC), and more recently FLISR self-healing — and it accepts them precisely because they are bounded, deterministic where it matters, and reversible. AI joins the loop on the same terms or not at all.

The central artefact is the operating envelope, and it deserves more engineering respect than it usually gets. A real envelope is not a min/max pair in a configuration screen. It is a versioned document and its enforcement code: hard bounds validated against power-flow studies of the actual network, rate limits so the loop cannot thrash equipment through excessive tap and switching operations, validity conditions that state when the model may act at all — topology as expected, telemetry fresh, no active switching in the area — and auto-revert triggers that return control to the previous scheme the moment any condition fails. Every element traces to evidence: the bounds to network studies, the validity conditions to the failure modes observed in the advisory year, the escalation thresholds to the override log.

The operational job changes accordingly. Supervising a closed loop is not watching its every action — that would just be stage 3 with worse ergonomics — it is watching its envelope statistics. The escalation rate (how often proposals fall outside bounds) is the leading indicator: rising escalations mean the network has moved away from the envelope's assumptions, through DER growth, reconfiguration or seasonal change, and the envelope needs review before an incident forces one. This is also where the security obligations bite hardest: a component that writes into a control scheme sits inside the OT boundary, and in North America the NERC CIP standards govern how it is built, patched, accessed and monitored. That is not an obstacle erected against AI — it is the same regime every control component lives under, and the roadmap's earlier stages exist partly so the assured path is built once, properly, before the loop needs it.

In practice

The envelope that caught the network changing

A distribution utility ran AI-recommended Volt/VAR schedules on a pilot feeder group through an existing VVO controller. The envelope held voltage inside statutory limits with margin, capped daily tap operations, and required as-operated topology confirmation before each schedule push. For two quarters the loop ran quietly, with escalations under two per cent. Then escalations tripled in a month. Investigation found no model fault: a wave of new rooftop solar connections had shifted the feeders' daytime voltage profile, and proposals that had been routine now pressed the upper bound. The envelope review brought the bounds and the model's training window up to date together — an afternoon of governance that, without the escalation signal, would have been an overvoltage investigation instead.

What it looks like

  • At least one closed loop is live: model setpoints or schedules flow into an engineered scheme such as VVO or DERMS dispatch
  • A written, versioned operating envelope exists, validated against power flow, with rate limits and validity conditions
  • Out-of-envelope proposals escalate to an operator; envelope breaches auto-revert and page
  • Reversion to the previous control source is drilled on a calendar, not discovered in an incident

Diagnostic signals you can check this week

  • Ask to see the envelope document and its version history. A settings screenshot is not an envelope
  • Ask when reversion to the previous control source was last drilled, and whether it was a calendar drill or an incident
  • Ask for the escalation-rate trend over six months, and what the last rise turned out to mean
  • Ask how a single automated setpoint from three months ago would be reconstructed — model version, envelope version, inputs, action

Anti-pattern · Widening the envelope because the loop has been quiet

After six clean months the pressure arrives from both directions: the model team wants headroom, operations has relaxed, and the envelope looks like conservatism tax. Widening it because nothing has gone wrong is inverted reasoning — nothing has gone wrong inside bounds that were validated; outside them is unvalidated by definition. Envelopes widen the way they were set: a power-flow study of the proposed bounds, a review of escalated proposals showing the widened region would have been safe, a new version with sign-off. Quiet operation is evidence the envelope works, not evidence it is unnecessary — the distinction funds the whole stage.

What holds you here

Each loop's envelope, assurance case and scheme integration is built as a one-off, so the second loop costs nearly what the first did and the portfolio stalls at one or two.

Highest-leverage next move

Extract the shared machinery — envelope validation service, scheme interfaces, action audit log, reversion pattern — so the next loop is an envelope and a model, not a programme.

Cost of leaving

Effort
18–30 months to a portfolio of loops; each additional loop far cheaper on the assured path
Team
Platform and integration engineers, protection/automation engineer, OT security assessor, envelope owner in operations, standing review forum
Risk
Concentrated — low frequency, high consequence; the governance and evidence obligations are the schedule, not the modelling
To next stage
18–30 months

If this is you, the next step is

We review bounds, validity conditions, reversion and audit trail against a real scenario on your network.

Stress-test an operating envelope

Stage 5

Coordinating

3% of operators sit here

Coordinating means several supervised loops and advisory services run on shared integration machinery, their envelopes are governed on a calendar keyed to network change, and AI-driven actions appear in reliability reporting like any other control action.

Coordinating is the stage vendor slides call the self-driving grid, and the honest version is smaller and more valuable than that. What actually exists at stage 5 is a portfolio: a handful of supervised closed loops in the contained, high-tempo corners of the network, a wider ring of advisory services, and — the part that earns the stage its name — shared machinery underneath them. One governed feed layer, one envelope-validation service, one action audit log, one reversion pattern, one assured OT path that every new component reuses. The marginal cost of the next loop collapses, which is the economic point, and every loop behaves identically under failure, which is the operational one.

The genuinely new engineering problem at this stage is interaction. Two loops that are each safe alone are not automatically safe together: a Volt/VAR loop and a DER dispatch loop both influence feeder voltage, and their envelopes must be validated jointly, not merely side by side — the same discipline protection engineers have always applied to interacting schemes, extended to learned controllers. This is also the honest boundary of the stage. System-level questions — can learned models one day assist constraint management across a whole transmission network? — are live research at bodies like EPRI, whose Open Power AI Consortium exists precisely because no single operator can validate such systems alone. A credible roadmap treats those as research to track, not line items to schedule.

What keeps a utility at stage 5 is governance rhythm rather than technology. Envelopes are reviewed on a calendar and re-validated when the network changes, because the network now changes constantly: every DER connection wave shifts the assumptions a bound was set under. Model risk management runs as a standing discipline with the shape that frameworks such as the NIST AI Risk Management Framework and ISO/IEC 42001 formalise — inventory, ownership, review cadence, evidence retention — applied with control-engineering seriousness. And the reporting closes the loop institutionally: AI-driven actions appear in reliability and regulatory reporting with the same reconstructability as any switching operation, which is what makes the capability durable across leadership changes, audits and control periods rather than dependent on the team that built it.

In practice

The joint envelope review that found the interaction

A utility running an AI-informed VVO loop brought a DER flexibility-dispatch loop to the same feeder areas. Each had a validated envelope; the commissioning gate required a joint study. The study found a plausible evening sequence — dispatch instructing battery export while the VVO held a high-voltage schedule set that morning — that could press upper voltage limits at two weakly monitored points. The fix was a shared constraint in the envelope-validation service: dispatch proposals in those areas are checked against the active VVO schedule before release. The sequence has since occurred twice in operation; both times the joint check reshaped the dispatch and no operator needed to intervene. The interaction would never have appeared in either loop's individual assurance case.

What it looks like

  • Multiple loops — VVO, DER dispatch, storm pre-positioning, ratings — run from shared feeds, envelope validation and audit machinery
  • Envelope and model reviews are keyed to network-change events: DER connection waves, reinforcements, reconfigurations
  • Loops are coordinated: their envelopes are checked jointly so two safe controllers cannot combine into an unsafe pair
  • AI actions are reconstructable years later and reported through the same reliability lens as any control action

Diagnostic signals you can check this week

  • Ask how many loops share one envelope-validation service and one audit log, versus carrying their own
  • Ask when envelopes were last reviewed jointly rather than one at a time
  • Pick an automated action from two years ago and time how long full reconstruction takes — model, envelope, inputs, outcome
  • Ask how the last major network change — a reinforcement, a DER wave — propagated into envelope and model reviews, and how long it took

Anti-pattern · Chasing loop count as the maturity metric

Once the machinery makes new loops cheap, the tempting scoreboard is how many are live, and the portfolio drifts toward decisions that were never good closed-loop candidates — high-consequence switching, storm restoration sequencing — because they are what remains. The matrix on this page exists for exactly this moment: some grid decisions earn a loop, some earn advisory support forever, and some belong to deterministic engineering permanently. A mature stage 5 holds the line and reports value per decision, not loops per year; the discipline of not closing a loop is the same discipline that made the closed ones safe.

What holds you here

Sustaining the stage is a standing governance cost — joint validation, review calendars, evidence retention — and it decays quietly the first year attention moves elsewhere.

Highest-leverage next move

Key every envelope and model review to network-change events, review interacting loops jointly, and report AI actions through the same reliability lens as any other control action.

Cost of leaving

Effort
Continuous — the work becomes governance rhythm, joint validation and evidence retention
Team
Platform team, envelope governance forum, protection/automation engineering, OT security, regulatory reporting partner
Risk
Concentrated and systemic — interaction effects and governance decay, not individual model failure

If this is you, the next step is

We stress-test interacting envelopes and the reconstruction trail across your live loops.

Review your loop portfolio jointly

Where utilities actually sit on the integration ladder

The distribution across the five stages, the side-screen plateau at stage 2, and the external evidence for what the far side is worth.

Most utilities sit at stage 2: models genuinely connected to live grid data, with output that never enters an operational system. The distribution below is illustrative — synthesised from the sector research linked beneath it rather than measured from a single survey — but its shape is the consistent finding of every serious study of utility AI: broad experimentation, a crowded read-only middle, and a small minority whose models can act on the grid under supervision.

Distribution of utilities across the five integration stages

Stage 2 is the mode and the plateau: the side screen needs no change approval, so programmes accumulate there. The sharpest drop is into stage 4 — the envelope, assurance and scheme-integration work that most programmes have never scoped. Distribution is illustrative, charted from the synthesis described in the source.

Share of utilities

  • 26% — 1 · Detached
  • 36% — 2 · Observing (the side-screen plateau)
  • 24% — 3 · Advising
  • 11% — 4 · Acting
  • 3% — 5 · Coordinating

Source: Illustrative distribution, synthesised from IEA Energy and AI and EPRI sector research

The context makes the plateau expensive. The grid the models must integrate with is being rebuilt around them: the IEA estimates that meeting national climate targets means adding or refurbishing 80 million km of grids by 2040 (opens in a new tab) — equal to the entire existing global grid — while Eurelectric's Grids for Speed (opens in a new tab) finds 30% of European grids already over forty years old and distribution investment needing to reach €67 billion a year to 2050. On the supply side, IRENA recorded 585 GW of renewable capacity added in 2024 alone (opens in a new tab) — 92.5% of all new capacity — most of it variable and much of it connecting at distribution voltage, and US electricity demand is growing again (opens in a new tab) after two flat decades. Every one of those trends multiplies the decisions the grid must make per hour; none of them adds operators to make them. That gap — more decisions, same people — is what the advisory and actuation stages exist to close, and research consortia such as EPRI's Open Power AI Consortium (opens in a new tab) exist because the sector knows it must close the gap collectively, with validation no single operator can produce alone.

The integration boundary map: what AI reads and writes, system by system

The grid control stack as an integration surface — what a model may consume from each system, what it may ever write back, at which stage, and the safeguard that makes the write safe.

AI integrates with the grid through a specific, enumerable set of systems, and each has its own read surface, write surface and safeguard. The map below is the page's working tool: find each system you run, read what a model may take from it and what — if anything — a model may ever write back, and note the stage at which each write becomes legitimate. Two features deserve emphasis before the table. First, the write column is short by design: most systems on the stack are read-rich and write-forbidden, and a roadmap that proposes writes everywhere has not understood the stack. Second, one row has no write stage at all — protection — and its permanence is a feature of the roadmap, not a limitation of it.

SystemAI reads (from stage)AI writes (from stage)The safeguard that makes the write safe
HistorianPoint history, event logs (stage 1)Never — the historian is the recordNot applicable; derived datasets live outside it
SCADA / EMSLive points, alarms, state estimates (stage 2)Display fields and advisories (stage 3); ratings and schedule inputs (stage 4)Assured OT path (CIP), display-only separation, envelope validation for anything a scheme consumes
ADMS / DMSAs-operated topology, loading, switching state (stage 2)Advisory fields and ranked suggestions (stage 3); VVO setpoint schedules (stage 4)Envelope with power-flow-validated bounds, rate limits, auto-revert to static schedules
OMSOutage events, restoration progress (stage 2)Predicted restoration times and damage estimates, flagged as predictions (stage 3)Uncertainty display, storm-mode scope honesty, operator override
DERMSDER registrations, availability, response history (stage 2)Dispatch and curtailment schedules within registered limits (stage 4)Registered device limits, joint envelope check against VVO, auto-revert
AMI head-endInterval loads, voltages, outage flags (stage 1–2)Never for control — read-only by designNot applicable; AMI-derived signals act only via ADMS/DERMS schemes
GIS & network modelAssets, connectivity, as-built record (stage 1)Suggested corrections routed to data stewards (stage 3)Human-reviewed correction queue — models never edit the model of record directly
Work & asset managementAsset condition, work history, inspection results (stage 1)Ranked inspection and maintenance queues with reason codes (stage 3)Planner override authority, logged re-ranking, periodic blind-spot review
Protection & UFLS relaysEvent records and disturbance data, for analysis only (stage 2)Never — settings change through protection engineeringDeterministic domain: AI may inform offline setting studies that engineers then apply through the certified process
The grid AI integration boundary map. 'Reads from stage' is when a model typically starts consuming the system; 'writes from stage' is when writing back becomes legitimate — through the named safeguard, never around it. Systems marked 'never' stay deterministic permanently.

Everything in the write column lives inside the OT security perimeter, and the perimeter has a name in each jurisdiction: in North America it is the NERC Critical Infrastructure Protection standards (opens in a new tab), and equivalent network-and-information-security regimes apply elsewhere. The practical consequence for the roadmap is that the first write is slow and every subsequent write is cheap — the zoning, patching, access and monitoring pattern is negotiated once, documented, and reused. Governance frameworks are converging on the same layered picture from the model side: the NIST AI Risk Management Framework (opens in a new tab) and the ISO/IEC 42001 management-system standard both formalise the inventory-ownership-review-evidence cycle that the envelope discipline on this page implements in control-engineering terms. Regulators are pushing the stack in the same direction from a third side: the US federal regulator FERC (opens in a new tab) now requires transmission providers to use ambient-adjusted line ratings under Order 881 — a mandated move from static assumptions to model-informed operations that lands squarely in this table's EMS row, and a preview of how model-informed values will keep arriving inside the control stack whether a utility has a roadmap or not.

Which grid decisions earn the closed loop — and which never should

Two axes — decision tempo and consequence of a wrong action — sort the grid's decisions into four quadrants, and only one of them is where the loop closes first.

The decisions that earn closed-loop AI first are the ones that are both fast and forgiving: high-tempo enough that human approval is the bottleneck, and contained enough that a wrong action is a correctable inefficiency rather than a safety event. Plotting tempo against consequence sorts the grid's decision landscape cleanly, and the sort does real work in a roadmap review — most disagreements about 'how far should AI go' dissolve once the specific decision under discussion is placed on the matrix rather than argued in the abstract.

The closed-loop candidacy matrix

Place each candidate decision by its tempo and the consequence of a wrong action. The top-left quadrant is where loops close first; the top-right stays deterministic permanently — and holding that line is what makes the rest of the roadmap trustworthy.

Close the loop here first

  • Volt/VAR setpoint schedules within statutory limits
  • DER dispatch and curtailment inside registered device limits
  • Battery charge/discharge scheduling against forecasts
  • Wrong is inefficient, visible and auto-revertible

Engineered schemes only — permanently

  • Protection operation and settings, under-frequency load shedding
  • AGC and primary frequency response
  • AI may inform offline studies; certified engineering applies them
  • Determinism here is what makes loops elsewhere acceptable

Advisory is enough

  • Maintenance and inspection prioritisation
  • Hosting-capacity screening and connection studies
  • Reinforcement and planning analysis
  • Value flows fully without any actuation path

Human decides, AI assembles

  • Switching plans and outage restoration sequencing
  • Storm crew pre-positioning and load transfers
  • AI drafts, checks and ranks; the operator commits
  • The accept/override log here is permanent, by design
Decision tempo — top: Seconds to minutes, bottom: Hours to seasons
Consequence of a wrong action — left: Contained, self-correcting, right: Cascading, safety-critical
  • Tempo decides where the human adds nothing

    A Volt/VAR schedule adjusting through the day, or DER dispatch tracking a five-minute market, outruns any approval workflow — the human in that loop is latency, not judgement. Conversely, a switching plan built over twenty minutes has room for a person, and the person carries context no model sees: the crew on the ground, the customer on life support, the substation with the known defect.

  • Consequence decides where the human is the point

    Restoration sequencing and load transfers are decisions the organisation must be able to defend afterwards, name by name. Advisory AI makes them faster and better-informed; closing the loop on them would trade accountability for a latency gain nobody asked for. The top-right quadrant goes further: protection and frequency response are certified deterministic domains, and their untouchability is a public commitment, not a technical shortfall.

  • The matrix is a sequence, not just a sorting

    Programmes that respect it build trust in a defensible order: advisory value in the bottom half funds the work, the top-left loop proves the envelope discipline, and the top-right line — never crossed — is what lets operations leadership, boards and regulators extend permission for the rest. Programmes that violate the order, usually by proposing automation in the bottom-right first because the savings look large, spend years rebuilding the trust one proposal cost.

What grid AI integration looks like in public

Two publicly reported programmes, read against the integration ladder. Neither is an Atomic Loops engagement — each links to the operator's own published material.

The public record illustrates the two halves of the roadmap unusually well: one operator shows what the grid's existing closed loops teach about actuation, and another shows what governed, versioned models look like when their output steers real operational decisions. In both cases the instructive variable is not the algorithm — it is the integration architecture around it: what acts, inside what bounds, with what published evidence.

Two programmes read against the ladder

Outcomes as reported by the operators themselves — verify figures against the linked source before reusing them; we have not independently audited them. Card images are generated industry scenes from this page's own image library, not operator photography, and imply no endorsement.

Generated scene: two engineers in hard hats and high-visibility vests reviewing a translucent grid-analytics overlay in front of transmission towers at duskDuke EnergyUS investor-owned utility · multi-state distribution networks34
Challenge
Storm-driven distribution outages at multi-state scale, where the minutes between fault and reroute decide whether thousands of customers experience a flicker or an extended outage — a decision tempo no human dispatch process can match.
Approach
A self-healing grid programme: sensors, communicating switches and reclosers, and automated logic deployed across its service territories, so the network can detect a fault, isolate the damaged section and reroute power around it without waiting for a human decision — the FLISR pattern, engineered and bounded.
Reported outcome
Duke Energy states that its self-healing technology can automatically detect outages and reroute power to restore service faster or avoid the outage altogether, can often restore power in less than a minute, and that it aims to serve most customers with some form of self-healing capability over the next few years.
What it shows about the curveThis is what stage 4 actuation looks like when the grid's own culture builds it: bounded, deterministic, reversible, and trusted precisely because of those properties. Today's self-healing logic is rules-based — which is the point. The telemetry, switching, communications and reversion machinery it required is exactly the machinery an AI loop needs, and the near-term role of learned models is tuning and coordinating such schemes inside envelopes, not replacing them.

Duke Energy — self-healing technology (opens in a new tab)

Generated scene: two analysts in a briefing room in front of a video wall showing a glowing network map of California with highlighted risk clustersPacific Gas and Electric (PG&E)Californian investor-owned utility · high-consequence wildfire operating environment23
Challenge
Wildfire ignition risk across tens of thousands of miles of distribution and transmission lines in extreme and shifting fire weather — an environment where model-informed operational decisions carry public-safety consequences and public scrutiny to match.
Approach
A layered Community Wildfire Safety Program combining a weather-station and high-definition camera network with wildfire risk models, whose documentation PG&E publishes and versions on the programme's own page — model governance practised in public — with risk outputs informing operational decisions such as where to harden the system and how to scope safety-driven shutoffs.
Reported outcome
PG&E publicly documents the programme's layers — weather stations, HD cameras and versioned risk-model documentation hosted on the programme page — and reports using risk-model output to target system hardening and inform Public Safety Power Shutoff decisions.
What it shows about the curveThis is stage 3 practised under the harshest possible spotlight, and its exportable discipline is the paper trail: versioned, published model documentation turns 'the model said so' into an auditable operational decision. Note also the actuation pattern hiding inside an advisory programme — risk models do not switch anything; they change the settings and scope of engineered safety measures that humans then operate. That is the envelope pattern, learned early.

PG&E — Community Wildfire Safety Program (opens in a new tab)

Read together, the two programmes bracket the roadmap's central claim. The grid will accept automation — at scale, in safety-critical territory — when it arrives engineered, bounded and reversible; and it will accept model-informed operations — under regulatory and public scrutiny — when the models are governed, versioned and documented where anyone can inspect them. Neither operator waited for a general-purpose grid intelligence, and neither needed one. The roadmap's stages are those two disciplines, built in order.

The integration architecture, layer by layer

The six layers a grid AI estate needs, each annotated with the stage that first requires it — and the two that are almost always missing when a loop is proposed.

A stage-4 capability requires six layers, and the two that determine the schedule are the two that look least like AI: the network model and state layer, and the actuation and safeguards layer. The architecture below is deliberately vendor-neutral — every layer is defined by what it must guarantee, not by what product provides it — and each is annotated with the stage that first requires it, so a programme can read off exactly which layers its target stage demands and which can wait.

Layers required by stage

Each layer is annotated with the integration stage that first requires it. A programme proposing a closed loop without layers 5 and 6 in place is proposing a stage-2 pilot with a write permission it should not be granted.

  1. Grid sensing & telemetry

    Stage 1+

    • SCADA / RTU points & historianThe operational record and its live edge
    • AMI interval dataLoading and voltage at the grid's periphery
    • Weather & environment feedsThe exogenous driver of most grid models
  2. Network model & state

    Stage 2+

    • As-operated topologyThe grid as switched now, not as built
    • State estimationCoherent electrical state from imperfect telemetry
    • DER registerWhat is connected, where, with what limits
  3. Model layer

    Stage 2+

    • Forecasting & condition modelsLoad, generation, thermal, asset risk
    • Serving & freshness gatesLatency matched to decision tempo; stale inputs suspend output
    • Evaluation in grid unitsExcursion minutes and avoided operations, not error metrics
  4. Decision integration

    Stage 3+

    • In-system writeEMS/ADMS/OMS fields and ranked work queues
    • Accept/override captureEvery judgement logged with context
    • Operator workflow designStorm-mode behaviour designed, not discovered
  5. Actuation & safeguards

    Stage 4+

    • Envelope validation serviceBounds, rate limits, validity conditions — checked before release
    • Engineered scheme interfacesVVO, DERMS, FLISR policy — never raw device writes
    • Auto-revert & kill switchPrevious control source one tested step away
  6. Assurance & governance overlay

    Stage 3+

    • Assured OT pathCIP-conformant zoning, access, patching — built once, reused
    • Model & envelope registryVersions, owners, review triggers keyed to network change
    • Action & decision audit logAny action reconstructable years later

Pipeline described

  1. Grid sensing & telemetry (stage 1+) — SCADA / RTU points & historian: The operational record and its live edge; AMI interval data: Loading and voltage at the grid's periphery; Weather & environment feeds: The exogenous driver of most grid models
  2. Network model & state (stage 2+) — As-operated topology: The grid as switched now, not as built; State estimation: Coherent electrical state from imperfect telemetry; DER register: What is connected, where, with what limits
  3. Model layer (stage 2+) — Forecasting & condition models: Load, generation, thermal, asset risk; Serving & freshness gates: Latency matched to decision tempo; stale inputs suspend output; Evaluation in grid units: Excursion minutes and avoided operations, not error metrics
  4. Decision integration (stage 3+) — In-system write: EMS/ADMS/OMS fields and ranked work queues; Accept/override capture: Every judgement logged with context; Operator workflow design: Storm-mode behaviour designed, not discovered
  5. Actuation & safeguards (stage 4+) — Envelope validation service: Bounds, rate limits, validity conditions — checked before release; Engineered scheme interfaces: VVO, DERMS, FLISR policy — never raw device writes; Auto-revert & kill switch: Previous control source one tested step away
  6. Assurance & governance overlay (stage 3+) — Assured OT path: CIP-conformant zoning, access, patching — built once, reused; Model & envelope registry: Versions, owners, review triggers keyed to network change; Action & decision audit log: Any action reconstructable years later
Step-by-step insights
Sensing — the cheapest layer to extend and the easiest to overspec
Stages 1–3 run almost entirely on telemetry a utility already owns: SCADA points, the historian, AMI intervals, weather. The recurring mistake is delaying the roadmap for a sensing programme — PMUs everywhere, LV monitoring everywhere — that the early stages do not need. Synchrophasors earn their place for specific applications (oscillation detection, model validation at transmission level); they are not an entry requirement. The entry requirement is governance of the sensing you have: freshness SLAs, gap monitoring, and one owner per feed who is paged when it stops.
Network model & state — the layer that decides whether recommendations are real
Every model output is computed against some picture of the network, and the roadmap's quiet dependency is that the picture is current. As-operated topology from the EMS/ADMS, state estimation that flags its own low-confidence regions, and a DER register that tracks the connection queue are what let a recommendation survive contact with the actual feeder. This layer is also where stage-4 validity conditions anchor: 'the model may act only when topology is confirmed as-operated' is enforceable only if as-operated topology is a maintained, machine-readable fact.
The model layer — freshness gates over accuracy points
The layer's differentiating discipline is not model choice but self-suspension: a served grid model must know when its inputs are stale or its validity conditions have failed, and must withdraw its output rather than guess. A forecast that silently degrades on a broken feed is worse than no forecast, because downstream consumers cannot tell the difference. Evaluation belongs in grid units for the same reason — a voltage model is not 'good at 2% error'; it is good if excursion minutes fall and tap operations do not rise, measured against held-out feeders.
Decision integration — where the roadmap's value is actually released
The write into the EMS, ADMS, OMS or work queue is small engineering with a long approval tail, and it is the single highest-leverage investment on the page: everything before it is preparation, everything after it is amplification. Design the operator workflow with operators, not for them — where the field renders, what happens to it during alarm floods, how uncertainty displays, when the output is suppressed. The accept/override capture is non-negotiable from day one, because it is simultaneously the trust metric, the blind-spot detector and the envelope calibration dataset.
Actuation & safeguards — the layer proposals always underscope
When a closed loop is proposed, the model exists and the scheme exists; what is missing is everything between them. The envelope validation service — the component that checks every proposal against bounds, rate limits and validity conditions before the scheme sees it — is a real system with its own assurance case, not an if-statement. So is the auto-revert machinery, which must return control to the previous source cleanly under every failure mode, and be drilled on a calendar. Budget these as first-class components: they are most of the stage-4 schedule and all of its trust.
The governance overlay — built from stage 3, priceless at stage 5
The overlay starts as bureaucratic-seeming hygiene — a registry of models and envelopes with versions, owners and review triggers, an audit log of decisions and actions — and becomes the estate's institutional memory. It is what lets a utility answer, years later, exactly which model version and envelope version produced an action, under what inputs, reviewed by whom. Frameworks like the NIST AI RMF and ISO/IEC 42001 describe this cycle in management-system language; the grid version simply implements it with control-engineering evidence standards, and reuses the assured OT path so each new component inherits security by construction.

The build order matters as much as the layers: sensing governance before network model, network model before serving, serving before the in-system write, the write and its logs before any envelope, and the assured OT path threaded through from the first write onward. Programmes that build in this order find each stage inheriting the one before; programmes that build a model first and integration later find themselves at stage 2 with a very good model and a two-year integration schedule they thought was already behind them.

A 90-day plan: AI Volt/VAR setpoints on one feeder group

The stage 3 → 4 transition made concrete on the grid's best first loop — model-recommended VVO schedules on a pilot feeder group, from baseline through shadow mode to supervised actuation. Contains no model development.

The first closed loop should be Volt/VAR optimisation on a bounded feeder group, and the quarter below assumes the common starting point: a utility with a working VVO or capacitor/regulator control scheme, an ADMS or equivalent telemetry, rising rooftop solar pushing daytime voltages around, and a load or voltage model already proven at the advisory stage. The problem is concrete and worth solving on its own — voltage excursions and losses on high-PV feeders — and the quarter contains no model development at all: every day is spent on baseline, shadow evidence, envelope engineering and the supervised write path.

Stage 3 → stage 4 on Volt/VAR, in one quarter

One feeder group, one envelope, one named owner. If any phase needs more than its window, narrow the scope — fewer feeders, a simpler schedule — rather than extending the plan. The OT assurance thread runs the full quarter and starts on day one, not day sixty.

  1. Days 1–15

    Baseline and ownership

    Pick a feeder group with modern VVO or regulator/capacitor controls and meaningful PV penetration. Pull six months of AMI voltage and loading history plus tap and capacitor operation counts. Compute the baseline: excursion minutes outside statutory limits, estimated losses, daily equipment operations. Name the envelope owner in distribution operations and open the OT assurance conversation for the write path now.

    A signed baseline and a named envelope owner

  2. Days 16–45

    Shadow mode

    The model recommends setpoint schedules daily; nothing is written. Each recommendation is logged beside what the incumbent scheme actually did and what the network then experienced. The comparison — would the model's schedule have reduced excursions and losses without added operations? — is reviewed weekly with the operations owner, building the evidence file the envelope will cite.

    A shadow log quantifying the gap between recommended and actual

  3. Days 46–70

    Envelope engineering and the drill

    Write the envelope: hard voltage bounds inside statutory limits with margin, validated by power-flow study; daily rate limits on tap and switching operations; validity conditions — as-operated topology confirmed, telemetry fresh, no active switching in the group; auto-revert to the incumbent schedule on any breach. Version it, sign it. Complete the assured OT path for the write, and drill the reversion on a quiet shift, twice.

    A versioned, validated envelope and a drilled reversion

  4. Days 71–90

    Supervised actuation and attribution

    The model writes schedules through the existing VVO controller inside the envelope, supervised daily. Out-of-envelope proposals escalate; escalation rate is tracked from day one. Report in grid units against the baseline and a comparable held-out feeder group: excursion minutes, losses, equipment operations, escalations. This number — not model accuracy — is what earns loop two.

    A supervised live loop and an attributed operational delta

The order is the discipline

  1. Shadow before writes, always

    Shadow mode looks like delay and is actually the fastest route: it produces the evidence that makes the envelope defensible and the write approvable in one pass. A loop proposed without a shadow log is a loop that will spend the same weeks in an approval queue — arguing from theory instead of from a comparison table.

  2. Envelope before accuracy

    The model at the start of this quarter is good enough, or it would not have survived the advisory stage. Every hour spent sharpening it is an hour not spent on the bounds, validity conditions and reversion machinery that actually gate go-live. Accuracy work resumes after the loop closes, when an accuracy point can be priced in excursion minutes.

  3. One feeder group before the network

    The pilot group's envelope will be wrong in instructive ways — bounds too tight at the solar peak, a validity condition that fires on a routine switching pattern. Finding that out on one group is an envelope revision; finding it out network-wide is an incident review. Scale by cloning the pattern, not by widening the pilot.

How grid AI integration fails — and the checklist that catches it early

Four failure modes account for most integration regressions, and none of them is a model failure. The checklist beneath them is the closed-loop gate: eight conditions, all observable in an afternoon.

Integration fails sideways, not head-on. Models rarely collapse; instead the network changes under them, trust evaporates in one bad storm, or a pilot's shortcuts quietly foreclose the production path. The four failure modes below are the ones that actually appear in post-mortems, each with the cheap preventive measure that would have caught it.

Likelihood: highImpact: high

The model is right and the network model is wrong

A recommendation valid against the recorded topology lands on a feeder that was reconfigured last week. On the advisory path this produces operator eye-rolls and eroded trust; on the actuation path it produces a schedule computed for a network that is not the one energised. The failure is invisible in every model metric, because the model did exactly what it should have with the state it was given.

PreventionValidity conditions tied to as-operated topology, and auto-suspension of output whenever the switching state is unconfirmed or stale.

Likelihood: mediumImpact: high

Trust collapses in one bad storm

A model trained on routine conditions advises confidently through a major event, is visibly wrong at the worst moment, and consultation ends — not for the model, for the programme. Operators extend trust slowly and withdraw it instantly, and the withdrawal outlives the incident by years. The root cause is scope dishonesty: nothing on the output said 'this model has never seen a storm like this'.

PreventionStorm-mode scope honesty designed in: widened uncertainty or outright suppression outside the training envelope, and storm-mode acceptance measured and reviewed after every event.

Likelihood: highImpact: medium

The pilot bypasses the OT boundary and cannot be kept

To move fast, the pilot reaches the data through a temporary exception — a jump host, a one-off firewall rule, an extract pretending to be a feed. The pilot succeeds; the exception cannot become permanent; the capability is rebuilt from scratch on the assured path, at full cost, a year later. The team reads it as bureaucracy defeating innovation, but the sequencing was the defect: the exception was always a dead end.

PreventionBuild the assured OT path for the first integration, however slow it feels — every later integration reuses it, which is where the speed actually comes from.

Likelihood: mediumImpact: high

Envelope rot

The envelope was validated at commissioning and never since. DER connections shift the voltage profile, a reinforcement changes impedances, seasonal patterns move — and the bounds silently stop matching the network. Too-tight bounds strangle the loop with escalations until someone widens them casually; too-loose bounds admit risk nobody has re-examined. Either way the envelope has become a historical document enforcing itself.

PreventionEnvelope reviews keyed to network-change events — connection waves, reinforcements, reconfigurations — with the escalation-rate trend watched as the early-warning signal between reviews.

The checklist below is the gate we apply before any loop goes live, and it doubles as a regression check on loops already running: every item is observable in an afternoon with the systems and documents you already have, and every unticked box maps to one of the failure modes above.

The closed-loop gate: eight conditions

Tick what is true of the decision you want to automate. Anything unticked is the work — and its absence is a reason the loop should not close yet. The list works without JavaScript.

0 of 8 ticked

Nothing ticked — you are earlier than you think, and that is fine

A blank list means the decision is at stage 2 or earlier, and proposing a loop for it would be the third failure mode in action. Start where the roadmap starts: the governed feed and the in-system write. The 90-day VVO plan above shows what the first properly-sequenced quarter looks like.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Operating envelope
The versioned control document — and its enforcement code — inside which a model may act on the grid: hard bounds validated by power-flow study, rate limits on equipment operations, validity conditions stating when the model may act at all, and auto-revert triggers returning control to the previous scheme on any breach.
As-operated network model
The representation of the grid as it is actually switched right now, maintained by the EMS/ADMS, as opposed to the as-built record in GIS. The currency of this model is what makes a recommendation real; validity conditions in operating envelopes anchor to it.
State estimation
The EMS computation that produces a coherent electrical state of the network from imperfect, incomplete telemetry. Its confidence regions matter to AI integration: model output should degrade or suspend where the estimated state is weak.
Shadow mode
Running a model's recommendations alongside the incumbent process without acting on them, logging both, and comparing outcomes. The evidence-generating step between advisory and actuation — a loop proposed without a shadow log is arguing from theory.
Accept/override log
The record of every model recommendation shown to an operator, whether it was accepted or overridden, and the context. Simultaneously the adoption metric, the model blind-spot detector, and the calibration dataset from which operating-envelope bounds are later derived.
Volt/VAR optimisation (VVO)
Coordinated control of voltage regulators, tap changers and capacitor banks to hold voltage within limits and reduce losses. The usual first closed loop for grid AI: high tempo, contained consequences, and an engineered controller already in place to actuate through.
FLISR
Fault location, isolation and service restoration — the self-healing distribution scheme that detects a fault, isolates the damaged section and reroutes supply automatically. The grid's proof that it accepts bounded closed loops; today rules-based, and the actuation machinery AI loops build on.
DERMS
The distributed energy resource management system that registers, monitors and dispatches DER — rooftop solar, batteries, flexible load — within each device's registered limits. A natural stage-4 write target because its dispatch engine is itself an engineered, bounded scheme.
Auto-revert
The safeguard that returns control to the previous decision source — a static schedule, the incumbent controller — the moment an envelope condition fails, without waiting for a human. Drilled on a calendar; an undrilled reversion is a hypothesis.
Escalation rate
The share of a loop's proposals that fall outside envelope bounds and route to an operator. The leading indicator of envelope health: a rising trend means the network has moved away from the envelope's assumptions and a review is due before an incident forces one.
Hosting capacity
The amount of DER a feeder can accommodate without violating voltage, thermal or protection limits. A classic advisory-stage AI product — and a perishable one, which is why detached hosting-capacity studies age into misinformation within months on fast-growing feeders.
Synchrophasor (PMU)
A sensor measuring voltage and current phasors many times per second, time-stamped to GPS. Valuable for oscillation detection and model validation at transmission level; deliberately not an entry requirement for the early integration stages, which run on SCADA and AMI.

Frequently asked questions

The questions utility engineering and operations leaders ask most often when placing their estate on the integration ladder.

What is a grid roadmap for AI integration?

It is the staged plan by which a utility deepens the connection between AI models and the grid: first a read path (live telemetry joined to an as-operated network model), then an advisory path (output landing inside the EMS, ADMS, OMS and work queues, with accepts and overrides logged), and finally an actuation path (setpoints flowing into engineered schemes like VVO or DERMS inside validated operating envelopes). Progress is measured in integration depth — what models may read, where output lands, what they may write — because that, not model accuracy, is what releases operational value.

How is this different from a general utility AI transformation roadmap?

A transformation roadmap sequences the programme — funding, phases, submission windows, organisation — while this roadmap sequences the wiring: which grid systems a model touches and what it is permitted to do there. The two are complementary and gate each other. A programme can be perfectly funded and still stuck at the side-screen stage because no write path exists; equally, integration ambitions beyond the advisory stage need the funded, multi-year horizon a transformation plan provides. This page deliberately owns the wires side and links the timeline page for the funding calendar.

Will AI ever control the grid directly?

Not in the sense the phrase suggests, and the roadmap's far stage is deliberately not that. What mature grid AI looks like is learned models adjusting the aims of engineered control schemes — VVO setpoints, DER dispatch, scheme policies — inside versioned, power-flow-validated envelopes, with supervision, escalation and auto-revert around them. Protection, under-frequency load shedding and frequency response remain deterministic, certified engineering permanently. That layering is not a temporary caution to be outgrown; it is the architecture that makes any grid automation acceptable, and AI joins it on the same terms as every previous generation of control intelligence.

What is an operating envelope for grid AI?

It is the versioned document and enforcement code defining where a model may act: hard bounds validated against power-flow studies of the actual network, rate limits so a loop cannot thrash tap changers and switches, validity conditions stating when the model may act at all (topology confirmed as-operated, telemetry fresh, no active switching), and auto-revert triggers returning control to the previous scheme on any breach. It is owned by a named person in operations, signed, versioned like code, and re-validated when the network changes. A min/max pair in a settings screen is not an envelope.

Does NERC CIP prohibit AI in the control room?

No. CIP and its international equivalents govern how components inside the OT security boundary are built, accessed, patched and monitored — they apply to an AI component exactly as they apply to any other control-adjacent system, and they date the work rather than blocking it. The practical pattern is to build the assured path once, properly, on the first in-system write: zoning, access, patching and monitoring negotiated and documented so every later integration reuses it. Programmes that bypass the boundary with pilot exceptions discover the exception cannot become permanent, and rebuild at full cost later.

What data does grid AI actually need — is AMI enough to start?

The early stages run almost entirely on data utilities already have: SCADA points and the historian, AMI interval loads and voltages, weather, and the asset register. The differentiator is not more sensors but governance of the existing ones — freshness SLAs, gap monitoring, and above all a maintained join to as-operated topology, because a model computing against the as-built record is analysing a grid that no longer exists. PMUs and dense LV monitoring earn their place later, for specific applications; treating them as an entry requirement is the most common way to delay a roadmap that could start this quarter.

Why did our grid AI pilot never reach the control room?

Almost always because its output had nowhere to land: the pilot shipped a dashboard, because a dashboard needs no change approval, and acting on it therefore depended on operators voluntarily consulting a separate screen — attention that disappears exactly during the events the model was built for. The fix is the stage-3 write: the same output rendered as a field or queue inside the EMS, ADMS, OMS or work-management system operators already use, with accepts and overrides logged. It is weeks of engineering with an approval tail, and it changes decisions the first week it lands.

What should AI never touch on the grid?

Protection operation and settings, under-frequency load shedding, and primary frequency response — the deterministic, certified core whose guaranteed behaviour is what makes everything else safe. AI contributes to these domains offline, by informing setting studies and analysing disturbance records, but the applied settings go through certified protection engineering, and no learned model sits in those loops. Holding that line is not a limitation of the roadmap; it is the commitment that earns permission for closed loops elsewhere, and a programme that proposes crossing it has misread both the engineering and the institution.

How long does it take to go from advisory to a first closed loop?

Twelve to eighteen months is the honest range for a first loop done properly, and the schedule is dominated by everything except the model: a full seasonal cycle of advisory evidence including storm behaviour, envelope engineering with power-flow validation, the assured OT path for the write, and drilled reversion. The 90-day plan on this page compresses the final transition for a utility that already has the advisory year behind it. Second and third loops are far faster, because the envelope pattern, the assurance case and the shared machinery are reused rather than rebuilt.

Which decision should close the loop first?

Volt/VAR optimisation on a bounded feeder group, in most estates. It sits in the ideal quadrant — high tempo, contained and self-correcting consequences — an engineered controller already exists to actuate through, statutory voltage limits give the envelope natural hard bounds, and rising PV penetration means the incumbent static schedules are visibly leaving value behind. DER dispatch within registered limits is the usual second. Restoration switching and load transfers, despite their apparent savings, belong in the human-decides quadrant: their consequence profile makes advisory support the right permanent depth.

How do we measure whether integration is actually working?

With paired adoption and value metrics at each stage, all readable from system telemetry. At the advisory stage: acceptance rate split by operating conditions (storm-mode acceptance especially), and the operational delta on covered versus held-out feeders — excursion minutes, losses, restoration estimates against actuals. At the actuation stage: escalation-rate trend as the envelope-health signal, reversion drills completed on calendar, and the attributed delta against baseline and holdout. A programme reporting model accuracy as its headline metric is measuring the model, not the integration — and the integration is where the value was hiding.

Where do dynamic line ratings and FERC Order 881 fit on this ladder?

They are the regulator-driven version of the same journey. Order 881 requires transmission providers to replace static assumptions with ambient-adjusted line ratings — model-informed values flowing into the EMS and market systems on a schedule — which is structurally a stage-3-to-4 integration: a model writes an operational value, inside engineered validity rules, with fallbacks. It is also a useful precedent internally: it demonstrates that model-informed values will keep arriving inside the control stack by mandate, and that utilities with a governed integration pattern absorb such requirements as configuration while others absorb them as programmes.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — forecasting, asset-condition prediction, network optimisation and decision support running against live operational data, integrated into the SCADA, historian, ADMS and EMS layer rather than delivered as dashboards.

  • · Delivery inside regulated network estates: historian, SCADA, ADMS and EMS integration
  • · Actuation-path engineering: operating envelopes, shadow modes, drilled reversion
  • · Evidence-first delivery: baselines, accept/override logs and audit-grade decision trails
  • · 13 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyEnergy and AI — executive summary (opens in a new tab)
  3. International Energy AgencyElectricity Grids and Secure Energy Transitions (opens in a new tab)
  4. International Energy AgencyGrids at risk of becoming the weak link of clean energy transitions (news release) (opens in a new tab)
  5. EurelectricGrids for Speed (opens in a new tab)
  6. IRENARecord-breaking annual growth in renewable power capacity (opens in a new tab)
  7. NERCCIP standards (opens in a new tab)
  8. NISTAI Risk Management Framework (opens in a new tab)
  9. FERCFederal Energy Regulatory Commission (opens in a new tab)
  10. EPRIOpen Power AI Consortium (opens in a new tab)
  11. US EIAElectricity data and analysis (opens in a new tab)
  12. Duke EnergySelf-healing technology (opens in a new tab)
  13. Pacific Gas and ElectricCommunity Wildfire Safety Program (opens in a new tab)

Find out how deep your integration actually goes — then take it one stage further

We run the assessment with your engineering and operations leads, mark the boundary map against your EMS, ADMS and DERMS estate, and leave you with a sequenced plan for the next stage — the write path, the envelope, or the first supervised loop. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.