Redefining Technology

Energy & UtilitiesFuture of AI & Visionary Thinking

Self-optimising utilities: the future of AI in energy & utilities, rung by rung

A self-optimising utility is one whose grid, generation and demand continuously rebalance themselves: AI adjusts dispatch, isolates faults and orchestrates distributed resources inside human-set operating envelopes. It is reached in five rungs — monitored, advisory, supervised closed-loop, domain-autonomous, self-optimising — and every rung below the last is already running somewhere in production today.

Utility control room with operators at consoles beneath a wall of live grid analytics, wind turbines and solar fields visible through the windows
Energy & Utilities · Future of AI & Visionary Thinking

Key takeaways

  1. A self-optimising utility is not a grid run by an unsupervised AI — it is an enumerated set of control loops that sense, decide and act continuously inside human-set operating envelopes, with everything outside those envelopes escalating to an operator.
  2. The end-state is reached in five rungs — monitored, advisory, supervised closed-loop, domain-autonomous, self-optimising — and the first four already run in production somewhere: National Grid ESO's ML forecasting, DeepMind's closed-loop cooling control, AEMO's virtual power plant demonstrations.
  3. The binding constraint on grid autonomy is not model quality but safety and override governance: written operating envelopes, a drilled one-switch reversion to manual, and control policies versioned and reviewed like protection settings.
  4. Sensing caps autonomy. A control loop can never be more trustworthy than the state estimate it acts on, which is why distribution-level observability — AMI, feeder sensors, a reconciled as-operated network model — is the first investment on the ladder, not the last.
  5. The honest horizon: bounded domain autonomy — battery dispatch, feeder FLISR, DER fleet orchestration — is a 3–5 year build for a utility starting from advisory today; cross-domain self-optimisation is a decade out and arrives domain by domain, not as a single switchover.

Abbreviations used on this page

SCADA
Supervisory control and data acquisition
EMS
Energy management system (transmission control)
ADMS
Advanced distribution management system
DERMS
Distributed energy resource management system
DER
Distributed energy resource (rooftop solar, batteries, EVs, flexible load)
BESS
Battery energy storage system
VPP
Virtual power plant (an aggregated DER fleet bid as one unit)
AGC
Automatic generation control (the classical frequency-regulation loop)
FLISR
Fault location, isolation and service restoration
PMU
Phasor measurement unit (high-resolution grid sensor)
AMI
Advanced metering infrastructure (smart meters)
SAIDI
System average interruption duration index (with SAIFI, the headline reliability KPI pair)

Free · 8 questions · ~3 minutes

Score your utility on the autonomy ladder

Eight questions, one at a time, about three minutes. Answer them and we build your personalised autonomy report — your rung on the ladder, your score on each of the four dimensions, and the specific blocker between you and the next rung — and send it to your inbox. Your result doubles as the baseline for your first closed-loop business case.

0 of 8 answered

Question 1 of 8Sensing & telemetry

How observable is your network below the transmission interface?

A control loop can never be more trustworthy than the telemetry it acts on. Distribution-level observability is what separates a grid that can be modelled from one that can only be studied.

How the score maps to a stage
  • 05 — Stage 1, Monitored. The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
  • 611 — Stage 2, Advisory. Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
  • 1216 — Stage 3, Supervised closed-loop. Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
  • 1721 — Stage 4, Domain-autonomous. Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
  • 2224 — Stage 5, Self-optimising. Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.

What a self-optimising utility is — and what it is not

A definition, the boundary that keeps it honest, and the ladder of five rungs that connects today's control room to the end-state.

A self-optimising utility is one whose core operating decisions — grid balancing, network reconfiguration, generation and storage dispatch, demand orchestration — run as continuous closed loops: sensors feed a live model of the system, optimisers decide, actuators act, and the results feed back, with humans setting the objectives and the bounds rather than taking each action. The concept has a research pedigree: NREL's autonomous energy systems programme describes exactly this — hierarchical, decentralised control with 'dynamic self-optimisation' — as the only tractable way to run a grid with hundreds of millions of controllable devices.

What it is not is a grid handed to an unsupervised intelligence. Utilities have run automation inside hard bounds for a century — protection relays clear faults with no human in the loop, and AGC has trimmed generator output against frequency since long before machine learning. The self-optimising utility extends that settled pattern to learned policies and far wider decision spaces. The fence moves; the principle that there is a fence does not — which is why every rung on this page is defined by what may act without a person, inside what envelope, on what evidence.

Value released against position on the autonomy ladder

The curve is not linear. Value stays close to flat through the monitored and advisory rungs — where most utilities are — and inflects at the supervised closed loop, when model output first reaches the control system that acts. This is why programmes that measure progress in models built rather than loops closed report activity without results.

Operational value released by stage

  • Stage 1 · Monitored — 34% of operators. The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
  • Stage 2 · Advisory — 38% of operators. Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
  • Stage 3 · Supervised closed-loop — 19% of operators. Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
  • Stage 4 · Domain-autonomous — 7% of operators. Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
  • Stage 5 · Self-optimising — 2% of operators. Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the IEA's Energy and AI analysis.

How a grid decision closes its loop at each rung

The control path, rung by rung. The rung is determined by where the arrow ends: rungs 1–2 terminate at a human reading a screen, rung 3 writes proposals into the EMS/ADMS for approval, and rungs 4–5 execute inside a versioned envelope with exceptions escalating. Most utilities are in the top lane.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • At rungs 1–2, SCADA and AMI data leave the operational systems as batch extracts, feed forecast models, and surface on an advisory screen beside the control systems. Whether the operator acts is optional and unlogged — and attention decays exactly when the model matters most, in storms and volatile settlement periods. This is where value leaks.
  • At rung 3, telemetry streams continuously into a reconciled state estimate; the optimiser writes its output into the EMS, ADMS or DERMS as a pre-filled proposal — a battery schedule, a switching plan — and the operator approves or overrides each one, with every decision logged. The approval log becomes the evidence base for autonomy.
  • At rungs 4–5, a versioned operating envelope owned by operations lets enumerated decision types — storage dispatch, FLISR, DER orchestration calls — execute unattended inside agreed bounds. Anything outside the envelope escalates to a human with full context, and every automated action carries a reconstructable audit trail.
Step-by-step insights
Batch extracts — the habit that caps every rung above
The day-late AMI extract and the hand-built feeder model are the artefacts that quietly cap a utility at rung 2. Everything downstream of a batch pull is stale on arrival and carries no lineage, so no control application can ever be more current than the export schedule. This is why the first investment on the ladder is almost never a model — it is streaming the telemetry the target loop needs and alarming its freshness, which is unglamorous and decisive in equal measure.
The advisory screen dead end
An advisory screen requires no change-board approval on a control system, which is exactly why pilots ship one — and exactly why they stall there. The recommendation lives beside the operator's real console, acting on it is a voluntary extra step, and voluntary steps are the first casualties of a storm shift. National Grid ESO's forecasting work shows this pattern's honest ceiling: genuine, measurable value in reserve and balancing decisions, capped by control-room bandwidth until the output moves into the control path itself.
The state estimate — the loop's real foundation
Setpoints are computed against the model, not against the field, so the as-operated state estimate is the component autonomy actually stands on. Transmission operators have trusted state estimation for decades; the frontier is extending it below the transmission interface, where DER growth has made the old fog operationally expensive. A utility that cannot state, with confidence bounds, what a named MV feeder is doing right now has found its rung — and its next investment — regardless of how good its models are.
The optimiser — commodity maths, scarce framing
The optimisation itself — unit commitment, optimal power flow, storage arbitrage, reinforcement-learning dispatch in research settings — is the most mature part of the stack. What distinguishes deployments that survive is the framing around it: objectives stated in operational units (imbalance cost, curtailed MWh, SAIDI minutes), constraints inherited from the envelope rather than hard-coded by the data team, and uncertainty carried through to the proposal so the operator sees confidence, not just a number.
The proposal and the approval log
Writing the proposal into the EMS or ADMS as a pre-filled action changes the default: accepting takes one click, declining is a logged choice with a reason. The log is not bureaucracy — it is the dataset autonomy is later built from. Six months of accept-and-override history tells you which proposals are always taken, which conditions drive overrides, and where the envelope actually binds. Skip the supervised period and you arrive at the autonomy conversation with opinions where the evidence should be.
Envelope, escalation and the audit trail
The rung 4–5 lane is a policy artefact more than a model artefact: an enumerated list of decision types that may execute unattended, the bounds per type, the escalation triggers, and the trail that reconstructs any action months later — model version, state estimate, envelope version, policy version. Utilities inherit this discipline from protection-settings governance, which is precisely why the operators who reach autonomy fastest treat the AI stack as an extension of operational governance rather than as an IT project.

The five rungs from monitored to self-optimising

For each rung: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what leaving costs.

Each rung below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a real control room, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that rung. A utility is at the rung its lowest-governed loop is at — a brilliant advisory forecast does not lift an estate whose fallbacks have never been drilled.

Select a rung

Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Monitored

34% of operators sit here

The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.

Rung 1 is not a backward place — it is where a century of engineering discipline lives. The grid at this rung is already automated in the classical sense: protection relays clear faults in milliseconds, AGC trims generator output against frequency, tap changers hold voltage. What is absent is any learned model in the loop, and — more importantly — any real-time observability below the transmission interface. The system is safe, but it is safe because humans keep it inside a wide margin.

The tell is the state estimate. At rung 1 the control room has a trustworthy picture of the transmission system and a fog below it: distribution feeders modelled from as-built drawings that have drifted from field reality, AMI interval data that lands in a billing warehouse a day later, rooftop solar that appears only as suppressed demand. Two engineers asked how much embedded generation is on a feeder will give two defensible answers, because nothing reconciles the model against the field.

This rung becomes expensive the moment the network stops being passive. DER growth, EV charging and electrified heat all move the load shape faster than manual processes can track it, and the operator's response — wider margins, more curtailment, more conservative connection limits — is paid for by every connected customer. The cost of staying at rung 1 is not visible on any dashboard; it is embedded in every margin.

In practice

The feeder nobody can see

A distribution operator receives a connection application for a 5 MW solar farm on a rural feeder. The planning engineer's hosting-capacity study takes six weeks, because the feeder model has to be rebuilt by hand from GIS records and a decade of switching logs before any power-flow study can be trusted. The answer comes back conservative — connection refused pending reinforcement — not because the network is full, but because nobody can prove it is not.

What it looks like

  • SCADA covers transmission and primary substations; below that, visibility fades
  • AMI data is collected for billing, not for operations
  • Forecasts are produced by planners in spreadsheets or vendor tools, offline
  • No model output reaches the control room in real time

Diagnostic signals you can check this week

  • Ask for the real-time state of a named MV feeder — if the answer involves a site visit or yesterday's AMI extract, you are here
  • Check whether any forecast is produced inside the dispatch cycle rather than the day before
  • Ask how much embedded solar is behind a given primary substation and compare two teams' answers
  • Count the control actions taken in a shift that a model informed: at rung 1 the honest count is zero

Anti-pattern · Buying the autonomous platform first

The instinctive move is a flagship procurement — an AI-enabled ADMS or a grid analytics platform — on the theory that autonomy is a product you install. It fails predictably: the optimisation modules sit dark because the network model beneath them is stale and the telemetry they need does not exist. Instrument one part of the network to the standard a control application needs, prove one forecast in the control room, and let the platform requirements fall out of that experience.

What holds you here

The network below the transmission interface is not observable enough to model, so no learned system can be trusted to advise on it.

Highest-leverage next move

Pick one operational forecast — regional demand, embedded solar, one constrained feeder — and deliver it into the control room on the operational cadence, with its accuracy tracked.

Cost of leaving

Effort
6–12 months
Team
One data engineer, one power systems engineer, part-time control-room sponsor
Risk
Low — the work is observational; nothing touches a control path yet
To next stage
6–12 months

If this is you, the next step is

A 2-week review: which feeders, which telemetry, which model gaps block the first advisory use case.

Scope your observability gap

Stage 2

Advisory

38% of operators sit here

Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.

Rung 2 is where most of the industry's visible AI progress lives, and it is genuinely valuable. National Grid ESO's machine-learning work with the Alan Turing Institute improved national solar forecasting accuracy by a reported 33%, and better forecasts translate directly into lower reserve holding and cheaper balancing — value without a single automated action. An operator can stay honest here: the models advise, the humans decide, and the accountability chain is exactly what it was before.

The structural weakness is that advisory value is capped by human bandwidth and decays under stress. A control room in a storm, or a trading desk in a volatile settlement period, is precisely when the recommendations are worth most and precisely when nobody has time to read them. The advisory screen is a second cockpit: everything on it is optional, and optional things are the first casualties of a busy shift.

The other quiet failure of rung 2 is that it generates no evidence for the next rung. A recommendation read but never logged against the action taken produces no acceptance data, no override reasons, and no basis for ever setting autonomy thresholds. Operators who want a closed loop in three years must start recording accept-and-override on advisory output today — the log, not the model, is the long-lead item.

In practice

The wind forecast and the redispatch that didn't happen

A system operator's ML wind forecast flags a probable 400 MW over-forecast for the evening peak, six hours out. The advisory screen shows it; the shift is mid-handover, the constraint desk is working a transformer outage, and the recommendation scrolls off. The imbalance materialises and is settled at peak prices. The model was right, the value was real, and none of it was captured — because capture depended on a person having a quiet afternoon.

What it looks like

  • ML forecasts — demand, wind, solar, prices — reach the control room in time to matter
  • Recommendations appear on advisory screens beside, not inside, the control systems
  • Forecast accuracy is measured and reported; action on it is voluntary
  • Every setpoint, switch and dispatch instruction is still human-initiated

Diagnostic signals you can check this week

  • Ask where a control-room operator sees model output: a separate screen or browser tab means rung 2
  • Check whether accepted and ignored recommendations are logged anywhere
  • Compare forecast usage on a quiet day against the last storm — the delta is the decay
  • Ask what the model's accuracy is worth in balancing cost or reserve — if nobody can say, value is still notional

Anti-pattern · Chasing accuracy instead of a control path

When advisory output is under-used, the reflex is to improve the model — more features, better ensembles, another point of forecast skill. But usage is a function of where the output lands, not of marginal accuracy. A moderately accurate setpoint proposal inside the EMS, pre-filled and one click from acceptance, changes more dispatch decisions than an excellent forecast on a side screen. Build the path into the control system first; buy accuracy when you can price it in balancing cost.

What holds you here

Model output stops at a screen: acting on it is voluntary, unlogged and first to be dropped under pressure, so value depends on the calmest hour of the shift.

Highest-leverage next move

Choose one bounded decision — battery schedule, feeder reconfiguration, reserve sizing — and write the model's proposal into the control system itself, operator-approved, with every accept and override logged.

Cost of leaving

Effort
9–18 months
Team
Integration engineer, ML engineer, a named control-room owner, protection/operations review
Risk
Medium — the first write into a control system needs an envelope, an approval step and a drilled reversion
To next stage
9–18 months

If this is you, the next step is

The advisory-to-closed-loop transition is our most common energy engagement. Typically one quarter, one loop.

Design your first closed loop

Stage 3

Supervised closed-loop

19% of operators sit here

Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.

Rung 3 is the pivotal rung, and it is defined by a change of address rather than a change of algorithm: the model's output moves from a screen beside the control system to a field inside it. A battery schedule arrives in the EMS as a proposed dispatch the operator confirms; a post-fault switching sequence arrives in the ADMS as a pre-built plan. The human still decides everything — but deciding is now one action inside their own console, and declining is a logged choice rather than a silent omission.

The disciplines that make this rung safe are borrowed from protection engineering, not data science. The operating envelope — which actions may be proposed, within what limits, under which network conditions — is written down, owned by operations, and treated with the seriousness of protection settings. The reversion to the previous scheme is a single switch, exercised deliberately on a quiet shift, because an undrilled fallback is a hypothesis. DeepMind's account of moving its data-centre cooling from recommendations to autonomous control describes exactly this frame: hard constraints, uncertainty-aware decisions, and operators who can take back control at any time.

What rung 3 buys, beyond its own operational value, is the evidence base for autonomy. Six months of accept-and-override logs on one loop tell you — with data, not judgement — which proposals are always accepted, which conditions drive overrides, and where the envelope is actually binding. That log is the raw material from which rung 4 thresholds are set, and the operators who skip it end up setting autonomy bounds by committee guesswork.

In practice

The battery that proposes its own schedule

A utility runs a 50 MW BESS against an intraday price and imbalance forecast. Each half-hour the optimiser writes a proposed charge/discharge schedule into the EMS; the control-room operator reviews and confirms it — usually in seconds, occasionally editing it around a known outage. Every confirmation and edit is logged. After two quarters the log shows 92% of proposals accepted unmodified, with overrides clustering in two named network conditions — which becomes the draft envelope for unattended operation.

What it looks like

  • Proposals arrive as pre-filled actions in the operator's own console, not on a side screen
  • A written operating envelope, owned by operations, bounds what may be proposed
  • Every accept and override is logged with context and reviewed
  • Reversion to the previous control scheme is one switch away and has been drilled

Diagnostic signals you can check this week

  • Open the operator's console: model proposals should be visible there, pre-filled, without a second login
  • Ask to see the operating envelope document and who signed it — operations, not the data team
  • Ask when the reversion to manual was last drilled; a date is a pass, a shrug is a fail
  • Pull one week of accept/override logs — if they don't exist, the loop is advisory with better plumbing

Anti-pattern · Skipping the supervised year

Vendors and pilots alike are tempted to jump from advisory straight to unattended operation — the demo works, the approvals feel like ceremony, and removing them is one config change. Resist it. The supervised period is not a trust ritual; it is the only source of the override data that makes autonomy bounds evidence-based, and the period in which operations learns the system's failure shapes cheaply. Every month skipped at rung 3 is paid back with interest after the first unattended mistake.

What holds you here

Every action still waits for a person, so loop throughput is capped by operator attention — and the case for removing the wait must be built from the approval log, which takes disciplined months to accumulate.

Highest-leverage next move

Run the loop supervised for at least two quarters, then let the override log — not opinion — define the conditions under which the loop may execute unattended.

Cost of leaving

Effort
12–24 months
Team
Platform engineer, ML engineer, control-room product owner, protection and compliance review
Risk
Medium-high — the envelope, audit and drill disciplines must hold through staff turnover
To next stage
12–24 months

If this is you, the next step is

We audit one live loop: envelope, override log, drill history, and what the data says about autonomy readiness.

Review your envelope and evidence

Stage 4

Domain-autonomous

7% of operators sit here

Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.

Rung 4 is autonomy with a fence around it. The utility has not handed the grid to a model; it has enumerated specific decision domains — intraday dispatch of a storage fleet, FLISR on instrumented feeder groups, orchestration of a contracted DER portfolio — and allowed each to run unattended strictly inside an envelope that operations wrote, signed and can revoke with one switch. The precedent is older than machine learning: AGC and protection relays have acted without per-action approval for decades. Rung 4 extends that settled pattern to learned policies and wider decision spaces — the governance frame is inherited, not invented.

The operator's job changes shape here. Instead of approving individual actions, the control room supervises populations of actions: watching override and escalation rates, reviewing the weekly excursion report, tuning envelopes as the network changes. The critical signal becomes the escalation rate — the share of decisions the loop hands back. A rise means the world has moved outside the policy's validity (a new interconnector, a cold snap, a tariff change reshaping evening load), and it should trigger an envelope review before it triggers an incident.

The hard engineering at rung 4 is mostly evidence engineering. A regulator, an auditor or a connection customer will eventually ask why the system took a specific action at 03:41 on a Tuesday, and the answer must be reconstructable: which model version, which state estimate, which envelope, which policy. Utilities already know how to do this — switching logs and protection-settings reviews are exactly this discipline — which is why the fastest climbers treat the autonomy stack as an extension of operational governance rather than as an IT system.

In practice

The feeder group that reconfigures itself

A distribution operator runs FLISR unattended across an instrumented feeder group: when a fault locks out a section, the system locates it from feeder sensors, isolates the faulted span and restores supply to healthy sections by closing ties — in under a minute, against a restoration plan a human would have taken twenty to build. Storm mode widens the envelope's escalation triggers; anything involving crew safety tags, backfeed risk or abnormal switching states goes straight to the operator with the proposed plan attached.

What it looks like

  • Enumerated decision types execute without per-action approval, inside written bounds
  • Out-of-envelope conditions escalate to an operator with full context
  • Envelopes and control policies are versioned, reviewed and drilled like protection settings
  • Every automated action carries a reconstructable audit trail

Diagnostic signals you can check this week

  • Ask for the list of decision types that execute unattended — rung 4 operators can hand you an enumerated, versioned document
  • Ask for last month's escalation rate on any autonomous loop, and whether it is trended
  • Pick one automated action from last quarter and ask for its full reconstruction — model, state, envelope, policy versions
  • Check whether the kill switch per domain has been exercised in the last six months

Anti-pattern · Inheriting thresholds across domains

The first autonomous domain works, and its envelope becomes the template: the battery loop's thresholds get copy-pasted onto the feeder loop, the summer envelope onto winter operation. But an envelope is evidence-shaped — it encodes one domain's override log under one range of conditions, and it transfers no better than a protection setting transfers to a different network. Every new domain re-earns autonomy from its own supervised period; every season and topology change gets an envelope review. The alternative is a bad automated action, and the usual response to that is a blanket switch-off — a two-rung regression from a single incident.

What holds you here

Each domain optimises alone: the battery fleet, the feeder group and the DER portfolio each run well inside their own envelope while leaving cross-domain value — and cross-domain conflicts — unmanaged.

Highest-leverage next move

Build the co-ordination layer: shared state, a hierarchy of envelopes, and an arbitration policy that decides which domain yields when their objectives collide.

Cost of leaving

Effort
24+ months
Team
Platform team, control-room product owner, standing governance forum, cyber and compliance partners
Risk
Concentrated — low-frequency, high-consequence, regulatory in nature; the audit trail is the deliverable
To next stage
24+ months

If this is you, the next step is

We run a scenario exercise against one live loop: envelope, escalation, reconstruction, kill switch.

Stress-test an autonomous domain

Stage 5

Self-optimising

2% of operators sit here

Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.

Rung 5 is the visionary end-state, and honesty about it matters: no utility operates here across its whole estate today, and none will for years. What exists are bounded previews — NREL's autonomous energy systems research explicitly targets 'dynamic self-optimisation' for networks with hundreds of millions of controllable devices and has moved into campus- and community-scale field demonstrations. The end-state is best understood as the co-ordination of many rung-4 domains, not as a new kind of intelligence.

What genuinely changes at rung 5 is who optimises the optimisers. At rung 4, humans tune each envelope; at rung 5, the system proposes its own re-tuning — retraining on drift, adjusting dispatch policy as DER uptake shifts the load shape, re-weighting objectives within bounds the governance forum sets. The human role concentrates into three functions that do not automate: setting the objective function, auditing that the system's behaviour matches it, and owning the escalation tiers when it does not.

The sceptic's question — why climb this far at all? — has a quantitative answer. The IEA estimates that applying existing AI tools to grid operations could unlock up to 175 GW of transmission capacity without building a single new line, and projects data-centre electricity demand more than doubling to around 945 TWh by 2030. A grid absorbing that load growth, plus electrified transport and heat, with a workforce that is not doubling, closes the gap with control-loop automation or does not close it at all. Self-optimisation is not a luxury end-state; it is the operating model the load curve is forcing.

In practice

The evening peak, arbitrated

On a winter evening the price signal, the constraint forecast and a cold-snap demand pickup collide. The storage fleet wants to discharge for price; the network layer wants reserve held against a feeder constraint; the DER portfolio can shift 80 MW of contracted flexibility. The arbitration policy — versioned, human-owned — resolves the conflict against this year's weighted objectives, publishes the reasoning to the control room, and escalates only the one decision that breaches the cross-domain envelope. The operator on shift reviews one decision, not three hundred.

What it looks like

  • Cross-domain optimisation: storage, network reconfiguration and DER flexibility are traded against each other continuously
  • Policies retrain and re-tune themselves inside human-set meta-envelopes
  • System-level objectives — cost, carbon, reliability — are explicit, weighted and versioned
  • Humans govern: setting objectives, auditing decisions, owning the escalation tiers

Diagnostic signals you can check this week

  • Ask whether any two autonomous domains share state and an arbitration policy, or merely coexist
  • Ask who owns the system-level objective weights and when they were last reviewed
  • Check whether policy re-tuning is itself logged, bounded and reversible
  • Ask what the operator on shift actually reviews in an evening peak — populations and exceptions, or individual actions

Anti-pattern · Declaring the end-state by press release

The gravitational pull at this altitude is narrative: 'the self-optimising grid' announced while three-quarters of the estate is still at rung 2. The damage is not embarrassment — it is that the claim redirects governance attention from the unglamorous work (envelope reviews, escalation tuning, evidence trails) to the story, and the first public incident lands on a system whose stated capability exceeds its audited one. Let the rung ladder describe reality; the estate is at the rung its lowest-governed autonomous loop is at.

What holds you here

Sustaining self-optimisation is a standing governance discipline — objectives drift, envelopes rot and audit expectations rise, so the rung is held by review cadence, not by any model.

Highest-leverage next move

Treat the objective function, the arbitration policy and every envelope as versioned, reviewable artefacts with the same rigour as protection settings — and staff the forum that reviews them.

Cost of leaving

Effort
Continuous
Team
Platform and control teams plus a standing socio-technical governance forum with regulatory engagement
Risk
Systemic — cross-domain coupling means failures propagate; governance and cyber-security are the binding disciplines

If this is you, the next step is

A working session on the honest gap between your architecture today and a governable rung-5 target.

Pressure-test your end-state roadmap

Where utilities sit on the ladder today

Illustrative distribution, synthesised from the adoption patterns in IEA and McKinsey electric-power research — to be replaced with measured values as assessments accumulate. The advisory rung is the mode and the plateau: forecasting is widespread, closed loops are rare, and the drop from rung 2 to rung 3 is the largest single transition loss on the ladder.

Share of utilities (illustrative)

  • 34% — 1 · Monitored
  • 38% — 2 · Advisory (the plateau)
  • 19% — 3 · Supervised closed-loop
  • 7% — 4 · Domain-autonomous
  • 2% — 5 · Self-optimising

Source: Illustrative distribution, anchored to IEA Energy and AI and McKinsey electric-power research

What already runs in production today

The end-state is speculative; its components are not. Each rung of the ladder is demonstrated by a named, publicly documented programme.

Every rung below self-optimisation already runs in production somewhere, publicly documented by the operator itself. That is what separates this end-state from most visionary AI claims: the argument is not that the technology will arrive, but that demonstrated components — ML forecasting in a national control room, closed-loop optimisation of critical infrastructure, market frameworks for orchestrated DER fleets — have simply not yet been assembled by one utility at estate scale. The table below reads the public record against the ladder.

ProgrammeOrganisationWhat it demonstratesRung
ML solar forecasting with the Alan Turing InstituteNational Grid ESOA reported 33% improvement in national solar forecast accuracy — advisory value in a live control room2
Solar nowcasting with Open Climate FixNational Grid ESOSatellite-driven minutes-ahead forecasts built for control-room use2
Wind output forecasting for market commitmentsGoogle DeepMind36-hour-ahead ML forecasts that raised the value of wind energy by roughly 20%2–3
Data-centre cooling: recommendations, then autonomous controlGoogle DeepMind40% cooling-energy reduction, later run closed-loop under safety constraints with operator override3–4
Virtual power plant demonstrationsAEMOAggregated residential batteries dispatched as one orchestrated unit in the national market3
Order No. 2222 — DER aggregations in wholesale marketsFERCThe market architecture that makes orchestrated DER fleets economic at scaleenabler
Autonomous energy systems research and field demonstrationsNRELHierarchical self-optimising control validated in campus- and community-scale deployments4–5
Publicly documented programmes mapped to the autonomy ladder. Each links from the sources list; outcomes are as reported by the named organisation.

Improved solar forecasts will help us run the system more efficiently, ultimately meaning lower bills for consumers.

The regulatory rail matters as much as the technology. In the United States, FERC (opens in a new tab) Order No. 2222 requires wholesale markets to admit aggregated distributed energy resources, which is what makes an orchestrated fleet of household batteries a market participant rather than a demonstration. In Australia, AEMO (opens in a new tab) has run virtual power plant demonstrations to establish how orchestrated DER fleets behave inside the national dispatch process. Neither framework mentions machine learning — but both define the arena in which predictive DER orchestration and autonomous demand response become businesses, and both assume exactly the envelope-and-escalation governance this page's ladder is built on.

Three operators, read against the ladder

Publicly reported programmes only, cited to the operator's own material. None is an Atomic Loops engagement — the value is in the shape of each climb.

The clearest evidence for the ladder is in what serious operators chose to build first. In each case below, the differentiator was not model sophistication — it was sequencing: observability before advice, advice before proposals, a supervised period before any autonomy, and a safety frame carried the whole way up. Read each one for the rung transition it demonstrates rather than the headline number.

Three public climbs

Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them.

Engineers reviewing a wall-sized national grid map with solar, demand and network-flow overlaysNational Grid ESOGB electricity system operator · national control room12
Challenge
Balancing a grid with fast-growing embedded solar the control room cannot meter directly: rooftop generation appears only as suppressed demand, so forecast error feeds straight into reserve holding and balancing costs.
Approach
ML forecasting built for the control room rather than the lab: a random-forest and ensemble approach developed with the Alan Turing Institute for national demand and solar, followed by satellite-driven solar nowcasting with Open Climate Fix for the minutes-ahead horizon.
Reported outcome
A reported 33% improvement in solar forecasting accuracy, with the operator publicly linking improved forecasts to running the system more efficiently and lowering consumer bills.
What it shows about the curveThis is the advisory rung done properly: value measured in operational units, delivered on the operational cadence, into the room where the decision happens — and trust built with the control room before any closed loop is proposed.

National Grid ESO — news (opens in a new tab)

Two engineers studying energy analytics displays in a control room overlooking wind turbines and solar panelsGoogle DeepMindHyperscale infrastructure operator · grid-scale load and renewables offtaker24
Challenge
Data-centre cooling plants with dozens of interacting setpoints and non-linear dynamics — energy-intensive, safety-critical, and beyond what static heuristics could optimise.
Approach
The canonical recommend-supervise-bound climb: 2016 recommendations implemented by human operators cut cooling energy; by 2018 the system operated the plant directly under a safety-first frame — hard-coded constraints, uncertainty-aware decisions, and operators able to take back control at any time. In parallel, 36-hour ML wind forecasts supported day-ahead market commitments.
Reported outcome
A reported 40% reduction in cooling energy, autonomous closed-loop operation under safety constraints, and a roughly 20% increase in the value of its wind energy — all published by the operator.
What it shows about the curveThe middle rungs of the ladder, demonstrated end to end on live critical infrastructure: autonomy was earned through a supervised period and bounded by an explicit safety frame — precisely the governance a grid control loop needs.

Google DeepMind — blog (opens in a new tab)

Sensor unit mounted on a high-voltage line with a power plant and transmission towers in the distanceEnelGlobal distribution operator · 69M+ end users across six countries13
Challenge
Making 1.9 million kilometres of distribution network observable and remotely operable — the precondition for any automated fault management or DER hosting at scale.
Approach
A decade-scale digitalisation of the network itself: smart metering at fleet scale, remote monitoring and control pushed into the grid, and a stated investment strategy of making networks increasingly digital, flexible and resilient — industrialised through its grid-technology arm.
Reported outcome
Enel reports 1.9 million km of power lines and 69.4 million active end users on increasingly digitalised networks, with 239.4 TWh distributed in the first half of 2026 — an observability-first platform on which automated reconfiguration becomes an increment rather than a leap.
What it shows about the curveThe unglamorous truth of the ladder: for a distribution operator, the longest lead item in self-healing networks is not the algorithm but the sensing and remote-control fabric — and Enel bought that fabric first, at full network scale.

Enel — Enel Grids (opens in a new tab)

Where autonomy lands first on a utility

Six operating domains, the decisions in each that can close their loop, the system of record each loop must write into, and the KPI it moves.

Autonomy lands domain by domain, and the domains are not equally ready. A decision is a good early candidate when three things are true: its consequences reverse on the loop's own timescale, its domain is observable enough to model honestly, and the system of record it writes into is one the utility controls. The map below is how we scope first and second loops with operators — read the sweet-spot column as where each domain's loop typically earns its keep, not as a promise.

DomainDecisions that can close the loopSystem of recordKPI it movesSweet spot
Balancing & system operationsReserve sizing, redispatch proposals, constraint managementEMS / market systemsBalancing cost, frequency deviationRung 3
Network operationsFLISR, feeder reconfiguration, volt/VAR optimisationADMS / OMSSAIDI & SAIFI, lossesRung 3–4
Generation & storage dispatchBESS charge/discharge, hybrid plant co-optimisation, AGC participationEMS / plant controllersImbalance cost, curtailed MWhRung 3–4
DER & demand orchestrationVPP dispatch, autonomous demand response, EV-charging shiftDERMS / aggregator platformPeak shaved, flexibility deliveredRung 3–4
Asset & maintenanceInspection triage, predictive maintenance scheduling, dynamic ratingsAsset management systemUnplanned outage rate, ratings headroomRung 2–3
Trading & market biddingRenewables commitment, storage arbitrage bids, imbalance positioningTrading / ETRM systemsCaptured spread, imbalance exposureRung 3
The utility decision landscape. 'Sweet spot' is the ladder rung at which the domain's closed loop typically becomes defensible — automating a rung-4 candidate from rung-1 telemetry is the fragile-autonomy quadrant described below.

What to automate first

Plot each candidate loop's observability against the reversibility of its actions. The quadrant tells you the honest next step — and three of the four answers are not 'automate it'.

Propose, human executes

  • Switching plans, redispatch, constraint actions
  • Well observed but consequence-heavy
  • Fix: supervised proposals with a drilled reversion

Close the loop

  • BESS dispatch, VPP calls, volt/VAR on instrumented feeders
  • Reversible next interval, fully observed
  • This is where rung 3 → 4 happens first

Keep humans deciding

  • Storm restoration priorities, safety-tagged switching
  • Low observability, irreversible consequences
  • Correctly held at advisory — possibly forever

Instrument first

  • DER-heavy feeders with billing-grade data only
  • Reversible actions, but the model is guessing
  • Fix: telemetry and state estimation before any loop
Observability of domain — top: Instrumented — trusted state, bottom: Fog — model drifts from field
Reversibility of action — left: Hard to undo (network-shaping), right: Easily reversed (next interval)

Storage is the natural first domain: a battery's actions reverse within the interval, its state of charge is perfectly observed, the system of record is the utility's own, and imbalance cost gives the loop a P&L nobody disputes. Feeder automation follows where the sensing fabric exists. Customer-facing orchestration — autonomous demand response, EV-charging shift — belongs later, not because the optimisation is harder but because wrong actions land on customer promises and the envelope negotiation crosses regulatory lines. Sequencing by approval friction, not model difficulty, is what separates a three-year ladder from a decade.

The horizon map: now, five years, ten years

The honest calendar for the self-optimising utility — what is deployable today, what is a 3–5 year build, and what genuinely needs a decade.

The self-optimising utility arrives in three horizons, not one leap — and the calendar is set by governance and load growth, not by algorithms. The pressure is quantified: the IEA projects data-centre electricity demand alone more than doubling to around 945 TWh by 2030, while the same analysis estimates AI-based tools could unlock up to 175 GW of transmission capacity from existing lines. The gap between those two numbers is the business case for every rung on this page.

Three horizons to a self-optimising estate

Horizon boundaries are judgements, not forecasts — each horizon's claim is anchored to a programme that already exists. What moves a utility between them is governance throughput: how fast envelopes, evidence and regulatory confidence accumulate.

  1. Today

    Now — proven and deployable

    ML forecasting in the control room (National Grid ESO's reported 33% solar-accuracy gain), closed-loop optimisation of bounded infrastructure under safety constraints (DeepMind's cooling control), orchestrated DER fleets in market frameworks (AEMO's VPP demonstrations, FERC Order 2222). Advisory value and first supervised loops are procurement decisions, not research bets.

    Rungs 2–3 available to any utility that sequences them

  2. 3–5 years

    Near horizon — bounded domain autonomy

    Storage fleets dispatching unattended inside envelopes; FLISR standard on instrumented feeder groups; predictive DER orchestration and autonomous demand response operating inside the market rails now being laid; dynamic ratings feeding constraint management continuously. The limiting reagents are distribution-level observability and the regulatory evidence base — both compounding now at the operators that started.

    Rung 4 in two or three domains at leading utilities

  3. ~10 years

    Far horizon — cross-domain self-optimisation

    Domains co-ordinating against explicit system objectives: storage, reconfiguration and flexibility traded off continuously, policies re-tuning inside human-set meta-envelopes — the pattern NREL's autonomous energy systems research is validating at community scale today, extended to utility estates. Arrives domain-pair by domain-pair, not as a switchover; the governance forum, not the platform, is the long-lead item.

    Rung 5 behaviour on enumerated domain clusters

  • Watch the escalation rates, not the demos

    The leading indicator of real autonomy is boring: published or auditable override and escalation rates on live loops. A utility that can show a falling override rate on a supervised battery loop is closer to the end-state than one announcing an AI platform.

  • Watch distribution observability spend

    Feeder sensors, AMI-to-operations pipelines and state-estimation programmes are the tell that an operator is building the floor autonomy stands on. Enel's network-scale digitalisation is the reference pattern.

  • Watch the market rails

    Each jurisdiction that operationalises DER aggregation — the FERC Order 2222 implementations, AEMO's evolving dispatch arrangements — converts predictive orchestration from a pilot into a revenue line, and pulls the DERMS layer up the ladder with it.

  • Watch who staffs the governance forum

    The utilities that reach rung 4 first will be the ones whose autonomy governance is chaired by operations with protection-settings discipline — not delegated to a data team. The org chart is a better predictor than the tech stack.

For the engineering culture this demands, the closest public reference is the systems literature rather than the AI literature — the coverage of grid autonomy in IEEE Spectrum's energy reporting (opens in a new tab) and the sector analyses from McKinsey's electric power practice (opens in a new tab) both converge on the same conclusion this page's ladder encodes: the constraint on the self-optimising utility is organisational metabolism — how fast envelopes, evidence and trust accumulate — not model capability.

The reference architecture, layer by layer

What actually has to exist at each rung — and why the safety layer is designed first, not bolted on.

A self-optimising capability is five layers, and the build order is the opposite of the demo order: sensing and safety are the long-lead items, while the optimiser — the part every pilot shows first — is the most replaceable component in the stack. The architecture below is deliberately vendor-neutral: every layer is defined by what it must guarantee, and each is annotated with the rung that first requires it.

Layers required by rung

Each layer is annotated with the ladder rung that first requires it. A programme reaching for rung 3 without the actuation and governance layers is building a rung-2 advisory tool with extra steps.

  1. Sensing & telemetry

    Stage 1+

    • SCADA / PMU / feeder sensorsGrid state at operational resolution
    • AMI-to-operations pipelineMeter data on the control cadence, not the billing cadence
    • DER & market telemetryBehind-the-meter fleets and price/constraint signals
  2. Network model & state estimation

    Stage 2+

    • As-operated connectivity modelReconciled against the field on a schedule
    • State estimation to MVTrusted enough to compute switching on
    • Freshness & quality alarmsStale state pages a named human
  3. Forecasting & optimisation

    Stage 2+

    • Demand / renewables / price forecastsOn the dispatch cadence, accuracy tracked
    • Dispatch & reconfiguration optimisersObjectives in operational units
    • Uncertainty carried to the outputProposals ship with confidence, not just numbers
  4. Control & actuation

    Stage 3+

    • Write-back into EMS/ADMS/DERMSProposals in the operator's own console
    • Approval workflow & override logEvery accept and decline recorded with context
    • One-switch reversionThe previous scheme, drilled, always available
  5. Safety & governance

    Stage 3+

    • Operating envelopesVersioned bounds owned by operations
    • Decision audit trailAny action reconstructable months later
    • Escalation & kill switchOut-of-bounds routes to a human; stop works and is exercised

Pipeline described

  1. Sensing & telemetry (stage 1+) — SCADA / PMU / feeder sensors: Grid state at operational resolution; AMI-to-operations pipeline: Meter data on the control cadence, not the billing cadence; DER & market telemetry: Behind-the-meter fleets and price/constraint signals
  2. Network model & state estimation (stage 2+) — As-operated connectivity model: Reconciled against the field on a schedule; State estimation to MV: Trusted enough to compute switching on; Freshness & quality alarms: Stale state pages a named human
  3. Forecasting & optimisation (stage 2+) — Demand / renewables / price forecasts: On the dispatch cadence, accuracy tracked; Dispatch & reconfiguration optimisers: Objectives in operational units; Uncertainty carried to the output: Proposals ship with confidence, not just numbers
  4. Control & actuation (stage 3+) — Write-back into EMS/ADMS/DERMS: Proposals in the operator's own console; Approval workflow & override log: Every accept and decline recorded with context; One-switch reversion: The previous scheme, drilled, always available
  5. Safety & governance (stage 3+) — Operating envelopes: Versioned bounds owned by operations; Decision audit trail: Any action reconstructable months later; Escalation & kill switch: Out-of-bounds routes to a human; stop works and is exercised
Step-by-step insights
Sensing — the layer that sets the ceiling
Every rung above monitored is capped by this layer, because a control loop inherits the blind spots of its telemetry. The practical priority for most utilities is not more transmission sensing — that is largely solved — but getting AMI out of the billing warehouse and onto the operational cadence, and instrumenting the specific feeders the first loops will run on. Instrument for the loop you intend to close, not for coverage statistics: one fully observed feeder group is worth more than 5% sensor coverage everywhere.
State estimation — autonomy's real foundation
The optimiser acts on the state estimate, not on the grid, so the estimate's trustworthiness is the loop's trustworthiness. Transmission operators have run trusted state estimation for decades; the frontier is below the interface, where DER growth breaks the old assumption that load is passive and predictable. The governance test is simple: would the control room compute a switching decision from this estimate? Until the answer is yes for the target feeders, the honest work is model reconciliation, not machine learning.
Forecasting & optimisation — strong maths, disciplined framing
This is the most commoditised layer, and the one where research is moving fastest — reinforcement-learning dispatch performs impressively in simulation, and NREL's autonomous energy systems work shows hierarchical optimisation scaling to community deployments. The discipline that matters in production is framing: objectives stated in the units the business already tracks (imbalance cost, curtailed MWh, SAIDI minutes), constraints inherited from the envelope layer rather than hard-coded, and every forecast shipping with its uncertainty so downstream decisions can be risk-aware.
Control & actuation — where rung 3 is won or lost
The write-back into the EMS, ADMS or DERMS is where most programmes stall, because it crosses from the data estate into the operational estate and inherits a control system's change governance. The component that unlocks the approval is invariably the reversion: operations leaders accept a new decision source they can instantly switch off. A write-back without a drilled fallback sits in a change queue indefinitely; one with it ships. Budget as much engineering for the fallback and the override log as for the optimiser.
Safety & governance — designed first, audited forever
At rung 3 the envelope and audit trail feel like process overhead; by rung 4 they are the product. What a regulator or connection customer will examine are the envelope versions, the escalation records and the reconstruction of specific actions — not the model card. Utilities hold an unfair advantage here: protection-settings governance and operational drills are exactly this discipline, already staffed. Extend those institutions to learned control rather than building a parallel 'AI governance' track that operations does not own.

Two build rules follow from the layer order. First, never let the optimiser outrun the envelope: a capability that can act must never exist, even in staging, before the bounds and kill switch that govern it. Second, make each layer pay its way at the rung that first needs it — streaming telemetry improves manual operations before any model exists, and a reconciled network model shortens every connection study. An architecture whose layers only pay at the end-state is a bet; one whose layers pay rent every rung is a programme.

A 90-day plan: closing the first loop on battery dispatch

The advisory-to-supervised transition made concrete on one utility decision — an intraday BESS schedule written into the EMS for operator approval. Contains no model development.

Closing the first loop takes about 90 days when scoped to a single decision, and multiple years when scoped to a transformation. The plan below makes that concrete on the most tractable decision in the utility estate: a grid-scale battery whose charge/discharge schedule is already recommended by a price and imbalance forecast but executed by hand from a side screen. The optimiser exists; the quarter contains no model development at all — only the control path, the envelope and the evidence.

Advisory to supervised closed-loop on one BESS, in one quarter

One battery, one control room, one named owner. If any phase needs more than its window, narrow the scope — fewer market services, tighter hours — rather than extending the plan.

  1. Days 1–15

    Baseline the battery and name the owner

    Pick one BESS and pull six months of dispatch history, imbalance settlement and market outcomes. Compute the baseline: revenue captured versus the perfect-foresight bound, imbalance cost incurred, curtailment absorbed. Name the control-room operations owner — the schedule is their number now. Agree the counterfactual for attribution: the same optimiser advisory-only on comparable days, or unoptimised intervals held out each week.

    A baseline and a counterfactual design the CFO accepts

  2. Days 16–45

    Write the proposed schedule into the EMS

    Deliver the optimiser's half-hourly schedule as a pre-filled proposal in the console the operator already uses — not a separate screen. The operator confirms or edits each schedule; static scheduling remains one switch away as the drilled fallback. Log every accept, edit and decline with a reason code. Draft the operating envelope with operations: charge/discharge limits, state-of-charge floors, market services in scope, network conditions that suspend proposals.

    Proposals live in the EMS; envelope v1 signed by operations

  3. Days 46–70

    Instrument the loop and drill the reversion

    Freshness alarms on every feed the schedule depends on — price, imbalance forecast, state of charge, network status — paging a named human. Drift monitoring keyed to the events that actually change the problem: tariff changes, new market services, seasonal load shape. Exercise the reversion to manual scheduling once, deliberately, on a quiet day, and record the drill. Review the first month's override log with the operators who produced it.

    Alarms live, reversion drilled, override log reviewed

  4. Days 71–90

    Attribute the value against the counterfactual

    Report the delta in captured revenue and imbalance cost against the agreed counterfactual — not forecast accuracy. Present the override analysis alongside it: acceptance rate, override clusters, envelope conditions that actually bound. Close the quarter with two artefacts: the attributed number that funds loop two, and the evidence pack — envelope, logs, drill record — that starts the autonomy conversation from data.

    An attributed delta and an evidence pack, both reusable

The order matters

  1. Control path before accuracy

    A moderately accurate schedule proposed inside the EMS changes more dispatch decisions than an excellent one on a side screen. Improve the model after the path exists, when an accuracy point can be priced in imbalance cost.

  2. Approval before autonomy

    Keep the operator's confirmation for at least two quarters even where unattended execution is technically trivial. The override log — which proposals were accepted, edited, declined, and why — is the dataset that sets defensible envelope bounds later. Skipping it means setting autonomy thresholds by committee.

  3. One battery before one fleet

    The shared platform — common telemetry, serving, envelope tooling — is worth building when the second and third loops are already asking for the same feeds. Building it before the first loop has an attributed number encodes guesses as architecture.

Safety and override governance: the discipline that decides the timeline

Four artefacts make grid autonomy defensible, and four failure modes send utilities back down the ladder. None of the eight is about model quality.

Safety governance is the pacing discipline of the whole ladder: a utility climbs exactly as fast as its evidence accumulates and falls exactly as far as its governance is shallow. Four artefacts carry that evidence — the same four an auditor, a regulator or an incident review will ask for — and none needs inventing: each extends a discipline the industry already runs for protection settings and switching authorisation.

  • The operating envelope

    A written, versioned statement of what the loop may do: decision types, bounds per type, network conditions that suspend it, escalation triggers. Owned and signed by operations, changed through the same board that changes protection settings. If the envelope lives in vendor configuration, it does not exist.

  • The tested reversion

    One switch back to the previous control scheme, exercised on a schedule and recorded. DeepMind's published account of autonomous cooling control makes the point at the source: operators accepted autonomy because taking back control was always available and always worked. An undrilled fallback is a hypothesis, and control rooms do not run on hypotheses.

  • The versioned control policy

    The model, its objectives and its thresholds treated as one reviewable artefact with a version history — who changed what, when, on what evidence. The question that will eventually arrive is 'why did the system do that at 03:41 on the 14th?', and the answer must name the policy version that produced the action.

  • The decision audit trail

    Every automated action and every escalation logged with the state estimate, envelope version and policy version that produced it, reconstructable months later without archaeology. Built as a by-product of the actuation layer, it turns regulatory questions into exports; assembled after the fact, it turns them into projects.

Likelihood: mediumImpact: high

One bad action switches everything off

An automated loop takes a visibly wrong action — a battery discharges into a negative-price interval, a reconfiguration islands a customer — and the organisational response is a blanket suspension of all automation. The programme regresses two rungs in an afternoon, and the memory outlasts the incident by years.

PreventionGraduated autonomy per decision type, envelopes that bound worst-case cost, and a pre-agreed incident playbook that suspends one loop, not the programme.

Likelihood: highImpact: medium

The network changes and the policy does not

A new interconnector, a step-change in DER uptake, an EV tariff that reshapes evening load — the input distribution moves, and a policy tuned on last year's grid quietly degrades. In utilities the drift is seasonal and structural, not random, so it evades generic statistical monitors until the envelope is breached.

PreventionDrift alarms keyed to named network events — commissioning dates, tariff changes, season starts — plus a scheduled envelope review that assumes the world has moved.

Likelihood: highImpact: medium

Envelope rot

The envelope is written once, at go-live, by the people who built the loop — and never revisited. Two years later it bounds a network that no longer exists: limits too loose where DER grew, too tight where reinforcement landed, escalation triggers tuned to a retired market service. The loop is now governed by a document about a different grid.

PreventionEnvelopes carry expiry dates like protection settings: review on calendar, on escalation-rate excursion, and on every material network change — whichever comes first.

Likelihood: lowImpact: high

The loop widens the attack surface

Closing a loop connects the data estate to the control estate, and every connection is a path an attacker can walk. A compromised forecast feed or poisoned market signal becomes, at rung 4, a hand on the switchgear. The cyber posture that protected an advisory stack is categorically insufficient for an actuating one.

PreventionSegregate the control path, authenticate and bound every setpoint at the receiving system, and keep the kill switch on infrastructure independent of the optimisation stack.

Proving the loop: the metrics that survive scrutiny

The KPI set for an autonomy programme, the telemetry each number comes from — and the readiness checklist for closing your first loop.

An autonomy claim you cannot attach telemetry to is a press release. Every KPI an autonomy programme should report reduces to timestamps, counts and settlement lines that the EMS, ADMS, DERMS or the serving layer already records — the instrumentation work is joining them, not creating them. The table below is the measurement sheet: what each metric reads, where it comes from, and the rung at which it first measures something real.

MetricWhat it readsSourceHonest from
Forecast skill vs operational baselineModel error against the method the control room used beforeServing log vs EMS historyRung 2
Decision latencySource event to proposal available at the point of controlServing log + control-system eventsRung 2
Proposal acceptance rateAccepted proposals ÷ proposals shown, with override reason codesApproval logRung 3
Attributed operational deltaImbalance cost, curtailed MWh or SAIDI minutes vs the agreed counterfactualSettlement + OMS vs holdoutRung 3
Loop availabilityShare of intervals the loop was in service and in envelopeControl-system + envelope logsRung 3
Escalation rateOut-of-envelope decisions handed to a human ÷ automated decisionsDecision logRung 4
Envelope excursion rateActions that approached or breached bounds, trended by causeEnvelope monitorRung 4
Reconstruction timeElapsed time to fully reconstruct one named automated actionAudit trail, tested by drillRung 4
The autonomy programme's KPI sheet. 'Honest from' is the ladder rung at which the metric first measures something real — quoting an automation rate before rung 4 is theatre.

Two disciplines keep the sheet honest. First, adoption and value metrics are always reported as a pair — an acceptance rate without an attributed delta is theatre, and a delta without an acceptance rate is unexplainable. Second, every value number is attributed against a counterfactual agreed before go-live. On a live grid, weather, tariffs and topology all move — without the counterfactual, they will claim the credit or take the blame.

First-loop readiness checklist

Seven conditions for closing your first supervised loop. If you cannot tick all seven, the gap is your next quarter's work — regardless of how good the optimiser is. Tick as you go; this list works without JavaScript.

0 of 7 ticked

Tick honestly — the blank list is data too

Most advisory-rung utilities can genuinely tick one or two of these, not zero. If none apply yet, don't start with tooling: pick one battery or one feeder group and run the 90-day plan above. Everything on this list falls out of doing that once.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Self-optimising utility
A utility whose core operating decisions run as continuous closed loops — sense, decide, act, learn — inside human-set envelopes, with operators governing objectives and exceptions rather than taking each action.
Operating envelope
The versioned, operations-owned statement of what an automated loop may do: decision types, bounds per type, conditions that suspend it, and the triggers that escalate a decision to a human.
Supervised closed loop
A control loop whose output is written into the EMS, ADMS or DERMS as a pre-filled proposal that an operator approves per action, with every accept and override logged. The pivotal rung between advisory AI and autonomy.
FLISR
Fault location, isolation and service restoration — the self-healing-network capability: detecting a feeder fault, isolating the faulted section and restoring healthy sections by closing ties, in seconds rather than crew-hours.
State estimation
The computation of the grid's most likely electrical state from telemetry and the network model. The trustworthiness of a control loop is capped by the trustworthiness of the state estimate it acts on.
DER orchestration
Co-ordinated dispatch of distributed energy resources — rooftop solar, batteries, EVs, flexible load — as a portfolio, typically through a DERMS or aggregator platform, against market and network objectives.
Virtual power plant
An aggregated fleet of distributed resources bid and dispatched as a single unit in the market. AEMO's demonstrations and FERC Order 2222 define the operating and market frameworks respectively.
Autonomous demand response
Demand-side flexibility that responds to price or network signals without per-event human action on either side — the customer's assets act within pre-agreed bounds, the utility's systems call them within an envelope.
Automatic generation control
The decades-old closed loop that continuously adjusts generator output to hold system frequency — the industry's own precedent that bounded, unattended control of the grid is a governance problem long since solved for narrow domains.
Nowcasting
Very-short-horizon forecasting — minutes to a few hours — typically from satellite or sensor data. National Grid ESO's solar nowcasting work with Open Climate Fix targets exactly this control-room horizon.
Escalation rate
The share of an autonomous loop's decisions that fall outside its envelope and route to a human. Monitored as a leading indicator: a rising rate means the world has moved outside the policy's validity.
Counterfactual attribution
Measuring a loop's value against what would have happened without it — holdout feeders on the old scheme, or unoptimised intervals rotated through the week — so the delta survives weather, tariff and topology changes.

Frequently asked questions

The questions utility leaders ask most often when testing the self-optimising vision against their own estate.

What is a self-optimising utility?

A self-optimising utility is one whose core operating decisions — balancing, network reconfiguration, dispatch, demand orchestration — run as continuous closed loops that sense the system, decide and act inside human-set operating envelopes, with operators governing objectives and exceptions rather than taking each action. It is the top rung of a five-rung ladder — monitored, advisory, supervised closed-loop, domain-autonomous, self-optimising — not a product that can be installed.

Does any self-optimising utility exist today?

Not at estate scale, and this page says so plainly. What exists are the components, publicly documented: ML forecasting in national control rooms (National Grid ESO's reported 33% solar-accuracy gain), closed-loop optimisation of critical infrastructure under safety constraints (DeepMind's data-centre cooling), orchestrated DER fleets inside market frameworks (AEMO's VPP demonstrations), and NREL's field demonstrations of hierarchical self-optimising control. The end-state is an assembly problem with demonstrated parts, which is precisely what makes it a credible decade rather than a fantasy.

How is this different from the automation grids already run, like AGC and protection relays?

It is the same governance pattern with a wider scope. AGC and protection prove that utilities already run unattended control inside hard bounds — the fence, the settings discipline, the drills. What changes is that learned policies extend bounded automation from fixed, narrow physics (frequency, fault clearing) to wide decision spaces (dispatch, reconfiguration, orchestration) that previously needed human judgement per action. The fence moves; the principle that there is a fence, owned by operations and reviewed like protection settings, does not.

How far away is the self-optimising grid?

Three horizons. Today: advisory forecasting and first supervised loops are procurement decisions, not research bets. Three to five years: bounded domain autonomy — battery fleets, FLISR on instrumented feeders, predictive DER orchestration — at utilities that start accumulating envelope evidence now. Roughly a decade: cross-domain self-optimisation, arriving domain-pair by domain-pair. The calendar is set by governance throughput — how fast envelopes, override evidence and regulatory confidence accumulate — not by model capability, which is already ahead of the industry's ability to govern it.

What should a utility automate first?

Grid-scale storage, in almost every estate. A battery's actions reverse within the interval, its state of charge is perfectly observed, the system of record is the utility's own, and imbalance cost gives the loop an undisputed P&L. Feeder automation (FLISR, volt/VAR) follows where sensing exists. Customer-facing orchestration comes later — not because the optimisation is harder, but because wrong actions land on customer promises and the envelope negotiation crosses regulatory lines. Sequence by approval friction and reversibility, not by model difficulty.

What is a supervised closed loop in grid operations?

A loop whose output is written into the control system — EMS, ADMS or DERMS — as a pre-filled proposal the operator approves per action, inside a written envelope, with every accept and override logged and a drilled one-switch reversion. It is the pivotal rung: the human still decides everything, but deciding is one action in their own console, and the approval log becomes the evidence from which autonomy bounds are later set. Most programmes that stall, stall immediately below this rung.

How do regulators view autonomous grid control?

The regulatory direction is enabling but evidence-hungry. FERC Order 2222 requires wholesale markets to admit aggregated DER — the market rail for orchestrated fleets — and AEMO's VPP demonstrations exist precisely to establish how orchestration behaves inside the dispatch process. What regulators consistently require is not the absence of automation but its accountability: reconstructable decisions, versioned policies, tested fallbacks. A utility that builds the envelope-and-audit discipline this page describes is building its regulatory case as a by-product of its engineering.

What data foundation does autonomous grid balancing need?

Observability at the resolution of the decisions being automated, which for most utilities means fixing the distribution fog: AMI data on the operational cadence rather than the billing cadence, sensors on the feeder groups the first loops will run on, and — decisively — a state estimate reconciled against field reality that the control room would trust for a switching decision. Transmission-level sensing is largely solved. The governance test is blunt: a loop can never be more trustworthy than the state estimate it acts on.

How do you keep an autonomous control loop safe?

Four artefacts, all extensions of disciplines utilities already run: a versioned operating envelope owned by operations, a one-switch reversion drilled on a schedule, a control policy with a reviewable version history, and an audit trail that reconstructs any action months later. DeepMind's published account of autonomous cooling control is the canonical template — hard constraints, uncertainty-aware decisions, operators able to take back control at any time. The safety case is governance, not model quality, and it is designed first, not bolted on.

How do you measure the value of closed-loop dispatch?

Against a counterfactual agreed before go-live, in operational units. For a battery loop: captured revenue and imbalance cost versus the same optimiser advisory-only on comparable days, or versus unoptimised intervals rotated through the week. For network loops: SAIDI minutes and losses against holdout feeders. Forecast skill is a model property and proves nothing about value on its own. The pairing discipline matters too: report acceptance rate and attributed delta together — either one alone is theatre.

Does reinforcement learning actually run grids today?

In production, almost nowhere; in the supporting layers, increasingly. RL performs impressively in grid simulation and in research programmes such as NREL's autonomous energy systems work, and DeepMind's cooling controller demonstrated learned closed-loop control on live critical infrastructure. But today's production grid loops overwhelmingly run classical optimisation — unit commitment, optimal power flow, model-predictive control — with ML supplying the forecasts. That division of labour is a strength, not a gap: the governance frame this page describes is agnostic to what sits inside the optimiser box.

What team does the first closed loop need?

Smaller and more operational than most utilities assume: one integration engineer who knows the EMS or ADMS estate, one ML or optimisation engineer, a named control-room owner whose number the loop moves, and part-time protection, compliance and cyber review. The scarce ingredient is not data science but the control-room product owner — someone with the standing to sign the envelope and the patience to review override logs monthly. Programmes with only a data-science owner stall at advisory; programmes with an operations owner climb.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — forecasting, dispatch optimisation, vision inspection and decision support running against live operational data, integrated into the EMS, ADMS and DERMS layer rather than delivered as dashboards.

  • · Production deployments across generation, networks and energy retail
  • · Autonomy and maturity assessments run jointly with control-room teams
  • · Integration-first delivery: setpoint write-back, operating envelopes, monitoring, rollback
  • · 12 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. NRELAutonomous energy systems research (opens in a new tab)
  3. Google DeepMindMachine learning can boost the value of wind energy (opens in a new tab)
  4. Google DeepMindDeepMind AI reduces Google data centre cooling bill by 40% (opens in a new tab)
  5. Google DeepMindSafety-first AI for autonomous data centre cooling and industrial control (opens in a new tab)
  6. National Grid ESO / NESOESO and The Alan Turing Institute use machine learning to help balance the GB electricity grid (opens in a new tab)
  7. National Grid ESO / NESOAI tool could help boost National Grid ESO's solar forecasts (opens in a new tab)
  8. McKinsey & CompanyElectric power & natural gas insights (opens in a new tab)
  9. IEEE SpectrumEnergy coverage (opens in a new tab)
  10. FERCFederal Energy Regulatory Commission (Order No. 2222) (opens in a new tab)
  11. AEMOAustralian Energy Market Operator (VPP demonstrations) (opens in a new tab)
  12. EnelEnel Grids — the future of the electric grid (opens in a new tab)

Find your rung — then close your first loop

We run the autonomy assessment with your engineering and control-room leads, benchmark the result against utilities of similar network shape, and leave you with a costed 90-day plan for your weakest dimension. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.