Energy & UtilitiesFuture of AI & Visionary Thinking
Self-optimising utilities: the future of AI in energy & utilities, rung by rung
A self-optimising utility is one whose grid, generation and demand continuously rebalance themselves: AI adjusts dispatch, isolates faults and orchestrates distributed resources inside human-set operating envelopes. It is reached in five rungs — monitored, advisory, supervised closed-loop, domain-autonomous, self-optimising — and every rung below the last is already running somewhere in production today.

Key takeaways
- A self-optimising utility is not a grid run by an unsupervised AI — it is an enumerated set of control loops that sense, decide and act continuously inside human-set operating envelopes, with everything outside those envelopes escalating to an operator.
- The end-state is reached in five rungs — monitored, advisory, supervised closed-loop, domain-autonomous, self-optimising — and the first four already run in production somewhere: National Grid ESO's ML forecasting, DeepMind's closed-loop cooling control, AEMO's virtual power plant demonstrations.
- The binding constraint on grid autonomy is not model quality but safety and override governance: written operating envelopes, a drilled one-switch reversion to manual, and control policies versioned and reviewed like protection settings.
- Sensing caps autonomy. A control loop can never be more trustworthy than the state estimate it acts on, which is why distribution-level observability — AMI, feeder sensors, a reconciled as-operated network model — is the first investment on the ladder, not the last.
- The honest horizon: bounded domain autonomy — battery dispatch, feeder FLISR, DER fleet orchestration — is a 3–5 year build for a utility starting from advisory today; cross-domain self-optimisation is a decade out and arrives domain by domain, not as a single switchover.
Abbreviations used on this page
- SCADA
- Supervisory control and data acquisition
- EMS
- Energy management system (transmission control)
- ADMS
- Advanced distribution management system
- DERMS
- Distributed energy resource management system
- DER
- Distributed energy resource (rooftop solar, batteries, EVs, flexible load)
- BESS
- Battery energy storage system
- VPP
- Virtual power plant (an aggregated DER fleet bid as one unit)
- AGC
- Automatic generation control (the classical frequency-regulation loop)
- FLISR
- Fault location, isolation and service restoration
- PMU
- Phasor measurement unit (high-resolution grid sensor)
- AMI
- Advanced metering infrastructure (smart meters)
- SAIDI
- System average interruption duration index (with SAIFI, the headline reliability KPI pair)
Free · 8 questions · ~3 minutes
Score your utility on the autonomy ladder
Eight questions, one at a time, about three minutes. Answer them and we build your personalised autonomy report — your rung on the ladder, your score on each of the four dimensions, and the specific blocker between you and the next rung — and send it to your inbox. Your result doubles as the baseline for your first closed-loop business case.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised autonomy report is ready
Tell us where to send it. Your rung appears on screen straight away, and the full report — dimension scores, how you compare with utilities of similar network shape, and the 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Monitored
The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
Your next movePick one operational forecast — regional demand, embedded solar, one constrained feeder — and deliver it into the control room on the operational cadence, with its accuracy tracked.
Stage 2 · Advisory
Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
Your next moveChoose one bounded decision — battery schedule, feeder reconfiguration, reserve sizing — and write the model's proposal into the control system itself, operator-approved, with every accept and override logged.
Stage 3 · Supervised closed-loop
Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
Your next moveRun the loop supervised for at least two quarters, then let the override log — not opinion — define the conditions under which the loop may execute unattended.
Stage 4 · Domain-autonomous
Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
Your next moveBuild the co-ordination layer: shared state, a hierarchy of envelopes, and an arbitration policy that decides which domain yields when their objectives collide.
Stage 5 · Self-optimising
Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.
Your next moveTreat the objective function, the arbitration policy and every envelope as versioned, reviewable artefacts with the same rigour as protection settings — and staff the forum that reviews them.
0 / 24
Sensing & telemetry
— / 6
Control-loop automation
— / 6
Safety & override governance
— / 6
Value measurement
— / 6
Your score maps to a rung on the autonomy ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your autonomy — a rung-4 optimiser on rung-1 telemetry is a hazard, not a capability — and it is where the next investment belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the autonomy ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your autonomy — a rung-4 optimiser on rung-1 telemetry is a hazard, not a capability — and it is where the next investment belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this benchmarked against comparable utilities?
We walk your engineering and control-room leads through the dimension scores, compare them against operators of similar network shape and DER penetration, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–5 — Stage 1, Monitored. The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
- 6–11 — Stage 2, Advisory. Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
- 12–16 — Stage 3, Supervised closed-loop. Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
- 17–21 — Stage 4, Domain-autonomous. Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
- 22–24 — Stage 5, Self-optimising. Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.
What a self-optimising utility is — and what it is not
A definition, the boundary that keeps it honest, and the ladder of five rungs that connects today's control room to the end-state.
A self-optimising utility is one whose core operating decisions — grid balancing, network reconfiguration, generation and storage dispatch, demand orchestration — run as continuous closed loops: sensors feed a live model of the system, optimisers decide, actuators act, and the results feed back, with humans setting the objectives and the bounds rather than taking each action. The concept has a research pedigree: NREL's autonomous energy systems programme describes exactly this — hierarchical, decentralised control with 'dynamic self-optimisation' — as the only tractable way to run a grid with hundreds of millions of controllable devices.
What it is not is a grid handed to an unsupervised intelligence. Utilities have run automation inside hard bounds for a century — protection relays clear faults with no human in the loop, and AGC has trimmed generator output against frequency since long before machine learning. The self-optimising utility extends that settled pattern to learned policies and far wider decision spaces. The fence moves; the principle that there is a fence does not — which is why every rung on this page is defined by what may act without a person, inside what envelope, on what evidence.
Value released against position on the autonomy ladder
The curve is not linear. Value stays close to flat through the monitored and advisory rungs — where most utilities are — and inflects at the supervised closed loop, when model output first reaches the control system that acts. This is why programmes that measure progress in models built rather than loops closed report activity without results.
Operational value released by stage
- Stage 1 · Monitored — 34% of operators. The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
- Stage 2 · Advisory — 38% of operators. Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
- Stage 3 · Supervised closed-loop — 19% of operators. Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
- Stage 4 · Domain-autonomous — 7% of operators. Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
- Stage 5 · Self-optimising — 2% of operators. Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the IEA's Energy and AI analysis.
How a grid decision closes its loop at each rung
The control path, rung by rung. The rung is determined by where the arrow ends: rungs 1–2 terminate at a human reading a screen, rung 3 writes proposals into the EMS/ADMS for approval, and rungs 4–5 execute inside a versioned envelope with exceptions escalating. Most utilities are in the top lane.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- At rungs 1–2, SCADA and AMI data leave the operational systems as batch extracts, feed forecast models, and surface on an advisory screen beside the control systems. Whether the operator acts is optional and unlogged — and attention decays exactly when the model matters most, in storms and volatile settlement periods. This is where value leaks.
- At rung 3, telemetry streams continuously into a reconciled state estimate; the optimiser writes its output into the EMS, ADMS or DERMS as a pre-filled proposal — a battery schedule, a switching plan — and the operator approves or overrides each one, with every decision logged. The approval log becomes the evidence base for autonomy.
- At rungs 4–5, a versioned operating envelope owned by operations lets enumerated decision types — storage dispatch, FLISR, DER orchestration calls — execute unattended inside agreed bounds. Anything outside the envelope escalates to a human with full context, and every automated action carries a reconstructable audit trail.
Step-by-step insights
- Batch extracts — the habit that caps every rung above
- The day-late AMI extract and the hand-built feeder model are the artefacts that quietly cap a utility at rung 2. Everything downstream of a batch pull is stale on arrival and carries no lineage, so no control application can ever be more current than the export schedule. This is why the first investment on the ladder is almost never a model — it is streaming the telemetry the target loop needs and alarming its freshness, which is unglamorous and decisive in equal measure.
- The advisory screen dead end
- An advisory screen requires no change-board approval on a control system, which is exactly why pilots ship one — and exactly why they stall there. The recommendation lives beside the operator's real console, acting on it is a voluntary extra step, and voluntary steps are the first casualties of a storm shift. National Grid ESO's forecasting work shows this pattern's honest ceiling: genuine, measurable value in reserve and balancing decisions, capped by control-room bandwidth until the output moves into the control path itself.
- The state estimate — the loop's real foundation
- Setpoints are computed against the model, not against the field, so the as-operated state estimate is the component autonomy actually stands on. Transmission operators have trusted state estimation for decades; the frontier is extending it below the transmission interface, where DER growth has made the old fog operationally expensive. A utility that cannot state, with confidence bounds, what a named MV feeder is doing right now has found its rung — and its next investment — regardless of how good its models are.
- The optimiser — commodity maths, scarce framing
- The optimisation itself — unit commitment, optimal power flow, storage arbitrage, reinforcement-learning dispatch in research settings — is the most mature part of the stack. What distinguishes deployments that survive is the framing around it: objectives stated in operational units (imbalance cost, curtailed MWh, SAIDI minutes), constraints inherited from the envelope rather than hard-coded by the data team, and uncertainty carried through to the proposal so the operator sees confidence, not just a number.
- The proposal and the approval log
- Writing the proposal into the EMS or ADMS as a pre-filled action changes the default: accepting takes one click, declining is a logged choice with a reason. The log is not bureaucracy — it is the dataset autonomy is later built from. Six months of accept-and-override history tells you which proposals are always taken, which conditions drive overrides, and where the envelope actually binds. Skip the supervised period and you arrive at the autonomy conversation with opinions where the evidence should be.
- Envelope, escalation and the audit trail
- The rung 4–5 lane is a policy artefact more than a model artefact: an enumerated list of decision types that may execute unattended, the bounds per type, the escalation triggers, and the trail that reconstructs any action months later — model version, state estimate, envelope version, policy version. Utilities inherit this discipline from protection-settings governance, which is precisely why the operators who reach autonomy fastest treat the AI stack as an extension of operational governance rather than as an IT project.
The five rungs from monitored to self-optimising
For each rung: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what leaving costs.
Each rung below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a real control room, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that rung. A utility is at the rung its lowest-governed loop is at — a brilliant advisory forecast does not lift an estate whose fallbacks have never been drilled.
Select a rung
Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Monitored
34% of operators sit here
The network is instrumented and watched; every control action is initiated by a human, and AI exists only in studies.
Rung 1 is not a backward place — it is where a century of engineering discipline lives. The grid at this rung is already automated in the classical sense: protection relays clear faults in milliseconds, AGC trims generator output against frequency, tap changers hold voltage. What is absent is any learned model in the loop, and — more importantly — any real-time observability below the transmission interface. The system is safe, but it is safe because humans keep it inside a wide margin.
The tell is the state estimate. At rung 1 the control room has a trustworthy picture of the transmission system and a fog below it: distribution feeders modelled from as-built drawings that have drifted from field reality, AMI interval data that lands in a billing warehouse a day later, rooftop solar that appears only as suppressed demand. Two engineers asked how much embedded generation is on a feeder will give two defensible answers, because nothing reconciles the model against the field.
This rung becomes expensive the moment the network stops being passive. DER growth, EV charging and electrified heat all move the load shape faster than manual processes can track it, and the operator's response — wider margins, more curtailment, more conservative connection limits — is paid for by every connected customer. The cost of staying at rung 1 is not visible on any dashboard; it is embedded in every margin.
In practice
The feeder nobody can see
A distribution operator receives a connection application for a 5 MW solar farm on a rural feeder. The planning engineer's hosting-capacity study takes six weeks, because the feeder model has to be rebuilt by hand from GIS records and a decade of switching logs before any power-flow study can be trusted. The answer comes back conservative — connection refused pending reinforcement — not because the network is full, but because nobody can prove it is not.
What it looks like
- SCADA covers transmission and primary substations; below that, visibility fades
- AMI data is collected for billing, not for operations
- Forecasts are produced by planners in spreadsheets or vendor tools, offline
- No model output reaches the control room in real time
Diagnostic signals you can check this week
- Ask for the real-time state of a named MV feeder — if the answer involves a site visit or yesterday's AMI extract, you are here
- Check whether any forecast is produced inside the dispatch cycle rather than the day before
- Ask how much embedded solar is behind a given primary substation and compare two teams' answers
- Count the control actions taken in a shift that a model informed: at rung 1 the honest count is zero
Anti-pattern · Buying the autonomous platform first
The instinctive move is a flagship procurement — an AI-enabled ADMS or a grid analytics platform — on the theory that autonomy is a product you install. It fails predictably: the optimisation modules sit dark because the network model beneath them is stale and the telemetry they need does not exist. Instrument one part of the network to the standard a control application needs, prove one forecast in the control room, and let the platform requirements fall out of that experience.
What holds you here
The network below the transmission interface is not observable enough to model, so no learned system can be trusted to advise on it.
Highest-leverage next move
Pick one operational forecast — regional demand, embedded solar, one constrained feeder — and deliver it into the control room on the operational cadence, with its accuracy tracked.
Cost of leaving
- Effort
- 6–12 months
- Team
- One data engineer, one power systems engineer, part-time control-room sponsor
- Risk
- Low — the work is observational; nothing touches a control path yet
- To next stage
- 6–12 months
If this is you, the next step is
A 2-week review: which feeders, which telemetry, which model gaps block the first advisory use case.
Stage 2
Advisory
38% of operators sit here
Models forecast and recommend on the operational cadence; operators read them, trust them unevenly, and take every action themselves.
Rung 2 is where most of the industry's visible AI progress lives, and it is genuinely valuable. National Grid ESO's machine-learning work with the Alan Turing Institute improved national solar forecasting accuracy by a reported 33%, and better forecasts translate directly into lower reserve holding and cheaper balancing — value without a single automated action. An operator can stay honest here: the models advise, the humans decide, and the accountability chain is exactly what it was before.
The structural weakness is that advisory value is capped by human bandwidth and decays under stress. A control room in a storm, or a trading desk in a volatile settlement period, is precisely when the recommendations are worth most and precisely when nobody has time to read them. The advisory screen is a second cockpit: everything on it is optional, and optional things are the first casualties of a busy shift.
The other quiet failure of rung 2 is that it generates no evidence for the next rung. A recommendation read but never logged against the action taken produces no acceptance data, no override reasons, and no basis for ever setting autonomy thresholds. Operators who want a closed loop in three years must start recording accept-and-override on advisory output today — the log, not the model, is the long-lead item.
In practice
The wind forecast and the redispatch that didn't happen
A system operator's ML wind forecast flags a probable 400 MW over-forecast for the evening peak, six hours out. The advisory screen shows it; the shift is mid-handover, the constraint desk is working a transformer outage, and the recommendation scrolls off. The imbalance materialises and is settled at peak prices. The model was right, the value was real, and none of it was captured — because capture depended on a person having a quiet afternoon.
What it looks like
- ML forecasts — demand, wind, solar, prices — reach the control room in time to matter
- Recommendations appear on advisory screens beside, not inside, the control systems
- Forecast accuracy is measured and reported; action on it is voluntary
- Every setpoint, switch and dispatch instruction is still human-initiated
Diagnostic signals you can check this week
- Ask where a control-room operator sees model output: a separate screen or browser tab means rung 2
- Check whether accepted and ignored recommendations are logged anywhere
- Compare forecast usage on a quiet day against the last storm — the delta is the decay
- Ask what the model's accuracy is worth in balancing cost or reserve — if nobody can say, value is still notional
Anti-pattern · Chasing accuracy instead of a control path
When advisory output is under-used, the reflex is to improve the model — more features, better ensembles, another point of forecast skill. But usage is a function of where the output lands, not of marginal accuracy. A moderately accurate setpoint proposal inside the EMS, pre-filled and one click from acceptance, changes more dispatch decisions than an excellent forecast on a side screen. Build the path into the control system first; buy accuracy when you can price it in balancing cost.
What holds you here
Model output stops at a screen: acting on it is voluntary, unlogged and first to be dropped under pressure, so value depends on the calmest hour of the shift.
Highest-leverage next move
Choose one bounded decision — battery schedule, feeder reconfiguration, reserve sizing — and write the model's proposal into the control system itself, operator-approved, with every accept and override logged.
Cost of leaving
- Effort
- 9–18 months
- Team
- Integration engineer, ML engineer, a named control-room owner, protection/operations review
- Risk
- Medium — the first write into a control system needs an envelope, an approval step and a drilled reversion
- To next stage
- 9–18 months
If this is you, the next step is
The advisory-to-closed-loop transition is our most common energy engagement. Typically one quarter, one loop.
Stage 3
Supervised closed-loop
19% of operators sit here
Model output writes proposed setpoints and switching actions into the EMS, ADMS or DERMS; operators approve each action inside a written envelope.
Rung 3 is the pivotal rung, and it is defined by a change of address rather than a change of algorithm: the model's output moves from a screen beside the control system to a field inside it. A battery schedule arrives in the EMS as a proposed dispatch the operator confirms; a post-fault switching sequence arrives in the ADMS as a pre-built plan. The human still decides everything — but deciding is now one action inside their own console, and declining is a logged choice rather than a silent omission.
The disciplines that make this rung safe are borrowed from protection engineering, not data science. The operating envelope — which actions may be proposed, within what limits, under which network conditions — is written down, owned by operations, and treated with the seriousness of protection settings. The reversion to the previous scheme is a single switch, exercised deliberately on a quiet shift, because an undrilled fallback is a hypothesis. DeepMind's account of moving its data-centre cooling from recommendations to autonomous control describes exactly this frame: hard constraints, uncertainty-aware decisions, and operators who can take back control at any time.
What rung 3 buys, beyond its own operational value, is the evidence base for autonomy. Six months of accept-and-override logs on one loop tell you — with data, not judgement — which proposals are always accepted, which conditions drive overrides, and where the envelope is actually binding. That log is the raw material from which rung 4 thresholds are set, and the operators who skip it end up setting autonomy bounds by committee guesswork.
In practice
The battery that proposes its own schedule
A utility runs a 50 MW BESS against an intraday price and imbalance forecast. Each half-hour the optimiser writes a proposed charge/discharge schedule into the EMS; the control-room operator reviews and confirms it — usually in seconds, occasionally editing it around a known outage. Every confirmation and edit is logged. After two quarters the log shows 92% of proposals accepted unmodified, with overrides clustering in two named network conditions — which becomes the draft envelope for unattended operation.
What it looks like
- Proposals arrive as pre-filled actions in the operator's own console, not on a side screen
- A written operating envelope, owned by operations, bounds what may be proposed
- Every accept and override is logged with context and reviewed
- Reversion to the previous control scheme is one switch away and has been drilled
Diagnostic signals you can check this week
- Open the operator's console: model proposals should be visible there, pre-filled, without a second login
- Ask to see the operating envelope document and who signed it — operations, not the data team
- Ask when the reversion to manual was last drilled; a date is a pass, a shrug is a fail
- Pull one week of accept/override logs — if they don't exist, the loop is advisory with better plumbing
Anti-pattern · Skipping the supervised year
Vendors and pilots alike are tempted to jump from advisory straight to unattended operation — the demo works, the approvals feel like ceremony, and removing them is one config change. Resist it. The supervised period is not a trust ritual; it is the only source of the override data that makes autonomy bounds evidence-based, and the period in which operations learns the system's failure shapes cheaply. Every month skipped at rung 3 is paid back with interest after the first unattended mistake.
What holds you here
Every action still waits for a person, so loop throughput is capped by operator attention — and the case for removing the wait must be built from the approval log, which takes disciplined months to accumulate.
Highest-leverage next move
Run the loop supervised for at least two quarters, then let the override log — not opinion — define the conditions under which the loop may execute unattended.
Cost of leaving
- Effort
- 12–24 months
- Team
- Platform engineer, ML engineer, control-room product owner, protection and compliance review
- Risk
- Medium-high — the envelope, audit and drill disciplines must hold through staff turnover
- To next stage
- 12–24 months
If this is you, the next step is
We audit one live loop: envelope, override log, drill history, and what the data says about autonomy readiness.
Stage 4
Domain-autonomous
7% of operators sit here
Named domains — a battery fleet, a feeder group, a DER portfolio — run unattended inside versioned envelopes; humans manage exceptions and policy.
Rung 4 is autonomy with a fence around it. The utility has not handed the grid to a model; it has enumerated specific decision domains — intraday dispatch of a storage fleet, FLISR on instrumented feeder groups, orchestration of a contracted DER portfolio — and allowed each to run unattended strictly inside an envelope that operations wrote, signed and can revoke with one switch. The precedent is older than machine learning: AGC and protection relays have acted without per-action approval for decades. Rung 4 extends that settled pattern to learned policies and wider decision spaces — the governance frame is inherited, not invented.
The operator's job changes shape here. Instead of approving individual actions, the control room supervises populations of actions: watching override and escalation rates, reviewing the weekly excursion report, tuning envelopes as the network changes. The critical signal becomes the escalation rate — the share of decisions the loop hands back. A rise means the world has moved outside the policy's validity (a new interconnector, a cold snap, a tariff change reshaping evening load), and it should trigger an envelope review before it triggers an incident.
The hard engineering at rung 4 is mostly evidence engineering. A regulator, an auditor or a connection customer will eventually ask why the system took a specific action at 03:41 on a Tuesday, and the answer must be reconstructable: which model version, which state estimate, which envelope, which policy. Utilities already know how to do this — switching logs and protection-settings reviews are exactly this discipline — which is why the fastest climbers treat the autonomy stack as an extension of operational governance rather than as an IT system.
In practice
The feeder group that reconfigures itself
A distribution operator runs FLISR unattended across an instrumented feeder group: when a fault locks out a section, the system locates it from feeder sensors, isolates the faulted span and restores supply to healthy sections by closing ties — in under a minute, against a restoration plan a human would have taken twenty to build. Storm mode widens the envelope's escalation triggers; anything involving crew safety tags, backfeed risk or abnormal switching states goes straight to the operator with the proposed plan attached.
What it looks like
- Enumerated decision types execute without per-action approval, inside written bounds
- Out-of-envelope conditions escalate to an operator with full context
- Envelopes and control policies are versioned, reviewed and drilled like protection settings
- Every automated action carries a reconstructable audit trail
Diagnostic signals you can check this week
- Ask for the list of decision types that execute unattended — rung 4 operators can hand you an enumerated, versioned document
- Ask for last month's escalation rate on any autonomous loop, and whether it is trended
- Pick one automated action from last quarter and ask for its full reconstruction — model, state, envelope, policy versions
- Check whether the kill switch per domain has been exercised in the last six months
Anti-pattern · Inheriting thresholds across domains
The first autonomous domain works, and its envelope becomes the template: the battery loop's thresholds get copy-pasted onto the feeder loop, the summer envelope onto winter operation. But an envelope is evidence-shaped — it encodes one domain's override log under one range of conditions, and it transfers no better than a protection setting transfers to a different network. Every new domain re-earns autonomy from its own supervised period; every season and topology change gets an envelope review. The alternative is a bad automated action, and the usual response to that is a blanket switch-off — a two-rung regression from a single incident.
What holds you here
Each domain optimises alone: the battery fleet, the feeder group and the DER portfolio each run well inside their own envelope while leaving cross-domain value — and cross-domain conflicts — unmanaged.
Highest-leverage next move
Build the co-ordination layer: shared state, a hierarchy of envelopes, and an arbitration policy that decides which domain yields when their objectives collide.
Cost of leaving
- Effort
- 24+ months
- Team
- Platform team, control-room product owner, standing governance forum, cyber and compliance partners
- Risk
- Concentrated — low-frequency, high-consequence, regulatory in nature; the audit trail is the deliverable
- To next stage
- 24+ months
If this is you, the next step is
We run a scenario exercise against one live loop: envelope, escalation, reconstruction, kill switch.
Stage 5
Self-optimising
2% of operators sit here
Domains co-ordinate against system-level objectives, and the system tunes its own policies inside a governance frame humans set and audit.
Rung 5 is the visionary end-state, and honesty about it matters: no utility operates here across its whole estate today, and none will for years. What exists are bounded previews — NREL's autonomous energy systems research explicitly targets 'dynamic self-optimisation' for networks with hundreds of millions of controllable devices and has moved into campus- and community-scale field demonstrations. The end-state is best understood as the co-ordination of many rung-4 domains, not as a new kind of intelligence.
What genuinely changes at rung 5 is who optimises the optimisers. At rung 4, humans tune each envelope; at rung 5, the system proposes its own re-tuning — retraining on drift, adjusting dispatch policy as DER uptake shifts the load shape, re-weighting objectives within bounds the governance forum sets. The human role concentrates into three functions that do not automate: setting the objective function, auditing that the system's behaviour matches it, and owning the escalation tiers when it does not.
The sceptic's question — why climb this far at all? — has a quantitative answer. The IEA estimates that applying existing AI tools to grid operations could unlock up to 175 GW of transmission capacity without building a single new line, and projects data-centre electricity demand more than doubling to around 945 TWh by 2030. A grid absorbing that load growth, plus electrified transport and heat, with a workforce that is not doubling, closes the gap with control-loop automation or does not close it at all. Self-optimisation is not a luxury end-state; it is the operating model the load curve is forcing.
In practice
The evening peak, arbitrated
On a winter evening the price signal, the constraint forecast and a cold-snap demand pickup collide. The storage fleet wants to discharge for price; the network layer wants reserve held against a feeder constraint; the DER portfolio can shift 80 MW of contracted flexibility. The arbitration policy — versioned, human-owned — resolves the conflict against this year's weighted objectives, publishes the reasoning to the control room, and escalates only the one decision that breaches the cross-domain envelope. The operator on shift reviews one decision, not three hundred.
What it looks like
- Cross-domain optimisation: storage, network reconfiguration and DER flexibility are traded against each other continuously
- Policies retrain and re-tune themselves inside human-set meta-envelopes
- System-level objectives — cost, carbon, reliability — are explicit, weighted and versioned
- Humans govern: setting objectives, auditing decisions, owning the escalation tiers
Diagnostic signals you can check this week
- Ask whether any two autonomous domains share state and an arbitration policy, or merely coexist
- Ask who owns the system-level objective weights and when they were last reviewed
- Check whether policy re-tuning is itself logged, bounded and reversible
- Ask what the operator on shift actually reviews in an evening peak — populations and exceptions, or individual actions
Anti-pattern · Declaring the end-state by press release
The gravitational pull at this altitude is narrative: 'the self-optimising grid' announced while three-quarters of the estate is still at rung 2. The damage is not embarrassment — it is that the claim redirects governance attention from the unglamorous work (envelope reviews, escalation tuning, evidence trails) to the story, and the first public incident lands on a system whose stated capability exceeds its audited one. Let the rung ladder describe reality; the estate is at the rung its lowest-governed autonomous loop is at.
What holds you here
Sustaining self-optimisation is a standing governance discipline — objectives drift, envelopes rot and audit expectations rise, so the rung is held by review cadence, not by any model.
Highest-leverage next move
Treat the objective function, the arbitration policy and every envelope as versioned, reviewable artefacts with the same rigour as protection settings — and staff the forum that reviews them.
Cost of leaving
- Effort
- Continuous
- Team
- Platform and control teams plus a standing socio-technical governance forum with regulatory engagement
- Risk
- Systemic — cross-domain coupling means failures propagate; governance and cyber-security are the binding disciplines
If this is you, the next step is
A working session on the honest gap between your architecture today and a governable rung-5 target.
Where utilities sit on the ladder today
Illustrative distribution, synthesised from the adoption patterns in IEA and McKinsey electric-power research — to be replaced with measured values as assessments accumulate. The advisory rung is the mode and the plateau: forecasting is widespread, closed loops are rare, and the drop from rung 2 to rung 3 is the largest single transition loss on the ladder.
Share of utilities (illustrative)
- 34% — 1 · Monitored
- 38% — 2 · Advisory (the plateau)
- 19% — 3 · Supervised closed-loop
- 7% — 4 · Domain-autonomous
- 2% — 5 · Self-optimising
Source: Illustrative distribution, anchored to IEA Energy and AI and McKinsey electric-power research
What already runs in production today
The end-state is speculative; its components are not. Each rung of the ladder is demonstrated by a named, publicly documented programme.
Every rung below self-optimisation already runs in production somewhere, publicly documented by the operator itself. That is what separates this end-state from most visionary AI claims: the argument is not that the technology will arrive, but that demonstrated components — ML forecasting in a national control room, closed-loop optimisation of critical infrastructure, market frameworks for orchestrated DER fleets — have simply not yet been assembled by one utility at estate scale. The table below reads the public record against the ladder.
| Programme | Organisation | What it demonstrates | Rung |
|---|---|---|---|
| ML solar forecasting with the Alan Turing Institute | National Grid ESO | A reported 33% improvement in national solar forecast accuracy — advisory value in a live control room | 2 |
| Solar nowcasting with Open Climate Fix | National Grid ESO | Satellite-driven minutes-ahead forecasts built for control-room use | 2 |
| Wind output forecasting for market commitments | Google DeepMind | 36-hour-ahead ML forecasts that raised the value of wind energy by roughly 20% | 2–3 |
| Data-centre cooling: recommendations, then autonomous control | Google DeepMind | 40% cooling-energy reduction, later run closed-loop under safety constraints with operator override | 3–4 |
| Virtual power plant demonstrations | AEMO | Aggregated residential batteries dispatched as one orchestrated unit in the national market | 3 |
| Order No. 2222 — DER aggregations in wholesale markets | FERC | The market architecture that makes orchestrated DER fleets economic at scale | enabler |
| Autonomous energy systems research and field demonstrations | NREL | Hierarchical self-optimising control validated in campus- and community-scale deployments | 4–5 |
Improved solar forecasts will help us run the system more efficiently, ultimately meaning lower bills for consumers.
The regulatory rail matters as much as the technology. In the United States, FERC (opens in a new tab) Order No. 2222 requires wholesale markets to admit aggregated distributed energy resources, which is what makes an orchestrated fleet of household batteries a market participant rather than a demonstration. In Australia, AEMO (opens in a new tab) has run virtual power plant demonstrations to establish how orchestrated DER fleets behave inside the national dispatch process. Neither framework mentions machine learning — but both define the arena in which predictive DER orchestration and autonomous demand response become businesses, and both assume exactly the envelope-and-escalation governance this page's ladder is built on.
Three operators, read against the ladder
Publicly reported programmes only, cited to the operator's own material. None is an Atomic Loops engagement — the value is in the shape of each climb.
The clearest evidence for the ladder is in what serious operators chose to build first. In each case below, the differentiator was not model sophistication — it was sequencing: observability before advice, advice before proposals, a supervised period before any autonomy, and a safety frame carried the whole way up. Read each one for the rung transition it demonstrates rather than the headline number.
Three public climbs
Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them.
National Grid ESOGB electricity system operator · national control room12
- Challenge
- Balancing a grid with fast-growing embedded solar the control room cannot meter directly: rooftop generation appears only as suppressed demand, so forecast error feeds straight into reserve holding and balancing costs.
- Approach
- ML forecasting built for the control room rather than the lab: a random-forest and ensemble approach developed with the Alan Turing Institute for national demand and solar, followed by satellite-driven solar nowcasting with Open Climate Fix for the minutes-ahead horizon.
- Reported outcome
- A reported 33% improvement in solar forecasting accuracy, with the operator publicly linking improved forecasts to running the system more efficiently and lowering consumer bills.
- What it shows about the curveThis is the advisory rung done properly: value measured in operational units, delivered on the operational cadence, into the room where the decision happens — and trust built with the control room before any closed loop is proposed.
Google DeepMindHyperscale infrastructure operator · grid-scale load and renewables offtaker24
- Challenge
- Data-centre cooling plants with dozens of interacting setpoints and non-linear dynamics — energy-intensive, safety-critical, and beyond what static heuristics could optimise.
- Approach
- The canonical recommend-supervise-bound climb: 2016 recommendations implemented by human operators cut cooling energy; by 2018 the system operated the plant directly under a safety-first frame — hard-coded constraints, uncertainty-aware decisions, and operators able to take back control at any time. In parallel, 36-hour ML wind forecasts supported day-ahead market commitments.
- Reported outcome
- A reported 40% reduction in cooling energy, autonomous closed-loop operation under safety constraints, and a roughly 20% increase in the value of its wind energy — all published by the operator.
- What it shows about the curveThe middle rungs of the ladder, demonstrated end to end on live critical infrastructure: autonomy was earned through a supervised period and bounded by an explicit safety frame — precisely the governance a grid control loop needs.
EnelGlobal distribution operator · 69M+ end users across six countries13
- Challenge
- Making 1.9 million kilometres of distribution network observable and remotely operable — the precondition for any automated fault management or DER hosting at scale.
- Approach
- A decade-scale digitalisation of the network itself: smart metering at fleet scale, remote monitoring and control pushed into the grid, and a stated investment strategy of making networks increasingly digital, flexible and resilient — industrialised through its grid-technology arm.
- Reported outcome
- Enel reports 1.9 million km of power lines and 69.4 million active end users on increasingly digitalised networks, with 239.4 TWh distributed in the first half of 2026 — an observability-first platform on which automated reconfiguration becomes an increment rather than a leap.
- What it shows about the curveThe unglamorous truth of the ladder: for a distribution operator, the longest lead item in self-healing networks is not the algorithm but the sensing and remote-control fabric — and Enel bought that fabric first, at full network scale.
Where autonomy lands first on a utility
Six operating domains, the decisions in each that can close their loop, the system of record each loop must write into, and the KPI it moves.
Autonomy lands domain by domain, and the domains are not equally ready. A decision is a good early candidate when three things are true: its consequences reverse on the loop's own timescale, its domain is observable enough to model honestly, and the system of record it writes into is one the utility controls. The map below is how we scope first and second loops with operators — read the sweet-spot column as where each domain's loop typically earns its keep, not as a promise.
| Domain | Decisions that can close the loop | System of record | KPI it moves | Sweet spot |
|---|---|---|---|---|
| Balancing & system operations | Reserve sizing, redispatch proposals, constraint management | EMS / market systems | Balancing cost, frequency deviation | Rung 3 |
| Network operations | FLISR, feeder reconfiguration, volt/VAR optimisation | ADMS / OMS | SAIDI & SAIFI, losses | Rung 3–4 |
| Generation & storage dispatch | BESS charge/discharge, hybrid plant co-optimisation, AGC participation | EMS / plant controllers | Imbalance cost, curtailed MWh | Rung 3–4 |
| DER & demand orchestration | VPP dispatch, autonomous demand response, EV-charging shift | DERMS / aggregator platform | Peak shaved, flexibility delivered | Rung 3–4 |
| Asset & maintenance | Inspection triage, predictive maintenance scheduling, dynamic ratings | Asset management system | Unplanned outage rate, ratings headroom | Rung 2–3 |
| Trading & market bidding | Renewables commitment, storage arbitrage bids, imbalance positioning | Trading / ETRM systems | Captured spread, imbalance exposure | Rung 3 |
What to automate first
Plot each candidate loop's observability against the reversibility of its actions. The quadrant tells you the honest next step — and three of the four answers are not 'automate it'.
Propose, human executes
- Switching plans, redispatch, constraint actions
- Well observed but consequence-heavy
- Fix: supervised proposals with a drilled reversion
Close the loop
- BESS dispatch, VPP calls, volt/VAR on instrumented feeders
- Reversible next interval, fully observed
- This is where rung 3 → 4 happens first
Keep humans deciding
- Storm restoration priorities, safety-tagged switching
- Low observability, irreversible consequences
- Correctly held at advisory — possibly forever
Instrument first
- DER-heavy feeders with billing-grade data only
- Reversible actions, but the model is guessing
- Fix: telemetry and state estimation before any loop
Storage is the natural first domain: a battery's actions reverse within the interval, its state of charge is perfectly observed, the system of record is the utility's own, and imbalance cost gives the loop a P&L nobody disputes. Feeder automation follows where the sensing fabric exists. Customer-facing orchestration — autonomous demand response, EV-charging shift — belongs later, not because the optimisation is harder but because wrong actions land on customer promises and the envelope negotiation crosses regulatory lines. Sequencing by approval friction, not model difficulty, is what separates a three-year ladder from a decade.
The horizon map: now, five years, ten years
The honest calendar for the self-optimising utility — what is deployable today, what is a 3–5 year build, and what genuinely needs a decade.
The self-optimising utility arrives in three horizons, not one leap — and the calendar is set by governance and load growth, not by algorithms. The pressure is quantified: the IEA projects data-centre electricity demand alone more than doubling to around 945 TWh by 2030, while the same analysis estimates AI-based tools could unlock up to 175 GW of transmission capacity from existing lines. The gap between those two numbers is the business case for every rung on this page.
Three horizons to a self-optimising estate
Horizon boundaries are judgements, not forecasts — each horizon's claim is anchored to a programme that already exists. What moves a utility between them is governance throughput: how fast envelopes, evidence and regulatory confidence accumulate.
Today
Now — proven and deployable
ML forecasting in the control room (National Grid ESO's reported 33% solar-accuracy gain), closed-loop optimisation of bounded infrastructure under safety constraints (DeepMind's cooling control), orchestrated DER fleets in market frameworks (AEMO's VPP demonstrations, FERC Order 2222). Advisory value and first supervised loops are procurement decisions, not research bets.
Rungs 2–3 available to any utility that sequences them
3–5 years
Near horizon — bounded domain autonomy
Storage fleets dispatching unattended inside envelopes; FLISR standard on instrumented feeder groups; predictive DER orchestration and autonomous demand response operating inside the market rails now being laid; dynamic ratings feeding constraint management continuously. The limiting reagents are distribution-level observability and the regulatory evidence base — both compounding now at the operators that started.
Rung 4 in two or three domains at leading utilities
~10 years
Far horizon — cross-domain self-optimisation
Domains co-ordinating against explicit system objectives: storage, reconfiguration and flexibility traded off continuously, policies re-tuning inside human-set meta-envelopes — the pattern NREL's autonomous energy systems research is validating at community scale today, extended to utility estates. Arrives domain-pair by domain-pair, not as a switchover; the governance forum, not the platform, is the long-lead item.
Rung 5 behaviour on enumerated domain clusters
Watch the escalation rates, not the demos
The leading indicator of real autonomy is boring: published or auditable override and escalation rates on live loops. A utility that can show a falling override rate on a supervised battery loop is closer to the end-state than one announcing an AI platform.
Watch distribution observability spend
Feeder sensors, AMI-to-operations pipelines and state-estimation programmes are the tell that an operator is building the floor autonomy stands on. Enel's network-scale digitalisation is the reference pattern.
Watch the market rails
Each jurisdiction that operationalises DER aggregation — the FERC Order 2222 implementations, AEMO's evolving dispatch arrangements — converts predictive orchestration from a pilot into a revenue line, and pulls the DERMS layer up the ladder with it.
Watch who staffs the governance forum
The utilities that reach rung 4 first will be the ones whose autonomy governance is chaired by operations with protection-settings discipline — not delegated to a data team. The org chart is a better predictor than the tech stack.
For the engineering culture this demands, the closest public reference is the systems literature rather than the AI literature — the coverage of grid autonomy in IEEE Spectrum's energy reporting (opens in a new tab) and the sector analyses from McKinsey's electric power practice (opens in a new tab) both converge on the same conclusion this page's ladder encodes: the constraint on the self-optimising utility is organisational metabolism — how fast envelopes, evidence and trust accumulate — not model capability.