Redefining Technology

Energy & UtilitiesAI Implementation & Best Practices

AI grid and layout optimisation in energy and utilities: from one-off studies to a continuously optimised network

AI grid and layout optimisation is the pairing of machine-learning models with physics-based solvers to decide how an electricity network is configured and where assets are placed — optimal power flow, feeder reconfiguration, substation and renewable-plant layout, storage siting, hosting capacity — with results delivered into the GIS, ADMS and SCADA systems utilities already operate.

Utility network operations view with AI-optimised grid topology and asset layout overlays across substations and feeders
Energy & Utilities · AI Implementation & Best Practices

Key takeaways

  1. AI grid and layout optimisation splits into two problem families — operating the network you already have (optimal power flow, feeder reconfiguration, Volt/VAR) and designing what you build next (substation siting, storage placement, wind and solar plant layout) — and both are solved by the same solver-plus-ML pattern.
  2. Machine learning does not replace the physics solver — it screens, warm-starts and accelerates it. Surrogates triage millions of candidate configurations; the AC power-flow engine certifies the shortlist. A recommendation that skips certification must never become a switching order.
  3. The implementation ladder runs Study-level → Pilot → Operationalised → Fleet-wide → Continuous. Most utilities sit at Pilot, where optimisation results live in engineering studies and slide decks rather than in the ADMS.
  4. Network model quality caps everything: an optimiser is only as good as the GIS/CIM model it runs on, and connectivity or phase errors in that model are the most common reason a technically excellent switching plan cannot be trusted.
  5. The prize is quantified: the IEA estimates up to 175 GW of transmission capacity could be unlocked by AI-enabled tools such as dynamic line rating without building a single new line, and reports AI-based fault detection cutting outage durations by 30–50%.

Abbreviations used on this page

ADMS
Advanced distribution management system
SCADA
Supervisory control and data acquisition
GIS
Geographic information system (the network asset map)
OPF
Optimal power flow
DER
Distributed energy resource (rooftop PV, batteries, EV chargers)
AMI
Advanced metering infrastructure (smart meters)
CIM
Common Information Model (IEC 61970/61968 network data standard)
VVO
Volt/VAR optimisation
DLR
Dynamic line rating
EMS
Energy management system (the transmission control room)
FLISR
Fault location, isolation and service restoration
SAIDI
System Average Interruption Duration Index

Free · 8 questions · ~3 minutes

Score your grid optimisation capability

Eight questions, one at a time, about three minutes. Answer them and we build your personalised report — your stage on the implementation ladder, your score on each of the four dimensions, and the specific blocker between you and the next stage — and send it to your inbox. Your result doubles as the baseline for your first optimisation pilot.

0 of 8 answered

Question 1 of 8Network data & GIS quality

How well does your GIS/CIM network model match the as-built network an optimiser would run on?

An optimiser is only as good as its map. Connectivity and phase errors turn optimal switching plans into plans that cannot be trusted.

How the score maps to a stage
  • 05 — Stage 1, Study-level. Optimisation exists as one-off engineering studies — planning tools, consultants and spreadsheets — with no repeatable path from network data to a decision.
  • 611 — Stage 2, Pilot. A solver-plus-ML pilot demonstrably beats current practice on a bounded slice of the network, but its output reaches operations as a report or dashboard.
  • 1216 — Stage 3, Operationalised. Optimisation runs on a schedule against a maintained network model, and its recommendations arrive as draft plans in the tools operators already use, with monitoring and a named owner.
  • 1721 — Stage 4, Fleet-wide. A shared, versioned network model and optimisation platform serve multiple decisions across the whole service territory, with value measured against holdout feeders.
  • 2224 — Stage 5, Continuous. Enumerated low-risk actions execute closed-loop inside a versioned operating envelope, while engineers manage policy and exceptions rather than individual decisions.

What AI grid and layout optimisation is — and how the solver + ML pattern works

A definition, the two problem families, and the pipeline that turns a network model into a certified switching plan or siting decision.

AI grid and layout optimisation is the use of machine learning alongside physics-based solvers to answer two families of question a utility faces constantly: how should the network we already have be configured — optimal power flow, feeder reconfiguration, Volt/VAR settings, dynamic line ratings — and where should the assets we are about to build go: substations, storage, wind turbines within a farm, solar blocks and the cables between them. Both families reduce to constrained optimisation over the same object, the network model, which is why they belong to one implementation programme rather than two.

The division of labour is strict. Classical solvers — AC power flow, mixed-integer programmes, interior-point OPF — remain the source of truth, because they respect the physics and the limits. Machine learning makes them usable at operational scale: surrogate models approximate power-flow outcomes thousands of times faster than the full solver, so millions of candidate configurations or layouts can be screened; learned warm starts cut solve times on the cases that matter; reinforcement-learning agents propose topology actions a human would not have enumerated (the approach explored in RTE's Learning to Run a Power Network challenge (opens in a new tab)). Methods such as DeepOPF (opens in a new tab) report order-of-magnitude speedups on security-constrained OPF with small optimality gaps — but in every serious deployment the shortlist those methods produce is re-certified by the full physics engine before anything reaches an operator.

Value released against time on the implementation ladder

The curve is not linear. Value stays near flat through the study and pilot stages — where most utilities are — and inflects when recommendations start reaching the ADMS and planning workflow as draft plans someone approves. Programmes that count studies produced rather than configurations enacted report activity without results.

Network value released by stage

  • Stage 1 · Study-level — 26% of operators. Optimisation exists as one-off engineering studies — planning tools, consultants and spreadsheets — with no repeatable path from network data to a decision.
  • Stage 2 · Pilot — 37% of operators. A solver-plus-ML pilot demonstrably beats current practice on a bounded slice of the network, but its output reaches operations as a report or dashboard.
  • Stage 3 · Operationalised — 23% of operators. Optimisation runs on a schedule against a maintained network model, and its recommendations arrive as draft plans in the tools operators already use, with monitoring and a named owner.
  • Stage 4 · Fleet-wide — 11% of operators. A shared, versioned network model and optimisation platform serve multiple decisions across the whole service territory, with value measured against holdout feeders.
  • Stage 5 · Continuous — 3% of operators. Enumerated low-risk actions execute closed-loop inside a versioned operating envelope, while engineers manage policy and exceptions rather than individual decisions.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with IEA analysis of AI in grid operations.

The solver + ML optimisation pipeline, end to end

The pipeline that makes optimisation operational. The top lane assembles a validated network model; the middle lane screens with ML and certifies with physics; the bottom lane is where most programmes fail — delivery into the switching and planning workflow with an engineer in the loop and a tested reversion.

  • Data & feeds
  • AI / model
  • System-of-record action
  • Human in the loop
  • Where value leaks

The process, in words

  • The network-model lane assembles the one object everything depends on: the GIS/CIM asset model, reconciled against SCADA and AMI telemetry until a power flow on the model reproduces what the instruments actually measure, with the DER register folded in. Corrections are written back into the GIS, not kept in study files.
  • The optimisation lane runs the two-speed engine: ML surrogates screen the combinatorial space — millions of switch states, siting grids or turbine layouts — in seconds, and the full AC power flow certifies the surviving shortlist against thermal limits, voltage bands and protection flags. Only certified candidates continue.
  • The operations lane is where value is realised or lost: certified recommendations arrive as draft switching plans or ranked siting candidates inside the ADMS and planning tools, an engineer approves or amends every one, the previous configuration stays one switch away, and model-mismatch monitoring pages a named owner when the map and the network drift apart.
Step-by-step insights
The GIS/CIM model — the asset the whole programme stands on
Every optimisation answer inherits the map's errors. Distribution GIS models routinely carry connectivity mistakes, unrecorded switch states and unknown phase assignments — harmless for asset accounting, fatal for optimisation, because a reconfiguration plan that assumes the wrong normally-open point is a plan for a different network. The unglamorous first investment is validation and write-back: fix the model where studies find it wrong, and make the fix permanent in the GIS rather than in a study file that dies with the study.
Telemetry reconciliation — the test that makes the model trustworthy
The acceptance test for a network model is simple to state: run a power flow on it under a measured operating state and compare the computed flows and voltages with what SCADA and AMI actually recorded. Small residuals mean the map is usable; large or drifting residuals localise exactly where it is wrong — a mis-recorded conductor size, a phase swap, a switch whose real state differs from the record. AMI voltage correlation is the cheap trick of the decade here: smart-meter voltage profiles cluster by phase, so phase-identification errors can be found statistically without sending a crew.
ML screening — why the surrogate exists
Feeder reconfiguration across a few dozen switches, a substation siting grid, or turbine positions on a hilly site all share a shape: a combinatorial space far too large to evaluate with a full power flow or wake simulation per candidate. The surrogate — a network model-conditioned regressor, a graph neural network, or a learned wake approximation — evaluates candidates thousands of times faster with known, bounded error, which converts an intractable search into an afternoon's compute. Its job is recall, not final judgement: it must not discard good candidates, and it must never be the last word.
Physics certification — the non-negotiable gate
Every candidate that could become a switching order or an investment decision is re-run through the full physics: AC power flow under the relevant operating states, thermal and voltage limit checks, radiality verification for distribution reconfiguration, and flags for anything with protection implications — because moving a normally-open point changes fault paths, and a configuration that is optimal for losses can be inoperable for protection. The certification harness, not the ML, is what an operations engineer is being asked to trust, and it is why they can.
Delivery and approval — the lane that decides everything
A certified recommendation still has to become an enacted change, and that path runs through people with other priorities. Delivering results as draft switching plans inside the ADMS — or ranked, pre-checked candidates inside the planning tool — removes the voluntary step that kills dashboards: the default action becomes the informed one. The approval log this produces is not overhead; it is the dataset from which any future closed-loop bounds will be derived, one action class at a time.
Mismatch monitoring — the alarm on the map itself
Networks change constantly — new connections, reconductoring, switch operations — and the model must change with them. Monitoring power-flow residuals against live telemetry turns model decay into an alarm rather than a quarterly surprise: when residuals on a feeder drift, the map and the network have diverged, and every optimisation result on that feeder is quarantined until the divergence is explained. This is the control that lets the rest of the pipeline run on a schedule without an engineer re-validating the world by hand each time.

The five stages in detail

For each stage: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what leaving costs.

Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a utility's estate, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Study-level

26% of operators sit here

Optimisation exists as one-off engineering studies — planning tools, consultants and spreadsheets — with no repeatable path from network data to a decision.

Stage 1 is not the absence of optimisation — utilities have run power-flow studies for decades. What is missing is a supply chain for them. Every study begins by rebuilding the world: exporting the network from GIS, patching the connectivity errors the last study also patched, guessing load allocations, and hand-assembling a model that is obsolete before the report is bound. The analysis is often genuinely good; none of it is reusable.

The tell is the network model. Ask where the authoritative model of a feeder lives and a stage-1 utility gives three answers: the GIS says one thing, the planning tool's last study file says another, and the protection engineer keeps a third copy with corrections nobody merged back. Two engineers asked for the losses on the same feeder will produce two defensible, different numbers, because there is no shared model to be right about.

This stage is expensive precisely because it looks cheap. No platform is being paid for, but every interconnection application, every reinforcement decision and every reconfiguration question pays the full model-rebuilding tax — and with DER connection volumes rising, the number of studies demanded per year is growing faster than any planning team.

In practice

The hosting-capacity study that never ends

A distribution utility receives a steady stream of solar and battery interconnection applications. Each one triggers a manual screening study: an engineer exports the feeder from GIS, fixes the same phase and connectivity errors as last time, allocates loads from billing data, and runs a power flow to check voltage rise. The study takes days, the queue grows, and the fixes never make it back into the GIS — so the next application starts from the same broken map.

What it looks like

  • Load-flow and reconfiguration studies run per request in planning tools, results delivered as PDFs
  • The GIS network model diverges from the as-built network and nobody measures by how much
  • Hosting-capacity and interconnection screening is redone manually for every application
  • No named owner for optimisation outcomes; each study has a different author

Diagnostic signals you can check this week

  • Ask for the losses on one named feeder from two different teams and compare the numbers
  • Check whether corrections made during studies are ever written back into the GIS
  • Measure elapsed days from interconnection application to screening result, and the trend
  • Ask who owns optimisation outcomes. If the answer is a rotating cast of study authors, you are here

Anti-pattern · Buying the digital twin before fixing the map

The instinctive fix is an enterprise digital-twin or ADMS-analytics programme: eighteen months, a platform, a steering committee. It fails for a boring reason — the platform inherits the GIS errors and simply serves the wrong network faster. Validate and repair the model for one substation cluster first, prove a power flow that matches SCADA telemetry, and let the platform's real requirements emerge from use. The map has to be right before anything built on it matters.

What holds you here

There is no validated, current network model to optimise against, so every study starts by rebuilding the network from a map that is partly wrong.

Highest-leverage next move

Pick one substation cluster and build the data path end to end — GIS export, model validation against SCADA/AMI telemetry, a power flow that matches reality. Not a platform.

Cost of leaving

Effort
3–6 months
Team
One power-systems engineer, one data engineer, part-time GIS support
Risk
Low — model validation is additive; nothing in operations depends on it yet
To next stage
3–6 months

If this is you, the next step is

A 2–3 week engagement: one substation cluster, model vs telemetry, a repair backlog you keep.

Scope a network-model validation

Stage 2

Pilot

37% of operators sit here

A solver-plus-ML pilot demonstrably beats current practice on a bounded slice of the network, but its output reaches operations as a report or dashboard.

Stage 2 is the most dangerous stage on the ladder because it looks like success. The optimiser finds feeder configurations that cut modelled losses; the micrositing study adds modelled energy yield; the deck lands well. And the network is operating exactly as it was before, because the recommendation lives in a report and the switching order that would enact it was never drafted, checked against protection settings, or scheduled into a maintenance window.

The structural reason is that a pilot is scoped to answer 'does the optimiser work?' while the organisation's real question is 'will operations act on it?'. Those have different builds. Proving optimality wants a clean, frozen snapshot of the network; earning a switching order wants live model refresh, protection review, a fallback configuration and an operations engineer who trusts the pipeline — none of which improves the objective value, all of which decides whether the objective value ever becomes real losses avoided.

Time at stage 2 is not neutral. Control-room and planning engineers learn that optimisation output is advisory; sponsors learn that AI produces studies rather than megawatt-hours; and the next proposal is funded against that memory. A utility that has sat at stage 2 for three years is usually harder to move than one at stage 1, because the organisational verdict — interesting, not operational — has already been reached.

In practice

The reconfiguration study nobody enacted

A distribution operator ran a reconfiguration optimisation across a 10-feeder cluster: mixed-integer programme, ML screening to keep the search tractable, every candidate certified with a full AC power flow. Modelled losses fell meaningfully, all constraints held. The result shipped as a 40-page study. Eighteen months later the tie and sectionalising switches were in exactly the positions the legacy operating diagram prescribed — no switching order was ever raised, because raising one was nobody's job.

What it looks like

  • One optimisation use case — reconfiguration, VVO or siting — proven on backtest against a validated model
  • Results ship as a study report or BI dashboard, not as draft switching plans
  • Scope is one substation cluster, one wind farm or one planning area
  • No switching order, setpoint change or investment decision has yet been executed from the pilot

Diagnostic signals you can check this week

  • Count switching orders, setpoint changes or siting decisions executed from the pilot. Usually zero
  • Ask who would draft, check and schedule the switching plan if the result were accepted tomorrow
  • Check whether the pilot's network model has been refreshed since the study snapshot
  • Ask the control room what they think of the pilot. If they have not heard of it, integration was never scoped

Anti-pattern · Chasing the optimality gap instead of the switching order

When a pilot stalls, the reflex is to improve the optimiser — a tighter formulation, a fancier learning architecture, another percent of modelled loss reduction. It rarely moves anything, because the constraint is not solution quality but the absent path from recommendation to enacted change. A merely good configuration that becomes a checked, approved switching plan beats an optimal one in a PDF. Spend the next quarter building the draft-plan workflow into the ADMS or planning tool, then return to the model when you can price an optimality point in real megawatt-hours.

What holds you here

The pilot proves optimality but never earns a route into the switching and planning workflow, so its value depends on someone volunteering to act on a report.

Highest-leverage next move

Deliver recommendations as draft switching plans or study candidates inside the ADMS or planning tool operators already use — with an engineer approving every one.

Cost of leaving

Effort
6–12 months
Team
One integration engineer, one optimisation/ML engineer, a named operations owner
Risk
Medium — the first machine-drafted switching plan needs protection review and a drilled fallback
To next stage
6–12 months

If this is you, the next step is

The stage 2→3 transition is our most common engagement. Typically 90 days on one substation cluster.

Get the pilot into the ADMS

Stage 3

Operationalised

23% of operators sit here

Optimisation runs on a schedule against a maintained network model, and its recommendations arrive as draft plans in the tools operators already use, with monitoring and a named owner.

Stage 3 is the first stage at which the capability survives its founders. The optimisation has a schedule, a destination, an owner and an alarm: the network model refreshes from GIS and telemetry, the solver runs, candidates are certified with a full power flow, and the top plans appear where a planning or operations engineer already works — who approves, edits or rejects each one, with the previous configuration always one switch away. That loop, not the solver, is the production asset.

The character of the work changes here. Stage-2 problems are mathematical; stage-3 problems are operational, and the discipline that solves them looks like site reliability engineering applied to power systems. What is the freshness SLA on the model? What pages whom when power-flow residuals against telemetry drift? What is the tested fallback when a recommended configuration must be reverted mid-shift? Utilities that import these patterns wholesale move faster than those that rediscover them.

The constraint that emerges is bespoke-ness. The reconfiguration pipeline, the VVO study stack and the storage-siting tool were each built as their own thing — three network-model extracts, three validation scripts, three monitoring crons. By the third or fourth optimisation product the team's capacity is absorbed by keeping the existing ones healthy, and velocity falls exactly when the programme's credibility is highest.

In practice

The weekly configuration run

A utility runs feeder reconfiguration weekly across two pilot substation clusters. Sunday night the model refreshes and the optimiser proposes configurations for the coming week; Monday morning a distribution engineer reviews the top three in the ADMS, checks the protection implications flagged by the pipeline, and approves or amends. Roughly one week in four the recommendation is 'no change'. Losses are tracked against a neighbouring holdout cluster that stays on the legacy configuration — and that delta, not the solver's objective value, is what the programme reports upward.

What it looks like

  • Scheduled optimisation runs against a refreshed, validated network model
  • Recommendations land as draft switching plans or ranked study candidates in the ADMS/planning workflow
  • Model mismatch — power-flow results vs SCADA/AMI telemetry — is monitored and alarmed
  • A named engineer owns the pipeline's operational behaviour and its rollback

Diagnostic signals you can check this week

  • Check whether the network model the optimiser uses refreshes automatically or by hand
  • Count distinct model-extract and validation implementations across optimisation use cases
  • Ask what fraction of recommendations are approved unchanged, and whether anyone tracks it
  • Ask what happens when power-flow results diverge from telemetry. If the answer is 'someone would notice', monitoring is aspirational

Anti-pattern · Scaling by cloning the pilot

The temptation at stage 3 is to replicate the working pilot per region — copy the pipeline, point it at the next cluster, hire another engineer. Each clone forks the network-model extract and the monitoring, so the estate's maintenance burden grows linearly while its consistency decays. The leverage is the opposite move: extract the shared layer — one validated model pipeline, one certification harness, one monitoring stack — so the next cluster is configuration. Do it before the fourth use case, with the team you already have.

What holds you here

Every optimisation product carries its own model extract, validation and monitoring, so the fourth use case costs as much as the first and the team becomes the bottleneck.

Highest-leverage next move

Extract the shared layer — one network-model pipeline, one power-flow certification harness, one monitoring stack — so the next use case is configuration rather than a project.

Cost of leaving

Effort
9–18 months
Team
Platform engineer, optimisation engineer, an SRE-style on-call arrangement
Risk
Medium — the shared-layer refactor competes with new use-case demand for the same people
To next stage
9–18 months

If this is you, the next step is

We map your existing pipelines and identify what is genuinely shareable.

Review your optimisation platform layer

Stage 4

Fleet-wide

11% of operators sit here

A shared, versioned network model and optimisation platform serve multiple decisions across the whole service territory, with value measured against holdout feeders.

At stage 4 the marginal cost of the next optimisation decision collapses. Because the network model, certification harness, delivery workflow and monitoring are shared, adding a use case — say, storage-siting screening for a new planning area — is mostly specification: which objective, which constraints, which metric, which approver. Utilities at this stage stop talking about 'AI projects' and start maintaining a roadmap of decisions, which is a structurally healthier conversation.

The measurement discipline is what separates stage 4 from a well-run stage 3. Every deployed optimisation has an operational metric a budget holder already tracks — technical losses in MWh, voltage-violation hours, hosting capacity released in MW, curtailment avoided — and where the network allows it, a holdout: comparable feeders left on legacy practice so the delta is attributable rather than asserted. This is unglamorous, and it is what keeps the programme funded through the budget cycle in which someone asks what the optimiser actually did.

The remaining constraint is deliberate: a human still approves every switching plan and every siting decision, so throughput is bounded by engineer capacity. For most decision classes that is the correct resting state — the move to closed-loop execution is a risk-appetite decision about a small, enumerated set of low-consequence actions, not a technical inevitability.

In practice

The three-week use case

A utility with a shared model and platform layer scoped storage-siting screening for congestion relief in one planning region. Agreeing the objective, the constraint set and the holdout design with the planning lead took two weeks; the engineering — a new objective function and a report template on the existing pipeline — took three days. That ratio, specification-heavy and build-light, is the stage-4 signature.

What it looks like

  • One versioned network model feeds reconfiguration, VVO, hosting capacity and siting studies alike
  • A new optimisation use case ships in weeks — mostly specification, not engineering
  • Every deployed use case has a named operational metric and, where feasible, a holdout
  • Planning and operations work from the same model, ending the two-map problem

Diagnostic signals you can check this week

  • Time from decision agreed to results in production, for the last three optimisation use cases
  • Whether one monitoring view covers every deployed optimisation
  • Whether any deployed use case has a live holdout, and who reads the delta
  • Whether a planning leader, not an engineer, can name what each optimisation is worth

Anti-pattern · Closing the loop because you can

Stage 4 makes closed-loop execution technically easy, which is exactly when it gets extended past the evidence. Approval logs from one action class — capacitor switching, say — get generalised into autonomy thresholds for reconfiguration actions that were never in those logs. The first automated switching mistake then gets every loop opened again, and the programme loses more ground than autonomy ever gained. Each action class re-earns autonomy from its own approval history, or it stays human-approved.

What holds you here

Every action still passes through an engineer, so optimisation throughput is bounded by approval capacity rather than by the platform.

Highest-leverage next move

Define the operating envelope — confidence, consequence and network-state bounds — under which specific low-risk action classes execute without approval, with the audit trail that makes it defensible.

Cost of leaving

Effort
18+ months
Team
Platform team, operations product owner, protection engineer, risk partner
Risk
Higher — governance, protection coordination and audit evidence become the binding constraints
To next stage
18+ months

If this is you, the next step is

Which action classes should execute unattended, and the evidence to prove each is safe.

Design your closed-loop bounds

Stage 5

Continuous

3% of operators sit here

Enumerated low-risk actions execute closed-loop inside a versioned operating envelope, while engineers manage policy and exceptions rather than individual decisions.

Stage 5 is narrower than the phrase 'self-optimising grid' suggests. It is a specific, enumerated set of actions — Volt/VAR setpoint updates, dynamic-line-rating-informed limits, routine loss-reducing reconfiguration on designated clusters — executing unattended inside a stated envelope, with everything outside that envelope escalating to an engineer. Anything touching protection philosophy, safety clearances or high-consequence customers is correctly held at stage 4 indefinitely.

By the time a utility arrives here, the engineering is largely solved; the hard artefact is evidence. A regulator, an auditor or a connections customer will eventually ask how a decision made without a human was made — and the answer must be reconstructable: which model version, which telemetry snapshot, which envelope version, which certification run. Treat the operating envelope with the same rigour as protection settings: versioned, reviewed, signed off, with a record of who changed what and why.

Sustaining stage 5 is a governance discipline, and it is the stage most prone to silent regression. The network changes underneath the envelope — a new DER cluster shifts flows, an EV corridor changes load shape — and thresholds set on last year's network quietly stop being valid. The leading indicator is the escalation rate: when the share of actions falling outside bounds trends up, the world has moved and the envelope needs review before an incident forces one.

In practice

The bounded loop

A utility runs closed-loop VVO across designated feeders and scheduled reconfiguration on two clusters, inside an envelope stating voltage bands, switching-frequency ceilings, excluded network states and customer classes. Roughly one action in twenty-five escalates for review. The escalation rate itself is on a dashboard the standing governance forum reads monthly; a sustained rise triggers an envelope review, and the reversion to legacy configurations is drilled twice a year on quiet shifts.

What it looks like

  • VVO setpoints, DLR-informed limits and routine reconfiguration execute automatically within bounds
  • Engineers set and review the operating envelope; exceptions escalate to a person
  • Every automated action carries a reconstructable audit trail
  • Reversion to the legacy configuration is a tested procedure, exercised on schedule

Diagnostic signals you can check this week

  • Whether the operating envelope is versioned and reviewed like protection settings
  • Whether the reversion procedure has been exercised in the last six months
  • Whether escalation rate is monitored as a leading indicator with a named reader
  • Whether any single automated action can be reconstructed end to end from logs

Anti-pattern · Treating the envelope as configuration

Bounds get tuned in a settings screen — no review, no version history, no record of who widened the voltage band on which date. The loop works right up until someone must explain an action taken eight months ago, and neither the model version nor the envelope that permitted it can be reconstructed. Version the envelope, review changes like protection-setting changes, and keep the trail.

What holds you here

Sustaining closed-loop operation is a governance problem — the constraint is envelope validity and regulatory evidence, not engineering.

Highest-leverage next move

Treat the operating envelope as a versioned, reviewable artefact with the same rigour as protection settings, and monitor escalation rate as its early-warning signal.

Cost of leaving

Effort
Continuous
Team
Platform team plus a standing governance forum including protection engineering
Risk
Concentrated — low frequency, high consequence, regulatory in nature

If this is you, the next step is

We stress-test the envelope, the trail and the reversion against a real scenario.

Audit a closed-loop decision path

Where utilities actually sit on the ladder

The distribution across the five stages, why the pilot plateau dominates, and the external evidence for what leaving it is worth.

Most utilities sit at the pilot stage: an optimisation has beaten current practice on a bounded study, and nothing in the control room or the planning workflow has changed. The distribution below is illustrative — synthesised from the adoption picture in IEA and EPRI grid-modernisation research rather than measured from a single survey — but the shape is consistent everywhere we look: a large pilot plateau, a thin operationalised band, and closed-loop operation confined to a handful of action classes at a handful of operators.

Distribution of utilities across the five stages

Illustrative distribution — synthesised from IEA and EPRI grid-modernisation research, not a single measured survey. The pilot stage is the mode and the plateau; the drop from Pilot to Operationalised is the largest single transition loss on the ladder.

Share of utilities (illustrative)

  • 26% — 1 · Study-level
  • 37% — 2 · Pilot (the plateau)
  • 23% — 3 · Operationalised
  • 11% — 4 · Fleet-wide
  • 3% — 5 · Continuous

Source: Illustrative, synthesised from IEA Energy and AI and EPRI grid-modernisation research

The external evidence for the far side of the plateau is unusually concrete for an AI field. The IEA's Energy and AI analysis (opens in a new tab) estimates that up to 175 GW of transmission capacity could be unlocked by AI-enabled tools such as dynamic line rating without building new lines, and reports AI-based fault detection reducing outage durations by 30–50%. EPRI (opens in a new tab) runs standing research programmes on grid analytics and DER integration that many of these methods graduate through, and the US DOE Office of Electricity (opens in a new tab) funds the ADMS and grid-modernisation work that is making optimisation-ready control rooms normal rather than exotic. On the market side, system operators publish the operational pain these tools target — CAISO (opens in a new tab) reports renewable curtailment volumes that grow with every interconnection cohort, and PJM (opens in a new tab) processes an interconnection queue measured in hundreds of gigawatts.

The optimisation problem map: what to solve, with what, moving which KPI

Eight problems cover most of the grid-and-layout space. Each has a natural formulation, a home system, a KPI a budget holder already tracks — and a stage at which it earns its keep.

AI grid and layout optimisation concentrates in eight problem classes, and choosing the first one well matters more than solving it brilliantly. A problem is a good first candidate when three things are true: the network model for its area can be validated in weeks, the decision cycle is short enough to measure inside a quarter, and the KPI it moves is one a regulator or budget holder already tracks. The map below is how we scope first and second use cases with utilities.

ProblemFormulation & methodData & systemsKPI it movesSweet spot
Optimal power flow (OPF) & dispatchAC/DC OPF, security-constrained; ML surrogates and warm starts for speedEMS, SCADA, network modelProduction cost, congestion, curtailmentStage 3–4
Feeder reconfigurationMixed-integer programme over switch states with radiality; ML screeningGIS/CIM, ADMS, SCADATechnical losses, voltage violationsStage 3–4
Volt/VAR optimisation (VVO)Setpoint optimisation over OLTCs, capacitors, smart invertersADMS, AMI voltagesViolation hours, CVR energy savingsStage 4–5
Hosting-capacity analysisRepeated power-flow sweeps per feeder; ML acceleration for scaleGIS/CIM, AMI, DER registerMW connectable, screening turnaroundStage 2–3
Storage siting & sizingPlacement/size co-optimisation vs reinforcement deferral and congestionPlanning tool, load forecasts, network modelReinforcement capex deferred, curtailment avoidedStage 3–4
Substation siting & network expansionExpansion-planning MILP over load-growth scenariosGIS, load forecasts, planning toolCapex per MW served, connection lead timeStage 2–3
Wind micrositing & solar layoutLayout optimisation under wake/terrain models; cable routing as constrained network designMet data, terrain/GIS, yield modelsEnergy yield, LCOE, balance-of-plant costStage 2–3
Dynamic line rating (DLR)Weather-conditioned thermal models setting real-time limitsWeather feeds, line sensors, EMSTransmission capacity released (MW)Stage 4–5
The grid and layout optimisation landscape. 'Sweet spot' is the ladder stage at which the problem typically earns its keep — attempting a stage-5 candidate from a stage-1 network model is the fragile-automation quadrant described later on this page.
  • Pick problems whose system of record you control

    Feeder reconfiguration and VVO live entirely inside your GIS, ADMS and SCADA estate; a transmission-level OPF change may involve a market operator's processes and timelines. Early use cases that stay inside your own four walls ship quarters faster.

  • Prefer decisions that recur

    A weekly reconfiguration run generates fifty-two attributable deltas a year and an approval log; a one-off substation siting generates one decision and no learning loop. Recurring decisions compound the pipeline investment — one-off layout studies are best run as consumers of the platform the recurring decisions justify.

  • Refuse problems the network model cannot yet support

    If power-flow results on a feeder do not reconcile with its telemetry, no optimisation on that feeder is decision-grade — whatever the solver says. Sequencing model validation ahead of ambition is the single most reliable predictor of a programme that survives contact with operations.

What the transitions look like in public

Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the operator's own published material.

The clearest evidence for the model-first thesis is in what large operators chose to build first. In each case below the differentiator was not a novel algorithm — it was that a shared, trustworthy representation of the network was treated as the foundation, and the optimisation and planning tools were built as consumers of it.

Three programmes read against the ladder

Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them.

High-voltage transmission infrastructure in the PJM regional gridPJM InterconnectionUS regional transmission organisation · 13 states + DC23
Challenge
The largest US grid operator faced an interconnection queue measured in hundreds of gigawatts, with each application studied against planning models spread across separate databases and tools — a per-study workflow that could not keep pace with demand growth.
Approach
In April 2025 PJM announced a multiyear collaboration with Google and Tapestry to apply AI to regional planning and generation interconnection — starting not with an optimiser but with a unified, cloud-based, version-controlled model of the PJM grid that brings the planning databases and tools into one collaborative representation, with AI models to verify applications and speed study processing.
Reported outcome
As reported by PJM, the collaboration aims to significantly cut processing times for new interconnection requests as PJM works through the roughly 67 GW remaining in its transition-phase queue into 2026 — applying AI first to the data and model layer that every downstream study depends on.
What it shows about the curveEven the operator with the deepest solver bench treats the shared network model as deliverable number one. Optimisation tooling is only as scalable as the model layer it runs on — which is exactly the stage 2→3 lesson at distribution scale.

PJM Inside Lines — PJM, Google & Tapestry announcement (opens in a new tab)

Distribution network infrastructure in a European service territoryE.ONEuropean energy networks group · distribution grids across Europe24
Challenge
Surging connection volumes — heat pumps, EV charging and rooftop PV — hitting low-voltage networks whose expansion planning was a manual, section-by-section engineering study, on network records that were never digitised for automated analysis.
Approach
E.ON has publicly described digitising its distribution-network data and applying AI-supported grid planning across its German network businesses — generating and comparing expansion options automatically so planners evaluate machine-produced candidates rather than drafting each study by hand.
Reported outcome
As communicated by E.ON, AI-supported planning is being rolled out across its distribution subsidiaries to keep pace with the energy transition's connection growth — turning grid planning from a per-request study into a repeatable, model-driven pipeline.
What it shows about the curveFleet-wide only works because planning became a pipeline on a shared digitised model rather than a craft performed per section. The layout problem — where reinforcement goes — industrialises exactly like the operational one.

E.ON — company communications (opens in a new tab)

Utility-scale renewable generation and distribution assets operated by a global utilityEnelGlobal utility · distribution networks in Europe & Latin America24
Challenge
Planning and operating distribution networks across many countries, with asset records of varying quality and investment decisions that had to be comparable across territories.
Approach
Enel has publicly described building digital representations of its distribution networks and applying AI across its grids business — network digital twins supporting inspection, planning and network development, so the same modelled asset base underpins decisions in every territory.
Reported outcome
As reported in Enel's own material, the digital-twin programme supports planning and investment decisions across the distribution networks it operates worldwide — one shared model layer consumed by many analyses, rather than one study at a time.
What it shows about the curveThe digital twin is the asset; every optimisation is a consumer of it. Multi-territory operators reach fleet-wide by standardising the model layer first — the analyses then travel between territories almost for free.

Enel — company publications (opens in a new tab)

The four dimensions that set your stage

Capability is not one number. Four dimensions gate each other, and the lowest is the real stage.

Your stage is set by the lowest of four dimensions, because each one gates the others: network data and GIS quality, solver and model integration, operational integration, and value measurement. A sophisticated solver estate running on an unvalidated network model produces confident answers about a network that does not exist — which is worse than no answer, because it spends trust that integration will later need.

  • Network data & GIS quality

    Whether a power flow on the model reproduces what SCADA and AMI measure, whether corrections are written back, and whether DER connections appear in the model at the pace they appear on the network. This dimension caps all three others, and it is the least glamorous to fund.

  • Solver & model integration

    Whether optimisation is a repeatable pipeline — versioned, re-runnable, ML-screened, physics-certified — or a sequence of hand-assembled studies. The certification harness matters more than the learning architecture: it is what an operations engineer is actually being asked to trust.

  • Operational integration

    Where results land and how fast the pipeline reflects network change. This is the dimension that separates Pilot from Operationalised, and in the programmes we review it is overwhelmingly the lowest-scoring of the four.

  • Value measurement

    Whether success is defined in units a budget holder tracks — losses in MWh, violation hours, MW of hosting capacity released, capex deferred — and attributed against holdout feeders rather than asserted from the solver's objective value. Programmes measured on objective values alone lose funding at the first budget review.

Diagnosing the real constraint

Plot your network-model quality against your operational integration. The quadrant names the next investment — and three of the four answers are not 'a better optimiser'.

Blocked at the control room

  • Trustworthy model, no route into operations
  • Highest-leverage position on the matrix
  • Fix: build the draft-plan workflow, not another study

Scaling

  • Both foundations in place
  • Constraint is now use-case throughput
  • Fix: extract the shared platform layer

Studying

  • Neither foundation in place
  • Normal at stage 1
  • Fix: validate the model for one substation cluster

Fragile automation

  • Integrated pipeline on an untrusted map
  • The most dangerous quadrant on the page
  • Fix: mismatch monitoring before any further rollout
Network model quality — top: Validated against telemetry, bottom: GIS diverges from as-built
Operational integration — left: Results land in reports, right: Results land in the ADMS

The reference architecture, layer by layer

What actually has to exist at each stage of the ladder — defined by what each layer must guarantee, not by any vendor's product map.

An operationalised grid-optimisation capability requires five layers, and the order in which you build them decides whether the programme compounds or stalls. The architecture below is deliberately unfashionable: nothing in it is vendor-specific, every layer is defined by the guarantee it must provide, and the two layers most often skipped — validation and observability — are the two that decide whether anyone in operations will act on the output.

Layers required by stage

Each layer is annotated with the ladder stage that first requires it. A programme reaching for Operationalised without the delivery and observability layers is building a longer pilot with extra steps.

  1. Source systems

    Stage 1+

    • GIS / CIM asset modelThe map: assets, connectivity, phases, switch states
    • SCADA / EMS + AMIFlows, voltages and load shapes the model must reproduce
    • DER register & connection queueWhat is connected, what is coming, and where
  2. Network model & data layer

    Stage 2+

    • Model validation & write-backPower flow vs telemetry; fixes land in the GIS
    • State & load profilesOperating states and AMI-derived load shapes, versioned
    • Phase identificationAMI voltage correlation finds phase errors statistically
  3. Optimisation & ML layer

    Stage 2+

    • Physics solversAC power flow, OPF, MILP reconfiguration — the source of truth
    • ML surrogates & warm startsScreening at scale with bounded, measured error
    • Layout optimisersSiting grids, micrositing under wake/terrain models, cable routing
  4. Decision delivery layer

    Stage 3+

    • Draft plans into the ADMSSwitching plans and setpoint changes, pre-checked
    • Study workbenchRanked siting and expansion candidates for planners
    • Approval & reversionEngineer approves; previous configuration one switch away
  5. Observability & governance

    Stage 3+

    • Mismatch & freshness alertingModel vs telemetry residuals page a named owner
    • Decision audit logReconstructable months later, per action
    • Versioned operating envelopeReviewed like protection settings (stage 5)

Pipeline described

  1. Source systems (stage 1+) — GIS / CIM asset model: The map: assets, connectivity, phases, switch states; SCADA / EMS + AMI: Flows, voltages and load shapes the model must reproduce; DER register & connection queue: What is connected, what is coming, and where
  2. Network model & data layer (stage 2+) — Model validation & write-back: Power flow vs telemetry; fixes land in the GIS; State & load profiles: Operating states and AMI-derived load shapes, versioned; Phase identification: AMI voltage correlation finds phase errors statistically
  3. Optimisation & ML layer (stage 2+) — Physics solvers: AC power flow, OPF, MILP reconfiguration — the source of truth; ML surrogates & warm starts: Screening at scale with bounded, measured error; Layout optimisers: Siting grids, micrositing under wake/terrain models, cable routing
  4. Decision delivery layer (stage 3+) — Draft plans into the ADMS: Switching plans and setpoint changes, pre-checked; Study workbench: Ranked siting and expansion candidates for planners; Approval & reversion: Engineer approves; previous configuration one switch away
  5. Observability & governance (stage 3+) — Mismatch & freshness alerting: Model vs telemetry residuals page a named owner; Decision audit log: Reconstructable months later, per action; Versioned operating envelope: Reviewed like protection settings (stage 5)
Step-by-step insights
Source systems — control of change is the timeline
The single biggest lever on a programme's schedule is whether you control change on the systems the pipeline reads and writes. A distribution utility's GIS, ADMS and AMI are its own; a transmission-level programme may run through a market operator's model-management processes and calendars. Sequence early use cases inside your own estate — reconfiguration, VVO, hosting capacity — and take externally-governed decisions later, once the platform has attributable wins to spend.
Network model layer — validation is a product, not a phase
The mistake is treating model validation as a one-off cleanse before the real work. Networks change weekly, so validation must be a standing product: scheduled reconciliation of power-flow results against telemetry, residuals tracked per feeder, corrections written back into the GIS under change control, and a quarantine rule that keeps optimisation off any feeder whose residuals are drifting. Utilities that fund this as a product find every later use case inherits a trustworthy map for free.
Optimisation layer — keep the solver replaceable
Bind use cases to a certification contract, not to a specific solver or learning architecture. The contract says: every candidate that could become an action is certified by a full physics run against the current model, with limit and protection checks recorded. Behind that contract you can swap a MILP for a heuristic, add a graph-network surrogate, or retire a model — without renegotiating trust with operations, because the thing they trust is the harness, not the algorithm.
Delivery layer — the reversion is the approval unlock
The component most often skipped is the tested reversion: the legacy configuration or setpoint set, restorable with one action. It looks like engineering pessimism; it is actually the political key that unlocks operational approval, because a control-room manager will accept a new decision source they can instantly undo. A draft-switching-plan proposal without a drilled reversion sits in a change queue for two quarters; one with it ships.
Observability & governance — evidence as a by-product
At stage 3 the decision log is an engineering convenience; by stage 5 it is the artefact a regulator or connections customer will actually examine. Building it as a by-product of the delivery layer — every recommendation, certification run, approval and reversion recorded — means the evidence for closed-loop operation accumulates for free, and the governance conversation becomes an export from the system rather than a project alongside it.

Read the layers bottom-up when auditing an existing programme: if power-flow results do not reconcile with telemetry, nothing above the second layer is decision-grade, whatever the demo shows. Read them top-down when scoping a new one: the delivery workflow you can realistically get approved — draft plans with engineer sign-off, almost always — defines how much of the lower stack the first quarter actually needs.

A 90-day plan: feeder reconfiguration on one substation cluster

The Pilot → Operationalised transition made concrete on one distribution problem — loss-reducing reconfiguration across 8–12 feeders, delivered as draft switching plans into the ADMS. Contains no new model development.

Moving one stage takes about 90 days when scoped to a single decision on a single substation cluster, and multiple years when scoped to a service territory. The plan below runs the transition on a specific, common problem: distribution feeders operating on switch configurations set years ago, carrying avoidable technical losses and voltage excursions because the normally-open points no longer match today's load and DER pattern. Most stage-2 utilities already have the optimisation working on a study snapshot — so this quarter contains integration, validation and attribution, and no new model development at all.

Pilot → Operationalised on feeder reconfiguration, in one quarter

One substation cluster of 8–12 feeders, one named owner. If any phase needs more than its window, narrow the scope — fewer feeders — rather than extending the plan.

  1. Days 1–15

    Validate the model and baseline the cluster

    Choose a cluster of 8–12 feeders across two or three substations, with SCADA-visible flows and a workable switch population. Reconcile a power flow on the GIS/CIM model against SCADA and AMI telemetry; fix connectivity and phase errors and write them back. Baseline six months of technical losses, voltage excursions and switching activity. Name the distribution operations manager as owner — losses and violations are their numbers.

    A model that reproduces telemetry, a baseline, one named owner

  2. Days 16–45

    Run the optimiser and deliver draft switching plans

    Run the reconfiguration optimisation — ML screening over switch states, full AC certification with radiality, limit and protection flags — against the validated model. Deliver the top candidates as draft switching plans in the ADMS or the operations workflow the engineers already use, each with its modelled delta and its reversion attached. The engineer approves, amends or rejects every plan; nothing executes automatically.

    Certified draft plans arriving where engineers already work

  3. Days 46–70

    Instrument the pipeline and drill the reversion

    Freshness alerting on the model refresh and telemetry feeds; power-flow-vs-telemetry residual monitoring per feeder with a named person paged on drift; a quarantine rule that pulls any feeder whose residuals exceed threshold. Execute the first approved reconfigurations in maintenance windows — and exercise the reversion to the legacy configuration once, deliberately, on a quiet shift.

    Monitoring live, first configurations enacted, reversion drilled

  4. Days 71–90

    Attribute in megawatt-hours, not objective values

    Hold out a comparable neighbouring cluster on its legacy configuration. Report the difference in technical losses (MWh), voltage-violation hours and switching burden between optimised and holdout clusters — not the solver's modelled improvement. Close the quarter with the delta, the approval statistics, and the costed proposal for cluster two.

    A loss and violation delta a regulator-facing budget holder accepts

The order matters

  1. Validation before optimisation

    A merely good configuration computed on a validated model beats an optimal one computed on a wrong map. Every hour spent reconciling the model against telemetry converts directly into switching plans an engineer can trust — hours spent tightening the solver on an unvalidated model convert into nothing.

  2. Certification before recommendation

    No candidate reaches the ADMS without a full physics run against the current model — limits, radiality, protection flags recorded. The certification harness is the artefact operations is being asked to trust; the ML behind it can then evolve freely without renegotiating that trust.

  3. Approval before autonomy

    Keep the engineer's approval on every plan through the first quarters, even where closed-loop execution is technically trivial. The approval log — what was accepted, amended, rejected, and why — is the dataset from which any future operating envelope will be derived, one action class at a time.

Instrumenting the KPIs: formula, source, cadence

Where each core optimisation KPI actually comes from — the formula, the system that produces it, and the stage at which it first measures something real. All telemetry, no self-report.

A KPI you cannot name a source system for is an opinion. Every core metric below reduces to quantities the GIS, SCADA, AMI, ADMS or the optimisation pipeline itself already records — the instrumentation work is joining them, not creating them. Read your stage's rows, wire those first, and treat 'honest from' as a hard rule: reporting a KPI before its stage produces a number with nothing underneath it.

KPIFormula / readSourceCadenceHonest from
Model mismatchPower-flow vs telemetry residuals, per feederValidation pipeline + SCADA/AMIPer refreshStage 1
Study turnaroundRequest received → screening result issued, elapsed daysStudy tracker / connection queuePer studyStage 1
Certified-candidate rateCandidates passing full physics certification ÷ candidates screenedOptimisation pipeline logsPer runStage 2
Plan acceptance rateDraft plans approved (incl. amended) ÷ plans presentedADMS approval logWeeklyStage 3
Technical-loss deltaLosses (MWh) on optimised cluster vs holdout, weather-normalisedSCADA/AMI energy balancesMonthlyStage 3
Voltage-violation hoursHours outside statutory band, optimised vs holdoutAMI / power-quality monitorsMonthlyStage 3
Hosting capacity releasedMW connectable after optimisation vs before, per feederHosting-capacity enginePer refreshStage 3
Time-to-productionUse case agreed → results in production, elapsed weeksDelivery trackerPer use caseStage 3
Automation rateClosed-loop actions ÷ all actions in automated classesDecision logWeeklyStage 5
Escalation rateOut-of-envelope escalations ÷ automated actionsDecision logWeeklyStage 5
Instrumentation build sheet for the core grid-optimisation KPIs in a utility estate. 'Honest from' is the ladder stage at which the KPI first measures something real.

Two disciplines make the sheet trustworthy. First, pipeline KPIs and value KPIs are always reported as a pair — a plan-acceptance rate without a loss delta is theatre, and a loss delta without acceptance statistics is unexplainable. Second, every value KPI is measured against a holdout — comparable feeders left on legacy practice — and weather-normalised, because in a live network, temperature and DER output will otherwise claim the credit or take the blame.

Operationalised readiness checklist

If you cannot tick all seven, you are still at Pilot regardless of how good the optimiser is. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Tick honestly — the blank list is data too

Most pilot-stage utilities can genuinely tick one or two of these, not zero. If none apply yet, don't start with tooling: pick one substation cluster and run the 90-day plan above. Everything on this list falls out of doing that once.

Failure modes that send programmes backwards

Capability is not monotonic. Four regressions account for almost all of it, and none of them announces itself.

Grid-optimisation programmes regress quietly, because the conditions that made a stage safe stop holding while the output keeps arriving. Four failure modes account for almost all of it, and each has a cheap preventive measure that costs a fraction of the incident it forestalls.

Likelihood: highImpact: high

Optimising on a map that has drifted

The network changes — new DER clusters, reconductoring, switch operations left in a non-standard state — and the model does not. Recommendations stay plausible while quietly describing a network that no longer exists; the failure surfaces as a switching plan that cannot be executed as written, and trust does not survive the second such plan.

PreventionPer-feeder mismatch monitoring with a quarantine rule: drifting feeders drop out of the optimisation scope automatically until reconciled.

Likelihood: mediumImpact: high

The uncertified shortcut

Under schedule pressure, a team starts acting on surrogate output directly — the ML said the configuration is fine, the full power-flow run takes too long. It works until it doesn't: surrogates carry bounded average error and unbounded tail error, and the tail case in a switching context is a limit violation on a real feeder.

PreventionMake certification structural, not procedural: the delivery layer physically refuses candidates without a recorded physics run.

Likelihood: highImpact: medium

The study that never becomes a switching order

The pipeline works, the recommendations are good, and month after month the configurations stay where they were — because drafting, protection-checking and scheduling the change is nobody's job. The programme reads as delivering while changing nothing, and is defunded in the next budget cycle with its acceptance log as the evidence.

PreventionA named operations owner whose metric moves with enacted changes, and plan-acceptance statistics reviewed monthly at the level that funds the programme.

Likelihood: mediumImpact: high

The envelope outlives the network

Closed-loop bounds set on last year's network stay in force while DER growth shifts flows underneath them. Actions remain technically in-envelope while the envelope itself has stopped describing safe operation; the first incident then gets every loop opened, a two-stage regression from a single event.

PreventionEscalation-rate trend as a standing governance input, and envelope review triggered by network-change events — DER cohort connections, reconfiguration of adjacent clusters — not just by the calendar.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Optimal power flow (OPF)
The optimisation problem of choosing generator setpoints and network controls to minimise cost or losses subject to power-flow physics and operating limits. AC OPF is non-convex and expensive; ML surrogates and warm starts are used to accelerate it at operational scale.
Feeder reconfiguration
Changing the open/closed states of tie and sectionalising switches so a distribution network — built meshed but operated radially — carries load along lower-loss, better-voltage paths. Formulated as a mixed-integer programme with a radiality constraint.
Radiality constraint
The requirement that a distribution network remain radial — every load fed along exactly one path — after reconfiguration, preserving protection coordination. A configuration that violates radiality is inoperable regardless of its losses.
Hosting capacity
The amount of DER (in MW) a feeder can accommodate without violating voltage, thermal or protection limits. Computed by repeated power-flow sweeps; the basis of interconnection screening and, increasingly, of published hosting-capacity maps.
Volt/VAR optimisation (VVO)
Coordinated optimisation of voltage-control assets — transformer taps, capacitor banks, smart-inverter reactive power — to hold voltage bands and reduce energy through conservation voltage reduction.
ML surrogate
A learned model that approximates an expensive physics computation — a power flow, a wake simulation — thousands of times faster, with measured, bounded error. Used to screen large candidate spaces; never the final word on a candidate that could become an action.
Warm start
Initialising a physics solver from an ML-predicted solution rather than from scratch, cutting solve time while keeping the solver's guarantees. The standard pattern for making security-constrained OPF fast enough for operational cadence.
Physics certification
Re-running every shortlisted candidate through the full solver — AC power flow with limit, radiality and protection checks — before it can be delivered as a draft plan. The gate that makes ML screening safe to use in a switching context.
Model mismatch
The residual between what a power flow on the network model computes and what SCADA/AMI telemetry measures. The health metric of the network model itself: drifting residuals localise where the map and the network have diverged.
Micrositing
Optimising the positions of individual turbines within a wind farm — or blocks and strings within a solar plant — under wake, terrain and constraint models, typically to maximise energy yield or minimise LCOE for the site.
Dynamic line rating (DLR)
Setting transmission line limits from actual weather and conductor conditions rather than conservative static assumptions, releasing capacity from existing lines. The IEA estimates AI-enabled tools of this class could unlock up to 175 GW without new construction.
Operating envelope
The versioned, reviewed set of bounds — voltage bands, switching-frequency ceilings, excluded network states — inside which enumerated action classes may execute closed-loop. The stage-5 artefact, governed with the same rigour as protection settings.

Frequently asked questions

The questions utility engineers and planning leads ask most often when scoping grid and layout optimisation.

What is AI grid optimisation, in practical terms?

AI grid optimisation is the use of machine learning to make classical power-system optimisation — optimal power flow, feeder reconfiguration, Volt/VAR control, siting studies — fast and repeatable enough to run operationally. The physics solvers remain the source of truth; ML screens huge candidate spaces, warm-starts the solvers and learns from telemetry. The practical output is not a smarter algorithm but a pipeline: a validated network model, a certified shortlist, and draft plans arriving inside the ADMS where an engineer approves them.

Does machine learning replace power-flow solvers?

No — and any architecture that proposes it should be rejected. Surrogate models approximate power-flow outcomes with bounded average error but unbounded tail error, and the tail case in a switching context is a limit violation on a real feeder. The production pattern is screen-then-certify: ML evaluates millions of candidates cheaply, and every candidate that could become an action is re-run through the full AC power flow with limit, radiality and protection checks. Methods like DeepOPF are valuable precisely because they accelerate the solver's workflow rather than replacing its guarantees.

What data does a grid optimisation programme need first?

Three things, in order: a GIS/CIM network model validated against telemetry — a power flow on the model must reproduce what SCADA and AMI measure; load and DER data fresh enough for the decision cadence — AMI-derived load shapes and an authoritative DER register; and an operational baseline for the target KPI, typically six months of losses, violation hours or screening turnaround. Everything else — feature stores, surrogates, RL agents — is downstream of those three and worthless without them.

How accurate does the GIS model need to be?

Accurate enough that power-flow residuals against telemetry are small and stable on the feeders in scope — that is the operational definition, and it is testable in weeks. Perfection across the territory is neither achievable nor required: the working pattern is to validate cluster by cluster, write corrections back into the GIS under change control, and quarantine any feeder whose residuals drift. AMI voltage correlation finds phase errors statistically, which removes the most expensive class of field verification.

What is hosting-capacity analysis and why start there?

Hosting-capacity analysis computes how much DER each feeder can accept without violating voltage, thermal or protection limits — the number behind interconnection screening. It is a strong first use case because it monetises the network model twice: the validated model it forces you to build is the same asset every later optimisation consumes, and the screening backlog it clears is a KPI both regulators and connection customers feel. NREL's advanced hosting-capacity work documents the methodology most utility implementations adapt.

How long does a feeder-reconfiguration pilot take?

About 90 days to go from a working study to an operationalised loop on one substation cluster of 8–12 feeders — provided the optimisation already works on a snapshot and the quarter is spent on validation, ADMS delivery, monitoring and attribution rather than model development. The elapsed time is dominated by two things: the state of the GIS on the chosen cluster, and the change-approval path for delivering draft switching plans into the operations workflow. Scoping to a whole territory instead of one cluster is what turns 90 days into two years.

What savings does reconfiguration actually deliver?

It depends on how far the current configuration has drifted from today's load and DER pattern, which is why the honest answer is measured, not quoted. The method is the number: baseline six months of technical losses and violation hours on the target cluster, hold out a comparable neighbouring cluster on its legacy configuration, and report the weather-normalised delta. Networks whose normally-open points have not been revisited in years tend to carry the most recoverable loss — and the same pipeline then keeps the configuration current as the network changes.

Where do wind micrositing and solar layout fit in the same programme?

They are the design-side family of the same discipline: constrained optimisation over a spatial model, screened by ML and certified by physics — wake and yield models instead of power flow. Micrositing and cable-routing decisions are made once per project rather than weekly, so they suit the study stage, but they should consume the same platform: the same data engineering, the same certification discipline, the same delivery into a planner's workbench. DeepMind's reported ~20% boost in wind-energy value shows the pattern's other half — the gain arrived when predictions were wired into commitment decisions.

Can reconfiguration and VVO run closed-loop safely?

Some action classes can, once the evidence exists. The path runs through the approval log: quarters of engineer decisions on draft plans establish which action classes are consistently approved unchanged, and those classes — typically VVO setpoints, then routine reconfiguration on designated clusters — graduate into a versioned operating envelope with escalation monitoring and a drilled reversion. Anything touching protection philosophy or high-consequence customers stays human-approved indefinitely. Autonomy is earned per action class from its own history, never inherited.

How does this interact with regulatory and audit obligations?

Favourably, if the pipeline is built with a decision log from the start. Network operators answer to regulators for losses, voltage quality and connection timelines, and reliability obligations require operational decisions to be explainable after the fact. A stage-3 pipeline produces exactly that evidence as a by-product: which model version, which certification run, who approved which plan, when. No framework we have encountered prohibits optimisation in the loop — they require that decisions be reconstructable and reversible, which is a maturity property, not a model property.

What team does the first quarter actually need?

Four roles, some part-time: a power-systems engineer who owns model validation and certification; a data or integration engineer who builds the refresh, delivery and monitoring path; an optimisation or ML engineer who runs the screening and solver stack; and — decisively — a named operations owner whose metric moves when configurations are enacted. Programmes fail more often from the missing fourth role than from any technical gap: without an operations owner, certified plans accumulate in a queue nobody is paid to empty.

Should we buy an ADMS optimisation module or build the pipeline?

Usually both, with the boundary drawn at the model layer. ADMS vendors ship capable VVO and FLISR modules, and using them is faster than rebuilding solved problems — but they inherit your network model's quality, and the validation, write-back and mismatch-monitoring layer is almost never part of the product. Build the model layer as your own asset, buy solvers and modules as consumers of it, and hold every component — bought or built — to the same certification contract. What you must own is the map and the evidence, not necessarily the algorithms.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — grid optimisation, forecasting, vision inspection and decision support running against live operational data, integrated into the GIS, ADMS and SCADA layer rather than delivered as dashboards.

  • · Production deployments across network operations, renewables and asset management
  • · Solver + ML hybrid systems: physics-certified optimisation, not black-box output
  • · Integration-first delivery: ADMS/SCADA write-back, monitoring, rollback
  • · 11 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyElectricity Grids and Secure Energy Transitions (opens in a new tab)
  3. EPRIGrid analytics and DER integration research (opens in a new tab)
  4. US Department of EnergyOffice of Electricity (opens in a new tab)
  5. PJM InterconnectionPJM, Google & Tapestry join forces to apply AI to regional planning and generation interconnection (opens in a new tab)
  6. PJMPJM Interconnection (opens in a new tab)
  7. CAISOCalifornia ISO (opens in a new tab)
  8. arXivDeepOPF: A Deep Neural Network Approach for Security-Constrained DC Optimal Power Flow (opens in a new tab)
  9. arXivLearning to run a power network challenge for training topology controllers (opens in a new tab)
  10. EnelEnel — company publications (opens in a new tab)
  11. E.ONE.ON — company communications (opens in a new tab)

Find out exactly where you are — then what to do about it

We run the assessment with your planning and operations leads, benchmark the result against comparable utilities, and leave you with a costed 90-day plan for your weakest dimension. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.