Redefining Technology

Energy & UtilitiesAI Adoption & Maturity Curve

AI adoption success factors in energy and utilities: what separates programmes that stick from pilots that stall

AI adoption success factors in energy and utilities are the conditions that decide whether an AI system reaches an operational decision and stays there. Seven recur: a named decision, an operations owner whose target moves, a baseline pulled before the build, a route into the ADMS, a drilled fallback, attribution someone will sign, and reuse by the second decision.

Generated scene: utility control room and field operations with AI decision-support overlays across grid displays
Energy & Utilities · AI Adoption & Maturity Curve

Key takeaways

  1. The strongest predictor of a successful utility AI programme is which decision it chose to attack, and that choice is fully verifiable before any money is spent. Score the candidate on system-of-record control, decision frequency, feedback latency, consequence, baseline availability, owner and benefit legibility — seven criteria, twenty-one points, one afternoon.
  2. Sponsorship is not who attended the kick-off; it is whose targets move. A use case whose only owner sits in an innovation or digital budget has no operational advocate when peak season, a storm or a reorganisation competes for attention, and it is the first thing dropped.
  3. The baseline is a one-way door. Five years of the outcome metric — customer minutes lost, deferred replacements, cost to serve — must be pulled from the system of record and its counterfactual agreed before the model runs, because a baseline assembled afterwards inherits the pilot's own effects and can never cleanly attribute them.
  4. In a regulated utility, a benefit only becomes permanent when it is legible to the body that sets allowed revenue. Express the value in units the price control or rate case already recognises — SAIDI minutes, TOTEX deferral, connections delivered — or the second programme will be funded from goodwill rather than evidence.
  5. Success compounds only when the second decision costs less than the first. If use case two rebuilds its own data path, monitoring and approval route, the programme is a series of successful projects rather than a capability, and its velocity falls as its maintenance load rises.

Abbreviations used on this page

SCADA
Supervisory control and data acquisition
EMS
Energy management system (transmission control)
ADMS
Advanced distribution management system
OMS
Outage management system
AMI
Advanced metering infrastructure (smart meters and head-end)
GIS
Geographic information system — the network model of record
CIS
Customer information system (billing and accounts)
WFM
Workforce management — field crew scheduling and dispatch
DER
Distributed energy resources (rooftop solar, batteries, EVs)
SAIDI
System average interruption duration index
RIIO
Ofgem's price-control framework: Revenue = Incentives + Innovation + Outputs
TOTEX
Total expenditure — capital and operating spend treated as one pot

Free · 8 questions · ~3 minutes

Score your preconditions, not your ambitions

Eight questions, one at a time, about three minutes — each scoring a precondition that either exists or does not. Answer them and we build your personalised report: your stage on the curve, your score on each of the four precondition dimensions, and the one missing factor most likely to stall your next use case. It lands in your inbox.

0 of 8 answered

Question 1 of 8Candidate selection

How was your current or next AI use case chosen?

Selection is the highest-leverage decision on the whole programme, and it is made before any money is spent.

How the score maps to a stage
  • 04 — Stage 1, Interested. AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
  • 510 — Stage 2, Committed. One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
  • 1116 — Stage 3, Operating. The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
  • 1721 — Stage 4, Attributed. The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
  • 2224 — Stage 5, Institutional. AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.

What AI adoption success factors are in energy and utilities

A definition, the seven factors that recur, and the mechanism that connects them — from a decision nobody has chosen yet to a capability nobody has to defend.

AI adoption success factors in energy and utilities are the conditions that determine whether an AI system reaches an operational decision and stays there. They are not qualities of the model, the vendor or the budget. They are properties of the decision you chose, the person who owns it, the system it writes into and the evidence you can produce about what changed — and every one of them can be checked before a build begins, which is what makes them useful rather than retrospective.

The distinction that organises this page is between factors and outcomes. "Executive buy-in", "a data-driven culture" and "the right talent" are outcomes: real, but not actionable on a Monday morning, and impossible to verify before funding. The seven factors below are verifiable. Each one is a yes or a no about a specific artefact — a scored candidate, a named objective, a pulled baseline, a field in the ADMS or works-management system, a dated drill log, a signed benefit statement, a reuse record — and a programme missing any of them fails in a predictable way at a predictable point.

  • 1 · A named decision, not a named technology

    The unit of work is one operational decision somebody makes on a schedule — which feeders to pre-stage crews on, which transformers to advance in the replacement plan, which contacts route to an assistant — not "AI for asset management". A decision has a frequency, an owner, a system of record and a metric. A domain has none of those, so nothing about it can be baselined or finished.

  • 2 · An operations owner whose target moves

    Sponsorship is measured by whose objectives change, not by who attended the kick-off. If the only owner sits in an innovation or digital budget, the use case has no advocate during a storm week, a regulatory submission or a reorganisation — and those are exactly the moments that decide whether it survives.

  • 3 · A baseline pulled before the build

    Five years of the outcome metric, at the granularity of the decision, extracted from the system of record with its definition written down and agreed. This is a one-way door: a baseline assembled after go-live contains the pilot's own effects, and no amount of later analysis removes them.

  • 4 · A route into the operational workflow

    The output appears as a field in the ADMS, OMS, GIS, works-management or customer system the decision-maker already has open — advisory is fine, absent is fatal. A portal behind a separate login shifts the burden of action onto a person whose existing process already works, and voluntary steps are the first thing dropped under pressure.

  • 5 · A live fallback the operation has drilled

    The previous method — the condition rule, the storm matrix, the seasonal profile, the manual triage — kept alive, one switch away, and exercised on a date somebody logged. This is not engineering pessimism; it is the political key that unlocks the change board, because operations will accept a new decision source they can instantly revert.

  • 6 · Attribution someone will sign

    The benefit read against a matched control population or a counterfactual agreed in writing before go-live, expressed in a unit the price control or rate case already recognises, and signed by finance. This is the factor most programmes discover they needed a year too late.

  • 7 · A second decision that reuses the first

    Feeds, freshness alerting, the write-back pattern, the override log, the drill calendar and the benefit ledger carry over, so decision two costs meaningfully less than decision one. Without reuse a programme is a sequence of successful projects whose velocity falls as its maintenance load rises.

How a chosen decision becomes a permanent capability

The mechanism, in three lanes. The top lane is the one most programmes skip entirely — they start at delivery, with a decision a vendor demonstration chose for them. The bottom lane is where capabilities quietly die: not from a failure, but from a benefit nobody attributed and an owner nobody replaced.

  • Data & feeds
  • Human in the loop
  • Where value leaks
  • AI / model
  • System-of-record action

The process, in words

  • Selection happens before any money moves. A standing register of candidate decisions is scored on seven criteria, the winner gets a named operations owner with an agreed target, and five years of the outcome metric are pulled from GIS, OMS or works management with the counterfactual written down. Skip this lane and the decision arrives with the vendor demonstration that suggested it.
  • Delivery is the first ninety days. Operational feeds run with freshness alerting, the model is served inside the cycle the decision actually runs on, and its output is written into the ADMS, OMS or works-management field the planner already reads. The operator acts and every override is logged with a reason code, while the previous method stays one switch away and the switch is drilled on a date somebody records.
  • Permanence is everything after go-live. The benefit is read against the matched control population held back at selection, the statement is signed by finance and regulation in units the business plan recognises, and the second decision inherits the feeds, the write-back pattern and the ledger. When the owner rotates and nobody inherits the benefit, the capability keeps running and stops counting.
Step-by-step insights
The candidate register — the artefact almost nobody has
Most utilities can produce a list of AI projects and cannot produce a list of decisions worth attacking. The two are very different objects: a project list records what was funded, a candidate register records what is available, including the decisions nobody has proposed because no vendor sells them. The register is cheap — a spreadsheet of decisions with frequency, owner, system of record and outcome metric — and it changes the sourcing of ideas from vendor roadmaps to the operation itself. Keeping it open and re-scored every planning cycle is the single practice that most reliably prevents a programme plateauing after its obvious candidates are used up.
Scoring before funding, and why an afternoon is enough
The seven selection criteria are all answerable from things a utility already knows: who owns the change board on the target system, how often the decision is made, how long before you learn whether it was right, what a wrong output costs, whether the outcome metric exists at the decision's granularity, whose objectives contain it, and whether the benefit is expressible in a unit the business plan uses. None of that needs a data-science assessment or a vendor engagement. The reason to insist on scoring is not the score itself but the conversations it forces — most disqualifying facts about a candidate surface in the first twenty minutes and would otherwise have surfaced in month nine.
The baseline is a one-way door
Everything else on this diagram can be retrofitted. A write-back can be added later, an owner can be reassigned, a drill can be scheduled next month. The baseline cannot. Once the model is influencing the operation, every subsequent extract of the outcome metric contains its effects, and separating them requires assumptions that reviewers are entitled to reject. Pull five years of the metric at the decision's granularity — per feeder, per asset, per job, per contact type — before anything runs, write the definition down, and store it somewhere a system migration will not quietly delete. A fortnight at the start buys a defensible number for the life of the capability.
Write-back is an approval problem wearing an engineering costume
Teams describe the route into operations as an integration task and estimate it in developer weeks. In a utility it is mostly an approval task: the ADMS, OMS and works-management systems are change-controlled, often vendor-supported, and sometimes inside a security boundary that treats a new write path as a design change. The variable that actually sets the timeline is who owns change on that system and how recently you shipped anything to it. That is why it is a selection criterion rather than a delivery detail — a decision whose system of record belongs to somebody else can be an excellent decision and still be the wrong first one.
The override log is the most under-used dataset in the sector
When operations accepts an advisory recommendation, or rejects it, and records why, the utility is generating something no vendor can supply: labelled evidence about how its own experts reason about its own network. That log tells you where the model is systematically wrong, which reason codes cluster on particular circuits or asset classes, whether trust is calibrated, and — much later — where the bounds of any automation could safely sit. Programmes that skip the advisory stage to save three months arrive at the automation conversation with nothing to set thresholds from except the vendor's defaults.
Why the bottom-right node is a risk node, not a failure
The most common ending for a utility AI capability is not an incident. It is a reorganisation. The model keeps producing output, operations keeps using it, and the person who knew what it was worth moves to another business unit. Two years later nobody can reconstruct the baseline, the control population has been switched over, and the benefit disappears from the plan — so when the next investment case is written, the evidence that would have funded it no longer exists. Naming a successor for the benefit, not just for the code, is a five-minute act that protects years of work.

The five stages, named for what a programme has proved

Interested, Committed, Operating, Attributed, Institutional. For each: what it looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.

The five stages below are proof states rather than capability levels — each is named for what the programme has actually earned, not for what it can build. That framing matters because a utility can hold impressive capability and still sit at stage 2: an excellent model, a real owner and a convincing back-test prove nothing about the operation until something in the operation is different. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Value released across the five proof states

The shape is not linear, and the flat part is longer than most plans assume. Value stays close to zero through Interested and Committed — where a majority of operators sit — and inflects at Operating, when the output first reaches the person who acts. The second inflection is financial rather than technical: at Attributed the value becomes defensible, which is what makes it survive a planning cycle.

Operational value released by stage

  • Stage 1 · Interested — 21% of operators. AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
  • Stage 2 · Committed — 34% of operators. One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
  • Stage 3 · Operating — 27% of operators. The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
  • Stage 4 · Attributed — 13% of operators. The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
  • Stage 5 · Institutional — 5% of operators. AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with IEA analysis of digitalisation in energy.

One reading discipline before the panels: score each live use case separately, then take the organisation's stage to be where the shared foundations sit rather than where the best individual example does. A single Attributed forecasting capability alongside four Interested experiments is a Committed organisation, because the next use case will inherit the foundations, not the exception.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Interested

21% of operators sit here

AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.

Stage 1 is not scepticism — most utilities at this stage are enthusiastic. What is missing is a unit of work. The programme is described in terms of a technology and a domain ("AI for asset management", "machine learning in the control room") rather than in terms of a decision somebody makes on a Tuesday morning with a deadline. Because there is no decision, there is no owner, no metric, no system of record and no baseline, which means none of the seven success factors is present or even checkable.

The tell is the language in the steering pack. Stage-1 material talks about capability, platforms and partnerships; stage-2 material talks about a feeder set, an asset class, a call type. Ask what will be different in the operation ninety days after go-live and a stage-1 answer describes an insight; a stage-2 answer describes an action taken differently by a named role.

This is a cheap stage to leave, and expensive to linger in for a reason that is not obvious. Every quarter spent here trains the organisation to treat AI as a discussion topic. When a real candidate finally arrives, it enters an environment where nobody has ever pulled a baseline, agreed a counterfactual or approved a write-back — so the first real use case pays the cost of building all of that, and gets judged on a timeline set by the demos.

In practice

The strategy that named a technology, not a decision

A distribution business published an AI strategy with five workstreams — asset intelligence, network optimisation, customer experience, workforce productivity, data foundations. Eighteen months later it had spent real money and could not name a single decision that was made differently because of it. The workstreams were all defensible; none of them was a decision with an owner, a frequency and a metric, so none could be baselined and none could be finished.

What it looks like

  • AI appears in the corporate strategy but not against a named decision
  • Vendor demos and conference visits are the main source of candidate ideas
  • Funding sits in an innovation or digital line, not an operational budget
  • No baseline has been pulled from any system of record

Diagnostic signals you can check this week

  • Ask for the candidate list: if it names technologies or domains rather than decisions with owners, you are here
  • Ask which system of record the first use case will write into — a hesitation is the stage-1 signature
  • Check whether anyone has pulled five years of any outcome metric from GIS, OMS or the works-management system
  • Look at where the budget line sits: innovation and digital lines fund experiments, operational lines fund changes

Anti-pattern · Running a discovery programme instead of choosing

The instinctive response to an empty candidate list is a discovery exercise — workshops across every function, a longlist of eighty ideas, a heat map. It feels rigorous and it postpones the only decision that matters. Eighty unscored ideas are less useful than three scored ones, because scoring forces the questions that actually predict success: who owns the decision, where does it execute, and is there a baseline. Run the scorecard on three candidates in an afternoon and fund the winner; the longlist can wait for the second planning cycle, when you will know far more about what your estate can absorb.

What holds you here

There is no named decision, so there is nothing to own, baseline, integrate or attribute — the success factors have no subject.

Highest-leverage next move

Name one operational decision, score it against the seven selection criteria, and give it an operations owner before any build is scoped.

Cost of leaving

Effort
2–6 weeks
Team
One operations manager and one analyst, part-time
Risk
Low — selection work, nothing in production changes
To next stage
1–2 months

If this is you, the next step is

A short session: we run the selection scorecard on your three strongest candidates and hand you the ranking.

Score three candidate decisions

Stage 2

Committed

34% of operators sit here

One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.

Stage 2 is where most of the sector is, and it is the most misleading place to be, because every visible indicator is green. The decision is named. The owner exists. The back-test is convincing — the storm model would have pre-staged crews better on the last three events, the health index would have caught two of the four transformers that failed. The steering group is pleased and the P&L is untouched, because a back-test is an argument about the past and an operation runs on the present.

The structural reason is that a proof of value is optimised to answer "does this work?" while the operation's question is "will the person on shift act on it?". Those have different builds. The fastest route to a demonstrable result is an analytics portal, and a portal is precisely the artefact that leaves the work instruction unchanged. Utilities compound this because their operational systems — ADMS, OMS, works management — sit behind change boards with queues measured in months, so the portal is not only faster but also the only thing the team can ship inside the proof-of-value window.

Time here is not neutral. Control-room and field staff learn that AI output is optional; asset planners learn that the health index is a second opinion; finance learns that AI produces slide decks rather than deferred capital. The next proposal is funded against that memory. A utility that has sat at stage 2 for three years is usually harder to move than one at stage 1, because the organisational antibodies are established and the candidate list has been picked over.

In practice

The health index that never reached the plan

A network operator built a transformer health index from dissolved-gas history, loading and age, and back-tested it against five years of failures. It ranked the population better than the existing condition rule and the asset-strategy team agreed. It shipped as a quarterly spreadsheet. Two replacement plans later, the plan was still built from the condition rule, because that rule was embedded in the works-management workflow and the spreadsheet was not. Nobody rejected the model; the plan simply had no field to put it in.

What it looks like

  • A specific decision is named, with an operations owner and a target metric
  • A baseline has been pulled from the system of record and its definition agreed
  • A model demonstrably beats the incumbent method on back-test
  • No standard operating procedure or work instruction has changed

Diagnostic signals you can check this week

  • Count the work instructions or standard operating procedures that changed because of the pilot — usually zero
  • Ask the person who makes the decision to show you where the model output appears in their normal day
  • Check whether the baseline was pulled before the model was built, and whether its definition is written down
  • Ask what happens to the pilot's output during a storm week or a peak period; if the honest answer is 'nobody looks', the route into operations is the gap

Anti-pattern · Improving accuracy to earn adoption

When a proof of value is not adopted, the reflex is to make it more accurate on the theory that trust follows precision. It rarely does in a utility. Adoption is a function of where the output appears and whether the previous method is still available if it is wrong: a moderately accurate ranking inside the works-management workflow changes more plans than an excellent one in a portal behind another login. Spend the next quarter on the write-back and the fallback drill, then revisit accuracy when you can price an accuracy point in deferred replacements or customer minutes lost.

What holds you here

The output has no home in an operational system, so acting on it depends on someone choosing to add a step to their day.

Highest-leverage next move

Write the output into the ADMS, OMS, GIS or works-management field the decision-maker already reads, keep the previous method one switch away, and drill that switch.

Cost of leaving

Effort
3–6 months
Team
One integration engineer, one ML engineer, a named operations owner, change-board sponsor
Risk
Medium — the first write into an operational system needs a rehearsed revert and change-board approval
To next stage
3–6 months

If this is you, the next step is

The stage 2→3 transition is our most common engagement — typically 90 days on one decision.

Get the pilot into the ADMS or works management

Stage 3

Operating

27% of operators sit here

The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.

Stage 3 is the first stage where the operation is genuinely different. A crew is pre-staged somewhere it would not have been, a transformer is advanced or deferred, a customer contact is answered in a way it would not have been. The character of the work changes with it: stage-1 and stage-2 problems are analytical, stage-3 problems are operational, and the disciplines that solve them — freshness alerting, drilled reversion, on-call ownership, change control — come from running systems rather than from building models.

What is missing at stage 3 is almost never engineering. It is evidence. Ask what the capability is worth and the honest stage-3 answer is a story: the team believes storm pre-staging improved, the asset planners find the ranking useful, complaints about estimated restoration times seem lower. All of that may be true and none of it survives a budget cycle, because a live distribution network changes for a dozen reasons at once — weather, DER growth, a reconfiguration, a vegetation programme, a new crew contract — and any of them will happily claim the credit or take the blame.

This is where the baseline pulled at stage 2 earns its keep, and where its absence is fatal in a way that cannot be repaired later. Attribution requires a comparison the operation agreed to in advance: a matched set of feeders, depots, asset cohorts or customer segments left on the previous method, or a documented counterfactual that finance accepted before the model ran. Retrofitting one after go-live is not a smaller version of the same job; it is a different and much weaker claim, and everybody in the room knows it.

In practice

The pre-staging that everyone believed in

A utility ran storm crew pre-staging from a model wired into its OMS for two seasons. Field leadership was convinced it helped, and it very likely did. When the capital review asked what it was worth, the team could produce restoration times for the two seasons — which were also the two seasons a new mutual-aid agreement and a large vegetation programme landed. No comparable districts had been held on the old matrix. The capability survived on sponsorship rather than evidence, and the second use case was funded a year late.

What it looks like

  • Output lands in an ADMS, OMS, GIS or works-management field, at least advisory
  • Overrides are logged with a reason code, not just accepted or ignored
  • A fallback to the previous method has been exercised and dated
  • The work instruction has been rewritten to reference the new field

Diagnostic signals you can check this week

  • Ask whether any part of the network, asset population or customer base is deliberately still on the previous method — if not, attribution is already compromised
  • Pull the override log: both a near-zero and a near-total override rate are alarms, and an absent log is worse than either
  • Ask for the date of the last fallback drill; an assurance without a log entry means the switch is theoretical
  • Ask finance what number they would put in a submission today, and watch whether the answer is a range, a story or a refusal

Anti-pattern · Declaring the benefit from the model's own metrics

With the output live and operations content, the temptation is to convert model performance into money directly: the ranking is X% better, replacements cost Y, therefore the benefit is X times Y. Every utility that has tried this has met the same reviewer question — compared with what? — and had no answer. The conversion is not wrong in principle; it is unfalsifiable in practice, and unfalsifiable numbers are removed from submissions rather than argued with. Hold a matched population, or agree a written counterfactual with finance before go-live. It costs a fortnight at the start and is unbuyable afterwards.

What holds you here

The value is asserted rather than attributed, so the capability is funded by sponsorship and falls with the sponsor.

Highest-leverage next move

Hold a matched control population or agree a written counterfactual, read the benefit against it, and get the statement signed by finance and regulation.

Cost of leaving

Effort
6–12 months
Team
Operations owner, data engineer, finance analyst, regulatory or business-plan contact
Risk
Medium — the control population must be genuinely comparable and genuinely left alone
To next stage
6–12 months

If this is you, the next step is

We design the control population or counterfactual with your finance and regulation leads, and leave you the method note.

Design the attribution before go-live

Stage 4

Attributed

13% of operators sit here

The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.

At stage 4 the argument stops being about AI. The conversation in the investment committee is about which decisions are on the roadmap and what each is worth, in the same units as every other proposal — customer minutes lost, TOTEX deferral, connections delivered, cost to serve. That is a healthier and much duller conversation, and its dullness is the point: capability that has to be re-sold every year is not capability, it is a campaign.

The engineering signature of this stage is that the marginal cost of the next decision has fallen. Feeds, freshness alerting, the write-back pattern into the ADMS or works-management system, the override log, the drill calendar and the benefit ledger are shared, so a new decision is mostly specification: which decision, which owner, which metric, which control population. When the third use case takes a third of the elapsed time of the first with the same team, the reuse factor is real. When it takes the same, the programme has built three projects and called it a platform.

The constraint that emerges here is not technical and not financial — it is the candidate pipeline. A programme that has proved it can convert a scored decision into a signed benefit will exhaust its obvious candidates within a couple of planning cycles, and the quality of the next tranche depends entirely on whether the register is being refreshed by the people closest to the work. Utilities that keep the register open, scored and visible keep compounding; utilities that treat selection as a one-off exercise plateau with a small portfolio of well-run models and a quietly ageing roadmap.

In practice

The third decision that cost a third

An operator that had wired outage prediction into its OMS added vegetation risk scoring and then LV load-at-risk ranking. The first took nine months, most of it spent proving the write-back pattern and negotiating the change board. The second took four. The third took eleven weeks, of which seven were spent agreeing the control population and the benefit unit with finance — the engineering was under three weeks. That ratio, specification-heavy and build-light, is the stage-4 signature.

What it looks like

  • A signed benefit statement exists in units the price control or rate case recognises
  • The second decision reused the first one's feeds, monitoring and approval route
  • Time from decision agreed to decision served is falling, use case by use case
  • The candidate register is scored and re-ranked on a planning cadence

Diagnostic signals you can check this week

  • Compare elapsed time from decision agreed to decision served across your last three use cases; if it is flat, nothing is being reused
  • Ask whether a single benefit ledger holds every live use case's baseline, control design and signed statement
  • Ask an operations leader — not an engineer — what each live model is worth; a confident answer in operational units is the stage-4 tell
  • Check when the candidate register was last re-scored and who contributed to it

Anti-pattern · Industrialising delivery while selection goes stale

Stage 4 makes delivery efficient, which makes it tempting to fill the pipeline with whatever is nearest — usually more variants of the decision you already solved, because the pattern is proven and the change board is friendly. The portfolio grows and the marginal benefit per use case falls, quietly, for two years. The discipline is to re-run selection against the whole register every planning cycle and accept that the highest-scoring candidate may sit in a function you have never worked with, needing a change board you have not met. Efficiency in delivery is not a reason to stop choosing well.

What holds you here

Success now depends on the individuals and sponsors who built it — nothing yet makes the capability survive a reorganisation, a vendor change or a price-control reset.

Highest-leverage next move

Move the benefits, the drill calendar and the candidate register onto standing operational and business-plan processes, so they have owners after the founders move on.

Cost of leaving

Effort
12–24 months
Team
Platform engineers, operations product owner, finance and regulatory partners
Risk
Medium — the shared layer competes with new-use-case demand for the same people
To next stage
18+ months

If this is you, the next step is

We map what your last three use cases actually shared, and what a fourth would still have to rebuild.

Measure your reuse factor

Stage 5

Institutional

5% of operators sit here

AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.

Stage 5 in a utility is not an autonomous grid; it is an unremarkable one. The distinguishing property is that nothing about the capability is special-cased. A model review sits in the same calendar as a relay test. A benefit statement is produced by the same process that produces every other business-plan output. Ownership transfer is a line on the leaver checklist beside system credentials. When AI stops needing its own governance forum, its own budget line and its own champion, it has become part of how the utility runs.

The risk that defines this stage is renewal rather than failure. Price controls reset, vendors are replaced, control rooms are merged, asset strategies change, and every one of those events silently invalidates something: a baseline definition, a control population, a data path, a drill owner. Institutional programmes survive them because the artefacts are on standing processes that force review — the business-plan cycle, the works-management release cycle, the operational calendar — rather than in a wiki that only the original team knows about.

The honest thing to say about stage 5 is that very few operators are here, and that arriving is less impressive than staying. The utilities that hold it treat the candidate register, the benefit ledger and the drill calendar as living operational artefacts with named owners and review dates, and they can produce all three inside an afternoon for an auditor, a regulator or an incoming executive. That reproducibility, not any particular model, is what makes the capability institutional.

In practice

The capability that survived a merger

Two control-room organisations merged and a large share of the combined analytics team left within a year. The forecasting and outage capabilities kept running, because each had a named operational owner in the receiving organisation, a runbook in the same repository as the rest of the operational documentation, benefits already booked in the business plan, and drill dates already in the operational calendar. Nothing heroic happened. Nobody had to rediscover what any of it was for, which is the whole test.

What it looks like

  • Benefits appear in the business plan or rate filing with a documented method
  • Fallback drills and model reviews sit on the same calendar as protection and relay tests
  • Ownership transfer is a defined step in the leaver and reorganisation process
  • The candidate register is refreshed by operations, not by the AI team

Diagnostic signals you can check this week

  • Ask for the benefit ledger, the candidate register and the drill calendar; if producing them takes more than an afternoon, they are not institutional
  • Check whether ownership transfer for a live model is a named step in the leaver and reorganisation process
  • Ask what happened to a live use case at the last reorganisation or vendor change, and who noticed
  • Ask whether an incoming operations leader could find out what each model is worth without asking the AI team

Anti-pattern · Assuming institutional means finished

Once the artefacts exist and the capability has survived a reorganisation, programmes stop investing in the parts nobody is currently complaining about: the register goes unrefreshed, drills slip past their dates, baselines quietly age out through a system migration. Decay is invisible for one planning cycle and expensive in the second, because the evidence chain breaks retrospectively — you discover the baseline is gone at the moment you need it. Put review dates on the artefacts themselves and treat a missed drill or a stale register with the seriousness of a missed statutory inspection.

What holds you here

Sustaining the capability is a renewal discipline — the constraint becomes keeping baselines, registers and drills current through resets, migrations and reorganisations.

Highest-leverage next move

Put explicit review dates and named successors on the benefit ledger, the candidate register and the drill calendar, and audit them on the business-plan cycle.

Cost of leaving

Effort
Continuous
Team
Operations owners, platform team, finance and regulatory partners on standing cycles
Risk
Low frequency, high consequence — decay is silent and only surfaces when evidence is demanded

If this is you, the next step is

We stress-test the artefacts against a real handover scenario and tell you what would not have survived.

Audit your capability against a reorganisation

Where energy and utilities operators actually sit today

The distribution across the five proof states, why the Committed crowd is so large, and what the external research says about the pressure now pushing against it.

Most energy and utilities operators are Committed rather than Operating: a decision has been chosen, a model beats the incumbent method on history, and nothing in the operation has changed. The crowding is not a failure of ambition. It is the predictable result of two structural facts — that the fastest route to a demonstrable result is an analytics portal, and that the operational systems where utility decisions actually execute sit behind change boards with queues measured in months. The proof of value can be finished inside its own window; the write-back cannot.

Distribution of energy and utilities operators across the five proof states

Illustrative distribution, synthesised from IEA, EPRI and Eurelectric adoption research — charted to show shape, not to report a survey. Committed is the mode. The drop from Committed to Operating is the largest transition loss on the curve, and it is an integration and approval gap rather than a modelling one.

Share of operators (illustrative)

  • 21% — 1 · Interested
  • 34% — 2 · Committed (the crowd)
  • 27% — 3 · Operating
  • 13% — 4 · Attributed
  • 5% — 5 · Institutional

Source: Illustrative; synthesised from IEA, EPRI and Eurelectric research

The pressure on the Committed crowd is rising from both directions. On the demand side, the IEA's Energy and AI report (opens in a new tab) records data centres consuming around 415 TWh in 2024 — about 1.5% of world electricity — and projects that more than doubling to roughly 945 TWh by 2030, with around one-tenth of global electricity demand growth to 2030 coming from that single sector. On the supply side the same report finds that AI tools could unlock up to 175 GW of transmission capacity on existing networks. Both numbers describe decisions — where to connect, what to reinforce, how far to push a circuit — that a utility currently makes with deterministic rules and engineering judgement.

The governance side has settled enough to stop being an excuse. NIST's AI Risk Management Framework (opens in a new tab) and ISO/IEC 42001 (opens in a new tab) give operators a management-system frame that predates any single regulator's position, Ofgem (opens in a new tab) and NERC (opens in a new tab) both ask for evidence rather than abstention, and sector bodies including Eurelectric (opens in a new tab), EPRI (opens in a new tab) and ENTSO-E (opens in a new tab) publish on adoption openly. Cross-industry research such as McKinsey's electric power and natural gas insights (opens in a new tab) continues to find experimentation running well ahead of impact. What separates the operators moving from those standing still is not access to any of this; it is whether the seven factors were engineered before the build.

The selection scorecard: score a decision before you fund it

The centrepiece. Seven criteria, twenty-one points, one afternoon — the instrument that decides more about a utility AI programme's outcome than anything that happens after funding.

Score a candidate decision on seven criteria before you fund it, because selection explains more variance in outcomes than delivery does. Two utilities with identical engineering capability, identical budgets and identical vendors will get different results if one attacks a decision made four hundred times a week in a system it controls with five years of history behind it, and the other attacks a decision made twice a year in a partner's platform with no comparable record. The scorecard below makes that difference visible in an afternoon rather than in month nine.

This is a different instrument from the maturity assessment further up the page. The assessment scores your organisation; the scorecard scores one candidate decision. A mature organisation can still choose badly, and an Interested organisation that picks well will out-deliver a Committed one that picks by demonstration. Run the scorecard on at least three candidates so the scores are comparative — a single score in isolation tells you almost nothing.

CriterionWhat you are actually testing0 — disqualifying3 — idealThe afternoon check
System-of-record controlWhether you can change the system where the decision executes, and how fastIt executes in a vendor-hosted or customer-facing system you cannot changeYour own ADMS, OMS, GIS or works-management system, and you shipped a change to it this quarterAsk who chairs that system's change board and when it last approved a new field
Decision frequencyHow much learning and how much value the decision can accumulate per quarterA handful of instances a year — an annual plan or a one-off studyHundreds or thousands of instances a week, per feeder, asset, job or contactCount the decision's instances in the system of record over the last 90 days
Feedback latencyHow long before you learn whether a given decision was rightYears — asset replacement outcomes, long-horizon investment choicesHours to days — crew routing, switching option ranking, contact triageAsk when the outcome field for that decision is populated, and by what process
Consequence and reversibilityWhat a wrong output costs, and how quickly it can be undoneSafety exposure or a supply interruption that cannot be unwoundContained and reversible within a shift, with the previous method availableWalk one plausible wrong output through to its worst realistic outcome
Baseline availabilityWhether the outcome metric exists in a system, at the decision's granularity, for long enoughNo historical record, or only at aggregate levelThree to five years, per decision instance, in a system you can query todayTry to pull twelve months of it yourself before the meeting ends
Owner with a moving targetWhether a named operations manager's objectives change if this worksOnly an innovation, digital or data team is accountableA named operations manager with the metric written into their objectivesRead the objectives; if the metric is not in them, the owner is nominal
Benefit legibilityWhether the value lands in a unit the price control, rate case or business plan already usesThe benefit is real but expressible only in internal or model unitsCustomer minutes lost, SAIDI minutes, TOTEX deferral, connections delivered, cost to serveFind the unit in the current business plan and check who owns that line
The selection scorecard. Score each criterion 0–3 against a single candidate decision; the maximum is 21. Every criterion is answerable from what the utility already knows, which is why the whole exercise fits in an afternoon.

The bands below convert the total into a decision. Their purpose is not precision — the difference between 14 and 15 is noise — but to make the two extremes unarguable: a candidate below 7 should not be the first thing you build no matter who is asking for it, and a candidate above 17 should not be waiting behind a longer discovery exercise.

ScoreVerdictWhat to do next
17–21Fund it nowName the owner, pull the baseline, book the change board. Nothing further is learned by studying it.
12–16Fund it once the named gap is closedOne or two criteria are dragging the score. Close those specifically — usually the baseline or the owner — rather than re-scoping the whole candidate.
7–11Not first — sequence it laterOften a genuinely valuable decision behind a change board you have not yet worked with. Build the relationship on an easier candidate and return in the next planning cycle.
0–6Do not build it as a benefit caseIf it is politically necessary, run it explicitly as a learning exercise with no benefit claim attached, and say so in writing at the start.
Verdict bands on the 0–21 total. Where two candidates are within two points of each other, break the tie on system-of-record control — it is the criterion that most often determines the timeline.

Two criteria deserve more weight than the arithmetic gives them. System-of-record control is the strongest single predictor of elapsed time, because a write-back into a system somebody else owns drags every change through a third party's release cycle — an excellent decision can be the wrong first decision for this reason alone. Benefit legibility is the strongest predictor of whether the capability survives its second year: a benefit that lands outside every unit in the business plan has to be re-explained to every new reviewer, and eventually one of them stops funding it.

Sequencing the shortlist

Plot the shortlist on benefit legibility against approval friction. The quadrant tells you the sequence, not the merit — a decision in the bottom-left may still be worth doing, just not as the one you are judged on.

Fund it this quarter

  • Benefit already speaks the plan's language
  • You control change on the system of record
  • This is the decision that funds the next three

Worth it — sequence the approvals first

  • High value, long approval path
  • Start the change-board conversation now, build later
  • Never make this your first use case

Cheap to build, hard to defend

  • Fast to ship, benefit nobody's plan recognises
  • Run it as a learning decision with no benefit claim
  • Excellent for proving the write-back pattern

The graveyard

  • Long approvals and an illegible benefit
  • Where enthusiastic programmes go to stall
  • Decline it explicitly rather than deferring it quietly
Benefit legibility — top: Customer minutes, TOTEX, connections, bottom: Value lands outside any plan unit
Approval friction — left: Your system, your change board, right: Vendor, regulator or customer-facing change

Where AI decisions live in a utility — and why each one succeeds or stalls

Seven operating domains, the decisions worth attacking in each, the system of record they execute in, the KPI they move, and the factor that most often decides their fate.

AI value in a utility concentrates in seven operating domains, and each has a characteristic success factor and a characteristic way of failing. The map below is how we shortlist candidates with operators: read the domain, check whether the system of record is one you control, confirm the KPI is a line somebody already owns in the business plan, and then read the last column — because in practice the same one or two factors decide the outcome within each domain, and they are predictable in advance.

DomainCandidate decisionsSystem of recordKPI it movesUsually decides it
Generation & renewablesWind and solar output forecasting, outage scheduling, thermal plant heat-rate tuningSCADA / historian / market systemsForecast error, availability factor, imbalance costFeedback latency — outcomes arrive within hours, so learning compounds fast
System operations & transmissionDay-ahead and intraday load forecasting, constraint cost ranking, dynamic line ratingEMSBalancing and constraint cost, marginConsequence and reversibility — advisory is the right posture, and that must be designed in
Distribution & network planningLV load-at-risk ranking, DER hosting capacity, reinforcement prioritisationADMS / GIS / DERMSConnections delivered, reinforcement TOTEXBaseline availability — network model history is often thinner than anyone expects
Outage & storm responseOutage prediction, crew pre-staging, restoration-time estimationOMS / WFMCustomer minutes lost, SAIDI, restoration-estimate accuracyRoute into operations — a portal is never opened during a storm
Assets & field workTransformer and cable health ranking, vegetation risk, inspection image triageWorks management / APM / GISDeferred replacements, TOTEX, failures avoidedBenefit legibility — capital deferral must be signed, not asserted
Customer & retailContact triage and drafting, billing anomaly detection, flexibility targetingCIS / AMI / contact platformCost to serve, complaint rate, satisfactionOwner with a moving target — the contact-centre lead's numbers must move
Safety & compliancePermit-to-work checks, evidence assembly, standards-change impact triageEHS and document systemsAudit findings, incident rate, submission effortDecision frequency — high volume makes an otherwise unglamorous domain compound
The energy and utilities decision map, indexed by success factor. 'Usually decides it' is the criterion that most often determines whether a candidate in that domain reaches production and stays there.

Outage and storm response is where most operators should start, and the reason is the last column rather than the value. The decision is made hundreds of times during an event, the system of record is one the utility controls, the outcome metric — customer minutes lost — is already in the regulatory business plan, and the previous method is a documented matrix that makes an obvious fallback. Distribution planning carries larger absolute value as electrification pressure grows, but its baselines are the thinnest in the estate, which is exactly the gap that cannot be closed retrospectively.

Customer and retail deserves more attention than it usually gets on an engineering-led roadmap. Contact volumes are enormous, feedback is immediate, the systems are usually more changeable than the operational ones, and the outcome metric is a cost line every finance function already tracks — which is why the sector's clearest published evidence of a measured, guarded AI benefit comes from that domain rather than from the control room. It is also the domain where the fallback is easiest: a human queue that already exists.

One caution across the whole map. The right first decision is rarely the most valuable decision; it is the most attributable one in a system you control, with a metric somebody already owns. The most valuable decision is usually the second or third, once the write-back pattern is proven and the change board has seen you deliver. Programmes that invert that order spend their first year negotiating rather than delivering, and get judged on the negotiation.

What success looks like in public

Three publicly reported programmes, read against the seven factors. None is an Atomic Loops engagement — each links to the organisation's own published material.

The clearest public evidence for the factors on this page is in the shape of what operators chose to publish. In each case below the notable thing is not the model: it is that the scope was bounded, the output landed where somebody already worked, the previous method stayed available, and the outcome was reported in a unit that means something to the operation — satisfaction against a human baseline, containment acres, installation cost. That is what a success factor looks like from the outside.

Three programmes read against the seven factors

Outcomes as reported by the organisations themselves — verify against the linked source before reusing them; we have not independently audited them. Card images are generated industry scenes and do not depict these organisations' facilities or people.

Generated scene: energy retailer customer operations floor with AI-assisted contact handling displaysOctopus EnergyEnergy retailer · UK-headquartered, Kraken-operated24
Challenge
Routine customer email — tariff renewals, payment dates, account details — consumes advisor time that would be better spent on complex and vulnerable-customer cases, but automating it risks degrading the experience precisely where a retailer is most visible.
Approach
Octopus built Arlo with its Kraken technology team to draft answers to routine emails only, inside stated guardrails: as Octopus describes it, Arlo "only deals with routine enquiries and works within strict guardrails" and "never handles vulnerable customers, sensitive cases or complex complaints". Every AI-written email is clearly labelled and customers can ask for a human advisor at any stage. During the trial it handled roughly 8,000 emails a week.
Reported outcome
As reported by Octopus Energy in July 2026, Arlo achieved a 76% customer satisfaction score, ahead of the 72% scored by comparable responses from human advisors, and the company said it planned to extend availability to more customers following the trial.
What it shows about the curveFactors 1, 5 and 6 in one programme. The decision was named narrowly rather than as 'AI for customer service'; the human queue remained the live fallback for everything outside the bounds; and the benefit was read against a comparable human baseline, which is why the number is quotable at all.

Octopus Energy — AI trial wins customer approval (opens in a new tab)

Generated scene: transmission corridor under wildfire-risk monitoring with detection overlaysXcel EnergyUS investor-owned utility · multi-state electric and gas23
Challenge
Wildfire ignition near power lines is a low-frequency, extremely high-consequence risk where minutes of detection delay change the outcome — and where the utility is not the body that responds, so any detection capability has to reach somebody else's dispatch process to be worth anything.
Approach
Xcel Energy deployed Pano AI camera systems, which combine high-definition cameras, AI-driven smoke detection and satellite data, each performing a 360-degree sweep every minute. Detections are verified by human analysts and triangulated for location before local fire agencies and dispatch centres are notified — the output routes into an existing emergency-response workflow rather than into a utility dashboard. Xcel has said 38 camera systems are planned across higher-risk areas of Minnesota, beginning with Mankato and Clear Lake.
Reported outcome
As reported by Xcel Energy, cameras in Douglas County, Colorado detected smoke following a lightning strike in June 2024 and firefighters contained the resulting fire to just three acres.
What it shows about the curveFactor 4 in its purest form: the output was designed to land in the responder's process, with a human verification step in between, so the measured outcome is containment rather than model precision. A detection capability that terminated in a utility screen would have been the same technology and a different result.

Xcel Energy newsroom — AI-driven wildfire detection (opens in a new tab)

Generated scene: utility-scale solar construction site with automated installation equipmentAESGlobal power company · generation and renewables developer34
Challenge
Utility-scale solar construction is dominated by a single repetitive decision-and-action loop — place the next module correctly — performed hundreds of thousands of times per project, with cost, schedule and manual-handling injury risk all concentrated in it.
Approach
AES developed Maximo, a field robot that uses lidar, cameras and computer-vision models to identify tracker structures and modules and to detect and self-correct installation inconsistencies. The system keeps installers out of repetitive heavy lifting and bending, and includes sensors that stop operation when workers are detected — safety designed as part of the capability rather than added afterwards.
Reported outcome
As reported by AES, Maximo "deploys solar panels in half the time at half the cost", installing modules roughly twice as fast as standard mechanical processes, with nearly 10 MW installed to date and plans for it to help build up to 5 GW of AES's solar pipeline over the following three years.
What it shows about the curveFactors 2 and 7. The benefit lands squarely in units a construction organisation already manages — installation cost, schedule, injury exposure — and the published pipeline commitment is what reuse looks like when it is real: the same capability applied across projects rather than rebuilt per site.

AES — Maximo (opens in a new tab)

None of the three is an argument that these organisations have solved AI adoption, and two of the three published a measured outcome rather than a projection — which is itself the point. A programme that can state what changed, in a unit its own operation already uses, against something it can compare with, has done the hard part. A programme that can only state what its model scores has not started it.

Attribution: the benefit case that survives a regulator

The success factor most programmes discover a year late. What to measure, where the baseline lives, how to attribute it, and who has to sign before it counts.

A utility AI benefit becomes permanent when the body that sets allowed revenue can follow the argument. That is a higher bar than an internal business case and a much lower bar than most teams fear: regulators do not require novelty or certainty, they require a stated method, a comparison, and a number in a unit the framework already uses. What they will not accept — and what a great many programmes offer — is a model-derived improvement multiplied by a unit cost, with no answer to the question 'compared with what?'.

The mechanics differ by jurisdiction and the discipline does not. Under Ofgem's (opens in a new tab) RIIO framework the benefit lands in TOTEX and output categories inside a price control; in the United States it lands in a state rate case, with transmission cost recovery running through FERC (opens in a new tab); European operators face the same question inside national regulatory accounts and, for some applications, an EU AI Act conformity conversation as well. In all of them the winning artefact is identical: a benefit statement with a written method, a named comparison, an owner and a date.

Benefit unitWhere the baseline livesAttribution methodSigns it
Customer minutes lost / SAIDI minutesOMS interruption records, per feeder and eventMatched feeder or district set held on the previous method, weather-normalisedNetwork operations director and the regulatory reporting lead
TOTEX deferral on asset replacementWorks management and asset registers, per asset and cohortMatched asset cohort on the incumbent condition rule, tracked for failures as well as spendAsset strategy manager and finance
Connections delivered / hosting capacity releasedGIS network model and the connections pipelineBefore-and-after on comparable application classes, with a documented counterfactualConnections lead and the business-plan owner
Cost to serve / contact handlingCIS and contact platform volumes, per contact typeHeld-back contact-type cohort or an A/B split, with satisfaction measured alongside costCustomer operations director
Imbalance and constraint costMarket and settlement systems, per settlement periodCounterfactual re-run of the incumbent method on the same periods, agreed in advanceTrading or system operations lead and finance
Safety exposure and manual handlingEHS incident and exposure records, per task typeTask-hours removed, evidenced against the incident and near-miss recordSafety director
Attribution build sheet by benefit unit. 'Signs it' is the role whose signature makes the number usable outside the programme — without it you have an estimate, not a benefit.

Building an attribution a reviewer will accept

  1. Agree the unit before the method

    Start from the business plan, not the model. Find the line the benefit will eventually land on — customer minutes, TOTEX, connections, cost to serve — and confirm who owns it. A benefit expressed in a unit nobody owns has no home and will be re-litigated by every new reviewer.

  2. Choose the comparison and write it down

    A matched control population is the strongest option: comparable feeders, districts, asset cohorts or contact types deliberately left on the previous method. Where holding a control is genuinely impossible — a single control room, a network-wide change — write the counterfactual instead: exactly what the incumbent method would have done, computed on the same periods, agreed with finance before go-live.

  3. Normalise for the things that will otherwise claim the credit

    Weather, storm severity, DER growth, reconfigurations and programme overlaps all move the same metrics. Agree the normalisation with whoever will review the number, in advance, and prefer a simple, defensible adjustment to a sophisticated one nobody outside the team can audit.

  4. Publish the method with the number, every time

    The benefit statement should be one page: the decision, the unit, the comparison, the normalisation, the result, the owner, the date. Numbers travel without their methods and lose their meaning on the way; a statement that carries its method survives the reviewer who was not in the room.

System operators have already shown that publishing the method is survivable and useful — NESO's digitalisation strategy and action plan (opens in a new tab) is an example of a body committing publicly to what it will deliver and how. Where an operator also runs a management-system frame such as ISO/IEC 42001 (opens in a new tab) or maps its practice to the NIST AI Risk Management Framework (opens in a new tab), the benefit statement and the assurance evidence come from the same artefacts, which is a large part of why attribution work pays for itself twice.

The enablement architecture, layer by layer

What actually has to exist for each proof state — including the layer most utilities have never built, which is the one that holds the evidence.

Reaching the Operating state needs four layers, and staying at Attributed needs a fifth that almost nobody builds. The architecture below is deliberately unfashionable: nothing in it is vendor-specific, and each layer is defined by what it must guarantee rather than by what product provides it. The order matters more than the contents — a programme that builds the model layer before the decision pipeline will produce excellent answers to questions nobody selected.

Layers required by proof state

Each layer is annotated with the stage that first requires it. The benefit ledger is the layer that separates Operating from Attributed, and it is the one most often absent — which is why so many working capabilities cannot say what they are worth.

  1. Decision pipeline

    Stage 1+

    • Candidate decision registerDecisions with frequency, owner, system of record, metric
    • Selection scorecardSeven criteria, scored before funding, re-run each cycle
    • Sequencing boardRanked by approval friction, not by enthusiasm
  2. Operational data

    Stage 2+

    • Historian and SCADA tagsContinuous, with explicit freshness SLAs
    • GIS network model and asset historyVersioned — a reconfiguration is a data event
    • AMI and works-management historyAt the decision's granularity, not aggregate
  3. Model and evaluation

    Stage 2+

    • Training pipelineScheduled, reproducible, versioned against the network model
    • ServingLatency budget matched to the decision cycle, not the reporting cycle
    • Evaluation in operational unitsCustomer minutes and TOTEX, never error alone
  4. Route into operations

    Stage 3+

    • Write-backInto the ADMS, OMS, GIS or works-management field in use
    • Override logEvery accept and reject, with a reason code
    • Drilled fallbackPrevious method one switch away, exercised on a date
  5. Benefit ledger

    Stage 4+

    • Baseline storePre-build metrics with definitions, immune to migrations
    • Control and counterfactual registryWhat is held back, since when, and by whom
    • Signed benefit statementsOne page each: unit, comparison, method, owner, date

Pipeline described

  1. Decision pipeline (stage 1+) — Candidate decision register: Decisions with frequency, owner, system of record, metric; Selection scorecard: Seven criteria, scored before funding, re-run each cycle; Sequencing board: Ranked by approval friction, not by enthusiasm
  2. Operational data (stage 2+) — Historian and SCADA tags: Continuous, with explicit freshness SLAs; GIS network model and asset history: Versioned — a reconfiguration is a data event; AMI and works-management history: At the decision's granularity, not aggregate
  3. Model and evaluation (stage 2+) — Training pipeline: Scheduled, reproducible, versioned against the network model; Serving: Latency budget matched to the decision cycle, not the reporting cycle; Evaluation in operational units: Customer minutes and TOTEX, never error alone
  4. Route into operations (stage 3+) — Write-back: Into the ADMS, OMS, GIS or works-management field in use; Override log: Every accept and reject, with a reason code; Drilled fallback: Previous method one switch away, exercised on a date
  5. Benefit ledger (stage 4+) — Baseline store: Pre-build metrics with definitions, immune to migrations; Control and counterfactual registry: What is held back, since when, and by whom; Signed benefit statements: One page each: unit, comparison, method, owner, date
Step-by-step insights
Decision pipeline — the layer that exists in a spreadsheet
This layer needs no technology at all, which is exactly why it gets skipped: there is nothing to procure and nobody to be excited about. Its value is that it changes where candidates come from. When decisions are sourced from a register the operation maintains, the programme attacks work the operation actually finds hard; when they are sourced from vendor roadmaps, the programme attacks work vendors find easy to demonstrate. Every subsequent layer inherits that choice, and no amount of engineering quality later compensates for it.
Operational data — versioned against the network model
Utility data has a property that catches teams from other sectors: the world it describes is itself under version control. A reconfiguration, a new connection, an AMI rollout or a GIS correction changes what a historical record means, not just what the current value is. Treat the network model as a first-class input with its own versions, and key data-quality alerts to network events rather than to statistical thresholds alone. A model trained across an undocumented reconfiguration is fitting two networks and reporting one.
Model and evaluation — evaluate in the unit you will be funded in
By the Operating state, training and serving are commodity concerns; evaluation is where the differentiating discipline sits. A health-index model is not 'good' at some ranking metric — it is good if replacements are deferred without failures rising. A storm model is good if crews arrive nearer the damage. Evaluating in customer minutes, deferred replacements and cost to serve from the start keeps the model honest and means the benefit conversation later requires translation rather than reconstruction.
Route into operations — the fallback is the approval unlock
The component most often deferred is the drilled fallback, and it is the one that determines whether anyone will let the system run. It reads as engineering pessimism; it is actually the political key to the change board, because operations leaders will accept a new decision source they can instantly revert. A write-back proposal without a rehearsed revert sits in a queue for two quarters. One with a dated drill log ships, and the drill costs an afternoon on a quiet shift.
Benefit ledger — the layer that survives the people
The benefit ledger is the only layer whose purpose is to be readable by people who were not there. It holds the pre-build baselines with their definitions, the registry of which feeders, cohorts or contact types are held back and since when, and the one-page signed statements. Its practical test is a handover: an incoming operations leader should be able to open it and learn what every live model is worth and how that was established, without asking the team that built it. That is what makes a capability institutional rather than merely working.

The two layers utilities most often skip are the first and the last, and they fail in mirror-image ways. Skipping the decision pipeline produces well-engineered answers to unselected questions. Skipping the benefit ledger produces real operational improvements that nobody can prove, which is indistinguishable from no improvement at the moment the next investment case is written.

A 90-day plan: earning an attributable benefit on transformer replacement deferral

The Committed → Operating transition made concrete on one decision — an asset health ranking written into the works-management plan, with the control cohort designed before anything runs.

Moving one proof state takes about ninety days when it is scoped to a single decision, and several years when it is scoped to a function. To make that concrete, the plan below runs the transition on a common distribution problem: ground-mounted transformer replacement prioritisation, where the plan is currently built from a condition rule and a health-index model already exists and back-tests well. Because the model exists, the quarter contains no model development at all — every day of it goes on selection, baseline, route into the plan, and the comparison that makes the number defensible.

Committed → Operating on transformer deferral, in one quarter

One licence area, one asset class, one owner. If any phase needs more than its window, narrow the scope — fewer substations, one asset cohort — rather than extending the plan.

  1. Days 1–15

    Choose the cohort and lock the baseline

    Pick one licence area and one asset class — ground-mounted distribution transformers above a given rating. Pull five years of health-index inputs, loading, dissolved-gas history, failures and replacement spend from the asset register, GIS and works-management system. Name the asset strategy manager as owner and confirm the deferral metric is in their objectives. Agree the matched control cohort with finance and regulation now, and write down the definition of every metric before anything is modelled.

    Baseline, control cohort and a named owner, all documented

  2. Days 16–45

    Put the ranking where the plan is built

    Write the risk ranking into the works-management or asset-planning field the planner already uses to build the replacement programme — not a separate portal. The incumbent condition rule stays live and visible beside it as the fallback, and the planner approves every change to the plan. Rewrite the work instruction to reference the new field, because an unchanged instruction means the change did not happen.

    Ranking visible in the planning workflow, work instruction updated

  3. Days 46–70

    Instrument the feeds, log the overrides, drill the revert

    Freshness alerting on the dissolved-gas, loading and asset-register feeds the ranking depends on, keyed to network events — a reconfiguration or a GIS correction changes what the history means. Capture every planner override with a reason code. Then exercise the revert to the condition rule once, deliberately, on a quiet week, and log the date.

    Alerting live, override log populated, fallback drilled and dated

  4. Days 71–90

    Read the benefit against the control cohort

    Compare the covered cohort against the held-back cohort on the two numbers that matter together: replacements deferred and failures experienced. A deferral benefit with rising failures is not a benefit. Write the one-page benefit statement — unit, comparison, normalisation, result, owner, date — and get it signed by the asset strategy manager and finance so it can enter the next business-plan cycle.

    A signed TOTEX deferral statement finance will defend

The order matters, and these three inversions are the common ones

  1. Baseline before build, always

    Everything else in this plan can be retrofitted; the baseline cannot. If the first fortnight slips, cut the asset class rather than the baseline work — a narrower cohort with a defensible comparison beats a broad one with an unfalsifiable claim.

  2. Route before accuracy

    A moderately accurate ranking in the planning workflow changes more replacement plans than an excellent one in a portal. Improve the model after the route exists, when you can price an accuracy point in deferred spend and avoided failures.

  3. Approval before automation

    Keep the planner's approval through the first quarter even where auto-population is technically trivial. The override log — which rankings were accepted, which were overruled, and why — is the only dataset that can later justify widening the model's remit, and it cannot be reconstructed afterwards.

Two things typically go wrong with this plan, and neither is technical. The first is that the control cohort quietly gets switched over halfway through, because it looks unfair to leave assets on the old rule — which is a legitimate concern to raise before day one and a fatal one to act on in week eight. The second is that the deferral number is reported without the failure number beside it; the pairing is what makes the claim credible, and separating them is how a good programme acquires a reputation for optimistic accounting.

Verifying the seven factors: what to read, and from where

Each factor reduces to something already recorded in a system. The instrumentation build sheet, and a checklist you can complete in an afternoon.

Every one of the seven factors is verifiable from telemetry or a dated artefact rather than from a self-report, which matters because self-assessment on this topic runs consistently optimistic. Teams remember the best use case and the intentions around it; the systems remember what actually happened. The build sheet below gives the read, the source and the threshold at which the factor counts as present.

FactorWhat you readSourceCadencePresent when
Named decisionDecision name, frequency, system of record, outcome metricCandidate registerPer planning cycleAll four fields populated for the funded candidate
Operations ownerWhether the outcome metric appears in a named manager's objectivesObjective and performance recordsAnnuallyThe metric is written into the objectives, not the project charter
Baseline before buildExtract date of the baseline versus first model-influenced decision dateBaseline store and deployment logOnce per use caseBaseline extract predates the first influenced decision
Route into operationsShare of in-scope decisions carrying the model fieldADMS / OMS / works managementWeeklyAbove 90% of in-scope decisions, and the work instruction references it
Drilled fallbackDate of the last exercised revert to the previous methodDrill logQuarterlyA dated log entry within the last six months
Signed attributionExistence and age of the one-page benefit statementBenefit ledgerPer planning cycleSigned by finance, method documented, reviewed within the cycle
ReuseElapsed days from decision agreed to decision served, per use caseDelivery trackerPer use caseFalling across the last three use cases with the same team
Instrumentation build sheet for the seven success factors in a utility estate. 'Present when' is the threshold at which the factor stops being an intention.

One measurement deserves its own note: the override rate on advisory recommendations, read from the override log. As an operating heuristic, sustained rates below roughly 40% acceptance mean the output arrives outside the workflow or the operator does not trust it — an integration or calibration problem before a model problem. A healthy band is roughly 60–85%, high enough to change decisions while overrides still carry information. Sustained acceptance above 95% usually means rubber-stamping, at which point the log has stopped generating the evidence that would later justify widening the model's remit.

The seven-factor check

One tick per factor, verified from a system or a dated artefact rather than from memory. If you cannot evidence it this afternoon, it is not present. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Nothing evidenced yet — start with selection, this month

Zero ticks is common and is not the same as zero progress; it usually means the work has been organised around technologies rather than decisions. Do not start with tooling. Build the candidate register, score three decisions against the seven criteria, and pick one. Every item on this list falls out of doing that once, properly.

How success factors decay, and what puts them back

Maturity is not monotonic. Four decays account for most of the ground utilities lose, and none of them announces itself.

Utilities lose ground on this curve quietly, because a decaying capability keeps producing output. The model still ranks, the field still populates, operations still uses it — and one of the conditions that made it defensible has stopped holding. The four decays below account for most of the regression we see, and each has a cheap preventive measure that costs far less than the recovery.

Likelihood: highImpact: high

The owner moves and the benefit becomes unowned

A reorganisation, a promotion or a merger moves the person whose objectives contained the metric. The capability keeps running and stops counting: within a year nobody can say what it was worth, and the evidence that would have funded the next investment case has no author.

PreventionName a successor for the benefit, not just for the code, and put ownership transfer on the leaver and reorganisation checklist beside system credentials.

Likelihood: mediumImpact: high

A system migration deletes the baseline

A works-management upgrade, a GIS re-platform or a historian consolidation drops history below the granularity the baseline needed, or changes a field's meaning without changing its name. The attribution breaks retrospectively — and you find out at the moment you need it.

PreventionKeep the baseline store outside the operational system's lifecycle, with definitions attached, and add 'AI baselines' to the data-migration impact checklist.

Likelihood: highImpact: medium

The fallback rots while accuracy work continues

The previous method decays out of use: the person who ran it retires, the rule stops being maintained, the switch is never exercised. The estate becomes more dependent on the model exactly as its safety net disappears, and the change board notices before the team does — usually by refusing the next write-back.

PreventionPut the fallback drill on the operational calendar with the same status as a statutory inspection, and treat a missed drill as an exception to be recorded.

Likelihood: mediumImpact: medium

The candidate register empties and selection stops

Delivery gets efficient, the obvious candidates are used up, and the pipeline fills with variants of the decision already solved because the pattern is proven and the change board is friendly. The portfolio grows while marginal benefit per use case falls, for two years, invisibly.

PreventionRe-score the whole register every planning cycle with contributions from operations, and accept that the best candidate may sit behind a change board you have never used.

The common thread is that all four decays are invisible for one planning cycle and expensive in the second. That asymmetry is why the artefacts on this page carry review dates rather than creation dates: a benefit statement, a baseline definition, a drill log and a candidate register are all claims about a world that keeps moving, and the only reliable defence is a standing process that forces somebody to look at them on a schedule.

The real opportunity with AI isn't replacing people, it's removing the repetitive things that get in the way of them doing their best work.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Candidate decision register
A maintained list of operational decisions worth attacking with AI, each with its frequency, owner, system of record and outcome metric. Distinct from a project list: it records what is available rather than what has been funded.
Selection scorecard
A seven-criterion, 0–21 rubric applied to a candidate decision before funding: system-of-record control, decision frequency, feedback latency, consequence and reversibility, baseline availability, owner with a moving target, and benefit legibility.
Baseline
The historical record of the outcome metric at the decision's granularity, extracted from the system of record before any model influences the operation, with its definition written down. It cannot be created retrospectively.
Counterfactual
A documented statement of what the incumbent method would have done over the same periods, agreed with finance in advance. The fallback attribution approach where holding a control population is genuinely impossible.
Matched control population
A comparable set of feeders, districts, asset cohorts or contact types deliberately left on the previous method so the operational improvement can be attributed rather than asserted. The strongest available evidence in a live network.
Benefit statement
A one-page record of a measured benefit: the decision, the unit, the comparison, the normalisation, the result, the owner and the date. Signed by finance, it is what lets a number travel without losing its meaning.
Benefit ledger
The store holding every live use case's baseline, control registry and signed benefit statement. The architectural layer that separates a working capability from a defensible one, and the one most often absent.
Write-back
Delivering model output into the operational system of record — an ADMS, OMS, GIS or works-management field the decision-maker already reads — rather than displaying it in a separate portal.
Override log
The record of every advisory recommendation an operator accepted or rejected, with a reason code. The most under-used dataset in the sector: it calibrates trust, locates systematic model error, and is the only basis for later widening a model's remit.
Fallback drill
A deliberate, dated exercise of the revert to the previous method, run on a quiet shift. It is the artefact that unlocks change-board approval for a write-back, because operations will accept a source it can instantly undo.
Customer minutes lost
The aggregate duration of supply interruptions experienced by customers, the standard European distribution reliability measure and the closest counterpart to SAIDI. The most commonly available benefit unit for outage-related use cases.
Reuse factor
The ratio of elapsed delivery time for the latest use case to the first, with the same team. A falling ratio is the observable signature of a capability; a flat one means the programme is building projects.

Frequently asked questions

The questions utility teams ask most often when deciding what to build, who should own it, and how to prove it worked.

What are the success factors for AI adoption in energy and utilities?

Seven recur. A named operational decision rather than a technology or a domain; an operations owner whose objectives contain the outcome metric; a baseline pulled from the system of record before the build; a route into the ADMS, OMS, GIS or works-management field the decision-maker already uses; a fallback to the previous method that has been drilled and dated; attribution against a control population or agreed counterfactual that finance will sign; and reuse, so the second decision costs less than the first. Every one is verifiable before funding.

What single factor best predicts whether a utility AI programme will succeed?

Which decision it chose. Two utilities with identical budgets, vendors and engineering capability get different outcomes if one attacks a decision made hundreds of times a week in a system it controls with five years of history, and the other attacks a decision made twice a year in a partner's platform with no comparable record. Selection is decided before any money moves and is fully checkable in an afternoon, which makes it the highest-leverage and cheapest thing to get right.

How do you choose the first AI use case in a utility?

Score at least three candidates on seven criteria: do you control change on the system where the decision executes, how often is the decision made, how quickly do you learn whether it was right, what does a wrong output cost and can it be undone, does the outcome metric exist historically at the decision's granularity, does a named operations manager's objectives contain it, and does the benefit land in a unit the business plan already uses. Pick the highest scorer, not the highest value — the most valuable decision is usually the second or third.

Why do utility AI pilots stall after a successful proof of concept?

Because the proof answered a different question from the one the operation asks. A back-test establishes that a model would have outperformed the incumbent method historically; the operation needs to know whether the person on shift will act on it. The fastest route to a demonstrable result is a portal, and a portal leaves the work instruction unchanged. In utilities this is compounded by change-controlled operational systems, so the portal is often the only thing deliverable inside the pilot window.

How do you prove an AI benefit in a price control or rate case?

State the method, name the comparison, and express the result in a unit the framework already recognises — customer minutes lost, TOTEX deferral, connections delivered, cost to serve. The strongest evidence is a matched population held on the previous method; where that is impossible, a counterfactual agreed with finance before go-live is acceptable. What reviewers reject is a model-derived improvement multiplied by a unit cost with no answer to 'compared with what?', because that claim cannot be falsified and therefore cannot be defended.

Should we build the data platform first or the first decision?

The decision. Building a platform first is the most reliable way to spend a year without changing anything operational, because the platform's requirements are being guessed rather than observed. Instrument the data behind one decision end to end — the historian tags, the GIS extract, the asset history it actually needs — ship it, then extract the shared layer at decision three or four when you know from experience which parts are genuinely common. The exception is the baseline: pull that early, because it expires.

Who should own an AI use case in a utility — innovation or operations?

Operations owns the outcome; engineering owns the system's behaviour. Concretely, a named operations manager should have the outcome metric in their objectives, and a named engineer should be accountable for freshness, drift, deployment and rollback. Programmes with only an engineering owner stall because nobody's targets improve if the model is used. Programmes with only an operational owner degrade quietly, because nobody is paged when a feed stops updating. Innovation teams can incubate; they cannot own a live operational decision.

How long should a utility AI pilot run before you decide?

Long enough to see the decision made enough times to be conclusive, which is a property of the decision rather than the calendar. A contact-triage decision reaches that in weeks; a storm-response decision needs a season; an asset-replacement decision may never reach it within a pilot at all, which is why feedback latency is a selection criterion. If a candidate cannot generate a conclusive number inside a year, do not run it as a pilot — run it as a permanent capability with a control cohort and read the benefit annually.

What is a healthy operator acceptance rate for AI recommendations?

As a heuristic, sustained acceptance below roughly 40% means the output is arriving outside the workflow or the operator does not trust it — an integration or calibration problem before a model problem. A healthy band is roughly 60–85%: high enough to change real decisions while overrides still carry information about where the model is wrong. Sustained acceptance above 95% usually indicates rubber-stamping, at which point the override log has stopped generating the evidence needed to justify widening the model's remit later.

How do we keep an AI benefit alive through a reorganisation or vendor change?

Name a successor for the benefit, not just for the code. Practically that means ownership transfer is a step in the leaver and reorganisation process; the benefit statement, baseline definitions and control registry live in a ledger outside the operational system's lifecycle; and 'AI baselines' appear on the data-migration impact checklist. The most common ending for a utility AI capability is not an incident but a reorganisation after which nobody can say what it was worth, so the evidence for the next investment case no longer exists.

Does buying an AI-enabled ADMS or APM module count as adoption?

Only if the same seven factors are present, and buying makes several of them harder rather than easier. A vendor module arrives with accuracy figures earned on another operator's network, so it needs acceptance testing against your own history before it influences anything. You control change on it less than on your own code, its baseline still has to be pulled by you, and its benefit still has to be attributed by you. Buying is a legitimate delivery choice; it is not a substitute for selection, baseline or attribution.

Do the success factors differ for a system operator, a network operator and a retailer?

The factors are identical; their difficulty is not. A system operator faces the highest consequence and lowest reversibility, so advisory posture and the drilled fallback dominate, and autonomy is correctly rare. A distribution network operator has the richest decision volume but the thinnest baselines, so the baseline work is the binding constraint. A retailer has changeable systems, immediate feedback and an obvious human fallback, which is why the sector's clearest measured benefits have come from customer operations first.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — load and generation forecasting, asset-health scoring, outage prediction and field decision support running against live operational data, integrated into the EMS, ADMS, OMS and works-management layer with the baselines, fallbacks and benefit attribution that let them stay switched on.

  • · Production deployments against live SCADA, AMI, GIS and historian data
  • · Use-case selection and maturity assessments run with utility engineering teams
  • · Attribution-first delivery: baselines, counterfactuals, signed benefit statements
  • · 18 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyEnergy and AI — executive summary (opens in a new tab)
  3. International Energy AgencyDigitalisation — decarbonisation enablers (opens in a new tab)
  4. NISTAI Risk Management Framework (opens in a new tab)
  5. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  6. OfgemOfgem — energy regulation and RIIO price controls (opens in a new tab)
  7. NERCNERC — reliability and security (opens in a new tab)
  8. FERCFederal Energy Regulatory Commission (opens in a new tab)
  9. EurelectricEurelectric — the European electricity industry (opens in a new tab)
  10. EPRIArtificial intelligence thought leadership (opens in a new tab)
  11. ENTSO-EENTSO-E — European transmission system operators (opens in a new tab)
  12. NESOJune 2025 Digitalisation Strategy and Action Plan (opens in a new tab)
  13. McKinsey & CompanyElectric power and natural gas insights (opens in a new tab)
  14. Octopus EnergyOctopus Energy's AI trial wins customer approval (opens in a new tab)
  15. KrakenOctopus Energy case study (opens in a new tab)
  16. Xcel EnergyXcel Energy brings AI-driven wildfire detection to Minnesota (opens in a new tab)
  17. Xcel EnergyUsing advanced technology to combat wildfire risk (opens in a new tab)
  18. AESMaximo — AI-enabled solar field robot (opens in a new tab)

Score the decision before you fund it

We run the precondition assessment with your operations, engineering and finance leads, score your strongest candidate decisions against all seven selection criteria, mark the baselines each one still needs, and leave you a ranked shortlist with a costed 90-day plan for the winner. You keep both whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.