Energy & UtilitiesAI Adoption & Maturity Curve
AI adoption success factors in energy and utilities: what separates programmes that stick from pilots that stall
AI adoption success factors in energy and utilities are the conditions that decide whether an AI system reaches an operational decision and stays there. Seven recur: a named decision, an operations owner whose target moves, a baseline pulled before the build, a route into the ADMS, a drilled fallback, attribution someone will sign, and reuse by the second decision.

Key takeaways
- The strongest predictor of a successful utility AI programme is which decision it chose to attack, and that choice is fully verifiable before any money is spent. Score the candidate on system-of-record control, decision frequency, feedback latency, consequence, baseline availability, owner and benefit legibility — seven criteria, twenty-one points, one afternoon.
- Sponsorship is not who attended the kick-off; it is whose targets move. A use case whose only owner sits in an innovation or digital budget has no operational advocate when peak season, a storm or a reorganisation competes for attention, and it is the first thing dropped.
- The baseline is a one-way door. Five years of the outcome metric — customer minutes lost, deferred replacements, cost to serve — must be pulled from the system of record and its counterfactual agreed before the model runs, because a baseline assembled afterwards inherits the pilot's own effects and can never cleanly attribute them.
- In a regulated utility, a benefit only becomes permanent when it is legible to the body that sets allowed revenue. Express the value in units the price control or rate case already recognises — SAIDI minutes, TOTEX deferral, connections delivered — or the second programme will be funded from goodwill rather than evidence.
- Success compounds only when the second decision costs less than the first. If use case two rebuilds its own data path, monitoring and approval route, the programme is a series of successful projects rather than a capability, and its velocity falls as its maintenance load rises.
Abbreviations used on this page
- SCADA
- Supervisory control and data acquisition
- EMS
- Energy management system (transmission control)
- ADMS
- Advanced distribution management system
- OMS
- Outage management system
- AMI
- Advanced metering infrastructure (smart meters and head-end)
- GIS
- Geographic information system — the network model of record
- CIS
- Customer information system (billing and accounts)
- WFM
- Workforce management — field crew scheduling and dispatch
- DER
- Distributed energy resources (rooftop solar, batteries, EVs)
- SAIDI
- System average interruption duration index
- RIIO
- Ofgem's price-control framework: Revenue = Incentives + Innovation + Outputs
- TOTEX
- Total expenditure — capital and operating spend treated as one pot
Free · 8 questions · ~3 minutes
Score your preconditions, not your ambitions
Eight questions, one at a time, about three minutes — each scoring a precondition that either exists or does not. Answer them and we build your personalised report: your stage on the curve, your score on each of the four precondition dimensions, and the one missing factor most likely to stall your next use case. It lands in your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the missing factor most likely to stall your next use case, and the 90-day plan for closing it — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Interested
AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
Your next moveName one operational decision, score it against the seven selection criteria, and give it an operations owner before any build is scoped.
Stage 2 · Committed
One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
Your next moveWrite the output into the ADMS, OMS, GIS or works-management field the decision-maker already reads, keep the previous method one switch away, and drill that switch.
Stage 3 · Operating
The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
Your next moveHold a matched control population or agree a written counterfactual, read the benefit against it, and get the statement signed by finance and regulation.
Stage 4 · Attributed
The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
Your next moveMove the benefits, the drill calendar and the candidate register onto standing operational and business-plan processes, so they have owners after the founders move on.
Stage 5 · Institutional
AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.
Your next movePut explicit review dates and named successors on the benefit ledger, the candidate register and the drill calendar, and audit them on the business-plan cycle.
0 / 24
Candidate selection
— / 6
Operational sponsorship
— / 6
Route into operations
— / 6
Benefit attribution
— / 6
Your score maps to a stage on the curve. The dimension breakdown matters more than the total: the lowest dimension is the precondition that caps everything else, and it is almost always cheaper to fix than the work it is currently blocking. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the curve. The dimension breakdown matters more than the total: the lowest dimension is the precondition that caps everything else, and it is almost always cheaper to fix than the work it is currently blocking.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want your candidate decisions scored with your operations leads?
We walk your operations, engineering and finance leads through the dimension scores, run the selection scorecard live against your three strongest candidate decisions, and leave you with a ranked shortlist and a costed 90-day plan for the winner. You keep both either way.
How the score maps to a stage
- 0–4 — Stage 1, Interested. AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
- 5–10 — Stage 2, Committed. One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
- 11–16 — Stage 3, Operating. The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
- 17–21 — Stage 4, Attributed. The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
- 22–24 — Stage 5, Institutional. AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.
What AI adoption success factors are in energy and utilities
A definition, the seven factors that recur, and the mechanism that connects them — from a decision nobody has chosen yet to a capability nobody has to defend.
AI adoption success factors in energy and utilities are the conditions that determine whether an AI system reaches an operational decision and stays there. They are not qualities of the model, the vendor or the budget. They are properties of the decision you chose, the person who owns it, the system it writes into and the evidence you can produce about what changed — and every one of them can be checked before a build begins, which is what makes them useful rather than retrospective.
The distinction that organises this page is between factors and outcomes. "Executive buy-in", "a data-driven culture" and "the right talent" are outcomes: real, but not actionable on a Monday morning, and impossible to verify before funding. The seven factors below are verifiable. Each one is a yes or a no about a specific artefact — a scored candidate, a named objective, a pulled baseline, a field in the ADMS or works-management system, a dated drill log, a signed benefit statement, a reuse record — and a programme missing any of them fails in a predictable way at a predictable point.
1 · A named decision, not a named technology
The unit of work is one operational decision somebody makes on a schedule — which feeders to pre-stage crews on, which transformers to advance in the replacement plan, which contacts route to an assistant — not "AI for asset management". A decision has a frequency, an owner, a system of record and a metric. A domain has none of those, so nothing about it can be baselined or finished.
2 · An operations owner whose target moves
Sponsorship is measured by whose objectives change, not by who attended the kick-off. If the only owner sits in an innovation or digital budget, the use case has no advocate during a storm week, a regulatory submission or a reorganisation — and those are exactly the moments that decide whether it survives.
3 · A baseline pulled before the build
Five years of the outcome metric, at the granularity of the decision, extracted from the system of record with its definition written down and agreed. This is a one-way door: a baseline assembled after go-live contains the pilot's own effects, and no amount of later analysis removes them.
4 · A route into the operational workflow
The output appears as a field in the ADMS, OMS, GIS, works-management or customer system the decision-maker already has open — advisory is fine, absent is fatal. A portal behind a separate login shifts the burden of action onto a person whose existing process already works, and voluntary steps are the first thing dropped under pressure.
5 · A live fallback the operation has drilled
The previous method — the condition rule, the storm matrix, the seasonal profile, the manual triage — kept alive, one switch away, and exercised on a date somebody logged. This is not engineering pessimism; it is the political key that unlocks the change board, because operations will accept a new decision source they can instantly revert.
6 · Attribution someone will sign
The benefit read against a matched control population or a counterfactual agreed in writing before go-live, expressed in a unit the price control or rate case already recognises, and signed by finance. This is the factor most programmes discover they needed a year too late.
7 · A second decision that reuses the first
Feeds, freshness alerting, the write-back pattern, the override log, the drill calendar and the benefit ledger carry over, so decision two costs meaningfully less than decision one. Without reuse a programme is a sequence of successful projects whose velocity falls as its maintenance load rises.
How a chosen decision becomes a permanent capability
The mechanism, in three lanes. The top lane is the one most programmes skip entirely — they start at delivery, with a decision a vendor demonstration chose for them. The bottom lane is where capabilities quietly die: not from a failure, but from a benefit nobody attributed and an owner nobody replaced.
- Data & feeds
- Human in the loop
- Where value leaks
- AI / model
- System-of-record action
The process, in words
- Selection happens before any money moves. A standing register of candidate decisions is scored on seven criteria, the winner gets a named operations owner with an agreed target, and five years of the outcome metric are pulled from GIS, OMS or works management with the counterfactual written down. Skip this lane and the decision arrives with the vendor demonstration that suggested it.
- Delivery is the first ninety days. Operational feeds run with freshness alerting, the model is served inside the cycle the decision actually runs on, and its output is written into the ADMS, OMS or works-management field the planner already reads. The operator acts and every override is logged with a reason code, while the previous method stays one switch away and the switch is drilled on a date somebody records.
- Permanence is everything after go-live. The benefit is read against the matched control population held back at selection, the statement is signed by finance and regulation in units the business plan recognises, and the second decision inherits the feeds, the write-back pattern and the ledger. When the owner rotates and nobody inherits the benefit, the capability keeps running and stops counting.
Step-by-step insights
- The candidate register — the artefact almost nobody has
- Most utilities can produce a list of AI projects and cannot produce a list of decisions worth attacking. The two are very different objects: a project list records what was funded, a candidate register records what is available, including the decisions nobody has proposed because no vendor sells them. The register is cheap — a spreadsheet of decisions with frequency, owner, system of record and outcome metric — and it changes the sourcing of ideas from vendor roadmaps to the operation itself. Keeping it open and re-scored every planning cycle is the single practice that most reliably prevents a programme plateauing after its obvious candidates are used up.
- Scoring before funding, and why an afternoon is enough
- The seven selection criteria are all answerable from things a utility already knows: who owns the change board on the target system, how often the decision is made, how long before you learn whether it was right, what a wrong output costs, whether the outcome metric exists at the decision's granularity, whose objectives contain it, and whether the benefit is expressible in a unit the business plan uses. None of that needs a data-science assessment or a vendor engagement. The reason to insist on scoring is not the score itself but the conversations it forces — most disqualifying facts about a candidate surface in the first twenty minutes and would otherwise have surfaced in month nine.
- The baseline is a one-way door
- Everything else on this diagram can be retrofitted. A write-back can be added later, an owner can be reassigned, a drill can be scheduled next month. The baseline cannot. Once the model is influencing the operation, every subsequent extract of the outcome metric contains its effects, and separating them requires assumptions that reviewers are entitled to reject. Pull five years of the metric at the decision's granularity — per feeder, per asset, per job, per contact type — before anything runs, write the definition down, and store it somewhere a system migration will not quietly delete. A fortnight at the start buys a defensible number for the life of the capability.
- Write-back is an approval problem wearing an engineering costume
- Teams describe the route into operations as an integration task and estimate it in developer weeks. In a utility it is mostly an approval task: the ADMS, OMS and works-management systems are change-controlled, often vendor-supported, and sometimes inside a security boundary that treats a new write path as a design change. The variable that actually sets the timeline is who owns change on that system and how recently you shipped anything to it. That is why it is a selection criterion rather than a delivery detail — a decision whose system of record belongs to somebody else can be an excellent decision and still be the wrong first one.
- The override log is the most under-used dataset in the sector
- When operations accepts an advisory recommendation, or rejects it, and records why, the utility is generating something no vendor can supply: labelled evidence about how its own experts reason about its own network. That log tells you where the model is systematically wrong, which reason codes cluster on particular circuits or asset classes, whether trust is calibrated, and — much later — where the bounds of any automation could safely sit. Programmes that skip the advisory stage to save three months arrive at the automation conversation with nothing to set thresholds from except the vendor's defaults.
- Why the bottom-right node is a risk node, not a failure
- The most common ending for a utility AI capability is not an incident. It is a reorganisation. The model keeps producing output, operations keeps using it, and the person who knew what it was worth moves to another business unit. Two years later nobody can reconstruct the baseline, the control population has been switched over, and the benefit disappears from the plan — so when the next investment case is written, the evidence that would have funded it no longer exists. Naming a successor for the benefit, not just for the code, is a five-minute act that protects years of work.
The five stages, named for what a programme has proved
Interested, Committed, Operating, Attributed, Institutional. For each: what it looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.
The five stages below are proof states rather than capability levels — each is named for what the programme has actually earned, not for what it can build. That framing matters because a utility can hold impressive capability and still sit at stage 2: an excellent model, a real owner and a convincing back-test prove nothing about the operation until something in the operation is different. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Value released across the five proof states
The shape is not linear, and the flat part is longer than most plans assume. Value stays close to zero through Interested and Committed — where a majority of operators sit — and inflects at Operating, when the output first reaches the person who acts. The second inflection is financial rather than technical: at Attributed the value becomes defensible, which is what makes it survive a planning cycle.
Operational value released by stage
- Stage 1 · Interested — 21% of operators. AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
- Stage 2 · Committed — 34% of operators. One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
- Stage 3 · Operating — 27% of operators. The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
- Stage 4 · Attributed — 13% of operators. The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
- Stage 5 · Institutional — 5% of operators. AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with IEA analysis of digitalisation in energy.
One reading discipline before the panels: score each live use case separately, then take the organisation's stage to be where the shared foundations sit rather than where the best individual example does. A single Attributed forecasting capability alongside four Interested experiments is a Committed organisation, because the next use case will inherit the foundations, not the exception.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Interested
21% of operators sit here
AI exists as intent — strategy slides, vendor demos, an innovation budget — but no operational decision has been named, so no success factor can even be tested.
Stage 1 is not scepticism — most utilities at this stage are enthusiastic. What is missing is a unit of work. The programme is described in terms of a technology and a domain ("AI for asset management", "machine learning in the control room") rather than in terms of a decision somebody makes on a Tuesday morning with a deadline. Because there is no decision, there is no owner, no metric, no system of record and no baseline, which means none of the seven success factors is present or even checkable.
The tell is the language in the steering pack. Stage-1 material talks about capability, platforms and partnerships; stage-2 material talks about a feeder set, an asset class, a call type. Ask what will be different in the operation ninety days after go-live and a stage-1 answer describes an insight; a stage-2 answer describes an action taken differently by a named role.
This is a cheap stage to leave, and expensive to linger in for a reason that is not obvious. Every quarter spent here trains the organisation to treat AI as a discussion topic. When a real candidate finally arrives, it enters an environment where nobody has ever pulled a baseline, agreed a counterfactual or approved a write-back — so the first real use case pays the cost of building all of that, and gets judged on a timeline set by the demos.
In practice
The strategy that named a technology, not a decision
A distribution business published an AI strategy with five workstreams — asset intelligence, network optimisation, customer experience, workforce productivity, data foundations. Eighteen months later it had spent real money and could not name a single decision that was made differently because of it. The workstreams were all defensible; none of them was a decision with an owner, a frequency and a metric, so none could be baselined and none could be finished.
What it looks like
- AI appears in the corporate strategy but not against a named decision
- Vendor demos and conference visits are the main source of candidate ideas
- Funding sits in an innovation or digital line, not an operational budget
- No baseline has been pulled from any system of record
Diagnostic signals you can check this week
- Ask for the candidate list: if it names technologies or domains rather than decisions with owners, you are here
- Ask which system of record the first use case will write into — a hesitation is the stage-1 signature
- Check whether anyone has pulled five years of any outcome metric from GIS, OMS or the works-management system
- Look at where the budget line sits: innovation and digital lines fund experiments, operational lines fund changes
Anti-pattern · Running a discovery programme instead of choosing
The instinctive response to an empty candidate list is a discovery exercise — workshops across every function, a longlist of eighty ideas, a heat map. It feels rigorous and it postpones the only decision that matters. Eighty unscored ideas are less useful than three scored ones, because scoring forces the questions that actually predict success: who owns the decision, where does it execute, and is there a baseline. Run the scorecard on three candidates in an afternoon and fund the winner; the longlist can wait for the second planning cycle, when you will know far more about what your estate can absorb.
What holds you here
There is no named decision, so there is nothing to own, baseline, integrate or attribute — the success factors have no subject.
Highest-leverage next move
Name one operational decision, score it against the seven selection criteria, and give it an operations owner before any build is scoped.
Cost of leaving
- Effort
- 2–6 weeks
- Team
- One operations manager and one analyst, part-time
- Risk
- Low — selection work, nothing in production changes
- To next stage
- 1–2 months
If this is you, the next step is
A short session: we run the selection scorecard on your three strongest candidates and hand you the ranking.
Stage 2
Committed
34% of operators sit here
One operational decision has been chosen, owned and baselined, and a model beats the current method on history — but nothing in the operation has changed yet.
Stage 2 is where most of the sector is, and it is the most misleading place to be, because every visible indicator is green. The decision is named. The owner exists. The back-test is convincing — the storm model would have pre-staged crews better on the last three events, the health index would have caught two of the four transformers that failed. The steering group is pleased and the P&L is untouched, because a back-test is an argument about the past and an operation runs on the present.
The structural reason is that a proof of value is optimised to answer "does this work?" while the operation's question is "will the person on shift act on it?". Those have different builds. The fastest route to a demonstrable result is an analytics portal, and a portal is precisely the artefact that leaves the work instruction unchanged. Utilities compound this because their operational systems — ADMS, OMS, works management — sit behind change boards with queues measured in months, so the portal is not only faster but also the only thing the team can ship inside the proof-of-value window.
Time here is not neutral. Control-room and field staff learn that AI output is optional; asset planners learn that the health index is a second opinion; finance learns that AI produces slide decks rather than deferred capital. The next proposal is funded against that memory. A utility that has sat at stage 2 for three years is usually harder to move than one at stage 1, because the organisational antibodies are established and the candidate list has been picked over.
In practice
The health index that never reached the plan
A network operator built a transformer health index from dissolved-gas history, loading and age, and back-tested it against five years of failures. It ranked the population better than the existing condition rule and the asset-strategy team agreed. It shipped as a quarterly spreadsheet. Two replacement plans later, the plan was still built from the condition rule, because that rule was embedded in the works-management workflow and the spreadsheet was not. Nobody rejected the model; the plan simply had no field to put it in.
What it looks like
- A specific decision is named, with an operations owner and a target metric
- A baseline has been pulled from the system of record and its definition agreed
- A model demonstrably beats the incumbent method on back-test
- No standard operating procedure or work instruction has changed
Diagnostic signals you can check this week
- Count the work instructions or standard operating procedures that changed because of the pilot — usually zero
- Ask the person who makes the decision to show you where the model output appears in their normal day
- Check whether the baseline was pulled before the model was built, and whether its definition is written down
- Ask what happens to the pilot's output during a storm week or a peak period; if the honest answer is 'nobody looks', the route into operations is the gap
Anti-pattern · Improving accuracy to earn adoption
When a proof of value is not adopted, the reflex is to make it more accurate on the theory that trust follows precision. It rarely does in a utility. Adoption is a function of where the output appears and whether the previous method is still available if it is wrong: a moderately accurate ranking inside the works-management workflow changes more plans than an excellent one in a portal behind another login. Spend the next quarter on the write-back and the fallback drill, then revisit accuracy when you can price an accuracy point in deferred replacements or customer minutes lost.
What holds you here
The output has no home in an operational system, so acting on it depends on someone choosing to add a step to their day.
Highest-leverage next move
Write the output into the ADMS, OMS, GIS or works-management field the decision-maker already reads, keep the previous method one switch away, and drill that switch.
Cost of leaving
- Effort
- 3–6 months
- Team
- One integration engineer, one ML engineer, a named operations owner, change-board sponsor
- Risk
- Medium — the first write into an operational system needs a rehearsed revert and change-board approval
- To next stage
- 3–6 months
If this is you, the next step is
The stage 2→3 transition is our most common engagement — typically 90 days on one decision.
Stage 3
Operating
27% of operators sit here
The output reaches the person who acts, in the system they already use, with the previous method drilled and every override logged — but the value is claimed, not attributed.
Stage 3 is the first stage where the operation is genuinely different. A crew is pre-staged somewhere it would not have been, a transformer is advanced or deferred, a customer contact is answered in a way it would not have been. The character of the work changes with it: stage-1 and stage-2 problems are analytical, stage-3 problems are operational, and the disciplines that solve them — freshness alerting, drilled reversion, on-call ownership, change control — come from running systems rather than from building models.
What is missing at stage 3 is almost never engineering. It is evidence. Ask what the capability is worth and the honest stage-3 answer is a story: the team believes storm pre-staging improved, the asset planners find the ranking useful, complaints about estimated restoration times seem lower. All of that may be true and none of it survives a budget cycle, because a live distribution network changes for a dozen reasons at once — weather, DER growth, a reconfiguration, a vegetation programme, a new crew contract — and any of them will happily claim the credit or take the blame.
This is where the baseline pulled at stage 2 earns its keep, and where its absence is fatal in a way that cannot be repaired later. Attribution requires a comparison the operation agreed to in advance: a matched set of feeders, depots, asset cohorts or customer segments left on the previous method, or a documented counterfactual that finance accepted before the model ran. Retrofitting one after go-live is not a smaller version of the same job; it is a different and much weaker claim, and everybody in the room knows it.
In practice
The pre-staging that everyone believed in
A utility ran storm crew pre-staging from a model wired into its OMS for two seasons. Field leadership was convinced it helped, and it very likely did. When the capital review asked what it was worth, the team could produce restoration times for the two seasons — which were also the two seasons a new mutual-aid agreement and a large vegetation programme landed. No comparable districts had been held on the old matrix. The capability survived on sponsorship rather than evidence, and the second use case was funded a year late.
What it looks like
- Output lands in an ADMS, OMS, GIS or works-management field, at least advisory
- Overrides are logged with a reason code, not just accepted or ignored
- A fallback to the previous method has been exercised and dated
- The work instruction has been rewritten to reference the new field
Diagnostic signals you can check this week
- Ask whether any part of the network, asset population or customer base is deliberately still on the previous method — if not, attribution is already compromised
- Pull the override log: both a near-zero and a near-total override rate are alarms, and an absent log is worse than either
- Ask for the date of the last fallback drill; an assurance without a log entry means the switch is theoretical
- Ask finance what number they would put in a submission today, and watch whether the answer is a range, a story or a refusal
Anti-pattern · Declaring the benefit from the model's own metrics
With the output live and operations content, the temptation is to convert model performance into money directly: the ranking is X% better, replacements cost Y, therefore the benefit is X times Y. Every utility that has tried this has met the same reviewer question — compared with what? — and had no answer. The conversion is not wrong in principle; it is unfalsifiable in practice, and unfalsifiable numbers are removed from submissions rather than argued with. Hold a matched population, or agree a written counterfactual with finance before go-live. It costs a fortnight at the start and is unbuyable afterwards.
What holds you here
The value is asserted rather than attributed, so the capability is funded by sponsorship and falls with the sponsor.
Highest-leverage next move
Hold a matched control population or agree a written counterfactual, read the benefit against it, and get the statement signed by finance and regulation.
Cost of leaving
- Effort
- 6–12 months
- Team
- Operations owner, data engineer, finance analyst, regulatory or business-plan contact
- Risk
- Medium — the control population must be genuinely comparable and genuinely left alone
- To next stage
- 6–12 months
If this is you, the next step is
We design the control population or counterfactual with your finance and regulation leads, and leave you the method note.
Stage 4
Attributed
13% of operators sit here
The benefit is measured against a control or agreed counterfactual, signed by finance and regulation, and the second and third decisions reuse the first one's foundations.
At stage 4 the argument stops being about AI. The conversation in the investment committee is about which decisions are on the roadmap and what each is worth, in the same units as every other proposal — customer minutes lost, TOTEX deferral, connections delivered, cost to serve. That is a healthier and much duller conversation, and its dullness is the point: capability that has to be re-sold every year is not capability, it is a campaign.
The engineering signature of this stage is that the marginal cost of the next decision has fallen. Feeds, freshness alerting, the write-back pattern into the ADMS or works-management system, the override log, the drill calendar and the benefit ledger are shared, so a new decision is mostly specification: which decision, which owner, which metric, which control population. When the third use case takes a third of the elapsed time of the first with the same team, the reuse factor is real. When it takes the same, the programme has built three projects and called it a platform.
The constraint that emerges here is not technical and not financial — it is the candidate pipeline. A programme that has proved it can convert a scored decision into a signed benefit will exhaust its obvious candidates within a couple of planning cycles, and the quality of the next tranche depends entirely on whether the register is being refreshed by the people closest to the work. Utilities that keep the register open, scored and visible keep compounding; utilities that treat selection as a one-off exercise plateau with a small portfolio of well-run models and a quietly ageing roadmap.
In practice
The third decision that cost a third
An operator that had wired outage prediction into its OMS added vegetation risk scoring and then LV load-at-risk ranking. The first took nine months, most of it spent proving the write-back pattern and negotiating the change board. The second took four. The third took eleven weeks, of which seven were spent agreeing the control population and the benefit unit with finance — the engineering was under three weeks. That ratio, specification-heavy and build-light, is the stage-4 signature.
What it looks like
- A signed benefit statement exists in units the price control or rate case recognises
- The second decision reused the first one's feeds, monitoring and approval route
- Time from decision agreed to decision served is falling, use case by use case
- The candidate register is scored and re-ranked on a planning cadence
Diagnostic signals you can check this week
- Compare elapsed time from decision agreed to decision served across your last three use cases; if it is flat, nothing is being reused
- Ask whether a single benefit ledger holds every live use case's baseline, control design and signed statement
- Ask an operations leader — not an engineer — what each live model is worth; a confident answer in operational units is the stage-4 tell
- Check when the candidate register was last re-scored and who contributed to it
Anti-pattern · Industrialising delivery while selection goes stale
Stage 4 makes delivery efficient, which makes it tempting to fill the pipeline with whatever is nearest — usually more variants of the decision you already solved, because the pattern is proven and the change board is friendly. The portfolio grows and the marginal benefit per use case falls, quietly, for two years. The discipline is to re-run selection against the whole register every planning cycle and accept that the highest-scoring candidate may sit in a function you have never worked with, needing a change board you have not met. Efficiency in delivery is not a reason to stop choosing well.
What holds you here
Success now depends on the individuals and sponsors who built it — nothing yet makes the capability survive a reorganisation, a vendor change or a price-control reset.
Highest-leverage next move
Move the benefits, the drill calendar and the candidate register onto standing operational and business-plan processes, so they have owners after the founders move on.
Cost of leaving
- Effort
- 12–24 months
- Team
- Platform engineers, operations product owner, finance and regulatory partners
- Risk
- Medium — the shared layer competes with new-use-case demand for the same people
- To next stage
- 18+ months
If this is you, the next step is
We map what your last three use cases actually shared, and what a fourth would still have to rebuild.
Stage 5
Institutional
5% of operators sit here
AI is a standing operating capability: benefits sit in the business plan, drills sit on the operational calendar, the candidate register is refreshed, and none of it depends on named individuals.
Stage 5 in a utility is not an autonomous grid; it is an unremarkable one. The distinguishing property is that nothing about the capability is special-cased. A model review sits in the same calendar as a relay test. A benefit statement is produced by the same process that produces every other business-plan output. Ownership transfer is a line on the leaver checklist beside system credentials. When AI stops needing its own governance forum, its own budget line and its own champion, it has become part of how the utility runs.
The risk that defines this stage is renewal rather than failure. Price controls reset, vendors are replaced, control rooms are merged, asset strategies change, and every one of those events silently invalidates something: a baseline definition, a control population, a data path, a drill owner. Institutional programmes survive them because the artefacts are on standing processes that force review — the business-plan cycle, the works-management release cycle, the operational calendar — rather than in a wiki that only the original team knows about.
The honest thing to say about stage 5 is that very few operators are here, and that arriving is less impressive than staying. The utilities that hold it treat the candidate register, the benefit ledger and the drill calendar as living operational artefacts with named owners and review dates, and they can produce all three inside an afternoon for an auditor, a regulator or an incoming executive. That reproducibility, not any particular model, is what makes the capability institutional.
In practice
The capability that survived a merger
Two control-room organisations merged and a large share of the combined analytics team left within a year. The forecasting and outage capabilities kept running, because each had a named operational owner in the receiving organisation, a runbook in the same repository as the rest of the operational documentation, benefits already booked in the business plan, and drill dates already in the operational calendar. Nothing heroic happened. Nobody had to rediscover what any of it was for, which is the whole test.
What it looks like
- Benefits appear in the business plan or rate filing with a documented method
- Fallback drills and model reviews sit on the same calendar as protection and relay tests
- Ownership transfer is a defined step in the leaver and reorganisation process
- The candidate register is refreshed by operations, not by the AI team
Diagnostic signals you can check this week
- Ask for the benefit ledger, the candidate register and the drill calendar; if producing them takes more than an afternoon, they are not institutional
- Check whether ownership transfer for a live model is a named step in the leaver and reorganisation process
- Ask what happened to a live use case at the last reorganisation or vendor change, and who noticed
- Ask whether an incoming operations leader could find out what each model is worth without asking the AI team
Anti-pattern · Assuming institutional means finished
Once the artefacts exist and the capability has survived a reorganisation, programmes stop investing in the parts nobody is currently complaining about: the register goes unrefreshed, drills slip past their dates, baselines quietly age out through a system migration. Decay is invisible for one planning cycle and expensive in the second, because the evidence chain breaks retrospectively — you discover the baseline is gone at the moment you need it. Put review dates on the artefacts themselves and treat a missed drill or a stale register with the seriousness of a missed statutory inspection.
What holds you here
Sustaining the capability is a renewal discipline — the constraint becomes keeping baselines, registers and drills current through resets, migrations and reorganisations.
Highest-leverage next move
Put explicit review dates and named successors on the benefit ledger, the candidate register and the drill calendar, and audit them on the business-plan cycle.
Cost of leaving
- Effort
- Continuous
- Team
- Operations owners, platform team, finance and regulatory partners on standing cycles
- Risk
- Low frequency, high consequence — decay is silent and only surfaces when evidence is demanded
If this is you, the next step is
We stress-test the artefacts against a real handover scenario and tell you what would not have survived.
Where energy and utilities operators actually sit today
The distribution across the five proof states, why the Committed crowd is so large, and what the external research says about the pressure now pushing against it.
Most energy and utilities operators are Committed rather than Operating: a decision has been chosen, a model beats the incumbent method on history, and nothing in the operation has changed. The crowding is not a failure of ambition. It is the predictable result of two structural facts — that the fastest route to a demonstrable result is an analytics portal, and that the operational systems where utility decisions actually execute sit behind change boards with queues measured in months. The proof of value can be finished inside its own window; the write-back cannot.
Distribution of energy and utilities operators across the five proof states
Illustrative distribution, synthesised from IEA, EPRI and Eurelectric adoption research — charted to show shape, not to report a survey. Committed is the mode. The drop from Committed to Operating is the largest transition loss on the curve, and it is an integration and approval gap rather than a modelling one.
Share of operators (illustrative)
- 21% — 1 · Interested
- 34% — 2 · Committed (the crowd)
- 27% — 3 · Operating
- 13% — 4 · Attributed
- 5% — 5 · Institutional
Source: Illustrative; synthesised from IEA, EPRI and Eurelectric research
The pressure on the Committed crowd is rising from both directions. On the demand side, the IEA's Energy and AI report (opens in a new tab) records data centres consuming around 415 TWh in 2024 — about 1.5% of world electricity — and projects that more than doubling to roughly 945 TWh by 2030, with around one-tenth of global electricity demand growth to 2030 coming from that single sector. On the supply side the same report finds that AI tools could unlock up to 175 GW of transmission capacity on existing networks. Both numbers describe decisions — where to connect, what to reinforce, how far to push a circuit — that a utility currently makes with deterministic rules and engineering judgement.
The governance side has settled enough to stop being an excuse. NIST's AI Risk Management Framework (opens in a new tab) and ISO/IEC 42001 (opens in a new tab) give operators a management-system frame that predates any single regulator's position, Ofgem (opens in a new tab) and NERC (opens in a new tab) both ask for evidence rather than abstention, and sector bodies including Eurelectric (opens in a new tab), EPRI (opens in a new tab) and ENTSO-E (opens in a new tab) publish on adoption openly. Cross-industry research such as McKinsey's electric power and natural gas insights (opens in a new tab) continues to find experimentation running well ahead of impact. What separates the operators moving from those standing still is not access to any of this; it is whether the seven factors were engineered before the build.
The selection scorecard: score a decision before you fund it
The centrepiece. Seven criteria, twenty-one points, one afternoon — the instrument that decides more about a utility AI programme's outcome than anything that happens after funding.
Score a candidate decision on seven criteria before you fund it, because selection explains more variance in outcomes than delivery does. Two utilities with identical engineering capability, identical budgets and identical vendors will get different results if one attacks a decision made four hundred times a week in a system it controls with five years of history behind it, and the other attacks a decision made twice a year in a partner's platform with no comparable record. The scorecard below makes that difference visible in an afternoon rather than in month nine.
This is a different instrument from the maturity assessment further up the page. The assessment scores your organisation; the scorecard scores one candidate decision. A mature organisation can still choose badly, and an Interested organisation that picks well will out-deliver a Committed one that picks by demonstration. Run the scorecard on at least three candidates so the scores are comparative — a single score in isolation tells you almost nothing.
| Criterion | What you are actually testing | 0 — disqualifying | 3 — ideal | The afternoon check |
|---|---|---|---|---|
| System-of-record control | Whether you can change the system where the decision executes, and how fast | It executes in a vendor-hosted or customer-facing system you cannot change | Your own ADMS, OMS, GIS or works-management system, and you shipped a change to it this quarter | Ask who chairs that system's change board and when it last approved a new field |
| Decision frequency | How much learning and how much value the decision can accumulate per quarter | A handful of instances a year — an annual plan or a one-off study | Hundreds or thousands of instances a week, per feeder, asset, job or contact | Count the decision's instances in the system of record over the last 90 days |
| Feedback latency | How long before you learn whether a given decision was right | Years — asset replacement outcomes, long-horizon investment choices | Hours to days — crew routing, switching option ranking, contact triage | Ask when the outcome field for that decision is populated, and by what process |
| Consequence and reversibility | What a wrong output costs, and how quickly it can be undone | Safety exposure or a supply interruption that cannot be unwound | Contained and reversible within a shift, with the previous method available | Walk one plausible wrong output through to its worst realistic outcome |
| Baseline availability | Whether the outcome metric exists in a system, at the decision's granularity, for long enough | No historical record, or only at aggregate level | Three to five years, per decision instance, in a system you can query today | Try to pull twelve months of it yourself before the meeting ends |
| Owner with a moving target | Whether a named operations manager's objectives change if this works | Only an innovation, digital or data team is accountable | A named operations manager with the metric written into their objectives | Read the objectives; if the metric is not in them, the owner is nominal |
| Benefit legibility | Whether the value lands in a unit the price control, rate case or business plan already uses | The benefit is real but expressible only in internal or model units | Customer minutes lost, SAIDI minutes, TOTEX deferral, connections delivered, cost to serve | Find the unit in the current business plan and check who owns that line |
The bands below convert the total into a decision. Their purpose is not precision — the difference between 14 and 15 is noise — but to make the two extremes unarguable: a candidate below 7 should not be the first thing you build no matter who is asking for it, and a candidate above 17 should not be waiting behind a longer discovery exercise.
| Score | Verdict | What to do next |
|---|---|---|
| 17–21 | Fund it now | Name the owner, pull the baseline, book the change board. Nothing further is learned by studying it. |
| 12–16 | Fund it once the named gap is closed | One or two criteria are dragging the score. Close those specifically — usually the baseline or the owner — rather than re-scoping the whole candidate. |
| 7–11 | Not first — sequence it later | Often a genuinely valuable decision behind a change board you have not yet worked with. Build the relationship on an easier candidate and return in the next planning cycle. |
| 0–6 | Do not build it as a benefit case | If it is politically necessary, run it explicitly as a learning exercise with no benefit claim attached, and say so in writing at the start. |
Two criteria deserve more weight than the arithmetic gives them. System-of-record control is the strongest single predictor of elapsed time, because a write-back into a system somebody else owns drags every change through a third party's release cycle — an excellent decision can be the wrong first decision for this reason alone. Benefit legibility is the strongest predictor of whether the capability survives its second year: a benefit that lands outside every unit in the business plan has to be re-explained to every new reviewer, and eventually one of them stops funding it.
Sequencing the shortlist
Plot the shortlist on benefit legibility against approval friction. The quadrant tells you the sequence, not the merit — a decision in the bottom-left may still be worth doing, just not as the one you are judged on.
Fund it this quarter
- Benefit already speaks the plan's language
- You control change on the system of record
- This is the decision that funds the next three
Worth it — sequence the approvals first
- High value, long approval path
- Start the change-board conversation now, build later
- Never make this your first use case
Cheap to build, hard to defend
- Fast to ship, benefit nobody's plan recognises
- Run it as a learning decision with no benefit claim
- Excellent for proving the write-back pattern
The graveyard
- Long approvals and an illegible benefit
- Where enthusiastic programmes go to stall
- Decline it explicitly rather than deferring it quietly
Where AI decisions live in a utility — and why each one succeeds or stalls
Seven operating domains, the decisions worth attacking in each, the system of record they execute in, the KPI they move, and the factor that most often decides their fate.
AI value in a utility concentrates in seven operating domains, and each has a characteristic success factor and a characteristic way of failing. The map below is how we shortlist candidates with operators: read the domain, check whether the system of record is one you control, confirm the KPI is a line somebody already owns in the business plan, and then read the last column — because in practice the same one or two factors decide the outcome within each domain, and they are predictable in advance.
| Domain | Candidate decisions | System of record | KPI it moves | Usually decides it |
|---|---|---|---|---|
| Generation & renewables | Wind and solar output forecasting, outage scheduling, thermal plant heat-rate tuning | SCADA / historian / market systems | Forecast error, availability factor, imbalance cost | Feedback latency — outcomes arrive within hours, so learning compounds fast |
| System operations & transmission | Day-ahead and intraday load forecasting, constraint cost ranking, dynamic line rating | EMS | Balancing and constraint cost, margin | Consequence and reversibility — advisory is the right posture, and that must be designed in |
| Distribution & network planning | LV load-at-risk ranking, DER hosting capacity, reinforcement prioritisation | ADMS / GIS / DERMS | Connections delivered, reinforcement TOTEX | Baseline availability — network model history is often thinner than anyone expects |
| Outage & storm response | Outage prediction, crew pre-staging, restoration-time estimation | OMS / WFM | Customer minutes lost, SAIDI, restoration-estimate accuracy | Route into operations — a portal is never opened during a storm |
| Assets & field work | Transformer and cable health ranking, vegetation risk, inspection image triage | Works management / APM / GIS | Deferred replacements, TOTEX, failures avoided | Benefit legibility — capital deferral must be signed, not asserted |
| Customer & retail | Contact triage and drafting, billing anomaly detection, flexibility targeting | CIS / AMI / contact platform | Cost to serve, complaint rate, satisfaction | Owner with a moving target — the contact-centre lead's numbers must move |
| Safety & compliance | Permit-to-work checks, evidence assembly, standards-change impact triage | EHS and document systems | Audit findings, incident rate, submission effort | Decision frequency — high volume makes an otherwise unglamorous domain compound |
Outage and storm response is where most operators should start, and the reason is the last column rather than the value. The decision is made hundreds of times during an event, the system of record is one the utility controls, the outcome metric — customer minutes lost — is already in the regulatory business plan, and the previous method is a documented matrix that makes an obvious fallback. Distribution planning carries larger absolute value as electrification pressure grows, but its baselines are the thinnest in the estate, which is exactly the gap that cannot be closed retrospectively.
Customer and retail deserves more attention than it usually gets on an engineering-led roadmap. Contact volumes are enormous, feedback is immediate, the systems are usually more changeable than the operational ones, and the outcome metric is a cost line every finance function already tracks — which is why the sector's clearest published evidence of a measured, guarded AI benefit comes from that domain rather than from the control room. It is also the domain where the fallback is easiest: a human queue that already exists.
One caution across the whole map. The right first decision is rarely the most valuable decision; it is the most attributable one in a system you control, with a metric somebody already owns. The most valuable decision is usually the second or third, once the write-back pattern is proven and the change board has seen you deliver. Programmes that invert that order spend their first year negotiating rather than delivering, and get judged on the negotiation.
What success looks like in public
Three publicly reported programmes, read against the seven factors. None is an Atomic Loops engagement — each links to the organisation's own published material.
The clearest public evidence for the factors on this page is in the shape of what operators chose to publish. In each case below the notable thing is not the model: it is that the scope was bounded, the output landed where somebody already worked, the previous method stayed available, and the outcome was reported in a unit that means something to the operation — satisfaction against a human baseline, containment acres, installation cost. That is what a success factor looks like from the outside.
Three programmes read against the seven factors
Outcomes as reported by the organisations themselves — verify against the linked source before reusing them; we have not independently audited them. Card images are generated industry scenes and do not depict these organisations' facilities or people.
Octopus EnergyEnergy retailer · UK-headquartered, Kraken-operated24
- Challenge
- Routine customer email — tariff renewals, payment dates, account details — consumes advisor time that would be better spent on complex and vulnerable-customer cases, but automating it risks degrading the experience precisely where a retailer is most visible.
- Approach
- Octopus built Arlo with its Kraken technology team to draft answers to routine emails only, inside stated guardrails: as Octopus describes it, Arlo "only deals with routine enquiries and works within strict guardrails" and "never handles vulnerable customers, sensitive cases or complex complaints". Every AI-written email is clearly labelled and customers can ask for a human advisor at any stage. During the trial it handled roughly 8,000 emails a week.
- Reported outcome
- As reported by Octopus Energy in July 2026, Arlo achieved a 76% customer satisfaction score, ahead of the 72% scored by comparable responses from human advisors, and the company said it planned to extend availability to more customers following the trial.
- What it shows about the curveFactors 1, 5 and 6 in one programme. The decision was named narrowly rather than as 'AI for customer service'; the human queue remained the live fallback for everything outside the bounds; and the benefit was read against a comparable human baseline, which is why the number is quotable at all.
Octopus Energy — AI trial wins customer approval (opens in a new tab)
Xcel EnergyUS investor-owned utility · multi-state electric and gas23
- Challenge
- Wildfire ignition near power lines is a low-frequency, extremely high-consequence risk where minutes of detection delay change the outcome — and where the utility is not the body that responds, so any detection capability has to reach somebody else's dispatch process to be worth anything.
- Approach
- Xcel Energy deployed Pano AI camera systems, which combine high-definition cameras, AI-driven smoke detection and satellite data, each performing a 360-degree sweep every minute. Detections are verified by human analysts and triangulated for location before local fire agencies and dispatch centres are notified — the output routes into an existing emergency-response workflow rather than into a utility dashboard. Xcel has said 38 camera systems are planned across higher-risk areas of Minnesota, beginning with Mankato and Clear Lake.
- Reported outcome
- As reported by Xcel Energy, cameras in Douglas County, Colorado detected smoke following a lightning strike in June 2024 and firefighters contained the resulting fire to just three acres.
- What it shows about the curveFactor 4 in its purest form: the output was designed to land in the responder's process, with a human verification step in between, so the measured outcome is containment rather than model precision. A detection capability that terminated in a utility screen would have been the same technology and a different result.
Xcel Energy newsroom — AI-driven wildfire detection (opens in a new tab)
AESGlobal power company · generation and renewables developer34
- Challenge
- Utility-scale solar construction is dominated by a single repetitive decision-and-action loop — place the next module correctly — performed hundreds of thousands of times per project, with cost, schedule and manual-handling injury risk all concentrated in it.
- Approach
- AES developed Maximo, a field robot that uses lidar, cameras and computer-vision models to identify tracker structures and modules and to detect and self-correct installation inconsistencies. The system keeps installers out of repetitive heavy lifting and bending, and includes sensors that stop operation when workers are detected — safety designed as part of the capability rather than added afterwards.
- Reported outcome
- As reported by AES, Maximo "deploys solar panels in half the time at half the cost", installing modules roughly twice as fast as standard mechanical processes, with nearly 10 MW installed to date and plans for it to help build up to 5 GW of AES's solar pipeline over the following three years.
- What it shows about the curveFactors 2 and 7. The benefit lands squarely in units a construction organisation already manages — installation cost, schedule, injury exposure — and the published pipeline commitment is what reuse looks like when it is real: the same capability applied across projects rather than rebuilt per site.
None of the three is an argument that these organisations have solved AI adoption, and two of the three published a measured outcome rather than a projection — which is itself the point. A programme that can state what changed, in a unit its own operation already uses, against something it can compare with, has done the hard part. A programme that can only state what its model scores has not started it.