Energy & UtilitiesReadiness & Transformation Roadmap
The AI readiness framework for utilities: scoring a candidate before you fund it
An AI readiness framework for utilities is a scored, evidence-backed test applied to one named decision on one named part of the estate, returning a verdict — proceed, proceed under conditions, shadow only, or refuse — that binds at the gates the utility already runs and carries an expiry date.

Key takeaways
- Readiness is not maturity. Maturity describes an organisation over time; a readiness verdict describes one named decision, on one named part of the estate, at one point in time — and it expires. A framework that produces a single organisational percentage cannot approve or refuse anything, which is why most of them change nothing.
- The framework's authority does not come from the quality of its questionnaire. It comes from binding to gates the utility already operates: the capital-approval record, the OT change and CIP applicability determination, and the operational readiness review before go-live. A verdict that is not a required field on one of those forms is advice.
- Score evidence, not opinion. Every test in a readiness dossier should name an artefact a reviewer can pull from GIS, ADMS, OMS, AMI, EAM or the historian without asking the delivery team. A claim with no artefact scores zero — not 'partially met'.
- Utilities already own the strongest readiness discipline of any industry: factory and site acceptance testing, commissioning, protection-settings verification, operational readiness reviews and certification intervals. Write the AI framework as an extension of that machinery rather than importing a generic IT maturity model.
- Readiness decays because the estate changes faster than the documents describing it. Give every verdict an expiry, re-score from telemetry the tests that telemetry can measure, and write the withdrawal trigger before go-live rather than during the first incident.
Abbreviations used on this page
- ADMS
- Advanced distribution management system
- SCADA
- Supervisory control and data acquisition (with the EMS, the control-room system pair)
- OMS
- Outage management system
- AMI
- Advanced metering infrastructure (the smart-meter estate and its head-end)
- GIS
- Geographic information system — the network connectivity and asset record
- EAM
- Enterprise asset management (work, maintenance and inspection records)
- OT
- Operational technology — the control and protection estate, as distinct from IT
- CIP
- Critical Infrastructure Protection — the NERC cyber-security standards family
- ORR
- Operational readiness review — the go-live gate before a change enters service
- SAT
- Site acceptance test (paired with the factory acceptance test, FAT)
- ETR
- Estimated time of restoration — the published restoration promise during an outage
- SAIDI
- System average interruption duration index — the headline reliability measure
Free · 8 questions · ~3 minutes
Score your readiness framework, not your organisation
Eight questions, one at a time, about three minutes. They score the instrument rather than the estate: how your organisation currently establishes whether an AI candidate is ready, what evidence that answer rests on, and whether the answer can refuse anything. Answer them and we build your personalised report — your stage on the ladder, your score on each of the four dimensions, and the specific gap between your verdicts and the gates they should be binding to — and send it to your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised readiness-framework report is ready
Tell us where to send it. Your stage appears on screen immediately, and the full report — dimension scores, the gates your verdicts are not yet binding to, and the shortest route from your current instrument to a binding one — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Opinion
Readiness is whatever the person you ask believes it is — no instrument, no evidence file, and no record that a readiness decision was ever made.
Your next movePick one named candidate decision and write down, on one page, what evidence would have to exist for it to be ready and which system that evidence lives in. Do not score anything yet.
Stage 2 · Survey
Readiness is a self-scored questionnaire about the organisation, producing a number that ranks nothing and can refuse nothing.
Your next moveRe-run the assessment on ONE named candidate decision, scoring only the tests whose evidence you can pull yourself from GIS, ADMS, OMS, AMI, EAM or the historian.
Stage 3 · Dossier
Readiness is an evidence file per candidate, scored by named assessors against artefacts pulled from the estate rather than supplied by the delivery team.
Your next moveGet the readiness verdict named as a required input on the artefacts that already control money and OT access — the capital-approval record, the change record, and the operational readiness review checklist.
Stage 4 · Gate
Readiness is a binding condition: no verdict, no funding release, no OT connection and no go-live, with conditions tracked to closure by name and date.
Your next moveStamp every verdict with an expiry date and a named re-assessor, and write the withdrawal criteria before go-live rather than during the first incident.
Stage 5 · Standing
Readiness is a live, expiring status: verdicts carry intervals, the measurable tests re-score from telemetry, and a stated trigger withdraws the decision from service.
Your next moveTreat the readiness verdict as a maintained artefact with an interval, an owner and a test record — the same discipline the estate already applies to protection settings and operator certification.
0 / 24
Evidence basis
— / 6
Scope resolution
— / 6
Control path
— / 6
Standing and renewal
— / 6
Your score maps to a stage on the readiness-framework ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the instrument, and for most utilities it is not evidence basis — it is control path or standing, the two dimensions a questionnaire never reaches. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the readiness-framework ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the instrument, and for most utilities it is not evidence basis — it is control path or standing, the two dimensions a questionnaire never reaches.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want a real candidate run through your framework by an outside assessor?
We take one live candidate from your portfolio, pull the evidence from your own systems rather than asking your team for it, score the twelve tests, and hand you the dossier and the verdict — including the refusal, if that is what the evidence says. You keep the dossier and the test definitions either way.
How the score maps to a stage
- 0–4 — Stage 1, Opinion. Readiness is whatever the person you ask believes it is — no instrument, no evidence file, and no record that a readiness decision was ever made.
- 5–10 — Stage 2, Survey. Readiness is a self-scored questionnaire about the organisation, producing a number that ranks nothing and can refuse nothing.
- 11–15 — Stage 3, Dossier. Readiness is an evidence file per candidate, scored by named assessors against artefacts pulled from the estate rather than supplied by the delivery team.
- 16–20 — Stage 4, Gate. Readiness is a binding condition: no verdict, no funding release, no OT connection and no go-live, with conditions tracked to closure by name and date.
- 21–24 — Stage 5, Standing. Readiness is a live, expiring status: verdicts carry intervals, the measurable tests re-score from telemetry, and a stated trigger withdraws the decision from service.
What an AI readiness framework for utilities actually is
A definition, the difference between readiness and maturity, and the path a readiness verdict travels through the assurance machinery a utility already runs.
An AI readiness framework for utilities is a scored, evidence-backed test applied to one named decision on one named part of the estate, which returns a verdict a decision-maker can act on: proceed, proceed under conditions, shadow only, or refuse. It has four working parts — a defined scope, a dossier of evidence pulled from the operational estate, a verdict issued by a named panel, and an expiry date. Remove any one of the four and what remains is a survey.
Readiness is routinely confused with maturity, and the confusion is expensive. Maturity describes an organisation's capability over time and is the right frame for a multi-year investment argument. Readiness describes one candidate at one moment and is the right frame for a funding release, an OT connection or a go-live. A utility can be genuinely mature and not ready for a specific candidate — the capability exists but the connectivity records on those particular feeders do not support it. It can also be immature and entirely ready for a narrow one, which is how most good programmes start.
It is also distinct from the two artefacts it is most often bundled with. A transformation blueprint says what the organisation intends to build; an integration roadmap says how deep into the grid a model may eventually reach. A readiness framework answers neither question. It answers a narrower and more immediately useful one: given this candidate, on this estate, today, may it proceed — and if not, exactly which artefact is missing and how long will it take to produce?
A defined scope, not a domain
"AI in outage management" is not a scope. "The published estimated time of restoration for distribution outages in the Eastern operating region, decided by the storm-room duty manager, written into the OMS" is a scope. Everything downstream — which evidence, which thresholds, which gate — is derived from that sentence, and a vague sentence produces a vague verdict.
Evidence pulled, never supplied
The assessor's read access to GIS, ADMS, OMS, AMI, EAM and the historian is infrastructure, not a favour requested per assessment. The moment evidence arrives as an attachment from the delivery team, the framework has quietly reverted to self-report while keeping the appearance of rigour.
A verdict with four possible values
Binary ready/not-ready loses the most useful answer in the set. "Shadow only" — run it read-only, against live data, writing nowhere — is frequently the correct verdict, and it is the one verdict that manufactures its own evidence while you wait for the slow gaps to close.
An expiry date and a withdrawal trigger
A utility already applies intervals to relay maintenance, protection-settings verification and system operator certification (opens in a new tab). A readiness verdict is the same class of artefact and decays for the same reason: the estate keeps moving after the paperwork stops.
The path a readiness verdict travels: evidence, assessment, and the gates that can refuse
The three lanes are the three things that have to be true for a readiness answer to be worth anything: the evidence was pulled rather than supplied, one panel scored one dossier against one scope, and the verdict binds at gates that already control money and OT access. Most frameworks build the first two lanes and stop before the third, which is why they change nothing.
- Data & feeds
- Where value leaks
- Human in the loop
- System-of-record action
The process, in words
- Evidence lane: the assessor pulls artefacts directly from GIS, ADMS, OMS, AMI, EAM and the historian, each carrying a timestamp, an owner and the date it was pulled. Any test with no artefact behind it scores zero and becomes part of what the verdict is about — it is not scored on the strength of the explanation.
- Assessment lane: one candidate, registered as a single named decision on a named asset boundary with a named decider, is scored across twelve tests in four dimensions. A standing panel — operations, OT security, the data owner, risk — issues one of four verdicts and attaches conditions with owners and dates.
- Authority lane: the consequence class is set before scoring begins, because it fixes the pass thresholds. The verdict is then a required input at three gates the utility already operates — capital release, the OT change and CIP applicability record, and the operational readiness review before go-live.
- The loop back: go-live stamps an expiry on the verdict and arms a withdrawal trigger. The tests that telemetry can measure are re-scored continuously against the same estate systems the original evidence came from, so a verdict cannot quietly outlive the network it described.
Step-by-step insights
- Pulled, not supplied — the difference the whole framework rests on
- The single highest-value design decision in a readiness framework is that the assessor holds standing read access to the operational estate. It sounds administrative and it is structural: an assessor who must request evidence can only ever grade the quality of the response, and the response is produced by the people whose programme is being assessed. Provision the read access once, as infrastructure, at the same time as you appoint the panel. Utilities that skip this step build excellent-looking dossiers that are, underneath, the same self-report the survey was — with the added disadvantage that everyone now believes the problem is solved.
- The unevidenced test — why zero, and not 'partially met'
- Scoring an unevidenced claim as partially met feels fair and is corrosive, because it converts the assessment back into a judgement of plausibility. Zero is not a punishment; it is a statement about what is currently knowable. It also has a useful practical effect: a dossier full of zeros is not an indictment of the team, it is a work list, and it is usually a much shorter one than anybody expected. The most common reaction to a first evidence-scored dossier is not defensiveness but relief, because for the first time the argument is about four specific missing artefacts rather than about whether the organisation is ready in general.
- Consequence class before evidence, always in that order
- The consequence class fixes the pass thresholds, so it has to be set before anyone knows how the evidence scored. Set afterwards, it quietly becomes the mechanism by which a desired answer is reached: a candidate that scored poorly is reclassified as low-consequence and passes. The class should name the worst credible outcome in the operator's own vocabulary — customer exposure, safety exposure, market exposure, regulatory reporting exposure — and state explicitly where none applies. A candidate whose worst credible outcome is a slightly worse inspection route is not the same artefact as one whose worst credible outcome is a published restoration promise during a storm, and it should not have to clear the same bar.
- Three gates, because they already exist
- Do not build a fourth gate. Utilities already operate the three that matter: the capital-approval record that releases money, the OT change record that grants connection and determines CIP applicability, and the operational readiness review that permits go-live. Each already has an owner, a calendar and a form, and each already knows how to refuse. Attaching the readiness verdict as a required field on those forms takes a fraction of the effort of standing up a new AI governance board, and it inherits authority the new board would spend two years trying to earn.
- Shadow only — the verdict most frameworks omit
- Binary ready/not-ready throws away the most productive answer available. A shadow-only verdict permits the candidate to run against live data, on the real cadence, writing nowhere, with its outputs logged and compared against what actually happened. It costs almost nothing, it is refusable by nobody because it touches no system of record, and it manufactures precisely the evidence the failed tests were asking for. For candidates blocked by a slow structural gap — a connectivity records programme, a CIP re-zoning, a metering rollout — shadow is not a consolation prize; it is the correct engineering answer and the fastest route to a later pass.
- The loop back — why standing is a maintenance discipline
- A verdict describes an estate at a moment. The estate then changes without telling anyone: re-conductoring, a new distributed-generation cluster, a meter firmware rollout, an ADMS vendor upgrade, a boundary redraw after reconfiguration. The re-score loop exists because those changes are invisible to a governance calendar and visible to telemetry. Pick the two or three tests that a machine can evaluate — feed freshness against SLA, reconciliation exception rate, scope drift against the in-scope asset list — and let them run weekly. The rest can wait for the interval. This is exactly the pattern the estate already applies to protection settings: some things are continuously monitored, others are verified on a maintenance interval, and nothing is assumed to hold indefinitely.
None of this machinery is novel to a utility, which is the argument worth making internally. The industry already runs the most rigorous readiness discipline of any sector: factory acceptance testing before equipment leaves a supplier, site acceptance testing before it is accepted, commissioning before energisation, protection-settings verification, and certification intervals for the people who operate the result. The gap is not that utilities lack a readiness culture. It is that AI candidates have been arriving through a channel where none of that discipline is attached, and the fix is extension rather than invention.
The five stages of readiness-framework maturity in detail
For each stage: what the instrument actually looks like in use, the signals a reviewer can check in an afternoon, the anti-pattern that traps organisations there, and what leaving costs.
Each stage below describes the instrument rather than the estate. That distinction matters when you place yourself: a utility with sophisticated AI in production can still sit at stage 2 on this ladder, because the way it decides whether the next candidate may proceed is a self-scored questionnaire. The hallmarks describe observable conditions, the diagnostic signals are checks you can run this week against your own records, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is present in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Opinion
21% of operators sit here
Readiness is whatever the person you ask believes it is — no instrument, no evidence file, and no record that a readiness decision was ever made.
Stage 1 is rarely ignorance, and it is almost never a shortage of assurance culture. The same utility that cannot say whether it is ready for a machine-learning model will not energise a substation without a signed site acceptance test, will not change a relay setting without a protection review, and will not put an operator on shift without a certification record. The discipline exists in abundance. The problem is that the AI candidate arrived through the innovation door rather than the engineering door, and nothing at that door asks for evidence.
The observable tell is asymmetric confidence. Ask three senior people whether a named candidate is ready and you get three confident, incompatible answers, each correct about the part of the estate that person can see. The asset manager knows the inspection history is digitised; the GIS lead knows the connectivity model has an unquantified error rate; the security lead knows nobody has determined whether the candidate lands in a CIP-scoped system. None of them is wrong. There is simply no frame that can hold all three answers at once and score them.
This is a cheap stage to leave and an expensive stage to stay in, because at stage 1 every AI proposal is re-litigated from first principles and none of the argument amortises. The first instrument does not need to be sophisticated; it needs to exist and be written down. A single page naming one candidate decision, the evidence that would have to exist for it to be ready, and where that evidence would come from, is already a stage-2 artefact — and it takes an afternoon, not a programme.
In practice
The three answers to one question
At a mid-sized distribution utility, an executive asked whether the business was ready to let a model prioritise wood-pole inspections. The asset manager said yes: condition data from the last three inspection cycles has been in the EAM for years. The GIS lead said no: pole-to-feeder connectivity has a known error rate on the rural circuits and nobody has ever counted it. The head of IT said it depended on which vendor was chosen. The meeting ended with a decision to proceed, taken on the strength of the most senior voice in the room. Six months later the pilot's biggest problem was rural connectivity error — the answer that had been in the room the whole time and had nowhere to be recorded.
What it looks like
- "Are we ready?" is answered in a meeting, differently by each person asked
- No written statement of what ready would mean for any specific decision
- Vendor demonstrations and reference visits are treated as readiness evidence
- No record exists of a readiness decision being made, deferred or reversed
Diagnostic signals you can check this week
- Ask three people whether one named candidate is ready, separately, and compare the answers
- Ask to see the last readiness decision. If there is no document, there was no decision
- Ask what evidence would change the answer. Silence means the question is not currently answerable
- Check whether any AI proposal has ever been deferred or refused on readiness grounds, ever
Anti-pattern · Commissioning a readiness study
The instinctive fix is to buy a readiness assessment: six weeks, an external team, a scored report with a radar chart at the front. It reliably produces a stage-2 artefact rather than a stage-3 one, because an external team brought in for six weeks has one realistic source of information — interviews with the people who own the systems. The output is therefore a well-presented self-report. Start at the other end: one candidate, one page of required evidence, and a genuine attempt to pull three of those artefacts yourself. What you learn from failing to pull them is worth more than the radar chart.
What holds you here
There is no instrument, so readiness cannot be argued about — only asserted, and the most senior assertion wins.
Highest-leverage next move
Pick one named candidate decision and write down, on one page, what evidence would have to exist for it to be ready and which system that evidence lives in. Do not score anything yet.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- One assessor with read access to the estate, one operations owner, part-time
- Risk
- Low — nothing in production changes; the only cost is the honesty of the first verdict
- To next stage
- 1–3 months
If this is you, the next step is
Two weeks: one candidate, the evidence pulled from your systems, one scored verdict.
Stage 2
Survey
44% of operators sit here
Readiness is a self-scored questionnaire about the organisation, producing a number that ranks nothing and can refuse nothing.
Stage 2 is the most dangerous stage on this ladder, because it produces an artefact that looks like governance and behaves like a newsletter. The survey has a scale, a scoring rubric, a comparison against last year and a slide that goes to the board. It is genuinely useful for one thing — opening a conversation about which parts of the business feel furthest behind — and useless for the thing it is invariably used for, which is deciding whether a specific piece of work should proceed.
Two structural defects cause this, and neither is fixable by improving the questionnaire. The first is self-report: the team most invested in a favourable answer supplies the answer, and does so without malice, because everyone scores their own house against the standard they know rather than the standard the assessor means. The second is resolution. An organisational readiness score cannot approve or refuse a project any more than an average annual temperature can tell you whether to take a coat out today. The unit of assessment and the unit of decision do not match.
Time at stage 2 is not neutral; it inoculates the organisation against the next attempt. Once the survey has run twice and changed nothing, the third round is completed carelessly, the scores drift upward for reasons nobody can name, and 'readiness assessment' acquires a reputation as an administrative exercise. Utilities that have run three annual readiness surveys are usually harder to move to stage 3 than utilities that have never run one, because the word itself now carries a meaning you have to argue against before you can start.
In practice
The 68% that approved everything
A utility ran an enterprise AI-readiness survey across eleven functions and scored 68%, presented to the executive as 'amber, improving'. The weakest single item, at 2 out of 5, was customer-to-asset linkage in GIS. Over the following year the same governance forum approved every AI proposal that reached it, including a customer-impact ranking model whose entire value depended on that linkage. Nobody connected the two documents, and nobody was being careless: the survey scored the organisation, the approvals were about projects, and there was no mechanism by which a number about the first could bind a decision about the second.
What it looks like
- A scored readiness questionnaire exists and is completed by the teams being assessed
- The output is an organisational percentage, a heat map or a radar chart
- Scores attach to the enterprise or a programme, never to a named decision
- No project has ever been stopped, reshaped or resequenced because of the score
Diagnostic signals you can check this week
- Ask who supplied the answers. If it was the team that owns the system, the score is a self-report
- Look for a project the score stopped or resequenced. Its absence is the finding
- Check the unit of assessment: organisation, function, or a named decision on a named asset set
- Ask what a score of 3 means in artefacts. If nobody can answer, the scale is decorative
Anti-pattern · Adding dimensions to fix the survey
When the survey does not change behaviour, the reflex is a better survey: more dimensions, weighted scoring, a peer benchmark, a maturity model with named levels. This makes the artefact more impressive and no more binding. Authority is not a property of instrument resolution; it is a property of attachment. A crude four-test dossier that is a required field on the capital-approval form outranks a forty-question weighted model that lives in a slide deck. Spend the next quarter on attachment, then improve the instrument once it has somewhere to land.
What holds you here
The score is self-reported and attached to the organisation, so it cannot rank candidates against each other or refuse any one of them.
Highest-leverage next move
Re-run the assessment on ONE named candidate decision, scoring only the tests whose evidence you can pull yourself from GIS, ADMS, OMS, AMI, EAM or the historian.
Cost of leaving
- Effort
- 3–6 months
- Team
- One assessor with read access to GIS, ADMS, OMS and EAM; one operations owner; one OT security reviewer
- Risk
- Medium — the first evidence-based verdict will contradict a self-report someone has already presented upward
- To next stage
- 3–6 months
If this is you, the next step is
We take one live candidate from your portfolio and re-score it on evidence you can pull.
Stage 3
Dossier
24% of operators sit here
Readiness is an evidence file per candidate, scored by named assessors against artefacts pulled from the estate rather than supplied by the delivery team.
Stage 3 is the first stage at which a readiness answer can be checked by somebody who was not in the room. The dossier's power is not its score; it is that the score is reconstructable. Six months later a different assessor can pull the same freshness report, count the same reconciliation exceptions, read the same interface record, and either agree or produce a specific disagreement. That property — reconstructability — is exactly what a utility's commissioning records already have, and it is the property that converts readiness from a conversation into an engineering artefact.
The character of the argument changes here, and the change is worth naming because it is the main practical benefit. Stage 1 and stage 2 disputes are about opinion and method, and they do not terminate. Stage 3 disputes are about artefacts, and artefact disputes terminate quickly: either the ninety-day freshness report exists for the named feeds or it does not; either the GIS-to-ADMS reconciliation exception count for the in-scope feeders has been produced or it has not. Meetings get shorter. Teams stop preparing rhetoric and start preparing evidence, which is a much cheaper thing to prepare.
The constraint that emerges at stage 3 is authority, and it is the reason most good frameworks stall here. The dossier is honest, well-researched and advisory. Proposals that fail it still get funded, because funding happens in a different room, on a different calendar, chaired by people who never see the dossier. What eventually kills the practice is not disagreement but timing: the dossier arrives after the money, at which point its only remaining function is to be filed as a risk item and referenced in a post-incident review.
In practice
The dossier that arrived after the money
A network operator built a genuinely rigorous dossier for an AI-assisted peak-load forecasting candidate: ninety days of AMI head-end feed freshness with the gaps counted, a GIS-to-ADMS reconciliation exception rate on the in-scope primary substations, a named decider in network planning, and a written statement that no CIP-scoped system was in the write path. It scored the candidate not-ready on two structural tests. The capital committee had approved the programme six weeks earlier on a business case that never mentioned data. The dossier was accepted, filed, and the work proceeded on the original schedule. The framework was excellent; it was simply on the wrong side of the calendar.
What it looks like
- Each candidate has its own dossier naming scope, decider and consequence class
- Every scored test names the artefact that evidences it and where it was pulled from
- Unevidenced claims score zero rather than being scored on the strength of the explanation
- A standing panel signs the verdict — operations, OT security, the data owner and risk
Diagnostic signals you can check this week
- Check whether any dossier was produced before, rather than after, a funding decision
- Ask whether a not-ready verdict has ever changed a funding outcome, and find the record
- Look at where each artefact came from: pulled by the assessor, or supplied by the delivery team
- Count the candidates in flight and the candidates with dossiers, and compare the two numbers
Anti-pattern · Making the dossier heavier to earn attention
When the dossier is politely ignored, the response is to make it more thorough — more tests, deeper analysis, a longer report, an appendix of evidence. Weight is not authority, and a heavier dossier is read by fewer people. The fix is constitutional and calendrical rather than editorial: get the readiness verdict named as a required field on the artefacts that already control money and OT access, then make the dossier shorter so that the people who must now read it actually do.
What holds you here
The dossier is advisory. It informs decisions that are actually made elsewhere, on a different calendar, by people who never see it.
Highest-leverage next move
Get the readiness verdict named as a required input on the artefacts that already control money and OT access — the capital-approval record, the change record, and the operational readiness review checklist.
Cost of leaving
- Effort
- 6–12 months
- Team
- Standing assessor panel with a chair outside the delivery line, plus the owner of the capital process
- Risk
- Medium — the first refusal is a political event and has to be survivable by design, not by luck
- To next stage
- 6–12 months
If this is you, the next step is
We map your capital, OT change and go-live gates and draft the readiness clause each one needs.
Stage 4
Gate
9% of operators sit here
Readiness is a binding condition: no verdict, no funding release, no OT connection and no go-live, with conditions tracked to closure by name and date.
Stage 4 is where the framework acquires the only property that ultimately matters: the ability to say no and be obeyed. Everything before this is diagnosis. The mechanics are trivial — a mandatory field on a form, a required attachment on a change record, a line on the operational readiness review checklist. The difficulty is entirely constitutional: somebody has to be willing to hold the first refusal against a sponsor who has already announced the programme.
What makes it survivable is that a utility already runs gates of exactly this shape, and everyone in the building accepts them. Nobody argues that a substation should be energised without a site acceptance test record, or that a protection setting should change without review, or that an operator should take a desk without a current certification. The AI gate is not a new kind of authority being invented; it is the existing kind, extended to a new class of change. Framing it that way — in the vocabulary of commissioning rather than the vocabulary of digital governance — is worth more than any amount of framework design, because it moves the argument from 'why are you slowing us down' to 'which acceptance test applies here'.
The failure that emerges at stage 4 is condition inflation. Refusal is expensive and visible, conditional passes are cheap and invisible, so panels drift toward passing everything with a list of conditions attached. Within a year the register holds dozens of open conditions, the median age is measured in quarters, and a conditional pass has become an unconditional one with paperwork. The metric that matters at this stage is not the pass rate; it is the condition closure rate and the age of the oldest open condition.
In practice
The condition register that became the programme plan
A utility that had wired readiness into its capital gate found that in the first year eleven of thirteen candidates passed conditionally, and that by month nine the register held forty-one open conditions with a median age of five months. The gate was technically binding and practically permissive. The panel introduced two rules: no candidate may carry more than three open conditions, and no team with an overdue condition may submit a new candidate. Submissions fell by a third in the following quarter and closure rates rose sharply — not because the rules were clever, but because closing a condition finally competed with starting something new.
What it looks like
- The capital-approval record cannot be completed without a readiness verdict reference
- The OT change record requires the dossier's CIP applicability determination
- Conditional passes carry named owners, due dates and the evidence that closes them
- At least one candidate has been refused in writing, and the refusal held for a full cycle
Diagnostic signals you can check this week
- Try to submit a capital request without a readiness reference and see whether the form permits it
- Read the condition register: count overdue conditions and measure the age of the oldest
- Find the last refusal and check whether it held for a full budget cycle or was reversed
- Ask whether a conditional pass has ever been converted back to a refusal for non-closure
Anti-pattern · Governing the pass rate instead of the closure rate
Refusals are visible and conditions are not, so panels optimise for the number they are asked about in review. A ninety per cent pass rate with a thirty per cent condition-closure rate is a materially worse position than a sixty per cent pass rate with everything closed, but only the first one looks like a functioning gate on a slide. Report closure rate and oldest-open-condition age at the same altitude as pass rate, and the incentive corrects itself without anyone having to be brave.
What holds you here
Verdicts bind at the moment of approval and then never expire, so a pass granted against last year's estate keeps authorising this year's system.
Highest-leverage next move
Stamp every verdict with an expiry date and a named re-assessor, and write the withdrawal criteria before go-live rather than during the first incident.
Cost of leaving
- Effort
- 12–18 months
- Team
- Standing panel, the owner of the capital process, the OT change board chair, and an executive sponsor prepared to hold one refusal
- Risk
- Higher — the real cost is the first held refusal, and it is paid in relationships rather than money
- To next stage
- 12–24 months
If this is you, the next step is
We run one of your real candidates through your gate and show you exactly where it leaks.
Stage 5
Standing
2% of operators sit here
Readiness is a live, expiring status: verdicts carry intervals, the measurable tests re-score from telemetry, and a stated trigger withdraws the decision from service.
Stage 5 is narrower than it sounds, and the narrowness is the point. It is not continuous assurance of everything; it is a small, enumerated set of live decisions whose readiness is maintained the way a utility already maintains relay settings and operator certifications — valid for a stated interval, re-earned rather than assumed, and revocable by a named person without a committee. Anything outside that enumerated set stays at stage 4, reviewed on a calendar, which is a perfectly respectable place for most of a portfolio to sit permanently.
The engineering is largely already in place by the time an operator reaches here: the feeds are monitored, the decision log exists, the revert route has been exercised. The genuinely hard part is deciding, in advance and in numbers, what 'no longer ready' means — and then accepting that the number will sometimes fire when nobody wants it to. A withdrawal trigger that has never fired is either an extraordinarily stable estate or a threshold set where it cannot reach. The second is far more common, and the way to tell them apart is to look at how the threshold was chosen.
Stage 5 is also the stage most likely to regress, and it regresses in silence rather than with an alarm. Expiries get extended in a governance meeting because nothing appears to have changed. Thresholds get widened after a run of alerts that turned out to be seasonal. The withdrawal authority migrates from a duty manager who is on shift to a committee that does not meet during a storm. None of these is a decision anybody would defend if it were written down as a decision; they happen because standing is administratively inconvenient and nobody is measuring its decay.
In practice
The expiry that caught the re-conductoring
A utility running AI-assisted estimated time of restoration held a twelve-month expiry on the verdict and one telemetry test that re-scored weekly: the GIS-to-ADMS reconciliation exception rate on the in-scope feeders. A re-conductoring programme changed connectivity across part of the operating region faster than the records were updated, and the exception rate crossed its threshold eight weeks before the annual review would have looked. The ETR model was returned to shadow for that operating area only, restored six weeks later once the records caught up, and the rest of the region ran throughout. The trigger did not prevent a failure so much as convert an invisible one into a scheduled one.
What it looks like
- Every live AI-supported decision has a readiness expiry date and a named re-assessor
- The tests telemetry can measure are re-scored continuously, not at an annual review
- Written withdrawal criteria exist and name the authority who can invoke them at 3am
- The withdrawal has been exercised — as a drill or for real — within the last twelve months
Diagnostic signals you can check this week
- Ask any live AI-supported decision for its readiness expiry date. If there is not one, the standing is fictional
- Check whether any expiry has ever been extended without re-assessment, and how routinely
- Read the withdrawal criteria and identify the person who can invoke them during a storm at 3am
- Look for evidence that the withdrawal has actually been exercised in the last year
Anti-pattern · Treating the expiry as an administrative date
Expiries get rolled forward in a governance meeting on the grounds that nothing has changed, and the sentence is almost always sincere and almost always wrong. The estate changes continuously and asymmetrically: a re-conductoring programme, a new distributed-generation cluster, a meter firmware rollout, a vendor upgrade to the ADMS, a boundary change after a network reconfiguration. None of those is announced to the readiness panel. Roll an expiry forward only on a re-score of the tests telemetry can measure — which takes minutes when the telemetry exists, and is the whole reason to build it.
What holds you here
Standing decays silently — the estate changes faster than the verdicts describing it, and nothing fires unless someone built the trigger and defended its threshold.
Highest-leverage next move
Treat the readiness verdict as a maintained artefact with an interval, an owner and a test record — the same discipline the estate already applies to protection settings and operator certification.
Cost of leaving
- Effort
- Continuous
- Team
- Standing panel plus a telemetry owner; roughly a day a month per live decision
- Risk
- Concentrated — low frequency, high consequence, and the failure mode is silence rather than an alarm
If this is you, the next step is
We test one live decision's expiry, telemetry and withdrawal trigger against a real scenario.
Where utilities actually sit on the readiness ladder
The distribution across the five stages, and why the survey stage is both the mode and the hardest one to leave.
Most utilities sit at stage 2, with a readiness instrument that is a self-scored organisational survey. The distribution below is weighted heavily toward that stage for a structural reason rather than a cultural one: a survey is the artefact that a governance function can produce without operational read access, and operational read access is exactly what an assessor needs and is rarely granted by default.
Distribution of utilities across the five readiness-framework stages
Illustrative distribution. The shape — a large survey-stage mode, a thin tail past the point where verdicts start to bind — is consistent with the gap between reported AI activity and reported AI governance depth in the cross-industry research linked below. The bar to watch is stage 4: everything before it diagnoses, only stage 4 decides.
Share of utilities
- 21% — 1 · Opinion
- 44% — 2 · Survey (the plateau)
- 24% — 3 · Dossier
- 9% — 4 · Gate
- 2% — 5 · Standing
Source: Illustrative distribution, synthesised from IEA, EPRI and cross-industry AI-governance research
The gap between activity and binding governance is not a utility-specific failure; it is the shape of AI adoption everywhere, and the energy sector's version of it is documented in the IEA's Energy and AI (opens in a new tab) analysis, which describes widespread pilot activity alongside a much narrower set of deployments that have reached operational systems. EPRI's research programmes (opens in a new tab) report the same pattern from the utility side: a great deal of applied research and piloting, and a much smaller population of deployments that have cleared the assurance route into an operational system. What is specific to utilities is not the size of that gap — it is that the machinery to close it is already installed and being used for other things, which is why the transition can be fast once somebody frames it as an acceptance test rather than as new governance.
| Stage | Can it rank two candidates? | Can it refuse one? | Can it lapse? | What the output is really for |
|---|---|---|---|---|
| 1 · Opinion | No — no common scale exists | Only by seniority | Not applicable | Settling an argument in one meeting |
| 2 · Survey | No — the score is organisational | No — nothing attaches to a project | No — it is re-run, not renewed | Reporting progress upward |
| 3 · Dossier | Yes — candidates share a scale and a test set | In principle; in practice the funding is elsewhere | No expiry is set | Informing a decision made in another room |
| 4 · Gate | Yes, and the ranking drives sequencing | Yes — the form will not complete without a pass | No — the pass, once granted, persists | Releasing money, OT access and go-live |
| 5 · Standing | Yes, continuously, as evidence moves | Yes, including after go-live | Yes — expiry, re-score and withdrawal trigger | Keeping a live decision authorised, or pulling it |
The drop from stage 3 to stage 4 is the largest single transition loss on this ladder, and it is not a technical one. Stage 3 costs analytical effort; stage 4 costs political capital exactly once, at the first refusal that holds. Utilities that have made the transition almost always describe the same enabling move — framing the readiness verdict as an acceptance test rather than as governance — because acceptance tests are already understood, already resourced, and already permitted to say no.
The readiness dossier: twelve tests and the artefact that proves each one
The instrument itself, printed. Four dimensions, twelve tests, and for each one the artefact a reviewer pulls, the system it comes from, and what a pass looks like.
The dossier is twelve tests across four dimensions, and every test names a specific artefact rather than a quality. That is the entire design principle: a test that asks "is the data good enough?" cannot be scored, while a test that asks "produce the ninety-day freshness record for these four named feeds, with the gaps counted" either can be satisfied or cannot. Twelve is not sacred, but it is close to the practical ceiling — beyond about fifteen tests, dossiers stop being read by the people whose signature makes them binding.
| Dimension | Test | The artefact a reviewer pulls | Where it comes from | What a pass looks like |
|---|---|---|---|---|
| Evidence basis | E1 · Input currency | Ninety-day freshness record for each named input feed, with gaps counted, against a stated SLA | Historian · AMI head-end · OMS event bus | Freshness measured, not asserted, and inside the decision's own cadence |
| Evidence basis | E2 · Network model currency | GIS-to-ADMS connectivity reconciliation exception count for the in-scope feeders | GIS · ADMS | Exception rate quantified and below the threshold set for the consequence class |
| Evidence basis | E3 · Label provenance | Record of where outcome labels come from, who validates them, and the dispute rate in review | EAM work orders · OMS restoration records | A named validator and a measured dispute rate, not an assumption that records are truth |
| Scope resolution | S1 · Decision statement | One page naming the decision, the system of record, the decider and the cadence | Candidate register | One decision, one decider. Two deciders means two candidates |
| Scope resolution | S2 · Consequence class | Written classification of the worst credible outcome, with safety-case applicability determined | Risk register · safety case | Classified before scoring began, with thresholds set by class |
| Scope resolution | S3 · Asset boundary | In-scope asset or feeder list exported from GIS, with the exclusions stated explicitly | GIS · candidate register | A list a reviewer can count, not a description of a region |
| Control path | C1 · Write path | Interface record naming the target field, the protocol and the change board that owns it | ADMS · OMS · EAM interface catalogue | The field a person actually reads is named, and the change board has seen it |
| Control path | C2 · Revert route | Test record for the revert to the previous decision source: date, duration, who ran it | Change record | Exercised, timed and signed — not described in a design document |
| Control path | C3 · OT security scope | CIP applicability determination with zone and conduit boundaries for the write path | OT security · network segmentation record | Determined before the funding gate, not discovered at connection |
| Standing and renewal | R1 · Expiry | Verdict record carrying an expiry date and a named re-assessor | Candidate register | A date and a name, both of which someone will still recognise in a year |
| Standing and renewal | R2 · Re-score telemetry | Definition of which tests re-score automatically, from which metric, against which threshold | Monitoring configuration | At least two tests machine-readable and running weekly |
| Standing and renewal | R3 · Withdrawal trigger | Written withdrawal criteria in the operating documentation, naming the invoking authority | Operating procedure · ORR record | Invocable by a person on shift, without convening anybody |
Two of the twelve carry disproportionate weight, and both sit in the control-path dimension. C2, the tested revert route, is the test that decides whether operations will accept the change at all — a written revert plan is a design intention, whereas a timed and signed revert test is the thing an operations manager will actually stake a shift on. C3, the CIP applicability determination, is the test whose absence most often stops a funded programme dead: applicability discovered at connection time typically costs two quarters, because the answer arrives after the architecture has been built around the assumption that it would not apply. Both are cheap to satisfy early and extremely expensive to satisfy late.
The scoring convention matters as much as the tests. Score 0 where no artefact exists; 1 where an artefact exists but was assembled for the assessment; 2 where it is produced by a running process but is not monitored; 3 where it is produced by a running process and a threshold on it is monitored. That scale rewards the thing you actually want — evidence that generates itself — and it makes the difference between a dossier that has to be rebuilt for every candidate and one where the second candidate inherits most of its evidence for free. By the third or fourth candidate on the same part of the estate, a well-built evidence pipeline reduces the assessment to the tests that are genuinely candidate-specific, which is usually four of the twelve.
One test is deliberately absent, and it is worth saying why. There is no test for model accuracy. Accuracy is a property of a model, measurable only once the model exists, and readiness is a question asked before the build is funded. A candidate with no model yet can score twelve out of twelve, and should: the framework is asking whether this estate can carry this decision, not whether a particular model is any good. Accuracy belongs in the acceptance criteria the verdict's conditions set — measured against a holdout, in the operational units the decision moves — not in the readiness score itself.
Four verdicts, and the three gates that make them binding
Which verdict the evidence supports, why 'shadow only' is the most underused of the four, and where each verdict has to land to stop being advice.
A readiness assessment should return one of four verdicts, and which one depends on two variables only: how strong the evidence is today, and how quickly the missing evidence can be manufactured. Everything else — the consequence class, the thresholds, the panel's judgement — has already been applied by the time the dossier is scored. The matrix below is how the score becomes a decision, and it is deliberately simple, because a verdict rule that needs a paragraph of explanation will be applied inconsistently by the second assessor.
Deriving the verdict from the dossier
Plot the scored dossier by how much evidence exists today and how fast the gaps can be closed. Three of the four answers are not 'no' — but only a framework that can reach the bottom-left quadrant and stay there is worth running at all.
Conditional pass
- Gaps are real but closable inside the build window
- Conditions carry a named owner, a date and closure evidence
- Cap open conditions per candidate — three is a workable limit
- Non-closure converts the pass back to a refusal, in writing
Proceed
- The dossier already carries the evidence the gates need
- Conditions, if any, are acceptance criteria rather than prerequisites
- Stamp the expiry and arm the withdrawal trigger at go-live
- This is where a mature evidence pipeline pays for itself
Refuse, and publish the gap list
- Weak evidence and slow structural gaps — the honest no
- The output is a work list, not a judgement on the team
- A framework that has never reached this quadrant is a survey
- Refusal with a dated gap list is more useful than a conditional pass
Shadow only
- Most tests pass; one slow structural gap blocks the write path
- Run read-only against live data, logging outputs, writing nowhere
- The shadow run manufactures the evidence the failed tests wanted
- Re-score when the structural gap closes, not on a calendar
A verdict only means something where it lands. The three gates below already exist in every utility, already have owners and calendars, and already know how to refuse things — which is precisely why the readiness verdict should attach to them rather than to a new committee built for the purpose. Each gate consumes a different subset of the dossier, and the sequencing is not arbitrary: the CIP applicability determination has to clear the OT gate before an architecture is built that assumes it will not apply.
| Gate | What it releases | Dossier evidence it consumes | What a refusal looks like here | Typical signatory |
|---|---|---|---|---|
| Capital or investment approval | Funding, headcount and a place in the delivery plan | S1 decision statement · S2 consequence class · E1–E3 evidence basis | The request is returned with a dated gap list and no funding line | Investment committee chair |
| OT change and CIP applicability | Connection to the operational estate and a write path | C1 write path · C3 CIP determination · C2 revert route | No connection is granted; the candidate runs shadow or not at all | OT change board chair, with the CIP compliance lead |
| Operational readiness review | Go-live: the decision enters service and people rely on it | C2 revert test record · R1 expiry · R3 withdrawal criteria | Go-live is deferred; the shadow run continues and re-scores | Operations manager for the affected area |
| Standing review (after go-live) | Continued authorisation of a live decision | R2 re-score telemetry · condition closure register · scope drift | The decision returns to shadow, or is withdrawn from service | Named re-assessor, with the duty manager able to invoke withdrawal |
Regulators already operate readiness gates of exactly this construction, which is a useful precedent to have in the room. Ofgem's Strategic Innovation Fund (opens in a new tab) releases funding to GB network licensees in phases — discovery, alpha, beta and deployment — with progression to each phase requiring evidence produced in the previous one rather than an argument about intent. The same logic sits inside the price-control regime (opens in a new tab) more broadly: expenditure is justified against evidence and reviewed at defined points, not authorised once and left. A utility proposing an internal readiness gate is not importing an unfamiliar governance concept; it is applying the funding discipline its own regulator already applies to it.
You can apply at any stage of your project but only when each funding round is open. Round 5 innovation challenges: discovery, alpha or beta and deployment phase.


