Energy & UtilitiesAI Implementation & Best Practices
AI operator assist in energy and utilities control rooms: from alarm noise to supervised autonomy
AI operator assist is decision support that works inside a live electricity control room — grouping an alarm flood into one event, ranking probable causes, drafting the next best action during switching and restoration. It does not replace the certified operator who authorises the action. It changes how fast that operator can see, understand and decide.

Key takeaways
- Operator assist is an HMI problem before it is a model problem. A suggestion that appears on a second monitor is a suggestion that is not read during the ten minutes it matters, because the operator's eyes are on the alarm summary and the single-line diagram they are accountable for.
- You cannot assist an unrationalised alarm system. EEMUA 191, quoted by the UK Health and Safety Executive, sets a long-term average of no more than one alarm every ten minutes per operator in normal operation and no more than ten in the first ten minutes after a major upset; below that discipline, AI simply summarises noise faster.
- Trust fails in both directions and both destroy value. Over-trust turns a certified operator into a button-presser who stops verifying; under-trust turns a funded programme into an ignored panel. The design variable that separates them is how cheaply the operator can check the suggestion on the screen they are already looking at.
- The ladder runs Alarm noise → Rationalised → Assisted → Guided → Supervised autonomy, and rungs two and three are where the operational value actually appears. Rung five is a small, enumerated set of bounded actions — not an unattended control room.
- Measure the control room, never the operator. Desk-level and event-level measures — alarms presented per flood, time to first correct action, suggestion acceptance and override reasons, customer minutes lost on covered feeders — answer the funding question without creating a surveillance instrument that certified professionals will rightly refuse.
Abbreviations used on this page
- SCADA
- Supervisory control and data acquisition — the telemetry and control layer
- EMS
- Energy management system (the transmission control room's master application)
- ADMS
- Advanced distribution management system (the distribution control room's equivalent)
- OMS
- Outage management system — the incident and restoration record
- HMI
- Human-machine interface — the console the operator actually works in
- SLD
- Single-line diagram, the network mimic on the operator's screen
- RTU
- Remote terminal unit — the substation device that reports to SCADA
- EEMUA
- Engineering Equipment and Materials Users Association, publisher of EEMUA 191 on alarm systems
- ISA
- International Society of Automation, publisher of ANSI/ISA-18.2 alarm management
- NERC
- North American Electric Reliability Corporation — certifies system operators and sets reliability standards
- FLISR
- Fault location, isolation and service restoration
- CML / CI
- Customer minutes lost and customer interruptions — the GB distribution reliability measures
Free · 8 questions · ~3 minutes
Score your control room on the assist ladder
Eight questions, one at a time, about three minutes. Answer them and we build your personalised report — your rung on the ladder, your score on each of the four dimensions, and the specific blocker between your desk and the next rung — and send it to your inbox. Every question is about the desk and the console, never about an individual operator.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised control-room report is ready
Tell us where to send it. Your rung appears on screen straight away, and the full report — dimension scores, the blocker between you and the next rung, and a 90-day plan for one desk — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Alarm noise
The console is saturated: alarms without defined responses, standing and chattering points, and floods that the operator survives on experience alone.
Your next moveMeasure one desk against the EEMUA 191 bands and rationalise the worst forty alarm points. No AI in this phase at all.
Stage 2 · Rationalised
Every alarm has a defined operator response, rates are measured against a published benchmark, and floods are characterised rather than endured.
Your next moveAdd correlation inside the existing alarm summary: group a flood into one event and offer ranked cause candidates with their evidence, changing nothing about who decides.
Stage 3 · Assisted
AI adds interpretation inside the operator's own screens — flood grouping, cause candidates, 'what changed' summaries — while every decision stays exactly where it was.
Your next moveMove from interpretation to proposal: draft the switching or restoration sequence, checked against tags, permits and limits, for the operator to accept, edit or reject.
Stage 4 · Guided
The assist proposes a next best action — a switching sequence, a restoration order, a relief action — checked against constraints, and the operator authorises, edits or rejects it.
Your next moveDefine the narrow, enumerated envelope in which a bounded action may execute unattended, with a tested inhibit, and prove it in the simulator before it is armed on the network.
Stage 5 · Supervised autonomy
A small, enumerated set of bounded actions executes without per-action approval inside a versioned envelope, with the operator supervising, inhibiting and handling everything outside it.
Your next moveTreat the envelope as a versioned, reviewed safety artefact with named owners, scheduled revalidation and a drilled inhibit.
0 / 24
Alarm and signal quality
— / 6
HMI integration
— / 6
Trust and override design
— / 6
Operator outcome measurement
— / 6
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the control room, and in our experience it is almost never the model — it is where the output renders and how overrides are handled. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the control room, and in our experience it is almost never the model — it is where the output renders and how overrides are handled.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this checked against the desk itself?
We sit with a shift, read a month of alarm history and override logs, and walk your control engineering and operations leads through what the numbers say — then leave you with a costed 90-day plan for one desk. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–5 — Stage 1, Alarm noise. The console is saturated: alarms without defined responses, standing and chattering points, and floods that the operator survives on experience alone.
- 6–11 — Stage 2, Rationalised. Every alarm has a defined operator response, rates are measured against a published benchmark, and floods are characterised rather than endured.
- 12–16 — Stage 3, Assisted. AI adds interpretation inside the operator's own screens — flood grouping, cause candidates, 'what changed' summaries — while every decision stays exactly where it was.
- 17–21 — Stage 4, Guided. The assist proposes a next best action — a switching sequence, a restoration order, a relief action — checked against constraints, and the operator authorises, edits or rejects it.
- 22–24 — Stage 5, Supervised autonomy. A small, enumerated set of bounded actions executes without per-action approval inside a versioned envelope, with the operator supervising, inhibiting and handling everything outside it.
What AI operator assist is in an electricity control room
A definition, the line it must not cross, and the path an alarm actually travels from a fault to an authorised action.
AI operator assist is decision support that runs inside a live electricity control room and works on the operator's own screens: it groups an alarm flood into a single event, ranks probable causes with their evidence, projects when a limit will be reached, and — higher up the ladder — drafts the switching or restoration sequence for a human to authorise. It is not autonomy, and it is not a dashboard. It is a change to what a certified operator sees and how quickly they can act on it.
The line it must not cross is defined by accountability rather than by technology. A control-room operator is a certified professional carrying personal responsibility for the network: in North America through NERC's System Operator Certification (opens in a new tab) credentials, in Great Britain through the authorised-person regime under the safety rules, and everywhere through the plain fact that their name is on the switching programme. Assist may change what that person perceives, comprehends and can project. It may not quietly become the thing that decides, because there is no mechanism by which accountability follows the suggestion.
That constraint is what makes this an implementation problem rather than a modelling one. The interesting questions are where a suggestion renders, how fast its evidence can be reached, what it does when its inputs are bad, how an override is captured, and how you prove any of it worked without building a surveillance instrument. Where the suggestion is computed — an optimiser, a classifier, a physics model — matters far less than the six inches of screen it has to arrive on. If you want the optimisation methods themselves, they are the subject of the sibling page on AI grid and layout optimisation; this page starts where that page's answer meets a human being.
How an alarm becomes an authorised action, by rung
The same fault, three control rooms. The rung is set by where the interpretation happens: in the operator's head from raw rows, in the console alongside the rows, or inside a bounded envelope that acts and then tells the desk what it did. Most control rooms are in the top lane.
- Data & feeds
- Where value leaks
- Human in the loop
- AI / model
- System-of-record action
The process, in words
- At rungs 1–2 a single feeder fault arrives as hundreds of alarm rows carrying equal visual weight. The operator performs the correlation in their head, at speed, from rows — and the quality of the first action depends on which experienced individual happens to be on shift.
- At rungs 3–4 the same fault arrives against a rationalised alarm set. A correlation layer groups it into one event with ranked cause candidates, a suggestion service attaches provenance, confidence and an expiry, and the result renders inside the alarm summary and on the mimic. The operator authorises, edits or rejects — and every rejection returns to the system as a reason code.
- At rung 5 a short, enumerated set of bounded actions executes inside a versioned envelope with a tested inhibit. Every execution is annunciated on the console with alarm-grade prominence, and anything outside the envelope stops and escalates to the desk rather than proceeding on a guess.
Step-by-step insights
- Why one fault produces hundreds of rows
- An electricity alarm flood is not a coincidence of unrelated events; it is one physical event observed by hundreds of instruments. A single earth fault operates protection, changes switch positions, trips downstream reclosers, drops voltage at connected substations, alarms low-voltage monitors, and raises communications alarms as devices lose supply. Every one of those is a true, correctly configured alarm. This is precisely why suppression is the wrong first instinct and correlation is the right one: nothing here is noise in the signal-processing sense — it is one story told by four hundred witnesses in arbitrary order.
- The rationalised alarm set is the substrate, not the ambition
- Correlation only produces interpretable output if each row means something agreed. ANSI/ISA-18.2 gives the lifecycle — philosophy, identification, rationalisation, design, implementation, operation, maintenance, monitoring and management of change — and EEMUA 191 gives the design guidance and performance targets that the HSE quotes. The reason this matters to an AI programme is not compliance: it is that 'these forty alarms are consequential' is a claim that can only be checked if somebody once wrote down what each of those forty alarms was for.
- Provenance is the difference between a suggestion and an assertion
- A suggestion that cannot show its evidence will be checked once, found unverifiable, and ignored thereafter. Provenance in a control room is specific: which telemetry points, which switch positions, which protection operations, which measurements, and how old each of them is. Data age is the item most often omitted and the one operators ask for first, because a confident statement built on a three-minute-old scan during a fast-moving event is a different object from the same statement built on live values.
- Rendering inside the console is an assurance question too
- Putting anything on a live operator console is a change to a safety-relevant human-machine interface, which is why the HSE treats control-room and HMI design as human-factors topics in their own right, and why NERC CIP change control applies to systems in the cyber boundary. The practical consequence is a longer approval path than teams expect and a shorter development path than they fear: most of the elapsed time is assurance, training and simulator work, not engineering.
- The override edge is the learning loop
- The dashed arrow back from the operator to the suggestion service is the most valuable line on the diagram. Overrides with reason codes are how a programme discovers that its loading assumption is wrong, that a proposal ignores a customer commitment, or that its ordering is unsafe with current permits. Without that edge, a control-room AI programme has no evidence base and its next release is guesswork — which is why rung 4 is unreachable from a rung-3 system that never captured dismissals.
- Rung 5 is a list, and the inhibit is the thing to look at
- Supervised autonomy earns its name from what the operator retains: alarm-grade annunciation of every automated action, and a single inhibit that takes the scheme out of service without asking anyone's permission. If either is missing, the control room has traded away situational awareness for speed. The rate of escalations out of the envelope is the leading indicator to watch — a rise means the network has moved, and the envelope needs revalidating before an incident forces the point.
The five rungs in detail
For each rung: what it looks like on a real desk, the signals a reviewer can check in an afternoon, the anti-pattern that traps control rooms there, and what leaving costs.
Each rung below is written for a practitioner rather than a buyer. The hallmarks are observable conditions on a desk, the diagnostic signals are checks you can run against your own historian and change records this week, and the anti-pattern is the specific mistake most often made trying to leave that rung. The ladder is deliberately front-loaded: rungs 1 and 2 contain no AI at all and carry a large share of the total value.
Select a rung
Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Alarm noise
24% of operators sit here
The console is saturated: alarms without defined responses, standing and chattering points, and floods that the operator survives on experience alone.
Rung 1 is not incompetence; it is accumulation. Every protection scheme commissioned, every new RTU point, every 'add an alarm for that' action item from an incident review adds rows to the same list, and nothing ever removes any. Over fifteen years the alarm system stops being a set of things that require a response and becomes an undifferentiated event stream that the operator is expected to interpret in real time.
The tell is the language the desk uses. At rung 1 operators talk about 'clearing the list' rather than about responding to alarms, and they will tell you — accurately, and without embarrassment — which alarms they ignore. That knowledge is real expertise, and it is also an uncontrolled risk register: it lives in individuals, it is not written down, and it walks out of the building with every retirement.
AI cannot be added usefully here, and adding it makes things worse in a specific way. A model trained on this alarm stream learns the noise as if it were signal, and any assist rendered on top of it either repeats the flood in a new format or hides parts of it by rules nobody agreed. The prerequisite is not a data platform; it is the unglamorous rationalisation work that the process industries have had a standard for since the 1990s.
In practice
The eleven minutes nobody could have handled
The Health and Safety Executive's account of the 1994 Texaco Milford Haven explosion is the reference case the whole alarm-management discipline is built on: in the last eleven minutes before the explosion, two operators had to recognise, acknowledge and act on 275 alarms. The plant was chemical rather than electrical, but the arithmetic is identical on a distribution desk during a storm — the human channel has a fixed bandwidth, and no amount of professionalism raises it.
What it looks like
- Nobody can state the desk's steady-state alarm rate per operator without a special study
- Standing alarms are permanently present and mentally filtered out
- A single feeder fault produces hundreds of alarm rows in under a minute
- Which alarms matter is knowledge held by long-serving operators, not by the system
Diagnostic signals you can check this week
- Ask the desk what its alarm rate is per operator per ten minutes in normal operation. A shrug is the answer
- Count standing alarms at the start of a shift; anything above single figures means the list is decorative
- Look at the top ten most frequent alarm points over a month — chattering points usually dominate
- Ask whether every alarm has a defined operator response written down. At rung 1 the honest answer is no
Anti-pattern · Buying an AI alarm-reduction product first
The obvious fix looks like a product: install something that suppresses or clusters alarms with machine learning and the noise problem disappears. It does not. Suppression rules that nobody rationalised become invisible policy — an alarm that a human decided mattered is now hidden by a model nobody can interrogate at three in the morning, and the first time a genuine event is inside a suppressed cluster, the tool is switched off permanently and the credibility of every later system goes with it. Rationalise first; the AI has to sit on top of an agreed set of alarms, not replace the agreement.
What holds you here
The alarm set has never been rationalised, so there is no agreed signal for an assist to reason over — only an event stream.
Highest-leverage next move
Measure one desk against the EEMUA 191 bands and rationalise the worst forty alarm points. No AI in this phase at all.
Cost of leaving
- Effort
- 3–6 months
- Team
- One control engineer, one protection or SCADA engineer, and operator time from every shift
- Risk
- Low — the work removes and re-tunes alarms under change control; nothing new is added to the console
- To next stage
- 3–6 months
If this is you, the next step is
Two to three weeks: rates by shift, standing and chattering inventory, flood characterisation from the historian.
Stage 2
Rationalised
34% of operators sit here
Every alarm has a defined operator response, rates are measured against a published benchmark, and floods are characterised rather than endured.
Rung 2 is where most of the safety value on this whole ladder is actually released, and it involves no artificial intelligence whatsoever. Rationalisation is a documented method — ANSI/ISA-18.2 gives it a lifecycle, EEMUA 191 gives it design guidance and performance targets — and the effect on a saturated desk is immediate and measurable in a way that later, cleverer work rarely is.
It is also slow, and honest planning matters. The HSE's own guidance notes that a quick first-pass review may cover perhaps fifty alarms per shift while a thorough review and redesign may take more than one shift per alarm. On a desk with four thousand configured alarms that arithmetic is the entire programme plan, which is why serious utilities rationalise by risk-ranked tranches rather than attempting the whole estate.
What rung 2 produces for the AI work that follows is not data but agreement. After rationalisation, an alarm means something specific that operations, protection and engineering have all signed. That shared definition is what makes a correlation model's output interpretable — 'these forty rows are one event' is only a useful sentence if everyone agrees what each row was for.
In practice
Forty points, one storm, three fewer screens of scrolling
A distribution control desk pulled twelve months of alarm history from the historian and found that forty configured points produced more than half of all alarm rows, almost all of them chattering analogue crossings and out-of-service plant. Re-tuning deadbands, filtering the chatter and suppressing out-of-service plant under an expiring shelve reduced the steady-state rate below the EEMUA manageable band. No model was involved; the operators noticed within one shift.
What it looks like
- An alarm rationalisation record exists: purpose, cause, consequence, response and priority per alarm
- Steady-state rate per operator is measured continuously and reported like any other KPI
- Shelving and suppression are explicit, time-limited and logged, not informal
- Flood events are counted, timed and reviewed — the desk knows its worst ten minutes of the year
Diagnostic signals you can check this week
- Ask to see the rationalisation record for a named alarm. Rung 2 produces it in a minute
- Check whether shelved alarms have expiry times and whether anyone reviews the shelved list at handover
- Look for a monthly alarm-performance report that reaches an accountable manager, not just the SCADA team
- Ask how a new alarm gets added — a change-control route with an impact-on-load assessment is the rung-2 signature
Anti-pattern · Declaring rationalisation done and letting it decay
Alarm systems regress silently. Every commissioning, every temporary configuration left in place, every incident action item that adds 'an alarm so this cannot happen again' pushes the rate back up, and eighteen months later the desk is at rung 1 with a rationalisation record that describes a system nobody has operated for a year. The management-of-change requirement in the ISA-18.2 lifecycle exists precisely for this, and it is the part most often skipped once the initial project ends.
What holds you here
The alarm set is clean but nothing yet interprets it — the operator still assembles the picture from rows during the exact minutes when assembling is hardest.
Highest-leverage next move
Add correlation inside the existing alarm summary: group a flood into one event and offer ranked cause candidates with their evidence, changing nothing about who decides.
Cost of leaving
- Effort
- 6–12 months for a full desk, less per tranche
- Team
- Control engineer, protection engineer, SCADA configuration, plus scheduled operator workshops
- Risk
- Medium — removing an alarm is a safety-relevant change and needs the same rigour as adding one
- To next stage
- 6–12 months
If this is you, the next step is
We sequence tranches by contribution to load and by consequence, so the first tranche is felt on the desk.
Stage 3
Assisted
27% of operators sit here
AI adds interpretation inside the operator's own screens — flood grouping, cause candidates, 'what changed' summaries — while every decision stays exactly where it was.
Rung 3 is the first rung where AI earns its place, and its contribution is narrow and precise: it compresses the perception and comprehension work that Endsley's model of situational awareness describes as the first two levels, so that the operator can spend their attention on the third — projecting what happens next. Grouping four hundred alarm rows into 'one earth fault on this feeder, these are the consequential alarms, these three are not explained by it' is worth more than any prediction, because it hands back minutes at the exact moment the desk has none.
The engineering discipline that distinguishes a rung-3 system from a demo is provenance. Every statement the assist makes must be traceable to the telemetry that produced it, timestamped, and legible in one glance — which switch position, which protection operation, which measurement, and how old each of them is. An operator with personal accountability for a switching decision will not, and should not, act on an unattributed assertion, and a system that cannot show its evidence gets ignored inside a fortnight.
The other rung-3 requirement is silence. The assist must be able to say nothing. When telemetry quality degrades, when the state estimator has not solved, when the event does not match anything the model has seen, the correct output is an explicit 'insufficient evidence' rather than a low-confidence guess presented in the same visual language as a confident one. Abstention is a feature that operators notice and reward with trust.
In practice
Four hundred rows, one sentence
A regional desk takes an 11 kV fault during a storm. Before assist, the operator sees several hundred alarm rows across three screens and reconstructs the picture from experience. After assist, the top of the alarm summary reads: one event, circuit breaker operated at a named substation, protection indication consistent with an earth fault on a named section, twelve downstream low-voltage alarms explained by it, and two alarms not explained — highlighted, because unexplained alarms are the ones that matter. The operator still makes every decision. They make the first one four minutes earlier.
What it looks like
- An alarm flood arrives as one grouped event with a ranked list of candidate causes and their evidence
- The assist renders inside the alarm summary and on the single-line diagram, not in a separate application
- Every suggestion carries provenance, a confidence statement and an expiry
- Suggestion views, accepts and dismissals are logged at desk level with reason codes
Diagnostic signals you can check this week
- Watch an operator during a flood: if they leave the alarm summary to consult the assist, integration has failed
- Ask the assist to explain one suggestion. If provenance takes more than a glance to reach, it will not be checked under pressure
- Check whether the assist has ever declined to answer, and whether operators can recall it doing so
- Look for the override log. If dismissals are not captured with reasons, rung 4 has no evidence base to build on
Anti-pattern · Improving the model to fix adoption
When operators do not use the assist, the reflex is to make it more accurate. Adoption in a control room is overwhelmingly a function of location, latency and legibility rather than accuracy: a moderately good cause ranking on the alarm summary changes more decisions than an excellent one behind a second login. Before touching the model, measure how many keystrokes and how many seconds separate the operator from the suggestion and from its evidence — that number, not the F1 score, is what is capping use.
What holds you here
The assist explains the present but proposes nothing, so the operator still constructs every action sequence by hand under time pressure.
Highest-leverage next move
Move from interpretation to proposal: draft the switching or restoration sequence, checked against tags, permits and limits, for the operator to accept, edit or reject.
Cost of leaving
- Effort
- 6–12 months
- Team
- Integration engineer with SCADA/ADMS experience, ML engineer, control-room human-factors input, named operations owner
- Risk
- Medium — anything rendered on a live console is a change to a safety-relevant interface and needs its own assurance
- To next stage
- 9–18 months
If this is you, the next step is
Where the suggestion renders, how it is dismissed, and what happens when it is unavailable.
Stage 4
Guided
12% of operators sit here
The assist proposes a next best action — a switching sequence, a restoration order, a relief action — checked against constraints, and the operator authorises, edits or rejects it.
Rung 4 changes the artefact. Instead of describing the network, the assist proposes an action on it, and that crosses a line the industry has always taken seriously: a switching programme is a controlled document, prepared and checked by authorised people, and a restoration order commits real customers to real outage minutes. A proposal is legitimate at this rung only if it arrives inside the existing preparation workflow, carries its constraint checks visibly, and is trivially rejectable.
The evidence engine of rung 4 is the override log built at rung 3. Acceptance rate by proposal type is the honest measure of whether the guidance is useful, and the reason codes attached to rejections are worth more than the acceptances: 'unsafe with current permits', 'ignores a customer commitment', 'right answer, wrong order' each point at a different defect, and the distribution across a quarter is the most reliable roadmap a team will get.
Training is a hard requirement rather than a nice-to-have here. NERC's PER-005-2 already obliges reliability coordinators, balancing authorities and transmission operators to train system operators through a systematic approach, and requires emergency-operations training using simulation technology where interconnection reliability operating limits are in play. A new decision aid that appears on the live console without ever having appeared in the simulator is, in that framework, an untrained tool in a trained environment.
In practice
The restoration order that arrives already checked
After a fault is isolated, the ADMS proposes three restoration sequences, ranked by customers restored per step and annotated with the constraint each one nearly violates: option A restores 1,400 customers in two steps but loads a transformer to 96% of rating on the forecast evening peak; option B restores 900 with headroom; option C waits for a field confirmation. The control engineer picks B, edits one step, and the edit is logged. Three months later, that edit pattern — repeated by four operators — is what tells the team the loading assumption was wrong.
What it looks like
- Draft switching and restoration sequences appear in the tool that already builds them, pre-checked against tags and permits
- Each proposal states what it assumes, what it would achieve, and what it does not know
- Acceptance and override rates are monitored per proposal type, with reason codes reviewed monthly
- Operators are trained on the assist in the simulator, including on deliberately degraded behaviour
Diagnostic signals you can check this week
- Check whether the proposal is pre-checked against live tags, permits and earths, or whether the operator has to do that themselves
- Read a month of override reason codes; an empty or single-value field means nothing is being learned
- Ask whether the assist has ever been exercised in the simulator with degraded or wrong inputs
- Ask an operator what happens if they reject every suggestion for a shift. If the answer involves anyone being asked why, the design is already coercive
Anti-pattern · Measuring the operator instead of the desk
Once acceptance data exists, someone will propose per-operator dashboards: who accepts, who overrides, who is slower. It is the fastest way to destroy a control-room programme. Operators hold personal, certified accountability for their decisions, and an instrument that appears to grade their judgement converts a decision aid into a surveillance tool — after which the honest overrides stop and the reason codes turn into whatever is safest to type. Aggregate to desk, shift type and proposal type; write the prohibition into the measurement plan and agree it with the operators' representatives before go-live.
What holds you here
Every action still waits for an operator keystroke, so throughput in a mass-event scenario is bounded by how many decisions a desk can physically authorise.
Highest-leverage next move
Define the narrow, enumerated envelope in which a bounded action may execute unattended, with a tested inhibit, and prove it in the simulator before it is armed on the network.
Cost of leaving
- Effort
- 12–24 months
- Team
- Platform and integration team, control engineering, simulator and training lead, safety-rules authority, human factors
- Risk
- Higher — proposals touch controlled documents, so assurance, change control and training evidence become the binding constraints
- To next stage
- 18+ months
If this is you, the next step is
From telemetry to constraint check to console to override log — where the evidence gaps are.
Stage 5
Supervised autonomy
3% of operators sit here
A small, enumerated set of bounded actions executes without per-action approval inside a versioned envelope, with the operator supervising, inhibiting and handling everything outside it.
Rung 5 is much narrower than the phrase suggests. It is not an unattended control room; it is a handful of decisions — fault isolation and restoration on a defined feeder class, dispatch of a pre-optimised set of balancing instructions, a Volt/VAR setpoint inside stated limits — that execute without waiting for a keystroke, while everything else stays exactly where it was. Utilities that describe rung 5 well always describe it as a list, never as a capability.
The operator's role changes shape rather than shrinking. Supervision means seeing what the system did as clearly as what it proposes, holding a working inhibit, and retaining the authority to take manual control without asking anyone. If the automated action is not annunciated on the same console with the same prominence as an alarm, the control room has lost situational awareness in exchange for speed, which is the single worst trade available in system operation.
Sustaining rung 5 is governance, not engineering. Envelopes drift out of validity as the network changes — new generation, new customers, a reconfigured feeder — and the evidence that satisfied last year's assurance review is not automatically adequate this year. The discipline that keeps it honest is the same one NERC applies to real-time monitoring in IRO-018: you must be told when the machinery you depend on has stopped working, because silent failure of a trusted system is worse than a visible one.
In practice
Bounded, annunciated, inhibitable
A distribution operator arms automated fault isolation and restoration on a defined class of feeders. Each execution posts to the alarm summary as a first-class event — what operated, what it restored, what remains isolated — and any case that falls outside the envelope stops and escalates to the desk with the reason. The control engineer keeps a single inhibit that takes the whole scheme out of service and does not require anyone's permission to use. Escalation frequency is reviewed monthly: a rise means the network has moved outside the envelope, not that operators are being difficult.
What it looks like
- The autonomous action set is written down, short, and reviewed like a safety document
- An operating envelope states the conditions under which each action may execute, with a tested inhibit
- Every automated action is annunciated to the desk and reconstructable from the decision log
- The rate of actions falling outside the envelope is monitored as a leading indicator
Diagnostic signals you can check this week
- Ask for the list of actions permitted to execute unattended. It should be short and printed
- Ask when the inhibit was last exercised deliberately, and whether the exercise was recorded
- Check whether automated actions appear on the operator's console with alarm-grade prominence
- Ask who reviews the envelope, how often, and what triggers a review outside the cycle
Anti-pattern · Extending the envelope by analogy
The envelope is proven on one feeder class, works well, and is then extended to a class that was never in the evidence — different protection philosophy, different customer mix, different telemetry quality. The first bad automated action results in the entire scheme being disarmed, and the control room's tolerance for anything automated drops for years. New action types re-earn autonomy from their own evidence, in the simulator first, with their own envelope and their own inhibit.
What holds you here
Sustaining autonomy is an assurance discipline: envelopes age as the network changes, and the evidence has to be rebuilt rather than inherited.
Highest-leverage next move
Treat the envelope as a versioned, reviewed safety artefact with named owners, scheduled revalidation and a drilled inhibit.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team, control engineering, safety and assurance, plus a standing review forum including operators
- Risk
- Concentrated — low frequency, high consequence, and squarely inside the regulator's field of view
If this is you, the next step is
We run the envelope, the inhibit and the annunciation against a real scenario with your operators.
Where control rooms actually sit on the ladder
The distribution across the five rungs, and why the largest transition loss is between a rationalised alarm system and a genuinely assisted one.
Most control rooms are on rungs 1 and 2 — that is, in alarm-management territory rather than AI territory. The distribution below is illustrative rather than surveyed: it is our synthesis of what the published alarm-management literature, the human-factors regulators and the energy research houses describe, and it is drawn to make one point that every number in it supports — the mass of the industry sits below the rung where a model does anything.
Illustrative distribution of electricity control rooms across the five rungs
Illustrative, not surveyed. Rungs 1 and 2 dominate: the binding constraint in most control rooms is alarm quality and console integration, not model capability. The steepest drop is 2 → 3, where interpretation has to move from the operator's head into the console without adding anything to the screen.
Share of control rooms
- 24% — 1 · Alarm noise
- 34% — 2 · Rationalised (the plateau)
- 27% — 3 · Assisted
- 12% — 4 · Guided
- 3% — 5 · Supervised autonomy
Source: Illustrative distribution, synthesised from IEA, HSE and ISA/EEMUA alarm-management material
The external picture is consistent. The IEA's Energy and AI report (opens in a new tab) frames AI's grid contribution largely in terms of faster fault identification and better use of existing assets rather than new autonomy; the HSE's human-factors guidance (opens in a new tab) has treated alarm handling, control-room design and shift handover as first-order safety topics for two decades; and the alarm-management standards themselves — ANSI/ISA-18.2 (opens in a new tab) and EEMUA 191 (opens in a new tab) — exist because the industry established, expensively, that operator attention is the scarce resource. Research bodies including EPRI (opens in a new tab) continue to publish on control-centre modernisation for the same reason.
Alarm rationalisation and flood management: the work before the AI
The published benchmarks, why an electricity flood is one event told by hundreds of witnesses, and the lifecycle that keeps a rationalised system rationalised.
Alarm rationalisation is the process of deciding, alarm by alarm, whether a signal deserves to interrupt a human being — and if it does, what the operator is supposed to do about it, how urgent it is, and how long they have. It is the prerequisite for every rung above it on this ladder, because an assist that reasons over an unrationalised alarm set inherits its incoherence and presents it with more authority.
| Alarm rate per operator | Published band | What the operator can actually do | What assist may legitimately do | Rung |
|---|---|---|---|---|
| Under 1 per 10 minutes | Very likely to be acceptable | Read every alarm, respond to each on its merits, keep a mental model of the network | Add projection: time-to-limit, forecast constraint, quiet context for the next shift | 3–4 |
| 1–2 per 10 minutes | Manageable | Keep up in normal operation; slip during concurrent events | Group consequential alarms; annotate the mimic; draft the routine switching sequence | 3–4 |
| 2–10 per 10 minutes | Over-demanding | Triage rather than respond; some alarms are acknowledged without being read | Only interpretation — grouping and cause ranking. No proposals until the rate comes down | 3 |
| More than 10 per 10 minutes | Unacceptable | Silence and scroll; the alarm system has stopped functioning as an alarm system | Nothing on the console. The work is rationalisation, not inference | 1–2 |
| Flood after a major event | No more than 10 displayed in the first 10 minutes is the design target | Reconstruct the event from rows while making the first decisions | Group the flood into one event with ranked causes and explicitly unexplained alarms | 3 |
Those bands come from EEMUA Publication 191 (opens in a new tab), the alarm-systems guide the HSE quotes in its own guidance (opens in a new tab) on better alarm handling: a long-term average alarm rate in normal operation of no more than one every ten minutes, and no more than ten alarms displayed in the first ten minutes following a major upset. They were written for process plant, and they transfer to electricity control desks without amendment because the constraint they describe is the operator, not the plant.
1 / 10 min
Long-term average alarm rate per operator that EEMUA 191 treats as the design target in normal operation
EEMUA 191, quoted by HSE
10 / 10 min
Maximum alarms displayed in the first ten minutes after a major upset, as a design target
EEMUA 191, quoted by HSE
> 1 shift
Time a thorough review and redesign of a single alarm can take; a quick first pass covers about 50 per shift
HSE, Better alarm handling
In the last 11 minutes before the explosion the two operators had to recognise, acknowledge and act on 275 alarms.
The electricity-specific point is that a flood is not noise. One earth fault operates protection, changes switch positions, trips downstream devices, drops voltage at connected substations, alarms low-voltage monitors, and raises communications alarms as equipment loses supply. Every row is a true alarm about a real state change. That is why the correct first move is correlation — collapsing one event into one line with its consequences attached — and why blanket suppression is the move that eventually hides something that mattered.
The rationalisation lifecycle, applied to one distribution desk
Write the philosophy before touching a point
A short, signed document that states what an alarm is on this system, the priority definitions and roughly how they should be distributed, the performance targets taken from EEMUA 191 (opens in a new tab), and who may add one. The ISA-18.2 lifecycle (opens in a new tab) starts here for a reason: without it, rationalisation becomes an argument per alarm.
Measure before you change anything
Pull twelve months from the historian: rate per operator per ten minutes by shift, standing alarms at shift start, the top fifty most frequent points, and every flood over the design target with its duration. This baseline is the only thing that will later prove the work was worth doing.
Rationalise by risk-ranked tranche
Take the points that contribute most to load first. For each: purpose, cause, consequence of no action, the defined operator response, the time available, and the priority that follows from the consequence. The HSE's guidance is blunt about the effort involved — plan in tranches, not in one campaign.
Fix the mechanical causes
Deadbands on repeating analogue crossings, filters on chattering points, suppression of alarms from out-of-service plant under an expiring shelve, and removal of alarms with no defined response. This is where the rate drops fastest and where operators notice within a shift.
Put management of change around it
A new alarm requires the same route as any other change to a safety-relevant interface: a stated response, a priority derived from consequence, and an assessment of its effect on the desk's alarm load. Without this step the system regresses to rung 1 within two years, which is the single most common outcome of a successful rationalisation project.
Report performance where it will be read
Monthly alarm performance — rate, standing alarms, floods, top contributors — to an accountable operations manager rather than to the SCADA team alone. Alarm systems decay when nobody with authority is looking at the number.
Only after this is there something for a model to work with. The order is not a preference: a correlation layer trained on chattering points learns that they are informative, and a cause ranker built on alarms with no defined response ranks candidates nobody can act on. Rationalisation is the cheapest capability on the ladder and the one that makes every later pound of AI spend defensible.
The control-room moment map: where assist belongs and where it does not
Eight moments an operator lives through, the assist that is legitimate in each, the console surface it must render on, and the outcome it should move.
Assist belongs to moments, not to systems. A control room is not a continuous state; it is a sequence of quite different situations — quiet monitoring, planned switching, a fault, a storm, a handover — each with its own attention profile, its own time budget and its own tolerance for a machine offering an opinion. The map below is how we scope assist work with utilities: pick a row, deliver that row properly on one desk, and only then pick another.
| Control-room moment | What the operator is doing | Legitimate assist | Where it must render | Earliest rung | Outcome measure |
|---|---|---|---|---|---|
| Steady-state monitoring | Scanning, maintaining a mental model, absorbing routine alarms | Suppression of out-of-service plant, chatter filtering, shelving with expiry, quiet 'what changed' summaries | Alarm summary and the shelved-alarm list | 2 | Rate per operator per 10 min; standing alarms at shift start |
| Alarm flood after a fault | Triaging hundreds of rows while making the first decisions | Group the flood into one event; rank cause candidates with evidence; highlight alarms not explained by the leading candidate | Top of the alarm summary, as a grouping of existing rows | 3 | Alarms presented per flood; time to first correct action |
| Fault location and restoration | Deciding the isolation and restoration sequence | Ranked restoration options with customers restored per step and the constraint each nearly breaches | Single-line diagram annotation and the OMS incident | 4 | CML/CI or SAIDI/SAIFI on covered feeders; restoration time |
| Planned switching | Preparing and checking a switching programme | Draft step sequence pre-checked against live tags, permits and earths; conflicting-permit detection | The switching schedule editor already in use | 4 | Preparation time; defects found at check stage |
| Constraint and limit management | Watching a boundary or a transformer approach its rating | Time-to-limit projection with the assumptions stated, plus candidate relief actions | Trend and limit panel beside the value being watched | 3 | Time spent in alert state; number of exceedances |
| Storm and major-event operation | Running many concurrent incidents with degraded telemetry | Event clustering, damage prediction, prioritisation by customers and vulnerability, ETR support | Event list and geographic view | 3 | Events per operator; estimated restoration time accuracy |
| Shift handover | Briefing the incoming shift on abnormalities and intent | Auto-drafted handover pack: open abnormalities, shelved alarms with expiry, active suggestions and overrides | The handover record itself | 3 | Handover completeness audit; repeat questions after handover |
| Post-event review | Reconstructing what happened and why | Event timeline with decisions, data quality at the time, and the assist's own provenance | Event journal and historian replay | 3 | Time to reconstruct an event; findings per review |
Read down the fourth column and the page's central claim becomes concrete: every legitimate assist renders on a surface the operator already looks at. Nothing in the map is delivered as a new application, and the two rows with the highest stakes — restoration and planned switching — are the two that must arrive inside controlled documents the utility already governs. Where those proposals are computed by an optimiser rather than a learned model is an implementation detail covered on the grid and layout optimisation page; what matters here is that the output lands in the switching schedule with its constraint checks visible.
Compress, don't add
The first family of assist takes what is already on the screen and makes it smaller: grouping, deduplication, explanation. It is the only family that is unambiguously welcome during a flood, because it reduces the number of things demanding attention rather than increasing it.
Project, don't predict theatrically
The second family answers 'when', not 'what if': time to a rating, time to a limit, time until a constraint binds. Projection is what Endsley's third level of situational awareness describes, and it is the level operators lose first under load. A projection with its assumptions rendered alongside it is trusted; a bare number is not.
Propose, with the checks attached
The third family drafts an action. It is legitimate only when the proposal arrives pre-checked against the constraints the operator would check anyway — tags, permits, earths, ratings, customer commitments — and only when rejecting it costs one keystroke. A proposal that is expensive to reject is a proposal that will be accepted for the wrong reason.
Record, because someone will ask
The fourth family produces no output for the operator at all: it writes the decision log — every suggestion made, seen, accepted, edited or rejected, with reasons and data ages. It costs almost nothing to build alongside the other three and it is the only reason the programme will survive its first post-incident review.
What operator assist looks like in public
Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the operator's own published material.
The clearest public evidence for the argument on this page is in what system operators chose to build. In each programme below, the differentiator is not the sophistication of the model but the fact that the output arrives where the operator already works — a dispatch list inside the balancing tools, an automated restoration annunciated to the desk, a real-time renewable picture inside the national control centre — and that the surrounding discipline was built at the same time.
Three programmes read against the ladder
Outcomes as reported by the operators themselves; verify figures against the linked source before reusing them, as we have not independently audited them. The card images are generated library illustrations of control-room scenes, not photographs of these operators or their facilities, and no endorsement is implied.
National Energy System Operator (NESO)GB electricity system operator · national balancing control room34
- Challenge
- Balancing a system with rapidly growing numbers of small balancing mechanism units and battery storage sites meant control-room engineers had to instruct each asset individually, which limited how much of that flexible capacity could realistically be used second by second.
- Approach
- The Open Balancing Platform introduced Bulk Dispatch: the platform presents control-room engineers with a pre-selected, optimised list of units that meets a network requirement, and lets them instruct hundreds of smaller units and battery sites in a single action — inside the balancing workflow rather than as a separate analytical tool.
- Reported outcome
- NESO reports that the tool greatly reduces the time taken to instruct balancing mechanism units and the number of manual instructions required from the control room, with further stages delivered across 2024 and 2025 and the platform set to replace the existing balancing systems by 2027.
- What it shows about the curveThis is rung 4 in its clearest form: the machine narrows the option set and prepares the action; the control-room engineer still chooses and commits it. The value came from putting the optimised list inside the dispatch path, not from the optimisation being novel.
NESO — first stages of the Open Balancing Platform go live (opens in a new tab)
Duke Energy FloridaUS investor-owned utility · distribution network, hurricane-exposed45
- Challenge
- Storm restoration on a large distribution network is bounded by how quickly faults can be located, isolated and switched around — work that, done manually, competes with everything else a control room is doing during a major event.
- Approach
- Self-healing technology detects an outage and reroutes power automatically on equipped parts of the network, isolating the faulted section without waiting for a per-action decision from the control room — a bounded, enumerated action set operating on a defined class of circuits.
- Reported outcome
- Duke Energy reported that during Hurricanes Helene and Milton the technology prevented more than 300,000 customer outages and saved customers more than 300 million minutes of total outage time, describing it as isolating the cause of an outage and reducing the number of customers affected by up to 75%, often restoring power in less than a minute.
- What it shows about the curveRung 5 done properly is narrow and specific: one class of circuits, one action type, executed within an envelope. It is also the rung where annunciation and the inhibit matter most — an automated restoration the desk cannot see is a loss of situational awareness bought with speed.
Duke Energy — self-healing technology during Helene and Milton (opens in a new tab)
Red Eléctrica (Redeia)Spanish transmission system operator · national control centre23
- Challenge
- Integrating very high volumes of variable renewable generation requires the national control centre to know, continuously and in real time, what every significant renewable facility is doing — a situational-awareness problem long before it is an optimisation problem.
- Approach
- Cecre, the control centre for renewable energies, sits inside Cecoel, the electricity control centre, and receives real-time information every twelve seconds from generating facilities via the generators' own control centres, covering connection status, production and voltage at the connection point — then supports real-time analysis of the system state and the operating measures required.
- Reported outcome
- Red Eléctrica reports that Cecre allows the maximum renewable production to be integrated while maintaining supply quality and security, and states that in 2024 it integrated more than 98% of renewable generation at peninsular level.
- What it shows about the curveSituational awareness is architecture, not decoration. The rung-3 lesson is that the assist layer is only as good as the telemetry cadence and the model beneath it — and that placing the renewable picture inside the existing control centre, rather than beside it, is what makes it operational.
Red Eléctrica — Cecre, the control centre for renewable energies (opens in a new tab)