Redefining Technology

Energy & UtilitiesReadiness & Transformation Roadmap

The AI readiness talent gap in utilities: closing both sides of the workforce equation

The AI readiness talent gap in utilities is the distance between the workforce a utility has and the workforce its AI programme needs — and it runs in both directions: scarce data and machine-learning skills flowing in too slowly, and decades of tacit grid knowledge retiring out faster than anyone is capturing it. Closing one side without the other closes nothing.

Utility engineers and data specialists working together in front of grid monitoring displays in a network operations room
Energy & Utilities · Readiness & Transformation Roadmap

Key takeaways

  1. The AI readiness talent gap in utilities is two-sided: data and ML skills the sector cannot hire fast enough, and tacit grid knowledge — transformer judgement, switching intuition, storm-response craft — leaving with a retirement wave. Programmes that only hire lose the domain; programmes that only document lose the capability.
  2. The scarcest role is not the data scientist. It is the person who can hold both worlds: the OT-cleared data engineer trusted near SCADA, and the ex-control-room engineer who can own an operator-facing product. Neither can be bought off the shelf; both take six months to a year to create.
  3. The talent ladder has five stages — Rented, Enclave, Blended, Pipelined, Self-renewing — and most utilities sit at the second, where a central data team exists but operations does not trust it. The move that matters is stage 2 to 3: dissolving the enclave into blended pods with shared targets.
  4. Build–buy–borrow is a per-role decision, not a programme philosophy. Grid-specific and scarce roles are built by converting engineers you already employ; generic and scarce roles are borrowed through partners and consortia; only genuinely transferable roles are bought.
  5. Knowledge capture only works as engineering, not as HR process. An exit interview preserves nothing a model can use; co-labelling historical cases with the expert before they leave turns thirty years of judgement into a training set — and the 90-day plan on this page does exactly that for one transformer fleet.

Abbreviations used on this page

OT
Operational technology — the control-system estate, as distinct from corporate IT
SCADA
Supervisory control and data acquisition
ADMS
Advanced distribution management system
EMS
Energy management system (transmission control room)
AMI
Advanced metering infrastructure — smart meters and the head-end system
DER
Distributed energy resources — rooftop solar, batteries, EV chargers, flexible load
EAM
Enterprise asset management system — the asset register and work-order estate
GIS
Geographic information system — the network's connectivity and location model
DGA
Dissolved gas analysis — the transformer-oil test behind most condition judgements
DNO
Distribution network operator (GB licensee)
CIP
Critical infrastructure protection — the NERC cyber-security standards family
MLOps
Machine-learning operations — serving, monitoring and retraining discipline

Free · 8 questions · ~3 minutes

Score your workforce on the talent ladder

Eight questions, one at a time, about three minutes — all of them about people, none about platforms. Answer them and we build your personalised talent-readiness report: your stage on the ladder, your score on each of the four workforce dimensions, and the specific gap — inflow or outflow — that is actually capping your AI programme. Sent to your inbox.

0 of 8 answered

Question 1 of 8Skills inventory & demographics

Which critical operational judgements depend on a single named person — and do you know who?

The single-point-of-knowledge register is the demographic exposure measure that matters. Retirement dates are known; what leaves with them usually is not.

How the score maps to a stage
  • 04 — Stage 1, Rented. Rented is the stage where every AI-relevant skill sits outside the organisation — vendors, consultants and integrators do the work, and the capability leaves when the contract ends.
  • 59 — Stage 2, Enclave. Enclave is the stage where an internal data team exists but operations does not trust it — capability has been hired into a central unit that the control room, depots and planning teams treat as a foreign country.
  • 1014 — Stage 3, Blended. Blended is the stage where paired pods of domain engineers and data specialists own operational outcomes together, and the first grid-specific hybrid roles — OT-cleared data engineers, model-literate asset engineers — exist and are trusted.
  • 1519 — Stage 4, Pipelined. Pipelined is the stage where talent supply is repeatable: a conversion pathway turns the utility's own engineers into hybrid roles on a schedule, external intake feeds the bottom, retention economics are solved deliberately, and knowledge capture runs against the retirement forecast rather than behind it.
  • 2024 — Stage 5, Self-renewing. Self-renewing is the stage where the workforce develops the workforce: AI literacy sits in ordinary role profiles, operational staff build and maintain models inside governed bounds, and competence is managed as an audited system rather than a programme — so the capability no longer depends on any individual, initiative or budget line.

What the AI readiness talent gap in utilities actually is

A definition, the two directions the gap runs in, and the flow diagram most readiness assessments never draw — the one with people in it.

The AI readiness talent gap in utilities is the distance between the workforce a utility employs and the workforce its AI programme requires — measured in roles it cannot fill, conversions it has not started, and expertise it is about to lose. It is the reason most utility AI readiness work fails its own test: assessments audit data platforms, integration estates and governance frameworks in detail, then compress the entire human question into a slide titled 'change management'. Yet every capped programme we have seen was capped by a person-shaped hole — the OT-cleared data engineer who took nine months to hire, the control-room product owner who cannot be hired at all, the transformer diagnostician who retired in March with thirty years of judgement uncaptured.

What makes the utility version of this gap different from every other industry's is that it runs in two directions at once. The inflow side is the familiar one: data engineers, ML engineers and MLOps specialists are scarce everywhere, and a regulated utility — with its pay bands, its security perimeter and its clock speed — competes for them at a structural disadvantage. The outflow side is the one the sector owns almost alone: the workforce that holds the network's tacit knowledge is old, and retiring. The IEA's World Energy Employment report (opens in a new tab) puts the global energy workforce at around 67 million people and documents the sector's persistent struggle to attract skilled workers — while inside any individual utility, the engineers who can read a dissolved-gas result against a loading history, or feel when a storm forecast means calling crews early, are disproportionately in their final decade of service. AI readiness depends on both flows: the new skills arriving, and the old judgement being captured before it walks out.

How AI capability enters a utility — and how grid knowledge leaves it

The two flows a talent-readiness plan must manage at once. The top lane is the inflow most programmes obsess over; the bottom lane is the outflow most programmes ignore. Both terminate in the blended core — or they terminate in value leaks.

  • Data & feeds
  • Human in the loop
  • Where value leaks
  • AI / model
  • System-of-record action

The process, in words

  • Capability flows in through three channels — external hires, partners, and internal conversion — and every external channel passes through the clearance-and-context bottleneck: months of security vetting and OT familiarisation before a hire is genuinely productive near SCADA. Conversion of the utility's own engineers bypasses that bottleneck, which is why it is the most underused high-leverage channel.
  • Both flows land in the blended core: pods pairing domain engineers with data specialists, owning one operational metric together, shipping models into the surfaces the operation already uses — the ADMS, the EMS, the EAM. The pod is simultaneously the delivery unit and the classroom in which hybrid people are made.
  • Knowledge flows out through retirement on a schedule nobody controls. The only intervention that works is capture while the expert is still on payroll — co-labelling historical cases so the judgement becomes a dataset a pod actively consumes. The dashed edge to retirement is the leak: expertise leaving uncaptured, which no later hire at any salary can restore.
Step-by-step insights
The external market — necessary, and structurally stacked against you
A utility hiring a data engineer competes with software firms on pay, with startups on excitement, and with consultancies on variety — and typically loses on all three. What it can win on is meaning and moat: the grid is the most consequential machine most engineers will ever touch, and grid context compounds into a career asset that generic data roles never build. Utilities that lead their recruitment with the mission and the craft — not the pension — measurably out-hire those that run the standard advert. But the market channel should still be reserved for the roles conversion cannot produce: the seed data engineer, the MLOps platform hand, the first governance lead.
Partners and consortia — renting the frontier without re-renting the basics
The borrow channel is legitimate and permanent: no single utility should train its own foundation models or chase the research frontier alone, and industry pooling — EPRI's Open Power AI Consortium is the visible example — exists precisely to share that cost. The discipline is the boundary: borrowed capability must never sit on the critical path of daily operations, and every engagement must leave residue — code the utility owns, definitions its people understand, at least one internal engineer who paired on the work end to end. Borrow the frontier; never re-rent the basics you already learned.
Internal conversion — the channel with the unfair advantage
A protection engineer who learns Python brings fifteen years of network judgement to every feature they build; a data-science graduate learns the same judgement over a decade, if ever. Conversion flips the scarcity: instead of competing in the world's tightest labour market for people who lack your context, you develop people who already hold the context, already hold the clearances, and already have the control room's trust. The costs are real — backfilling their old role, paying for the new one properly, tolerating a slower ramp — but they are a fraction of the fully loaded cost of the external alternative, and the retention profile is dramatically better.
The clearance bottleneck — why the first hire starts six months early
Working OT-adjacent is gated, and rightly: in North America the NERC CIP standards make personnel risk assessment and training a compliance obligation around the bulk electric system, and equivalent security regimes apply elsewhere. Sourcing, vetting, CIP-style training and supervised familiarisation stack into a six-to-nine-month runway between requisition and genuine productivity near SCADA. Programmes that discover this on the critical path lose two quarters; programmes that start the hire at business-case stage — before the funding lands — hide the runway entirely. It is the single most schedule-critical fact in this entire page.
Co-labelling — the only capture that survives contact with a model
Most knowledge capture produces documents, and documents do not train models or successors — they decay unread in the EAM. Co-labelling produces data: the retiring diagnostician works through hundreds of historical cases — DGA panels, thermal images, failure records — recording the call they would have made and why, with a pod engineer beside them encoding the reasoning into features and labels. The output is threefold: a labelled corpus the current model trains on, a validation set future models are judged against, and a successor who spent two hundred cases' worth of hours inside the expert's head. It is the highest-value engineering work most utilities have never scheduled.

The five stages of the utility AI talent ladder

Rented, Enclave, Blended, Pipelined, Self-renewing — for each stage: what it looks like from the control room, the signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what leaving costs.

The talent ladder runs from capability you rent to capability that renews itself, and every utility we have worked with sits identifiably on one of its five rungs. Each stage below is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own organisation this week — most of them answerable from the ATS, the HR system and one conversation with a shift engineer — and the anti-pattern is the specific mistake most often made trying to climb out. Note what the ladder is not: it is not a technology maturity model. A utility can run a modern data platform from stage 2, and a spreadsheet estate from stage 3. The ladder measures people, trust and renewal — the things platforms cannot substitute for.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Rented

26% of operators sit here

Rented is the stage where every AI-relevant skill sits outside the organisation — vendors, consultants and integrators do the work, and the capability leaves when the contract ends.

Rented is not a failure state — it is where almost every utility rationally starts, because the first pilots arrive attached to vendor products and consultancy engagements. The failure is staying there. A rented capability produces demonstrations, reports and occasionally a working model, but it produces no residue: no internal engineer who understands the feature pipeline, no code the utility can modify, no judgement about what worked that survives the supplier rotating its team. The tenth engagement costs what the first did, and teaches the organisation exactly as little.

The tell is the renewal meeting. At Rented, the conversation about year two of an AI contract contains no utility voice that can challenge the supplier technically — procurement negotiates price because nobody can negotiate substance. That asymmetry compounds: the supplier's people learn the network, the feeder quirks, the data gotchas, and then take that learning to their next client. The utility has paid to train the market's engineers on its own grid.

What makes Rented dangerous rather than merely slow is the outflow side. While the capability is rented, the workforce that holds the network's tacit knowledge — the protection engineer who can read a DGA result like a novel, the control-room veteran who knows which alarms lie — keeps ageing. Every year at Rented is a year of retirements uncaptured, and that loss, unlike a delayed hire, cannot be bought back at any price.

In practice

The model the supplier took home

A mid-sized DNO commissioned a consultancy to build a fault-prediction proof of concept on overhead lines. The model worked; the final presentation impressed the board. Eighteen months later a new asset-management director asked to extend it to cable networks and discovered the utility owned a slide deck: the code sat in the supplier's repository, the feature definitions in a departed contractor's head, and the two engineers who had fed the project data had both retired. Extending it meant starting again — at full price, with a new supplier.

What it looks like

  • All modelling and data engineering is delivered by vendors or consultants
  • No internal role profile mentions data or ML competence
  • Nobody in-house can reproduce, retrain or even rerun what a supplier built
  • Retirement dates are tracked by HR; the knowledge behind them is tracked by nobody

Diagnostic signals you can check this week

  • Ask who in-house could retrain the most recent vendor-built model. If the answer is a company name, you are at Rented
  • Check whether any internal job description requires data or ML competence. At Rented, none does
  • Ask what artefacts the last AI engagement left behind — code, data definitions, runbooks — and who can use them
  • Ask for the list of judgements that depend on one named person. At Rented the list does not exist

Anti-pattern · Hiring a head of AI first

The instinctive exit from Rented is a senior hire — a head of AI or chief data officer — brought in to 'build the capability'. Landed into an organisation with no engineers to lead, no data plumbing and no operational sponsor, the hire spends a year writing strategy, loses the confidence of the executive that hired them, and leaves. The sequence is backwards: the first hires are doers with a bounded problem — one data engineer, one converted domain engineer, one use case — and leadership is added once there is something to lead.

What holds you here

There is no internal skill to accumulate learning on, so every engagement resets to zero — and the retirement clock keeps running.

Highest-leverage next move

Build the single-point-of-knowledge register, write the first two role profiles, and start the OT data-engineer hire now — it is the longest lead item on the whole roadmap.

Cost of leaving

Effort
6–9 months to a defensible Enclave-or-better position
Team
One OT-adjacent data engineer (external hire), one seconded domain engineer, a named operations sponsor
Risk
Low — the work is additive; the real risk is another year of uncaptured retirements
To next stage
6–9 months

If this is you, the next step is

A two-week engagement: the single-point-of-knowledge register, the first role profiles, and the sourcing plan.

Scope your first two hires

Stage 2

Enclave

37% of operators sit here

Enclave is the stage where an internal data team exists but operations does not trust it — capability has been hired into a central unit that the control room, depots and planning teams treat as a foreign country.

Enclave is the most common stage and the least stable. The utility has done the hard fundraising work — headcount approved, a data science team hired, often at salary bands that required an uncomfortable HR exception — and the team is genuinely capable. What it lacks is a route into the operation. Its members sit in a corporate office, work from data extracts, and present findings to committees. The models are competent; the network carries on exactly as before. Six months of this teaches operations that the team is decorative, and teaches the team that the utility was not serious — which is when the CVs go out.

The attrition mechanics deserve more attention than they get. A data scientist inside a utility enclave is being paid below the market that recruited them, working on problems that never ship, and watching former colleagues at software firms deploy weekly. The utility's traditional retention levers — stability, pension, a thirty-year career arc — are precisely the ones this cohort discounts. The result is a two-year revolving door: hire at a premium, lose before the first production deployment, re-hire at a bigger premium. The enclave does not just underdeliver; it actively burns the employer brand needed to staff stage 3.

The exit is not more hiring — it is dissolution. The enclave dissolves into blended pods: one data engineer and one modeller paired with a protection engineer, an asset manager or a shift engineer, sitting with the operational team, owning one operational number together. Utilities resist this because it feels like breaking up the expensive thing they just built. It is actually the first moment the expensive thing starts to pay, and — the part nobody expects — the first moment the data specialists stop resigning, because their work finally lands on a real grid.

In practice

The forecasting team the control room never called

A vertically integrated utility built a five-person analytics team that produced, among other things, a genuinely good day-ahead demand forecast. The control room continued running on the incumbent vendor forecast, because the analytics team's version arrived by email, carried no accountability, and nobody standing a shift had ever met its authors. When the team's lead modeller resigned for a trading house, her handover note observed that in two years no operational decision had changed because of anything the team produced. The forecast was excellent. It was also irrelevant.

What it looks like

  • A central data science or innovation team exists, typically 3–10 people
  • Its backlog is set by an innovation board, not by operational pain
  • Control-room and field engineers describe the team as 'them', not 'us'
  • Attrition in the data team runs well above the utility's average

Diagnostic signals you can check this week

  • Ask a shift engineer to name one member of the data team. At Enclave, they cannot
  • Compare the data team's attrition rate with the engineering average — at Enclave it is typically double or worse
  • Check where the data team sits — physically and organisationally. Corporate floor, corporate reporting line: Enclave
  • Count operational decisions changed by the team's output in the last quarter. Zero is the Enclave signature

Anti-pattern · Fixing the enclave with a bigger backlog

When an enclave underdelivers, the reflex is to feed it better projects — an innovation funnel, hackathons, an executive-sponsored use-case pipeline. This treats the problem as demand when the problem is trust. Operations does not withhold cooperation because it lacks ideas; it withholds cooperation because the team has no skin in operational outcomes and no member who has ever stood a storm shift. No backlog fixes that. Pairing fixes it — one pod, one operational metric, shared on-call for the thing they ship.

What holds you here

Capability and credibility live in different buildings: the data team has no route into operations, and operations has no reason to open one.

Highest-leverage next move

Dissolve the enclave into its first blended pod — one data engineer, one modeller, one domain engineer, one operational number — and move their desks to the operational team.

Cost of leaving

Effort
9–15 months to blend the first two pods and prove one operational result
Team
Existing data team, two seconded domain engineers at 50%+, one control-room or depot 'landlord' per pod
Risk
Medium — the first pod must land a visible win or the blending model is discredited with it
To next stage
9–15 months

If this is you, the next step is

Half a day with your data lead and operations lead: the problem, the pairing, the metric, the seating plan.

Design your first blended pod

Stage 3

Blended

22% of operators sit here

Blended is the stage where paired pods of domain engineers and data specialists own operational outcomes together, and the first grid-specific hybrid roles — OT-cleared data engineers, model-literate asset engineers — exist and are trusted.

Blended is where the talent programme starts producing compound interest. The pod structure does two jobs at once: it ships operational results, and it manufactures exactly the hybrid people the market cannot sell you. Twelve months of pairing turns a data engineer into someone who understands why the feeder naming convention lies, and turns a protection engineer into someone who can read a confusion matrix. Neither conversion happens in a classroom. The pod is the classroom, and its output — beyond the model — is the first generation of staff who hold both worlds.

The binding constraint at this stage is access, not skill. A data engineer cannot build governed feeds from SCADA and the historian without crossing the OT security perimeter, and utilities are rightly conservative about who crosses it — in North America the NERC CIP obligations make that caution a compliance matter, not a preference. Clearing a data engineer to work OT-adjacent takes months of vetting, training and supervised access, which is why the role has to be hired and cleared long before the programme needs it, and why poaching someone already cleared commands the premium it does.

Blended is also the stage where the outflow side of the gap must be wired in, because pods create the first genuine consumers for captured knowledge. A retiring diagnostic engineer's judgement is worth little as a document in the EAM; it is worth a great deal as a labelled dataset a pod is actively training on. The discipline is co-labelling: sitting the expert with the pod over historical cases — DGA results, thermal images, failure records — and recording their calls and their reasoning while they are still on payroll. Utilities that skip this at stage 3 arrive at stage 4 with pipelines and platforms and nobody left who can tell the model when it is wrong.

In practice

The pod that out-recruited the enclave

A transmission operator moved two of its four data scientists to sit with the substation asset team, paired with a senior asset engineer a year from retirement, with one shared target: cut unplanned transformer outages on a named fleet. In nine months the pod shipped a condition-ranking model into the EAM's maintenance-planning screen, co-labelled eleven years of DGA history with the retiring engineer, and — the result nobody projected — received four internal transfer requests from engineers who wanted the next pod seat. The same utility had spent two years failing to fill a data science vacancy by advertisement.

What it looks like

  • At least one pod pairs data specialists with domain engineers on a shared operational metric
  • An OT-cleared data engineer works inside the control-network perimeter, with security's blessing
  • Converted domain engineers — protection, planning, asset — do model work part-time and are paid for it
  • The control room consults model output because a named engineer it knows answers for it

Diagnostic signals you can check this week

  • Ask whether any data specialist has OT-network access agreed with security. Blended requires at least one
  • Ask a pod's domain engineer what precision and recall mean, and its data engineer what a DGA result is. Both should manage
  • Check whether converted domain engineers were re-graded or re-titled. Unpaid conversion is unfinished conversion
  • Look for a co-labelling artefact — a labelled historical dataset with the expert's reasoning attached. Its absence means the outflow side is still open

Anti-pattern · Scaling pods faster than the trust that feeds them

After the first pod lands a win, the temptation is to stamp out five more by reorganisation — assign names to pods on a slide and declare the operating model transformed. But a pod is not a seating plan; it is a trust relationship with a specific operational team, built through one delivered result. Pods declared faster than results are pods in name only, and the operational teams assigned to them learn the initiative is theatre. Grow pods at the rate results land — roughly one new pod per proven pod per year — and staff each new one with at least one veteran of a working pod.

What holds you here

Everything runs through a handful of hybrid people the market cannot replace — one resignation from the first pod can undo a year.

Highest-leverage next move

Turn the conversions the pods produced by accident into a deliberate pathway — named curriculum, re-grading, a cohort a year — and start the graduate and apprentice intake that feeds it.

Cost of leaving

Effort
12–24 months to a repeatable pod model and the first internal conversion cohort
Team
Two to four pods; one platform/MLOps engineer shared across them; an L&D partner for the conversion pathway
Risk
Medium — key-person risk concentrates in the first OT-cleared engineer and first pod leads
To next stage
12–24 months

If this is you, the next step is

We review pairing, access, pay and capture against what has actually made pods survive elsewhere.

Pressure-test your pod design

Stage 4

Pipelined

11% of operators sit here

Pipelined is the stage where talent supply is repeatable: a conversion pathway turns the utility's own engineers into hybrid roles on a schedule, external intake feeds the bottom, retention economics are solved deliberately, and knowledge capture runs against the retirement forecast rather than behind it.

Pipelined is the stage most talent strategies claim and few operations can evidence. The test is unglamorous: when a pod loses its data engineer, how long until a replacement of equal usefulness is productive? At stage 3 the honest answer is six months to a year of external search plus clearance plus familiarisation. At Pipelined it is weeks, because the replacement already exists inside the organisation — a conversion-cohort graduate who knows the estate, holds the clearance, and has sat in a pod as an apprentice member. The pipeline's product is not headcount; it is replaceability without loss.

The economics need to be faced squarely rather than managed by exception. A utility cannot and should not match trading-house or big-tech pay for pure ML talent — but it does not need to, because its hybrid roles are not fungible with those markets. An OT-cleared data engineer with three years of grid context has a career moat a generic data engineer lacks, and the utility can pay a defensible premium for the hybrid — benchmarked against what it would actually cost to replace them, not against the engineering band next door. Utilities that benchmark hybrid pay against other utilities' generic bands are running a subsidised training scheme for consultancies.

On the outflow side, Pipelined means capture has become boring — which is the goal. The register of single points of knowledge is reviewed quarterly alongside the retirement forecast; every critical judgement has a capture plan with a date; co-labelling sessions are scheduled work with work orders, not favours extracted in notice periods. The wider industry has begun building shared scaffolding for exactly this — EPRI's Open Power AI Consortium pools utilities' effort on power-sector AI models and practices — but the tacit knowledge of a specific network can only be captured by the utility that runs it, while the people who hold it are still on payroll.

In practice

The cohort that closed a vacancy in three weeks

A distribution utility ran its second annual conversion cohort — six engineers from protection, planning and asset management, each spending two days a week for nine months on data engineering and ML fundamentals, each attached to a working pod, each re-graded on completion. When the storm-response pod's data engineer left for a software firm, the vacancy was filled in three weeks by a cohort graduate from network planning, already CIP-trained, already known to the control room. The external advert that would once have run for eight months was never posted.

What it looks like

  • A named conversion pathway — curriculum, mentoring pod seat, re-grading — graduates a cohort a year
  • Graduate and apprentice intake includes data-and-grid hybrid schemes, not just traditional disciplines
  • Pay for hybrid roles is benchmarked against the market that actually poaches them, and refreshed annually
  • The single-point-of-knowledge register drives scheduled capture, prioritised by retirement date and criticality

Diagnostic signals you can check this week

  • Ask for the conversion pathway's completion and retention numbers. A pathway without numbers is a slide
  • Check the last pay-benchmark refresh for hybrid roles, and what market it was benchmarked against
  • Ask how the next twelve months of retirements map to capture plans. Pipelined has a dated answer per name
  • Time-to-fill for pod vacancies: weeks from internal supply is Pipelined; months of external search is not

Anti-pattern · Outsourcing the pipeline to a training vendor

At this stage a procurement instinct kicks in: buy a 'data academy' from a training vendor, put two hundred staff through it, report the throughput to the board. Coursework without a pod seat and a re-graded role to land in converts nobody — it produces certificates, a brief morale bump, and cynicism when the promised transformation does not follow. The pathway's scarce ingredient was never content, which is abundant; it is the pod seat with mentoring and the re-grading at the end. Size cohorts to the pod seats available, not to the licence count the vendor discounts at.

What holds you here

The pipeline is funded as an initiative rather than as infrastructure, so it survives exactly until the first difficult budget round.

Highest-leverage next move

Move competence from programme to management system: role profiles, training records and capture obligations written into the operating model and audited like any other licence obligation.

Cost of leaving

Effort
18+ months to make supply demonstrably repeatable across two cohort cycles
Team
Programme lead, L&D partner, pod leads as mentors, HR reward partner for benchmarking and re-grading
Risk
Medium — pipelines are quietly starved in cost-cutting rounds because their loss shows up two years later
To next stage
18+ months

If this is you, the next step is

Two days: conversion throughput, benchmark drift, capture backlog — and where the next cut would actually land.

Audit your pipeline economics

Stage 5

Self-renewing

4% of operators sit here

Self-renewing is the stage where the workforce develops the workforce: AI literacy sits in ordinary role profiles, operational staff build and maintain models inside governed bounds, and competence is managed as an audited system rather than a programme — so the capability no longer depends on any individual, initiative or budget line.

Self-renewing is narrower and more procedural than the name suggests. It does not mean every lineworker writes Python; it means the organisation has stopped treating AI competence as a special commodity held by special people. A planning engineer extending a load-forecast feature, a substation engineer flagging a mislabelled training case through the same system they would use for a defect, an apprentice rotating through a pod as routinely as through a depot — these are the textures of stage 5. The specialist roles still exist; what has changed is that the organisation around them can absorb their departure without a programme review.

The governance scaffolding is what makes this safe rather than chaotic. Management-system standards for AI — ISO/IEC 42001 is the reference point — treat competence the way safety management always has: the organisation must determine what competence each AI-touching role requires, ensure it exists, and keep the evidence. For a utility this is familiar territory; it already runs authorisation regimes for switching, for working near live equipment, for control-room roles certified under NERC's operator-certification scheme in North America. Stage 5 simply extends a discipline the industry has practised for a century — competence as a controlled, audited property of roles — to the model estate.

Sustaining stage 5 is a renewal discipline, and regression is quiet. The signature failure is competence rot: role profiles written in the stage-4 push slowly drifting out of date while the model estate grows, until an incident review discovers that the person who approved a model change held no current competence for it. The countermeasure is the same as for any management system — periodic review keyed to change, not to calendar piety: new model classes, new platform capabilities and reorganisations each trigger a competence-profile review, exactly as a network change triggers a protection-settings review.

In practice

The retirement that was a calendar entry

At a utility operating at stage 5, the retirement of the senior cable-diagnostics engineer — the kind of departure that at stage 1 erases a capability — was a calendar entry. His judgement on partial-discharge interpretation had been co-labelled into the diagnostic corpus over three years; two conversion-cohort engineers held current competence sign-off on the model family; the leaving presentation included a chart of his labelled cases as a proud artefact. The following month, the model caught a cable-joint signature one of his successors confirmed — using the criteria he had taught the dataset.

What it looks like

  • Model-literacy expectations appear in ordinary engineering role profiles, not specialist ones
  • Operational teams extend and retrain models themselves, within a governed platform and policy
  • Competence, training and capture obligations are part of the management system and get audited
  • Leaving events — resignation or retirement — are routine, because succession and capture are continuous

Diagnostic signals you can check this week

  • Pick three AI-touching roles and ask for their competence profiles and evidence. Stage 5 produces them like maintenance records
  • Ask when a competence profile was last updated because the model estate changed — not because a year passed
  • Check whether any operational team retrained or extended a model without the central team in the loop, legitimately
  • Ask what happened at the last senior technical retirement. 'Routine' is the stage-5 answer

Anti-pattern · Mistaking a mature platform for a mature workforce

Stage-5 platforms make model work look easy, and the anti-pattern is reading platform maturity as workforce maturity — thinning the specialist core, pausing the cohorts, letting the capture cadence slip, because 'the platform handles it'. The platform automates yesterday's competence; it cannot renew it. Two years of this leaves an estate of models nobody currently employed can deeply defend — a Rented stage with better tooling. The pipeline and the capture discipline are not scaffolding to remove at maturity; they are the maturity.

What holds you here

Nothing external blocks stage 5 — the threat is internal: competence rot, quietly, while the dashboards stay green.

Highest-leverage next move

Key competence review to change: every new model class, platform capability or reorganisation triggers a profile review, exactly as a network change triggers a protection review.

Cost of leaving

Effort
Continuous — renewal keyed to change in the model estate and the workforce
Team
Standing competence authority (as for switching authorisations), pod network, cohort programme as business-as-usual
Risk
Concentrated in complacency — the failure mode is slow, silent competence rot, discovered by incident

If this is you, the next step is

We audit AI-role competence profiles and evidence the way a safety auditor reads authorisation records.

Review your competence management

Where utilities sit on the ladder today

The distribution across the five stages, why the Enclave is the mode, and what the wider labour market does to every rung.

Most utilities sit at the Enclave stage — a hired central team that operations has not adopted — and the distribution falls away steeply above it. That shape is worth dwelling on, because it differs from the technology-maturity distributions the industry is used to seeing: utilities have collectively bought a great deal of AI tooling, but tooling migrates between stages in a procurement cycle while people migrate in career-lengths. The result is a sector whose platform maturity now runs a full stage ahead of its workforce maturity almost everywhere — and the workforce number is the binding one.

Distribution of utilities across the five talent stages

Illustrative distribution. The Enclave is the mode and the trap: it is where investment has already happened but value has not, which makes it both the most common and the most politically fragile place to be stuck.

Share of utilities

  • 26% — 1 · Rented
  • 37% — 2 · Enclave (the trap)
  • 22% — 3 · Blended
  • 11% — 4 · Pipelined
  • 4% — 5 · Self-renewing

Source: Illustrative distribution, synthesised from IEA workforce research and McKinsey AI adoption surveys

The macro forces behind the distribution all push the same way. AI adoption has gone mainstream across industries — McKinsey's State of AI research (opens in a new tab) has tracked the share of organisations using AI in at least one function climbing to more than three-quarters — which means utilities now compete for the same talent as every other sector, rather than against each other. Meanwhile the energy transition is expanding the sector's own headcount needs: IRENA and the ILO count 16.2 million renewable-energy jobs worldwide (opens in a new tab) and rising, and European industry bodies such as Eurelectric (opens in a new tab) have made workforce and skills a standing policy theme for exactly this reason. A utility talent strategy written as if the competition were the neighbouring utility is benchmarking against the one market that is not the threat.

6–9 mo

realistic runway from requisition to a productive OT-cleared data engineer, once vetting and familiarisation are counted (illustrative; see the role ledger)

typical attrition in enclave-stage data teams relative to the utility's engineering average (illustrative pattern across published attrition research)

37%

of utilities sitting at the Enclave stage in the illustrative distribution above — investment made, value not yet landed

The role ledger: ten roles, their scarcity, and the build–buy–borrow call

The roles a utility AI programme actually runs on — what each one really does, how hard it is to get, how long until it is productive, and which sourcing channel is the honest default.

A utility AI programme runs on about ten roles, and the talent gap is really ten separate gaps of very different depths. Treating them as one 'AI skills' shortage produces the classic failure: a hiring plan that buys what should be built, builds what could be bought cheaply, and never notices that its two most critical roles cannot be sourced externally at all. The ledger below is the planning tool we use with utility leadership teams — for each role, what the work actually is, how scarce it genuinely is in the market, how long from start date to real productivity inside a regulated estate, and the sourcing call that usually survives contact with reality.

RoleWhat they actually doMarket scarcityTime to productiveDefault call
OT-cleared data engineerBuilds governed feeds from SCADA, the historian and AMI into the analytics estate — inside the security perimeter, with the OT team's trustSevere — cleared-and-experienced is a seller's market6–9 monthsBuy one to seed, then build the rest
Power-systems ML engineerModels load, DER output and network constraints with the physics respected — knows why a naive forecast violates thermal limitsSevere — the intersection is tiny6–12 monthsBuild: convert a power engineer, add ML
Asset-health data scientistCondition models from DGA, thermal imagery, partial discharge and load history; works to failure modes, not just featuresHigh3–6 monthsBuild from asset engineering
MLOps / platform engineerServing, monitoring, retraining pipelines; makes model operations boring — the discipline transfers from any industryHigh, but transferable~3 monthsBuy
Control-room product ownerOwns the operator-facing decision surface; an ex-shift engineer the room still trusts — decides what an operator sees and whenSevere — cannot be bought externally at all~12 months (internal)Build only
GIS / network-model stewardKeeps the as-operated network model joined to the asset register and the EAM — the join every grid model silently depends onModerate3–6 monthsBuild
Forecasting analystDemand, generation and imbalance forecasting into operations and trading; owns forecast bias as a number with a name on itModerate~3 monthsBuild or buy
AI governance leadModel risk, audit evidence, alignment to NIST AI RMF and ISO/IEC 42001; speaks regulator fluentlyHigh~6 monthsBuy or borrow
Knowledge engineerRuns the capture programme: co-labelling sessions, judgement elicitation, the single-point-of-knowledge registerHigh — the title barely exists yet3–6 monthsBuild from L&D or engineering
Data product managerTurns use cases into owned, funded products inside a regulated estate; manages the seam between pods and the businessHigh~6 monthsBuy
The utility AI role ledger. Time-to-productive is elapsed time from start date (or pathway start, for built roles) to unsupervised useful work inside the estate — including clearance and OT familiarisation where relevant. Sourcing calls are defaults, not laws; the matrix below shows how to re-derive them for your market.

Two of the ten gate everything else, and neither is the data scientist. The OT-cleared data engineer gates the inflow: until governed feeds exist from SCADA, the historian and AMI, every other role is working on exports — and the six-to-nine-month runway to create this person makes them the schedule-critical hire of the entire programme, which is why the requisition belongs at business-case stage, not at kick-off. The control-room product owner gates the outcome: models change nothing until an operator acts on them, and the person who can design that surface — and answer for it to a room they once stood shifts in — exists only inside your own organisation, a year of deliberate development away. In North America there is an instructive precedent for how seriously the industry can take role-competence when it chooses to: NERC (opens in a new tab) runs a formal certification regime for system operators. The hybrid roles on this ledger deserve the same seriousness, and at stage 5 they get it.

Build, buy or borrow — deriving the call for any role

Plot any role by how scarce it is in the market and how grid-specific its value is. The quadrant gives the honest sourcing default — and explains why the two most critical utility AI roles can only be built.

Build (steady state)

  • Asset-health data scientist, GIS steward, forecasting analyst
  • Context matters more than market heat
  • Conversion pathway at cohort pace

Build long, borrow the bridge

  • OT-cleared data engineer, power-systems ML engineer, control-room product owner
  • The market cannot sell you these — grow them
  • Partner cover while the pathway runs, never on the critical path

Buy

  • MLOps engineer, data product manager
  • Transferable craft, normal market
  • Standard hire; pay the going rate and move on

Borrow

  • Frontier ML, LLM specialists, one-off research
  • Scarce everywhere, generic everywhere
  • Partners and consortia — with a residue clause in every contract
Grid-specificity of the role — top: Worthless without network context, bottom: Transfers from any industry
Market scarcity — left: Hirable in weeks, right: Seller's market

The outflow side: capturing grid knowledge before it retires

Why the retirement wave is an AI-readiness problem, why conventional knowledge capture fails, and the engineering discipline that actually preserves judgement.

Grid knowledge leaves through retirement on a schedule no programme controls, and in most utilities it leaves uncaptured — which makes the retirement wave a first-order AI readiness problem, not an HR talking point. The judgement at risk is precisely the judgement AI systems need most: which historical failures looked like this DGA signature, which alarms in this substation are known liars, what the network actually does — as opposed to what the GIS says it does — on the rural spurs rebuilt after the 2007 storms. Models trained without this judgement learn the paper network; operators who hold the judgement then correctly distrust the models, and the programme stalls with everyone behaving reasonably. Regulators have seen the workforce risk clearly enough that GB network price controls have required licensees to evidence workforce resilience plans (opens in a new tab) as part of their submissions — but a resilience plan that ends at succession charts still captures nothing a successor or a model can use.

  • Capture fails when it is run as an HR process

    Exit interviews, handover documents and leaver checklists are designed to close accounts, not to preserve judgement. Thirty years of diagnostic craft cannot be written down in a notice period, and what does get written lands in a document-management system where no model and few successors will ever read it. The failure is structural: the process is owned by the wrong function, scheduled at the wrong time, and produces the wrong artefact.

  • Capture fails when it produces documents instead of data

    A document describes judgement; a labelled dataset embodies it. The test for any capture exercise is brutal and simple: could a model train on the output, and could a successor be assessed against it? Prose in the EAM fails both. Two hundred historical cases labelled with the expert's call and reasoning pass both — and take roughly the same number of expert-hours to produce.

  • Capture fails when nothing consumes it

    Capture without a consumer is archaeology in advance. The knowledge-capture exercises that survive are the ones wired into a live use case: a pod actively training a condition model on the labelled corpus, a conversion cohort being assessed against the expert's calls, a validation set gating the next model release. The consumer creates the pull that keeps capture funded — which is why capture belongs inside the AI programme, not beside it.

Knowledge capture that feeds a model, not a binder

  1. Build the single-point-of-knowledge register

    One workshop per operational domain — asset management, control room, field operations — listing every judgement that currently depends on a named individual, scored by criticality and years-to-retirement. This register, not the org chart, is the real map of your outflow risk, and it doubles as the prioritisation queue for everything below. Most utilities are genuinely surprised by what surfaces: the register typically finds two or three single points of knowledge nobody in leadership had identified.

  2. Schedule co-labelling as engineering work

    For each critical judgement, sit the expert with a pod engineer over the historical record — DGA panels, protection operations, storm logs — labelling case by case: the call, the confidence, the reasoning, the tells. Budget real hours against work orders; a meaningful corpus is two hundred cases and change, which at a disciplined half-day cadence is a quarter's work alongside the day job. The corpus becomes training data, validation gold-standard and assessment material in one artefact.

  3. Pair the successor before the leaving date, not after

    Co-labelling with a successor in the room is succession planning that actually transfers the judgement: every disagreement between expert and successor is a teaching moment recorded as data. The successor does not need to be a data specialist — they need to be the person who will answer the questions the expert answers today, with the corpus behind them and a model beside them.

  4. Make the model the living archive

    The final step inverts the usual relationship: instead of the expert's knowledge being archived and forgotten, it lives inside a model that operations consults daily, with the labelled corpus as its audit trail. When the model flags a transformer and a successor confirms it using the criteria the corpus taught, the retired engineer's judgement is still doing shifts. That — not a leaving presentation — is what captured knowledge looks like.

What talent-led programmes look like in public

Three publicly reported approaches, read against the talent ladder. None is an Atomic Loops engagement — each links to the organisation's own published material.

The clearest public evidence for the talent thesis is in where capability actually came from at the organisations that made AI stick. In each case below the differentiating move was a people decision — who was converted, who was paired with whom, what was pooled rather than duplicated — and the technology followed. Outcomes are as reported by each organisation in its own material; where the organisation has not published verifiable figures, we keep the outcome qualitative rather than invent precision.

Three sourcing strategies, read against the ladder

Read against the talent ladder: build (Duke), pair (PG&E), pool (EPRI's consortium). Outcomes as reported by the organisations themselves — verify against the linked source before reusing figures. The EPRI card uses an industry-scene image from our generated library, not an EPRI photograph.

Power generation monitoring and diagnostics operations sceneDuke EnergyUS investor-owned utility · millions of electric and gas customers14
Challenge
Predictive analytics across a very large generation fleet demanded people who understood both the machines and the models — a profile the external market does not sell.
Approach
Duke built its fleet monitoring-and-diagnostics capability around experienced plant and equipment engineers, centralised into a monitoring centre and equipped with analytics — converting deep domain judgement into analytical roles rather than hiring analytics and hoping the domain would follow.
Reported outcome
Duke Energy has publicly credited its centralised monitoring-and-diagnostics approach with early catches of developing equipment problems across its fleet, avoiding major failures and unplanned outages — as described in its own newsroom and fleet-technology communications.
What it shows about the curveThe build channel works at scale: engineers who already held the machine judgement became the analysts, and the trust problem that sinks enclave teams never arose — the fleet already knew them. That is the tl-quadrant of the sourcing matrix, executed over a decade.

Duke Energy newsroom (opens in a new tab)

Transmission line inspection imagery review operations scenePacific Gas and ElectricCalifornia investor-owned utility · ~16 million people served24
Challenge
Wildfire-driven inspection programmes produced millions of asset images — far beyond what inspection engineers could review manually, and far beyond what data teams could interpret without them.
Approach
PG&E has publicly described pairing its inspection and engineering expertise with machine-learning image triage — models surfacing candidate defects, inspection specialists making the calls, their corrections feeding back — and has since extended AI assistance into standards lookup and customer operations.
Reported outcome
PG&E has publicly reported using AI-assisted review to triage inspection imagery at a scale manual review could not reach, with expert review retained on the safety-critical calls — as described in the company's own innovation and wildfire-safety communications.
What it shows about the curveThe pairing pattern is the Blended stage in production: the model does the volume, the domain expert does the judgement, and every correction is training data. The inspection specialists were not displaced by the model — they became its teachers, which is co-labelling under operational pressure.

PG&E (opens in a new tab)

Industry scene: utility engineers collaborating around shared grid analytics displaysEPRI Open Power AI ConsortiumIndustry research institute · utilities and technology partners, pooling AI capability12
Challenge
No individual utility can staff frontier AI work — domain-adapted models, evaluation benchmarks, emerging practice — and duplicating the attempt across hundreds of utilities multiplies the same scarce-talent problem industry-wide.
Approach
EPRI convened utilities and technology companies into the Open Power AI Consortium to develop open, power-sector-specific AI models and shared practice — pooling exactly the work that is scarce everywhere and grid-specific nowhere, so member utilities can point their own people at their own networks.
Reported outcome
The consortium operates as a named, public membership programme developing domain-specific open models and resources for the power sector — as described in EPRI's own consortium material.
What it shows about the curvePooling is the br-quadrant of the sourcing matrix institutionalised: borrow the frontier collectively. The boundary the ladder insists on still applies — consortium membership raises a utility's floor, but the tacit knowledge of a specific network can only be captured and embedded by the utility that runs it.

EPRI — Open Power AI Consortium (opens in a new tab)

The operating model: where each role sits, stage by stage

The talent architecture as a layered system — which layer each stage of the ladder requires, and which roles live in it.

A talent-ready utility is organised in five layers, and each rung of the ladder requires one more of them. Reading your own organisation against this architecture is the fastest structural diagnostic on this page: find the highest layer that genuinely exists — staffed, funded, producing — and you have found your stage, usually one rung below where the strategy deck says you are. The layers are deliberately vendor-free and title-flexible; what matters is that the function exists and reports somewhere sane, not what it is called.

The talent stack, layer by layer

Each layer is annotated with the ladder stage that first requires it. A utility building layer 4 (pipeline) on a missing layer 3 (blended pods) is running a training scheme for other employers — the pathway produces people the operating model cannot hold.

  1. Operational workforce

    Stage 1+

    • Control-room and field teamsThe people whose decisions AI must eventually change
    • Protection, planning and asset engineersThe conversion pathway's raw material — and the domain judgement
    • Veteran expertsThe outflow risk: the register names them, capture preserves them
  2. Data and platform core

    Stage 2+

    • OT-cleared data engineer(s)Governed feeds from SCADA, historian, AMI — the schedule-critical seat
    • MLOps / platform engineerServing, monitoring, retraining — bought, not built
    • GIS / network-model stewardKeeps the as-operated model joined to the EAM
  3. Blended delivery pods

    Stage 3+

    • Domain + data pairsOne shared operational metric, one seating plan, shared on-call
    • Control-room product ownerThe unbuyable role: an ex-shift engineer who owns the operator surface
    • Asset-health and forecasting specialistsConverted engineers doing model work at the network's pace
  4. Pipeline and development

    Stage 4+

    • Conversion pathwayCohorts, curriculum, pod seats, re-grading — supply on a schedule
    • Graduate and apprentice intakeHybrid data-and-grid schemes feeding the bottom of the funnel
    • Partner and consortium managementBorrowed capability with residue clauses — never on the critical path
  5. Governance and renewal

    Stage 5+

    • Competence authorityProfiles, evidence and sign-off for AI-touching roles — run like switching authorisations
    • AI governance leadModel risk and audit evidence, aligned to NIST AI RMF and ISO/IEC 42001
    • Capture programmeThe register, the co-labelling cadence, the corpus — as business-as-usual

Pipeline described

  1. Operational workforce (stage 1+) — Control-room and field teams: The people whose decisions AI must eventually change; Protection, planning and asset engineers: The conversion pathway's raw material — and the domain judgement; Veteran experts: The outflow risk: the register names them, capture preserves them
  2. Data and platform core (stage 2+) — OT-cleared data engineer(s): Governed feeds from SCADA, historian, AMI — the schedule-critical seat; MLOps / platform engineer: Serving, monitoring, retraining — bought, not built; GIS / network-model steward: Keeps the as-operated model joined to the EAM
  3. Blended delivery pods (stage 3+) — Domain + data pairs: One shared operational metric, one seating plan, shared on-call; Control-room product owner: The unbuyable role: an ex-shift engineer who owns the operator surface; Asset-health and forecasting specialists: Converted engineers doing model work at the network's pace
  4. Pipeline and development (stage 4+) — Conversion pathway: Cohorts, curriculum, pod seats, re-grading — supply on a schedule; Graduate and apprentice intake: Hybrid data-and-grid schemes feeding the bottom of the funnel; Partner and consortium management: Borrowed capability with residue clauses — never on the critical path
  5. Governance and renewal (stage 5+) — Competence authority: Profiles, evidence and sign-off for AI-touching roles — run like switching authorisations; AI governance lead: Model risk and audit evidence, aligned to NIST AI RMF and ISO/IEC 42001; Capture programme: The register, the co-labelling cadence, the corpus — as business-as-usual
Step-by-step insights
Operational workforce — the layer everyone already has and nobody counts
Every utility has layer 1; almost none inventories it as AI capability. The protection engineer who automated her own settings checks in a spreadsheet macro, the planning analyst who taught himself Python for load-flow batch runs — these people are the conversion pathway's first cohort, discoverable in an afternoon with the right question. The layer's other asset is negative: it is where the trust deficit lives, and no amount of investment in layers 2–5 substitutes for this layer deciding the programme is credible.
Data and platform core — small, senior, and hired before it is needed
The core's steady-state size surprises people: two to four people serve a mid-sized utility well into stage 4, provided they are senior and the pods do the domain work. The scheduling truth is less comfortable — the OT-cleared seat takes six to nine months to fill and clear, so the requisition belongs in the business case, not the mobilisation plan. A core hired late becomes the excuse for a platform procurement to 'bridge the gap', and the bridge becomes the building.
Blended delivery pods — the layer that manufactures its own staff
Pods are the only layer that produces more capability than it consumes: every year of pod operation yields conversions the market cannot sell — data engineers with feeder intuition, protection engineers who can read a learning curve. This is why the architecture fails when built out of order. A pipeline (layer 4) without pods produces certificates; pods without a pipeline produce heroes who burn out. The pod comes first, proves the model, then the pipeline industrialises what the pod proved.
Pipeline and development — supply as infrastructure, not initiative
The pipeline layer converts talent from a project risk into a utility function — planned, funded and boring, like vegetation management. Its components are unglamorous: a curriculum that respects shift patterns, cohort sizes matched to pod seats, re-grading agreed with reward before the first cohort starts (retro-fitting pay is how utilities lose their first graduates), and partner contracts that leave residue. Its enemy is the budget cycle: the pipeline's loss materialises two years after its funding is cut, which is why stage-5 organisations write it into the operating model where single budget rounds cannot reach it.
Governance and renewal — competence managed like the industry already manages safety
The top layer extends a century-old utility discipline — authorisation regimes for switching, live work and control-room roles — to the model estate. Management-system standards for AI (ISO/IEC 42001 is the anchor; NIST's AI RMF the widely used framework) ask organisations to determine, ensure and evidence the competence behind AI decisions, which is precisely what an authorisation regime does. Utilities are better placed than any other sector to operate this layer, because they alone already run its analogue at scale — the work is translation, not invention.

The architecture also settles a governance question that stalls many programmes: where AI-role competence requirements should live. The emerging answer from the standards world — ISO/IEC 42001 (opens in a new tab), the AI management-system standard, and NIST's AI Risk Management Framework (opens in a new tab) — is that competence is a managed property of the organisation, with defined requirements, maintained evidence and periodic review. Utilities should recognise the shape instantly: it is an authorisation regime, and the industry has run those for generations. The practical move at stages 3 and 4 is to draft competence profiles for the hybrid roles as they are created, so that stage 5's audit discipline inherits artefacts instead of inventing them retrospectively.

A 90-day plan: staffing transformer health before the experts retire

The talent ladder made concrete on one asset-health problem — a blended pod stood up, a retiring diagnostician's judgement co-labelled into a corpus, and the schedule-critical hire started. Contains no platform procurement.

Closing a talent gap takes years in general and 90 days in particular — the trick is refusing to solve it in general. The plan below runs the first Enclave-to-Blended move on one concrete, common problem: a DNO's power-transformer condition assessment depends on two diagnostic engineers, one retiring in under a year, while the asset-health model the strategy promises has no one to build it and no one to teach it. One quarter, one pod, one fleet, one corpus — and the long-lead hire started on day one rather than at some future kick-off. Everything in it generalises; nothing in it is generic.

Enclave → Blended on one transformer fleet, in one quarter

One fleet, one pod, one retiring expert, one requisition. If any phase needs more than its window, narrow the scope — fewer transformers, one voltage class — rather than extending the plan.

  1. Days 1–15

    Register, pod and the long-lead requisition

    Build the single-point-of-knowledge register for the transformer fleet: every judgement — DGA interpretation, thermal assessment, end-of-life calls — mapped to its holder and their horizon. Form the pod: one asset engineer at 50%, one data specialist from the central team, the retiring diagnostician contracted for two half-days a week, a named asset-management sponsor. Open the OT-cleared data engineer requisition today — it is the longest lead item on the entire roadmap and nothing else in this plan depends on waiting for it.

    Register drafted, pod seated together, requisition live

  2. Days 16–45

    Co-labelling sprints on the historical record

    Three structured half-days per week: the diagnostician works through the fleet's historical DGA panels, oil results, thermal images and failure records with the pod's data specialist — recording the call, the confidence, the reasoning and the tells, case by case, into a labelled corpus. The asset engineer drafts the competence profile for the future asset-health role from what the sessions surface. Target: 150+ labelled cases by day 45, disagreements between expert and successor logged as their own data.

    A growing labelled corpus; the judgement becoming data

  3. Days 46–70

    First model against the corpus, expert in the loop

    The pod trains a first condition-ranking model on the corpus and runs it across the live fleet — output surfaced as a ranked review list inside the EAM maintenance-planning screen the asset team already uses, never as a separate dashboard. The diagnostician reviews every model call, and every agreement or correction extends the corpus. In parallel: pay and re-grading conversation for the pod's converting asset engineer, benchmarked against the market that would actually poach them.

    Model in the EAM, expert-review loop running, conversion re-grade agreed

  4. Days 71–90

    Attribution, succession and the scale case

    Measure and report in the operation's own units: review-list precision against the expert's calls, knowledge-coverage ratio for the fleet's critical judgements, successor assessment against the corpus, requisition status on the OT hire. Write the scale case in people terms — the next fleet, the next expert on the register, the first conversion-cohort proposal — and put the capture cadence into business-as-usual work orders so it survives the quarter's end.

    A defended result, a working succession, and the case for cohort one

The order matters

  1. Requisition before everything

    The OT-cleared data engineer takes six to nine months from advert to productive; every other item in this plan takes less. Opening the requisition on day one costs nothing if priorities change and saves two quarters if they do not. It is the cheapest schedule insurance on this page.

  2. Capture before modelling

    The corpus is the model's ceiling: a condition model trained on unlabelled history learns averages, while one trained on the diagnostician's labels learns judgement. Thirty co-labelling sessions before the first training run is not a delay to the model — it is most of the model.

  3. The EAM screen before any new screen

    The asset team plans maintenance from the EAM, so that is where the ranked list lands — the same integration-first rule the maturity literature applies to systems, applied here to attention. A new dashboard would demand the team change its habits to consume the pod's work; a ranked list in their existing screen only asks them to read.

  4. Re-grade during, not after

    The converting asset engineer becomes market-visible the moment the pod ships anything — recruiters read LinkedIn too. Settling the re-grade and progression inside the quarter, while the work is being done, retains the person the plan just created. Retro-fitting pay after a resignation letter costs more and often fails anyway.

Measuring the gap: talent KPIs and the blended-pod checklist

The workforce numbers that make talent readiness arguable rather than asserted — each with its formula, source system and the stage at which it first means anything.

A talent gap you cannot measure is a talent anecdote, and people programmes run on anecdotes get cut first. Every KPI below reduces to records that already exist — in the applicant-tracking system, the HR system, the learning records and the pod's own delivery tracker — so instrumenting them is a joining exercise, not a data-collection programme. The 'honest from' column matters as much as the formula: a conversion throughput reported before a pathway exists, or a knowledge-coverage ratio before a register exists, is theatre, exactly as an automation rate reported before autonomy exists would be.

KPIFormula / readSourceCadenceHonest from
Time-to-fill, OT-adjacent rolesRequisition opened → offer accepted, elapsed daysATSPer requisitionStage 1
Offer-acceptance rate, data rolesOffers accepted ÷ offers madeATSQuarterlyStage 1
Time-to-productiveStart date → first unsupervised production contributionDelivery tracker + access recordsPer hireStage 2
Regretted attrition, data rolesRegretted leavers ÷ average headcount, vs engineering baselineHR systemQuarterlyStage 2
Control-room acceptanceOperator actions taken on model advisories ÷ advisories shownADMS / EMS advisory logWeeklyStage 3
Knowledge-coverage ratioCritical judgements with a live capture artefact ÷ register totalSingle-point-of-knowledge registerQuarterlyStage 3
Conversion throughputEngineers completing the pathway and re-graded ÷ cohort intakeL&D records + HR systemPer cohortStage 4
Pipeline yieldCohort and intake graduates still in post at 24 monthsHR systemAnnualStage 4
Pod survival ratePods intact (staffed, delivering) at 12 months ÷ pods formedProgramme recordsAnnualStage 4
The talent-readiness KPI sheet for a utility AI programme. 'Honest from' is the ladder stage at which the KPI first measures something real rather than something aspirational.

Two of these deserve board-level attention because they predict everything else. Regretted attrition in data roles, read against the engineering baseline, is the earliest reliable signal of an enclave failing — it moves quarters before delivery metrics do. And the knowledge-coverage ratio is the only number on the sheet that measures the outflow side at all: a programme reporting healthy hiring metrics while coverage sits at 20% is winning the battle it chose to fight and losing the one it did not. The checklist below converts the stage-3 transition — the one that matters most — into seven conditions you can verify this week.

Blended-pod readiness checklist

Seven conditions for a first pod that survives. If you cannot tick at least the first three, the pod will be an enclave with a nicer seating plan. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Zero ticks is a starting position, not a verdict

Most utilities at the Enclave stage genuinely tick none of these — the pod concept is new to them. Do not start with tooling or hiring: pick the operational problem (item one) and the register workshop that finds your capture scope (item four). Both are workshops, not programmes, and everything else falls out of them.

Failure modes that reopen the gap

Talent readiness is not monotonic either. Five regressions account for most of the ground utilities lose — and every one is cheaper to prevent than to repair.

Talent maturity regresses more quietly than technology maturity, because the dashboards stay green while the people who made them green update their CVs. The five failure modes below account for most of the regression we see, and they share a signature: each one looks, in the quarter it happens, like a reasonable economy.

Likelihood: highImpact: high

The enclave collapse

The lead data scientist resigns; within two quarters, half the team follows — the market prices enclave experience generously, and the survivors inherit the workload that drove the leavers out. A five-person capability built over three years unwinds in six months, and the organisational memory of the failure poisons the next attempt's hiring.

PreventionBlend early. Pods with operational ownership are the retention mechanism — meaningful work retains where pay alone cannot.

Likelihood: highImpact: medium

Training with nowhere to land

A broad AI-literacy programme puts hundreds through coursework with no pod seats, no re-grading and no changed roles at the end. Certificates are issued, nothing changes, and the workforce learns that AI training is a tick-box — which raises the cost of the pathway that comes later and actually needs volunteers.

PreventionAttach every course to a destination: a pod seat, a re-graded role, a named responsibility. Size training to seats, not licences.

Likelihood: mediumImpact: high

The contractor cliff

Programme knowledge concentrates in contractors and a delivery partner; at contract end or rate-card renegotiation, it walks. The utility discovers that its 'internal' capability was a staffing arrangement — models nobody employed can retrain, pipelines nobody employed can fix — and re-procures the same knowledge at a premium.

PreventionResidue clauses in every engagement: paired delivery, code ownership, handover artefacts — and at least one employee who worked every stream end to end.

Likelihood: highImpact: high

Benchmarking pay against the wrong market

Reward benchmarks hybrid roles against other utilities' generic engineering bands, because that is what the survey covers. The market that actually poaches — software firms, trading houses, consultancies — pays against a different curve, and the utility loses exactly its most converted, most productive people while its benchmark says pay is competitive.

PreventionBenchmark against the destination of your last three regretted leavers, not against the sector survey. Refresh annually.

Likelihood: highImpact: high

Capture postponed to the notice period

The capture plan for a critical expert is deferred — busy quarter, storm season, budget round — until the retirement letter arrives, and four weeks of handover is asked to do the work of two years of co-labelling. The documents produced are sincere and useless, and the loss surfaces eighteen months later as a model nobody can correct and a successor nobody finished training.

PreventionCapture runs against the register's retirement forecast, as scheduled work orders — never as a leaving process.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Talent ladder
The five-stage maturity model for utility AI workforce capability — Rented, Enclave, Blended, Pipelined, Self-renewing — measuring people, trust and renewal rather than platforms.
Blended pod
A small delivery team pairing domain engineers with data specialists, seated with the operational team and owning one operational metric together. The unit that both ships results and manufactures hybrid people.
OT clearance
The security vetting, training and supervised-access process required before a data specialist may work against control-system estates (SCADA, historian, ADMS) — a compliance matter under regimes such as NERC CIP, and a six-month-plus lead time in practice.
Time-to-productive
Elapsed time from a hire's start date (or a conversion's pathway start) to their first unsupervised, useful contribution in production. The honest hiring metric — time-to-fill flatters roles that then spend months in clearance and familiarisation.
Conversion pathway
The structured route by which a utility's own engineers — protection, planning, asset — become hybrid data roles: curriculum, a mentored pod seat, and re-grading on completion. The sourcing channel with the unfair advantage of pre-existing domain judgement.
Single-point-of-knowledge register
The maintained list of critical operational judgements that depend on one named individual, scored by criticality and years-to-retirement. The real map of a utility's outflow risk, and the prioritisation queue for capture.
Co-labelling
Knowledge capture as engineering: a veteran expert works through historical cases with a pod engineer, recording calls, confidence and reasoning as a labelled dataset that trains models, validates successors and outlives the expert's tenure.
Knowledge-coverage ratio
The share of register-listed critical judgements that have a live capture artefact — a labelled corpus, a validated successor, a model in use. The only common KPI that measures the outflow side of the talent gap.
Build–buy–borrow
The per-role sourcing decision: build (convert internally), buy (hire externally), or borrow (partners and consortia) — derived from a role's market scarcity and its grid-specificity, not chosen as a programme-wide philosophy.
Residue clause
Contract terms ensuring an external engagement leaves internal capability behind: paired delivery, code and definition ownership, handover artefacts, and a named employee who worked the stream end to end.
Competence rot
The stage-5 failure mode in which role competence profiles drift out of date while the model estate grows, discovered by incident review. Countered by keying competence review to change — new model classes, platforms, reorganisations — rather than the calendar.

Frequently asked questions

The questions utility leaders ask most often when the workforce side of AI readiness finally gets on the agenda.

What is the AI readiness talent gap in utilities?

It is the distance between the workforce a utility employs and the workforce its AI programme needs — and it runs in two directions at once. On the inflow side, data engineering, ML and MLOps skills are scarce everywhere and utilities compete for them at a structural disadvantage. On the outflow side, the engineers holding the network's tacit judgement — transformer diagnostics, switching intuition, storm craft — are disproportionately close to retirement, and that knowledge leaves uncaptured unless it is deliberately engineered into datasets and successors. A readiness plan that addresses only hiring, or only knowledge capture, closes neither gap.

How many AI specialists does a utility actually need to start?

Fewer than most plans assume: two to four people can seed a credible programme. The minimum viable core is one OT-adjacent data engineer, one modeller, and — the part usually missed — one seconded domain engineer at 50% or more with a real backfill. The domain engineer is not support staff; they are half the pod and the source of the judgement the models need. Programmes fail far more often from missing domain time than from missing data scientists, and a ten-person data team with no operational pairing is an expensive way to stay at stage 2.

Should a utility hire data scientists or retrain its own engineers?

Both, but in the opposite ratio to common practice: buy the small transferable core, build the grid-specific majority. MLOps and platform skills transfer from any industry and should be hired. But the roles that create most value — power-systems ML, asset-health modelling, the control-room product owner — depend on network judgement that takes a decade to acquire externally and already exists in your protection, planning and asset engineers. A conversion pathway that adds ML skills to fifteen years of grid context produces a better hybrid, faster and with better retention, than the reverse ever does.

How long does it take to get a data engineer productive near OT systems?

Plan on six to nine months from requisition to genuine productivity, and treat anything faster as a bonus. The runway stacks: sourcing a candidate willing and able to work in a regulated estate, security vetting and training — a compliance obligation under regimes such as NERC CIP — then supervised OT familiarisation before unsupervised work near SCADA or the historian is sensible. This makes the role the longest lead item on most utility AI roadmaps, which is why the requisition belongs in the business case rather than the mobilisation plan. Poaching someone already cleared shortens the runway and raises the price accordingly.

How can a utility compete with technology-sector pay for AI talent?

By refusing to fight on the open market's terms. For generic ML talent a utility will usually lose a pay race — but its key roles are not generic: an OT-cleared data engineer with grid context has a career moat no software firm offers, and the grid itself is a mission that measurably out-recruits another advertising optimisation role. The practical policy is a defensible hybrid premium benchmarked against the market that actually poaches — the destinations of your last three regretted leavers, not the utility-sector salary survey — plus the retention lever pay cannot buy: work that visibly ships into a real network.

What is a blended pod and why does it work better than a central team?

A blended pod pairs one or two data specialists with a domain engineer, seats them with the operational team, and gives them one shared operational metric — unplanned outages on a fleet, forecast bias on a region. It works because it fixes the two failures central teams cannot: trust, since the control room will consult a model when a named engineer it knows answers for it, and learning, since twelve months of pairing produces hybrid people the market cannot sell. The pod is simultaneously the delivery unit and the classroom, which is why the ladder treats it as the pivotal stage-3 structure.

How do you capture the knowledge of retiring grid engineers for AI?

Through co-labelling, not documentation. Sit the expert with a pod engineer over the historical record — DGA panels, thermal images, protection operations, storm logs — and record their call, confidence and reasoning case by case into a labelled dataset. Two hundred cases is a meaningful corpus and roughly a quarter's work at a half-day cadence. The corpus then does three jobs: training data for condition models, a gold-standard validation set, and assessment material for successors. The test any capture method must pass: could a model train on the output, and could a successor be examined against it? Documents fail both; corpora pass both.

Does ISO/IEC 42001 say anything about workforce competence?

Yes — competence is a core requirement, not a footnote. As a management-system standard, ISO/IEC 42001 requires organisations to determine the competence needed by people whose work affects the AI management system, ensure they have it, and retain evidence. NIST's AI Risk Management Framework makes the same point through its governance function. For a utility this shape is familiar: it is an authorisation regime, structurally identical to the ones already run for switching, live working and control-room roles. The practical move is to draft competence profiles as hybrid roles are created at stages 3–4, so stage 5's audit discipline inherits artefacts rather than reconstructing them.

Can consortium membership replace in-house AI hiring for a utility?

No — it changes what you must hire for, not whether. Industry pooling such as EPRI's Open Power AI Consortium is the rational way to access frontier, grid-generic work no single utility should fund alone: domain-adapted models, benchmarks, shared practice. What no consortium can supply is the capability that is grid-specific by definition — your network's data plumbing, your control room's trust, your retiring experts' judgement. The honest division: borrow the frontier collectively, build the network-specific core internally, and require every borrowed engagement to leave residue — code, definitions and at least one internal engineer who paired on it end to end.

What talent KPIs should a utility board track for AI readiness?

Five cover the board level. Time-to-fill and time-to-productive for OT-adjacent roles measure the inflow pipe's real diameter. Regretted attrition in data roles, read against the engineering baseline, is the earliest signal of an enclave failing — it moves quarters before delivery metrics. Conversion throughput measures whether the pathway is real or a slide. And the knowledge-coverage ratio — critical judgements with a live capture artefact, over the register total — is the only number that tracks the outflow side at all. A programme reporting four healthy hiring metrics and no coverage number is fighting half the war.

When should the first AI hires start relative to the wider programme?

Before the programme formally exists. The schedule-critical hire — the OT-cleared data engineer — carries a six-to-nine-month runway of sourcing, vetting and familiarisation, which is longer than most business-case cycles; opening the requisition at business-case stage means the person lands roughly when funding does. The same logic applies to conversions: a pathway cohort started this quarter graduates into the delivery phase, not after it. Utilities that sequence hiring after approval spend the first funded year waiting for people, which reads afterwards as the technology being slow. The sequencing of the funding itself is a different question — covered in the transformation-timeline page in this series.

Is the talent gap different for a small municipal or cooperative utility?

The mechanics are identical; the arithmetic is harsher and the answer leans harder on borrow and build. A small utility cannot carry a data platform team, which makes three moves decisive: conversion of its own engineers, because even two hybrid people transform a small operation; shared capability through consortia, joint programmes and — where the regulatory framework allows — services pooled with neighbouring utilities; and ruthless residue discipline with vendors, since a small utility that lets knowledge sit with suppliers is permanently at stage 1. The register-and-capture work is, if anything, more urgent: in a fifty-engineer utility, a single retirement can be a capability extinction event.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for energy, manufacturing and logistics operators — forecasting, asset-condition prediction, network optimisation and decision support running against live operational data, integrated into the ADMS, EMS and historian layer rather than delivered as dashboards.

  • · Delivery inside regulated network estates: SCADA, historian, ADMS and EMS integration
  • · Blended-team delivery: our engineers pair with utility domain engineers by design
  • · Knowledge-capture engineering: expert judgement turned into labelled datasets, not binders
  • · 12 cited sources on this page

Sources

  1. International Energy AgencyEnergy and AI (opens in a new tab)
  2. International Energy AgencyWorld Energy Employment 2024 (opens in a new tab)
  3. IRENA / International Labour OrganizationRenewable energy and jobs: Annual review 2024 (opens in a new tab)
  4. McKinsey & CompanyThe state of AI (opens in a new tab)
  5. EPRI (Electric Power Research Institute)Open Power AI Consortium (opens in a new tab)
  6. EurelectricWorkforce and skills policy (opens in a new tab)
  7. NERCReliability and certification programmes (opens in a new tab)
  8. NISTAI Risk Management Framework (opens in a new tab)
  9. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  10. OfgemEnergy network regulation (opens in a new tab)
  11. Duke EnergyNewsroom — fleet monitoring and diagnostics (opens in a new tab)
  12. Pacific Gas and ElectricInnovation and wildfire-safety communications (opens in a new tab)

Find out which side of the gap is capping you — then close it

We run the talent assessment with your engineering, HR and operations leads, verify the workforce numbers against your own systems, and leave you a costed plan: the roles to open this month, the pod to form this quarter, and the capture to schedule before the next retirement. You keep the plan whether or not we build it with you.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.