Energy & UtilitiesReadiness & Transformation Roadmap
The AI readiness talent gap in utilities: closing both sides of the workforce equation
The AI readiness talent gap in utilities is the distance between the workforce a utility has and the workforce its AI programme needs — and it runs in both directions: scarce data and machine-learning skills flowing in too slowly, and decades of tacit grid knowledge retiring out faster than anyone is capturing it. Closing one side without the other closes nothing.

Key takeaways
- The AI readiness talent gap in utilities is two-sided: data and ML skills the sector cannot hire fast enough, and tacit grid knowledge — transformer judgement, switching intuition, storm-response craft — leaving with a retirement wave. Programmes that only hire lose the domain; programmes that only document lose the capability.
- The scarcest role is not the data scientist. It is the person who can hold both worlds: the OT-cleared data engineer trusted near SCADA, and the ex-control-room engineer who can own an operator-facing product. Neither can be bought off the shelf; both take six months to a year to create.
- The talent ladder has five stages — Rented, Enclave, Blended, Pipelined, Self-renewing — and most utilities sit at the second, where a central data team exists but operations does not trust it. The move that matters is stage 2 to 3: dissolving the enclave into blended pods with shared targets.
- Build–buy–borrow is a per-role decision, not a programme philosophy. Grid-specific and scarce roles are built by converting engineers you already employ; generic and scarce roles are borrowed through partners and consortia; only genuinely transferable roles are bought.
- Knowledge capture only works as engineering, not as HR process. An exit interview preserves nothing a model can use; co-labelling historical cases with the expert before they leave turns thirty years of judgement into a training set — and the 90-day plan on this page does exactly that for one transformer fleet.
Abbreviations used on this page
- OT
- Operational technology — the control-system estate, as distinct from corporate IT
- SCADA
- Supervisory control and data acquisition
- ADMS
- Advanced distribution management system
- EMS
- Energy management system (transmission control room)
- AMI
- Advanced metering infrastructure — smart meters and the head-end system
- DER
- Distributed energy resources — rooftop solar, batteries, EV chargers, flexible load
- EAM
- Enterprise asset management system — the asset register and work-order estate
- GIS
- Geographic information system — the network's connectivity and location model
- DGA
- Dissolved gas analysis — the transformer-oil test behind most condition judgements
- DNO
- Distribution network operator (GB licensee)
- CIP
- Critical infrastructure protection — the NERC cyber-security standards family
- MLOps
- Machine-learning operations — serving, monitoring and retraining discipline
Free · 8 questions · ~3 minutes
Score your workforce on the talent ladder
Eight questions, one at a time, about three minutes — all of them about people, none about platforms. Answer them and we build your personalised talent-readiness report: your stage on the ladder, your score on each of the four workforce dimensions, and the specific gap — inflow or outflow — that is actually capping your AI programme. Sent to your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your talent-readiness report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the role-ledger priorities for your stage, and a 90-day capture-and-hiring plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Rented
Rented is the stage where every AI-relevant skill sits outside the organisation — vendors, consultants and integrators do the work, and the capability leaves when the contract ends.
Your next moveBuild the single-point-of-knowledge register, write the first two role profiles, and start the OT data-engineer hire now — it is the longest lead item on the whole roadmap.
Stage 2 · Enclave
Enclave is the stage where an internal data team exists but operations does not trust it — capability has been hired into a central unit that the control room, depots and planning teams treat as a foreign country.
Your next moveDissolve the enclave into its first blended pod — one data engineer, one modeller, one domain engineer, one operational number — and move their desks to the operational team.
Stage 3 · Blended
Blended is the stage where paired pods of domain engineers and data specialists own operational outcomes together, and the first grid-specific hybrid roles — OT-cleared data engineers, model-literate asset engineers — exist and are trusted.
Your next moveTurn the conversions the pods produced by accident into a deliberate pathway — named curriculum, re-grading, a cohort a year — and start the graduate and apprentice intake that feeds it.
Stage 4 · Pipelined
Pipelined is the stage where talent supply is repeatable: a conversion pathway turns the utility's own engineers into hybrid roles on a schedule, external intake feeds the bottom, retention economics are solved deliberately, and knowledge capture runs against the retirement forecast rather than behind it.
Your next moveMove competence from programme to management system: role profiles, training records and capture obligations written into the operating model and audited like any other licence obligation.
Stage 5 · Self-renewing
Self-renewing is the stage where the workforce develops the workforce: AI literacy sits in ordinary role profiles, operational staff build and maintain models inside governed bounds, and competence is managed as an audited system rather than a programme — so the capability no longer depends on any individual, initiative or budget line.
Your next moveKey competence review to change: every new model class, platform capability or reorganisation triggers a profile review, exactly as a network change triggers a protection review.
0 / 24
Skills inventory & demographics
— / 6
Sourcing & pipeline
— / 6
Embedding & operational trust
— / 6
Retention & knowledge renewal
— / 6
Your score maps to a stage on the talent ladder, but the dimension breakdown is the useful part: inflow problems (sourcing, embedding) and outflow problems (renewal, capture) need different interventions, and the lowest dimension names yours. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the talent ladder, but the dimension breakdown is the useful part: inflow problems (sourcing, embedding) and outflow problems (renewal, capture) need different interventions, and the lowest dimension names yours.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the workforce numbers verified rather than self-reported?
We run the same assessment as a structured review with your engineering, HR and operations leads — time-to-fill from your ATS, attrition from your HR system, the single-point-of-knowledge register built live in the room — and leave you a costed plan for the weakest dimension. You keep the plan either way.
How the score maps to a stage
- 0–4 — Stage 1, Rented. Rented is the stage where every AI-relevant skill sits outside the organisation — vendors, consultants and integrators do the work, and the capability leaves when the contract ends.
- 5–9 — Stage 2, Enclave. Enclave is the stage where an internal data team exists but operations does not trust it — capability has been hired into a central unit that the control room, depots and planning teams treat as a foreign country.
- 10–14 — Stage 3, Blended. Blended is the stage where paired pods of domain engineers and data specialists own operational outcomes together, and the first grid-specific hybrid roles — OT-cleared data engineers, model-literate asset engineers — exist and are trusted.
- 15–19 — Stage 4, Pipelined. Pipelined is the stage where talent supply is repeatable: a conversion pathway turns the utility's own engineers into hybrid roles on a schedule, external intake feeds the bottom, retention economics are solved deliberately, and knowledge capture runs against the retirement forecast rather than behind it.
- 20–24 — Stage 5, Self-renewing. Self-renewing is the stage where the workforce develops the workforce: AI literacy sits in ordinary role profiles, operational staff build and maintain models inside governed bounds, and competence is managed as an audited system rather than a programme — so the capability no longer depends on any individual, initiative or budget line.
What the AI readiness talent gap in utilities actually is
A definition, the two directions the gap runs in, and the flow diagram most readiness assessments never draw — the one with people in it.
The AI readiness talent gap in utilities is the distance between the workforce a utility employs and the workforce its AI programme requires — measured in roles it cannot fill, conversions it has not started, and expertise it is about to lose. It is the reason most utility AI readiness work fails its own test: assessments audit data platforms, integration estates and governance frameworks in detail, then compress the entire human question into a slide titled 'change management'. Yet every capped programme we have seen was capped by a person-shaped hole — the OT-cleared data engineer who took nine months to hire, the control-room product owner who cannot be hired at all, the transformer diagnostician who retired in March with thirty years of judgement uncaptured.
What makes the utility version of this gap different from every other industry's is that it runs in two directions at once. The inflow side is the familiar one: data engineers, ML engineers and MLOps specialists are scarce everywhere, and a regulated utility — with its pay bands, its security perimeter and its clock speed — competes for them at a structural disadvantage. The outflow side is the one the sector owns almost alone: the workforce that holds the network's tacit knowledge is old, and retiring. The IEA's World Energy Employment report (opens in a new tab) puts the global energy workforce at around 67 million people and documents the sector's persistent struggle to attract skilled workers — while inside any individual utility, the engineers who can read a dissolved-gas result against a loading history, or feel when a storm forecast means calling crews early, are disproportionately in their final decade of service. AI readiness depends on both flows: the new skills arriving, and the old judgement being captured before it walks out.
How AI capability enters a utility — and how grid knowledge leaves it
The two flows a talent-readiness plan must manage at once. The top lane is the inflow most programmes obsess over; the bottom lane is the outflow most programmes ignore. Both terminate in the blended core — or they terminate in value leaks.
- Data & feeds
- Human in the loop
- Where value leaks
- AI / model
- System-of-record action
The process, in words
- Capability flows in through three channels — external hires, partners, and internal conversion — and every external channel passes through the clearance-and-context bottleneck: months of security vetting and OT familiarisation before a hire is genuinely productive near SCADA. Conversion of the utility's own engineers bypasses that bottleneck, which is why it is the most underused high-leverage channel.
- Both flows land in the blended core: pods pairing domain engineers with data specialists, owning one operational metric together, shipping models into the surfaces the operation already uses — the ADMS, the EMS, the EAM. The pod is simultaneously the delivery unit and the classroom in which hybrid people are made.
- Knowledge flows out through retirement on a schedule nobody controls. The only intervention that works is capture while the expert is still on payroll — co-labelling historical cases so the judgement becomes a dataset a pod actively consumes. The dashed edge to retirement is the leak: expertise leaving uncaptured, which no later hire at any salary can restore.
Step-by-step insights
- The external market — necessary, and structurally stacked against you
- A utility hiring a data engineer competes with software firms on pay, with startups on excitement, and with consultancies on variety — and typically loses on all three. What it can win on is meaning and moat: the grid is the most consequential machine most engineers will ever touch, and grid context compounds into a career asset that generic data roles never build. Utilities that lead their recruitment with the mission and the craft — not the pension — measurably out-hire those that run the standard advert. But the market channel should still be reserved for the roles conversion cannot produce: the seed data engineer, the MLOps platform hand, the first governance lead.
- Partners and consortia — renting the frontier without re-renting the basics
- The borrow channel is legitimate and permanent: no single utility should train its own foundation models or chase the research frontier alone, and industry pooling — EPRI's Open Power AI Consortium is the visible example — exists precisely to share that cost. The discipline is the boundary: borrowed capability must never sit on the critical path of daily operations, and every engagement must leave residue — code the utility owns, definitions its people understand, at least one internal engineer who paired on the work end to end. Borrow the frontier; never re-rent the basics you already learned.
- Internal conversion — the channel with the unfair advantage
- A protection engineer who learns Python brings fifteen years of network judgement to every feature they build; a data-science graduate learns the same judgement over a decade, if ever. Conversion flips the scarcity: instead of competing in the world's tightest labour market for people who lack your context, you develop people who already hold the context, already hold the clearances, and already have the control room's trust. The costs are real — backfilling their old role, paying for the new one properly, tolerating a slower ramp — but they are a fraction of the fully loaded cost of the external alternative, and the retention profile is dramatically better.
- The clearance bottleneck — why the first hire starts six months early
- Working OT-adjacent is gated, and rightly: in North America the NERC CIP standards make personnel risk assessment and training a compliance obligation around the bulk electric system, and equivalent security regimes apply elsewhere. Sourcing, vetting, CIP-style training and supervised familiarisation stack into a six-to-nine-month runway between requisition and genuine productivity near SCADA. Programmes that discover this on the critical path lose two quarters; programmes that start the hire at business-case stage — before the funding lands — hide the runway entirely. It is the single most schedule-critical fact in this entire page.
- Co-labelling — the only capture that survives contact with a model
- Most knowledge capture produces documents, and documents do not train models or successors — they decay unread in the EAM. Co-labelling produces data: the retiring diagnostician works through hundreds of historical cases — DGA panels, thermal images, failure records — recording the call they would have made and why, with a pod engineer beside them encoding the reasoning into features and labels. The output is threefold: a labelled corpus the current model trains on, a validation set future models are judged against, and a successor who spent two hundred cases' worth of hours inside the expert's head. It is the highest-value engineering work most utilities have never scheduled.
The five stages of the utility AI talent ladder
Rented, Enclave, Blended, Pipelined, Self-renewing — for each stage: what it looks like from the control room, the signals a reviewer can check in an afternoon, the anti-pattern that traps utilities there, and what leaving costs.
The talent ladder runs from capability you rent to capability that renews itself, and every utility we have worked with sits identifiably on one of its five rungs. Each stage below is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own organisation this week — most of them answerable from the ATS, the HR system and one conversation with a shift engineer — and the anti-pattern is the specific mistake most often made trying to climb out. Note what the ladder is not: it is not a technology maturity model. A utility can run a modern data platform from stage 2, and a spreadsheet estate from stage 3. The ladder measures people, trust and renewal — the things platforms cannot substitute for.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Rented
26% of operators sit here
Rented is the stage where every AI-relevant skill sits outside the organisation — vendors, consultants and integrators do the work, and the capability leaves when the contract ends.
Rented is not a failure state — it is where almost every utility rationally starts, because the first pilots arrive attached to vendor products and consultancy engagements. The failure is staying there. A rented capability produces demonstrations, reports and occasionally a working model, but it produces no residue: no internal engineer who understands the feature pipeline, no code the utility can modify, no judgement about what worked that survives the supplier rotating its team. The tenth engagement costs what the first did, and teaches the organisation exactly as little.
The tell is the renewal meeting. At Rented, the conversation about year two of an AI contract contains no utility voice that can challenge the supplier technically — procurement negotiates price because nobody can negotiate substance. That asymmetry compounds: the supplier's people learn the network, the feeder quirks, the data gotchas, and then take that learning to their next client. The utility has paid to train the market's engineers on its own grid.
What makes Rented dangerous rather than merely slow is the outflow side. While the capability is rented, the workforce that holds the network's tacit knowledge — the protection engineer who can read a DGA result like a novel, the control-room veteran who knows which alarms lie — keeps ageing. Every year at Rented is a year of retirements uncaptured, and that loss, unlike a delayed hire, cannot be bought back at any price.
In practice
The model the supplier took home
A mid-sized DNO commissioned a consultancy to build a fault-prediction proof of concept on overhead lines. The model worked; the final presentation impressed the board. Eighteen months later a new asset-management director asked to extend it to cable networks and discovered the utility owned a slide deck: the code sat in the supplier's repository, the feature definitions in a departed contractor's head, and the two engineers who had fed the project data had both retired. Extending it meant starting again — at full price, with a new supplier.
What it looks like
- All modelling and data engineering is delivered by vendors or consultants
- No internal role profile mentions data or ML competence
- Nobody in-house can reproduce, retrain or even rerun what a supplier built
- Retirement dates are tracked by HR; the knowledge behind them is tracked by nobody
Diagnostic signals you can check this week
- Ask who in-house could retrain the most recent vendor-built model. If the answer is a company name, you are at Rented
- Check whether any internal job description requires data or ML competence. At Rented, none does
- Ask what artefacts the last AI engagement left behind — code, data definitions, runbooks — and who can use them
- Ask for the list of judgements that depend on one named person. At Rented the list does not exist
Anti-pattern · Hiring a head of AI first
The instinctive exit from Rented is a senior hire — a head of AI or chief data officer — brought in to 'build the capability'. Landed into an organisation with no engineers to lead, no data plumbing and no operational sponsor, the hire spends a year writing strategy, loses the confidence of the executive that hired them, and leaves. The sequence is backwards: the first hires are doers with a bounded problem — one data engineer, one converted domain engineer, one use case — and leadership is added once there is something to lead.
What holds you here
There is no internal skill to accumulate learning on, so every engagement resets to zero — and the retirement clock keeps running.
Highest-leverage next move
Build the single-point-of-knowledge register, write the first two role profiles, and start the OT data-engineer hire now — it is the longest lead item on the whole roadmap.
Cost of leaving
- Effort
- 6–9 months to a defensible Enclave-or-better position
- Team
- One OT-adjacent data engineer (external hire), one seconded domain engineer, a named operations sponsor
- Risk
- Low — the work is additive; the real risk is another year of uncaptured retirements
- To next stage
- 6–9 months
If this is you, the next step is
A two-week engagement: the single-point-of-knowledge register, the first role profiles, and the sourcing plan.
Stage 2
Enclave
37% of operators sit here
Enclave is the stage where an internal data team exists but operations does not trust it — capability has been hired into a central unit that the control room, depots and planning teams treat as a foreign country.
Enclave is the most common stage and the least stable. The utility has done the hard fundraising work — headcount approved, a data science team hired, often at salary bands that required an uncomfortable HR exception — and the team is genuinely capable. What it lacks is a route into the operation. Its members sit in a corporate office, work from data extracts, and present findings to committees. The models are competent; the network carries on exactly as before. Six months of this teaches operations that the team is decorative, and teaches the team that the utility was not serious — which is when the CVs go out.
The attrition mechanics deserve more attention than they get. A data scientist inside a utility enclave is being paid below the market that recruited them, working on problems that never ship, and watching former colleagues at software firms deploy weekly. The utility's traditional retention levers — stability, pension, a thirty-year career arc — are precisely the ones this cohort discounts. The result is a two-year revolving door: hire at a premium, lose before the first production deployment, re-hire at a bigger premium. The enclave does not just underdeliver; it actively burns the employer brand needed to staff stage 3.
The exit is not more hiring — it is dissolution. The enclave dissolves into blended pods: one data engineer and one modeller paired with a protection engineer, an asset manager or a shift engineer, sitting with the operational team, owning one operational number together. Utilities resist this because it feels like breaking up the expensive thing they just built. It is actually the first moment the expensive thing starts to pay, and — the part nobody expects — the first moment the data specialists stop resigning, because their work finally lands on a real grid.
In practice
The forecasting team the control room never called
A vertically integrated utility built a five-person analytics team that produced, among other things, a genuinely good day-ahead demand forecast. The control room continued running on the incumbent vendor forecast, because the analytics team's version arrived by email, carried no accountability, and nobody standing a shift had ever met its authors. When the team's lead modeller resigned for a trading house, her handover note observed that in two years no operational decision had changed because of anything the team produced. The forecast was excellent. It was also irrelevant.
What it looks like
- A central data science or innovation team exists, typically 3–10 people
- Its backlog is set by an innovation board, not by operational pain
- Control-room and field engineers describe the team as 'them', not 'us'
- Attrition in the data team runs well above the utility's average
Diagnostic signals you can check this week
- Ask a shift engineer to name one member of the data team. At Enclave, they cannot
- Compare the data team's attrition rate with the engineering average — at Enclave it is typically double or worse
- Check where the data team sits — physically and organisationally. Corporate floor, corporate reporting line: Enclave
- Count operational decisions changed by the team's output in the last quarter. Zero is the Enclave signature
Anti-pattern · Fixing the enclave with a bigger backlog
When an enclave underdelivers, the reflex is to feed it better projects — an innovation funnel, hackathons, an executive-sponsored use-case pipeline. This treats the problem as demand when the problem is trust. Operations does not withhold cooperation because it lacks ideas; it withholds cooperation because the team has no skin in operational outcomes and no member who has ever stood a storm shift. No backlog fixes that. Pairing fixes it — one pod, one operational metric, shared on-call for the thing they ship.
What holds you here
Capability and credibility live in different buildings: the data team has no route into operations, and operations has no reason to open one.
Highest-leverage next move
Dissolve the enclave into its first blended pod — one data engineer, one modeller, one domain engineer, one operational number — and move their desks to the operational team.
Cost of leaving
- Effort
- 9–15 months to blend the first two pods and prove one operational result
- Team
- Existing data team, two seconded domain engineers at 50%+, one control-room or depot 'landlord' per pod
- Risk
- Medium — the first pod must land a visible win or the blending model is discredited with it
- To next stage
- 9–15 months
If this is you, the next step is
Half a day with your data lead and operations lead: the problem, the pairing, the metric, the seating plan.
Stage 3
Blended
22% of operators sit here
Blended is the stage where paired pods of domain engineers and data specialists own operational outcomes together, and the first grid-specific hybrid roles — OT-cleared data engineers, model-literate asset engineers — exist and are trusted.
Blended is where the talent programme starts producing compound interest. The pod structure does two jobs at once: it ships operational results, and it manufactures exactly the hybrid people the market cannot sell you. Twelve months of pairing turns a data engineer into someone who understands why the feeder naming convention lies, and turns a protection engineer into someone who can read a confusion matrix. Neither conversion happens in a classroom. The pod is the classroom, and its output — beyond the model — is the first generation of staff who hold both worlds.
The binding constraint at this stage is access, not skill. A data engineer cannot build governed feeds from SCADA and the historian without crossing the OT security perimeter, and utilities are rightly conservative about who crosses it — in North America the NERC CIP obligations make that caution a compliance matter, not a preference. Clearing a data engineer to work OT-adjacent takes months of vetting, training and supervised access, which is why the role has to be hired and cleared long before the programme needs it, and why poaching someone already cleared commands the premium it does.
Blended is also the stage where the outflow side of the gap must be wired in, because pods create the first genuine consumers for captured knowledge. A retiring diagnostic engineer's judgement is worth little as a document in the EAM; it is worth a great deal as a labelled dataset a pod is actively training on. The discipline is co-labelling: sitting the expert with the pod over historical cases — DGA results, thermal images, failure records — and recording their calls and their reasoning while they are still on payroll. Utilities that skip this at stage 3 arrive at stage 4 with pipelines and platforms and nobody left who can tell the model when it is wrong.
In practice
The pod that out-recruited the enclave
A transmission operator moved two of its four data scientists to sit with the substation asset team, paired with a senior asset engineer a year from retirement, with one shared target: cut unplanned transformer outages on a named fleet. In nine months the pod shipped a condition-ranking model into the EAM's maintenance-planning screen, co-labelled eleven years of DGA history with the retiring engineer, and — the result nobody projected — received four internal transfer requests from engineers who wanted the next pod seat. The same utility had spent two years failing to fill a data science vacancy by advertisement.
What it looks like
- At least one pod pairs data specialists with domain engineers on a shared operational metric
- An OT-cleared data engineer works inside the control-network perimeter, with security's blessing
- Converted domain engineers — protection, planning, asset — do model work part-time and are paid for it
- The control room consults model output because a named engineer it knows answers for it
Diagnostic signals you can check this week
- Ask whether any data specialist has OT-network access agreed with security. Blended requires at least one
- Ask a pod's domain engineer what precision and recall mean, and its data engineer what a DGA result is. Both should manage
- Check whether converted domain engineers were re-graded or re-titled. Unpaid conversion is unfinished conversion
- Look for a co-labelling artefact — a labelled historical dataset with the expert's reasoning attached. Its absence means the outflow side is still open
Anti-pattern · Scaling pods faster than the trust that feeds them
After the first pod lands a win, the temptation is to stamp out five more by reorganisation — assign names to pods on a slide and declare the operating model transformed. But a pod is not a seating plan; it is a trust relationship with a specific operational team, built through one delivered result. Pods declared faster than results are pods in name only, and the operational teams assigned to them learn the initiative is theatre. Grow pods at the rate results land — roughly one new pod per proven pod per year — and staff each new one with at least one veteran of a working pod.
What holds you here
Everything runs through a handful of hybrid people the market cannot replace — one resignation from the first pod can undo a year.
Highest-leverage next move
Turn the conversions the pods produced by accident into a deliberate pathway — named curriculum, re-grading, a cohort a year — and start the graduate and apprentice intake that feeds it.
Cost of leaving
- Effort
- 12–24 months to a repeatable pod model and the first internal conversion cohort
- Team
- Two to four pods; one platform/MLOps engineer shared across them; an L&D partner for the conversion pathway
- Risk
- Medium — key-person risk concentrates in the first OT-cleared engineer and first pod leads
- To next stage
- 12–24 months
If this is you, the next step is
We review pairing, access, pay and capture against what has actually made pods survive elsewhere.
Stage 4
Pipelined
11% of operators sit here
Pipelined is the stage where talent supply is repeatable: a conversion pathway turns the utility's own engineers into hybrid roles on a schedule, external intake feeds the bottom, retention economics are solved deliberately, and knowledge capture runs against the retirement forecast rather than behind it.
Pipelined is the stage most talent strategies claim and few operations can evidence. The test is unglamorous: when a pod loses its data engineer, how long until a replacement of equal usefulness is productive? At stage 3 the honest answer is six months to a year of external search plus clearance plus familiarisation. At Pipelined it is weeks, because the replacement already exists inside the organisation — a conversion-cohort graduate who knows the estate, holds the clearance, and has sat in a pod as an apprentice member. The pipeline's product is not headcount; it is replaceability without loss.
The economics need to be faced squarely rather than managed by exception. A utility cannot and should not match trading-house or big-tech pay for pure ML talent — but it does not need to, because its hybrid roles are not fungible with those markets. An OT-cleared data engineer with three years of grid context has a career moat a generic data engineer lacks, and the utility can pay a defensible premium for the hybrid — benchmarked against what it would actually cost to replace them, not against the engineering band next door. Utilities that benchmark hybrid pay against other utilities' generic bands are running a subsidised training scheme for consultancies.
On the outflow side, Pipelined means capture has become boring — which is the goal. The register of single points of knowledge is reviewed quarterly alongside the retirement forecast; every critical judgement has a capture plan with a date; co-labelling sessions are scheduled work with work orders, not favours extracted in notice periods. The wider industry has begun building shared scaffolding for exactly this — EPRI's Open Power AI Consortium pools utilities' effort on power-sector AI models and practices — but the tacit knowledge of a specific network can only be captured by the utility that runs it, while the people who hold it are still on payroll.
In practice
The cohort that closed a vacancy in three weeks
A distribution utility ran its second annual conversion cohort — six engineers from protection, planning and asset management, each spending two days a week for nine months on data engineering and ML fundamentals, each attached to a working pod, each re-graded on completion. When the storm-response pod's data engineer left for a software firm, the vacancy was filled in three weeks by a cohort graduate from network planning, already CIP-trained, already known to the control room. The external advert that would once have run for eight months was never posted.
What it looks like
- A named conversion pathway — curriculum, mentoring pod seat, re-grading — graduates a cohort a year
- Graduate and apprentice intake includes data-and-grid hybrid schemes, not just traditional disciplines
- Pay for hybrid roles is benchmarked against the market that actually poaches them, and refreshed annually
- The single-point-of-knowledge register drives scheduled capture, prioritised by retirement date and criticality
Diagnostic signals you can check this week
- Ask for the conversion pathway's completion and retention numbers. A pathway without numbers is a slide
- Check the last pay-benchmark refresh for hybrid roles, and what market it was benchmarked against
- Ask how the next twelve months of retirements map to capture plans. Pipelined has a dated answer per name
- Time-to-fill for pod vacancies: weeks from internal supply is Pipelined; months of external search is not
Anti-pattern · Outsourcing the pipeline to a training vendor
At this stage a procurement instinct kicks in: buy a 'data academy' from a training vendor, put two hundred staff through it, report the throughput to the board. Coursework without a pod seat and a re-graded role to land in converts nobody — it produces certificates, a brief morale bump, and cynicism when the promised transformation does not follow. The pathway's scarce ingredient was never content, which is abundant; it is the pod seat with mentoring and the re-grading at the end. Size cohorts to the pod seats available, not to the licence count the vendor discounts at.
What holds you here
The pipeline is funded as an initiative rather than as infrastructure, so it survives exactly until the first difficult budget round.
Highest-leverage next move
Move competence from programme to management system: role profiles, training records and capture obligations written into the operating model and audited like any other licence obligation.
Cost of leaving
- Effort
- 18+ months to make supply demonstrably repeatable across two cohort cycles
- Team
- Programme lead, L&D partner, pod leads as mentors, HR reward partner for benchmarking and re-grading
- Risk
- Medium — pipelines are quietly starved in cost-cutting rounds because their loss shows up two years later
- To next stage
- 18+ months
If this is you, the next step is
Two days: conversion throughput, benchmark drift, capture backlog — and where the next cut would actually land.
Stage 5
Self-renewing
4% of operators sit here
Self-renewing is the stage where the workforce develops the workforce: AI literacy sits in ordinary role profiles, operational staff build and maintain models inside governed bounds, and competence is managed as an audited system rather than a programme — so the capability no longer depends on any individual, initiative or budget line.
Self-renewing is narrower and more procedural than the name suggests. It does not mean every lineworker writes Python; it means the organisation has stopped treating AI competence as a special commodity held by special people. A planning engineer extending a load-forecast feature, a substation engineer flagging a mislabelled training case through the same system they would use for a defect, an apprentice rotating through a pod as routinely as through a depot — these are the textures of stage 5. The specialist roles still exist; what has changed is that the organisation around them can absorb their departure without a programme review.
The governance scaffolding is what makes this safe rather than chaotic. Management-system standards for AI — ISO/IEC 42001 is the reference point — treat competence the way safety management always has: the organisation must determine what competence each AI-touching role requires, ensure it exists, and keep the evidence. For a utility this is familiar territory; it already runs authorisation regimes for switching, for working near live equipment, for control-room roles certified under NERC's operator-certification scheme in North America. Stage 5 simply extends a discipline the industry has practised for a century — competence as a controlled, audited property of roles — to the model estate.
Sustaining stage 5 is a renewal discipline, and regression is quiet. The signature failure is competence rot: role profiles written in the stage-4 push slowly drifting out of date while the model estate grows, until an incident review discovers that the person who approved a model change held no current competence for it. The countermeasure is the same as for any management system — periodic review keyed to change, not to calendar piety: new model classes, new platform capabilities and reorganisations each trigger a competence-profile review, exactly as a network change triggers a protection-settings review.
In practice
The retirement that was a calendar entry
At a utility operating at stage 5, the retirement of the senior cable-diagnostics engineer — the kind of departure that at stage 1 erases a capability — was a calendar entry. His judgement on partial-discharge interpretation had been co-labelled into the diagnostic corpus over three years; two conversion-cohort engineers held current competence sign-off on the model family; the leaving presentation included a chart of his labelled cases as a proud artefact. The following month, the model caught a cable-joint signature one of his successors confirmed — using the criteria he had taught the dataset.
What it looks like
- Model-literacy expectations appear in ordinary engineering role profiles, not specialist ones
- Operational teams extend and retrain models themselves, within a governed platform and policy
- Competence, training and capture obligations are part of the management system and get audited
- Leaving events — resignation or retirement — are routine, because succession and capture are continuous
Diagnostic signals you can check this week
- Pick three AI-touching roles and ask for their competence profiles and evidence. Stage 5 produces them like maintenance records
- Ask when a competence profile was last updated because the model estate changed — not because a year passed
- Check whether any operational team retrained or extended a model without the central team in the loop, legitimately
- Ask what happened at the last senior technical retirement. 'Routine' is the stage-5 answer
Anti-pattern · Mistaking a mature platform for a mature workforce
Stage-5 platforms make model work look easy, and the anti-pattern is reading platform maturity as workforce maturity — thinning the specialist core, pausing the cohorts, letting the capture cadence slip, because 'the platform handles it'. The platform automates yesterday's competence; it cannot renew it. Two years of this leaves an estate of models nobody currently employed can deeply defend — a Rented stage with better tooling. The pipeline and the capture discipline are not scaffolding to remove at maturity; they are the maturity.
What holds you here
Nothing external blocks stage 5 — the threat is internal: competence rot, quietly, while the dashboards stay green.
Highest-leverage next move
Key competence review to change: every new model class, platform capability or reorganisation triggers a profile review, exactly as a network change triggers a protection review.
Cost of leaving
- Effort
- Continuous — renewal keyed to change in the model estate and the workforce
- Team
- Standing competence authority (as for switching authorisations), pod network, cohort programme as business-as-usual
- Risk
- Concentrated in complacency — the failure mode is slow, silent competence rot, discovered by incident
If this is you, the next step is
We audit AI-role competence profiles and evidence the way a safety auditor reads authorisation records.
Where utilities sit on the ladder today
The distribution across the five stages, why the Enclave is the mode, and what the wider labour market does to every rung.
Most utilities sit at the Enclave stage — a hired central team that operations has not adopted — and the distribution falls away steeply above it. That shape is worth dwelling on, because it differs from the technology-maturity distributions the industry is used to seeing: utilities have collectively bought a great deal of AI tooling, but tooling migrates between stages in a procurement cycle while people migrate in career-lengths. The result is a sector whose platform maturity now runs a full stage ahead of its workforce maturity almost everywhere — and the workforce number is the binding one.
Distribution of utilities across the five talent stages
Illustrative distribution. The Enclave is the mode and the trap: it is where investment has already happened but value has not, which makes it both the most common and the most politically fragile place to be stuck.
Share of utilities
- 26% — 1 · Rented
- 37% — 2 · Enclave (the trap)
- 22% — 3 · Blended
- 11% — 4 · Pipelined
- 4% — 5 · Self-renewing
Source: Illustrative distribution, synthesised from IEA workforce research and McKinsey AI adoption surveys
The macro forces behind the distribution all push the same way. AI adoption has gone mainstream across industries — McKinsey's State of AI research (opens in a new tab) has tracked the share of organisations using AI in at least one function climbing to more than three-quarters — which means utilities now compete for the same talent as every other sector, rather than against each other. Meanwhile the energy transition is expanding the sector's own headcount needs: IRENA and the ILO count 16.2 million renewable-energy jobs worldwide (opens in a new tab) and rising, and European industry bodies such as Eurelectric (opens in a new tab) have made workforce and skills a standing policy theme for exactly this reason. A utility talent strategy written as if the competition were the neighbouring utility is benchmarking against the one market that is not the threat.
6–9 mo
realistic runway from requisition to a productive OT-cleared data engineer, once vetting and familiarisation are counted (illustrative; see the role ledger)
2×
typical attrition in enclave-stage data teams relative to the utility's engineering average (illustrative pattern across published attrition research)
37%
of utilities sitting at the Enclave stage in the illustrative distribution above — investment made, value not yet landed
The role ledger: ten roles, their scarcity, and the build–buy–borrow call
The roles a utility AI programme actually runs on — what each one really does, how hard it is to get, how long until it is productive, and which sourcing channel is the honest default.
A utility AI programme runs on about ten roles, and the talent gap is really ten separate gaps of very different depths. Treating them as one 'AI skills' shortage produces the classic failure: a hiring plan that buys what should be built, builds what could be bought cheaply, and never notices that its two most critical roles cannot be sourced externally at all. The ledger below is the planning tool we use with utility leadership teams — for each role, what the work actually is, how scarce it genuinely is in the market, how long from start date to real productivity inside a regulated estate, and the sourcing call that usually survives contact with reality.
| Role | What they actually do | Market scarcity | Time to productive | Default call |
|---|---|---|---|---|
| OT-cleared data engineer | Builds governed feeds from SCADA, the historian and AMI into the analytics estate — inside the security perimeter, with the OT team's trust | Severe — cleared-and-experienced is a seller's market | 6–9 months | Buy one to seed, then build the rest |
| Power-systems ML engineer | Models load, DER output and network constraints with the physics respected — knows why a naive forecast violates thermal limits | Severe — the intersection is tiny | 6–12 months | Build: convert a power engineer, add ML |
| Asset-health data scientist | Condition models from DGA, thermal imagery, partial discharge and load history; works to failure modes, not just features | High | 3–6 months | Build from asset engineering |
| MLOps / platform engineer | Serving, monitoring, retraining pipelines; makes model operations boring — the discipline transfers from any industry | High, but transferable | ~3 months | Buy |
| Control-room product owner | Owns the operator-facing decision surface; an ex-shift engineer the room still trusts — decides what an operator sees and when | Severe — cannot be bought externally at all | ~12 months (internal) | Build only |
| GIS / network-model steward | Keeps the as-operated network model joined to the asset register and the EAM — the join every grid model silently depends on | Moderate | 3–6 months | Build |
| Forecasting analyst | Demand, generation and imbalance forecasting into operations and trading; owns forecast bias as a number with a name on it | Moderate | ~3 months | Build or buy |
| AI governance lead | Model risk, audit evidence, alignment to NIST AI RMF and ISO/IEC 42001; speaks regulator fluently | High | ~6 months | Buy or borrow |
| Knowledge engineer | Runs the capture programme: co-labelling sessions, judgement elicitation, the single-point-of-knowledge register | High — the title barely exists yet | 3–6 months | Build from L&D or engineering |
| Data product manager | Turns use cases into owned, funded products inside a regulated estate; manages the seam between pods and the business | High | ~6 months | Buy |
Two of the ten gate everything else, and neither is the data scientist. The OT-cleared data engineer gates the inflow: until governed feeds exist from SCADA, the historian and AMI, every other role is working on exports — and the six-to-nine-month runway to create this person makes them the schedule-critical hire of the entire programme, which is why the requisition belongs at business-case stage, not at kick-off. The control-room product owner gates the outcome: models change nothing until an operator acts on them, and the person who can design that surface — and answer for it to a room they once stood shifts in — exists only inside your own organisation, a year of deliberate development away. In North America there is an instructive precedent for how seriously the industry can take role-competence when it chooses to: NERC (opens in a new tab) runs a formal certification regime for system operators. The hybrid roles on this ledger deserve the same seriousness, and at stage 5 they get it.
Build, buy or borrow — deriving the call for any role
Plot any role by how scarce it is in the market and how grid-specific its value is. The quadrant gives the honest sourcing default — and explains why the two most critical utility AI roles can only be built.
Build (steady state)
- Asset-health data scientist, GIS steward, forecasting analyst
- Context matters more than market heat
- Conversion pathway at cohort pace
Build long, borrow the bridge
- OT-cleared data engineer, power-systems ML engineer, control-room product owner
- The market cannot sell you these — grow them
- Partner cover while the pathway runs, never on the critical path
Buy
- MLOps engineer, data product manager
- Transferable craft, normal market
- Standard hire; pay the going rate and move on
Borrow
- Frontier ML, LLM specialists, one-off research
- Scarce everywhere, generic everywhere
- Partners and consortia — with a residue clause in every contract
The outflow side: capturing grid knowledge before it retires
Why the retirement wave is an AI-readiness problem, why conventional knowledge capture fails, and the engineering discipline that actually preserves judgement.
Grid knowledge leaves through retirement on a schedule no programme controls, and in most utilities it leaves uncaptured — which makes the retirement wave a first-order AI readiness problem, not an HR talking point. The judgement at risk is precisely the judgement AI systems need most: which historical failures looked like this DGA signature, which alarms in this substation are known liars, what the network actually does — as opposed to what the GIS says it does — on the rural spurs rebuilt after the 2007 storms. Models trained without this judgement learn the paper network; operators who hold the judgement then correctly distrust the models, and the programme stalls with everyone behaving reasonably. Regulators have seen the workforce risk clearly enough that GB network price controls have required licensees to evidence workforce resilience plans (opens in a new tab) as part of their submissions — but a resilience plan that ends at succession charts still captures nothing a successor or a model can use.
Capture fails when it is run as an HR process
Exit interviews, handover documents and leaver checklists are designed to close accounts, not to preserve judgement. Thirty years of diagnostic craft cannot be written down in a notice period, and what does get written lands in a document-management system where no model and few successors will ever read it. The failure is structural: the process is owned by the wrong function, scheduled at the wrong time, and produces the wrong artefact.
Capture fails when it produces documents instead of data
A document describes judgement; a labelled dataset embodies it. The test for any capture exercise is brutal and simple: could a model train on the output, and could a successor be assessed against it? Prose in the EAM fails both. Two hundred historical cases labelled with the expert's call and reasoning pass both — and take roughly the same number of expert-hours to produce.
Capture fails when nothing consumes it
Capture without a consumer is archaeology in advance. The knowledge-capture exercises that survive are the ones wired into a live use case: a pod actively training a condition model on the labelled corpus, a conversion cohort being assessed against the expert's calls, a validation set gating the next model release. The consumer creates the pull that keeps capture funded — which is why capture belongs inside the AI programme, not beside it.
Knowledge capture that feeds a model, not a binder
Build the single-point-of-knowledge register
One workshop per operational domain — asset management, control room, field operations — listing every judgement that currently depends on a named individual, scored by criticality and years-to-retirement. This register, not the org chart, is the real map of your outflow risk, and it doubles as the prioritisation queue for everything below. Most utilities are genuinely surprised by what surfaces: the register typically finds two or three single points of knowledge nobody in leadership had identified.
Schedule co-labelling as engineering work
For each critical judgement, sit the expert with a pod engineer over the historical record — DGA panels, protection operations, storm logs — labelling case by case: the call, the confidence, the reasoning, the tells. Budget real hours against work orders; a meaningful corpus is two hundred cases and change, which at a disciplined half-day cadence is a quarter's work alongside the day job. The corpus becomes training data, validation gold-standard and assessment material in one artefact.
Pair the successor before the leaving date, not after
Co-labelling with a successor in the room is succession planning that actually transfers the judgement: every disagreement between expert and successor is a teaching moment recorded as data. The successor does not need to be a data specialist — they need to be the person who will answer the questions the expert answers today, with the corpus behind them and a model beside them.
Make the model the living archive
The final step inverts the usual relationship: instead of the expert's knowledge being archived and forgotten, it lives inside a model that operations consults daily, with the labelled corpus as its audit trail. When the model flags a transformer and a successor confirms it using the criteria the corpus taught, the retired engineer's judgement is still doing shifts. That — not a leaving presentation — is what captured knowledge looks like.
What talent-led programmes look like in public
Three publicly reported approaches, read against the talent ladder. None is an Atomic Loops engagement — each links to the organisation's own published material.
The clearest public evidence for the talent thesis is in where capability actually came from at the organisations that made AI stick. In each case below the differentiating move was a people decision — who was converted, who was paired with whom, what was pooled rather than duplicated — and the technology followed. Outcomes are as reported by each organisation in its own material; where the organisation has not published verifiable figures, we keep the outcome qualitative rather than invent precision.
Three sourcing strategies, read against the ladder
Read against the talent ladder: build (Duke), pair (PG&E), pool (EPRI's consortium). Outcomes as reported by the organisations themselves — verify against the linked source before reusing figures. The EPRI card uses an industry-scene image from our generated library, not an EPRI photograph.
Duke EnergyUS investor-owned utility · millions of electric and gas customers14
- Challenge
- Predictive analytics across a very large generation fleet demanded people who understood both the machines and the models — a profile the external market does not sell.
- Approach
- Duke built its fleet monitoring-and-diagnostics capability around experienced plant and equipment engineers, centralised into a monitoring centre and equipped with analytics — converting deep domain judgement into analytical roles rather than hiring analytics and hoping the domain would follow.
- Reported outcome
- Duke Energy has publicly credited its centralised monitoring-and-diagnostics approach with early catches of developing equipment problems across its fleet, avoiding major failures and unplanned outages — as described in its own newsroom and fleet-technology communications.
- What it shows about the curveThe build channel works at scale: engineers who already held the machine judgement became the analysts, and the trust problem that sinks enclave teams never arose — the fleet already knew them. That is the tl-quadrant of the sourcing matrix, executed over a decade.
Pacific Gas and ElectricCalifornia investor-owned utility · ~16 million people served24
- Challenge
- Wildfire-driven inspection programmes produced millions of asset images — far beyond what inspection engineers could review manually, and far beyond what data teams could interpret without them.
- Approach
- PG&E has publicly described pairing its inspection and engineering expertise with machine-learning image triage — models surfacing candidate defects, inspection specialists making the calls, their corrections feeding back — and has since extended AI assistance into standards lookup and customer operations.
- Reported outcome
- PG&E has publicly reported using AI-assisted review to triage inspection imagery at a scale manual review could not reach, with expert review retained on the safety-critical calls — as described in the company's own innovation and wildfire-safety communications.
- What it shows about the curveThe pairing pattern is the Blended stage in production: the model does the volume, the domain expert does the judgement, and every correction is training data. The inspection specialists were not displaced by the model — they became its teachers, which is co-labelling under operational pressure.
EPRI Open Power AI ConsortiumIndustry research institute · utilities and technology partners, pooling AI capability12
- Challenge
- No individual utility can staff frontier AI work — domain-adapted models, evaluation benchmarks, emerging practice — and duplicating the attempt across hundreds of utilities multiplies the same scarce-talent problem industry-wide.
- Approach
- EPRI convened utilities and technology companies into the Open Power AI Consortium to develop open, power-sector-specific AI models and shared practice — pooling exactly the work that is scarce everywhere and grid-specific nowhere, so member utilities can point their own people at their own networks.
- Reported outcome
- The consortium operates as a named, public membership programme developing domain-specific open models and resources for the power sector — as described in EPRI's own consortium material.
- What it shows about the curvePooling is the br-quadrant of the sourcing matrix institutionalised: borrow the frontier collectively. The boundary the ladder insists on still applies — consortium membership raises a utility's floor, but the tacit knowledge of a specific network can only be captured and embedded by the utility that runs it.