Manufacturing (Non-Automotive)AI-Driven Disruptions & Innovations
AI factory continuous learning: how non-automotive manufacturers keep models accurate in production
AI factory continuous learning is the discipline of keeping deployed factory models accurate as the plant changes — monitoring decay, retraining on governed production data, validating every candidate against golden sets, and releasing it under change control. In non-automotive manufacturing the constraint is rarely the algorithm; it is whether the plant can prove each model change safe.

Key takeaways
- A continuously learning factory is not a factory that changes its own mind — it is a factory whose models are retrained on governed production data, validated against versioned golden sets, and released under change control fast enough to track the plant. The loop is an operations artefact, not a research one.
- Factory models decay on plant events, not on the calendar: SKU changeovers, raw-material lot changes, tool wear, maintenance resets and sensor recalibration. A retraining cadence keyed to those events beats any fixed schedule, which over-retrains stable models and under-serves volatile ones.
- Label capture is the loop's fuel line. If operator dispositions, re-inspection verdicts and lab results do not flow back into training data automatically, every retrain starts with a relabelling project — and the loop's cycle time is set by labelling, not by compute.
- The gap between today's supervised loops and tomorrow's autonomous ones is change control, not algorithms. Even the flagship WEF lighthouse factories publicly document human promotion gates; a 'learning licence' — a versioned policy stating what a model may change without sign-off — is how autonomy arrives without abandoning the QMS.
- Future-readiness is present-readiness: the plant that can safely retrain, validate and release a model this month is the plant positioned for bounded autonomy later. Nothing about waiting makes the licence easier to earn.
Abbreviations used on this page
- MES
- Manufacturing execution system
- SCADA
- Supervisory control and data acquisition
- PLC
- Programmable logic controller
- DCS
- Distributed control system
- APC
- Advanced process control
- CMMS
- Computerised maintenance management system
- OEE
- Overall equipment effectiveness
- SPC
- Statistical process control
- QMS
- Quality management system
- MOC
- Management of change (the quality-system change-control process)
- AOI
- Automated optical inspection
- OPC UA
- OPC Unified Architecture (the industrial data-exchange standard)
Free · 8 questions · ~3 minutes
Score your plant's learning loop
Eight questions, one at a time, about three minutes. Answer them and we build your personalised learning-loop report — your stage on the cadence ladder, your score on each of the four dimensions, and the specific blocker between you and the next stage — and send it to your inbox. Your result doubles as the baseline for your first instrumented model.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised loop report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, how your loop compares with plants of similar product churn, and the 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Frozen
Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
Your next movePick the one model whose errors cost most and instrument its decay — model verdicts against operator dispositions, weekly, per SKU. Measurement first, retraining second.
Stage 2 · Manually refreshed
Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
Your next moveVersion the machinery from the last refresh — the export queries, the labelling rules, the test set — so the next one is a run, not a project. The golden set comes first.
Stage 3 · Scheduled
A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
Your next moveKey retraining to plant events: recipe changes in the MES, lot changes in the ERP, overhaul closures in the CMMS, SPC rule breaches. The calendar becomes the fallback, not the driver.
Stage 4 · Event-driven
Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
Your next moveTurn the approval history into policy: mine a year of promotion decisions for the bounds within which quality has always said yes, and draft the learning licence from them.
Stage 5 · Licensed
Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.
Your next moveReview the licence on the plant's change calendar, not the audit calendar: every new product introduction, substrate change or camera refresh is a licence question before it is a retraining trigger.
0 / 24
Decay visibility
— / 6
Retraining pipeline
— / 6
Validation & release
— / 6
Learning governance
— / 6
Your score maps to a stage on the cadence ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your loop, and it is where the next quarter's investment belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the cadence ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your loop, and it is where the next quarter's investment belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the loop reviewed against the plant, not the answers?
Self-assessment runs a stage optimistic, because the best-tended model is easier to recall than the frozen ones. We review the registry, the golden sets and the MOC records with your engineering and quality leads, and leave you with a costed 90-day plan for the weakest dimension. You keep the plan either way.
How the score maps to a stage
- 0–4 — Stage 1, Frozen. Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
- 5–10 — Stage 2, Manually refreshed. Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
- 11–16 — Stage 3, Scheduled. A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
- 17–21 — Stage 4, Event-driven. Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
- 22–24 — Stage 5, Licensed. Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.
What AI factory continuous learning is — and what it is not
A definition, the loop that implements it, and the boundary between bounded learning (real) and self-evolution (speculative).
AI factory continuous learning is the discipline of keeping deployed models accurate as the plant changes: measuring post-deployment performance against live labels, retraining on governed production data when the plant's own events demand it, validating every candidate against a versioned golden set and a shadow run, and releasing the update under the same change control that governs any other process change. It is a closed loop between the line and the model — and every element of it is ordinary engineering, deployed today.
What it is not is a factory that rewrites its own rules. The phrase 'continuously learning factory' invites an image of self-evolving production systems, and it is worth being blunt: no published, named production deployment exists of a plant whose models change their own objectives or expand their own authority without human-set bounds. What the frontier actually looks like is documented by the World Economic Forum's Global Lighthouse Network (opens in a new tab) — factories with instrumented learning loops, shadow validation and human promotion gates — and by research such as acatech's Industrie 4.0 Maturity Index (opens in a new tab), whose top stages, 'predictability' and 'adaptability', describe exactly this bounded, evidence-gated form of self-optimisation. The distance between an ordinary plant and that frontier is not an algorithm; it is the loop.
Value released against learning-loop maturity
The curve is not linear. A frozen model releases its value once and then leaks it as the plant drifts away from the training data; value inflects when retraining becomes routine (stage 3) and compounds when learning transfers across lines (stage 4). This is why plants that measure progress in models deployed rather than models maintained report activity without results.
Model value retained and compounded by stage
- Stage 1 · Frozen — 34% of operators. Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
- Stage 2 · Manually refreshed — 31% of operators. Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
- Stage 3 · Scheduled — 22% of operators. A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
- Stage 4 · Event-driven — 10% of operators. Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
- Stage 5 · Licensed — 3% of operators. Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with WEF Global Lighthouse Network findings.
Three ways a factory model lives after deployment
The learning loop, lane by lane. The frozen lane is the industry default: the model decays silently until a crisis funds a rescue. The supervised loop is today's best practice — labels flow back, drift is watched, candidates are validated in shadow and promoted through a human gate. The licensed loop closes automatically, but only inside a versioned policy; it is bounded autonomy, not self-evolution.
- Data & feeds
- AI / model
- Where value leaks
- Human in the loop
- System-of-record action
The process, in words
- In the frozen lane, a model is trained once on the commissioning dataset and deployed. The plant then moves — new SKUs, raw-material lot changes, tool wear — while the model still answers the plant it was trained on. Decay is silent because nothing measures post-deployment performance, and the lane ends in a crisis-funded rescue project. This is where most factory AI lives.
- In the supervised loop, operator dispositions and lab results flow back as labels, a decay monitor watches performance against them — keyed to changeovers and lot changes, not the calendar — and a versioned pipeline retrains on trigger. Candidates must pass golden-set replay and a shadow run against live production before a named person promotes them, with an MOC record and a one-click rollback.
- In the licensed loop, a versioned learning licence agreed with quality states exactly what may change without sign-off and what evidence each release must write. Updates inside the licence promote automatically; anything outside escalates to a person, and every decision on the line logs the model version that made it.
Step-by-step insights
- Why frozen is the default, not the exception
- Most factory models arrive inside purchased equipment, trained by the vendor or integrator against the product mix present at commissioning and accepted at the FAT like any other machine capability. The commercial incentives all point to freezing: the vendor warranties the shipped behaviour, the contract rarely includes retraining or raw-data export, and the plant has no line item for model maintenance because the model was bought as a feature, not as a living system. Nothing in that arrangement fails loudly — the model just answers last year's plant with steadily less relevance, and the workarounds accumulate quietly on the line.
- Label capture is the loop's fuel line
- Every retrain is only as good as its labels, and in a plant the labels already exist — they are just thrown away. The operator at the verify station who overturns a false reject has labelled an image; the QC lab result that grades a batch has labelled a historian window; the rework ticket that comes back 'no fault found' has labelled an escape. Designing the loop means routing those dispositions back against their source records automatically, with latency measured in hours. A plant that captures labels as a by-product of normal work retrains from a running start; one that doesn't begins every refresh with a relabelling project, and its loop cycle time is set by labelling, not compute.
- The decay monitor listens to the plant, not the calendar
- Factory models decay on events: a recipe change in the MES shifts the imaging distribution, a new raw-material lot in the ERP shifts the process distribution, a closed overhaul work order in the CMMS resets a vibration signature, an SPC rule breach flags that the process itself has moved. A decay monitor keyed to those events catches drift within a shift of its cause, while a purely statistical monitor waits for enough degraded output to cross a threshold — and a calendar waits for the quarter to end. The event stream the plant already produces is the correct clock for the loop.
- The validation harness: golden sets and shadow runs
- The golden set is the plant's institutional memory — every defect class that ever mattered, every SKU in proportion, versioned like code and refreshed at every product introduction. Replay against it catches the classic retraining injury: a candidate that improves on average while regressing on a rare, expensive defect class. The shadow run answers the question the golden set cannot: how does the candidate behave on today's live production, including the new lot or recipe that triggered it? A challenger that scores live traffic without acting on it converts promotion from a judgement call into an evidence comparison.
- The promotion gate is data collection for the licence
- Every promotion decision a quality engineer makes — approved at these golden-set scores, queried at that shadow delta, refused for that regression — is a labelled example of the plant's real risk appetite. A year of such decisions, recorded against written criteria, is the empirical basis for a learning licence: the bounds within which approval was never withheld are the bounds a policy can safely automate. Plants that keep informal gates accumulate nothing and face the autonomy question cold; plants that record their gates are writing the licence without noticing.
- The licence is a quality artefact, not an engineering one
- The learning licence belongs to the QMS, next to the control plan: a versioned document stating, for one named model, what may change without sign-off (weights, not architecture; thresholds within a band), what evidence each automatic release must write, what triggers auto-rollback, and when the licence itself is reviewed. Regulators in adjacent domains have formalised the same idea as predetermined change-control plans for machine-learning systems, which is a useful template precisely because it was built to satisfy auditors. The licence's job is to make machine-speed learning and change control the same motion rather than opposing ones.
The five stages of the cadence ladder in detail
For each stage: what it looks like on the plant floor, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving costs.
Each stage below is written for a practitioner rather than a buyer. The ladder climbs by cadence — how fast, and on whose clock, a deployed model can safely learn: never (frozen), after a crisis (manually refreshed), on the calendar (scheduled), on the plant's events (event-driven), and automatically within policy (licensed). The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own registry and QMS this week, and the anti-pattern is the specific mistake most often made trying to leave.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Frozen
34% of operators sit here
Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
Stage 1 is the default state of factory AI, and it is worth being precise about why. Most plant models arrive inside equipment: the vision system commissioned with the filler, the anomaly detector bundled with the compressor, the soft sensor the integrator tuned at start-up. They are accepted at the factory acceptance test against the product mix of that month, and from that day the model is treated like any other machine setting — fixed, warrantied, and nobody's job to revisit.
The tell is not poor performance; it is unmeasured performance. A frozen model's accuracy at deployment is usually genuine. What is missing is any record of its accuracy since — no live comparison against operator verdicts, no false-reject trend, no escape tracking tied to model version. Ask what the inspection system's false-reject rate was last month and this month, and a stage-1 plant cannot answer from data. Meanwhile the plant has moved: new SKUs, a resin supplier change, a rebuilt gearbox, recalibrated cameras. The model still answers the plant it was trained on.
The cost surfaces as workarounds rather than incidents. Operators learn the system over-rejects the new label film, so sensitivity gets turned down until it passes everything — at which point the plant is paying for inspection it no longer receives. Frozen models rarely fail loudly; they fade into expensive decoration, and the budget conversation three years later is about replacing the system rather than about the missing loop that would have kept it alive.
In practice
The inspection system commissioned at FAT
A beverage packer's label-inspection cameras were trained on the twelve SKUs running at commissioning. Three years later the line runs thirty-eight SKUs, two label-film suppliers have changed, and the false-reject storm after every new launch is handled the same way each time: the vendor engineer visits, sensitivity comes down a notch, and the QMS records a 'settings adjustment'. Nobody can say what the system still catches, because nothing compares its verdicts with the verify station's.
What it looks like
- Model files carry the commissioning date; nothing has changed since
- No post-deployment performance record exists anywhere
- Operators have built workarounds — sensitivity turned down, stations bypassed
- The vendor contract has no retraining clause and no data-export path
Diagnostic signals you can check this week
- Check the model file or vendor firmware date against the commissioning date — if they match, the model has never learned
- Ask for last month's false-reject rate versus the rate at acceptance; a stage-1 plant answers from memory, not from data
- Walk the line and count workarounds: sensitivity reductions, bypassed stations, 'ignore that alarm' notes
- Read the vendor contract for a retraining clause and a raw-data export path — absence of both is the commercial signature of frozen AI
Anti-pattern · Replacing the model instead of building the loop
The instinctive response to a decayed system is procurement: rip out the three-year-old vision system and buy this year's, which demos brilliantly against the current product mix — exactly as the old one did against the old mix. Without a loop, the new system starts decaying on day one, and the plant is on a three-year replacement treadmill that costs more than the retraining capability would. Before replacing anything, instrument the decay: compare model verdicts with operator dispositions for one month. That number funds the loop, and sometimes it saves the replacement.
What holds you here
Nothing measures post-deployment performance, so decay is invisible and no case for retraining can be made from data.
Highest-leverage next move
Pick the one model whose errors cost most and instrument its decay — model verdicts against operator dispositions, weekly, per SKU. Measurement first, retraining second.
Cost of leaving
- Effort
- 2–4 months
- Team
- One process or quality engineer plus one data engineer, part-time
- Risk
- Low — measurement is additive; nothing in production changes yet
- To next stage
- 2–4 months
If this is you, the next step is
A short engagement: pick the model, wire verdicts against dispositions, measure the decay curve.
Stage 2
Manually refreshed
31% of operators sit here
Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
Stage 2 plants have proven the most important thing: retraining works. A refresh rescued the vision system after the packaging change, or rebuilt the soft sensor after the raw-material supplier switch, and performance recovered. What they have not built is the machinery that makes the second refresh cheaper than the first — so every refresh is a bespoke project with a bespoke export, a bespoke relabelling effort and a bespoke definition of 'good enough to ship'.
The economics are quietly punishing. A refresh that takes eight weeks means the plant runs a degraded model for two months per incident — and because the trigger is a crisis, the degradation ran unmeasured for months before that. Worse, the refresh is validated against whatever test data the engineer assembled that week. Nothing guarantees the new model still catches the rare defect class the old one was originally bought for; regressions on low-frequency defects are the classic manual-refresh injury, and they surface as field escapes with a customer's name attached.
The stage is also fragile to people. The engineer who ran the last refresh encoded a dozen decisions — which weeks to export, which images to exclude, what counts as a scratch versus a mark — in their head and their laptop. When a different engineer runs the next one, those decisions get remade differently, and the two models bracket a definition drift nobody chose. A stage-2 plant does not have a retraining process; it has retraining incidents with local heroes.
In practice
The relabelling fortnight
A speciality-chemicals plant runs a soft sensor predicting batch viscosity from in-process measurements. When a solvent supplier changed, predictions drifted and off-spec batches reached the QC lab before anyone connected the two. The fix took six weeks: export eighteen months of historian data, relabel against lab results, rebuild, argue about acceptance. Two years later a second supplier change repeated the whole exercise — run by a different engineer, from scratch, with a different test set, because nothing from the first refresh had been kept as reusable machinery.
What it looks like
- Every retrain is triggered by a complaint, an escape or a false-reject storm
- Training data is re-exported and relabelled by hand for each refresh
- Each refresh invents its own test set, so no two are comparable
- Lead time from 'the model is wrong' to 'the model is fixed' is 6–12 weeks
Diagnostic signals you can check this week
- Ask when the last retrain happened and what triggered it — 'a customer complaint' is the stage-2 answer
- Ask to see the current validation set; if the answer is 'which one?', each refresh invented its own
- Time the loop: date of first degradation evidence to date the fixed model shipped. Weeks means stage 2
- Check whether the last refresh's data-selection choices are written anywhere a second engineer could follow
Anti-pattern · The heroic refresh
After a successful rescue, the plant declares the problem solved and disbands. The refresh treated a symptom; the disease is that decay is only visible through crises. The anti-pattern is spending the post-rescue quarter on nothing, when the same quarter could wire the monitoring that makes the next decay visible in week one instead of month four — and turn the rescue's one-off scripts into a pipeline. A plant that only retrains when customers complain has outsourced its drift monitoring to its customers.
What holds you here
Every refresh is rebuilt from scratch, so retraining is too slow and too expensive to run before a crisis forces it.
Highest-leverage next move
Version the machinery from the last refresh — the export queries, the labelling rules, the test set — so the next one is a run, not a project. The golden set comes first.
Cost of leaving
- Effort
- 4–9 months
- Team
- One ML engineer, one data engineer, a named quality-side owner
- Risk
- Medium — the first standing golden set forces the plant to agree what 'good' means, which is politics as much as engineering
- To next stage
- 4–9 months
If this is you, the next step is
We take your last manual refresh and rebuild it as versioned, repeatable machinery.
Stage 3
Scheduled
22% of operators sit here
A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
Stage 3 is where continuous learning stops being a slogan and becomes a standing capability. The pipeline exists: training data flows from the historian and MES through governed queries, the candidate model is evaluated against a golden set that is itself versioned and stratified, and promotion is a decision a named person takes against written criteria, recorded in the QMS. The refresh that took eight weeks at stage 2 takes days, and — more importantly — it happens before the crisis rather than after it.
The discipline that defines this stage is the golden set. It is the plant's institutional memory of what the model must never forget: every defect class that ever mattered, every SKU in proportion, the near-misses and the edge cases, each with an agreed label. Candidates that improve on average but regress on a rare, expensive defect class get caught here — the exact failure the manual refresh shipped. Golden sets rot, though, and a stage-3 plant treats additions to the set as seriously as changes to the model: every new product introduction adds its rows before launch, not after the first escape.
The constraint that emerges is cadence mismatch. The calendar says quarterly; the plant does not decay quarterly. A stable tablet-compression model gets retrained four times a year for nothing, while the vision model on the fast-churning packaging line is six weeks stale the day it ships. Fixed schedules are the training wheels of continuous learning — the right cadence is the plant's own event stream, and noticing that is what pushes an operator toward stage 4.
In practice
The quarterly retrain that missed the changeover
A pharma packaging operation retrains its blister-inspection model quarterly, with golden-set replay and QA sign-off — genuine stage 3. Mid-quarter, a new foil supplier came in. The foil's reflectivity shifted the imaging distribution, false rejects tripled, and for five weeks operators re-inspected an extra bin per shift while the calendar counted down. The scheduled retrain fixed it on schedule — five weeks after the plant needed it. The post-mortem question was the right one: why does the model learn on the calendar's clock instead of the plant's?
What it looks like
- Retraining runs on a schedule from a versioned pipeline, historian to registry
- A versioned golden set gates every promotion, stratified by SKU and defect class
- Decay is tracked on a dashboard reviewed alongside SPC and quality metrics
- Every promotion leaves an MOC record: what changed, evidence, approver
Diagnostic signals you can check this week
- Ask to see the golden set's version history — a stage-3 plant shows commits; a stage-2 plant shows a folder
- Compare the retrain calendar with the changeover calendar; if they never reference each other, cadence is arbitrary
- Check whether any promotion was ever refused on golden-set evidence — a gate that has never said no is not a gate
- Look for one model retrained often and another rarely, with a written reason; uniform cadence for every model is the stage-3 tell
Anti-pattern · One cadence for every model
The pipeline makes retraining cheap, so the plant standardises: everything quarterly. But cadence is a property of the decay driver, not of the pipeline. Vision models on high-churn packaging lines decay per changeover; process models decay per raw-material lot; maintenance models decay when the CMMS closes an overhaul work order and resets a vibration signature. Setting one calendar for all of them over-trains the stable and abandons the volatile. Set cadence per model from its measured decay curve — the curve stage 1 taught you to draw.
What holds you here
Retraining runs on the calendar while the plant decays on events — changeovers, lots, overhauls — so the cadence is always wrong for someone.
Highest-leverage next move
Key retraining to plant events: recipe changes in the MES, lot changes in the ERP, overhaul closures in the CMMS, SPC rule breaches. The calendar becomes the fallback, not the driver.
Cost of leaving
- Effort
- 9–18 months
- Team
- ML engineer, data engineer, quality engineer as standing approver
- Risk
- Medium — wiring triggers to MES and ERP events crosses team boundaries, and label capture at the verify station changes operator workflow
- To next stage
- 9–18 months
If this is you, the next step is
We map your decay drivers to MES, ERP and CMMS events and wire the first trigger.
Stage 4
Event-driven
10% of operators sit here
Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
Stage 4 inverts the direction of attention. At stage 3, people decide when the model should learn; at stage 4, the plant tells them. A recipe change in the MES, a new resin lot booked in the ERP, a gearbox overhaul closed in the CMMS, an SPC rule breach on a critical characteristic — each fires an evaluation, and where the champion degrades, a challenger trains on the freshest labelled window and enters shadow. The loop's raw material is the label stream: verify-station dispositions, QC lab results and rework outcomes land against their source records within hours, because label capture was designed into the operator workflow rather than bolted on.
Shadow deployment is the stage's signature discipline. The challenger scores every live part or batch but acts on none; its verdicts are logged beside the champion's and compared over a defined window that must include the conditions that triggered it — the new lot, the changed recipe. Promotion is then an evidence-backed decision: here is the golden-set replay, here is two weeks of shadow against live production, here is the delta by defect class. The human gate remains, and it is worth being honest that it now sets the loop's cycle time: a challenger validated in three days can wait two weeks for the MOC record and the quality sign-off.
This is also the stage where learning compounds across the estate. The defect-detection model proven on line 3 transfers to line 5 with a fine-tune on local imagery rather than a from-scratch build; the sister plant inherits the pipeline, the golden-set structure and the trigger map. Fleet learning is why stage-4 operators pull away from stage-3 ones on cost: the second line's loop costs a fraction of the first's, and the tenth is configuration. The frontier operators documented by the World Economic Forum's Global Lighthouse Network are recognisable stage-4 shops — instrumented loops, shadow validation, human gates — which is itself the strongest public evidence of where the real frontier currently sits.
In practice
The changeover that retrained the model
An electronics manufacturer's SMT line runs AOI models per assembly family. A new assembly's recipe activation in the MES fires an evaluation: the champion's predicted false-call rate on the new family exceeds its band, so a challenger fine-tunes overnight on the first shift's operator-verified labels. It shadows the champion for three days across two thousand boards, the comparison lands in the quality engineer's queue with golden-set replay attached, and promotion happens inside the week — with the previous version one click away. Total operator-visible disruption: none.
What it looks like
- Triggers fire from plant events: MES recipe changes, ERP lot changes, CMMS overhauls, SPC breaches
- A champion–challenger pair runs continuously; the challenger scores live traffic in shadow
- Operator dispositions and lab results flow back as labels with measured latency
- Models transfer between lines and sites with local fine-tuning — fleet learning
Diagnostic signals you can check this week
- Ask what fired the last retrain; the stage-4 answer names a plant event, not a date or a crisis
- Check for a challenger scoring live traffic right now, and where its shadow log lives
- Measure label latency: production event to usable labelled record. Hours is stage 4; weeks is not
- Ask how the second line got its model — 'transferred and fine-tuned from line 3' is the fleet-learning tell
Anti-pattern · Shipping around the paperwork
The loop now outruns the QMS: a challenger is validated in days while the change record takes a fortnight, and the engineering team starts quietly promoting 'minor' model updates outside MOC to keep pace. It works until the day it defines the plant's audit finding — a nonconformance investigation asks which model version passed the affected lot, and the honest answer is that the registry and the QMS disagree. The fix is not more discipline but less friction: make the pipeline generate the change-control evidence automatically, so the compliant path is also the fast path. That automation is precisely the groundwork for stage 5.
What holds you here
Every promotion still waits for a human gate, so the loop's cycle time is set by sign-off and MOC paperwork rather than by training and validation.
Highest-leverage next move
Turn the approval history into policy: mine a year of promotion decisions for the bounds within which quality has always said yes, and draft the learning licence from them.
Cost of leaving
- Effort
- 18+ months
- Team
- Platform engineer, ML engineer, quality partner with delegated approval authority
- Risk
- Higher — trigger wiring spans MES, ERP and CMMS ownership boundaries, and approval-lead-time politics surface between engineering and quality
- To next stage
- 18+ months
If this is you, the next step is
We automate the evidence a promotion needs so sign-off takes a day, not a fortnight.
Stage 5
Licensed
3% of operators sit here
Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.
Stage 5 is narrower than the phrase 'autonomous learning' suggests, and the narrowness is the point. A learning licence is a quality-system artefact: a versioned document, owned jointly by engineering and quality, that enumerates for one named model what may change without human sign-off — retraining on new data with a fixed architecture, decision thresholds within a stated band — and what evidence each automatic promotion must generate: golden-set score floors, shadow-window minimums, per-defect-class regression limits, an auto-rollback trigger. Inside the licence, the loop closes at machine speed. Outside it, nothing moves without a person, exactly as at stage 4.
The honest context is that almost no manufacturer operates here today, and the page's own distribution reflects that. The intellectual scaffolding exists — regulators in adjacent domains have formalised predetermined-change frameworks for machine-learning systems, and the same logic maps cleanly onto a manufacturing MOC — but public, named examples of auto-promoted model updates in production plants are scarce. That is not a reason to dismiss the stage; it is the reason to treat everything below it as the qualification. A licence is only as credible as the promotion history behind it, and the only way to accumulate that history is to run a supervised loop well, for a long time.
Sustaining stage 5 is renewal discipline. The licence encodes assumptions — this product mix, this camera fleet, this defect taxonomy — and the plant changes underneath all of them. The escalation rate is the canary: when the share of candidates falling outside the licence rises, the world has left the licence's validity envelope and the review should happen before an incident forces it. The failure mode is never the dramatic rogue model; it is licence creep — bounds quietly widened after each escalation until the licence describes what the system does rather than what quality agreed. Version the licence, review it like code, and let it shrink as willingly as it grows.
In practice
The bounded licence on the packaging line
A packaging operation runs its carton-print inspection model under a licence: retrains may promote automatically if golden-set recall on every defect class stays above its floor, false-reject rate on a three-day shadow window stays inside its ceiling, and the decision threshold moves less than a stated delta. Each auto-promotion writes an evidence pack into the QMS; roughly one candidate in eight falls outside and lands in the quality engineer's queue. When a new carton substrate pushed escalations from one-in-eight towards one-in-three, the licence review happened that month — the point of the metric is that the review pre-empted the incident.
What it looks like
- A versioned learning licence states what may change without sign-off, by how much, evidenced how
- Updates inside the licence promote automatically; every release writes its evidence pack
- Out-of-bounds candidates escalate to a person — and the escalation rate is itself monitored
- Auto-rollback triggers and the kill switch are drilled, with the licence reviewed on a fixed cycle
Diagnostic signals you can check this week
- Ask to read the licence — a stage-5 plant hands you a versioned document with named owners, not a description of engineering culture
- Check that every automatic promotion wrote an evidence pack the QMS can produce on demand
- Ask when the kill switch and auto-rollback were last exercised deliberately; 'never' means untested
- Look at the escalation-rate trend and who reviews it — an unmonitored escalation rate means the licence is drifting unobserved
Anti-pattern · Licence creep
Each escalation is resolved by widening the bound that caught it — the golden-set floor nudged down, the threshold band nudged out — with no review of whether the world changed or the licence was simply inconvenient. Eighteen months later the licence permits nearly everything and certifies nothing, and the first serious audit finds a policy that was rewritten by its own exceptions. Treat every proposed widening as a licence change with quality sign-off, and record the narrowing decisions too: a licence that only ever grows is not being governed.
What holds you here
Keeping the licence valid as the plant changes — renewal, escalation review and the discipline to narrow bounds — is a permanent governance cost, not a project.
Highest-leverage next move
Review the licence on the plant's change calendar, not the audit calendar: every new product introduction, substrate change or camera refresh is a licence question before it is a retraining trigger.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team plus a standing engineering–quality review forum
- Risk
- Concentrated — low-frequency, high-consequence, and regulatory in character; the licence itself is what an auditor will examine
If this is you, the next step is
We red-team the bounds, the evidence packs and the rollback against a real excursion scenario.
Where manufacturing plants actually sit on the ladder
The distribution across the five stages, and why the frozen–refreshed plateau holds two thirds of the industry.
Most plants sit in the first two stages: their models are frozen, or retrained only when a crisis forces it. That is the honest reading of the adoption research — AI use is now near-universal at the organisational level, while the operational machinery that keeps models accurate remains rare. The distribution below shows the shape: a large frozen base, a plateau of crisis-driven refreshers, and a thin frontier of event-driven and licensed loops.
Distribution of manufacturers across the cadence ladder
Illustrative distribution — charted for orientation, not measurement. Synthesised from McKinsey's State of AI adoption research and the operating patterns documented across WEF Global Lighthouse Network sites; the frozen–refreshed plateau is where roughly two thirds of plants sit.
Share of plants (illustrative)
- 34% — 1 · Frozen (the silent majority)
- 31% — 2 · Manually refreshed
- 22% — 3 · Scheduled
- 10% — 4 · Event-driven
- 3% — 5 · Licensed
Source: Illustrative, synthesised from McKinsey State of AI and WEF Global Lighthouse Network research
The frontier is public. The Global Lighthouse Network (opens in a new tab) — the World Economic Forum's programme identifying factories deploying fourth-industrial-revolution technology at production scale — documents, site by site, what the top of the ladder looks like in practice: closed-loop quality systems, models maintained across whole fleets of lines, and measured impact on OEE, yield and energy. Two details in those write-ups matter for this page. First, the lighthouse use cases are overwhelmingly supervised loops — the human promotion gate appears everywhere. Second, the lighthouses report their capabilities in loop terms (how fast a model tracks the plant) rather than deployment terms (how many models exist), which is exactly the shift the cadence ladder measures. MHI's annual industry survey (opens in a new tab) tracks the same adoption-versus-impact gap from the supply-chain side of manufacturing.
What decays, where: the model families of a manufacturing plant
Six model families, the system each lives in, the plant event that makes each one decay — and the cadence each actually needs.
Every model family in a plant decays for a different reason, and the decay driver — not the model type — is what sets the right retraining cadence. A vision model on a high-churn packaging line decays at every changeover; a predictive-maintenance model decays when the overhaul it predicts actually happens and resets the signature it learned; a demand model decays with the season. The map below is how we scope loops with operators: read your most expensive model's row, and its cadence column is the stage the ladder says that model needs — whatever stage the rest of the estate is at.
| Model family | What it decides | System of record | What makes it decay | Cadence sweet spot |
|---|---|---|---|---|
| Vision inspection / AOI | Accept, reject, rework routing | MES + inspection station | New SKUs, label and substrate changes, camera and lighting drift | Stage 4 |
| Process soft sensors & APC setpoints | Setpoints and quality predictions within the control window | DCS / APC | Raw-material lot variation, catalyst ageing, fouling, sensor recalibration | Stage 3–4 |
| Predictive maintenance | Intervention timing, spares staging | CMMS + historian | Overhauls resetting signatures, operating-regime change, replaced components | Stage 3 |
| Demand & production scheduling | Sequence, batch sizing, campaign length | ERP / APS + MES | Seasonality, portfolio churn, promotions, customer mix | Stage 3 |
| Energy optimisation | Load shifting, compressor and utilities scheduling | SCADA / BMS | Tariff changes, weather seasonality, line reconfiguration | Stage 3 |
| Batch golden-profile monitoring | Deviation detection against the reference trajectory | Historian + QMS | Recipe changes, vessel maintenance, instrument recalibration | Stage 4 |
Two readings of the map are worth making explicit. First, the families that decay on discrete, loggable events — vision models at changeovers, batch profiles at recipe changes — are the natural first candidates for event-driven learning, because the trigger already exists as an MES or ERP record and merely needs wiring. Families that decay on slow, continuous drivers — fouling, seasonality — are well served by scheduled cadence and gain less from triggers. Second, the map is why 'what stage is your plant?' is really a portfolio question: a sensible estate runs its packaging-line AOI at stage 4 and its energy model at stage 3, and the assessment on this page scores the shared machinery — pipeline, golden sets, governance — that all of them draw on.
One family deserves a boundary note: models inside safety-instrumented functions do not belong on this ladder at all. A trip function or interlock is certified against frozen, verified behaviour under functional-safety standards, and 'continuously learning' is precisely the property certification excludes. The correct pattern — visible in every credible deployment — is that learning systems advise and optimise up to the boundary of the safety envelope, and the certified layer beneath them stays frozen. A vendor pitching a learning model inside the safety loop is pitching a recertification programme, and probably does not know it.
The change-control wall: what separates real from speculative
Every claim about the self-optimising factory sorts into deployed, demonstrated, or speculative — and the sorting variable is change control, not algorithms.
The wall between today's supervised loops and the imagined self-evolving factory is change control, and it is a load-bearing wall rather than an obstacle. A manufacturing plant runs on the discipline that process changes are proposed, risk-assessed, evidenced and approved before they touch product — the MOC process of an ISO 9001-style QMS, hardened further in regulated sectors where process validation demands that any change to a validated state be requalified. A model update is a process change. That single sentence explains most of the distance between AI marketing and plant reality: the constraint on learning speed was never training compute; it is the rate at which change evidence can be honestly generated and reviewed.
| Capability | Status | What bounds it |
|---|---|---|
| Drift monitoring and scheduled retraining with a human promotion gate | Deployed today, widely — ordinary engineering | Cost and discipline only |
| Event-triggered retraining with shadow validation and fleet transfer | Deployed at leading sites; documented across WEF lighthouse factories | Label capture and MOC throughput |
| Closed-loop setpoint adjustment inside a fixed, engineered envelope | Deployed for decades as APC; AI now proposes better envelopes | The envelope is engineered and certified by humans |
| Automatic promotion of retrained models under a versioned learning licence | Nascent — the policy template exists; few public production examples | Quality-system change control and accumulated approval evidence |
| Continuous learning inside safety-instrumented functions | Blocked — certification requires frozen, verified behaviour | Functional-safety standards (IEC 61511-class certification) |
| A factory whose models rewrite their own objectives and authority | Speculative — no published, named production deployment | Evidence cannot be generated faster than product can be measured |
The last row's bound deserves one more sentence, because it is physics rather than caution: a model update is validated by comparing predicted quality against measured quality, and measurement takes as long as the process takes. A batch that cures for eight hours yields one label every eight hours, however fast the GPU. Learning speed in a factory is throughput-limited by the plant's own measurement cycle — which is why the credible frontier is a loop that wastes none of that scarce evidence, not a system that somehow transcends it.
Read the ledger bottom-up and the strategic conclusion falls out: everything above the wall is earned by running everything below it. The licence bounds of row four are mined from the approval history of row two; the approval history exists only if the loop exists; the loop exists only if decay is measured. This is the sense in which future-readiness is present-readiness — a plant cannot buy its way to licensed autonomy, because the purchase price is denominated in its own accumulated promotion evidence. acatech's maturity research (opens in a new tab) reaches the same conclusion from the organisational side: its 'adaptability' stage is defined as the outcome of the preceding stages' data and process discipline, not as a technology acquisition.


