Redefining Technology

Manufacturing (Non-Automotive)AI-Driven Disruptions & Innovations

AI factory continuous learning: how non-automotive manufacturers keep models accurate in production

AI factory continuous learning is the discipline of keeping deployed factory models accurate as the plant changes — monitoring decay, retraining on governed production data, validating every candidate against golden sets, and releasing it under change control. In non-automotive manufacturing the constraint is rarely the algorithm; it is whether the plant can prove each model change safe.

Factory production line with model-performance and retraining overlays across inspection and process stations — illustrative scene
Manufacturing (Non-Automotive) · AI-Driven Disruptions & Innovations

Key takeaways

  1. A continuously learning factory is not a factory that changes its own mind — it is a factory whose models are retrained on governed production data, validated against versioned golden sets, and released under change control fast enough to track the plant. The loop is an operations artefact, not a research one.
  2. Factory models decay on plant events, not on the calendar: SKU changeovers, raw-material lot changes, tool wear, maintenance resets and sensor recalibration. A retraining cadence keyed to those events beats any fixed schedule, which over-retrains stable models and under-serves volatile ones.
  3. Label capture is the loop's fuel line. If operator dispositions, re-inspection verdicts and lab results do not flow back into training data automatically, every retrain starts with a relabelling project — and the loop's cycle time is set by labelling, not by compute.
  4. The gap between today's supervised loops and tomorrow's autonomous ones is change control, not algorithms. Even the flagship WEF lighthouse factories publicly document human promotion gates; a 'learning licence' — a versioned policy stating what a model may change without sign-off — is how autonomy arrives without abandoning the QMS.
  5. Future-readiness is present-readiness: the plant that can safely retrain, validate and release a model this month is the plant positioned for bounded autonomy later. Nothing about waiting makes the licence easier to earn.

Abbreviations used on this page

MES
Manufacturing execution system
SCADA
Supervisory control and data acquisition
PLC
Programmable logic controller
DCS
Distributed control system
APC
Advanced process control
CMMS
Computerised maintenance management system
OEE
Overall equipment effectiveness
SPC
Statistical process control
QMS
Quality management system
MOC
Management of change (the quality-system change-control process)
AOI
Automated optical inspection
OPC UA
OPC Unified Architecture (the industrial data-exchange standard)

Free · 8 questions · ~3 minutes

Score your plant's learning loop

Eight questions, one at a time, about three minutes. Answer them and we build your personalised learning-loop report — your stage on the cadence ladder, your score on each of the four dimensions, and the specific blocker between you and the next stage — and send it to your inbox. Your result doubles as the baseline for your first instrumented model.

0 of 8 answered

Question 1 of 8Decay visibility

How do you find out that a deployed model's performance has degraded?

Decay that is only visible through complaints has already cost you — the loop cannot be faster than its detection.

How the score maps to a stage
  • 04 — Stage 1, Frozen. Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
  • 510 — Stage 2, Manually refreshed. Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
  • 1116 — Stage 3, Scheduled. A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
  • 1721 — Stage 4, Event-driven. Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
  • 2224 — Stage 5, Licensed. Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.

What AI factory continuous learning is — and what it is not

A definition, the loop that implements it, and the boundary between bounded learning (real) and self-evolution (speculative).

AI factory continuous learning is the discipline of keeping deployed models accurate as the plant changes: measuring post-deployment performance against live labels, retraining on governed production data when the plant's own events demand it, validating every candidate against a versioned golden set and a shadow run, and releasing the update under the same change control that governs any other process change. It is a closed loop between the line and the model — and every element of it is ordinary engineering, deployed today.

What it is not is a factory that rewrites its own rules. The phrase 'continuously learning factory' invites an image of self-evolving production systems, and it is worth being blunt: no published, named production deployment exists of a plant whose models change their own objectives or expand their own authority without human-set bounds. What the frontier actually looks like is documented by the World Economic Forum's Global Lighthouse Network (opens in a new tab) — factories with instrumented learning loops, shadow validation and human promotion gates — and by research such as acatech's Industrie 4.0 Maturity Index (opens in a new tab), whose top stages, 'predictability' and 'adaptability', describe exactly this bounded, evidence-gated form of self-optimisation. The distance between an ordinary plant and that frontier is not an algorithm; it is the loop.

Value released against learning-loop maturity

The curve is not linear. A frozen model releases its value once and then leaks it as the plant drifts away from the training data; value inflects when retraining becomes routine (stage 3) and compounds when learning transfers across lines (stage 4). This is why plants that measure progress in models deployed rather than models maintained report activity without results.

Model value retained and compounded by stage

  • Stage 1 · Frozen — 34% of operators. Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.
  • Stage 2 · Manually refreshed — 31% of operators. Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.
  • Stage 3 · Scheduled — 22% of operators. A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.
  • Stage 4 · Event-driven — 10% of operators. Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.
  • Stage 5 · Licensed — 3% of operators. Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with WEF Global Lighthouse Network findings.

Three ways a factory model lives after deployment

The learning loop, lane by lane. The frozen lane is the industry default: the model decays silently until a crisis funds a rescue. The supervised loop is today's best practice — labels flow back, drift is watched, candidates are validated in shadow and promoted through a human gate. The licensed loop closes automatically, but only inside a versioned policy; it is bounded autonomy, not self-evolution.

  • Data & feeds
  • AI / model
  • Where value leaks
  • Human in the loop
  • System-of-record action

The process, in words

  • In the frozen lane, a model is trained once on the commissioning dataset and deployed. The plant then moves — new SKUs, raw-material lot changes, tool wear — while the model still answers the plant it was trained on. Decay is silent because nothing measures post-deployment performance, and the lane ends in a crisis-funded rescue project. This is where most factory AI lives.
  • In the supervised loop, operator dispositions and lab results flow back as labels, a decay monitor watches performance against them — keyed to changeovers and lot changes, not the calendar — and a versioned pipeline retrains on trigger. Candidates must pass golden-set replay and a shadow run against live production before a named person promotes them, with an MOC record and a one-click rollback.
  • In the licensed loop, a versioned learning licence agreed with quality states exactly what may change without sign-off and what evidence each release must write. Updates inside the licence promote automatically; anything outside escalates to a person, and every decision on the line logs the model version that made it.
Step-by-step insights
Why frozen is the default, not the exception
Most factory models arrive inside purchased equipment, trained by the vendor or integrator against the product mix present at commissioning and accepted at the FAT like any other machine capability. The commercial incentives all point to freezing: the vendor warranties the shipped behaviour, the contract rarely includes retraining or raw-data export, and the plant has no line item for model maintenance because the model was bought as a feature, not as a living system. Nothing in that arrangement fails loudly — the model just answers last year's plant with steadily less relevance, and the workarounds accumulate quietly on the line.
Label capture is the loop's fuel line
Every retrain is only as good as its labels, and in a plant the labels already exist — they are just thrown away. The operator at the verify station who overturns a false reject has labelled an image; the QC lab result that grades a batch has labelled a historian window; the rework ticket that comes back 'no fault found' has labelled an escape. Designing the loop means routing those dispositions back against their source records automatically, with latency measured in hours. A plant that captures labels as a by-product of normal work retrains from a running start; one that doesn't begins every refresh with a relabelling project, and its loop cycle time is set by labelling, not compute.
The decay monitor listens to the plant, not the calendar
Factory models decay on events: a recipe change in the MES shifts the imaging distribution, a new raw-material lot in the ERP shifts the process distribution, a closed overhaul work order in the CMMS resets a vibration signature, an SPC rule breach flags that the process itself has moved. A decay monitor keyed to those events catches drift within a shift of its cause, while a purely statistical monitor waits for enough degraded output to cross a threshold — and a calendar waits for the quarter to end. The event stream the plant already produces is the correct clock for the loop.
The validation harness: golden sets and shadow runs
The golden set is the plant's institutional memory — every defect class that ever mattered, every SKU in proportion, versioned like code and refreshed at every product introduction. Replay against it catches the classic retraining injury: a candidate that improves on average while regressing on a rare, expensive defect class. The shadow run answers the question the golden set cannot: how does the candidate behave on today's live production, including the new lot or recipe that triggered it? A challenger that scores live traffic without acting on it converts promotion from a judgement call into an evidence comparison.
The promotion gate is data collection for the licence
Every promotion decision a quality engineer makes — approved at these golden-set scores, queried at that shadow delta, refused for that regression — is a labelled example of the plant's real risk appetite. A year of such decisions, recorded against written criteria, is the empirical basis for a learning licence: the bounds within which approval was never withheld are the bounds a policy can safely automate. Plants that keep informal gates accumulate nothing and face the autonomy question cold; plants that record their gates are writing the licence without noticing.
The licence is a quality artefact, not an engineering one
The learning licence belongs to the QMS, next to the control plan: a versioned document stating, for one named model, what may change without sign-off (weights, not architecture; thresholds within a band), what evidence each automatic release must write, what triggers auto-rollback, and when the licence itself is reviewed. Regulators in adjacent domains have formalised the same idea as predetermined change-control plans for machine-learning systems, which is a useful template precisely because it was built to satisfy auditors. The licence's job is to make machine-speed learning and change control the same motion rather than opposing ones.

The five stages of the cadence ladder in detail

For each stage: what it looks like on the plant floor, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving costs.

Each stage below is written for a practitioner rather than a buyer. The ladder climbs by cadence — how fast, and on whose clock, a deployed model can safely learn: never (frozen), after a crisis (manually refreshed), on the calendar (scheduled), on the plant's events (event-driven), and automatically within policy (licensed). The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own registry and QMS this week, and the anti-pattern is the specific mistake most often made trying to leave.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Frozen

34% of operators sit here

Models are trained once — usually at commissioning — deployed, and never retrained; decay is invisible because nothing measures post-deployment performance.

Stage 1 is the default state of factory AI, and it is worth being precise about why. Most plant models arrive inside equipment: the vision system commissioned with the filler, the anomaly detector bundled with the compressor, the soft sensor the integrator tuned at start-up. They are accepted at the factory acceptance test against the product mix of that month, and from that day the model is treated like any other machine setting — fixed, warrantied, and nobody's job to revisit.

The tell is not poor performance; it is unmeasured performance. A frozen model's accuracy at deployment is usually genuine. What is missing is any record of its accuracy since — no live comparison against operator verdicts, no false-reject trend, no escape tracking tied to model version. Ask what the inspection system's false-reject rate was last month and this month, and a stage-1 plant cannot answer from data. Meanwhile the plant has moved: new SKUs, a resin supplier change, a rebuilt gearbox, recalibrated cameras. The model still answers the plant it was trained on.

The cost surfaces as workarounds rather than incidents. Operators learn the system over-rejects the new label film, so sensitivity gets turned down until it passes everything — at which point the plant is paying for inspection it no longer receives. Frozen models rarely fail loudly; they fade into expensive decoration, and the budget conversation three years later is about replacing the system rather than about the missing loop that would have kept it alive.

In practice

The inspection system commissioned at FAT

A beverage packer's label-inspection cameras were trained on the twelve SKUs running at commissioning. Three years later the line runs thirty-eight SKUs, two label-film suppliers have changed, and the false-reject storm after every new launch is handled the same way each time: the vendor engineer visits, sensitivity comes down a notch, and the QMS records a 'settings adjustment'. Nobody can say what the system still catches, because nothing compares its verdicts with the verify station's.

What it looks like

  • Model files carry the commissioning date; nothing has changed since
  • No post-deployment performance record exists anywhere
  • Operators have built workarounds — sensitivity turned down, stations bypassed
  • The vendor contract has no retraining clause and no data-export path

Diagnostic signals you can check this week

  • Check the model file or vendor firmware date against the commissioning date — if they match, the model has never learned
  • Ask for last month's false-reject rate versus the rate at acceptance; a stage-1 plant answers from memory, not from data
  • Walk the line and count workarounds: sensitivity reductions, bypassed stations, 'ignore that alarm' notes
  • Read the vendor contract for a retraining clause and a raw-data export path — absence of both is the commercial signature of frozen AI

Anti-pattern · Replacing the model instead of building the loop

The instinctive response to a decayed system is procurement: rip out the three-year-old vision system and buy this year's, which demos brilliantly against the current product mix — exactly as the old one did against the old mix. Without a loop, the new system starts decaying on day one, and the plant is on a three-year replacement treadmill that costs more than the retraining capability would. Before replacing anything, instrument the decay: compare model verdicts with operator dispositions for one month. That number funds the loop, and sometimes it saves the replacement.

What holds you here

Nothing measures post-deployment performance, so decay is invisible and no case for retraining can be made from data.

Highest-leverage next move

Pick the one model whose errors cost most and instrument its decay — model verdicts against operator dispositions, weekly, per SKU. Measurement first, retraining second.

Cost of leaving

Effort
2–4 months
Team
One process or quality engineer plus one data engineer, part-time
Risk
Low — measurement is additive; nothing in production changes yet
To next stage
2–4 months

If this is you, the next step is

A short engagement: pick the model, wire verdicts against dispositions, measure the decay curve.

Instrument one model's decay

Stage 2

Manually refreshed

31% of operators sit here

Retraining happens — but as a rescue project after a crisis, rebuilt from scratch each time, with no standing validation set and a lead time measured in weeks.

Stage 2 plants have proven the most important thing: retraining works. A refresh rescued the vision system after the packaging change, or rebuilt the soft sensor after the raw-material supplier switch, and performance recovered. What they have not built is the machinery that makes the second refresh cheaper than the first — so every refresh is a bespoke project with a bespoke export, a bespoke relabelling effort and a bespoke definition of 'good enough to ship'.

The economics are quietly punishing. A refresh that takes eight weeks means the plant runs a degraded model for two months per incident — and because the trigger is a crisis, the degradation ran unmeasured for months before that. Worse, the refresh is validated against whatever test data the engineer assembled that week. Nothing guarantees the new model still catches the rare defect class the old one was originally bought for; regressions on low-frequency defects are the classic manual-refresh injury, and they surface as field escapes with a customer's name attached.

The stage is also fragile to people. The engineer who ran the last refresh encoded a dozen decisions — which weeks to export, which images to exclude, what counts as a scratch versus a mark — in their head and their laptop. When a different engineer runs the next one, those decisions get remade differently, and the two models bracket a definition drift nobody chose. A stage-2 plant does not have a retraining process; it has retraining incidents with local heroes.

In practice

The relabelling fortnight

A speciality-chemicals plant runs a soft sensor predicting batch viscosity from in-process measurements. When a solvent supplier changed, predictions drifted and off-spec batches reached the QC lab before anyone connected the two. The fix took six weeks: export eighteen months of historian data, relabel against lab results, rebuild, argue about acceptance. Two years later a second supplier change repeated the whole exercise — run by a different engineer, from scratch, with a different test set, because nothing from the first refresh had been kept as reusable machinery.

What it looks like

  • Every retrain is triggered by a complaint, an escape or a false-reject storm
  • Training data is re-exported and relabelled by hand for each refresh
  • Each refresh invents its own test set, so no two are comparable
  • Lead time from 'the model is wrong' to 'the model is fixed' is 6–12 weeks

Diagnostic signals you can check this week

  • Ask when the last retrain happened and what triggered it — 'a customer complaint' is the stage-2 answer
  • Ask to see the current validation set; if the answer is 'which one?', each refresh invented its own
  • Time the loop: date of first degradation evidence to date the fixed model shipped. Weeks means stage 2
  • Check whether the last refresh's data-selection choices are written anywhere a second engineer could follow

Anti-pattern · The heroic refresh

After a successful rescue, the plant declares the problem solved and disbands. The refresh treated a symptom; the disease is that decay is only visible through crises. The anti-pattern is spending the post-rescue quarter on nothing, when the same quarter could wire the monitoring that makes the next decay visible in week one instead of month four — and turn the rescue's one-off scripts into a pipeline. A plant that only retrains when customers complain has outsourced its drift monitoring to its customers.

What holds you here

Every refresh is rebuilt from scratch, so retraining is too slow and too expensive to run before a crisis forces it.

Highest-leverage next move

Version the machinery from the last refresh — the export queries, the labelling rules, the test set — so the next one is a run, not a project. The golden set comes first.

Cost of leaving

Effort
4–9 months
Team
One ML engineer, one data engineer, a named quality-side owner
Risk
Medium — the first standing golden set forces the plant to agree what 'good' means, which is politics as much as engineering
To next stage
4–9 months

If this is you, the next step is

We take your last manual refresh and rebuild it as versioned, repeatable machinery.

Turn the rescue into a pipeline

Stage 3

Scheduled

22% of operators sit here

A standing pipeline retrains on a calendar, validates every candidate against a versioned golden set, and promotes through a human sign-off — the first stage where learning is routine.

Stage 3 is where continuous learning stops being a slogan and becomes a standing capability. The pipeline exists: training data flows from the historian and MES through governed queries, the candidate model is evaluated against a golden set that is itself versioned and stratified, and promotion is a decision a named person takes against written criteria, recorded in the QMS. The refresh that took eight weeks at stage 2 takes days, and — more importantly — it happens before the crisis rather than after it.

The discipline that defines this stage is the golden set. It is the plant's institutional memory of what the model must never forget: every defect class that ever mattered, every SKU in proportion, the near-misses and the edge cases, each with an agreed label. Candidates that improve on average but regress on a rare, expensive defect class get caught here — the exact failure the manual refresh shipped. Golden sets rot, though, and a stage-3 plant treats additions to the set as seriously as changes to the model: every new product introduction adds its rows before launch, not after the first escape.

The constraint that emerges is cadence mismatch. The calendar says quarterly; the plant does not decay quarterly. A stable tablet-compression model gets retrained four times a year for nothing, while the vision model on the fast-churning packaging line is six weeks stale the day it ships. Fixed schedules are the training wheels of continuous learning — the right cadence is the plant's own event stream, and noticing that is what pushes an operator toward stage 4.

In practice

The quarterly retrain that missed the changeover

A pharma packaging operation retrains its blister-inspection model quarterly, with golden-set replay and QA sign-off — genuine stage 3. Mid-quarter, a new foil supplier came in. The foil's reflectivity shifted the imaging distribution, false rejects tripled, and for five weeks operators re-inspected an extra bin per shift while the calendar counted down. The scheduled retrain fixed it on schedule — five weeks after the plant needed it. The post-mortem question was the right one: why does the model learn on the calendar's clock instead of the plant's?

What it looks like

  • Retraining runs on a schedule from a versioned pipeline, historian to registry
  • A versioned golden set gates every promotion, stratified by SKU and defect class
  • Decay is tracked on a dashboard reviewed alongside SPC and quality metrics
  • Every promotion leaves an MOC record: what changed, evidence, approver

Diagnostic signals you can check this week

  • Ask to see the golden set's version history — a stage-3 plant shows commits; a stage-2 plant shows a folder
  • Compare the retrain calendar with the changeover calendar; if they never reference each other, cadence is arbitrary
  • Check whether any promotion was ever refused on golden-set evidence — a gate that has never said no is not a gate
  • Look for one model retrained often and another rarely, with a written reason; uniform cadence for every model is the stage-3 tell

Anti-pattern · One cadence for every model

The pipeline makes retraining cheap, so the plant standardises: everything quarterly. But cadence is a property of the decay driver, not of the pipeline. Vision models on high-churn packaging lines decay per changeover; process models decay per raw-material lot; maintenance models decay when the CMMS closes an overhaul work order and resets a vibration signature. Setting one calendar for all of them over-trains the stable and abandons the volatile. Set cadence per model from its measured decay curve — the curve stage 1 taught you to draw.

What holds you here

Retraining runs on the calendar while the plant decays on events — changeovers, lots, overhauls — so the cadence is always wrong for someone.

Highest-leverage next move

Key retraining to plant events: recipe changes in the MES, lot changes in the ERP, overhaul closures in the CMMS, SPC rule breaches. The calendar becomes the fallback, not the driver.

Cost of leaving

Effort
9–18 months
Team
ML engineer, data engineer, quality engineer as standing approver
Risk
Medium — wiring triggers to MES and ERP events crosses team boundaries, and label capture at the verify station changes operator workflow
To next stage
9–18 months

If this is you, the next step is

We map your decay drivers to MES, ERP and CMMS events and wire the first trigger.

Move from calendar to event triggers

Stage 4

Event-driven

10% of operators sit here

Retraining is triggered by the plant's own events, a challenger runs in shadow against live production, and labels flow back automatically — but every promotion still waits for a human gate.

Stage 4 inverts the direction of attention. At stage 3, people decide when the model should learn; at stage 4, the plant tells them. A recipe change in the MES, a new resin lot booked in the ERP, a gearbox overhaul closed in the CMMS, an SPC rule breach on a critical characteristic — each fires an evaluation, and where the champion degrades, a challenger trains on the freshest labelled window and enters shadow. The loop's raw material is the label stream: verify-station dispositions, QC lab results and rework outcomes land against their source records within hours, because label capture was designed into the operator workflow rather than bolted on.

Shadow deployment is the stage's signature discipline. The challenger scores every live part or batch but acts on none; its verdicts are logged beside the champion's and compared over a defined window that must include the conditions that triggered it — the new lot, the changed recipe. Promotion is then an evidence-backed decision: here is the golden-set replay, here is two weeks of shadow against live production, here is the delta by defect class. The human gate remains, and it is worth being honest that it now sets the loop's cycle time: a challenger validated in three days can wait two weeks for the MOC record and the quality sign-off.

This is also the stage where learning compounds across the estate. The defect-detection model proven on line 3 transfers to line 5 with a fine-tune on local imagery rather than a from-scratch build; the sister plant inherits the pipeline, the golden-set structure and the trigger map. Fleet learning is why stage-4 operators pull away from stage-3 ones on cost: the second line's loop costs a fraction of the first's, and the tenth is configuration. The frontier operators documented by the World Economic Forum's Global Lighthouse Network are recognisable stage-4 shops — instrumented loops, shadow validation, human gates — which is itself the strongest public evidence of where the real frontier currently sits.

In practice

The changeover that retrained the model

An electronics manufacturer's SMT line runs AOI models per assembly family. A new assembly's recipe activation in the MES fires an evaluation: the champion's predicted false-call rate on the new family exceeds its band, so a challenger fine-tunes overnight on the first shift's operator-verified labels. It shadows the champion for three days across two thousand boards, the comparison lands in the quality engineer's queue with golden-set replay attached, and promotion happens inside the week — with the previous version one click away. Total operator-visible disruption: none.

What it looks like

  • Triggers fire from plant events: MES recipe changes, ERP lot changes, CMMS overhauls, SPC breaches
  • A champion–challenger pair runs continuously; the challenger scores live traffic in shadow
  • Operator dispositions and lab results flow back as labels with measured latency
  • Models transfer between lines and sites with local fine-tuning — fleet learning

Diagnostic signals you can check this week

  • Ask what fired the last retrain; the stage-4 answer names a plant event, not a date or a crisis
  • Check for a challenger scoring live traffic right now, and where its shadow log lives
  • Measure label latency: production event to usable labelled record. Hours is stage 4; weeks is not
  • Ask how the second line got its model — 'transferred and fine-tuned from line 3' is the fleet-learning tell

Anti-pattern · Shipping around the paperwork

The loop now outruns the QMS: a challenger is validated in days while the change record takes a fortnight, and the engineering team starts quietly promoting 'minor' model updates outside MOC to keep pace. It works until the day it defines the plant's audit finding — a nonconformance investigation asks which model version passed the affected lot, and the honest answer is that the registry and the QMS disagree. The fix is not more discipline but less friction: make the pipeline generate the change-control evidence automatically, so the compliant path is also the fast path. That automation is precisely the groundwork for stage 5.

What holds you here

Every promotion still waits for a human gate, so the loop's cycle time is set by sign-off and MOC paperwork rather than by training and validation.

Highest-leverage next move

Turn the approval history into policy: mine a year of promotion decisions for the bounds within which quality has always said yes, and draft the learning licence from them.

Cost of leaving

Effort
18+ months
Team
Platform engineer, ML engineer, quality partner with delegated approval authority
Risk
Higher — trigger wiring spans MES, ERP and CMMS ownership boundaries, and approval-lead-time politics surface between engineering and quality
To next stage
18+ months

If this is you, the next step is

We automate the evidence a promotion needs so sign-off takes a day, not a fortnight.

Design the promotion evidence pack

Stage 5

Licensed

3% of operators sit here

Qualifying model updates promote automatically inside a versioned learning licence agreed with quality — bounded autonomy with auto-generated evidence, not self-evolution.

Stage 5 is narrower than the phrase 'autonomous learning' suggests, and the narrowness is the point. A learning licence is a quality-system artefact: a versioned document, owned jointly by engineering and quality, that enumerates for one named model what may change without human sign-off — retraining on new data with a fixed architecture, decision thresholds within a stated band — and what evidence each automatic promotion must generate: golden-set score floors, shadow-window minimums, per-defect-class regression limits, an auto-rollback trigger. Inside the licence, the loop closes at machine speed. Outside it, nothing moves without a person, exactly as at stage 4.

The honest context is that almost no manufacturer operates here today, and the page's own distribution reflects that. The intellectual scaffolding exists — regulators in adjacent domains have formalised predetermined-change frameworks for machine-learning systems, and the same logic maps cleanly onto a manufacturing MOC — but public, named examples of auto-promoted model updates in production plants are scarce. That is not a reason to dismiss the stage; it is the reason to treat everything below it as the qualification. A licence is only as credible as the promotion history behind it, and the only way to accumulate that history is to run a supervised loop well, for a long time.

Sustaining stage 5 is renewal discipline. The licence encodes assumptions — this product mix, this camera fleet, this defect taxonomy — and the plant changes underneath all of them. The escalation rate is the canary: when the share of candidates falling outside the licence rises, the world has left the licence's validity envelope and the review should happen before an incident forces it. The failure mode is never the dramatic rogue model; it is licence creep — bounds quietly widened after each escalation until the licence describes what the system does rather than what quality agreed. Version the licence, review it like code, and let it shrink as willingly as it grows.

In practice

The bounded licence on the packaging line

A packaging operation runs its carton-print inspection model under a licence: retrains may promote automatically if golden-set recall on every defect class stays above its floor, false-reject rate on a three-day shadow window stays inside its ceiling, and the decision threshold moves less than a stated delta. Each auto-promotion writes an evidence pack into the QMS; roughly one candidate in eight falls outside and lands in the quality engineer's queue. When a new carton substrate pushed escalations from one-in-eight towards one-in-three, the licence review happened that month — the point of the metric is that the review pre-empted the incident.

What it looks like

  • A versioned learning licence states what may change without sign-off, by how much, evidenced how
  • Updates inside the licence promote automatically; every release writes its evidence pack
  • Out-of-bounds candidates escalate to a person — and the escalation rate is itself monitored
  • Auto-rollback triggers and the kill switch are drilled, with the licence reviewed on a fixed cycle

Diagnostic signals you can check this week

  • Ask to read the licence — a stage-5 plant hands you a versioned document with named owners, not a description of engineering culture
  • Check that every automatic promotion wrote an evidence pack the QMS can produce on demand
  • Ask when the kill switch and auto-rollback were last exercised deliberately; 'never' means untested
  • Look at the escalation-rate trend and who reviews it — an unmonitored escalation rate means the licence is drifting unobserved

Anti-pattern · Licence creep

Each escalation is resolved by widening the bound that caught it — the golden-set floor nudged down, the threshold band nudged out — with no review of whether the world changed or the licence was simply inconvenient. Eighteen months later the licence permits nearly everything and certifies nothing, and the first serious audit finds a policy that was rewritten by its own exceptions. Treat every proposed widening as a licence change with quality sign-off, and record the narrowing decisions too: a licence that only ever grows is not being governed.

What holds you here

Keeping the licence valid as the plant changes — renewal, escalation review and the discipline to narrow bounds — is a permanent governance cost, not a project.

Highest-leverage next move

Review the licence on the plant's change calendar, not the audit calendar: every new product introduction, substrate change or camera refresh is a licence question before it is a retraining trigger.

Cost of leaving

Effort
Continuous
Team
Platform team plus a standing engineering–quality review forum
Risk
Concentrated — low-frequency, high-consequence, and regulatory in character; the licence itself is what an auditor will examine

If this is you, the next step is

We red-team the bounds, the evidence packs and the rollback against a real excursion scenario.

Stress-test a learning licence

Where manufacturing plants actually sit on the ladder

The distribution across the five stages, and why the frozen–refreshed plateau holds two thirds of the industry.

Most plants sit in the first two stages: their models are frozen, or retrained only when a crisis forces it. That is the honest reading of the adoption research — AI use is now near-universal at the organisational level, while the operational machinery that keeps models accurate remains rare. The distribution below shows the shape: a large frozen base, a plateau of crisis-driven refreshers, and a thin frontier of event-driven and licensed loops.

Distribution of manufacturers across the cadence ladder

Illustrative distribution — charted for orientation, not measurement. Synthesised from McKinsey's State of AI adoption research and the operating patterns documented across WEF Global Lighthouse Network sites; the frozen–refreshed plateau is where roughly two thirds of plants sit.

Share of plants (illustrative)

  • 34% — 1 · Frozen (the silent majority)
  • 31% — 2 · Manually refreshed
  • 22% — 3 · Scheduled
  • 10% — 4 · Event-driven
  • 3% — 5 · Licensed

Source: Illustrative, synthesised from McKinsey State of AI and WEF Global Lighthouse Network research

The frontier is public. The Global Lighthouse Network (opens in a new tab) — the World Economic Forum's programme identifying factories deploying fourth-industrial-revolution technology at production scale — documents, site by site, what the top of the ladder looks like in practice: closed-loop quality systems, models maintained across whole fleets of lines, and measured impact on OEE, yield and energy. Two details in those write-ups matter for this page. First, the lighthouse use cases are overwhelmingly supervised loops — the human promotion gate appears everywhere. Second, the lighthouses report their capabilities in loop terms (how fast a model tracks the plant) rather than deployment terms (how many models exist), which is exactly the shift the cadence ladder measures. MHI's annual industry survey (opens in a new tab) tracks the same adoption-versus-impact gap from the supply-chain side of manufacturing.

What decays, where: the model families of a manufacturing plant

Six model families, the system each lives in, the plant event that makes each one decay — and the cadence each actually needs.

Every model family in a plant decays for a different reason, and the decay driver — not the model type — is what sets the right retraining cadence. A vision model on a high-churn packaging line decays at every changeover; a predictive-maintenance model decays when the overhaul it predicts actually happens and resets the signature it learned; a demand model decays with the season. The map below is how we scope loops with operators: read your most expensive model's row, and its cadence column is the stage the ladder says that model needs — whatever stage the rest of the estate is at.

Model familyWhat it decidesSystem of recordWhat makes it decayCadence sweet spot
Vision inspection / AOIAccept, reject, rework routingMES + inspection stationNew SKUs, label and substrate changes, camera and lighting driftStage 4
Process soft sensors & APC setpointsSetpoints and quality predictions within the control windowDCS / APCRaw-material lot variation, catalyst ageing, fouling, sensor recalibrationStage 3–4
Predictive maintenanceIntervention timing, spares stagingCMMS + historianOverhauls resetting signatures, operating-regime change, replaced componentsStage 3
Demand & production schedulingSequence, batch sizing, campaign lengthERP / APS + MESSeasonality, portfolio churn, promotions, customer mixStage 3
Energy optimisationLoad shifting, compressor and utilities schedulingSCADA / BMSTariff changes, weather seasonality, line reconfigurationStage 3
Batch golden-profile monitoringDeviation detection against the reference trajectoryHistorian + QMSRecipe changes, vessel maintenance, instrument recalibrationStage 4
The decay map for a non-automotive manufacturing estate. 'Cadence sweet spot' is the ladder stage at which the model family's decay driver is actually covered; running a family at a lower stage means its decay outruns its learning.

Two readings of the map are worth making explicit. First, the families that decay on discrete, loggable events — vision models at changeovers, batch profiles at recipe changes — are the natural first candidates for event-driven learning, because the trigger already exists as an MES or ERP record and merely needs wiring. Families that decay on slow, continuous drivers — fouling, seasonality — are well served by scheduled cadence and gain less from triggers. Second, the map is why 'what stage is your plant?' is really a portfolio question: a sensible estate runs its packaging-line AOI at stage 4 and its energy model at stage 3, and the assessment on this page scores the shared machinery — pipeline, golden sets, governance — that all of them draw on.

One family deserves a boundary note: models inside safety-instrumented functions do not belong on this ladder at all. A trip function or interlock is certified against frozen, verified behaviour under functional-safety standards, and 'continuously learning' is precisely the property certification excludes. The correct pattern — visible in every credible deployment — is that learning systems advise and optimise up to the boundary of the safety envelope, and the certified layer beneath them stays frozen. A vendor pitching a learning model inside the safety loop is pitching a recertification programme, and probably does not know it.

The change-control wall: what separates real from speculative

Every claim about the self-optimising factory sorts into deployed, demonstrated, or speculative — and the sorting variable is change control, not algorithms.

The wall between today's supervised loops and the imagined self-evolving factory is change control, and it is a load-bearing wall rather than an obstacle. A manufacturing plant runs on the discipline that process changes are proposed, risk-assessed, evidenced and approved before they touch product — the MOC process of an ISO 9001-style QMS, hardened further in regulated sectors where process validation demands that any change to a validated state be requalified. A model update is a process change. That single sentence explains most of the distance between AI marketing and plant reality: the constraint on learning speed was never training compute; it is the rate at which change evidence can be honestly generated and reviewed.

CapabilityStatusWhat bounds it
Drift monitoring and scheduled retraining with a human promotion gateDeployed today, widely — ordinary engineeringCost and discipline only
Event-triggered retraining with shadow validation and fleet transferDeployed at leading sites; documented across WEF lighthouse factoriesLabel capture and MOC throughput
Closed-loop setpoint adjustment inside a fixed, engineered envelopeDeployed for decades as APC; AI now proposes better envelopesThe envelope is engineered and certified by humans
Automatic promotion of retrained models under a versioned learning licenceNascent — the policy template exists; few public production examplesQuality-system change control and accumulated approval evidence
Continuous learning inside safety-instrumented functionsBlocked — certification requires frozen, verified behaviourFunctional-safety standards (IEC 61511-class certification)
A factory whose models rewrite their own objectives and authoritySpeculative — no published, named production deploymentEvidence cannot be generated faster than product can be measured
The reality ledger for the continuously learning factory. Every capability on this page sorts into one of these rows; the sorting variable is what bounds it, and in every case the bound is evidential or regulatory rather than algorithmic.

The last row's bound deserves one more sentence, because it is physics rather than caution: a model update is validated by comparing predicted quality against measured quality, and measurement takes as long as the process takes. A batch that cures for eight hours yields one label every eight hours, however fast the GPU. Learning speed in a factory is throughput-limited by the plant's own measurement cycle — which is why the credible frontier is a loop that wastes none of that scarce evidence, not a system that somehow transcends it.

Read the ledger bottom-up and the strategic conclusion falls out: everything above the wall is earned by running everything below it. The licence bounds of row four are mined from the approval history of row two; the approval history exists only if the loop exists; the loop exists only if decay is measured. This is the sense in which future-readiness is present-readiness — a plant cannot buy its way to licensed autonomy, because the purchase price is denominated in its own accumulated promotion evidence. acatech's maturity research (opens in a new tab) reaches the same conclusion from the organisational side: its 'adaptability' stage is defined as the outcome of the preceding stages' data and process discipline, not as a technology acquisition.

What the loop looks like in public

Three publicly reported programmes, read against the cadence ladder. None is an Atomic Loops engagement — each links to the operator's own material or the research programme that documented it.

The clearest public evidence for the ladder is in what the most-studied factories actually run. In each case below, the differentiator is not model sophistication — it is that retraining, validation and release are wired into plant operations, and that a human gate still sits on promotion. Read them for the shape of the loop; none of the three claims licensed autonomy, which is itself the honest headline about where the industry's frontier sits.

Three programmes read against the ladder

Outcomes as reported by the operators themselves or by the WEF's lighthouse documentation. Verify figures against the linked source before reusing them; we have not independently audited them. Images are generated industry scenes, not operator photography.

Electronics manufacturing line with automated optical inspection and analytics overlays — illustrative sceneSiemens — Electronics Works AmbergElectronics manufacturing · PCB assembly at ~one product per second24
Challenge
Sustaining extreme quality on high-mix PCB assembly — hundreds of product variants, continuous product change — where X-ray and optical inspection capacity is expensive and every new variant shifts the data the inspection models see.
Approach
Siemens has publicly presented Amberg as its showcase for AI embedded in production: inspection models maintained against live process data, closed-loop analytics deciding where physical inspection effort is spent, and model updates handled inside the plant's engineering and quality process rather than as external data-science projects.
Reported outcome
Siemens reports process quality at Amberg in the region of 99.9989% sustained alongside decades of product churn, and presents the plant's AI-supported X-ray inspection as a headline example of models reducing physical inspection load. The plant is among the sites recognised in the WEF's Global Lighthouse Network.
What it shows about the curveAmberg's quality figure is not a model property — no frozen model survives that product churn. It is a loop property: the models are maintained at the cadence of the plant's own change, which is what stage 4 means.

Siemens press (landing — Amberg coverage) (opens in a new tab)

Darkened automated electronics assembly hall with robotic stations and monitoring screens — illustrative sceneFoxconn — Shenzhen 'lights-out' factoryContract electronics manufacturing · precision components at extreme volume24
Challenge
Running a highly automated precision-machining and assembly operation with minimal on-floor staffing — where no operator is present to notice a decaying model, so machine-driven quality decisions must be monitored and corrected systemically.
Approach
Foxconn's Shenzhen site was among the earliest factories admitted to the WEF's Global Lighthouse Network, documented for combining automated production with self-adjusting machining and quality systems that are tuned from accumulated process data across the fleet of identical stations — learning applied at fleet scale rather than machine by machine.
Reported outcome
Through the lighthouse programme's published documentation, the site reported material production-efficiency gains and reduced inventory cycles from its automation-plus-learning approach; the network's write-ups present it as a reference case for what sustained, systemic optimisation of an automated line looks like.
What it shows about the curveLights-out operation raises the loop's stakes rather than lowering them: with no human noticing decay at the station, decay visibility and fleet-level learning become the safety net. Automation without a loop is a frozen plant that merely fails faster.

WEF Global Lighthouse Network (opens in a new tab)

Semiconductor cleanroom with wafer-handling automation and process-monitoring displays — illustrative sceneBosch — Dresden wafer fabSemiconductor manufacturing · 300mm wafer fabrication34
Challenge
Ramping a new 300mm fab to mature yield — a process with hundreds of steps where deviations must be caught from data long before end-of-line test, and where every process adjustment is a controlled change in a certified environment.
Approach
Bosch designed Dresden from the start as what it calls an AIoT factory: comprehensive process data collection, AI models classifying wafer-surface defect patterns from inspection imagery, and predictive analysis of production equipment — with the models improved continuously from the fab's accumulated data, as Bosch describes in its own account of the plant.
Reported outcome
Bosch reports that AI evaluation of production data at Dresden detects process anomalies and defect signatures early enough to correct processes before defective wafers progress, and credits the data-driven approach with accelerating the fab's ramp — publicly positioning Dresden as one of the world's first AI-native fabs.
What it shows about the curveA greenfield plant can be born at stage 3 — pipelines and monitoring designed in from day one — but stage 4 is still earned in operation, because event-driven learning needs the plant's own accumulated events. The ladder has no skippable rungs, only cheaper ones.

Bosch — chip factory Dresden (opens in a new tab)

The learning-loop architecture, layer by layer

What has to exist at each stage of the ladder — defined by what each layer must guarantee, not by any vendor's product.

A supervised loop requires five layers, and the build order matters more than the tooling. The architecture below is deliberately vendor-neutral: each layer is defined by what it must guarantee to the layer above, and each is annotated with the ladder stage that first requires it. The data path follows the plant's existing standards — OPC UA (opens in a new tab) for contextualised data out of the control layer, the ISA-95 (opens in a new tab) level model for where the MES sits between controls and enterprise systems, and the MESA model (opens in a new tab) for how execution functions interconnect. A loop that respects those boundaries inherits the plant's existing integration discipline; one that bypasses them becomes the next generation's legacy.

Layers required by stage

Each layer is annotated with the ladder stage that first requires it. A plant attempting event-driven learning without the label-capture layer is building a stage-2 refresh habit with more expensive tooling.

  1. Line and control systems

    Stage 1+

    • PLC / DCS and field sensorsThe physical truth the models describe
    • SCADA and historianTime-series memory, OPC UA out
    • MES and quality stationsOrders, recipes, dispositions — the decision context
  2. Data and label capture

    Stage 2+

    • Contextualised ingestionProcess data joined to order, recipe and lot context
    • Label routingOperator dispositions, lab results and rework outcomes flow back automatically
    • Feature and label storeOne governed definition per signal, with lineage
  3. Training and registry

    Stage 3+

    • Versioned retraining pipelineHistorian to candidate, reproducible on demand
    • Model registryEvery version, its training window and its lineage
    • Golden validation setsVersioned, stratified by SKU and defect class
  4. Release and serving

    Stage 3+

    • Shadow / champion–challengerCandidates score live traffic without acting
    • Staged rollout and rollbackOne line first; previous version one click away
    • Write-back with approvalInto the MES or APC field the operator already reads
  5. Loop governance

    Stage 4+

    • Decay and trigger monitoringKeyed to MES, ERP and CMMS events; pages a named owner
    • Decision audit logModel version recorded on every accept/reject
    • Learning licenceVersioned policy for auto-promotion (stage 5)

Pipeline described

  1. Line and control systems (stage 1+) — PLC / DCS and field sensors: The physical truth the models describe; SCADA and historian: Time-series memory, OPC UA out; MES and quality stations: Orders, recipes, dispositions — the decision context
  2. Data and label capture (stage 2+) — Contextualised ingestion: Process data joined to order, recipe and lot context; Label routing: Operator dispositions, lab results and rework outcomes flow back automatically; Feature and label store: One governed definition per signal, with lineage
  3. Training and registry (stage 3+) — Versioned retraining pipeline: Historian to candidate, reproducible on demand; Model registry: Every version, its training window and its lineage; Golden validation sets: Versioned, stratified by SKU and defect class
  4. Release and serving (stage 3+) — Shadow / champion–challenger: Candidates score live traffic without acting; Staged rollout and rollback: One line first; previous version one click away; Write-back with approval: Into the MES or APC field the operator already reads
  5. Loop governance (stage 4+) — Decay and trigger monitoring: Keyed to MES, ERP and CMMS events; pages a named owner; Decision audit log: Model version recorded on every accept/reject; Learning licence: Versioned policy for auto-promotion (stage 5)
Step-by-step insights
Line and controls — the loop starts below the model
The loop's ceiling is set in the control layer: if the historian samples a critical signal at the wrong resolution, or dispositions are recorded on paper at the verify station, no downstream architecture recovers the loss. The first audit of any loop build is unglamorous — signal inventory, sampling rates, disposition capture points — and it routinely finds that the single highest-leverage investment is a barcode scanner or an extra disposition code at a quality station, not anything with 'AI' in its name.
Data and label capture — context is what makes plant data trainable
Raw historian tags are almost untrainable: the same temperature trace means different things on different recipes, lots and tool states. The capture layer's real work is contextualisation — joining process data to the MES order, recipe version and material lot it belongs to, so a training window can be selected by meaning rather than by timestamp. This is also where label latency is won or lost: the elapsed time from a production event to a usable labelled record should be measured in hours, and it is the single number that best predicts how fast the whole loop can turn.
Training and registry — reproducibility is a quality requirement, not a nicety
In a plant, 'which data trained the model that passed this lot?' is a quality question that will eventually be asked by an auditor or a customer, and the registry is what makes the answer take minutes instead of days. Every model version links to its training window, pipeline version and golden-set result; every golden set is itself versioned, and additions arrive with product introductions rather than after escapes. The registry is to models what the device history record is to product — and it earns the same respect.
Release and serving — shadow first, one line first, rollback always
The release layer encodes the plant's humility: no candidate touches product until it has scored live traffic in shadow, no rollout goes fleet-wide before one line has run it, and the previous version is one click away with the click drilled on a quiet shift. The write-back target matters as much as the gate — output lands in the MES or APC field the operator already reads, because a recommendation behind a separate login decays into a dashboard nobody opens exactly when the plant is busiest.
Loop governance — the audit trail is a by-product, or it is a project
Built correctly, governance evidence accrues automatically: every trigger, shadow comparison, promotion decision and rollback is a log entry the QMS can produce on demand, and the eventual learning licence is drafted from that accumulated history. Built as an afterthought, the same evidence is a quarterly assembly project that gets skipped under pressure — and its absence is discovered during the one investigation where it mattered. The design rule is simple: if a promotion needs a document, the pipeline writes the document.

The layer most often skipped is label capture, and it is the one that decides the loop's cycle time. Plants reliably fund the visible layers — training pipelines and serving are demonstrable — and defer the workflow change that makes operator dispositions flow back automatically. The result is a stage-3 pipeline fed by stage-2 labelling: retraining is push-button, but assembling its training data still takes six weeks, and the loop's advertised cadence is a fiction. Fund the fuel line before the engine.

A 90-day plan: one vision model, frozen to supervised

The stage 2 → 3 transition made concrete on one specific problem — a packaging-line inspection model whose false rejects climb after every SKU changeover. Contains no new model development.

Moving one stage takes about 90 days when scoped to a single model, and multiple years when scoped to an estate. The plan below runs the transition on a specific, common problem: a vision inspection model on a high-speed packaging line — cartons, labels, fill levels — whose false-reject rate climbs after every SKU changeover, quietly costing re-inspection labour and yield until the next crisis-funded refresh. The model itself already works; the quarter builds the loop around it, and contains no new model development at all.

Frozen to supervised on one inspection model, in one quarter

One line, one model, one named owner. If any phase needs more than its window, narrow the scope — fewer SKUs, one shift — rather than extending the plan.

  1. Days 1–15

    Instrument the decay

    Wire the comparison that stage 1 never had: every model verdict logged against the verify station's operator disposition, per SKU, per shift. Pull six months of QMS re-inspection records to reconstruct the false-reject history around past changeovers. Name the line's quality engineer as owner — false rejects and escapes are already their numbers.

    A decay curve per SKU, and one named owner

  2. Days 16–45

    Build the golden set and the label path

    Assemble the versioned golden set: every defect class in the inspection specification, every active SKU in proportion, the historical escapes, the borderline images operators argue about — each with an agreed label, signed off by quality. In parallel, make the verify station's dispositions flow back automatically as labels, with latency measured. This is the fuel line; it outlasts every model.

    A signed-off golden set v1, labels flowing at known latency

  3. Days 46–70

    Script the retrain and run the first supervised cycle

    Turn the last manual refresh into a versioned pipeline: training window selected by MES context (SKU, lot, recipe version), candidate evaluated by golden-set replay, results in a one-page evidence pack. Run one full cycle: retrain on the recent labelled window, shadow the candidate against the live model for one week of production, and promote through a real MOC record with the quality engineer's sign-off. Drill the rollback once, deliberately, on a quiet shift.

    One evidence-gated promotion completed; rollback drilled

  4. Days 71–90

    Attribute, and wire the first trigger

    Report the delta in the units the plant already tracks: false-reject rate by SKU, re-inspection labour hours per shift, first-pass yield against the pre-loop baseline — with the decay curve from phase 1 as the counterfactual. Then wire the first event trigger: a new SKU's recipe activation in the MES automatically fires a golden-set evaluation and, on breach, a retrain. Write the cadence into the control plan so the loop survives its builders.

    An attributed delta the budget accepts, and the first trigger live

The order matters

  1. Measurement before retraining

    The decay curve from days 1–15 is what funds everything after it — and occasionally it shows the model is fine and the problem is lighting or a lens. Retraining before measuring risks solving the wrong problem with the most expensive tool available.

  2. The golden set before the pipeline

    A pipeline without a golden set automates the promotion of unvalidated models — it makes the plant worse, faster. The golden set is also the politically hard part, because it forces quality and engineering to agree what 'good' means; do it while goodwill is high.

  3. One supervised cycle before any trigger

    Run the full loop by hand once — retrain, shadow, evidence pack, sign-off, rollback drill — before automating any of it. The manual cycle surfaces every gap in the evidence pack while the stakes are low, and gives quality a concrete artefact to approve rather than a concept.

Verifying the loop — and the four ways it goes wrong

The loop KPIs that prove learning is real, read from telemetry rather than self-report — and the failure modes that quietly send plants back down the ladder.

A learning loop is verified from its own telemetry, not from anyone's assessment answers. Every metric below reduces to timestamps and counts that the registry, the QMS and the MES already record once the loop exists — and each has a threshold that separates one ladder stage from the next. The measurement discipline here is the same one national metrology programmes bring to manufacturing generally (see NIST's manufacturing programmes (opens in a new tab)): a number nobody can trace to an instrument is an opinion.

Loop KPIFormula / readSourceHonest from
Model ageToday − training-data cutoff of the live versionModel registryStage 1
Decay ratePerformance delta per month vs live labels, per SKUDecision log + dispositionsStage 1
Label latencyProduction event → usable labelled record, elapsedMES / QMS timestampsStage 2
Retrain lead timeTrigger (or request) → validated candidate, elapsedPipeline + registryStage 3
Promotion lead timeValidated candidate → live in production, incl. MOCRegistry + QMSStage 3
Shadow coverageLive decisions scored by a challenger ÷ all decisionsServing logStage 4
Rollback drill recencyDays since the revert was last exercised deliberatelyChange logStage 3
Escalation rateOut-of-licence candidates ÷ all candidatesLicence logStage 5
The loop instrumentation sheet. 'Honest from' is the ladder stage at which the KPI first measures something real — quoting model age at stage 1 is easy (it is the commissioning date); quoting retrain lead time honestly requires a pipeline to exist.

Two of these deserve a habit, not just a dashboard. Model age is the estate-level health metric: plot it for every model in production, and the frozen tail of the histogram is your quiet risk register. Promotion lead time is the governance health metric: when it grows while retrain lead time shrinks, the loop is outrunning the QMS, and the fix is the auto-generated evidence pack — before the engineering team finds the workaround that becomes an audit finding.

Likelihood: highImpact: high

Feedback contamination — the model starts grading its own homework

Once operators trust the model, they stop re-inspecting what it passes, so escapes vanish from the label stream and retraining data increasingly contains only the model's own confident verdicts. Performance metrics improve while real performance quietly degrades — the loop is learning from a mirror.

PreventionA standing audit sample: a small random fraction of passed product is always independently re-inspected, and those labels are flagged as the only unbiased ground truth.

Likelihood: highImpact: medium

Golden-set rot — the exam stops matching the plant

The golden set assembled at loop launch slowly stops representing the product mix: new SKUs are missing, retired defect classes dominate, and candidates pass replay while failing the actual line. Every promotion is then validated against a museum.

PreventionGolden-set additions are a mandatory line item in every new product introduction — the set is versioned, and its SKU coverage is reviewed against the active catalogue quarterly.

Likelihood: mediumImpact: high

Retraining on an excursion — the model learns the deviation as normal

A trigger fires during an unflagged process excursion — a drifting sensor, an off-spec lot that QC has not yet dispositioned — and the pipeline dutifully trains the deviation in. The new model now considers the excursion normal, and will not flag its recurrence.

PreventionTraining windows are gated on process state: periods under SPC violation, open deviations or pending dispositions are excluded from training data automatically.

Likelihood: mediumImpact: high

The loop outruns the change record

Retraining becomes cheap, sign-off stays expensive, and 'minor' updates start shipping outside MOC to keep pace. The registry and the QMS drift apart until a nonconformance investigation asks which model version passed the affected lot — and gets two answers.

PreventionThe pipeline writes the change-control evidence automatically, so the compliant path costs no more than the shortcut — and a registry-to-QMS reconciliation runs monthly.

Supervised-loop readiness checklist

If you cannot tick all seven, the model is still effectively frozen or crisis-refreshed, whatever the tooling around it. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Tick honestly — the blank list is data too

Zero ticks is the frozen estate, and it is the most common honest answer. Don't start with tooling: pick the one model whose errors cost most and run days 1–15 of the 90-day plan above — the decay curve makes the case for everything else on this list.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Model decay
The degradation of a deployed model's real-world performance as the plant moves away from the conditions it was trained on. Distinct from a broken model: a decayed model still runs, still returns confident answers, and is wrong more often than anyone measures.
Drift
The movement of live input or label distributions away from the training distribution — in a plant, driven by SKU changeovers, raw-material lot variation, tool wear, maintenance resets and sensor recalibration rather than by randomness.
Golden set
A versioned, quality-approved validation dataset — every defect class that matters, every SKU in proportion, historical escapes included — that every retrained candidate must pass before promotion. The plant's institutional memory of what a model must never forget.
Shadow deployment
Running a candidate model against live production without letting it act: its verdicts are logged beside the live model's for comparison. The evidence bridge between offline validation and production trust.
Champion–challenger
The serving pattern in which the live model (champion) is continuously compared against a retrained candidate (challenger) scoring the same traffic in shadow, with promotion decided from the logged comparison.
Label latency
The elapsed time from a production event to a usable labelled record — an operator disposition, a lab result, a rework outcome landing against its source data. The single best predictor of how fast a learning loop can turn.
Retrain lead time
Elapsed time from a retraining trigger (or request) to a validated candidate model. The pipeline half of loop cadence; promotion lead time is the governance half.
Feedback contamination
The failure mode in which a model influences its own training data — operators stop re-inspecting what it passes, escapes vanish from the labels, and the loop learns from a mirror. Prevented by a standing, independently inspected audit sample.
Learning licence
A versioned quality-system policy stating, for one named model, what may change without human sign-off, within what bounds, evidenced how, and with what auto-rollback triggers. The artefact that makes stage-5 autonomy compatible with change control.
Fleet learning
Transferring a model proven on one line or site to others with local fine-tuning rather than from-scratch rebuilds — the mechanism by which a second loop costs a fraction of the first.
Management of change (MOC)
The quality-system process by which plant changes are proposed, risk-assessed, evidenced and approved. A model update is a process change; the loop's release path either integrates with MOC or accumulates audit risk.

Frequently asked questions

The questions plant engineering and quality teams ask most often when putting a learning loop around production models.

What is a continuously learning factory?

A continuously learning factory is one whose deployed AI models are kept accurate as the plant changes: post-deployment performance is measured against live labels, retraining is triggered by the plant's own events, every candidate passes golden-set replay and a shadow run, and releases go through change control. It is not a factory that changes its own rules — every credible deployment keeps human-set bounds, and most keep a human promotion gate.

How often should factory AI models be retrained?

On the cadence of the model's decay driver, not a universal calendar. Vision models on high-churn packaging lines decay per SKU changeover and suit event-triggered retraining; process soft sensors decay per raw-material lot; maintenance models decay when overhauls reset the signatures they learned; demand models decay seasonally. Measure the decay curve first — the retraining interval should be shorter than the time the curve takes to cross the cost threshold you care about.

Is continuous learning the same as online learning?

No, and the difference matters in a plant. Online learning updates a model incrementally on each new observation, which makes behaviour hard to validate and impossible to tie to a change record — a poor fit for a QMS. Factory continuous learning is batch learning run often: discrete retrains on governed data windows, each producing a versioned candidate that is validated and released under change control. The loop is continuous; every individual change is discrete, evidenced and reversible.

How do you know a deployed model has decayed?

By comparing its verdicts against ground truth the plant already produces: operator dispositions at verify stations, QC lab results, rework outcomes. Log the comparison per SKU and per shift, chart it like any SPC metric, and alert a named owner on threshold breach — with extra scrutiny after changeovers, new lots and maintenance. A plant that cannot produce last month's false-reject rate from data has no decay visibility, whatever its tooling.

Can models retrain automatically without breaking ISO 9001 or GxP change control?

Yes, but only under a predefined policy — what this page calls a learning licence. Change-control regimes require that changes to a validated state be assessed and evidenced; they do not require that a human perform each assessment manually. A versioned policy that states what may change, within what bounds, with what auto-generated evidence and rollback, satisfies the intent — regulators in adjacent domains have formalised the same idea as predetermined change-control plans for machine-learning systems. What breaks compliance is silent retraining outside any policy.

What data do we need before starting a learning loop?

Less than a platform programme suggests. For one model you need its live verdicts, the ground-truth dispositions the plant already generates, and enough context — SKU, recipe, lot — to slice performance by the events that drive decay. Most plants have all three in the historian, MES and QMS; the missing piece is usually the join, plus disposition capture at one workstation. Instrument one model end to end before generalising anything.

What is a golden set, and how big does it need to be?

A golden set is the versioned validation dataset every retrained candidate must pass: all defect classes in the inspection specification, all active SKUs in realistic proportion, historical escapes, and the borderline cases operators dispute — each with a quality-approved label. Size follows coverage, not a round number: enough examples per defect class per SKU for the pass thresholds to be statistically meaningful, which typically lands in the low thousands of items for a packaging vision model. Its refresh discipline matters more than its size.

What is shadow deployment and why does it matter in a plant?

Shadow deployment runs a candidate model on live production without letting it act — verdicts are logged and compared against the live model's over a defined window. It matters because a plant's offline validation can never fully contain today's conditions: the new lot, the changed recipe, the recalibrated camera. The shadow run is the evidence that converts a promotion from a judgement call into a comparison — and it is what quality signs with confidence.

Does continuous learning apply to safety functions?

No. Safety-instrumented functions — trips, interlocks, protective stops — are certified against frozen, verified behaviour under functional-safety standards, and a model that changes is a model whose certification evidence no longer describes it. The credible pattern keeps learning systems on the optimisation side of the boundary: they advise and adjust within an engineered envelope, while the certified safety layer beneath them stays frozen. Any proposal to put learning inside the safety loop is a recertification programme wearing an AI costume.

What does building a learning loop cost in team terms?

For the first model: roughly one ML engineer and one data engineer for a quarter, plus real time from a named quality engineer — the 90-day plan on this page is that shape. The recurring cost is smaller but permanent: monitoring review, golden-set maintenance at each product introduction, and the promotion gate's time. The second loop costs a fraction of the first if the pipeline and label capture were built as shared machinery, which is the strongest argument for building the first one properly.

How does this ladder relate to acatech's Industrie 4.0 Maturity Index?

They measure different axes and agree at the top. The acatech index measures an organisation's overall digital-industrial capability across six stages, culminating in 'predictability' and 'adaptability' — the self-optimising factory. This page's ladder measures one specific capability inside that journey: how fast, and under whose control, deployed models can safely learn. A plant can be acatech-mature and still run frozen models; the ladder is the instrument for that specific gap, and its licensed stage is what acatech's adaptability looks like at the level of a single model's change control.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for manufacturing, logistics and energy operators — vision inspection, soft sensors, process optimisation and decision support running against live plant data, integrated into the MES and control layer rather than delivered as dashboards.

  • · Production deployments across discrete, batch and continuous manufacturing
  • · Retraining and release pipelines built jointly with plant quality teams
  • · Integration-first delivery: MES write-back, drift monitoring, drilled rollback
  • · 11 cited sources on this page

Sources

  1. McKinsey & CompanyThe State of AI (opens in a new tab)
  2. McKinsey & CompanyOperations insights (AI-enabled maintenance research) (opens in a new tab)
  3. World Economic ForumGlobal Lighthouse Network (opens in a new tab)
  4. acatech — National Academy of Science and EngineeringIndustrie 4.0 Maturity Index (2020 update) (opens in a new tab)
  5. NISTManufacturing programmes (opens in a new tab)
  6. International Society of AutomationISA-95 Enterprise-Control System Integration (opens in a new tab)
  7. OPC FoundationOPC Unified Architecture (opens in a new tab)
  8. MESA InternationalThe MESA Model (opens in a new tab)
  9. SiemensSiemens press (Electronic Works Amberg coverage) (opens in a new tab)
  10. BoschBosch chip factory Dresden (opens in a new tab)
  11. MHIAnnual Industry Report (opens in a new tab)

Find out where your loop actually is — then close the gap

We run the assessment with your engineering and quality leads, review the registry and MOC records against the answers, and leave you with a costed 90-day plan for your weakest dimension. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.