Redefining Technology

Silicon Wafer EngineeringReadiness & Transformation Roadmap

Fab AI transformation phases: sequencing an AI programme in silicon wafer engineering

Fab AI transformation phases are the qualification states an AI system passes through inside a silicon wafer fab — off-tool, shadow, engineer-approved, closed loop within limits, and copy-exact release. Each phase ends at a gate, and a gate is closed by signed evidence, not by a date on a programme plan.

Generated scene: wafer fab cleanroom with process tools and gowned engineers at tool consoles across a bay
Silicon Wafer Engineering · Readiness & Transformation Roadmap

Key takeaways

  1. A fab AI phase is a qualification state, not a calendar period. The five states are off-tool, shadow, engineer-approved, closed loop within limits, and copy-exact release — and a programme is in the lowest phase any of its decisions actually occupies, not the highest one it has demonstrated.
  2. Four gates separate the five phases, and each is closed by an artefact the fab already knows how to issue: a data qualification report, a marathon evidence pack, an approval-log analysis with a control envelope, and a propagation package with chamber-matching evidence. If a phase ended without one of those, it did not end.
  3. Shadow is the phase that consumes years. A model running beside the incumbent EWMA controller costs almost nothing to keep and generates almost nothing that closes a gate, so it survives every budget review and never advances. The exit condition is a marathon window covering the fab's known regime changes — PM cycles, chamber swaps, product mix — not a better backtest.
  4. Skipping a phase is more expensive than repeating one. The recognisable failures — a closed loop qualified on one chamber propagated to an unmatched one, an approval log nobody analysed before setting limits, a change made outside management of change — all cost scrapped material and a suspended programme, and all are cheaper to prevent than to explain.
  5. Phase 5 is bounded by the customer, not by the model. On an automotive- or medical-qualified line, changing the way a setpoint is chosen can trigger a process change notification and a requalification, so copy-exact release is a commercial and quality decision that engineering can prepare for but cannot make alone.

Abbreviations used on this page

MES
Manufacturing execution system — the fab's lot and route system of record
APC
Advanced process control — the layer that adjusts recipe setpoints
R2R
Run-to-run control, typically an EWMA controller on a setpoint
FDC
Fault detection and classification, running on tool trace data
VM
Virtual metrology — a predicted measurement standing in for a physical one
SPC
Statistical process control — the control charts the fab already runs
OCAP
Out-of-control action plan — the scripted response to an SPC alarm
MOC
Management of change — the fab's controlled change procedure
ECN
Engineering change notice — the document an MOC issues
RMS
Recipe management system — where qualified recipes and limits live
PCN
Process change notification — the notice a customer is owed before a qualified process changes
CMP
Chemical mechanical planarisation — the worked example used throughout this page

Free · 8 questions · ~3 minutes

Score how your programme is phased

Eight questions, one at a time, about three minutes. They score sequencing discipline rather than ambition: whether your phase transitions produce artefacts, whether AI change runs inside the fab's own change control, whether decisions advance one at a time or accumulate in shadow, and whether the previous control state can actually be restored. The result names the phase you are genuinely in and the gate that is holding you there.

0 of 8 answered

Question 1 of 8Gate evidence

When a decision last moved from one phase to the next, what marked the transition?

A phase that ended without an artefact did not end. The artefact is what a quality reviewer, a customer or your successor will ask for.

How the score maps to a stage
  • 04 — Stage 1, Off-tool. Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
  • 510 — Stage 2, Shadow. Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
  • 1116 — Stage 3, Engineer-approved. Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
  • 1721 — Stage 4, Closed loop in limits. Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
  • 2224 — Stage 5, Copy-exact release. Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.

What fab AI transformation phases are — and why a date never closes one

A definition, the five qualification states, and the round trip one CMP lot's data makes at each of them.

Fab AI transformation phases are the qualification states an AI decision occupies inside a wafer fab, and a date never closes one because a phase is left by evidence rather than by elapsed time. The five states are off-tool, where the analysis touches nothing; shadow, where the model runs live beside the incumbent controller but binds nothing; engineer-approved, where it writes a proposal a named engineer signs off per lot; closed loop in limits, where it acts unattended inside a versioned envelope on a qualified toolset; and copy-exact release, where the whole assembly propagates under matching evidence and revalidation.

The important structural point is that a fab did not need AI to invent this ladder. It already runs one. Tool and process qualification, marathon runs, split lots, statistical process control with a written out-of-control action plan, management of change issuing engineering change notices, copy-exact propagation between sites, and process change notification to externally qualified customers — that machinery exists, is understood by every module engineer, and is the reason a 1,500-step process holds together at all. An AI programme that builds a parallel governance structure alongside it has doubled the paperwork and halved the authority. The programmes that move build their phases inside it.

This is why phase language borrowed from software — pilot, MVP, rollout, scale — sits so badly in a fab. Those words describe how much of the organisation is exposed to a change. Fab phases describe how much authority a decision has been granted over material that cannot be un-processed. A pilot in a fab is not a small rollout; it is a qualification state with a defined evidence requirement, and the statistical part of that requirement is set out in the NIST/SEMATECH e-Handbook's process-control chapters (opens in a new tab), which most fabs already treat as the reference for their control charting.

Authority granted against evidence accumulated

The curve is deliberately not smooth. Nothing is granted through off-tool and shadow no matter how much analysis accumulates, because neither phase produces the kind of evidence a gate accepts. The step happens at engineer-approved, when the decision first enters the controlled process, and the slope after it is set by how fast the approval log accumulates rather than by model quality. Illustrative shape, consistent with the published deployment reports cited on this page.

Authority granted over the controlled process by stage

  • Stage 1 · Off-tool — 22% of operators. Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
  • Stage 2 · Shadow — 34% of operators. Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
  • Stage 3 · Engineer-approved — 27% of operators. Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
  • Stage 4 · Closed loop in limits — 13% of operators. Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
  • Stage 5 · Copy-exact release — 4% of operators. Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.

Curve shape: logistic, plotted from the stage data above. Distribution: Illustrative; shape consistent with GlobalFoundries' and Flexciton's published deployment accounts.

One CMP lot's data round trip, phase by phase

The same physical loop — trace out, decision back — drawn at each phase. What changes between lanes is not the model but where the arrow is allowed to terminate: at a scoreboard, at an engineer's approval, or at the controller itself. Most fab decisions are in the top lane.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • At phases 1–2, trace data leaves the FDC historian as an extract, is joined to metrology by hand, and produces a model that scores lots on a scoreboard nothing consumes. No engineering change notice exists because nothing in the controlled process has changed, so the run generates accuracy history rather than gate evidence — and accuracy history is exactly what a qualification review will not accept.
  • At phase 3, trace and context stream together — chamber, pad life, conditioner state, lot genealogy — into a frozen model version whose output is written as a setpoint proposal into the run-to-run controller's queue. The process engineer approves or overrides each one with a mandatory reason code, and post-CMP metrology feeds back. The approval log produced here is the raw material for the next gate.
  • At phases 4–5, a versioned envelope derived from that log lets the loop act unattended on named chambers and products. Anything outside the envelope holds the lot, escalates and reverts to the incumbent controller rather than being clamped and continued. Propagation to another chamber or site is a separate qualification with its own matching evidence, not a configuration copy.
Step-by-step insights
The trace extract — why the context join, not the model, caps phase 1
A hand-built join between trace and metrology encodes one engineer's assumptions about which timestamp to trust, how to attribute a lot to a chamber on a multi-platen tool, and what to do with rework. None of it is written down, so a second study is not comparable to the first even when both are correct. The fix that closes gate 1 is unglamorous: carry chamber, consumable state, PM cycle and lot genealogy as owned, lineage-tracked columns, and let the model be whatever it is.
The shadow scoreboard — a comfortable dead end with a real cost
Shadow is safe, cheap and survives every budget review, which is why decisions accumulate there. The cost is not the compute; it is that process engineers learn AI output is decorative and module owners learn the programme changes nothing. Three years of that is harder to reverse than never having started, because the next proposal is heard against the memory of the last one. The only defence is a written exit criterion agreed at the moment shadow starts, expressed in operating regimes rather than accuracy — and a frozen model version, because retraining mid-window silently restarts the window.
Streamed context — what has to travel with the trace
For a CMP decision the trace alone is nearly useless: removal rate depends on pad age, conditioner disc wear, slurry batch, platen and incoming thickness, and a prediction that cannot see those is fitting noise it will later be blamed for. The context set is decided by the process, not by the data platform, which is why the module engineer has to specify it. This is also the point where an honest fab discovers which of those facts exists only on a whiteboard in the sub-fab, and fixing that is part of gate 1, not a later concern.
The setpoint proposal — write into the queue, not onto a screen
The single highest-leverage design decision at phase 3 is where the proposal appears. Written into the run-to-run controller's queue beside the incumbent value, it is on the path the engineer already walks and the default action becomes the informed one. Displayed on a separate dashboard, acting on it is a voluntary extra step, and voluntary steps are the first thing dropped during a ramp or a qualification crunch — which is precisely when the decision is worth most. This is why write-path work usually deserves the quarter that teams want to spend on accuracy.
The approval log — the instrument, not the paperwork
Every accept and override, with its reason code and the context at the time, is the dataset from which the next gate's envelope is derived. It shows where the model and its engineers disagree and, crucially, in which regimes — first lots after a pad change, a particular product layer, a specific chamber. A log analysed quarterly turns into limits with evidence behind them. A log collected and never read leaves a fab with a full database, no envelope, and the choice between staying at phase 3 forever or guessing.
The envelope boundary — hold and escalate, never clamp and continue
The behaviour at the edge of the envelope is what a quality reviewer will examine, because it is where the design's honesty shows. Clamping the adjustment to the limit and letting the lot proceed hides the fact that the world has moved outside the qualified region; holding the lot and escalating surfaces it while the material is still recoverable. Trending the breach rate then gives the fab a leading indicator with real meaning: a rise says the process has drifted out of the envelope's validity and the envelope needs review before an excursion forces one.

The five phases in detail

For each phase: what it looks like on the floor, the signals a reviewer can check in an afternoon, the anti-pattern that traps fabs there, and what leaving costs in team terms.

Each phase below is written for a process or automation engineer rather than for a buyer. The hallmarks describe observable conditions on the floor, the diagnostic signals are checks you can run against your own document system and logs this week, and the anti-pattern is the specific mistake most often made trying to leave that phase.

Select a phase

Every phase's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Off-tool

22% of operators sit here

Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.

Off-tool is where nearly every fab AI idea legitimately starts, and where a surprising number of them are still sitting three years later. The work itself is often excellent: a process engineer who has lived with a chamber for a decade builds a model of removal rate against pad life, or of etch depth against chamber-wall condition, and it explains variation the existing control scheme does not. What is missing is not insight. What is missing is a route from that insight into a lot's history.

The diagnostic is the data path. At this phase the authoritative version of last quarter's chamber behaviour is an extract someone pulled with a filter they chose, on a day they remember approximately, joined to metrology by a lot ID and a timestamp that may or may not be in the same time zone. Two engineers asked the same question will produce two defensible and different answers, because the context — which chamber, which pad, which conditioner disc, which product layer, which preventive-maintenance cycle — was reconstructed by hand each time rather than carried by the data.

This is a cheap phase to leave and an expensive one to occupy. The cost is not the models that go nowhere; it is that nothing accumulates. The tenth study costs what the first one did, because the context join is rebuilt from scratch every time and none of the runs is comparable to any other. The gate out of off-tool is therefore not about modelling at all — it is a data qualification exercise, and it is the only gate on this ladder that a data team can close largely on its own.

In practice

The removal-rate study that runs twice a year

An oxide CMP module engineer is asked, roughly every six months, why post-polish thickness range widens toward the end of pad life. Each time, an engineer exports six weeks of trace from the FDC historian, joins it by hand to post-CMP metrology and to the pad-change log kept in a maintenance spreadsheet, and produces a well-argued deck showing the drift. The deck is correct. Nothing in the recipe, the R2R controller or the pad-change interval changes, and six months later the same export is pulled again — with a slightly different filter, so the two studies cannot be compared.

What it looks like

  • Trace and metrology data arrives as manual extracts from the FDC historian or a data-warehouse query
  • The model runs on an engineer's workstation or a shared analytics server, on demand
  • No management-of-change record exists, because nothing in the controlled process has changed
  • Results are presented at a yield or module meeting and then stop moving

Diagnostic signals you can check this week

  • Ask where the model's training data came from. If the answer is a filename or a saved query, you are in this phase
  • Ask whether chamber, pad life, conditioner state and PM cycle are columns in the dataset or facts someone remembered
  • Check whether any AI output has ever been referenced in an OCAP, an ECN or a qualification report
  • Ask two engineers for the same statistic — for example mean removal rate on one chamber last quarter — and compare the two numbers

Anti-pattern · Building the fab data lake before qualifying one dataset

The instinctive response to a hand-built context join is a platform programme: ingest every tool, every parameter, every trace channel, and sort out meaning later. It is the most reliable way to spend eighteen months without closing a single gate, because the requirements are being guessed rather than observed, and because a lake with no qualified dataset in it is indistinguishable from an expensive export. Qualify one decision's data end to end — one module, one toolset, one metrology step, with lineage and owners — and let the platform's real shape emerge from the second and third of those.

What holds you here

No qualified, context-complete dataset exists, so every study rebuilds the world by hand and no two studies are comparable.

Highest-leverage next move

Close the data qualification gate on one decision: one toolset, one metrology step, with chamber, consumable state, PM cycle and lot genealogy carried as data rather than remembered.

Cost of leaving

Effort
2–4 months
Team
One data engineer, one process engineer at roughly a day a week
Risk
Low — nothing in the controlled process is touched, so nothing can be scrapped
To next stage
2–4 months

If this is you, the next step is

A short engagement: pick the decision, fix the context join, produce the data qualification report.

Qualify one dataset end to end

Stage 2

Shadow

34% of operators sit here

Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.

Shadow is the most comfortable phase on the ladder and therefore the most dangerous. Running a predictor beside the existing controller is cheap, technically satisfying and entirely safe: the model sees production reality, its errors are visible, and no wafer is at risk. Every one of those properties is a reason to stay. Fabs routinely keep decisions in shadow for two or three years, and the programme reports steadily improving accuracy the whole time without a single gate closing.

The structural reason is that shadow answers a question nobody needs answered. A shadow run demonstrates that the model would have been right; the gate to the next phase requires evidence that it will be right through the regimes the process actually visits — the end of a pad's life, the first lots after a chamber wet clean, a product mix shift, the quarter a second-source consumable is introduced. That is a marathon question, and it is answered by designing the shadow window to cover those events deliberately rather than by leaving the model running and hoping the events show up.

There is also an organisational cost to a long shadow. Process engineers learn that the AI output is decorative, module owners learn that the programme does not change anything, and the next proposal is funded against that memory. A fab that has kept six decisions in shadow for three years is harder to move than a fab that has never tried, because the antibodies are established and the phrase people use is not sceptical, it is bored.

In practice

The virtual metrology screen nobody has ever acted on

A 200mm fab stood up a virtual metrology model for post-etch critical dimension, running on chamber trace after every lot. On the shadow scoreboard it tracks the measured CD closely on the sampled lots, and it has been doing so since 2024. It is displayed on a screen in the module office. Sampling has not been reduced by a single wafer, because reducing sampling requires a qualification argument and a change to the sampling plan, and no one has ever written one. The model has been correct for two years in a way that has cost the fab metrology time rather than saving it.

What it looks like

  • The model scores every lot in near-real time and its output is stored, not applied
  • A comparison against the incumbent EWMA controller or dispatch rule is visible to engineers
  • No recipe, setpoint, dispatch decision or hold is affected by the output
  • Nobody has opened a management-of-change record, because formally nothing has changed

Diagnostic signals you can check this week

  • Ask how long the decision has been in shadow. Anything past nine months without a written gate plan is a stall, not a study
  • Check whether the shadow window deliberately spans a PM cycle, a wet clean and a consumable change, or whether it is simply elapsed time
  • Ask what the exit criterion is, in a sentence. If the answer contains an accuracy number and no operating regime, the gate is not defined
  • Look for a frozen model version. If the model has been retrained during the shadow window, the window has been restarted and nobody noticed

Anti-pattern · Improving accuracy to earn the right to act

When shadow does not convert, the reflex is to improve the model, on the theory that authority follows precision. It rarely does. What gates ask for is evidence of behaviour under known regimes, a frozen version, a documented failure mode and a defined response when the prediction is wrong — none of which is an accuracy problem. A predictor that is moderately accurate and demonstrably well behaved across pad life and post-clean lots will pass a gate that a more accurate one with an undocumented failure mode will not. Spend the next quarter designing the marathon and writing the OCAP branch, then revisit accuracy when you can price an accuracy point in metrology hours or scrapped wafers.

What holds you here

There is no defined exit criterion, so the shadow run continues indefinitely and produces accuracy history instead of gate evidence.

Highest-leverage next move

Freeze the model version, define the marathon window by the regimes it must cover, and write the prediction qualification report that the module owner will sign.

Cost of leaving

Effort
3–6 months
Team
One ML engineer, one process engineer as owner, module engineering time for the marathon review
Risk
Low to medium — the risk is elapsed time and credibility, not material
To next stage
3–6 months

If this is you, the next step is

We define the regimes the shadow window must cover and the evidence pack that comes out of it.

Design the marathon that closes the gate

Stage 3

Engineer-approved

27% of operators sit here

Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.

Engineer-approved is the first phase in which the programme survives the person who built it. There is a document trail, a named owner, a defined fallback and a logged human decision on every action, which together mean the capability keeps working when the engineer moves module. It is also the first phase in which the fab gets something back: metrology time, cycle time, scrapped material or engineer attention, in units the module already reports.

The character of the work changes sharply here. Off-tool and shadow problems are analytical; engineer-approved problems are procedural. Which field does the proposal write to. What does the engineer see when they approve. What is the reason code taxonomy, and who maintains it. What does the OCAP say when the model and the SPC chart disagree. What happens on a shift where the data feed stops. These have well-established answers in a fab's own control discipline — the NIST/SEMATECH handbook's process-control chapters are still the clearest public statement of the statistical part — and importing them wholesale is far faster than rediscovering them under a different vocabulary.

The trap that emerges is the unread approval log. The log is the single most valuable artefact this phase produces: it is the dataset from which the next gate's control envelope is derived, and it is the only place the fab can see where the model and its engineers disagree and why. Fabs that treat approval as a formality rather than as instrumentation reach the end of a year with a full log, no analysis, and no basis for setting limits — which means they either stay here indefinitely or set an envelope by guesswork.

In practice

The setpoint proposal on the APC screen

An oxide CMP fleet runs a pad-life-aware removal-rate prediction that proposes a polish-time setpoint into the R2R controller's queue. The process engineer sees the proposal, the incumbent EWMA value and the delta, and approves or overrides with one of six reason codes. In the first two months the override rate on freshly conditioned pads runs far higher than on mid-life pads — a pattern nobody predicted, visible only because the reason codes were mandatory. That pattern later becomes the first exclusion in the control envelope: the loop does not act on the first lots after a pad change.

What it looks like

  • Output is written into the APC, RMS or MES queue the engineer already works in, not a separate screen
  • Every accept and override is logged with a reason code and the context at the time
  • An ECN exists, the OCAP has a branch for the new signal, and the incumbent controller is one switch away
  • A named module owner is accountable for the decision's behaviour, not the data team

Diagnostic signals you can check this week

  • Ask to see the approval log. If reason codes are optional or a free-text box, the phase is running but not instrumented
  • Check whether the incumbent controller can be restored by an engineer on shift, without a ticket to another team
  • Read the OCAP. If it does not mention the model at all, the model is outside the fab's control discipline
  • Ask who is paged when the prediction feed stops. If the answer is a data team with no on-shift presence, the fallback is theoretical

Anti-pattern · Treating approval as a formality on the way to automation

Because approval feels like a temporary inconvenience, teams optimise it away: a single accept button, no reason codes, no context capture, engineers clicking through a queue at shift change. The phase then generates no evidence, and the gate to closed loop has to be argued from model accuracy — which is exactly the argument that will not survive a quality review. The approval step is not a concession to caution. It is the instrument that produces the envelope, and a fab that rubber-stamps for a year has to spend another year re-earning what it threw away.

What holds you here

The approval log is collected but never analysed, so there is no evidential basis for the limits a closed loop would run inside.

Highest-leverage next move

Analyse the approval log by context — chamber, consumable state, product, shift — and derive an explicit control envelope with named exclusions before proposing any unattended action.

Cost of leaving

Effort
6–12 months
Team
Integration engineer, ML engineer, module process owner, an automation or MES contact for the write path
Risk
Medium — the first write into a controlled system needs an ECN, an OCAP revision and a drilled revert
To next stage
6–12 months

If this is you, the next step is

Reason-code taxonomy, context capture and the analysis that turns the log into a control envelope.

Instrument the approval log properly

Stage 4

Closed loop in limits

13% of operators sit here

Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.

Closed loop in limits sounds like the destination and is better understood as a narrow, carefully bounded permission. It is not an autonomous fab and it is not an autonomous module: it is one decision, on an enumerated set of chambers, for an enumerated set of products, within a stated adjustment magnitude, with a written response for everything else. Fabs that describe this phase in broader terms than that are usually describing an intention rather than a qualified state.

The engineering here is mostly finished by the time a fab arrives. What is not finished is the envelope, and the envelope is the deliverable. It answers four questions in writing — where may this act, on what, by how much, and what happens at the boundary — and it is versioned and reviewed like a recipe because that is exactly what it is. Published research is instructive about how much of this is still open: reinforcement-learning controllers have been compared favourably with EWMA and general harmonic rule controllers, and demonstrated on a nonlinear CMP process without an explicit model, but that work is numerical and simulation-based. It tells you the control idea is sound; it does not tell you your chamber is in the envelope.

The constraint that binds at this phase is not technical confidence but material consequence. A closed loop on a CMP setpoint that drifts in the wrong direction removes material that cannot be put back. That is why the boundary behaviour — hold and escalate, never clamp and continue — matters more than the loop's average performance, and why the excursion review is a standing item rather than an incident response. Fabs that get this right treat the envelope breach rate the way they treat an SPC chart: as a signal about the world, not as a nuisance.

In practice

The envelope that names three chambers and excludes two

A fab qualifies a closed-loop removal-rate adjustment on five oxide CMP chambers. Chamber matching evidence supports three of them; the remaining two show a consistent offset traced to a different platen supplier. The envelope names the three, excludes the two by serial number, caps the adjustment at a stated fraction of the recipe's polish time, and excludes the first lots after any pad change. The two excluded chambers stay at engineer-approved. That is not a failure of the programme — it is the phase working correctly, and the document is what makes it defensible a year later.

What it looks like

  • The envelope is an explicit, versioned document: which chambers, which products, which consumable states, which magnitude of adjustment
  • Excursions outside the envelope hold the lot and escalate rather than being clamped silently
  • The revert to the incumbent controller has been exercised deliberately, not just documented
  • Envelope breach rate is monitored as a leading indicator and reviewed on a defined cadence

Diagnostic signals you can check this week

  • Ask to see the envelope as a document with a version number. If it is a configuration screen, it is not a controlled artefact
  • Check when the revert to the incumbent controller was last exercised on a live shift, not simulated
  • Ask what happens at the boundary. If the answer is that the adjustment is capped and the lot continues, the escalation path does not exist
  • Check whether envelope breach rate is trended and reviewed, or only looked at after an excursion

Anti-pattern · Widening the envelope because nothing has gone wrong

Six quiet months read as evidence that the limits were too conservative, and the envelope grows — another chamber, another product, a larger adjustment — usually in a configuration change rather than a controlled document revision. The new scope has no approval-log history behind it, so the first genuine excursion happens in a region the evidence never covered. The predictable outcome is that the loop is switched off entirely and the fab regresses two phases from one event. Widen the envelope the way you earned it: from the approval log for that chamber and that product, through management of change.

What holds you here

The envelope holds for the chambers and products it was qualified on, and every extension needs its own evidence — so scope grows slowly and pressure to widen it informally is constant.

Highest-leverage next move

Produce the propagation package: chamber-matching evidence, a per-site revalidation plan, the envelope as a controlled document, and the customer notification assessment.

Cost of leaving

Effort
9–18 months from first approval
Team
Module process owner, equipment engineering, quality, plus the platform team
Risk
Higher — the consequence of a wrong action is material, and the evidence burden is a quality matter
To next stage
12–24 months

If this is you, the next step is

We stress-test the limits, the boundary behaviour and the revert against a real excursion scenario.

Review a control envelope before it goes live

Stage 5

Copy-exact release

4% of operators sit here

Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.

Copy-exact release borrows its discipline from a practice the industry has run for decades: propagate a proven process by reproducing its conditions exactly rather than re-optimising at each site, and treat deviation as something to be justified rather than assumed harmless. Intel's Copy Exactly! methodology is the best-known articulation of it, and the logic transfers directly to a learned controller — with one important amendment. A recipe copied exactly onto matched equipment behaves the same way. A model copied exactly onto a differently-behaving chamber does not, because the model encodes the statistical fingerprint of the equipment it learned on.

That amendment is an active research problem rather than a solved one. Work on domain adaptation and equipment matching exists precisely because deep-learning-based process monitoring does not transfer cleanly between non-identical tools, and proposes methods to align them. Read that literature as a warning about sequencing: propagation is a qualification exercise per target, and a fab that budgets it as a deployment task will discover the difference on the first mismatched chamber.

The other binding constraint at this phase is commercial. On a line qualified for automotive or medical parts, changing how a setpoint is chosen can constitute a process change, which can oblige the fab to issue a process change notification and, depending on the customer's agreement, to requalify. That makes copy-exact release a decision the quality organisation and the customer-facing account team make jointly with engineering. It is entirely normal for a technically ready loop to be held at closed loop in limits for a year on one line while it runs released on another — and a programme that has not modelled that is not planning, it is hoping.

In practice

The release that shipped to two of four fabs

A qualified etch-depth loop is packaged as a controlled release: model version, envelope, OCAP branch, matching criteria and revalidation protocol. Two fabs take it inside a quarter — same tool generation, comparable chamber population, no externally qualified products on the affected layers. The third fab's chamber population fails the matching criteria and enters its own engineer-approved phase to build local evidence. The fourth runs automotive-qualified product on that layer and holds pending a customer notification assessment. One release, four correct and different outcomes.

What it looks like

  • The model, its envelope and its OCAP branch ship together as one versioned, controlled release
  • Propagation to a new chamber or site requires matching evidence and site revalidation, not a copy of a configuration file
  • Change notification obligations to externally qualified customers are assessed before, not after, release
  • Model and envelope versions are reconstructable for any lot, months later, for a customer or quality audit

Diagnostic signals you can check this week

  • Ask whether the model version and envelope version that governed a specific lot three months ago can be recovered from the record
  • Check whether propagation to a new chamber requires documented matching evidence or only an engineer's judgement
  • Ask who assesses customer change-notification obligations, and whether they are consulted before release or after
  • Check whether a released loop has ever been withdrawn, and whether the withdrawal path is written down

Anti-pattern · Treating the envelope as configuration once the release exists

The release is version-controlled, and then the limits are tuned in a settings screen with no revision history and no signature. The system works until someone has to explain a decision made eight months earlier, at which point neither the model version nor the limit that produced it can be reconstructed, and the fab cannot demonstrate to a customer that the qualified process was in force. Version the envelope with the model, require a signature on every change, and keep the trail for as long as the product's qualification requires — which on automotive parts is a long time.

What holds you here

Propagation is gated by equipment matching and by external qualification obligations, so the constraint becomes evidence and change notification rather than engineering.

Highest-leverage next move

Treat the envelope as part of the controlled release, sign every change, and keep model-to-lot traceability for the full qualification life of the affected products.

Cost of leaving

Effort
Continuous
Team
Platform and module teams plus a standing change board with quality and customer-quality representation
Risk
Concentrated — low frequency, high consequence, and increasingly a customer-contractual matter

If this is you, the next step is

We reconstruct a specific lot's governing model and envelope version from the record, the way an auditor would.

Audit a released loop end to end

Where wafer fabs actually sit across the five phases

The distribution is heavily weighted toward shadow, and the shadow-to-approved step is the largest single loss on the ladder.

Most fab AI decisions are in shadow. The pattern that shows up repeatedly is a fab with several models running live against production data, none of them affecting a lot, and a programme narrative built on accuracy improvements — with a much smaller number of decisions that have actually entered the controlled process, and a very small number running unattended inside a written envelope.

Illustrative distribution of fab AI decisions across the five phases

Illustrative distribution, synthesised from the published deployment accounts and research cited on this page — not a survey. The shape is the argument: shadow is the mode, and the step down to engineer-approved is the largest single loss on the ladder.

Share of fab AI decisions

  • 22% — 1 · Off-tool
  • 34% — 2 · Shadow (the stall)
  • 27% — 3 · Engineer-approved
  • 13% — 4 · Closed loop in limits
  • 4% — 5 · Copy-exact release

Source: Illustrative; synthesised from the published sources listed on this page

The published record is consistent with that shape in an instructive way: what operators announce is nearly always a portfolio, and what they quantify is nearly always the far end of it. GlobalFoundries reports having deployed over 60 smart manufacturing solutions since 2020 (opens in a new tab) across its fabs, with its Singapore 300mm site named to the World Economic Forum's Global Lighthouse Network in September 2025 — five years of accumulation, not a programme quarter. Its own manufacturing pages describe custom-built engines that classify wafer patterns automatically and speed troubleshooting by up to 10× (opens in a new tab), which is a phase-3-and-beyond claim about engineer time, not an accuracy claim.

That distinction between simulated and qualified runs through the whole research literature and is worth internalising before setting a programme's expectations. Long-horizon reinforcement-learning control for fabs is evaluated against high-fidelity simulations of industry-real scenarios (opens in a new tab); photolithography cluster-tool scheduling algorithms are described by their authors as showing promise for real-world implementation (opens in a new tab); wafer-map defect classifiers report F1 scores above 98% on public benchmark datasets (opens in a new tab). All three are good work. None of them is a closed gate in your fab, and treating a published result as though it were is how a programme ends up promising phase 4 outcomes on a phase 2 evidence base.

The gate register: four gates, and the artefact that closes each

The centre of this page. One row per gate — what it opens, what has to be true to enter it, the document that closes it, who signs, and how long it honestly takes.

Four gates separate the five phases, and each is closed by an artefact the fab already knows how to issue. That is the whole design principle: a gate that requires a new kind of document will be argued about, deferred and eventually skipped, while a gate that requires a qualification report, an ECN, an OCAP revision or a propagation package slots into machinery that already has signatories, a review cadence and a filing location. The register below is the page's working instrument — everything after it is either an application of a row or a consequence of skipping one.

GatePhase it opensEntry conditionExit evidence — the artefactSigned byTypical elapsed
1 · Data qualification2 · ShadowOne decision named, with its metrology step and its toolset. The context set — chamber, consumable state, PM cycle, lot genealogy, product layer — specified by the module engineer, not by the data team.A data qualification report: lineage for every field, owner per field, known gaps and their treatment, and a reproducible dataset that two engineers can query to the same answer.Module process owner and the data owner jointly2–4 months
2 · Prediction qualification3 · Engineer-approvedA frozen model version and a marathon window specified by regime — a full pad or consumable life, a wet clean, a PM cycle, a product changeover — rather than by elapsed weeks.A prediction qualification report: performance by regime not in aggregate, the documented failure modes, the fallback behaviour, and the proposed OCAP branch and reason-code taxonomy.Module engineering, with equipment engineering on the failure-mode section3–6 months
3 · Action qualification4 · Closed loop in limitsAt least two quarters of reason-coded approval log covering the regimes in scope, analysed by context rather than in aggregate.A control envelope as a versioned document — named chambers, named products, named exclusions, maximum adjustment, boundary behaviour — plus a drilled revert and the ECN that puts it into the controlled process.Module owner, equipment engineering and quality6–12 months
4 · Propagation qualification5 · Copy-exact releaseA stable envelope on the source toolset, and a candidate target with a documented matching hypothesis.A propagation package: chamber-matching evidence, the per-site revalidation protocol, the model and envelope as one controlled release, and a completed customer change-notification assessment.Change board, including quality and customer-quality representation3–9 months per target
The gate register for a wafer fab AI programme. Elapsed times are typical for one decision on one toolset with a named owner; they are dominated by evidence accumulation and review cadence, not by engineering. Extending scope resets the clock for the new scope, not for the whole decision.

Two properties of the register do most of the work. First, the exit evidence for every gate is a document with a signatory who is not the person who built the model — which is what stops a programme grading its own homework. Second, the entry condition for each gate is the exit evidence of the last one, so the register is a chain rather than a menu: there is no legitimate route from a shadow run to a closed loop, because the envelope in gate 3 is derived from the approval log that only exists once gate 2 has been closed and phase 3 has actually been run.

How to write a gate that actually closes

  1. Name the artefact before the work starts

    The single most effective intervention available is to write the title and the signatory of the closing document on the day the phase begins. It converts an open-ended study into a piece of work with a definition of done, and it surfaces immediately whether anyone senior is willing to put their name on the result — which is the real question and is much cheaper to answer at the start.

  2. Specify the window by regime, never by duration

    "Twelve weeks" samples whatever happened in twelve weeks. "A full pad life, one wet clean, one PM cycle and one product changeover, with the model version frozen" samples the conditions the process actually visits. The second takes about as long and produces evidence the first cannot, because the failures worth knowing about live in the transitions.

  3. Make the signatory someone with material at risk

    A gate signed by the programme sponsor tests enthusiasm. A gate signed by the module owner whose scrap number moves tests the evidence. If the module owner will not sign, the interesting information is why — and it is almost always a specific, addressable objection about a regime the evidence did not cover.

  4. Route the change through management of change, always

    Every AI-driven change to a setpoint, a sampling plan or a dispatch rule goes through the same MOC and ECN route as any other process change, and revises the OCAP where it introduces a new signal. This is not bureaucratic tribute. It is what makes the change visible to the excursion review that will eventually look at it, and what makes the programme legible to a customer audit.

  5. Keep one gate open at a time per decision

    Fabs that run several decisions through several gates simultaneously discover that review capacity — module engineering, equipment engineering, quality — is the binding constraint, not engineering capacity. One decision moving cleanly through consecutive gates finishes faster than four moving in parallel, and it produces a template the next four inherit.

  • Skipping gate 1 costs you every later comparison

    Without a qualified dataset, the shadow run's performance cannot be attributed to a regime, because the regime was never recorded. The gate 2 report then has nothing to say about behaviour under pad-life or post-clean conditions, and the reviewer's only available question is about aggregate accuracy — which is exactly the question that does not justify granting authority.

  • Skipping gate 2 puts an unqualified prediction in front of an engineer

    The engineer is then doing the qualification themselves, per lot, without the evidence to do it well. Override rates run high, trust erodes fast, and the approval log fills with noise instead of signal — which quietly disqualifies gate 3 as well, since the envelope has to come from that log.

  • Skipping gate 3 is the one that scraps material

    An envelope set from model confidence rather than from approval history has no evidence about the regimes where engineers actually disagreed with the model. The first excursion lands in one of them. The typical outcome is not a tuned envelope but a switched-off loop and a two-phase regression from a single event.

  • Skipping gate 4 breaks on the first unmatched chamber

    A model carries the statistical fingerprint of the equipment it learned on, which is why equipment matching and domain adaptation are an open research problem rather than a deployment detail (opens in a new tab). Copying an envelope to a chamber with a different platen supplier, a different age or a different clean history is the fastest way to turn a good loop into a scrap event on a line that was working yesterday.

Where AI lands in a wafer fab, and the phase each decision can realistically reach

Seven process areas, the decisions worth wiring, the system that owns each one, and the gate that limits how far it can go.

AI value in a wafer fab concentrates in seven operating areas, and each one has a different ceiling — set not by model difficulty but by what the decision touches. A decision is a good candidate for early phases when three things are true: the system of record is inside the fab's own change control, the feedback loop closes inside a shift or two rather than at final test, and the metric it moves is one a module owner already reports. The map below is how a first and second decision get chosen.

Process areaHigh-value decisionsSystem of recordMetric it movesRealistic ceilingLimiting gate
LithoOverlay and focus feed-forward, reticle and job sequencing on the bottleneck toolsetScanner job control, APC, MES dispatcherOverlay residual, rework rate, litho toolset movesPhase 4 on overlay; phase 3 on sequencingGate 3 — scanner-adjacent control is tightly governed and the envelope argument is hard
Etch and depositionEndpoint and depth prediction, chamber-condition-aware setpoint adjustment, seasoning cadenceAPC / R2R, RMS, FDCDepth or thickness range, chamber-to-chamber spread, requal frequencyPhase 4 per matched fleetGate 4 — chamber matching, not the model
CMP and planarisationRemoval-rate prediction over pad life, polish-time setpoint, conditioner and pad-change timingAPC / R2R, maintenance systemPost-CMP thickness range, dishing and erosion, consumable cost per waferPhase 4Gate 3 — the approval log has to cover post-pad-change regimes
Diffusion, thermal and implantBatch composition, queue-time risk scoring, furnace profile compensationMES dispatcher, APCQueue-time violations, batch utilisation, uniformityPhase 3, phase 4 on queue-time holdsGate 2 — queue-time effects are slow and marathon windows are long
Metrology and defect inspectionVirtual metrology for sampling reduction, wafer-map classification, review-image triageMetrology host, YMS, MES sampling planMetrology hours per lot, time to disposition, escape ratePhase 4 on sampling; phase 3 on classificationGate 2 — reducing sampling requires a qualification argument, not a model
WIP flow and dispatchLot prioritisation, toolset scheduling, AMHS routing, bottleneck protectionMES dispatcher, scheduler, AMHS controllerCycle time per layer, x-factor, on-time delivery to commitPhase 4 fab-wide, demonstrated publiclyGate 3 — the envelope is the set of rules the scheduler may override
Sub-fab and facilitiesAbatement and pump health prediction, chiller and UPW load management, exhaust balancingFacilities SCADA / BMSUnplanned tool downtime from facilities, energy per waferPhase 4, sometimes faster than the fab floorGate 1 — facilities historians rarely carry fab-side context
The wafer fab decision map. "Realistic ceiling" is the phase a decision can typically reach with current practice; "limiting gate" is the one that actually binds it, which is rarely the first one a programme worries about.

The two rows most often misread are metrology and WIP flow. Virtual metrology looks like the easiest phase-4 decision in the fab because the model is comparatively tractable — published work on chemical-vapour-deposition virtual metrology in a mass-produced process reports prediction accuracy around 0.7 with roughly 70% data availability (opens in a new tab), and argues that this supports reduced physical metrology frequency. But the value only arrives when the sampling plan changes, and changing a sampling plan is a qualification argument about escape risk, not a modelling result. Fabs that skip that argument end up with a permanently accurate model and a permanently unchanged sampling plan — the phase-2 stall in its purest form.

WIP flow runs the other way. Scheduling and dispatch look institutionally harder because they touch every module, yet they are the area with the clearest public evidence of a fab reaching phase 4 across an entire site, and the practical literature on machine learning for fab scheduling (opens in a new tab) has been discussing production deployments for years. The reason is structural: a dispatch decision is reversible on the next lot, it does not remove material, and the fallback is a rule set the fab already runs. Reversibility, not tractability, is what sets how far a decision can travel.

The compliance frame is the other ceiling, and in a foundry it is the harder one. The industry's own standards define the vocabulary these phases live in — SEMI E10 for equipment reliability and availability states, E30 GEM for the equipment communication layer, E133 for the advanced process control framework itself, and E164 for the common metadata that makes tool data interpretable across a fleet (the SEMI standards programme (opens in a new tab) is the reference, though its site is bot-walled to automated fetching). Above those sit the customer-facing obligations: a line producing automotive parts under IATF 16949 carries change-notification duties, and a process change notification can convert a technically finished phase-4 loop into a nine-month commercial conversation. Assess that at gate 3, not at gate 4 — the answer occasionally changes which toolset you should have started on.

Two of the phase ceilings above are set by infrastructure rather than by evidence — whether trace can leave a tool at the rate a control loop needs, and where the compute has to sit to meet the deadline. That is a different question from sequencing and it has its own page: see AI readiness infrastructure for wafer fabs for the interfaces, clock budgets and qualified write paths underneath this map.

Why phases get run out of order — and what it costs

Three structural patterns account for most out-of-order programmes, and none of them is a modelling problem.

Phases get run out of order because the fab's incentives and the programme's incentives point at different artefacts. A programme is funded against a roadmap with dates on it, and dates are met by demonstrating capability; a fab grants authority against evidence, and evidence accumulates at the speed of pad lives and PM cycles. When those two clocks disagree, the programme reaches for the phase that can be shown on the date rather than the phase the evidence supports, and the sequence breaks in one of three recognisable ways.

  • The demonstration jumps the ladder

    A convincing closed-loop demonstration is built on a single chamber for a review, without gate 2 or gate 3 behind it. It works, because a single chamber over a short window is a benign environment. The demonstration then becomes the reference point for what the programme can do, and every later request for a marathon window or an approval log reads internally as engineering caution rather than as the actual gate. The fix is procedural and costs nothing: label the demonstration as a phase-1 artefact in the room where it is shown.

  • Shadow is used as a substitute for a decision

    When nobody wants to own the risk of granting authority, shadow is the answer that offends no one: the model stays live, the programme stays funded, and the decision is deferred indefinitely. This is the most common cause of a multi-year stall and it is not a technical failure at all — it is an unassigned owner. The tell is that nobody can name who would sign the gate 2 report, which is a question worth asking at the start of every shadow run.

  • Propagation is planned as deployment

    Once a loop works on one fleet, the plan says roll it out, and the roll-out is scheduled as an IT activity with a duration per site. Gate 4 is not an IT activity: it is a qualification per target, with matching evidence and revalidation, and its duration is set by the target's own evidence accumulation. Programmes that discover this halfway through a multi-site plan usually lose more credibility than the loop ever earned.

Diagnosing an out-of-order programme

Plot how phases end against how change enters the controlled process. The quadrant names the next investment — and in three of the four cases it is not a better model.

Shelf-qualified

  • Good evidence, no route into the controlled process
  • Reports and marathon data exist; no ECN ever issued
  • Fix: take one closed gate through MOC and let it be the template

Releasable

  • Evidence and change control both in place
  • The only quadrant from which phase 4 is reachable safely
  • Fix: shorten the queue — one decision through consecutive gates

Roadmap theatre

  • Phases are calendar periods and nothing is signed
  • Common where the programme reports to IT rather than to manufacturing
  • Fix: name the artefact and the signatory for one decision

Rubber-stamped

  • MOC records exist; the evidence behind them is a date
  • The most dangerous quadrant — the paperwork is protective in appearance only
  • Fix: audit one recent ECN back to the evidence that justified it
Gate evidence — top: Phases end on signed artefacts, bottom: Phases end on dates
Change-control fit — left: AI change runs beside MOC, right: AI change runs through MOC, ECN and OCAP

Quantifying the benefit of an alternative scheduling approach remains a challenging task.

That sentence, from a vendor writing about its own successful production deployment, is worth more than most fab AI marketing. A running fab cannot hold conditions still for an A/B test: product mix moves, tools go down, demand changes. Which is precisely why the gate register asks for regime-covered evidence and approval-log analysis rather than for a controlled trial — the fab cannot supply the trial, and a phase model that demands one will simply be ignored.

Three phased programmes, read against the register

Publicly reported deployments, each cited to the operator's or vendor's own published material. None is an Atomic Loops engagement.

The clearest public evidence for the gate register is in what operators chose to publish about how long things took. In each case below the interesting detail is not the outcome but the sequence — a bounded trial before enablement, a five-year accumulation before an external designation, or a portfolio described in terms of engineer time rather than model accuracy.

Three deployments read against the four gates

Outcomes as reported by the named organisations themselves; verify figures against the linked source before reusing them, and note where the source itself flags measurement difficulty. Card images are generated industry scenes from our image library — not photographs of these operators' facilities, and not an endorsement.

Generated scene: cleanroom bay with process tools and an overlaid scheduling network diagramSeagate Springtown, with FlexcitonWafer fab, recording heads · Northern Ireland34
Challenge
Wafer flow through the fab was steered by a large and growing population of ad hoc operational rules, applied manually to control WIP. The rules worked, but they made the fab's behaviour dependent on which rules were in force that week and on the people maintaining them.
Approach
An optimisation-based scheduler was trialled over a bounded window — March to May 2022 — before being enabled fab-wide, running 24/7 from June 2022. The published account describes both a toolset-level scheduler and a fab-wide scheduler, and is explicit that a conventional A/B comparison was not possible in a live fab, so the benefit was argued from several analytical angles instead.
Reported outcome
Flexciton reports a large fall in manual intervention after deployment, with ad hoc rule transactions averaging fewer than 150 per week, and states that further analysis suggests substantial throughput and cycle-time improvements — while explicitly noting that quantifying the benefit of an alternative scheduling approach remains a challenging task. The work was also presented as a technical paper at ASMC 2023.
What it shows about the curveThis is gate 3 done properly and said out loud. The exit evidence was a change in operating behaviour — interventions collapsing — rather than a clean causal number, because a running fab cannot supply one. A programme whose gate criteria demand a controlled trial will never close a gate in a real fab.

Flexciton — fab-wide scheduling case study (opens in a new tab)

Generated scene: wafer inspection tool in a cleanroom with a wafer-map defect overlay and gowned engineersGlobalFoundriesGlobal foundry · 300mm and 200mm fabs35
Challenge
Making machine learning a repeatable capability across a multi-site foundry estate rather than a set of per-site projects, where a bespoke build per fab would never amortise and where customer-qualified processes constrain what may change.
Approach
A portfolio approach accumulated over years rather than a single programme: predictive tool-health modelling informing maintenance decisions, custom-built engines classifying wafer patterns automatically, and a proprietary factory control tower monitoring production processes across the global manufacturing network.
Reported outcome
GlobalFoundries reports over 60 smart manufacturing solutions deployed since 2020, states that automatic wafer-pattern classification speeds troubleshooting by up to 10× and reduces wafer scrap, and had its Singapore 300mm fab named to the World Economic Forum's Global Lighthouse Network in September 2025.
What it shows about the curveThe gate-4 signature is a portfolio measured in years and an external designation arriving five years after the accumulation started. Note also what is quantified and what is not: engineer troubleshooting time carries a number, tool uptime and scrap are described qualitatively. That is what honest reporting from a large estate looks like.

GlobalFoundries — Lighthouse announcement and digital manufacturing (opens in a new tab)

Generated scene: memory wafer fab interior with automated handling and process equipmentMicronMemory manufacturer · front-end wafer fabs24
Challenge
A memory process runs to roughly 1,500 steps over months, so the cost of relying on human vigilance to spot flaws and equipment trouble is high and the feedback loop from a mistake to its discovery is long.
Approach
Micron describes applying AI across the process rather than at a single step, with the stated aim of improving accuracy and coverage, and frames the payoff in terms of how quickly new products can be brought up rather than in terms of model metrics.
Reported outcome
Koen de Backer, Micron's vice-president of smart manufacturing and artificial intelligence, is quoted in Micron's own published account saying the company can now launch products twice as fast while improving productivity by 10%.
What it shows about the curveThe choice of headline metric is the lesson. Ramp speed and productivity are outcomes a manufacturing organisation already owns and reports; model accuracy is not. A programme whose phase-4 business case is written in the fab's existing metrics survives a budget review that one written in F1 scores will not.

Micron — Smart sight: how Micron uses AI to enhance yield and quality (opens in a new tab)

The reference architecture, by the phase that first requires it

Six layers, each annotated with the phase that makes it non-optional — and the one layer teams reliably build too early.

Each phase makes a different layer non-optional, and building layers ahead of the phase that needs them is the commonest way a fab programme spends a year without moving. The architecture below is deliberately vendor-neutral: every layer is defined by what it must guarantee at a given gate, not by what product supplies it, because the products differ between a 200mm speciality fab and a 300mm logic fab while the guarantees do not.

Layers required by phase

Read the annotations as thresholds, not as a build order preference: a programme trying to reach phase 4 without the envelope and OCAP layer is running an unqualified closed loop, whatever it is called internally.

  1. Tool and process layer

    Stage 1+

    • Equipment trace and FDCThe signals the decision will actually depend on
    • Metrology and inspectionThe measurement the prediction is answerable to
    • Recipe management (RMS)Where qualified recipes and limits already live
  2. Context and lineage layer

    Stage 1+

    • Lot genealogy from MESRoute, layer, product, rework history
    • Equipment stateChamber, consumable age, PM and clean history
    • Field ownership and lineageOne definition per field, with a named owner
  3. Model layer

    Stage 2+

    • Frozen versioningA marathon window is void if the model changed inside it
    • Regime-aware evaluationPerformance reported by regime, never only in aggregate
    • Drift monitors keyed to fab eventsPM, wet clean, consumable change, mix shift
  4. Proposal and approval layer

    Stage 3+

    • Write path into APC or the dispatcherThe queue the engineer already reads
    • Reason-coded approval logAccept, override and context, captured per lot
    • Incumbent fallbackThe previous controller, one switch away, on shift
  5. Envelope and OCAP layer

    Stage 4+

    • Versioned control envelopeChambers, products, exclusions, maximum adjustment
    • Boundary behaviourHold and escalate, never clamp and continue
    • Breach-rate monitoringTrended as a leading indicator, reviewed on cadence
  6. Release and propagation layer

    Stage 5+

    • Controlled release packageModel, envelope and OCAP branch versioned together
    • Matching and revalidationEvidence per target chamber, fleet or site
    • Change-notification assessmentCustomer obligations resolved before release

Pipeline described

  1. Tool and process layer (stage 1+) — Equipment trace and FDC: The signals the decision will actually depend on; Metrology and inspection: The measurement the prediction is answerable to; Recipe management (RMS): Where qualified recipes and limits already live
  2. Context and lineage layer (stage 1+) — Lot genealogy from MES: Route, layer, product, rework history; Equipment state: Chamber, consumable age, PM and clean history; Field ownership and lineage: One definition per field, with a named owner
  3. Model layer (stage 2+) — Frozen versioning: A marathon window is void if the model changed inside it; Regime-aware evaluation: Performance reported by regime, never only in aggregate; Drift monitors keyed to fab events: PM, wet clean, consumable change, mix shift
  4. Proposal and approval layer (stage 3+) — Write path into APC or the dispatcher: The queue the engineer already reads; Reason-coded approval log: Accept, override and context, captured per lot; Incumbent fallback: The previous controller, one switch away, on shift
  5. Envelope and OCAP layer (stage 4+) — Versioned control envelope: Chambers, products, exclusions, maximum adjustment; Boundary behaviour: Hold and escalate, never clamp and continue; Breach-rate monitoring: Trended as a leading indicator, reviewed on cadence
  6. Release and propagation layer (stage 5+) — Controlled release package: Model, envelope and OCAP branch versioned together; Matching and revalidation: Evidence per target chamber, fleet or site; Change-notification assessment: Customer obligations resolved before release
Step-by-step insights
Context and lineage layer — the layer that decides everything above it
This is where consumable age, PM history and lot genealogy stop being tribal knowledge. It is also where fabs discover that some of the context they need exists only as a maintenance spreadsheet or a whiteboard in the sub-fab, and closing that gap is part of gate 1 rather than a later data-quality project. Every phase above inherits the quality of this layer: a regime-aware evaluation at gate 2 is impossible if the regime was never recorded.
Model layer — freezing is a governance feature, not a limitation
The instinct to keep improving a model during its evaluation window is exactly wrong at gate 2, because it makes the window unusable as evidence. Freeze the version, run the marathon, report by regime, and put improvements in the next version behind the next window. Drift monitoring should likewise be keyed to the events that actually change a fab's input distribution — preventive maintenance, wet cleans, consumable lot changes, product mix shifts — rather than to statistical thresholds alone.
Proposal and approval layer — the fallback is what unlocks the write
The component most often skipped is the incumbent fallback: the previous controller, restorable by an engineer on shift without a ticket to another team. It looks like engineering pessimism and is actually the political key to the write-path approval, because module owners will accept a new decision source they can instantly withdraw. A write-path proposal without a drilled revert sits in a change queue for two quarters; one with it goes through.
Envelope and OCAP layer — the artefact a quality reviewer will read
An envelope that lives in a configuration screen is not a controlled artefact, whatever the access controls are. Write it as a versioned document with named chambers, named products, named exclusions and a stated maximum adjustment, put the boundary behaviour in the OCAP, and require a signature on every revision. This is also the layer that makes an excursion investigation tractable: without a versioned envelope, reconstructing why the system acted as it did six months ago is guesswork.
Release and propagation layer — matching evidence is a deliverable
Propagation fails on equipment differences the model can see and people cannot, which is why domain adaptation and equipment matching remain active research topics rather than solved deployment steps. Budget matching evidence per target as work, define what disqualifies a target in advance, and accept that some chambers will stay at phase 3 indefinitely. A release that names its exclusions is stronger than one that claims universal applicability.

The layer built too early, almost universally, is the model layer's tooling — experiment tracking, feature platforms, serving infrastructure — while the context and lineage layer is still a set of manual joins. The symptom is a fab with an impressive MLOps stack and no qualified dataset in it. The rule that avoids it is simple: build a layer when the gate you are currently trying to close requires it, and not before, because until then you are guessing at requirements that phase 3 will tell you for free.

A 90-day plan: closing gate 2 on one oxide CMP fleet

The shadow-to-engineer-approved transition made concrete on one decision — a pad-life-aware removal-rate prediction proposing the polish-time setpoint. Contains no model development.

Closing one gate takes about 90 days when it is scoped to a single decision on a single fleet, and several years when it is scoped to a module. To make that concrete, the plan below runs gate 2 on a specific and very common wafer fab problem: post-CMP thickness range widening across pad life on an oxide CMP fleet, where the run-to-run controller compensates after the fact rather than anticipating the trend. The prediction usually already exists in shadow at this point, so the quarter contains no model development at all — it is evidence, integration and paperwork.

Gate 2 on one oxide CMP fleet, in one quarter

One fleet, one product family, one named module owner. If any phase needs longer than its window, narrow the scope — fewer chambers, one product — rather than extending the plan. The marathon window is the one item that cannot be compressed, because it is set by pad life.

  1. Days 1–15

    Name the artefact, freeze the model, specify the regimes

    Agree the title and signatory of the prediction qualification report on day one. Freeze the model version under evaluation. With the module engineer, write the marathon specification by regime: at least one full pad life per chamber in scope, one wet clean, one PM cycle, one product changeover, and the first lots after each pad change called out separately because that is where disagreement concentrates.

    A named report, a named signatory, a frozen version and a regime specification

  2. Days 16–60

    Run the marathon and record by regime

    The frozen model scores every lot alongside the incumbent EWMA controller. Performance is recorded per regime and per chamber, not in aggregate. Failure modes are catalogued as they occur — feed interruptions, missing consumable-age data, chamber events — together with what the prediction did in each case, because the report's failure-mode section is the part equipment engineering will actually read.

    Regime-resolved performance and a catalogue of real failure modes

  3. Days 45–75

    Design the write path, the reason codes and the OCAP branch

    In parallel with the tail of the marathon: build the proposal write into the run-to-run controller's queue so the engineer sees the proposed setpoint beside the incumbent value and the delta. Define a short reason-code taxonomy with the engineers who will use it — six codes, not twenty. Draft the OCAP branch stating precedence when the prediction and the SPC chart disagree, and configure the incumbent controller as a one-switch fallback.

    A write path, a reason-code taxonomy and a drafted OCAP revision

  4. Days 76–90

    Review, sign, and raise the change

    Present the qualification report to module and equipment engineering. Expect the failure-mode section to generate the substantive discussion and expect at least one regime to be excluded — that is the review working. Raise the ECN, revise the OCAP, and drill the revert to the incumbent controller once on a live shift before the first approved proposal is written.

    A signed report, an issued ECN, a revised OCAP and a drilled revert

The order matters

  1. Evidence before accuracy

    A moderately accurate removal-rate prediction with a documented failure mode and a defined fallback passes a gate that a more accurate one with neither will not. Improve the model in the next version, behind the next window, once you can price an accuracy point in thickness range or consumable cost.

  2. Approval before autonomy, for at least two quarters

    Keep the engineer's approval on every proposal even where the write path could act on its own. The reason-coded log is the only evidential basis for the gate-3 envelope, and two quarters is roughly what it takes to cover the regimes that matter on a CMP fleet.

  3. One fleet before one module

    Extending to a second fleet is a second qualification, not a copy. Do it after the first has produced an approval log worth analysing, so the second inherits a reason-code taxonomy, an OCAP pattern and a report template rather than repeating their design.

  4. Assess change notification early, not at release

    If any product on the fleet is externally qualified, ask the quality organisation at gate 2 what a later closed loop would oblige you to notify. The answer occasionally moves the whole plan to a different fleet, and finding that out in month three is far cheaper than in month eighteen.

The evidence pack: what each gate actually asks for

The artefacts, what has to be inside each one, where the content comes from, and the telemetry that proves a gate is genuinely shut.

Each gate asks for one document, and each document has a small number of sections that reviewers actually read. Everything else in an evidence pack is context. The table below is the build sheet: what the artefact is called, what must be inside it, which system supplies that content, and which gate it closes — so a team can start writing the document on day one rather than assembling it retrospectively from whatever the study happened to produce.

ArtefactSections that get readWhere the content comes fromGate it closes
Data qualification reportField-by-field lineage; owner per field; known gaps and how they are treated; a reproducible query two engineers agree onMES, FDC historian, metrology host, maintenance recordsGate 1
Prediction qualification reportPerformance by regime; documented failure modes and observed behaviour in each; fallback behaviour; proposed OCAP branch and reason codesMarathon window logs, frozen model version, incident notesGate 2
Approval-log analysisOverride rate by chamber, product and consumable state; reason-code distribution; the regimes where model and engineer disagreeApproval log from the proposal layerGate 3 (input)
Control envelopeNamed chambers and products; named exclusions; maximum adjustment; boundary behaviour; version and signatureApproval-log analysis plus equipment engineering judgementGate 3
Revert drill recordDate, shift, who performed it, elapsed time to restore the incumbent controller, what did not go to planThe live drill itself — not a simulationGate 3
Propagation packageMatching hypothesis and evidence per target; revalidation protocol; release version; disqualification criteriaTarget chamber data, source envelope, change boardGate 4
Change-notification assessmentWhich products are externally qualified; whether the change is notifiable; customer agreements that apply; timingQuality organisation and customer-quality contactsGate 4
Evidence-pack build sheet for a wafer fab AI programme. The sections listed are the ones reviewers engage with; a pack missing one of them tends to be sent back regardless of its length.

Alongside the documents, four measurements tell you whether a gate is genuinely shut rather than administratively closed. All four are readable from systems the fab already runs, and all four are more informative than model accuracy at this stage.

MeasurementHow to read itPhase 2Phase 3Phase 4
Regime coverageShare of specified regimes actually observed inside the frozen evaluation windowOften under half100% before the report is signedRe-established for each new scope
Reason-code completenessShare of approvals and overrides carrying a code and captured contextNot applicableAbove 95%, or the log is not evidenceRetained for envelope revision
Time to restore the incumbentElapsed minutes in the last live revert drillUntestedTested off-lineDrilled on shift within six months
Envelope breach rateShare of decisions falling outside the envelope, trendedNot applicableNot applicableTrended and reviewed on cadence
Telemetry that verifies a phase transition. Read these instead of asking the team whether the gate is closed — a gate that is genuinely shut shows up in the logs.

Gate 3 readiness checklist

If you cannot tick all eight, you are still at engineer-approved regardless of how well the prediction performs. Tick as you go — this list works without JavaScript.

0 of 8 ticked

Nothing ticked yet — start with the artefact, not the tooling

Zero ticks usually means the decision is still in shadow with no written exit criterion. Name the qualification report and its signatory, specify the marathon window by regime, and freeze the model version. Those three acts take an afternoon.

How programmes fall back down the ladder

Phases are not monotonic. Four regressions account for almost all of it, and three are preventable with paperwork.

Phases are not monotonic, and a fab AI programme falls further per incident than most industrial software does, because the response to a bad automated action is nearly always to switch off everything of that class. Four regressions account for almost all observed cases, and three of the four are prevented by an artefact rather than by engineering.

Likelihood: highImpact: high

The envelope is widened in a settings screen

Six quiet months read as evidence that the limits were conservative, so another chamber or product is added through configuration rather than through a document revision. The new scope has no approval-log history behind it, and the first genuine excursion lands in a region the evidence never covered. The typical response is a full switch-off, which is a two-phase regression from a single event.

PreventionEnvelope changes go through management of change with a signature, exactly like a recipe change.

Likelihood: highImpact: medium

The model is retrained and the qualification is not repeated

A routine retrain ships because accuracy improved on backtest, and the version that was qualified is quietly no longer the version running. Nothing visible changes until the new version behaves differently in a regime the old one handled — usually post-clean or end-of-pad-life lots, which are underrepresented in any backtest.

PreventionTie the envelope's version to the model's version; a new model version enters at engineer-approved until it has its own window.

Likelihood: mediumImpact: high

The loop is propagated to an unmatched chamber

A working loop is copied to a chamber with a different consumable supplier, a different age or a different clean history. The model carries the statistical fingerprint of the equipment it learned on, so it acts confidently and wrongly, and the resulting scrap is attributed to AI in general rather than to the propagation step.

PreventionDocumented matching evidence per target, with disqualification criteria agreed in advance of the request.

Likelihood: highImpact: medium

The owner moves module and the loop keeps running

A phase-4 capability becomes an unowned one the moment nobody is accountable for its behaviour, because it continues acting and people continue trusting it. In a fab this bites hardest around ramps and requalifications, when the conditions the envelope assumed are exactly what is changing.

PreventionOwnership transfer is on the module handover checklist alongside SPC charts and OCAP ownership.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Gate
The transition between two phases, closed by a named artefact with a signatory outside the build team. A gate that was closed by a date or a steering-group decision has not been closed in any sense a quality reviewer will accept.
Shadow mode
Running a model continuously on live production data alongside the incumbent controller or rule, with no lot's outcome depending on its output. Safe, cheap, and the phase where fab AI programmes most often stall for years.
Marathon window
An extended evaluation run specified by the operating regimes it must cover — a full consumable life, a wet clean, a preventive-maintenance cycle, a product changeover — rather than by elapsed weeks. The model version must be frozen for the duration or the window is void.
Control envelope
The versioned document stating where a closed loop may act: named chambers, named products, named exclusions, maximum adjustment and behaviour at the boundary. Derived from the approval log, signed, and revised only through management of change.
OCAP
Out-of-control action plan — the fab's scripted response to a control-chart alarm. An AI signal that is not named in the OCAP is outside the fab's control discipline, and the shift will resolve model-versus-chart conflicts differently each time.
Management of change (MOC)
The controlled procedure through which any change to a qualified process enters production, issuing an engineering change notice. Routing AI-driven changes through it is what makes them visible to the excursion review that will eventually examine them.
Reason code
The mandatory, short-list category an engineer selects when accepting or overriding a model proposal. The reason-code distribution by chamber, product and consumable state is what turns an approval log into a defensible control envelope.
Chamber matching
Establishing that two chambers behave equivalently for the purpose at hand. Required before a qualified loop can be propagated, because a learned model encodes the statistical fingerprint of the equipment it was trained on.
Copy exact
The discipline of propagating a proven process by reproducing its conditions rather than re-optimising at each site — best known from Intel's Copy Exactly! methodology. Applied to a learned controller it requires matching evidence, because the model is not indifferent to the equipment underneath it.
Virtual metrology (VM)
A predicted measurement standing in for a physical one, computed from tool trace data. Its value is realised only when the sampling plan changes, which is a qualification argument about escape risk rather than a modelling result.
Process change notification (PCN)
The formal notice a fab owes customers before changing a qualified process, and a common phase-5 constraint. On automotive- or medical-qualified lines, changing how a setpoint is chosen can be notifiable and can trigger requalification.

Frequently asked questions

The questions fab teams ask most often when placing an AI programme on the phase ladder.

How long does it take to move a fab AI decision from shadow to engineer-approved?

About 90 days when it is scoped to one decision on one fleet with a named signatory, and several years when it is scoped to a module. The work is evidence, integration and paperwork rather than model development, because a decision that has been in shadow already has a model that performs. The one element that cannot be compressed is the marathon window, since it is set by the physical regimes — a full consumable life, a wet clean, a preventive-maintenance cycle — that the evidence has to cover.

Can different decisions be in different phases at the same time?

Yes, and they always are. Score each decision separately. The programme's phase is the level at which the shared foundations sit — qualified data, change-control fit, an analysable approval log — which is usually lower than the most advanced individual decision. A fab with one closed loop on an etch fleet and six models in shadow is a phase-2 programme with a phase-4 exception, and planning it as a phase-4 organisation is how the next four decisions get sequenced badly.

Do we need a fab data lake before starting AI?

No, and building one first is a reliable way to spend eighteen months without closing a gate. Qualify one decision's data end to end — one toolset, one metrology step, with chamber, consumable state, preventive-maintenance history and lot genealogy carried as data — and let the platform's shape emerge from the second and third decisions. The requirements a platform actually has to meet are visible after phase 3 has run once; before that they are guesses encoded as architecture.

What actually closes a gate in a wafer fab?

A named document with a signatory outside the build team. Gate 1 closes with a data qualification report; gate 2 with a prediction qualification report covering performance by regime, documented failure modes and the proposed OCAP branch; gate 3 with a versioned control envelope derived from the approval log plus a drilled revert and an issued engineering change notice; gate 4 with a propagation package containing matching evidence, a revalidation protocol and a completed change-notification assessment. Anything else is a status update.

How long should a shadow run last before it becomes a stall?

Judge it by coverage rather than duration. A shadow run is finished when it has observed the regimes specified in advance — typically a full consumable life, a wet clean, a preventive-maintenance cycle and a product changeover — with the model version frozen throughout. In practice that is three to six months on a CMP or etch fleet and longer on slow loops like diffusion queue-time effects. Any shadow run past nine months without a written exit criterion and a named signatory is a stall, not a study.

Why is the approval log more important than model accuracy at phase 3?

Because it is the only evidential basis for the control envelope that gate 3 requires. The log shows where the model and its engineers disagree and in which regimes — first lots after a pad change, a particular product layer, one chamber — and those disagreements become the envelope's named exclusions. Accuracy tells you how the model performs on average; the log tells you where it should not be allowed to act, which is the question a quality reviewer will ask.

Can a fab skip shadow if the model is clearly better on historical data?

It can, and it will pay for it at the next gate. Historical performance says nothing about behaviour when a feed stops, when consumable-age data is missing, or in the first lots after a chamber event — and those failure modes are what equipment engineering will interrogate in the qualification report. Skipping shadow also means the engineer approving each proposal is performing the qualification themselves, per lot, without evidence, which drives override rates up and fills the approval log with noise rather than signal.

Who should sign a fab AI qualification gate?

Someone with material at risk. The module process owner whose thickness range, scrap number or cycle time moves is the right signatory for gates 1 to 3, with equipment engineering co-signing the failure-mode section and quality joining at gate 3. Gate 4 belongs to the change board with quality and customer-quality representation. A gate signed by the programme sponsor tests enthusiasm rather than evidence, and it will not stand up in an excursion investigation.

How does phasing differ for a foundry versus an integrated device manufacturer?

The ladder is identical; the ceiling on phase 5 differs. An integrated device manufacturer owns its own product qualification, so propagating a closed loop is an internal quality decision. A foundry runs customer-qualified processes and may owe a process change notification before changing how a setpoint is chosen, with requalification implications on automotive or medical parts. Foundries should therefore sequence early decisions where the affected layers carry no externally qualified product, and assess notification obligations at gate 3 rather than gate 4.

Where do fab standards like SEMI E133 or E164 fit into the phases?

They define the vocabulary the phases operate in rather than gating them directly. SEMI E133 describes the advanced process control framework a phase-3 write path targets, E164 addresses the common equipment metadata that makes trace interpretable across a fleet — which is a gate 1 concern — E30 GEM covers the equipment communication layer, and E10 defines the equipment state model most availability metrics rest on. None of them says anything about machine learning. Their value is that a proposal expressed in their terms is legible to the automation organisation immediately.

What does a phase transition cost in team terms?

Gate 1 is roughly a data engineer and a process engineer at a day a week for two to four months. Gate 2 adds an ML engineer and a slice of module engineering for the marathon review. Gate 3 is the expensive one: an integration engineer for the write path, sustained module engineering attention to the approval log, plus equipment engineering and quality for the envelope. Gate 4 is mostly quality and change-board time per target. The dominant cost across all four is review capacity, which is why running one decision through consecutive gates finishes faster than four in parallel.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for wafer fabs and other high-mix process manufacturers — trace-data pipelines, virtual metrology, run-to-run and dispatch decision support — integrated into the APC, MES and recipe layer under the site's own change control, rather than delivered as dashboards.

  • · Control-loop integration against APC, RMS and MES rather than parallel tooling
  • · Qualification-first delivery: shadow harness, marathon evidence, reason-coded approval logs
  • · Rollback and OCAP design treated as a deliverable, not an afterthought
  • · 16 cited sources on this page

Sources

  1. FlexcitonFab-wide scheduling of semiconductor plants: a large-scale industrial deployment case study (opens in a new tab)
  2. FlexcitonTechnical paper, ASMC 2023 (opens in a new tab)
  3. GlobalFoundriesGlobalFoundries joins the World Economic Forum's Global Lighthouse Network (opens in a new tab)
  4. GlobalFoundriesDigital manufacturing (opens in a new tab)
  5. Micron TechnologySmart sight: how Micron uses AI to enhance yield and quality (opens in a new tab)
  6. NIST / SEMATECHe-Handbook of Statistical Methods — process or product monitoring and control (opens in a new tab)
  7. SEMIStandards programme (E10, E30 GEM, E133, E164) — landing page, bot-walled to automated fetching (opens in a new tab)
  8. Semiconductor EngineeringUsing ML for improved fab scheduling (opens in a new tab)
  9. arXivHeterogeneous domain adaptation and equipment matching (DBACS) (opens in a new tab)
  10. arXivEvent-driven reinforcement learning enables long-horizon control in semiconductor fabrication (opens in a new tab)
  11. arXivLearning to optimize capacity planning in semiconductor manufacturing (opens in a new tab)
  12. arXivReinforcement learning for process control with application in semiconductor manufacturing (opens in a new tab)
  13. arXivMFRL-BI: a model-free reinforcement learning process control scheme, demonstrated on CMP (opens in a new tab)
  14. arXivMachine learning based CVD virtual metrology in a mass-produced semiconductor process (opens in a new tab)
  15. arXivOn scheduling a photolithography process containing cluster tools (opens in a new tab)
  16. arXivSemiconductor wafer map defect classification with tiny vision transformers (opens in a new tab)

Find out which gate is holding you — then close it

We run the phasing assessment with your module, automation and quality leads, identify the decision with the shortest path to a closed gate, and leave you with a costed 90-day plan and a draft evidence pack. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.