Silicon Wafer EngineeringReadiness & Transformation Roadmap
Fab AI transformation phases: sequencing an AI programme in silicon wafer engineering
Fab AI transformation phases are the qualification states an AI system passes through inside a silicon wafer fab — off-tool, shadow, engineer-approved, closed loop within limits, and copy-exact release. Each phase ends at a gate, and a gate is closed by signed evidence, not by a date on a programme plan.

Key takeaways
- A fab AI phase is a qualification state, not a calendar period. The five states are off-tool, shadow, engineer-approved, closed loop within limits, and copy-exact release — and a programme is in the lowest phase any of its decisions actually occupies, not the highest one it has demonstrated.
- Four gates separate the five phases, and each is closed by an artefact the fab already knows how to issue: a data qualification report, a marathon evidence pack, an approval-log analysis with a control envelope, and a propagation package with chamber-matching evidence. If a phase ended without one of those, it did not end.
- Shadow is the phase that consumes years. A model running beside the incumbent EWMA controller costs almost nothing to keep and generates almost nothing that closes a gate, so it survives every budget review and never advances. The exit condition is a marathon window covering the fab's known regime changes — PM cycles, chamber swaps, product mix — not a better backtest.
- Skipping a phase is more expensive than repeating one. The recognisable failures — a closed loop qualified on one chamber propagated to an unmatched one, an approval log nobody analysed before setting limits, a change made outside management of change — all cost scrapped material and a suspended programme, and all are cheaper to prevent than to explain.
- Phase 5 is bounded by the customer, not by the model. On an automotive- or medical-qualified line, changing the way a setpoint is chosen can trigger a process change notification and a requalification, so copy-exact release is a commercial and quality decision that engineering can prepare for but cannot make alone.
Abbreviations used on this page
- MES
- Manufacturing execution system — the fab's lot and route system of record
- APC
- Advanced process control — the layer that adjusts recipe setpoints
- R2R
- Run-to-run control, typically an EWMA controller on a setpoint
- FDC
- Fault detection and classification, running on tool trace data
- VM
- Virtual metrology — a predicted measurement standing in for a physical one
- SPC
- Statistical process control — the control charts the fab already runs
- OCAP
- Out-of-control action plan — the scripted response to an SPC alarm
- MOC
- Management of change — the fab's controlled change procedure
- ECN
- Engineering change notice — the document an MOC issues
- RMS
- Recipe management system — where qualified recipes and limits live
- PCN
- Process change notification — the notice a customer is owed before a qualified process changes
- CMP
- Chemical mechanical planarisation — the worked example used throughout this page
Free · 8 questions · ~3 minutes
Score how your programme is phased
Eight questions, one at a time, about three minutes. They score sequencing discipline rather than ambition: whether your phase transitions produce artefacts, whether AI change runs inside the fab's own change control, whether decisions advance one at a time or accumulate in shadow, and whether the previous control state can actually be restored. The result names the phase you are genuinely in and the gate that is holding you there.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised phasing report is ready
Tell us where to send it. Your phase appears on screen straight away, and the full report — dimension scores, the gate that is holding you, and a 90-day plan to close it on one toolset — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Off-tool
Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
Your next moveClose the data qualification gate on one decision: one toolset, one metrology step, with chamber, consumable state, PM cycle and lot genealogy carried as data rather than remembered.
Stage 2 · Shadow
Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
Your next moveFreeze the model version, define the marathon window by the regimes it must cover, and write the prediction qualification report that the module owner will sign.
Stage 3 · Engineer-approved
Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
Your next moveAnalyse the approval log by context — chamber, consumable state, product, shift — and derive an explicit control envelope with named exclusions before proposing any unattended action.
Stage 4 · Closed loop in limits
Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
Your next moveProduce the propagation package: chamber-matching evidence, a per-site revalidation plan, the envelope as a controlled document, and the customer notification assessment.
Stage 5 · Copy-exact release
Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.
Your next moveTreat the envelope as part of the controlled release, sign every change, and keep model-to-lot traceability for the full qualification life of the affected products.
0 / 24
Gate evidence
— / 6
Change-control fit
— / 6
Sequencing discipline
— / 6
Reversion readiness
— / 6
Your score maps to a phase. The dimension breakdown matters more than the total: the lowest dimension is the gate that is actually holding you, and it is where the next quarter's work belongs — regardless of how good the underlying models are. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a phase. The dimension breakdown matters more than the total: the lowest dimension is the gate that is actually holding you, and it is where the next quarter's work belongs — regardless of how good the underlying models are.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the gate plan written against your own toolsets?
We will walk your module and automation leads through the dimension scores, pick the decision with the shortest path to a closed gate, and leave you with a costed 90-day plan and a draft evidence pack. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–4 — Stage 1, Off-tool. Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
- 5–10 — Stage 2, Shadow. Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
- 11–16 — Stage 3, Engineer-approved. Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
- 17–21 — Stage 4, Closed loop in limits. Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
- 22–24 — Stage 5, Copy-exact release. Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.
What fab AI transformation phases are — and why a date never closes one
A definition, the five qualification states, and the round trip one CMP lot's data makes at each of them.
Fab AI transformation phases are the qualification states an AI decision occupies inside a wafer fab, and a date never closes one because a phase is left by evidence rather than by elapsed time. The five states are off-tool, where the analysis touches nothing; shadow, where the model runs live beside the incumbent controller but binds nothing; engineer-approved, where it writes a proposal a named engineer signs off per lot; closed loop in limits, where it acts unattended inside a versioned envelope on a qualified toolset; and copy-exact release, where the whole assembly propagates under matching evidence and revalidation.
The important structural point is that a fab did not need AI to invent this ladder. It already runs one. Tool and process qualification, marathon runs, split lots, statistical process control with a written out-of-control action plan, management of change issuing engineering change notices, copy-exact propagation between sites, and process change notification to externally qualified customers — that machinery exists, is understood by every module engineer, and is the reason a 1,500-step process holds together at all. An AI programme that builds a parallel governance structure alongside it has doubled the paperwork and halved the authority. The programmes that move build their phases inside it.
This is why phase language borrowed from software — pilot, MVP, rollout, scale — sits so badly in a fab. Those words describe how much of the organisation is exposed to a change. Fab phases describe how much authority a decision has been granted over material that cannot be un-processed. A pilot in a fab is not a small rollout; it is a qualification state with a defined evidence requirement, and the statistical part of that requirement is set out in the NIST/SEMATECH e-Handbook's process-control chapters (opens in a new tab), which most fabs already treat as the reference for their control charting.
Authority granted against evidence accumulated
The curve is deliberately not smooth. Nothing is granted through off-tool and shadow no matter how much analysis accumulates, because neither phase produces the kind of evidence a gate accepts. The step happens at engineer-approved, when the decision first enters the controlled process, and the slope after it is set by how fast the approval log accumulates rather than by model quality. Illustrative shape, consistent with the published deployment reports cited on this page.
Authority granted over the controlled process by stage
- Stage 1 · Off-tool — 22% of operators. Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
- Stage 2 · Shadow — 34% of operators. Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
- Stage 3 · Engineer-approved — 27% of operators. Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
- Stage 4 · Closed loop in limits — 13% of operators. Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
- Stage 5 · Copy-exact release — 4% of operators. Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.
Curve shape: logistic, plotted from the stage data above. Distribution: Illustrative; shape consistent with GlobalFoundries' and Flexciton's published deployment accounts.
One CMP lot's data round trip, phase by phase
The same physical loop — trace out, decision back — drawn at each phase. What changes between lanes is not the model but where the arrow is allowed to terminate: at a scoreboard, at an engineer's approval, or at the controller itself. Most fab decisions are in the top lane.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- At phases 1–2, trace data leaves the FDC historian as an extract, is joined to metrology by hand, and produces a model that scores lots on a scoreboard nothing consumes. No engineering change notice exists because nothing in the controlled process has changed, so the run generates accuracy history rather than gate evidence — and accuracy history is exactly what a qualification review will not accept.
- At phase 3, trace and context stream together — chamber, pad life, conditioner state, lot genealogy — into a frozen model version whose output is written as a setpoint proposal into the run-to-run controller's queue. The process engineer approves or overrides each one with a mandatory reason code, and post-CMP metrology feeds back. The approval log produced here is the raw material for the next gate.
- At phases 4–5, a versioned envelope derived from that log lets the loop act unattended on named chambers and products. Anything outside the envelope holds the lot, escalates and reverts to the incumbent controller rather than being clamped and continued. Propagation to another chamber or site is a separate qualification with its own matching evidence, not a configuration copy.
Step-by-step insights
- The trace extract — why the context join, not the model, caps phase 1
- A hand-built join between trace and metrology encodes one engineer's assumptions about which timestamp to trust, how to attribute a lot to a chamber on a multi-platen tool, and what to do with rework. None of it is written down, so a second study is not comparable to the first even when both are correct. The fix that closes gate 1 is unglamorous: carry chamber, consumable state, PM cycle and lot genealogy as owned, lineage-tracked columns, and let the model be whatever it is.
- The shadow scoreboard — a comfortable dead end with a real cost
- Shadow is safe, cheap and survives every budget review, which is why decisions accumulate there. The cost is not the compute; it is that process engineers learn AI output is decorative and module owners learn the programme changes nothing. Three years of that is harder to reverse than never having started, because the next proposal is heard against the memory of the last one. The only defence is a written exit criterion agreed at the moment shadow starts, expressed in operating regimes rather than accuracy — and a frozen model version, because retraining mid-window silently restarts the window.
- Streamed context — what has to travel with the trace
- For a CMP decision the trace alone is nearly useless: removal rate depends on pad age, conditioner disc wear, slurry batch, platen and incoming thickness, and a prediction that cannot see those is fitting noise it will later be blamed for. The context set is decided by the process, not by the data platform, which is why the module engineer has to specify it. This is also the point where an honest fab discovers which of those facts exists only on a whiteboard in the sub-fab, and fixing that is part of gate 1, not a later concern.
- The setpoint proposal — write into the queue, not onto a screen
- The single highest-leverage design decision at phase 3 is where the proposal appears. Written into the run-to-run controller's queue beside the incumbent value, it is on the path the engineer already walks and the default action becomes the informed one. Displayed on a separate dashboard, acting on it is a voluntary extra step, and voluntary steps are the first thing dropped during a ramp or a qualification crunch — which is precisely when the decision is worth most. This is why write-path work usually deserves the quarter that teams want to spend on accuracy.
- The approval log — the instrument, not the paperwork
- Every accept and override, with its reason code and the context at the time, is the dataset from which the next gate's envelope is derived. It shows where the model and its engineers disagree and, crucially, in which regimes — first lots after a pad change, a particular product layer, a specific chamber. A log analysed quarterly turns into limits with evidence behind them. A log collected and never read leaves a fab with a full database, no envelope, and the choice between staying at phase 3 forever or guessing.
- The envelope boundary — hold and escalate, never clamp and continue
- The behaviour at the edge of the envelope is what a quality reviewer will examine, because it is where the design's honesty shows. Clamping the adjustment to the limit and letting the lot proceed hides the fact that the world has moved outside the qualified region; holding the lot and escalating surfaces it while the material is still recoverable. Trending the breach rate then gives the fab a leading indicator with real meaning: a rise says the process has drifted out of the envelope's validity and the envelope needs review before an excursion forces one.
The five phases in detail
For each phase: what it looks like on the floor, the signals a reviewer can check in an afternoon, the anti-pattern that traps fabs there, and what leaving costs in team terms.
Each phase below is written for a process or automation engineer rather than for a buyer. The hallmarks describe observable conditions on the floor, the diagnostic signals are checks you can run against your own document system and logs this week, and the anti-pattern is the specific mistake most often made trying to leave that phase.
Select a phase
Every phase's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Off-tool
22% of operators sit here
Off-tool is the phase in which the analysis exists but touches nothing: extracted trace data, a model on an engineer's workstation, and no path into any lot's history.
Off-tool is where nearly every fab AI idea legitimately starts, and where a surprising number of them are still sitting three years later. The work itself is often excellent: a process engineer who has lived with a chamber for a decade builds a model of removal rate against pad life, or of etch depth against chamber-wall condition, and it explains variation the existing control scheme does not. What is missing is not insight. What is missing is a route from that insight into a lot's history.
The diagnostic is the data path. At this phase the authoritative version of last quarter's chamber behaviour is an extract someone pulled with a filter they chose, on a day they remember approximately, joined to metrology by a lot ID and a timestamp that may or may not be in the same time zone. Two engineers asked the same question will produce two defensible and different answers, because the context — which chamber, which pad, which conditioner disc, which product layer, which preventive-maintenance cycle — was reconstructed by hand each time rather than carried by the data.
This is a cheap phase to leave and an expensive one to occupy. The cost is not the models that go nowhere; it is that nothing accumulates. The tenth study costs what the first one did, because the context join is rebuilt from scratch every time and none of the runs is comparable to any other. The gate out of off-tool is therefore not about modelling at all — it is a data qualification exercise, and it is the only gate on this ladder that a data team can close largely on its own.
In practice
The removal-rate study that runs twice a year
An oxide CMP module engineer is asked, roughly every six months, why post-polish thickness range widens toward the end of pad life. Each time, an engineer exports six weeks of trace from the FDC historian, joins it by hand to post-CMP metrology and to the pad-change log kept in a maintenance spreadsheet, and produces a well-argued deck showing the drift. The deck is correct. Nothing in the recipe, the R2R controller or the pad-change interval changes, and six months later the same export is pulled again — with a slightly different filter, so the two studies cannot be compared.
What it looks like
- Trace and metrology data arrives as manual extracts from the FDC historian or a data-warehouse query
- The model runs on an engineer's workstation or a shared analytics server, on demand
- No management-of-change record exists, because nothing in the controlled process has changed
- Results are presented at a yield or module meeting and then stop moving
Diagnostic signals you can check this week
- Ask where the model's training data came from. If the answer is a filename or a saved query, you are in this phase
- Ask whether chamber, pad life, conditioner state and PM cycle are columns in the dataset or facts someone remembered
- Check whether any AI output has ever been referenced in an OCAP, an ECN or a qualification report
- Ask two engineers for the same statistic — for example mean removal rate on one chamber last quarter — and compare the two numbers
Anti-pattern · Building the fab data lake before qualifying one dataset
The instinctive response to a hand-built context join is a platform programme: ingest every tool, every parameter, every trace channel, and sort out meaning later. It is the most reliable way to spend eighteen months without closing a single gate, because the requirements are being guessed rather than observed, and because a lake with no qualified dataset in it is indistinguishable from an expensive export. Qualify one decision's data end to end — one module, one toolset, one metrology step, with lineage and owners — and let the platform's real shape emerge from the second and third of those.
What holds you here
No qualified, context-complete dataset exists, so every study rebuilds the world by hand and no two studies are comparable.
Highest-leverage next move
Close the data qualification gate on one decision: one toolset, one metrology step, with chamber, consumable state, PM cycle and lot genealogy carried as data rather than remembered.
Cost of leaving
- Effort
- 2–4 months
- Team
- One data engineer, one process engineer at roughly a day a week
- Risk
- Low — nothing in the controlled process is touched, so nothing can be scrapped
- To next stage
- 2–4 months
If this is you, the next step is
A short engagement: pick the decision, fix the context join, produce the data qualification report.
Stage 2
Shadow
34% of operators sit here
Shadow is the phase in which the model runs continuously on live data alongside the incumbent controller or rule, and no lot's outcome depends on what it says.
Shadow is the most comfortable phase on the ladder and therefore the most dangerous. Running a predictor beside the existing controller is cheap, technically satisfying and entirely safe: the model sees production reality, its errors are visible, and no wafer is at risk. Every one of those properties is a reason to stay. Fabs routinely keep decisions in shadow for two or three years, and the programme reports steadily improving accuracy the whole time without a single gate closing.
The structural reason is that shadow answers a question nobody needs answered. A shadow run demonstrates that the model would have been right; the gate to the next phase requires evidence that it will be right through the regimes the process actually visits — the end of a pad's life, the first lots after a chamber wet clean, a product mix shift, the quarter a second-source consumable is introduced. That is a marathon question, and it is answered by designing the shadow window to cover those events deliberately rather than by leaving the model running and hoping the events show up.
There is also an organisational cost to a long shadow. Process engineers learn that the AI output is decorative, module owners learn that the programme does not change anything, and the next proposal is funded against that memory. A fab that has kept six decisions in shadow for three years is harder to move than a fab that has never tried, because the antibodies are established and the phrase people use is not sceptical, it is bored.
In practice
The virtual metrology screen nobody has ever acted on
A 200mm fab stood up a virtual metrology model for post-etch critical dimension, running on chamber trace after every lot. On the shadow scoreboard it tracks the measured CD closely on the sampled lots, and it has been doing so since 2024. It is displayed on a screen in the module office. Sampling has not been reduced by a single wafer, because reducing sampling requires a qualification argument and a change to the sampling plan, and no one has ever written one. The model has been correct for two years in a way that has cost the fab metrology time rather than saving it.
What it looks like
- The model scores every lot in near-real time and its output is stored, not applied
- A comparison against the incumbent EWMA controller or dispatch rule is visible to engineers
- No recipe, setpoint, dispatch decision or hold is affected by the output
- Nobody has opened a management-of-change record, because formally nothing has changed
Diagnostic signals you can check this week
- Ask how long the decision has been in shadow. Anything past nine months without a written gate plan is a stall, not a study
- Check whether the shadow window deliberately spans a PM cycle, a wet clean and a consumable change, or whether it is simply elapsed time
- Ask what the exit criterion is, in a sentence. If the answer contains an accuracy number and no operating regime, the gate is not defined
- Look for a frozen model version. If the model has been retrained during the shadow window, the window has been restarted and nobody noticed
Anti-pattern · Improving accuracy to earn the right to act
When shadow does not convert, the reflex is to improve the model, on the theory that authority follows precision. It rarely does. What gates ask for is evidence of behaviour under known regimes, a frozen version, a documented failure mode and a defined response when the prediction is wrong — none of which is an accuracy problem. A predictor that is moderately accurate and demonstrably well behaved across pad life and post-clean lots will pass a gate that a more accurate one with an undocumented failure mode will not. Spend the next quarter designing the marathon and writing the OCAP branch, then revisit accuracy when you can price an accuracy point in metrology hours or scrapped wafers.
What holds you here
There is no defined exit criterion, so the shadow run continues indefinitely and produces accuracy history instead of gate evidence.
Highest-leverage next move
Freeze the model version, define the marathon window by the regimes it must cover, and write the prediction qualification report that the module owner will sign.
Cost of leaving
- Effort
- 3–6 months
- Team
- One ML engineer, one process engineer as owner, module engineering time for the marathon review
- Risk
- Low to medium — the risk is elapsed time and credibility, not material
- To next stage
- 3–6 months
If this is you, the next step is
We define the regimes the shadow window must cover and the evidence pack that comes out of it.
Stage 3
Engineer-approved
27% of operators sit here
Engineer-approved is the phase in which the model writes a proposal into the system that runs the decision — an APC setpoint, a sampling choice, a dispatch priority — and a named engineer approves or overrides each one.
Engineer-approved is the first phase in which the programme survives the person who built it. There is a document trail, a named owner, a defined fallback and a logged human decision on every action, which together mean the capability keeps working when the engineer moves module. It is also the first phase in which the fab gets something back: metrology time, cycle time, scrapped material or engineer attention, in units the module already reports.
The character of the work changes sharply here. Off-tool and shadow problems are analytical; engineer-approved problems are procedural. Which field does the proposal write to. What does the engineer see when they approve. What is the reason code taxonomy, and who maintains it. What does the OCAP say when the model and the SPC chart disagree. What happens on a shift where the data feed stops. These have well-established answers in a fab's own control discipline — the NIST/SEMATECH handbook's process-control chapters are still the clearest public statement of the statistical part — and importing them wholesale is far faster than rediscovering them under a different vocabulary.
The trap that emerges is the unread approval log. The log is the single most valuable artefact this phase produces: it is the dataset from which the next gate's control envelope is derived, and it is the only place the fab can see where the model and its engineers disagree and why. Fabs that treat approval as a formality rather than as instrumentation reach the end of a year with a full log, no analysis, and no basis for setting limits — which means they either stay here indefinitely or set an envelope by guesswork.
In practice
The setpoint proposal on the APC screen
An oxide CMP fleet runs a pad-life-aware removal-rate prediction that proposes a polish-time setpoint into the R2R controller's queue. The process engineer sees the proposal, the incumbent EWMA value and the delta, and approves or overrides with one of six reason codes. In the first two months the override rate on freshly conditioned pads runs far higher than on mid-life pads — a pattern nobody predicted, visible only because the reason codes were mandatory. That pattern later becomes the first exclusion in the control envelope: the loop does not act on the first lots after a pad change.
What it looks like
- Output is written into the APC, RMS or MES queue the engineer already works in, not a separate screen
- Every accept and override is logged with a reason code and the context at the time
- An ECN exists, the OCAP has a branch for the new signal, and the incumbent controller is one switch away
- A named module owner is accountable for the decision's behaviour, not the data team
Diagnostic signals you can check this week
- Ask to see the approval log. If reason codes are optional or a free-text box, the phase is running but not instrumented
- Check whether the incumbent controller can be restored by an engineer on shift, without a ticket to another team
- Read the OCAP. If it does not mention the model at all, the model is outside the fab's control discipline
- Ask who is paged when the prediction feed stops. If the answer is a data team with no on-shift presence, the fallback is theoretical
Anti-pattern · Treating approval as a formality on the way to automation
Because approval feels like a temporary inconvenience, teams optimise it away: a single accept button, no reason codes, no context capture, engineers clicking through a queue at shift change. The phase then generates no evidence, and the gate to closed loop has to be argued from model accuracy — which is exactly the argument that will not survive a quality review. The approval step is not a concession to caution. It is the instrument that produces the envelope, and a fab that rubber-stamps for a year has to spend another year re-earning what it threw away.
What holds you here
The approval log is collected but never analysed, so there is no evidential basis for the limits a closed loop would run inside.
Highest-leverage next move
Analyse the approval log by context — chamber, consumable state, product, shift — and derive an explicit control envelope with named exclusions before proposing any unattended action.
Cost of leaving
- Effort
- 6–12 months
- Team
- Integration engineer, ML engineer, module process owner, an automation or MES contact for the write path
- Risk
- Medium — the first write into a controlled system needs an ECN, an OCAP revision and a drilled revert
- To next stage
- 6–12 months
If this is you, the next step is
Reason-code taxonomy, context capture and the analysis that turns the log into a control envelope.
Stage 4
Closed loop in limits
13% of operators sit here
Closed loop in limits is the phase in which the model acts without per-lot human approval, but only inside a versioned envelope on a qualified toolset, with everything outside it routed to the OCAP.
Closed loop in limits sounds like the destination and is better understood as a narrow, carefully bounded permission. It is not an autonomous fab and it is not an autonomous module: it is one decision, on an enumerated set of chambers, for an enumerated set of products, within a stated adjustment magnitude, with a written response for everything else. Fabs that describe this phase in broader terms than that are usually describing an intention rather than a qualified state.
The engineering here is mostly finished by the time a fab arrives. What is not finished is the envelope, and the envelope is the deliverable. It answers four questions in writing — where may this act, on what, by how much, and what happens at the boundary — and it is versioned and reviewed like a recipe because that is exactly what it is. Published research is instructive about how much of this is still open: reinforcement-learning controllers have been compared favourably with EWMA and general harmonic rule controllers, and demonstrated on a nonlinear CMP process without an explicit model, but that work is numerical and simulation-based. It tells you the control idea is sound; it does not tell you your chamber is in the envelope.
The constraint that binds at this phase is not technical confidence but material consequence. A closed loop on a CMP setpoint that drifts in the wrong direction removes material that cannot be put back. That is why the boundary behaviour — hold and escalate, never clamp and continue — matters more than the loop's average performance, and why the excursion review is a standing item rather than an incident response. Fabs that get this right treat the envelope breach rate the way they treat an SPC chart: as a signal about the world, not as a nuisance.
In practice
The envelope that names three chambers and excludes two
A fab qualifies a closed-loop removal-rate adjustment on five oxide CMP chambers. Chamber matching evidence supports three of them; the remaining two show a consistent offset traced to a different platen supplier. The envelope names the three, excludes the two by serial number, caps the adjustment at a stated fraction of the recipe's polish time, and excludes the first lots after any pad change. The two excluded chambers stay at engineer-approved. That is not a failure of the programme — it is the phase working correctly, and the document is what makes it defensible a year later.
What it looks like
- The envelope is an explicit, versioned document: which chambers, which products, which consumable states, which magnitude of adjustment
- Excursions outside the envelope hold the lot and escalate rather than being clamped silently
- The revert to the incumbent controller has been exercised deliberately, not just documented
- Envelope breach rate is monitored as a leading indicator and reviewed on a defined cadence
Diagnostic signals you can check this week
- Ask to see the envelope as a document with a version number. If it is a configuration screen, it is not a controlled artefact
- Check when the revert to the incumbent controller was last exercised on a live shift, not simulated
- Ask what happens at the boundary. If the answer is that the adjustment is capped and the lot continues, the escalation path does not exist
- Check whether envelope breach rate is trended and reviewed, or only looked at after an excursion
Anti-pattern · Widening the envelope because nothing has gone wrong
Six quiet months read as evidence that the limits were too conservative, and the envelope grows — another chamber, another product, a larger adjustment — usually in a configuration change rather than a controlled document revision. The new scope has no approval-log history behind it, so the first genuine excursion happens in a region the evidence never covered. The predictable outcome is that the loop is switched off entirely and the fab regresses two phases from one event. Widen the envelope the way you earned it: from the approval log for that chamber and that product, through management of change.
What holds you here
The envelope holds for the chambers and products it was qualified on, and every extension needs its own evidence — so scope grows slowly and pressure to widen it informally is constant.
Highest-leverage next move
Produce the propagation package: chamber-matching evidence, a per-site revalidation plan, the envelope as a controlled document, and the customer notification assessment.
Cost of leaving
- Effort
- 9–18 months from first approval
- Team
- Module process owner, equipment engineering, quality, plus the platform team
- Risk
- Higher — the consequence of a wrong action is material, and the evidence burden is a quality matter
- To next stage
- 12–24 months
If this is you, the next step is
We stress-test the limits, the boundary behaviour and the revert against a real excursion scenario.
Stage 5
Copy-exact release
4% of operators sit here
Copy-exact release is the phase in which a qualified loop becomes a controlled, versioned release propagated to other chambers, fleets and sites under revalidation, with customer notification handled where the process is externally qualified.
Copy-exact release borrows its discipline from a practice the industry has run for decades: propagate a proven process by reproducing its conditions exactly rather than re-optimising at each site, and treat deviation as something to be justified rather than assumed harmless. Intel's Copy Exactly! methodology is the best-known articulation of it, and the logic transfers directly to a learned controller — with one important amendment. A recipe copied exactly onto matched equipment behaves the same way. A model copied exactly onto a differently-behaving chamber does not, because the model encodes the statistical fingerprint of the equipment it learned on.
That amendment is an active research problem rather than a solved one. Work on domain adaptation and equipment matching exists precisely because deep-learning-based process monitoring does not transfer cleanly between non-identical tools, and proposes methods to align them. Read that literature as a warning about sequencing: propagation is a qualification exercise per target, and a fab that budgets it as a deployment task will discover the difference on the first mismatched chamber.
The other binding constraint at this phase is commercial. On a line qualified for automotive or medical parts, changing how a setpoint is chosen can constitute a process change, which can oblige the fab to issue a process change notification and, depending on the customer's agreement, to requalify. That makes copy-exact release a decision the quality organisation and the customer-facing account team make jointly with engineering. It is entirely normal for a technically ready loop to be held at closed loop in limits for a year on one line while it runs released on another — and a programme that has not modelled that is not planning, it is hoping.
In practice
The release that shipped to two of four fabs
A qualified etch-depth loop is packaged as a controlled release: model version, envelope, OCAP branch, matching criteria and revalidation protocol. Two fabs take it inside a quarter — same tool generation, comparable chamber population, no externally qualified products on the affected layers. The third fab's chamber population fails the matching criteria and enters its own engineer-approved phase to build local evidence. The fourth runs automotive-qualified product on that layer and holds pending a customer notification assessment. One release, four correct and different outcomes.
What it looks like
- The model, its envelope and its OCAP branch ship together as one versioned, controlled release
- Propagation to a new chamber or site requires matching evidence and site revalidation, not a copy of a configuration file
- Change notification obligations to externally qualified customers are assessed before, not after, release
- Model and envelope versions are reconstructable for any lot, months later, for a customer or quality audit
Diagnostic signals you can check this week
- Ask whether the model version and envelope version that governed a specific lot three months ago can be recovered from the record
- Check whether propagation to a new chamber requires documented matching evidence or only an engineer's judgement
- Ask who assesses customer change-notification obligations, and whether they are consulted before release or after
- Check whether a released loop has ever been withdrawn, and whether the withdrawal path is written down
Anti-pattern · Treating the envelope as configuration once the release exists
The release is version-controlled, and then the limits are tuned in a settings screen with no revision history and no signature. The system works until someone has to explain a decision made eight months earlier, at which point neither the model version nor the limit that produced it can be reconstructed, and the fab cannot demonstrate to a customer that the qualified process was in force. Version the envelope with the model, require a signature on every change, and keep the trail for as long as the product's qualification requires — which on automotive parts is a long time.
What holds you here
Propagation is gated by equipment matching and by external qualification obligations, so the constraint becomes evidence and change notification rather than engineering.
Highest-leverage next move
Treat the envelope as part of the controlled release, sign every change, and keep model-to-lot traceability for the full qualification life of the affected products.
Cost of leaving
- Effort
- Continuous
- Team
- Platform and module teams plus a standing change board with quality and customer-quality representation
- Risk
- Concentrated — low frequency, high consequence, and increasingly a customer-contractual matter
If this is you, the next step is
We reconstruct a specific lot's governing model and envelope version from the record, the way an auditor would.
Where wafer fabs actually sit across the five phases
The distribution is heavily weighted toward shadow, and the shadow-to-approved step is the largest single loss on the ladder.
Most fab AI decisions are in shadow. The pattern that shows up repeatedly is a fab with several models running live against production data, none of them affecting a lot, and a programme narrative built on accuracy improvements — with a much smaller number of decisions that have actually entered the controlled process, and a very small number running unattended inside a written envelope.
Illustrative distribution of fab AI decisions across the five phases
Illustrative distribution, synthesised from the published deployment accounts and research cited on this page — not a survey. The shape is the argument: shadow is the mode, and the step down to engineer-approved is the largest single loss on the ladder.
Share of fab AI decisions
- 22% — 1 · Off-tool
- 34% — 2 · Shadow (the stall)
- 27% — 3 · Engineer-approved
- 13% — 4 · Closed loop in limits
- 4% — 5 · Copy-exact release
Source: Illustrative; synthesised from the published sources listed on this page
The published record is consistent with that shape in an instructive way: what operators announce is nearly always a portfolio, and what they quantify is nearly always the far end of it. GlobalFoundries reports having deployed over 60 smart manufacturing solutions since 2020 (opens in a new tab) across its fabs, with its Singapore 300mm site named to the World Economic Forum's Global Lighthouse Network in September 2025 — five years of accumulation, not a programme quarter. Its own manufacturing pages describe custom-built engines that classify wafer patterns automatically and speed troubleshooting by up to 10× (opens in a new tab), which is a phase-3-and-beyond claim about engineer time, not an accuracy claim.
That distinction between simulated and qualified runs through the whole research literature and is worth internalising before setting a programme's expectations. Long-horizon reinforcement-learning control for fabs is evaluated against high-fidelity simulations of industry-real scenarios (opens in a new tab); photolithography cluster-tool scheduling algorithms are described by their authors as showing promise for real-world implementation (opens in a new tab); wafer-map defect classifiers report F1 scores above 98% on public benchmark datasets (opens in a new tab). All three are good work. None of them is a closed gate in your fab, and treating a published result as though it were is how a programme ends up promising phase 4 outcomes on a phase 2 evidence base.
The gate register: four gates, and the artefact that closes each
The centre of this page. One row per gate — what it opens, what has to be true to enter it, the document that closes it, who signs, and how long it honestly takes.
Four gates separate the five phases, and each is closed by an artefact the fab already knows how to issue. That is the whole design principle: a gate that requires a new kind of document will be argued about, deferred and eventually skipped, while a gate that requires a qualification report, an ECN, an OCAP revision or a propagation package slots into machinery that already has signatories, a review cadence and a filing location. The register below is the page's working instrument — everything after it is either an application of a row or a consequence of skipping one.
| Gate | Phase it opens | Entry condition | Exit evidence — the artefact | Signed by | Typical elapsed |
|---|---|---|---|---|---|
| 1 · Data qualification | 2 · Shadow | One decision named, with its metrology step and its toolset. The context set — chamber, consumable state, PM cycle, lot genealogy, product layer — specified by the module engineer, not by the data team. | A data qualification report: lineage for every field, owner per field, known gaps and their treatment, and a reproducible dataset that two engineers can query to the same answer. | Module process owner and the data owner jointly | 2–4 months |
| 2 · Prediction qualification | 3 · Engineer-approved | A frozen model version and a marathon window specified by regime — a full pad or consumable life, a wet clean, a PM cycle, a product changeover — rather than by elapsed weeks. | A prediction qualification report: performance by regime not in aggregate, the documented failure modes, the fallback behaviour, and the proposed OCAP branch and reason-code taxonomy. | Module engineering, with equipment engineering on the failure-mode section | 3–6 months |
| 3 · Action qualification | 4 · Closed loop in limits | At least two quarters of reason-coded approval log covering the regimes in scope, analysed by context rather than in aggregate. | A control envelope as a versioned document — named chambers, named products, named exclusions, maximum adjustment, boundary behaviour — plus a drilled revert and the ECN that puts it into the controlled process. | Module owner, equipment engineering and quality | 6–12 months |
| 4 · Propagation qualification | 5 · Copy-exact release | A stable envelope on the source toolset, and a candidate target with a documented matching hypothesis. | A propagation package: chamber-matching evidence, the per-site revalidation protocol, the model and envelope as one controlled release, and a completed customer change-notification assessment. | Change board, including quality and customer-quality representation | 3–9 months per target |
Two properties of the register do most of the work. First, the exit evidence for every gate is a document with a signatory who is not the person who built the model — which is what stops a programme grading its own homework. Second, the entry condition for each gate is the exit evidence of the last one, so the register is a chain rather than a menu: there is no legitimate route from a shadow run to a closed loop, because the envelope in gate 3 is derived from the approval log that only exists once gate 2 has been closed and phase 3 has actually been run.
How to write a gate that actually closes
Name the artefact before the work starts
The single most effective intervention available is to write the title and the signatory of the closing document on the day the phase begins. It converts an open-ended study into a piece of work with a definition of done, and it surfaces immediately whether anyone senior is willing to put their name on the result — which is the real question and is much cheaper to answer at the start.
Specify the window by regime, never by duration
"Twelve weeks" samples whatever happened in twelve weeks. "A full pad life, one wet clean, one PM cycle and one product changeover, with the model version frozen" samples the conditions the process actually visits. The second takes about as long and produces evidence the first cannot, because the failures worth knowing about live in the transitions.
Make the signatory someone with material at risk
A gate signed by the programme sponsor tests enthusiasm. A gate signed by the module owner whose scrap number moves tests the evidence. If the module owner will not sign, the interesting information is why — and it is almost always a specific, addressable objection about a regime the evidence did not cover.
Route the change through management of change, always
Every AI-driven change to a setpoint, a sampling plan or a dispatch rule goes through the same MOC and ECN route as any other process change, and revises the OCAP where it introduces a new signal. This is not bureaucratic tribute. It is what makes the change visible to the excursion review that will eventually look at it, and what makes the programme legible to a customer audit.
Keep one gate open at a time per decision
Fabs that run several decisions through several gates simultaneously discover that review capacity — module engineering, equipment engineering, quality — is the binding constraint, not engineering capacity. One decision moving cleanly through consecutive gates finishes faster than four moving in parallel, and it produces a template the next four inherit.
Skipping gate 1 costs you every later comparison
Without a qualified dataset, the shadow run's performance cannot be attributed to a regime, because the regime was never recorded. The gate 2 report then has nothing to say about behaviour under pad-life or post-clean conditions, and the reviewer's only available question is about aggregate accuracy — which is exactly the question that does not justify granting authority.
Skipping gate 2 puts an unqualified prediction in front of an engineer
The engineer is then doing the qualification themselves, per lot, without the evidence to do it well. Override rates run high, trust erodes fast, and the approval log fills with noise instead of signal — which quietly disqualifies gate 3 as well, since the envelope has to come from that log.
Skipping gate 3 is the one that scraps material
An envelope set from model confidence rather than from approval history has no evidence about the regimes where engineers actually disagreed with the model. The first excursion lands in one of them. The typical outcome is not a tuned envelope but a switched-off loop and a two-phase regression from a single event.
Skipping gate 4 breaks on the first unmatched chamber
A model carries the statistical fingerprint of the equipment it learned on, which is why equipment matching and domain adaptation are an open research problem rather than a deployment detail (opens in a new tab). Copying an envelope to a chamber with a different platen supplier, a different age or a different clean history is the fastest way to turn a good loop into a scrap event on a line that was working yesterday.
Where AI lands in a wafer fab, and the phase each decision can realistically reach
Seven process areas, the decisions worth wiring, the system that owns each one, and the gate that limits how far it can go.
AI value in a wafer fab concentrates in seven operating areas, and each one has a different ceiling — set not by model difficulty but by what the decision touches. A decision is a good candidate for early phases when three things are true: the system of record is inside the fab's own change control, the feedback loop closes inside a shift or two rather than at final test, and the metric it moves is one a module owner already reports. The map below is how a first and second decision get chosen.
| Process area | High-value decisions | System of record | Metric it moves | Realistic ceiling | Limiting gate |
|---|---|---|---|---|---|
| Litho | Overlay and focus feed-forward, reticle and job sequencing on the bottleneck toolset | Scanner job control, APC, MES dispatcher | Overlay residual, rework rate, litho toolset moves | Phase 4 on overlay; phase 3 on sequencing | Gate 3 — scanner-adjacent control is tightly governed and the envelope argument is hard |
| Etch and deposition | Endpoint and depth prediction, chamber-condition-aware setpoint adjustment, seasoning cadence | APC / R2R, RMS, FDC | Depth or thickness range, chamber-to-chamber spread, requal frequency | Phase 4 per matched fleet | Gate 4 — chamber matching, not the model |
| CMP and planarisation | Removal-rate prediction over pad life, polish-time setpoint, conditioner and pad-change timing | APC / R2R, maintenance system | Post-CMP thickness range, dishing and erosion, consumable cost per wafer | Phase 4 | Gate 3 — the approval log has to cover post-pad-change regimes |
| Diffusion, thermal and implant | Batch composition, queue-time risk scoring, furnace profile compensation | MES dispatcher, APC | Queue-time violations, batch utilisation, uniformity | Phase 3, phase 4 on queue-time holds | Gate 2 — queue-time effects are slow and marathon windows are long |
| Metrology and defect inspection | Virtual metrology for sampling reduction, wafer-map classification, review-image triage | Metrology host, YMS, MES sampling plan | Metrology hours per lot, time to disposition, escape rate | Phase 4 on sampling; phase 3 on classification | Gate 2 — reducing sampling requires a qualification argument, not a model |
| WIP flow and dispatch | Lot prioritisation, toolset scheduling, AMHS routing, bottleneck protection | MES dispatcher, scheduler, AMHS controller | Cycle time per layer, x-factor, on-time delivery to commit | Phase 4 fab-wide, demonstrated publicly | Gate 3 — the envelope is the set of rules the scheduler may override |
| Sub-fab and facilities | Abatement and pump health prediction, chiller and UPW load management, exhaust balancing | Facilities SCADA / BMS | Unplanned tool downtime from facilities, energy per wafer | Phase 4, sometimes faster than the fab floor | Gate 1 — facilities historians rarely carry fab-side context |
The two rows most often misread are metrology and WIP flow. Virtual metrology looks like the easiest phase-4 decision in the fab because the model is comparatively tractable — published work on chemical-vapour-deposition virtual metrology in a mass-produced process reports prediction accuracy around 0.7 with roughly 70% data availability (opens in a new tab), and argues that this supports reduced physical metrology frequency. But the value only arrives when the sampling plan changes, and changing a sampling plan is a qualification argument about escape risk, not a modelling result. Fabs that skip that argument end up with a permanently accurate model and a permanently unchanged sampling plan — the phase-2 stall in its purest form.
WIP flow runs the other way. Scheduling and dispatch look institutionally harder because they touch every module, yet they are the area with the clearest public evidence of a fab reaching phase 4 across an entire site, and the practical literature on machine learning for fab scheduling (opens in a new tab) has been discussing production deployments for years. The reason is structural: a dispatch decision is reversible on the next lot, it does not remove material, and the fallback is a rule set the fab already runs. Reversibility, not tractability, is what sets how far a decision can travel.
The compliance frame is the other ceiling, and in a foundry it is the harder one. The industry's own standards define the vocabulary these phases live in — SEMI E10 for equipment reliability and availability states, E30 GEM for the equipment communication layer, E133 for the advanced process control framework itself, and E164 for the common metadata that makes tool data interpretable across a fleet (the SEMI standards programme (opens in a new tab) is the reference, though its site is bot-walled to automated fetching). Above those sit the customer-facing obligations: a line producing automotive parts under IATF 16949 carries change-notification duties, and a process change notification can convert a technically finished phase-4 loop into a nine-month commercial conversation. Assess that at gate 3, not at gate 4 — the answer occasionally changes which toolset you should have started on.
Two of the phase ceilings above are set by infrastructure rather than by evidence — whether trace can leave a tool at the rate a control loop needs, and where the compute has to sit to meet the deadline. That is a different question from sequencing and it has its own page: see AI readiness infrastructure for wafer fabs for the interfaces, clock budgets and qualified write paths underneath this map.
Why phases get run out of order — and what it costs
Three structural patterns account for most out-of-order programmes, and none of them is a modelling problem.
Phases get run out of order because the fab's incentives and the programme's incentives point at different artefacts. A programme is funded against a roadmap with dates on it, and dates are met by demonstrating capability; a fab grants authority against evidence, and evidence accumulates at the speed of pad lives and PM cycles. When those two clocks disagree, the programme reaches for the phase that can be shown on the date rather than the phase the evidence supports, and the sequence breaks in one of three recognisable ways.
The demonstration jumps the ladder
A convincing closed-loop demonstration is built on a single chamber for a review, without gate 2 or gate 3 behind it. It works, because a single chamber over a short window is a benign environment. The demonstration then becomes the reference point for what the programme can do, and every later request for a marathon window or an approval log reads internally as engineering caution rather than as the actual gate. The fix is procedural and costs nothing: label the demonstration as a phase-1 artefact in the room where it is shown.
Shadow is used as a substitute for a decision
When nobody wants to own the risk of granting authority, shadow is the answer that offends no one: the model stays live, the programme stays funded, and the decision is deferred indefinitely. This is the most common cause of a multi-year stall and it is not a technical failure at all — it is an unassigned owner. The tell is that nobody can name who would sign the gate 2 report, which is a question worth asking at the start of every shadow run.
Propagation is planned as deployment
Once a loop works on one fleet, the plan says roll it out, and the roll-out is scheduled as an IT activity with a duration per site. Gate 4 is not an IT activity: it is a qualification per target, with matching evidence and revalidation, and its duration is set by the target's own evidence accumulation. Programmes that discover this halfway through a multi-site plan usually lose more credibility than the loop ever earned.
Diagnosing an out-of-order programme
Plot how phases end against how change enters the controlled process. The quadrant names the next investment — and in three of the four cases it is not a better model.
Shelf-qualified
- Good evidence, no route into the controlled process
- Reports and marathon data exist; no ECN ever issued
- Fix: take one closed gate through MOC and let it be the template
Releasable
- Evidence and change control both in place
- The only quadrant from which phase 4 is reachable safely
- Fix: shorten the queue — one decision through consecutive gates
Roadmap theatre
- Phases are calendar periods and nothing is signed
- Common where the programme reports to IT rather than to manufacturing
- Fix: name the artefact and the signatory for one decision
Rubber-stamped
- MOC records exist; the evidence behind them is a date
- The most dangerous quadrant — the paperwork is protective in appearance only
- Fix: audit one recent ECN back to the evidence that justified it
Quantifying the benefit of an alternative scheduling approach remains a challenging task.
That sentence, from a vendor writing about its own successful production deployment, is worth more than most fab AI marketing. A running fab cannot hold conditions still for an A/B test: product mix moves, tools go down, demand changes. Which is precisely why the gate register asks for regime-covered evidence and approval-log analysis rather than for a controlled trial — the fab cannot supply the trial, and a phase model that demands one will simply be ignored.


