Silicon Wafer EngineeringAI-Driven Disruptions & Innovations
AI-driven fab resilience: absorbing disruption in silicon wafer engineering
AI-driven fab resilience is a wafer fab's ability to detect a disruption, bound the material at risk and re-plan capacity inside the clock that disruption starts, using models and the MES rather than a war room. It is measured in wafers held in time, Q-time windows kept and committed dates met.

Key takeaways
- Fab resilience is a clock problem, not a recovery-speed problem. Every disruption starts a physical clock — a Q-time window closing, a chamber losing its seasoning, a bulk-gas day tank draining, a customer commit date — and the only wafers you save are the ones you act on before that clock expires.
- The clocks differ by six orders of magnitude, so a single 'resilience programme' is meaningless. A voltage sag is decided in 200 milliseconds by hardware; an unplanned tool down is decided in the Q-time windows already open behind it; an export or substrate shock is decided in quarters. Each class needs its own mechanism and its own bound.
- One shared primitive underlies all of them: a material-at-risk query that can say, for any state change, exactly which lots are affected, how long each has left, and which qualified alternatives exist. Most fabs answer that question with a spreadsheet and a phone call, which is why detection improvements do not turn into saved wafers.
- Almost every published AI-for-resilience result is a detection or planning result obtained offline or in simulation, not an autonomous-recovery result on a live line. Treating the difference honestly is what separates a credible resilience programme from a self-healing-fab pitch.
- Future-readiness here is present-readiness. The fab that will be able to use an autonomous recovery loop in five years is the one that today has Q-time windows in the MES rather than in a document, a machine-readable qualification register, and a rehearsed rollback to manual dispatch.
Abbreviations used on this page
- MES
- Manufacturing execution system — the fab's system of record for lots, routes and steps
- FDC
- Fault detection and classification — tool-trace monitoring against known-good signatures
- SPC
- Statistical process control
- VM
- Virtual metrology — predicting a wafer measurement from tool sensor data instead of measuring it
- R2R
- Run-to-run control — feeding measured results back into the next run's recipe offsets
- WIP
- Work in progress — the wafers currently inside the line
- MAR
- Material at risk — the wafers that may have been affected between the last known-good point and the hold
- Q-time
- Queue-time constraint — the maximum permitted elapsed time between two linked process steps
- RTD
- Real-time dispatcher — the engine that decides which lot runs next on which tool
- CMMS
- Computerised maintenance management system
- UPW
- Ultrapure water — a fab utility whose interruption stops wet processing within minutes
- OT
- Operational technology — the tool, controller and facilities network, as distinct from corporate IT
Free · 8 questions · ~3 minutes
Score your fab's disruption response
Eight questions, one at a time, about three minutes. Answer them against one line — not the whole site — and we build your personalised resilience report: your rung, your score on each of the four dimensions, and the specific constraint standing between you and the next rung. Your result doubles as the baseline for every later claim about wafers saved.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised resilience report is ready
Tell us where to send it. Your rung appears on screen straight away, and the full report — dimension scores, the shock classes your current response leaves unbounded, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Reactive
The fab learns about a disruption when a tool alarms or a metrology gate fails, and the response is a war room working from exports.
Your next moveRecord the Q-time windows on one critical-path link as machine-readable data in the MES, and build the query that lists the lots inside them.
Stage 2 · Instrumented
The fab can see a shock and bound it quickly — material at risk is computed automatically — but a human still decides and executes every response.
Your next moveLet detection create the hold and let the dispatcher receive a Q-time-aware re-dispatch recommendation, with the previous rule one switch away.
Stage 3 · Contained
The fab acts inside the clock: bounded holds and Q-time-aware re-dispatch execute through the MES and the dispatcher, with a human approving rather than assembling.
Your next moveTreat second-source qualification coverage on the critical path as a resilience KPI with a target, and fund the tool time to move it.
Stage 4 · Rehearsed
The fab has already run its responses before the shock: scenarios are replayed against the real route, playbooks are pre-approved and capacity re-planning respects qualification and Q-time.
Your next movePut the rehearsal library on a schedule tied to route and tool-set changes, and treat scenario staleness as a tracked metric with an owner.
Stage 5 · Adaptive
A named, bounded set of responses executes without approval inside stated limits, and the fab's response policy is a versioned artefact updated from what each event taught it.
Your next moveTreat the response policy as a versioned, reviewable artefact under the same change control as a recipe, and monitor escalation rate as the signal that it has expired.
0 / 24
Shock detection
— / 6
Material at risk
— / 6
Recovery orchestration
— / 6
Rehearsal and learning
— / 6
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your response, and in fabs it is almost never detection. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your response, and in fabs it is almost never detection.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this checked against your own event history?
We take your last twelve months of unplanned downs, excursions and facilities events, replay them against your current route, and show you what your present response policy would actually have done. You keep the replay and the findings either way.
How the score maps to a stage
- 0–4 — Stage 1, Reactive. The fab learns about a disruption when a tool alarms or a metrology gate fails, and the response is a war room working from exports.
- 5–10 — Stage 2, Instrumented. The fab can see a shock and bound it quickly — material at risk is computed automatically — but a human still decides and executes every response.
- 11–15 — Stage 3, Contained. The fab acts inside the clock: bounded holds and Q-time-aware re-dispatch execute through the MES and the dispatcher, with a human approving rather than assembling.
- 16–20 — Stage 4, Rehearsed. The fab has already run its responses before the shock: scenarios are replayed against the real route, playbooks are pre-approved and capacity re-planning respects qualification and Q-time.
- 21–24 — Stage 5, Adaptive. A named, bounded set of responses executes without approval inside stated limits, and the fab's response policy is a versioned artefact updated from what each event taught it.
What AI-driven fab resilience is — and what starts the clock
A definition, the four things a fab has to do when a shock lands, and the clock that decides how many wafers survive it.
AI-driven fab resilience is the use of models and automated decisions to shorten the interval between a disruption occurring in a wafer fab and the fab acting correctly on the material that disruption put at risk. It is not disaster recovery, and it is not business continuity planning: those are about restoring a site. Resilience in this sense is about the wafers already inside the line when the shock lands, and about the commit dates attached to them.
Four things have to happen, in order, and they map onto four different engineering problems. The fab has to anticipate — know that something is going wrong before the measurement says so; bound — state precisely which wafers are affected and how long each has; act — hold, re-dispatch or re-plan inside that time; and adapt — change the policy so the next occurrence is cheaper. Most fab programmes invest almost everything in the first and almost nothing in the second, which is why better detection so rarely shows up as fewer scrapped wafers.
The reason ordering matters is that a fab is a re-entrant flow shop with an unusual property: many of its steps are joined by hard time limits. A wafer that has been cleaned must reach the next deposition step within a bounded window or the surface is no longer what the process assumed. Published work on scheduling realistic semiconductor flows describes the shape of the problem plainly — hundreds of operations, taking several months from lot release to completion (opens in a new tab), on machines that require product-specific setups and specialised maintenance. In a line like that, a disruption does not cost you downtime hours. It costs you the material that was mid-window when the disruption started, plus the cycle time of everything queued behind it.
Disruption absorbed against progress up the rungs
The curve is not linear. Absorption stays close to flat through rungs 1 and 2 — where most fabs are — because seeing a shock faster does not save material on its own. It inflects at rung 3, when the bounded response starts executing through the MES and the dispatcher rather than through a person. This is why fabs that measure progress in models deployed rather than in wafers held in time report activity without results.
Disruption absorbed without scrap or missed commits by stage
- Stage 1 · Reactive — 22% of operators. The fab learns about a disruption when a tool alarms or a metrology gate fails, and the response is a war room working from exports.
- Stage 2 · Instrumented — 37% of operators. The fab can see a shock and bound it quickly — material at risk is computed automatically — but a human still decides and executes every response.
- Stage 3 · Contained — 26% of operators. The fab acts inside the clock: bounded holds and Q-time-aware re-dispatch execute through the MES and the dispatcher, with a human approving rather than assembling.
- Stage 4 · Rehearsed — 12% of operators. The fab has already run its responses before the shock: scenarios are replayed against the real route, playbooks are pre-approved and capacity re-planning respects qualification and Q-time.
- Stage 5 · Adaptive — 3% of operators. A named, bounded set of responses executes without approval inside stated limits, and the fab's response policy is a versioned artefact updated from what each event taught it.
Curve shape: logistic, plotted from the stage data above. Distribution: Shape consistent with published fab capacity-planning and rescheduling research.
What happens in a fab in the first 15 minutes, the first shift and the first week after a shock
The same event, drawn across three time horizons. Each lane is a different engineering problem with a different owner, and the lane a fab is genuinely good at determines its rung. Most fabs are competent in the first lane, improvising in the second, and absent in the third — which is why the same event costs the same amount every time it happens.
- Data & feeds
- AI / model
- System-of-record action
- Where value leaks
- Human in the loop
The process, in words
- In the first fifteen minutes the only question that matters is bounding. A tool state change, an FDC trip or a facilities alarm has to become a precise list of affected lots, each with the time remaining on any Q-time window it sits inside, and a hold placed on the smallest defensible set. A fab that answers this with a person is placing wide holds late.
- In the first shift the question is where the material may go. The re-dispatch proposal is only safe if the qualification register is current, so this lane depends less on the model than on whether qualification data is written by the qualification process itself. The dispatcher approves each move with the previous rule one switch away, and the capacity re-plan and the customer commit dates follow from what was actually done.
- In the first week the question is whether the fab is now different. The event is replayed against the real route, the response policy is updated as a versioned artefact rather than as a post-mortem document, anything touching qualification or product change notification goes through the normal sign-off gate, and the scenario joins the rehearsal library so that the next occurrence is a selection rather than a design.
- The dashed return arrow is the whole argument. A fab becomes resilient when the third lane feeds the first — when the response to the next event is already computed, already approved and already bounded before the alarm goes off.
Step-by-step insights
- The tool state change is a fact; the material at risk is a query
- SEMI's E10 equipment-state model gives every tool an unambiguous vocabulary for productive, standby, engineering, scheduled and unscheduled downtime, and every modern fab logs those transitions. That is the easy half. The hard half is that the state change alone tells you nothing about consequence: which lots ran on that chamber since the last known-good point, which are inside an open Q-time window, and how long each has left. That join — tool history against lot history against route constraints — is the single most valuable object in a resilience programme, and it needs no machine learning at all. Fabs that build it first find that most of their later models become useful; fabs that build models first find the models have nowhere to land.
- Why a wide hold is the rational response to an unbounded question
- When nobody can say quickly which wafers are affected, holding everything plausible is not incompetence, it is the correct decision under uncertainty. The cost is real but diffuse: cycle time on lots that were never at risk, rework on lots that breach a window while waiting to be cleared, and an erosion of trust in holds that makes the next one slower to place. Narrowing the hold is therefore not primarily a yield intervention. It is a cycle-time and credibility intervention, and it is measured in lots released within the shift rather than in defects found.
- Q-time links: the constraint that turns downtime into scrap
- A queue-time constraint is a maximum permitted elapsed time between two process steps — post-clean to gate deposition, post-etch to metal fill, develop to hard bake. The chemistry does not care that a tool is down. When the window closes, the lot goes to rework if a rework path exists and to scrap if it does not. This is the mechanism by which an eight-hour tool down becomes a scrap event rather than a schedule event, and it is why the useful unit of fab resilience is a link rather than a tool. A fab that knows its windows as data can act inside them; a fab that keeps them in a process specification can only discover breaches afterwards.
- The qualification register is load-bearing infrastructure
- Every re-dispatch decision is a claim that a particular chamber may legally run a particular recipe today. In most fabs that claim is held in a spreadsheet updated after qualification events rather than by them, and it is wrong for days at a time — usually just after preventive maintenance, which is exactly when chambers are being brought back and re-dispatch is most likely. A response loop reading a stale register will confidently route material to a chamber that lost its qualification. The fix is unglamorous and organisational: make the qualification process the writer of record, and block dispatch on an expired entry rather than warning about it.
- Replay is how you find out your playbook has expired
- A wafer fab's route changes constantly — steps move, tool groups get re-dedicated, links are added when a process is tightened. A response playbook written eighteen months ago is a statement about a line that no longer exists, and nothing on the floor reveals the mismatch until a shock lands on the drifted part. Replaying real historical events against the current route is cheap, entirely offline, and reliably embarrassing the first time. Treat scenario staleness as a metric with an owner, tied to route and tool-set changes rather than to the calendar.
- The escalation path is what makes autonomy defensible
- Nothing that touches a recipe, a qualification decision, a product change notification or an export classification should ever execute unattended, regardless of how confident a model is. This is not conservatism for its own sake: those actions carry obligations to customers and to certification bodies that a fab cannot discharge with a log entry. Designing the escalation path first — enumerating what may execute and routing everything else to a named human — is what allows the small, bounded set of automated responses to be defended in an audit rather than switched off after the first surprise.
The five rungs of fab disruption response
For each rung: what the first hour of an event actually looks like, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps fabs there, and what leaving costs.
The ladder below measures one thing: what a fab does between a shock landing and the material at risk being correctly handled. It is deliberately not an organisational maturity model — a fab can be excellent at yield analytics and sit at rung 1 on disruption response, because the two capabilities share data but not deadlines. Score one line rather than a site, and score it on the median event rather than the one you handled best.
Each rung is written for a practitioner. The hallmarks describe observable conditions, the diagnostic signals are checks you can run against your own MES and trace history this week, and the anti-pattern is the specific mistake most often made trying to leave that rung. If your interest is the tool group rather than the event — chamber matching, drift between nominally identical chambers, and the control loops that close it — that is a different problem with a different ladder, covered in the autonomous wafer fleets page.
Select a rung
Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Reactive
22% of operators sit here
The fab learns about a disruption when a tool alarms or a metrology gate fails, and the response is a war room working from exports.
Rung 1 is not the absence of data. A modern fab at rung 1 is already drowning in it: every chamber streams trace, every step writes to the MES, facilities historians hold years of chiller and ultrapure-water telemetry. What is missing is any path from that data to a decision that has to be made inside a clock. When something goes wrong, the fab convenes people and reconstructs the world.
The tell is the shape of the first hour. Someone opens a query tool, someone else calls the module owner, a third person walks the floor to see which lots are physically where, and the answer to 'which wafers are affected' arrives after the answer would have been useful. By the time the hold is placed it is placed wide, because a wide hold is the only defensible response to an unbounded question.
This is an expensive stage to stay in and, unlike most maturity problems, its cost is lumpy rather than continuous. Nothing looks wrong for weeks; then one unplanned down on a tool with open Q-time links behind it produces a scrap and rework bill that would have funded three years of the instrumentation that would have prevented it.
In practice
The hold that was placed three lots too wide
An etch chamber drifts. The FDC system flags it, but the flag sits in a queue with two hundred others from the same shift, so the drift is confirmed at the next in-line critical-dimension measurement — four hours and eleven lots later. Nobody can say quickly which of those eleven lots ran on the affected chamber versus the other five in the group, so all eleven are held. Nine are eventually released after measurement; two are reworked. The two-lot answer took three days to reach.
What it looks like
- The first reliable signal of an excursion is a failed measurement, not a trace
- Material at risk is reconstructed by hand from MES exports and memory
- Q-time windows live in a process document, not in the MES
- Post-event reviews produce a slide deck nobody reads before the next event
Diagnostic signals you can check this week
- Ask how long it takes to list every lot processed on one chamber since a given timestamp. If the answer involves a person, you are here
- Ask where the maximum Q-time for a named link is recorded. A PDF or a process spec is the rung-1 answer
- Count how many holds in the last quarter were later released without a finding — wide holds are the signature of an unbounded query
- Ask what happened to the WIP the last time a bulk-gas or exhaust event stopped a module. If nobody can reconstruct it, it was never recorded
Anti-pattern · Buying a resilience platform
The instinctive fix is a programme: a data lake, a fab-wide anomaly platform, a control tower with a big screen. It reliably consumes a year without changing the first hour of a single event, because the platform's requirements are being guessed rather than observed. Instrument one clock end to end instead — one Q-time link, one tool group — and let the second and third clock tell you what is genuinely shared.
What holds you here
The question 'which wafers are affected, and how long do they have' cannot be answered faster than a person can answer it, so every response is placed wide and late.
Highest-leverage next move
Record the Q-time windows on one critical-path link as machine-readable data in the MES, and build the query that lists the lots inside them.
Cost of leaving
- Effort
- 2-4 months
- Team
- One data engineer, one process-integration engineer, part-time
- Risk
- Low — the work is read-only against MES and trace data, and nothing in production depends on it yet
- To next stage
- 2-4 months
If this is you, the next step is
A 2-week engagement: pick the link, map the data path, size the query.
Stage 2
Instrumented
37% of operators sit here
The fab can see a shock and bound it quickly — material at risk is computed automatically — but a human still decides and executes every response.
Rung 2 is where most fabs with a serious data programme arrive, and it is a genuine achievement: the first hour of an event stops being archaeology. The material-at-risk query answers while the tool is still coming down, the hold is placed on the right lots, and the post-event review has real timestamps in it rather than recollections.
It is also where the returns quietly stall, because seeing faster does not by itself save wafers. The bounded answer arrives on a screen and then waits for a human to act on it — a dispatcher who is currently handling three other things, or an engineer who has to be paged, found and briefed. The clock does not pause while that happens, and the clocks that matter most in a fab are the short ones.
Time spent at rung 2 is not neutral. Detection coverage tends to expand faster than response capacity, so the alarm count per shift rises, operators triage by habit, and the organisation learns that most alerts do not need action. That learning is rational and it is exactly what makes the fab slower to respond to the alert that does.
In practice
The alert that arrived during the shift handover
A deposition module trips at 06:52, eight minutes before handover. The detection system correctly identifies the affected chamber, the material-at-risk query returns four lots, two of them inside an open Q-time window with roughly forty minutes left. The alert lands in a queue that the outgoing shift has stopped working and the incoming shift has not started. At 07:31 the dispatcher sees it. One lot has already exceeded the window and goes to rework — a decision that a rule could have made at 06:53.
What it looks like
- Trace-based detection runs continuously and raises a bounded alert
- A material-at-risk query answers in seconds from MES and tool history
- Q-time windows are machine-readable and visible on the floor
- Response is still a person deciding and typing into the MES
Diagnostic signals you can check this week
- Measure the gap between the material-at-risk answer and the first action taken on it, across the last twenty events
- Count alerts raised per area per shift and compare it with the number of holds actually placed
- Ask a dispatcher to show you where a re-dispatch recommendation would appear in their normal screen. If it appears nowhere, integration is the gap
- Check whether any detection alert can create a hold automatically, or whether every hold is typed by a person
Anti-pattern · Adding more detection to fix a response problem
When bounded alerts do not turn into saved wafers, the reflex is to widen coverage — more chambers, more signals, more models — on the theory that earlier detection buys more time. It does not, because the binding constraint is the queue between the alert and the action. A moderately good detector wired into an automatic bounded hold saves more material than an excellent detector that pages a human at 06:52. Spend the next quarter on the write path.
What holds you here
The bounded answer reaches a screen rather than the system that executes, so the response is capped by whoever is free to act on it.
Highest-leverage next move
Let detection create the hold and let the dispatcher receive a Q-time-aware re-dispatch recommendation, with the previous rule one switch away.
Cost of leaving
- Effort
- 4-8 months
- Team
- One integration engineer, one ML engineer, a named process-integration owner and an MES change owner
- Risk
- Medium — the first automatic write into the MES needs an approved rollback and a drilled manual path
- To next stage
- 4-8 months
If this is you, the next step is
The rung 2 to 3 transition is the most common fab engagement. Typically 90 days.
Stage 3
Contained
26% of operators sit here
The fab acts inside the clock: bounded holds and Q-time-aware re-dispatch execute through the MES and the dispatcher, with a human approving rather than assembling.
Rung 3 is the first rung where the fab's response survives the absence of its best people. The path from signal to action has a destination, a schedule, an owner and an alarm, so a night shift with a new dispatcher produces roughly the same first hour as a day shift with the module owner on the floor. That is what containment means, and it is a narrower claim than it sounds — the fab is not recovering by itself, it is holding the line while people decide.
The character of the engineering changes here. Rung 1 and 2 problems are analytical; rung 3 problems are operational, and the discipline that solves them looks more like site reliability engineering than data science. What is the error budget for an automatic hold? What is the false-hold rate the fab will tolerate before it turns the loop off? What is the rollback, and when was it last exercised? These have well-known answers elsewhere and importing them is faster than rediscovering them.
The constraint that emerges is qualification coverage. A Q-time-aware re-dispatch is only useful if there is somewhere qualified to send the lot, and in most fabs the honest answer for critical layers is that there is not. The next investment is therefore usually not a better model but a deliberate widening of second-source qualification on the critical path — which is a process-integration and capacity decision with a real cost.
In practice
The night the loop held and the qual register did not
A chamber goes down at 02:10 with six lots inside an open post-clean window. The hold is automatic and correct; the re-dispatch proposal names two alternate chambers. One of them lost its qualification for that recipe eleven days earlier following a preventive-maintenance event, and the qualification register — a spreadsheet updated weekly — still says it is qualified. Two lots run on a chamber that should not have taken them. Nothing is scrapped, but the fab spends a fortnight on a deviation investigation, and the loop is switched off for six weeks.
What it looks like
- Detection creates a bounded hold automatically, on the smallest defensible set
- The dispatcher receives a re-dispatch proposal that respects Q-time and qualification
- Every proposal has a logged accept or override, with a reason code
- A tested one-switch fallback to the previous dispatch rule exists and has been used
Diagnostic signals you can check this week
- Check whether the qualification register the dispatcher relies on is written by the qualification process itself or maintained by hand
- Measure the false-hold rate: automatic holds later released with no finding, as a share of automatic holds
- Ask when the fallback to manual dispatch was last exercised deliberately, and on which shift
- Look at the override reason codes for a month. If most overrides say 'not qualified', qualification coverage is your constraint, not the model
Anti-pattern · Widening automation before widening qualification
Containment works, so the natural next step is to extend it to more tool groups and more shock classes. But the loop's usefulness is bounded by where a lot can legally go, and extending it across layers with single-chamber qualification produces confident recommendations that resolve to 'hold everything' — the rung-1 outcome with better logging. Widen second-source qualification on the critical path first, then extend the loop into the space that creates.
What holds you here
The response can only route material to somewhere qualified, and on critical layers most fabs have single-chamber qualification, so the loop resolves to a wide hold.
Highest-leverage next move
Treat second-source qualification coverage on the critical path as a resilience KPI with a target, and fund the tool time to move it.
Cost of leaving
- Effort
- 6-12 months
- Team
- Process integration, dispatch or industrial engineering, an MES owner, plus qualification capacity
- Risk
- Medium to high — qualification work consumes tool time that the fab is selling, so it competes directly with output
- To next stage
- 6-12 months
If this is you, the next step is
We map critical-path recipes against qualified chambers and price the gap in tool hours.
Stage 4
Rehearsed
12% of operators sit here
The fab has already run its responses before the shock: scenarios are replayed against the real route, playbooks are pre-approved and capacity re-planning respects qualification and Q-time.
Rung 4 changes the question from 'how fast can we respond' to 'have we already decided'. A fab cannot predict which tool will fail on which shift, but it can pre-compute the response to each of a small number of shock shapes and pre-approve the actions inside each one. When the event lands, the decision is a selection rather than a design, and the time that used to go into designing it goes into executing it.
The engineering here is replay rather than prediction. Take the last two years of actual unplanned downs, utility events and excursions, re-run each against the route and tool set as they exist today, and measure what the current response policy would have done. This is where most fabs discover that the policy they wrote eighteen months ago no longer matches the line: a step has moved, a tool group has been re-dedicated, a Q-time link has been added, and the playbook routes lots to a chamber that no longer exists.
Rung 4 is also where capacity re-planning stops being a monthly spreadsheet exercise and becomes part of the response. Published work on learned capacity planning is explicit that the interesting effects are the slow ones — bottlenecks forming gradually along the process flow rather than at the tool that failed — which is exactly the effect a human re-planning under pressure cannot see and a replayed scenario can.
In practice
The replay that found the playbook pointing at a decommissioned chamber
A fab replays fourteen months of unplanned downs on its photolithography cluster against the current route. The response policy performs well on eleven of the thirteen scenarios. On the other two it routes lots to a track that was re-dedicated to a different product family in the spring, and the replay shows those lots sitting in a queue long enough to breach two Q-time links. Nothing had gone wrong on the floor, because neither scenario had happened since the re-dedication. The finding cost an afternoon of compute and would have cost a week of scrap.
What it looks like
- Historical downs and utility events are replayed against the current route, not a generic model
- Each shock class has a pre-approved playbook naming who may do what without further sign-off
- Capacity re-planning proposals respect dedication, qualification and Q-time constraints
- Rehearsal coverage and scenario staleness are tracked as metrics
Diagnostic signals you can check this week
- Ask when the response playbooks were last replayed against the current route, not the route they were written for
- Count shock classes with a rehearsed scenario in the last six months. Most fabs can name one, usually tool down
- Check whether the capacity re-plan produced during the last real event respected qualification and dedication, or ignored them
- Ask who is permitted to execute each playbook action without further sign-off, and whether that is written down
Anti-pattern · Rehearsing the events you already handle well
Rehearsal budgets get spent on the familiar shock — the unplanned tool down — because the data is richest there and the exercise runs cleanly. The classes that actually hurt are the ones with no recent history in the fab: a bulk-gas interruption, a loss of confidence in MES data, an export-licence change that invalidates a routing. Deliberately rehearse the classes you have no data for, accept that the scenario is coarse, and treat the coarseness as the finding.
What holds you here
Response policies drift out of alignment with a line that changes every quarter, and nobody notices until a shock lands on the drifted part.
Highest-leverage next move
Put the rehearsal library on a schedule tied to route and tool-set changes, and treat scenario staleness as a tracked metric with an owner.
Cost of leaving
- Effort
- 12-18 months
- Team
- Industrial engineering, process integration, a simulation or scheduling engineer, facilities and security partners
- Risk
- Medium — the work is offline and safe, but it competes for the same scarce process-integration people as new product introductions
- To next stage
- 12-18 months
If this is you, the next step is
We take your last twelve months of events and re-run them against today's line.
Stage 5
Adaptive
3% of operators sit here
A named, bounded set of responses executes without approval inside stated limits, and the fab's response policy is a versioned artefact updated from what each event taught it.
Rung 5 is much narrower than the phrase 'self-healing fab' suggests, and the narrowness is the point. It is a specific, enumerated list of responses that may execute without a person — placing a bounded hold, re-dispatching a lot inside an open Q-time window to an alternate chamber that is currently qualified, releasing a lot when a measurement clears it — with everything outside that list escalating. Recipe changes, qualification decisions, product change notifications and anything with export-control implications are correctly held at human sign-off permanently.
By the time a fab arrives here the engineering is largely solved. The hard part is evidence: demonstrating to a customer's quality auditor, to a certification body under ISO 9001 or IATF 16949, or to an internal export-control officer that an action taken without a human was taken correctly, under a policy that was reviewed and versioned, and that the whole thing can be reconstructed months later. Treat the response policy with the same rigour as the model, because it is the artefact that gets examined.
Sustaining rung 5 is a governance discipline and it is the rung most likely to regress. Bounds derived from one shock class get quietly applied to another. A tool group is re-dedicated and the policy is not re-replayed. Escalation rate is the leading indicator worth watching: when the share of automated responses falling outside bounds rises, the world has moved outside the policy's validity, and a review is cheaper than the incident that would otherwise force one.
In practice
The bounded response set
A fab runs three unattended responses across two tool groups: automatic bounded hold on a confirmed trace excursion, automatic re-dispatch of Q-time-linked lots to a currently qualified alternate chamber, and automatic release when the confirming measurement clears. Roughly one response in fifteen escalates — usually because no qualified alternate exists. The escalation rate is monitored weekly; a rise triggers a review of qualification coverage rather than a widening of the bounds.
What it looks like
- An enumerated set of responses executes unattended within explicit bounds
- Every automated action carries a reconstructable audit trail
- The response policy is versioned, reviewed and change-controlled like code
- Anything touching qualification, product change notification or export status escalates by construction
Diagnostic signals you can check this week
- Ask whether the response policy has a version history and a named reviewer, or lives in a settings screen
- Ask when the kill switch was last exercised, on which shift, and what the fab lost while it was off
- Check whether escalation rate is monitored as a leading indicator or only reported after an incident
- Ask whether an auditor could reconstruct one specific automated hold from twelve months ago, end to end, from logs alone
Anti-pattern · Treating the response policy as configuration
Bounds get tuned in a UI with no review, no version history and no record of who changed what. It works until someone has to explain an automated hold taken eight months ago during a customer audit, at which point neither the policy nor the model that acted under it can be reconstructed. Version the policy, review changes through the same gate as a recipe change, and keep the trail — the cost is trivial next to a single unexplainable deviation.
What holds you here
Sustaining bounded autonomy is a governance problem: the constraint is audit evidence and change control, not engineering.
Highest-leverage next move
Treat the response policy as a versioned, reviewable artefact under the same change control as a recipe, and monitor escalation rate as the signal that it has expired.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team plus a standing forum with quality, process integration and export compliance
- Risk
- Concentrated — low frequency, high consequence, and regulatory rather than technical in nature
If this is you, the next step is
We stress-test the policy, the trail and the rollback against a real historical event.
Where wafer fabs actually sit on the rungs today
The distribution across the ladder, and why the rung 2 to rung 3 step is the largest single loss.
Most fabs with an active data programme sit at rung 2: they can see a shock quickly and bound it, and then a human has to act. The distribution below is weighted heavily toward that band, with a long thin tail at rungs 4 and 5 where responses are rehearsed before the event and a small enumerated set executes unattended.
Illustrative distribution of wafer fabs across the five rungs
Rung 2 is the mode and the plateau. The step from rung 2 to rung 3 — from a bounded answer on a screen to a bounded response executing through the MES — is the largest single transition loss on the ladder, and it is an integration and qualification problem rather than a modelling one.
Share of fabs
- 22% — 1 · Reactive
- 37% — 2 · Instrumented (the plateau)
- 26% — 3 · Contained
- 12% — 4 · Rehearsed
- 3% — 5 · Adaptive
The plateau is not a semiconductor-specific failure; it is the standard shape of an analytics programme that has not yet earned a write path into the system that executes. What is specific to wafer fabs is the character of the blocker. In most industries the constraint on acting automatically is approval; in a fab it is qualification. A re-dispatch is a statement about where a lot may legally go, and on critical layers the honest answer is often 'one chamber'. Widening second-source qualification costs tool hours the fab is currently selling, which is why the rung 2 to 3 step is usually funded as a capacity decision rather than as a software project.
The equipment and metrology vendors have been building toward the same conclusion from the other side. Process control, inspection and equipment intelligence portfolios from KLA (opens in a new tab), Lam Research (opens in a new tab) and Tokyo Electron (opens in a new tab) increasingly ship with the assumption that trace data leaves the tool continuously and lands somewhere that can act on it, and research institutes such as imec (opens in a new tab) publish on the same integration question. The trade press tracks the gap between what is technically available and what is actually wired in; Semiconductor Engineering's manufacturing coverage (opens in a new tab) is the most useful running record of it.
The disruption clock: what each shock class gives you before the loss is permanent
Six shock classes, the clock each one starts, the material at risk while it runs, the decision that must happen inside it, and the rung at which that decision stops needing a war room.
Every disruption in a fab starts a clock, and the clock — not the disruption — decides the loss. A voltage sag is over in a fraction of a second and its consequences are fixed by hardware that was specified years earlier. An unplanned tool down is decided in the Q-time windows already open behind it, which are typically measured in tens of minutes to a few hours. A substrate or export shock is decided in quarters. Treating these as one problem called 'resilience' is the most common structural mistake in fab continuity programmes, because it produces one governance forum, one dashboard and no mechanism that fits any of the six.
| Shock class | What starts the clock | Time you actually have | Material at risk | The decision inside the clock | Stops needing a war room at |
|---|---|---|---|---|---|
| Utility transient — voltage sag, chiller or exhaust trip, UPW interruption | The event itself; equipment ride-through ends within milliseconds | 200 ms to a few minutes | Every wafer in a chamber mid-recipe on tools that did not ride through | Nothing in the loop: ride-through is a hardware and facilities decision. The only AI-addressable job is reconstructing, within minutes, which lots aborted mid-step and which are recoverable | Rung 2 |
| Unplanned tool down | The equipment-state transition to unscheduled downtime | Minutes to hours — set by the Q-time windows already open behind the tool | Lots inside an open window, plus everything queued behind the tool on a dedicated path | Re-dispatch the Q-time-linked lots to a currently qualified alternate chamber, or route them to the rework path deliberately, before the window closes | Rung 3 |
| Process excursion | The first out-of-control trace, not the failed measurement | One to several lots — the gap between the excursion and the next measurement step | Every wafer processed between the last known-good point and the hold | Bound the material at risk with virtual metrology and trace history, and hold at the smallest defensible boundary rather than the widest safe one | Rung 3 |
| Consumable, bulk-gas or precursor interruption | The day tank, cylinder bank or bulk supply crossing its reserve level | Hours to days, set by on-site storage | The whole queue for every layer that depends on the constrained consumable | Re-sequence WIP so the constrained step is spent on the highest-value lots, and park the rest before their windows open rather than after | Rung 4 |
| OT security or data-integrity incident | Loss of confidence in MES or tool-network data, which usually precedes confirmation of the intrusion | Hours to weeks | All WIP whose routing, recipe and process history may be unverifiable | Fall back to a signed, offline copy of route and recipe state and re-establish what each lot actually ran, before restarting rather than after | Rung 4 |
| Supply, substrate or export shock | The commit date, not the shipment date | Weeks to quarters | The loaded plan and every customer commitment inside it | Re-plan capacity across qualified sites and tell customers which commitments move, before they ask | Rung 5 |
The clocks span roughly six orders of magnitude, from milliseconds to quarters, and that spread is the practical content of the table. The top row is genuinely outside AI's reach in the moment: semiconductor process equipment is specified against SEMI's F47 standard (opens in a new tab), which requires tools to ride through voltage sags to 50% of nominal for 200 milliseconds, 70% for half a second and 80% for a full second. Whether a given tool rides through a given sag was decided when it was bought and installed. What is not decided in advance is the next twenty minutes — which lots aborted mid-step, which are re-runnable, which just started a Q-time clock they cannot now finish — and that reconstruction is exactly the material-at-risk query.
One primitive underlies all six rows
Every row's decision depends on the same object: a query that takes a state change and returns the affected lots, the time remaining on each open window, and the currently qualified alternatives. Build it once against the MES, the tool history and the qualification register, and it serves the utility row, the tool-down row and the excursion row unchanged. It is also the single cheapest artefact on this page, because it contains no model.
The clock length sets the automation argument, not the risk appetite
A decision that must be made in ninety seconds cannot be made by a person who has to be paged, so the utility and tool-down rows argue for automation on physics rather than on efficiency. A decision that has three weeks — the supply row — should stay with people, because the value there is in judgement about customers and commitments, not in latency. Fabs that automate in clock order rather than in enthusiasm order end up with a small, defensible automated set.
Rework paths are part of the clock, not an afterthought
For several links, breaching the window means rework rather than scrap — a re-clean, a strip and re-coat, a re-measure. Whether a rework path exists, how much capacity it has and how much cycle time it costs changes the correct decision inside the clock. Fabs that model the rework path explicitly make better hold decisions than fabs that treat every window breach as a loss.
The bottom two rows are planning problems wearing a resilience badge
Consumable and supply shocks are decided by inventory policy, second-source qualification and the loaded plan, all of which are set months earlier. The AI contribution is in the re-planning: producing a capacity plan that respects dedication, qualification and Q-time in hours rather than weeks, so the customer conversation happens while options still exist.
Each of those decisions lives in a different system, owned by a different part of the fab, and measured by a different number. That is the second half of the centrepiece: knowing the clock is useless if you cannot say which system has to act inside it. The map below is how we scope resilience work with fab teams — it names the operating domain, the decisions worth wiring, the system of record that has to accept the write, the KPI that will move, and the rung at which the decision typically earns automation.
| Operating domain | Decisions worth wiring | System of record | KPI it moves | Earns automation at |
|---|---|---|---|---|
| Facilities and utilities | Ride-through classification, load shedding order, restart sequencing for wet and vacuum tools | Building management system and the facilities historian | Sags ridden through vs sags that stopped a tool; excursion-free chiller and UPW hours | Rung 2 |
| Equipment and maintenance | Down prediction, maintenance deferral under load, restart and re-qualification order | CMMS plus the SEMI E10 equipment state log | Unscheduled down hours, mean time to restore, deferred-maintenance backlog | Rung 3 |
| Process control and metrology | Excursion confirmation, material-at-risk bounding, sampling rate under suspicion | FDC and SPC engines plus the virtual metrology service | Material at risk per excursion, time to bound, false-hold rate | Rung 3 |
| Planning and dispatch | Q-time-aware re-dispatch, hot-lot re-prioritisation, capacity re-plan after a down | MES plus the real-time dispatcher and the planning system | Q-time violations, cycle-time X-factor, on-time delivery to commit | Rung 3 |
| Materials and supply | Consumable reserve triggers, second-source activation, substrate and mask allocation | ERP and supply planning plus gas and chemical inventory systems | Days of cover, second-source qualification coverage, expedite spend | Rung 4 |
| Security and data integrity | Anomalous OT traffic triage, recipe and route integrity verification, safe restart order | OT network monitoring and the MES audit log | Mean time to detect, recipe integrity verification time, verified-restart lead time | Rung 4 |
Two of these domains are usually orphaned. Facilities telemetry sits with a team that does not attend process meetings, so the fab cannot answer 'was that excursion a chiller event?' without a phone call — an answer a joined historian gives in seconds. And security sits with corporate IT, whose incident playbook assumes systems can be isolated and restarted, which is not a safe assumption for a tool network holding wafers mid-process. NIST's SP 800-82 guide to operational technology security (opens in a new tab) is explicit that OT environments invert the usual priority ordering — availability and safety ahead of confidentiality — and it is the right starting document for a fab writing its own restart sequence rather than inheriting IT's.
Real today, research-stage, or still speculation: bounding the self-healing fab
The claims made for autonomous, self-healing fabs, separated into what runs on production lines now, what published research actually demonstrates and in what setting, and what physics, qualification and export control will not allow.
Most of what is sold as a 'self-healing fab' is either twenty-five years old or has never left simulation, and the two categories need to be told apart before any of it goes into a capital plan. The oldest capability on the list — predicting a wafer measurement from tool sensor data rather than measuring it — was demonstrated on a production plasma etch reactor under a SEMATECH programme and is now routine. The newest — a fab that changes its own process and re-qualifies itself — has no published instance anywhere, and the reasons are regulatory and physical rather than computational.
| Claim | Status | What the published evidence actually says | What bounds it |
|---|---|---|---|
| Virtual metrology substituting for measurement on the recovery path | Real today, bounded | The SEMATECH J-88-E programme used virtual sensors on a Lam 9600 aluminium plasma etch reactor to predict recipe setpoints and wafer-state characteristics — line-width reduction and remaining oxide — from station, optical-emission and RF sensors, wafer to wafer, in a production setting | A prediction never clears a qualification gate. Measurement uncertainty is also larger than it looks: combining methods with a common-mean model can under-state total uncertainty by up to a factor of five |
| Anomaly prediction ahead of the FDC trip | Research-stage, promising | Forecast-based anomaly prediction on multivariate fab process traces held stable roughly 50 time points ahead, with a graph neural network outperforming a univariate forecaster at lower computational cost | The horizon is finite and short. Fifty samples is not a shift, so prediction buys minutes, not maintenance windows — useful for bounding, not for planning |
| Automatic detection and isolation of specific equipment faults | Research-stage, simulated | A hybrid model-plus-classifier scheme isolated two wafer-handler robot faults — an abrupt broken belt and an incipient arm tilt — outperforming purely data-driven methods, in a realistic simulation-based case study | Abrupt and incipient faults behave differently and need different detectors. No published live-line result; the simulation used derived motion models rather than fleet data |
| Learned capacity re-planning during a disruption | Research-stage, benchmark-scale | A deep-reinforcement-learning policy over a heterogeneous graph of machines and steps improved throughput and cycle time by about 1.8% each in the largest tested scenario, on Intel's Minifab model and the SMT2020 testbed. Separately, distributed multi-agent rescheduling with explicit risk assessment improved throughput while cutting computation against a centralised scheduler | Benchmark models are far smaller than a 40,000-wafer-start line, and the constraints that bind in a real fab — dedication, qualification, reticle availability, Q-time — are simplified in both |
| Generative models proposing recipes, masks or layouts during recovery | Research-stage, hard-constrained | Published analysis argues that in semiconductor manufacturing physically invalid generated samples are unusable rather than merely low quality, and that physics must be enforced by construction rather than by filtering afterwards | Lithography, transport, reaction and device physics. Any recipe change is also a change-control event with customer and certification consequences, whoever proposed it |
| A fab that detects, repairs and re-qualifies itself without human sign-off | Speculation | No published account of any fab qualifying a process change without human approval. Every documented autonomy claim stops at proposing, bounding or executing pre-approved actions | Customer qualification and product change notification obligations, ISO 9001 and IATF 16949 change control, and export-licence conditions attached to process technology and tooling |
1.8%
throughput gain and cycle-time reduction from a learned capacity-planning policy — on benchmark fab models, not a live line
arXiv 2509.15767
~50 steps
horizon over which published fab anomaly prediction stayed stable ahead of the event
arXiv 2510.20718
5×
how far a common-mean model can under-state hybrid-metrology uncertainty when inconsistent results are ignored
arXiv 2602.23131
Physically invalid samples are not merely low quality but unusable.
That sentence is the whole discipline for a fab. In most domains a generative model that is wrong 5% of the time is a useful model with a review step. In a fab, a proposed action that violates the process physics is not a lower-quality action — it is not an action at all, and the review step that catches it is the expensive part. This is why the credible programmes bound their models by construction: the re-dispatch engine cannot propose a chamber that is not qualified because the qualification register is a hard filter, not a scoring feature; the hold engine cannot release a lot whose window has closed because the window is a constraint on the route.
The uncertainty finding deserves its own attention, because it undermines the intuition that more measurement always tightens a bound. Work on hybrid metrology reports that an IEEE roadmap target of ±0.17 nm at 95% coverage becomes ±0.8 nm once dark uncertainty between inconsistent measurement methods (opens in a new tab) is accounted for with a random-effects model — and that ignoring it under-states total uncertainty by as much as five times. A material-at-risk boundary computed from an over-confident measurement is a boundary that will occasionally be too narrow, which is the one failure mode a resilience system must not have. Bound with uncertainty explicit, and prefer a slightly wider automatic hold to a narrow one you cannot defend.
The honest conclusion is unglamorous and it is the thesis of this page: future-readiness in fab resilience is almost entirely present-readiness. Every capability in the research column consumes the same three inputs — machine-readable Q-time windows, a current qualification register, and a joined trace-and-lot history with reliable timestamps. A fab that builds those has already bought most of the option value on whatever arrives next. A fab that skips them and buys a platform will find that the platform has nothing to stand on, and the equipment vendors and research institutes publishing in this space, from ASML (opens in a new tab) to imec (opens in a new tab), are describing the same precondition from their own side of the interface.
What fab resilience looks like in public
Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the operator's own published material.
The clearest public evidence for the argument on this page is not in AI announcements but in what large operators chose to build long before AI was involved: replication between sites, qualification breadth, and the ability to move capacity when one site cannot run. Those are resilience mechanisms, and the AI contribution in each case is narrower and more useful than the marketing around it — detecting divergence sooner, bounding material faster, and re-planning capacity in hours rather than weeks.
Three programmes read against the ladder
Outcomes as reported by the operators themselves. Card images are generated industry scenes from our own library, not photographs of these operators' facilities, and no endorsement is implied. Verify figures against the linked source before reusing them; we have not independently audited them.
IntelIDM and foundry · multi-site 300 mm network34
- Challenge
- Making capacity fungible between sites, so that a disruption at one facility does not strand the products qualified there. Transferring a process between fabs is normally slow because tiny differences in equipment set, parameters and facilities produce different results on the same recipe.
- Approach
- Intel has publicly described replicating process, equipment set and parameters between development and high-volume sites as an operating discipline rather than a per-transfer project, and reports on its manufacturing network, node ramps and capacity moves through its own newsroom.
- Reported outcome
- Intel publishes ongoing reporting on its manufacturing network and the movement of technologies across it; the practice is the clearest public example of replication treated as standing infrastructure rather than as a transfer exercise.
- What it shows about the curveFungible capacity is the strongest resilience mechanism a multi-site operator has, and it is built years before it is needed. AI does not create fungibility — it shortens the loop that detects divergence from the replicated baseline, which is the thing that quietly erodes it.
GlobalFoundriesSpecialty foundry · fabs in the US, Germany and Singapore34
- Challenge
- Customers on long-lifecycle parts — automotive, industrial, aerospace — need assurance that a device can still be built if one site or one region becomes unavailable, and that assurance has to be technical rather than contractual.
- Approach
- Qualifying platform technologies across a geographically distributed manufacturing footprint and publishing those qualifications, so that the resilience claim made to customers is footprint plus qualification coverage rather than inventory.
- Reported outcome
- GlobalFoundries reports platform qualifications and multi-site programmes through its own news channel; the pattern it makes visible is that second-source capability is a qualification asset that has to be maintained, not a contractual promise.
- What it shows about the curveSecond-source qualification coverage on the critical path is the resilience KPI most fabs do not track, and it is the ceiling on every automated re-dispatch decision. A response loop cannot send a lot anywhere the qualification register does not already permit.
SamsungIDM · memory and foundry, 300 mm23
- Challenge
- Utility-side shocks stop a line in seconds and the loss is decided in the following hours rather than by the outage itself. It is a matter of public record that a February 2021 grid emergency in Texas forced curtailment at semiconductor plants in Austin, Samsung's among them, and that restarting a large fab after an unplanned power loss takes far longer than restoring power.
- Approach
- Samsung reports on automation, analytics and smart-manufacturing programmes across its semiconductor operations through its own newsroom, including the application of machine learning to process and yield analysis.
- Reported outcome
- Samsung publishes ongoing reporting on smart-manufacturing and automation work in its semiconductor operations; the wider public record of the 2021 Austin curtailment illustrates the asymmetry between restoring a utility and restoring a line.
- What it shows about the curveHardening the utility is a facilities capital decision that AI does not touch. The AI-addressable part of a utility shock is entirely downstream: reconstructing within minutes which lots aborted mid-recipe, which are re-runnable and which just started a window they cannot finish.
Read together, the three make a single point. Resilience is bought in advance — in replication, in qualification breadth, in facilities specification — and spent in the first hour. Nothing on the ladder above rung 3 is achievable without the advance purchase, and nothing below rung 3 makes use of it when the hour arrives. Other operators publish in the same territory; Micron's newsroom (opens in a new tab) is another useful running record of a multi-site memory manufacturer's disclosures on manufacturing continuity.