Redefining Technology

Silicon Wafer EngineeringFuture of AI & Visionary Thinking

Decentralised autonomy in silicon wafer engineering: which fab decisions can safely leave the centre

Decentralised autonomy in a wafer fab is the deliberate placement of decision authority at the smallest scope that holds every constraint the decision touches — a chamber, a tool, a bay — instead of at a central scheduler. It is a governance design before it is an AI design, and the physics of the process decides most of it.

Illustrative cleanroom scene: process tools and operator stations with local decision displays beside each tool rather than one central console
Silicon Wafer Engineering · Future of AI & Visionary Thinking

Key takeaways

  1. Decentralised autonomy is a question about authority, not about intelligence. The interesting variable is not how clever the local model is but which decisions it is permitted to make alone, inside what bounds, and who finds out afterwards. A fab can buy every model on the market and change nothing about where decisions are made.
  2. A decision can only be devolved to a scope that holds all of its constraints. Chamber-level run-to-run correction is local physics and devolves cleanly. Reticle movement, Q-time rescue and hot-lot priority couple across bays, so devolving them without an arbitration layer does not distribute authority — it distributes conflict.
  3. The fab is already the most decentralised factory in manufacturing, and always was. Interlocks, equipment state machines, FDC holds and AMHS routing have executed without a central decision for decades. What has not moved in thirty years is the layer above them: the cross-fab production-control decisions that, as Flexciton puts it in its published material, no MES handles and no scheduler touches.
  4. The published record does not show negotiating agents running fabs. It shows centralised data architectures and control towers — Bosch’s Dresden fab describes fully connected systems within a centralised data architecture, GlobalFoundries publishes a global factory control tower — with autonomy at the execution layer beneath them. Multi-agent control is a mature research literature and a thin deployment record.
  5. Future-readiness here is almost entirely present-readiness. Nothing in a 2035 vision of a self-coordinating fab is reachable without three artefacts a fab can build this quarter: a written authority envelope per decision, a live constraint feed that tells the edge what it must respect, and a reconciliation trail good enough for a customer’s process-change audit.

Abbreviations used on this page

MES
Manufacturing execution system — the fab’s system of record for lots and steps
APC
Advanced process control
R2R
Run-to-run control — the per-run feedback loop on a tool
FDC
Fault detection and classification
EDA
Equipment Data Acquisition — the SEMI “Interface A” standards suite
GEM
Generic Equipment Model (SEMI E30, carried over SECS-II)
AMHS
Automated material handling system — the overhead carrier transport
FOUP
Front-opening unified pod — the 300 mm wafer carrier
WIP
Work in progress — lots currently on the line
Q-time
Queue time — the maximum permitted interval between two linked process steps
SPC
Statistical process control
PCN
Process change notification — the contractual notice a fab owes a customer before a qualified process changes

Free · 8 questions · ~3 minutes

Score where your fab’s decision authority actually sits

Eight questions, one at a time, about three minutes. Answer them and we build your personalised devolution report — your rung on the ladder, your score on each of the four dimensions, and the specific constraint holding authority at the centre — and send it to your inbox. Score one bay or one toolset rather than the whole fab; the answers are sharper.

0 of 8 answered

Question 1 of 8Decision rights

Is there a written statement of which decisions a bay or toolset may make without asking?

Devolution is a delegation of authority. Where the delegation is unwritten, it does not exist — it is a habit that changes with the shift supervisor.

How the score maps to a stage
  • 04 — Stage 1, Central command. Every non-safety decision is made on the central plane’s cadence; the bay and the tool execute and report.
  • 510 — Stage 2, Sensing edge. Local intelligence can see but cannot act: tool and bay models raise signals that a human or the central plane must act on.
  • 1116 — Stage 3, Bounded local authority. Named decisions execute locally inside a written envelope, with global constraints fed down live and every local action reconciled into the record.
  • 1721 — Stage 4, Negotiated autonomy. Local decision-makers hold rights over scarce shared resources and obtain them through an explicit arbitration protocol rather than a central plan.
  • 2224 — Stage 5, Federated line. Authority is federated across bays and sites under versioned policy, with defined and rehearsed behaviour when the central plane is unreachable.

What decentralised autonomy in a fab actually means

A definition, the difference between devolving execution and devolving authority, and the three planes a fab decision can be made on.

Decentralised autonomy in a wafer fab means moving the authority to decide — not the ability to compute — down to the smallest scope that holds every constraint the decision touches. That is a deliberately narrow definition, and the narrowness is the point. A tool that runs a large model locally has decentralised computation. A tool that may place a hold without asking has decentralised authority. Only the second changes how the fab behaves, and only the second requires anyone to write anything down.

The distinction matters because fabs have been quietly decentralising execution for thirty years while centralising authority just as steadily. Interlocks abort a run without consulting anything. Equipment state machines transition on their own. Automated material handling routes carriers continuously without a person choosing a path — GlobalFoundries publishes (opens in a new tab) that its overhead transport delivers as many as 10,000 carrier moves a day in a single fab. Meanwhile the decisions above that layer — what to start, what to hold, what to expedite, what to measure — have concentrated into central planning and dispatch systems, and into the heads of the people who operate them.

So the vision usually described as “the decentralised fab” is not a break with fab history; it is the resumption of an old trend at a layer that has been stuck. It also has a hard ceiling that no amount of model capability moves, because a decision cannot be devolved to a scope that cannot see its own constraints. The diagram below is the whole argument in one picture: three planes, and the observation that what travels down from the centre in a devolved fab is a constraint, not an instruction.

Where a fab decision is actually made — the three planes

The vertical axis is authority, not time. At rung 1 the central plane sends instructions and the bay plane is empty. From rung 3 the central plane sends envelopes and constraints instead, the bay plane decides inside them, and the central plane’s remaining jobs are arbitration, invariants and the record. The tool plane has been autonomous since long before anyone called it that.

  • Data & feeds
  • Human in the loop
  • AI / model
  • System-of-record action
  • Where value leaks

The process, in words

  • The central plane holds the fab state of record — WIP, queue-time clocks, dedication rules, the reticle map — and, in a devolved fab, stops using it to write instructions. Instead it issues two things downward: a versioned authority envelope stating what each scope may decide and within what bounds, and a live constraint feed carrying the global facts a local decision must respect. It keeps three jobs for itself: arbitrating contended resources, enforcing invariants no local scope may cross, and reconciling every devolved action back into the record.
  • The bay plane is where rungs 3 to 5 actually live, and at rungs 1 and 2 it is empty. A bay decision agent takes tool traces from below and constraints from above and decides in seconds rather than on the central re-plan cadence — sampling, holding, releasing, routing. When it needs something scarce it asks the arbiter for a bounded, expiring lease rather than escalating to a person. When a decision falls outside its envelope it escalates, and the escalation is the safety valve rather than a failure of the design.
  • The tool plane has been autonomous since long before the word was fashionable. Equipment controllers execute moves and hold run-to-run offsets; interlocks and equipment safety systems can veto anything above them, locally and instantly, and are never part of a negotiation. Any architecture that proposes to make safety interlocks smart, remote or negotiable has misunderstood which decisions were already decentralised and why.
  • Read across the diagram and the shape of the whole page appears: what moves down is constraints, what moves up is evidence, and the central plane trades the power to instruct for the power to bound. That trade is the entire content of decentralised autonomy in a fab.
Step-by-step insights
Fab state of record — the thing you cannot devolve
Every devolved decision is a bet that the local scope knows enough. The fab state of record is what makes that bet reasonable: one authoritative version of where lots are, which queue timers are running, which tools are qualified for which layers, and where each reticle physically is. Fabs that try to devolve before this exists end up with bays deciding confidently against divergent pictures of the same line, which is worse than a slow central plan because the disagreement is invisible. Nothing on the ladder above rung 2 is reachable without a single, current, authoritative state.
The authority envelope — a delegation, written like one
An envelope names the decision, the scope permitted to make it, the numeric bounds, the preconditions that suspend it, the budget it consumes and the revert. It reads like a delegation of authority in an organisation because it is one, and the useful test is whether a new shift supervisor could read it and know exactly what the bay may do at 3am without calling anyone. Envelopes are short — a page — and their value comes entirely from being versioned and reviewed, because the question a customer audit asks is not what the bound is now but what it was then.
The constraint feed — where devolution usually fails
Devolving a decision blinds the deciding scope to everything outside it, so the global facts it must respect have to be pushed down continuously with an explicit freshness guarantee. Queue timers are the canonical example and the canonical failure: a bay choosing what to run next cannot know a lot is forty minutes from a queue-timer breach unless the clock is fed to it. Flexciton has published technical material specifically on why queue timers defeat conventional schedulers, and the reason generalises — Q-time couples two steps that no single scope owns. Stale feed, suspended decision: that rule is cheap to implement and it is the difference between devolution and an unmonitored guess.
The arbiter — small, boring and load-bearing
The arbiter is deliberately not a planner. It does one thing: grants bounded, expiring leases on genuinely scarce shared resources, and refuses with a reason and an expiry so the refused scope can plan around it. Leases degrade well under failure — an expired lease returns the resource automatically, and a partitioned agent cannot hold one indefinitely. Keeping the arbiter small also keeps it auditable: its entire behaviour is a log of grants and refusals, which is a far easier artefact to defend than a full optimisation trace.
Escalation as a designed output, not an exception
The most common design error in devolved systems is treating escalation as failure and tuning it toward zero. The escalation rate is a signal: it measures how often the world falls outside the envelope, and a rising rate means conditions have moved beyond the envelope’s validity — a new product mix, a requalified toolset, a changed dedication map. Fabs that suppress escalations lose their early-warning system and discover the drift as an excursion instead. Budget for escalations, staff for them, and watch the trend rather than the count.
Reconciliation — the audit trail as a by-product
Every devolved action must land in the MES record with its inputs, its envelope version and, where relevant, the arbiter’s grant. At rung 3 this looks like engineering hygiene. By rung 5 it is the artefact a customer’s change-control audit examines, because in a qualified process the rule that decided is part of the process. Building reconciliation as a by-product of the action path — rather than as a reporting job that runs later — means the evidence accumulates for free and the audit becomes an export rather than a project.

One consequence follows immediately and it is worth stating before the ladder, because it reframes the whole topic. If what the centre sends down is a constraint rather than an instruction, then the quality of a fab’s decentralisation is bounded by the quality of its central state. Decentralised autonomy is not the opposite of a strong central plane — it is what a strong central plane buys you. That is why the fabs with the most publicly documented automation are also the ones describing, in their own material, a highly centralised data architecture.

The five rungs of the devolution ladder

For each rung: what it looks like on the floor, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps fabs there, and what leaving costs.

The ladder measures one thing only: how much decision authority has left the centre, safely. It is not a measure of how much AI a fab runs, and a fab can score highly on every other maturity model in this knowledge base while sitting at rung 2 here. Each rung below is written for a practitioner — the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that rung.

Decision authority safely devolved, against time on the ladder

The curve is flat through rungs 1 and 2 because sensing is not deciding: a fab can add local models for years and devolve nothing. It inflects at rung 3, when the first written envelope makes a bay’s decision real, and again at rung 4, when arbitration lets several devolved decisions coexist without a person resolving them. It also has an asymptote, which most visions of the autonomous fab omit — a permanent set of decisions that correctly stay central forever.

Decision authority safely devolved by stage

  • Stage 1 · Central command — 31% of operators. Every non-safety decision is made on the central plane’s cadence; the bay and the tool execute and report.
  • Stage 2 · Sensing edge — 42% of operators. Local intelligence can see but cannot act: tool and bay models raise signals that a human or the central plane must act on.
  • Stage 3 · Bounded local authority — 19% of operators. Named decisions execute locally inside a written envelope, with global constraints fed down live and every local action reconciled into the record.
  • Stage 4 · Negotiated autonomy — 7% of operators. Local decision-makers hold rights over scarce shared resources and obtain them through an explicit arbitration protocol rather than a central plan.
  • Stage 5 · Federated line — 1% of operators. Authority is federated across bays and sites under versioned policy, with defined and rehearsed behaviour when the central plane is unreachable.

Curve shape: logistic, plotted from the stage data above. Distribution: Illustrative shape, consistent with published material on the path to fab autonomy.

Select a rung

Every rung’s full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Central command

31% of operators sit here

Every non-safety decision is made on the central plane’s cadence; the bay and the tool execute and report.

Rung 1 is not primitive. A great many high-yielding fabs run here deliberately, and the reasoning is sound: a central plane can see the whole line, so it can trade off Q-time links, bottleneck loading and customer commitments in one place, and one place is far easier to qualify than fifty. Centralisation is a legitimate architecture, not a failure state.

What makes rung 1 costly is not centralisation but cadence. The central plane re-decides on a fixed rhythm — every fifteen minutes, every half hour, every shift — while the fab throws up events on the timescale of seconds. A chamber goes down, a carrier arrives early, a metrology result comes back out of trend, and the fab spends the remainder of the interval executing a plan that was true when it was computed and is not true now. The gap between those two clocks is where every rung-1 fab loses time, and it is invisible on any dashboard because the plan was followed correctly.

The tell is not technological, it is documentary. Ask for the list of decisions a bay may make without asking. At rung 1 there is no such list — not because nobody has written it down, but because nobody has ever needed to: every answer is “ask production control”. That absence is the thing that has to change first, and it costs nothing but argument.

In practice

The twenty-minute-old dispatch list

A 200 mm fab issues a fab-wide dispatch list on a twenty-minute cycle. Nine minutes in, an etch chamber goes down on an FDC alarm. The list still shows four lots routed to it. The bay technician knows within a minute; production control learns when the next cycle recomputes; the lots sit. Nothing in the system is broken, no rule was violated, and the fab has just spent eleven minutes executing a plan built for a fab that no longer exists.

What it looks like

  • The dispatch list is the decision — tools and operators execute it as issued
  • Local models, where they exist at all, are reporting artefacts rather than actors
  • The only decisions genuinely made at the edge are interlocks and equipment state transitions
  • Nobody can produce a written statement of which decisions a bay is allowed to make

Diagnostic signals you can check this week

  • Ask for the written list of decisions a bay may make alone. If there is none, you are here
  • Measure the interval between central re-plans, then measure the mean time between events that invalidate a plan. Compare the two numbers
  • Watch what happens on a tool-down: count the minutes before the plan reflects it
  • Ask a shift supervisor how a priority conflict is resolved at 3am. If the answer names a person rather than a rule, authority is undocumented

Anti-pattern · Buying agents before writing the decision rights

The instinctive response to rung 1 is to procure something agentic — an edge platform, a multi-agent scheduler, an on-tool inference stack. It fails in a predictable way: the software arrives able to decide, and the organisation has no statement of what it is allowed to decide, so it is deployed in advisory mode “until we are comfortable”. Advisory mode is rung 2, and fabs stay there for years. Write the decision rights first. It is a document, it costs a fortnight of arguments between industrial engineering, process and quality, and it is the artefact every later rung depends on.

What holds you here

No decision has a written owner or a written bound, so nothing can be devolved without an argument that has to be had from scratch every time.

Highest-leverage next move

Enumerate the fab’s recurring decisions and, for each, record who makes it today, on what cadence, and what it couples to. Do not automate anything yet.

Cost of leaving

Effort
4–8 weeks
Team
One industrial engineer, one production-control lead, one quality representative — part time
Risk
Very low — the output is a document, and nothing in production changes
To next stage
3–6 months

If this is you, the next step is

A two-week working session: enumerate the decisions, agree the scopes, mark what is already devolved.

Draft your first decision-rights register

Stage 2

Sensing edge

42% of operators sit here

Local intelligence can see but cannot act: tool and bay models raise signals that a human or the central plane must act on.

Rung 2 is the mode of the industry and the most comfortable place on the ladder to get stuck. Local models exist, they are often genuinely good, and they produce a stream of signals nobody disputes. What they do not have is standing. The model can observe that chamber 4’s endpoint trace has shifted; it cannot hold chamber 4, because holding a chamber is a capacity decision and capacity decisions belong to production control.

The structural problem is that rung 2 optimises the wrong quantity. Every quarter of effort goes into raising the sensitivity and specificity of local detection, because that is the tractable engineering problem, and none goes into the question that determines whether detection matters: how long between the signal and the action, and who is in the loop. A fab can halve its false-alarm rate and change nothing about the time-to-action, because the time-to-action is set by shift handover and a ticket queue.

Time spent at rung 2 is not neutral. Engineers learn that local signals are advisory, so they stop reading them; production control learns that the models cry wolf, so it discounts them; and the next proposal for local authority is argued against that memory. The organisational antibodies are the real cost, and they are why a fab that has sat at rung 2 for five years is usually harder to move than one still at rung 1.

In practice

The alarm that everyone had already seen

A diffusion bay runs an FDC model that flags a slowly drifting temperature profile on one furnace tube. It fires on a Tuesday. The alarm goes to a queue reviewed at the Thursday process meeting, where an engineer confirms the drift and raises a work request; the tube is corrected the following Monday. Six days of wafers went through with a known, detected, unacted-upon deviation. Nobody behaved incorrectly. The model’s accuracy was never the constraint.

What it looks like

  • FDC, health and quality models run at the tool or the bay and generate alarms
  • Every model output terminates in a screen, a ticket or a chart
  • Model quality is debated far more often than model authority
  • The escalation path from a local signal to an action is human and undocumented

Diagnostic signals you can check this week

  • Take one local alarm and time it end to end: signal raised, human aware, action taken. Do it for three alarms and take the worst
  • Ask what the model is allowed to do if nobody responds. If the answer is “nothing”, authority is the gap, not accuracy
  • Count how many local models write anything at all into the MES. At rung 2 the answer is usually zero
  • Check whether alarm volume per shift exceeds what one engineer can triage. Above that line, the signals are decorative

Anti-pattern · Improving detection to earn authority

When local signals are ignored, the reflex is to make them better, on the theory that trust follows precision. It does not, because the bottleneck is not belief — it is standing. A model with an 80% precision that is permitted to place a two-hour engineering hold changes more wafers than a 97%-precision model that can only raise a ticket. Spend the next quarter negotiating a narrow, bounded, revertible authority for one existing model, and revisit detection quality once you can price a false alarm in lost tool hours rather than in reviewer irritation.

What holds you here

Local intelligence has no standing: every signal must pass through a human queue, so the fab’s response time is set by shift rhythm rather than by the physics.

Highest-leverage next move

Pick one existing local model and negotiate the narrowest possible authority envelope for it — one decision, one bay, a hard budget, and a one-switch revert.

Cost of leaving

Effort
3–6 months
Team
One integration engineer, one process engineer, a named quality owner for the envelope
Risk
Medium — the first bounded local action needs a documented revert and a qualification review
To next stage
6–12 months

If this is you, the next step is

We take a model you already run and design the narrowest envelope that makes it act.

Give one existing model a bounded authority

Stage 3

Bounded local authority

19% of operators sit here

Named decisions execute locally inside a written envelope, with global constraints fed down live and every local action reconciled into the record.

Rung 3 is the first rung at which anything has genuinely been decentralised, and the artefact that makes it real is not the model — it is the envelope. An envelope is a short, versioned, reviewable document that says: this bay may place engineering holds on this toolset, up to four hours, up to two chambers concurrently, provided no lot in the bay has less than ninety minutes of Q-time remaining, and every hold is written to the MES within thirty seconds. That is the whole idea. It reads like a delegation of authority because that is exactly what it is.

The second artefact is the constraint feed, and it is the one fabs consistently underestimate. Devolving a decision means the deciding scope no longer sees the whole line, so anything global it must respect has to be pushed to it, live, with a freshness guarantee. Q-time remaining is the canonical case: a bay agent choosing what to run next cannot know that the lot in front of it is forty minutes from a queue-timer breach unless the clock is fed to it. A devolved decision made against a stale constraint is not autonomy, it is an unmonitored guess with a fast response time.

The character of the work at rung 3 is closer to operations engineering than to data science. What is the freshness SLA on the constraint feed? What happens if it goes stale — does the agent fail open, fail closed, or fail to the previous rule? Who is paged? These are boring, answerable questions with well-established answers, and answering them is what distinguishes a fab that devolves safely from one that devolves loudly.

In practice

The four-hour hold envelope

An etch bay is granted authority to hold a chamber for up to four hours on an FDC excursion, capped at two chambers at once, suspended automatically whenever the bay’s queue-time exposure exceeds a stated threshold. In the first quarter the bay used it forty-one times. Nine of those holds were later judged unnecessary by process engineering — and the argument that settled the review was not whether the model was right, but that the total capacity cost of the nine false holds was a fraction of the cost of one excursion caught six days late.

What it looks like

  • A versioned envelope states which decisions a bay may make and within what numeric bounds
  • The constraint feed pushes live Q-time remaining, dedication and budget state to the edge
  • Every devolved action lands in the MES record with its inputs and its envelope version
  • There is a tested revert to the previous decision source, and it has been exercised

Diagnostic signals you can check this week

  • Ask to see the envelope. If it is a slide rather than a versioned document with an owner, it is not an envelope
  • Check the constraint feed’s freshness alarm. If there is none, the devolved decision is running blind on its worst day
  • Sample ten devolved actions from last month and try to reconstruct each from the MES record alone
  • Ask when the revert was last exercised. “We have never needed to” means it is untested, not that it works

Anti-pattern · Widening the envelope by ticket

Envelopes creep. A hold cap of four hours becomes eight because a specific incident made eight seem sensible, and the change is made in a configuration screen by a person with the access to make it. Six months later nobody can say what the bound was in March, which is precisely the question a customer’s process-change audit asks. Treat the envelope as a controlled document: versioned, reviewed by the same forum that reviews a recipe change, with the version identifier stamped on every action taken under it.

What holds you here

Each devolved decision is negotiated, documented and monitored separately, so the fourth costs as much as the first and contention between them is still resolved by people.

Highest-leverage next move

Identify the shared resources your devolved decisions are already competing for, and give them an explicit arbitration primitive instead of a phone call.

Cost of leaving

Effort
6–12 months
Team
Integration engineer, process engineer, industrial engineer, quality owner, plus a standing review forum
Risk
Medium — the exposure is qualification and change control, not model error
To next stage
12–24 months

If this is you, the next step is

One decision, one bay: the bounds, the constraint feed, the revert and the audit record.

Design an authority envelope

Stage 4

Negotiated autonomy

7% of operators sit here

Local decision-makers hold rights over scarce shared resources and obtain them through an explicit arbitration protocol rather than a central plan.

Rung 4 is where decentralisation stops being an efficiency argument and becomes an architecture. Once several bays hold real authority, they start wanting the same things at the same time: the same reticle, the same metrology slot, the same technician, the same batch furnace load, the same overhead vehicle on the same track segment. At rung 3 those collisions are resolved by escalation, which is a polite word for a person on a radio. At rung 4 they are resolved by a protocol.

The protocol does not have to be exotic. The useful primitive is a lease: a scope requests a contended resource, an arbiter grants it for a bounded period under stated conditions, and the lease expires whether or not it was used. Leases are attractive precisely because they degrade well — an expired lease returns the resource, a partitioned agent cannot hold one forever, and the arbiter’s decisions are a small, auditable log rather than a full plan. The academic lineage here is long: holonic manufacturing control has been formalising exactly this trade-off between local autonomy and global coherence for over two decades.

What is genuinely hard at rung 4 is not the protocol but the objective. Local scopes optimise what they can see, and what they can see is their own tool group. The published evidence that this is a real effect rather than a worry is unusually direct: recent work analysing joint versus modular learning for job-shop scheduling with transport resources measures the coordination gap between them explicitly. A fab arriving at rung 4 has to decide what the arbiter is protecting — line balance, Q-time integrity, on-time delivery — and encode that as invariants the local scopes cannot bid their way around.

In practice

The reticle nobody could book

Two litho bays in a fab with devolved lot selection both hold work needing the same reticle within the same hour. Under rung 3 the conflict surfaces as two engineers escalating simultaneously and a supervisor picking. Under rung 4 both bays request a lease; the arbiter grants ninety minutes to the bay whose lots carry the tighter Q-time exposure, refuses the other with a stated reason and an expiry it can plan around, and logs both. The second bay does not wait on a person; it schedules around a known return time.

What it looks like

  • Contended resources — reticles, metrology capacity, AMHS segments, qual slots — are allocated by lease rather than by plan
  • The central plane’s job has shifted from instructing to enforcing invariants and resolving conflict
  • Priority is a negotiable, expiring claim rather than a static flag on a lot
  • Contention itself is measured: wait time per resource, lease starvation, escalation rate

Diagnostic signals you can check this week

  • List your genuinely scarce shared resources and ask, for each, what allocates it. If the answer is a human, you are below rung 4
  • Check whether priority claims expire. Non-expiring priority is how a fab accumulates twelve simultaneous hot lots
  • Measure lease wait time per resource and look for starvation — one scope repeatedly refused is a broken objective, not bad luck
  • Ask what the arbiter protects. If nobody can name the invariants, local optimisation will eventually eat the line

Anti-pattern · Letting the agents set the objective

The seductive version of rung 4 is emergent: give every scope a reward, let them bid, and trust that good global behaviour appears. It does not, reliably, and the failure mode is quiet — the line balances worse, cycle time variance rises, and the cause is distributed across a hundred locally defensible choices. Global objectives belong in the arbiter as hard invariants, not in the agents as soft incentives. Bidding decides who gets a scarce resource; it must never decide whether a Q-time link may be broken.

What holds you here

Arbitration works inside one bay boundary and one shift, but nothing defines behaviour when the central plane is unreachable or when a second site is involved.

Highest-leverage next move

Define degraded-mode behaviour explicitly — what each scope may decide when it cannot reach the arbiter, and for how long — and rehearse it.

Cost of leaving

Effort
18+ months
Team
Platform team, industrial engineering, production control, quality, plus a standing arbitration-policy forum
Risk
Higher — the exposure is systemic behaviour under contention, which is hard to test and easy to miss
To next stage
24+ months

If this is you, the next step is

Which resources need leases, what the invariants are, and how contention is measured.

Design your arbitration layer

Stage 5

Federated line

1% of operators sit here

Authority is federated across bays and sites under versioned policy, with defined and rehearsed behaviour when the central plane is unreachable.

Rung 5 is narrower than the phrase “autonomous fab” suggests, and it is worth being precise about what it is not. It is not a fab without people, and it is not a fab in which everything is decided locally. It is a fab in which a specific, enumerated set of decisions executes under policy that is distributed rather than centrally instructed, in which the failure of the central plane degrades the fab gracefully instead of stopping it, and in which the whole arrangement can be explained to a customer’s auditor.

The distinguishing discipline is partition behaviour. Every distributed system eventually faces the question of what to do when the parts cannot talk to each other, and a fab’s answer must be per decision, not global. Interlocks keep working — they always did. A bay may keep placing holds under its existing envelope for a bounded period, because a hold is conservative. It may not keep granting itself reticle leases, because leases are the thing the arbiter exists to prevent double-issuing. Writing that table down, per decision, is most of the work of rung 5.

The second discipline is evidence, and it is the one that will actually gate a real fab. A qualified process serving automotive or medical customers carries change-control obligations: the rule that decided is part of the process, so the envelope version, the constraint values it saw and the arbiter’s grant are all part of the record. GlobalFoundries publishes its quality and certification frame openly, and every fab operating to comparable standards faces the same question — not “was the AI right?” but “can you reconstruct, eighteen months later, what it was permitted to do and why it did this?”. Rung 5 is sustainable only where the answer is yes by construction.

In practice

The forty-minute partition

A multi-site operator loses the link between one fab’s bay network and its central plane for forty minutes during a network change. Nothing stops. Interlocks are unaffected. Bays continue placing holds under envelopes cached with an explicit expiry, and stop issuing anything that consumes a shared resource. Two reticle requests queue rather than resolve. When the link returns, the reconciliation job replays forty minutes of local actions into the record and flags three that fell outside the cached envelope for human review. The event is a rehearsal, not an incident, because the behaviour was specified.

What it looks like

  • Envelopes and invariants are distributed as versioned policy, not configured per site
  • Degraded-mode behaviour is specified per decision and exercised on a schedule
  • Cross-site decisions are federated by policy, with data-sharing constraints made explicit
  • Any devolved action from any site is reconstructable to customer-audit standard

Diagnostic signals you can check this week

  • Ask for the per-decision partition table. Its absence means degraded mode is whatever the code happens to do
  • Check when the partition drill was last run. Annual is thin; quarterly is defensible
  • Pick one devolved action from a year ago and reconstruct the envelope version and constraint values it ran under
  • Ask how a policy change propagates to a second site, and how you would prove which version was live at a given time

Anti-pattern · Treating the policy as configuration

At rung 5 the policy — envelopes, invariants, partition rules — is the most consequential artefact in the system, and it is the one most likely to be edited in a settings screen with no review, no version history and no record of who changed what. The system works until someone has to explain an action taken eight months ago, at which point neither the bound nor the rule that produced it can be recovered. Policy is code. Review it like code, version it like code, and stamp its version on every action taken under it.

What holds you here

Sustaining federation is a governance and evidence problem — the constraint is customer change control and policy versioning, not engineering.

Highest-leverage next move

Treat envelopes, invariants and partition rules as one versioned, reviewable policy artefact, and rehearse the partition on a schedule.

Cost of leaving

Effort
Continuous
Team
Platform team, a standing policy forum, quality and customer-facing change control
Risk
Concentrated — low frequency, high consequence, and contractual rather than technical in nature

If this is you, the next step is

We run a partition scenario against your real envelopes and reconciliation trail.

Stress-test a federated decision path

Where wafer fabs actually sit on the ladder today

The distribution, why rung 2 is the plateau, and the one distributed system almost every fab already runs without noticing.

Most fabs are at rung 2 — local intelligence that can see and cannot act. The distribution below is weighted heavily toward the sensing edge, because adding a local model requires no delegation of authority and therefore no argument with quality, industrial engineering or a customer. Adding an envelope requires all three, which is why the drop from rung 2 to rung 3 is the largest single transition loss on the ladder and the one worth planning for explicitly.

Illustrative distribution of wafer fabs across the devolution ladder

An illustrative distribution, not a measurement. It synthesises what published fab smart-manufacturing material reports about adoption and its barriers, and it is used here only to show the shape of the plateau — treat the individual percentages as indicative rather than as a benchmark.

Share of fabs

  • 31% — 1 · Central command
  • 42% — 2 · Sensing edge (the plateau)
  • 19% — 3 · Bounded local authority
  • 7% — 4 · Negotiated autonomy
  • 1% — 5 · Federated line

Source: Illustrative distribution, synthesised from published smart-manufacturing adoption reporting in fabs

10×

Faster troubleshooting from AI wafer-pattern classification, as published by GlobalFoundries

GlobalFoundries

~500k

3D objects in the plant digital twin at Bosch’s Dresden fab, mapping infrastructure and machines

Bosch

50+

Fab employees surveyed on smart-manufacturing adoption in fabs and the barriers to it

Flexciton

The barriers reported in that survey are worth reading against the ladder rather than against a technology roadmap. Flexciton’s 2025 front-end fab insights report (opens in a new tab), drawn from a survey of more than fifty fab employees worldwide, names integration with existing systems, data inadequacies and risk-averse leadership among the reasons advanced technologies move slowly into fabs. Every one of those is a rung-2-to-rung-3 blocker rather than a modelling one: integration is the constraint feed, data inadequacy is the fab state of record, and risk aversion is the entirely rational response to being asked to delegate authority without a written bound or a tested revert.

It is also worth being clear about what the plateau is not. It is not a shortage of algorithms. Learned dispatch and fab-scale scheduling have an active literature — self-supervised and reinforcement-learning approaches to semiconductor fab scheduling (opens in a new tab) and hybrid answer-set-programming formulations of the same problem (opens in a new tab) both address realistic fab constraints, and the latter is explicit that existing practice schedules locally with greedy heuristics or by optimising machine groups independently. The methods exist. What is missing in most fabs is the authority for anything to act on them without a person in the loop.

The authority map: which fab decisions can actually be devolved

Row per decision: the latency the physics demands, what the decision couples to, the smallest scope that holds those constraints — and therefore whether it can leave the centre at all.

A fab decision can be devolved to a scope if and only if that scope can see every constraint the decision touches, fast enough to matter. That single rule decides most of what follows, and it is a physics-and-coupling test rather than a technology test — which is why the map below does not change if the models get better. It is the page’s centrepiece: work down it with your own decision list, and the argument about what to automate stops being a matter of ambition and becomes a matter of arithmetic.

Fab decisionLatency the physics demandsWhat it couples toSmallest scope holding the constraintsDevolvable?
Abort or interlock trip on an out-of-spec chamber stateMillisecondsThe wafer in the chamber; personnel and equipment safetyThe tool controllerAlready devolved — hard-wired, and never negotiable
Run-to-run offset for the next run on this chamberOne run (minutes)Metrology feedback for this layer on this chamberThe tool’s own control loopDevolves cleanly — the classic bounded envelope
Hold a chamber out of service on an FDC excursionSeconds to minutesDownstream WIP, qualification status, the shift’s capacityThe bay, given a hold budgetDevolvable at rung 3 with a budget and a revert
Which wafers on this lot to send to measurementMinutes — before the lot moves onMetrology capacity, SPC power, product risk rulesThe bay, inside a measurement budgetDevolvable at rung 3; the SPC rules must move with it
Next lot onto a non-bottleneck toolsetSecondsQ-time links upstream and downstream, dedication, batch shapeThe bay — but only if queue-time clocks are fed liveRung 3–4, and only with a live constraint feed
Batch composition and release on a furnace or wet benchMinutesBatch efficiency, cycle time, every Q-time link in the batchThe toolset, with segment-level constraints pushed downRung 4 — the three objectives conflict and must be arbitrated
Reticle movement between litho toolsMinutesEvery lot needing that reticle, across baysFab-wide — a reticle is a physical singletonNot devolvable; arbitrate by lease at rung 4
Q-time rescue when a tool goes down mid-linkMinutesTwo or more process steps and the whole segment’s WIPThe segment, not the tool or the bayRung 4 — needs a negotiated priority, not a local one
Hot-lot priority and its expiryHoursCustomer commitments, every queue the lot will pass throughFab-wide, with expiry enforced centrallyNot devolvable; the claim expires, the authority does not move
Preventive-maintenance window placement across a toolsetHours to daysThe week’s capacity, spares, technician availabilityToolset plus maintenance planningPartially devolvable at rung 4; the calendar stays central
Wafer starts and the week’s product mixDaysCustomer commitments, capacity, capital, revenue recognitionThe fab and its plannersCorrectly central, permanently
The authority map for a wafer fab. “Smallest scope that holds the constraints” is the test: a decision devolves cleanly only when one scope can see everything it couples to, within the latency the physics demands. Rows marked not devolvable are not waiting for better AI — they are waiting for arbitration, which is rung 4.

Three patterns fall out of the map. First, everything at the millisecond and single-run end of the latency column is already devolved and has been for years — the industry solved local autonomy long ago and called it control, not AI. Second, the rows that resist devolution resist for one reason: the resource or the constraint is shared. A reticle is a singleton; a queue timer links two steps that no single scope owns; a hot lot passes through queues belonging to bays that will never meet. Third, the rows at the bottom are not failures of ambition. A fab that devolved its start plan to individual bays would not be more autonomous; it would have no commercial control over what it manufactures.

The queue-timer row deserves separate attention because it is where most well-intentioned devolution programmes come apart. Queue-time constraints link steps across tool groups: after a wet clean, wafers must reach the next furnace within a stated window or be reworked or scrapped. A bay agent optimising its own toolset’s utilisation will happily accept a lot whose downstream link it cannot see, and the cost of that acceptance appears somewhere else, later, as scrap that nobody attributes to a dispatch decision. Flexciton has published technical material specifically on why queue timers defeat conventional schedulers (opens in a new tab) and on the yield-versus-capacity trade-off they force (opens in a new tab). The engineering conclusion for a devolution programme is blunt: if you cannot feed queue-time remaining to the edge, do not devolve lot selection at all.

Where to devolve first — latency demand against constraint span

Plot a decision by how fast it decays and how far it couples. Only one quadrant is a good first move; one is a trap; and the bottom-right quadrant is a permanent, correct home for a large part of what a fab decides.

Devolve first

  • Fast and locally contained — FDC holds, R2R correction, in-bay sampling
  • The scope already sees everything the decision touches
  • Move: write one envelope, feed the local constraints, ship it

Negotiate, don’t just devolve

  • Fast and globally coupled — reticle contention, Q-time rescue, AMHS congestion
  • Devolving without arbitration distributes conflict, not authority
  • Move: build the lease primitive before moving the decision

Devolution buys nothing

  • Slow and contained — routine consumable replacement, local calibration cadence
  • A central plan is fast enough, and one place is cheaper to qualify
  • Move: leave it alone and spend the effort elsewhere

Correctly central, permanently

  • Slow and globally coupled — starts, mix, capital, PM calendar
  • These are commercial decisions wearing operational clothes
  • Move: improve the central plan; do not try to distribute it
Latency the physics demands — top: Seconds — the decision decays fast, bottom: Hours — a shift is fine
Constraint span of the decision — left: Contained: one tool or one bay, right: Coupled: crosses bays, links steps, shares a scarce resource

The trap quadrant is the top right, and it is where visionary programmes reliably land, because those decisions are the ones that hurt most and therefore attract attention first. Reticle contention and queue-time rescue are genuinely fast, genuinely painful and genuinely global — which is exactly why they need an arbitration primitive before they need a local agent. A fab that devolves the top-right quadrant without the arbiter has not decentralised authority; it has replaced one queue at production control with several simultaneous ones and removed the person who used to resolve them.

Real today, published research, and speculation — kept apart

The discipline this topic demands: three columns, and a refusal to let a research claim borrow the credibility of a production one.

Most writing about the autonomous fab fails at one specific point: it moves between what is running in production, what a paper demonstrated in simulation, and what somebody hopes will be true in 2040, without ever signalling the change. The table below keeps those three apart deliberately, claim by claim. It is the most useful section of this page for anyone building an internal business case, because a capital committee will not fund a roadmap whose evidence classes are mixed — and it will fund one whose author separated them before being asked.

Claim you will hearWhat is genuinely in productionWhat published work actually showsWhat remains speculation
“The fab runs itself”Material handling moves carriers continuously without human dispatch — GlobalFoundries publishes up to 10,000 carrier moves a day in one fabLearned dispatch and fab-scale scheduling are formulated against realistic fab constraints and evaluated largely in simulationAn end-to-end fab in which no human owns a decision class
“Agents will negotiate with each other”No publicly documented fab-scale agent negotiation; contention is resolved by plan or by peopleHolonic and multi-agent control architectures have a two-decade literature, including benchmarking frameworks for comparing themEmergent negotiation replacing an explicit, auditable protocol
“Edge AI decides at the tool”Run-to-run control, FDC and on-tool health models — decades old, and genuinely localMulti-agent approaches to dynamic dispatching in material-handling systems are demonstrated on enterprise dataOn-tool models altering a qualified recipe without change control
“The fab is one distributed brain”Centralised data architectures and control towers — Bosch describes fully connected systems within a centralised data architecture at DresdenFederated learning addresses sharing model updates without sharing data; it is about privacy and data locality, not decision authorityCross-company federated control of production decisions
“Swarm behaviour will optimise the line”Nothing in a production fabBio-inspired flocking has been proposed as an optimisation inspiration for production plants — an ideas paper, not a deployment reportSelf-organising WIP flow with no arbiter and no invariants
“Autonomy removes the qualification burden”Nothing — qualified processes still carry change-control and customer-notification obligationsNo published work claims otherwise; the deciding rule is part of the processCustomers and auditors accepting an unversioned autonomous policy
Three evidence classes for the decentralised-fab vision. Column two is production reality with a published source; column three states what research actually claims, including its evaluation setting; column four is speculation, labelled as such. Nothing moves left without new evidence.

The multi-agent row is the one most often overstated, so it is worth being exact about what the literature does and does not say. Holonic manufacturing control architectures (opens in a new tab) have been formalising the trade-off between local autonomy and global coherence since well before the current AI cycle, and there is published work on how to benchmark holonic systems against each other (opens in a new tab) precisely because comparing them has been hard. That is a mature research programme, not a shipping product line. More recent work is sharper still about the cost of modularity: an analysis of the coordination gap between joint and modular learning for job-shop scheduling with transport resources (opens in a new tab) measures what is lost when decisions are learned separately rather than together, and separate work on graph-structured experience reuse for multi-agent adaptation in dynamic manufacturing (opens in a new tab) addresses how such agents adapt. Both are evidence that coordination is the hard part — which is the argument for an arbiter, not against decentralisation.

  • Physics is the first bound, and it does not move

    Queue-time windows, thermal budgets, resist shelf life and chamber seasoning are chemistry and physics. No architecture makes a linked step less linked. Every claim about a self-coordinating fab has to survive the observation that some of its couplings are covalent.

  • Qualification is the second bound, and it is contractual

    A qualified process is qualified as a whole, and the rule that decides is part of it. Fabs serving automotive and medical customers publish their certification frames openly — GlobalFoundries’ quality and certifications page (opens in a new tab) is a representative example — and the obligations behind them do not distinguish between a human decision and an automated one. A change in an autonomous decision rule is a process change, with everything that follows.

  • Interfaces are the third bound, and they are where the work is

    Devolved decisions need equipment data and they need it in a normalised form. The SEMI standards suite — GEM for equipment communication, the EDA or “Interface A” standards for structured data collection — is what makes that possible, and the state of a fab’s conformance to it is usually the real schedule driver. SEMI’s standards programme (opens in a new tab) is the reference point; the fab-side work is normalisation, not procurement.

  • Evidence is the fourth bound, and it is the one that decides

    Whatever a fab devolves, it must be able to reconstruct months later: the action, its inputs, the envelope version, the arbiter’s grant. The fabs that will devolve most in 2035 are the ones that made reconciliation a by-product of the action path in 2026, rather than a reporting project scheduled after the models.

These cross-fab and multi-step decisions sit between dispatching and planning. No MES handles them. No scheduler touches them. They live in the heads of experienced engineers — and the response depends on who is on shift.

That quotation is the honest starting point for any decentralisation programme, and it reframes the whole vision. The gap in a modern fab is not that decisions are made centrally; it is that a large class of consequential decisions is made by nobody in particular, at a scope nobody wrote down, with an answer that varies by shift. Decentralised autonomy is one legitimate response to that gap. Better central automation is another, and for the top-right quadrant of the matrix it is usually the correct one. What is not a response is leaving the decisions unassigned and adding models around them.

What the published record actually shows

Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the material it is drawn from.

The public record on fab autonomy points somewhere slightly awkward for the decentralisation thesis, and it is more useful for that. The most documented programmes are not distributed at all: they are centralised data architectures, global control towers and enterprise AI functions, sitting above an execution layer that has been autonomous for decades. Read the three below for what leading fabs actually chose to build, not for a benchmark to apply to your own line.

Three programmes read against the devolution ladder

Outcomes as reported in the linked material; we have not independently audited the figures, and where a result comes from a vendor webinar rather than a peer-reviewed study it is stated. Card images are generated industry scenes from our library — none depicts the named operator’s facility, and none implies an endorsement.

Illustrative cleanroom scene: a wafer on a stage under an inspection head with analytics overlays and engineers at consoles behindGlobalFoundriesGlobal foundry · fabs in the US, Europe and Asia23
Challenge
Scaling AI-supported manufacturing decisions across a footprint of geographically separated fabs, where a per-site build never amortises and per-site divergence in data definitions makes any cross-site comparison meaningless.
Approach
GlobalFoundries publishes a digital-manufacturing programme built around centralised capability: a proprietary factory control tower described as a virtual fabric monitoring key production processes and performance metrics across all its global manufacturing with 24/7 support from manufacturing hubs, a global AI centre of excellence based in Singapore that pilots and scales solutions across sites, and heavy automation at the execution layer including robotics and an overhead material-handling system.
Reported outcome
GlobalFoundries publishes that its overhead transport delivers as many as 10,000 carrier moves a day in a single fab, that custom AI engines classifying wafer patterns speed troubleshooting time by up to 10× and reduce wafer scrap, and that in 2025 the World Economic Forum designated its 300 mm fab in Singapore part of the Global Lighthouse Network of advanced manufacturers.
What it shows about the curveCoordination centralises as execution devolves, and that is not a contradiction. The material handling is autonomous; the decision layer above it is a control tower. This is rung 2 moving to rung 3 with an unusually strong central spine — which is exactly the spine any later devolution will need.

GlobalFoundries — digital manufacturing (opens in a new tab)

Illustrative cleanroom scene viewed through observation glass, with a plant model rendered as an overlay and engineers at monitoring stationsBosch (Dresden)300 mm automotive and power semiconductor fab · greenfield13
Challenge
Bringing up a greenfield 300 mm fab for automotive ICs and power semiconductors — a product class with among the most demanding traceability and change-control obligations in the industry — with the opportunity to choose its data and decision architecture from scratch rather than inherit one.
Approach
Bosch describes the Dresden plant as data-centric from the start: systems fully connected within a centralised data architecture giving detailed data capture for each chip, AI applied to data-based process control, real-time monitoring and machine learning for early deviation detection, and a digital twin of the plant consisting of around half a million 3D objects mapping infrastructure and machines for planning, remote maintenance and simulating process changes.
Reported outcome
Bosch reports that this combination supports traceability and consistent quality control per chip, earlier deviation detection and higher manufacturing yields, alongside optimised planning and simulation of process changes through the digital twin.
What it shows about the curveThe most interesting sentence in Bosch’s own description is “centralised data architecture”. Given a blank sheet and every reason to be radical, a modern greenfield fab built the central spine first. Devolution is what you can do once that spine exists — the ordering is not optional.

Bosch — Dresden semiconductor plant (opens in a new tab)

Illustrative cleanroom scene: engineers in coveralls beside a process tool with production-control analytics rendered as an overlayUnnamed production fab (reported by Flexciton)Front-end wafer fab · production control13
Challenge
A class of cross-fab, multi-step production-control decisions — when to start WIP, how to respond to a tool going down, how to balance competing priorities when capacity is tight, how to keep bottleneck tools fed, how to manage queue timers, whether to batch aggressively or protect cycle time — that no MES handled and no scheduler touched, and which in practice depended on which experienced engineer was on shift.
Approach
Rather than devolving these decisions to individual toolsets, the reported approach automates them across the whole fab in real time, every cycle, sitting above both rule-based dispatching and any toolset scheduler already in place — that is, assigning the unowned decisions an explicit owner rather than distributing them.
Reported outcome
Flexciton reports a cycle-time reduction of over 10% alongside a throughput improvement at a production fab, and describes more consistent performance across shifts with less reactive firefighting. This is published as a vendor webinar and case description rather than a peer-reviewed study, and should be read at that weight.
What it shows about the curveThe largest reported win in this space came from assigning authority, not from decentralising it. Before asking which decisions should move to the edge, ask which decisions currently have no owner at all — that list is usually longer, cheaper to fix and worth more.

Flexciton — “The decisions that nobody owns” (opens in a new tab)

Read together, the three make one uncomfortable and useful point: nothing in the published record describes a fab in which authority has been genuinely distributed to negotiating local agents. What the record describes is strong central state, autonomous execution, and a slowly closing gap in between. That is not an argument against the decentralised vision — it is a statement about sequence. For wider context on where fab automation is heading, Semiconductor Engineering’s manufacturing coverage (opens in a new tab) tracks the field continuously, imec (opens in a new tab) publishes on advanced process and equipment research, and TSMC (opens in a new tab), Intel (opens in a new tab), Micron (opens in a new tab) and GlobalFoundries (opens in a new tab) all publish periodically on smart-manufacturing programmes.

Contention: what devolved decisions fight over, and how to settle it

The scarce shared resources of a wafer fab, what allocates each one today, and the arbitration primitive that replaces the phone call.

Devolved decisions collide over a small, well-known set of scarce resources, and the whole difficulty of rung 4 is contained in that list. A fab does not have a general contention problem; it has eight or nine specific ones, each with a physical reason for being scarce and each currently allocated either by a static central plan or by whoever escalates first. Naming them individually is what turns an abstract multi-agent architecture into a piece of engineering with a scope.

Contended resourceWho wants itWhat allocates it todayArbitration primitiveFailure mode if unarbitrated
A specific reticleEvery litho bay with lots on that layerThe dispatch plan, or an escalation to production controlExpiring lease with a stated return timeTwo bays hold work for the same reticle and both wait on a person
Metrology capacityEvery bay wanting to sample, plus engineering’s special requestsA static sampling plan set quarterlyMeasurement budget per bay, refilled on a cadenceAdaptive sampling everywhere means measurement queues everywhere
AMHS vehicles and track segmentsEvery move request in the fab, continuouslyThe material-handling controller — already autonomousAlready arbitrated; treat as the reference implementationCongestion and deadlock around busy stockers and lifts
Qualification slots on a shared toolProcess engineering, plus any bay wanting a chamber requalifiedA qual calendar owned by process engineeringReservation with a priority class and an expiryRequal backlog silently caps how freely lots can be routed
Batch capacity on a furnace or wet benchEvery lot arriving at the batch step within the windowLocal batching rules, tuned by handSegment-level batch objective as a central invariantBatches optimised locally break Q-time links downstream
Maintenance technicians and sparesEvery bay holding a tool for a PM or a repairThe maintenance schedule and a shift supervisorWork-request queue with an explicit priority policyDevolved holds create more work than the crew can absorb
Hot-lot prioritySales, planning and anyone with a committed customer dateA flag on the lot, rarely removedExpiring claim with a fab-wide budget on concurrent claimsPriority inflation: everything is hot, so nothing is
Stocker and buffer space near a bottleneckEvery upstream bay pushing WIP toward the constraintPhysical capacity and whoever notices firstSpace budget per feeding bay, enforced centrallyThe bottleneck is starved while its buffers are full of the wrong work
The contention table for a wafer fab. The arbitration primitive column is deliberately conservative: leases, budgets and expiring claims are chosen because they degrade safely under partition, not because they are the most sophisticated option available.

The material-handling row is the most instructive on the table, because it is the one already solved. An AMHS arbitrates vehicles and track segments continuously, at a latency no human process could match, under rules that are conservative by design and that fail safe when a segment goes down. Nobody describes it as multi-agent AI and nobody needs to. It is the existence proof that a fab can run an arbitrated, distributed, safety-critical allocation system in production — and the model to copy when designing the reticle or metrology equivalent.

  • Leases, not assignments

    A lease is granted for a bounded period and expires whether or not it was used. That single property is what makes contention safe under failure: an agent that loses contact with the arbiter cannot hold a reticle indefinitely, and the resource returns without anyone intervening. Assignments, by contrast, have to be revoked, and revocation is exactly the operation a partitioned system cannot perform.

  • Refusals carry a reason and a time

    An arbiter that says no without saying when is functionally an escalation. A refusal that carries an expected availability lets the refused scope plan around it — run something else, pre-stage the batch, adjust the sampling — which is the whole operational point of arbitration over queuing.

  • Invariants live in the arbiter, incentives live in the agents

    Anything a local scope must never do — break a queue-time link, exceed a concurrent-hold budget, route to an unqualified chamber — belongs in the arbiter as a hard limit, not in the agent as a penalty term. The reason is auditability as much as safety: a hard limit is a rule you can show a customer, and a penalty term is a number you have to explain.

  • Measure contention, not just throughput

    Lease wait time per resource, refusal rate by requesting scope, and starvation — one scope repeatedly refused — are the health metrics of rung 4. A fab that only measures moves and cycle time will see a contention problem as unexplained variance, and will attribute it to the models rather than to the allocation.

Photolithography deserves a note of its own, because it concentrates almost every contention type in one area: a scarce, physically singular reticle, cluster tools whose internal scheduling is itself a hard problem, and the fab’s usual bottleneck. There is published work on scheduling a photolithography process containing cluster tools (opens in a new tab) that shows how quickly the combinatorics escalate even before any of it is distributed. The practical implication for a devolution programme is to sequence litho last: it is where local decisions have the most global consequence, and where a wrong arbitration policy is most expensive.

The reference architecture for devolved authority, layer by layer

What has to exist for each rung — and which layer you can honestly defer.

Devolving one decision safely requires five layers, and the order in which they are built decides whether the programme compounds or stalls. The architecture below is deliberately unfashionable: nothing in it is vendor-specific, every layer is defined by what it must guarantee rather than by what product provides it, and the layer most often skipped — reconciliation — is the one that determines whether anyone will let the arrangement run.

Layers required by rung

Each layer is annotated with the rung that first requires it. A programme aiming at rung 3 without the constraint and reconciliation layers is building a rung-2 sensing edge with extra latency.

  1. Equipment and safety layer

    Stage 1+

    • Interlocks and equipment safetyLocal, hard-wired, outside every negotiation
    • Equipment controllerExecutes moves and recipes; holds the run-to-run offset
    • GEM / EDA interfacesStructured equipment data, normalised across vendors
  2. State and constraint layer

    Stage 2+

    • Fab state of recordOne authoritative view of lots, steps, qualification and reticles
    • Live constraint feedQ-time remaining, dedication, budgets — with a freshness SLA
    • Trace and FDC streamsTool signals delivered at the rate the decision needs
  3. Local decision layer

    Stage 3+

    • Bay or tool decision agentDecides inside the envelope, in seconds
    • Envelope evaluatorChecks bounds and preconditions before every action
    • Deterministic fallback ruleThe prior rule, one switch away and always available
  4. Arbitration layer

    Stage 4+

    • Resource leasesReticles, metrology capacity, qual slots, buffer space
    • Global invariantsHard limits no local scope can bid past
    • Contention telemetryWait time, refusal rate, starvation by scope
  5. Policy, reconciliation and evidence

    Stage 3+

    • Versioned envelope policyReviewed like a recipe change; stamped on every action
    • Reconciliation into the MESAction, inputs, envelope version, arbiter grant
    • Partition and degraded modePer-decision behaviour when the plane is unreachable

Pipeline described

  1. Equipment and safety layer (stage 1+) — Interlocks and equipment safety: Local, hard-wired, outside every negotiation; Equipment controller: Executes moves and recipes; holds the run-to-run offset; GEM / EDA interfaces: Structured equipment data, normalised across vendors
  2. State and constraint layer (stage 2+) — Fab state of record: One authoritative view of lots, steps, qualification and reticles; Live constraint feed: Q-time remaining, dedication, budgets — with a freshness SLA; Trace and FDC streams: Tool signals delivered at the rate the decision needs
  3. Local decision layer (stage 3+) — Bay or tool decision agent: Decides inside the envelope, in seconds; Envelope evaluator: Checks bounds and preconditions before every action; Deterministic fallback rule: The prior rule, one switch away and always available
  4. Arbitration layer (stage 4+) — Resource leases: Reticles, metrology capacity, qual slots, buffer space; Global invariants: Hard limits no local scope can bid past; Contention telemetry: Wait time, refusal rate, starvation by scope
  5. Policy, reconciliation and evidence (stage 3+) — Versioned envelope policy: Reviewed like a recipe change; stamped on every action; Reconciliation into the MES: Action, inputs, envelope version, arbiter grant; Partition and degraded mode: Per-decision behaviour when the plane is unreachable
Step-by-step insights
Equipment and safety — the layer you inherit, and must not touch
The safety layer is already decentralised and already correct, and the only sensible relationship a devolution programme has with it is deference. Interlocks veto; they are not consulted. What does need work at this layer is data: GEM and the EDA standards define how equipment exposes structured information, and conformance across a mixed-vendor estate is uneven enough that trace normalisation is usually the longest single task in any fab AI programme. Budget for it explicitly, because it is invisible in a business case and it sets the schedule.
State and constraint — the layer that decides how far you can go
A devolved scope is blind to everything outside itself, so its ceiling is set by what the constraint feed carries and how fresh it is. Two design rules pay for themselves repeatedly: every constraint carries a timestamp and a validity, and staleness suspends the dependent decision rather than degrading it silently. Fabs that instrument freshness discover within a week that some feeds they assumed were live are hours old — and that discovery is worth more than the model it was meant to support.
Local decision — the agent is the smallest part of the work
The decision agent itself is usually modest: a model or a rule, an envelope check, an action, a log line. What surrounds it is the engineering. The deterministic fallback rule matters most and is skipped most often — it is the previous behaviour, always available, one switch away, and it is the political key that unlocks approval, because production control will accept a new decision source it can instantly revert and will not accept one it cannot. A devolution proposal without a drilled revert sits in a change queue for two quarters; one with it ships.
Arbitration — build it before you need it, keep it small
The arbiter should be the least clever component in the architecture. Its entire job is granting and refusing bounded leases and enforcing invariants, and its entire behaviour should be reconstructable from a log a person can read. Resist the pull to make it a planner: the moment it starts optimising, it becomes a central scheduler with extra steps, and the fab has walked back to rung 1 through the side door while calling it rung 4.
Policy and reconciliation — the layer that makes it defensible
Envelopes, invariants and partition rules are one artefact and should be versioned as one. Every devolved action carries the version it ran under, and reconciliation writes action, inputs, version and grant back into the MES record as part of the action path rather than as a later job. This is what converts a customer’s change-control audit from a project into an export, and it is the difference between a fab that can extend autonomy and one that has to defend the autonomy it already has.

The layer most commonly deferred is reconciliation, and deferring it is the most expensive scheduling decision in the whole programme. Devolved actions taken without a reconstructable record are not merely undocumented; they are unqualifiable, which means the fab cannot extend the envelope, cannot answer a customer question about them, and in the worst case has to switch the capability off during an audit. Build reconciliation with the first envelope, when it costs a fortnight, rather than after the fourth, when it costs a re-architecture.

A 90-day plan: devolving metrology sampling to one bay

The rung 2 → 3 transition made concrete on one wafer-fab decision — which wafers to measure — chosen because it is bounded, reversible and immediately measurable. Contains no model development.

Moving one rung takes about 90 days when it is scoped to a single decision, and several years when it is scoped to a function. To make that concrete, the plan below devolves one specific decision: which wafers on a lot to send to measurement, in one bay, inside a fixed measurement budget. It is the right first move for three reasons — the decision is contained within a bay, it is fully reversible by reverting to the static sampling plan, and its cost and benefit are both measurable within the quarter. The models required already exist in most fabs; the quarter contains no model development at all.

Rung 2 → rung 3 on metrology sampling, in one quarter

One bay, one measurement step, one owner. If any phase needs more than its window, narrow the scope — fewer products, one layer — rather than extending the plan.

  1. Days 1–20

    Write the envelope and baseline the sampling

    Choose one bay and one measurement step. Pull six months of sampling history, metrology queue times and the excursions actually caught, and compute the current static plan’s measurement rate, detection lag and metrology-tool loading. Then write the envelope: the bay may vary which wafers are measured within a fixed weekly measurement budget, never below a stated floor per lot, never on product classes marked mandatory-measure, suspended automatically if the constraint feed goes stale. Name the bay process owner as accountable, and get the envelope through the same review forum that approves a recipe change.

    A versioned envelope and a baseline nobody disputes

  2. Days 21–45

    Stand up the constraint feed and the fallback

    Feed the bay the three things it cannot see: remaining measurement budget, live metrology queue state, and the product risk classes that force a measurement. Implement the deterministic fallback — the existing static sampling plan — as a one-switch revert, and wire staleness on any feed to suspend the devolved decision automatically. Nothing decides autonomously yet; the feed and the fallback are proven first.

    Live constraint feed with a freshness alarm, and a tested revert

  3. Days 46–70

    Devolve the decision and reconcile every action

    Let the bay decide sampling inside the envelope. Every decision writes to the MES record with its inputs, the envelope version and the budget state at the time. Keep the SPC rules honest: because adaptive sampling changes the sampling distribution, the control rules must move with it, and the review forum should sign off the revised rules rather than inheriting the old ones by accident. Exercise the revert once, deliberately, on a quiet shift.

    Devolved sampling live, fully reconciled, revert drilled

  4. Days 71–90

    Attribute in detection lag and metrology hours

    Hold out a comparable bay or product family on the static plan. Report the difference in three numbers: mean lag from excursion onset to detection, metrology tool-hours consumed, and wafers at risk between two measurements. Report the escalation rate alongside them — how often the decision fell outside the envelope — because that number is what sizes the next envelope, and it is the honest measure of whether the bounds were right.

    A detection-lag and metrology-hours delta with a holdout behind it

The order matters

  1. The envelope before the agent

    Write and approve the bounds before anything decides. A team that builds first and negotiates afterwards ends up in advisory mode, which is rung 2 with more infrastructure. The envelope is also the artefact that makes the quality conversation tractable, because it converts “can the AI decide?” into “may this bay vary sampling within this budget?”, which is a question a fab already knows how to answer.

  2. The constraint feed before the decision

    Prove the feed and its freshness alarm while the static plan is still in charge. A devolved decision running against a stale budget or an out-of-date risk class is the specific failure that makes a fab switch the whole capability off, and it is entirely preventable by sequencing.

  3. The reconciliation before the second decision

    Do not devolve a second decision until the first one’s actions can be reconstructed from the record alone, including the envelope version. Reconciliation built once, at the action path, is cheap; retrofitted across four devolved decisions, it is a re-architecture — and until it exists, none of the devolved authority can be extended.

Proving authority actually moved: formula, source, cadence

Where each metric comes from, how often to read it, and the rung at which it starts measuring something real. All telemetry, no self-report.

A devolution claim you cannot name a source system for is an opinion. Every metric below reduces to timestamps, counts and version identifiers that the MES, the decision log, the constraint feed or the arbiter already produces — the instrumentation work is joining them, not inventing them. The table is the build sheet: read your rung’s rows, and treat “honest from” as a filter, because a metric quoted before its rung is measuring a thing that does not yet exist.

MetricFormula / readSourceCadenceHonest from
Envelope coverageDecisions with a versioned envelope ÷ decisions in the devolution scopeDecision-rights registerMonthlyRung 2
Local decision latencyAction timestamp − triggering-event timestampDecision log + tool tracePer decisionRung 3
Constraint-feed stalenessNow − last update, per constraint, at decision timeConstraint feedContinuousRung 3
Escalation rateOut-of-envelope escalations ÷ devolved decisionsDecision logWeeklyRung 3
Reconciliation completenessDevolved actions reconstructable from the MES record ÷ all devolved actionsMES + decision logWeeklyRung 3
Revert exercise recencyDays since the fallback was last deliberately exercisedChange recordMonthlyRung 3
Lease wait timeGrant timestamp − request timestamp, per contended resourceArbiter logPer requestRung 4
StarvationConsecutive refusals to one requesting scopeArbiter logWeeklyRung 4
Q-time breach rate on devolved stepsBreaches ÷ linked lots passing through the devolved decisionMES step timestampsWeeklyRung 3
Partition drill recencyDays since degraded-mode behaviour was last rehearsedChange recordQuarterlyRung 5
Instrumentation build sheet for a devolution programme in a wafer fab. “Honest from” is the rung at which the metric first measures something real.

Two of those metrics deserve a note because they are routinely misread. Escalation rate is not a defect rate — a rate that falls to zero usually means the envelope is too wide rather than that the world became simpler, and a rising rate is an early warning that conditions have moved outside the envelope’s validity. Reconciliation completeness is the one that gates everything above rung 3: a fab cannot extend an envelope it cannot audit, so a completeness figure below the high nineties is a hard stop on further devolution regardless of how well the decisions are performing.

Rung 3 readiness checklist

If you cannot tick all eight for a specific decision, that decision has not been devolved — it has been delegated informally, which is a different and less durable thing. Tick as you go; this list works without JavaScript.

0 of 8 ticked

Nothing ticked — start with the register, not the platform

Zero ticks is the normal starting point and it is not a technology problem. Spend a fortnight enumerating the decisions in one bay, who makes each today and what each couples to. The register is free, it is the artefact every later rung depends on, and it usually reveals two decisions that are already devolvable with no new software at all.

Failure modes that send devolved authority backwards

Devolution is not monotonic. Four regressions account for nearly all of it, and three of the four are silent.

Devolution regresses more often than it advances, and it usually regresses quietly, because a devolved decision that has gone wrong keeps producing actions that everyone keeps trusting. The four patterns below account for nearly all of it in fabs, and each has a cheap preventive measure that has to be in place before the failure rather than after it.

Likelihood: highImpact: high

Local optimisation eats the line

A bay agent maximises what it can see — its own tool utilisation, its own moves — and pays for it downstream in broken queue-time links, a starved bottleneck or batches that suit the furnace and nobody else. Every individual decision is defensible; the aggregate is worse than the central plan it replaced, and the cause is distributed across hundreds of choices nobody logged as a mistake.

PreventionPut the global objective in the arbiter as hard invariants, never in the agent as a reward term, and monitor segment-level Q-time breach rate rather than bay-level throughput.

Likelihood: highImpact: high

The constraint feed goes stale and the decision keeps deciding

A dedication map, a risk class or a budget counter stops updating. The devolved decision continues, confidently, against yesterday’s facts. This is more dangerous than an outage because nothing looks wrong: the actions arrive on time, in the right format, and are wrong in a way that only shows up in yield or in a queue-timer breach weeks later.

PreventionEvery constraint carries a timestamp and a validity; staleness suspends the dependent decision automatically and reverts to the deterministic fallback rather than degrading silently.

Likelihood: mediumImpact: high

Envelope creep by ticket

Bounds are widened one incident at a time, in a configuration screen, by whoever has access. Six months later nobody can state what the bound was in March — which is precisely the question a customer’s change-control audit asks. The capability is then either withdrawn or defended with a reconstruction exercise that costs more than the whole programme saved.

PreventionEnvelopes are controlled documents reviewed by the forum that approves recipe changes, and every action carries the envelope version it ran under.

Likelihood: lowImpact: high

Partition with no defined degraded mode

The central plane becomes unreachable during a network change or a system upgrade. Some devolved decisions stop, some carry on, and which is which is whatever the code happens to do. The recovery is worse than the outage: local actions taken during the partition were never reconciled, so the record has a hole exactly where the interesting events are.

PreventionA per-decision partition table — what may continue, for how long, and what must stop — rehearsed quarterly, with a reconciliation job that replays local actions on reconnect and flags anything outside the cached envelope.

The through-line is that none of these four is a model failure. Every one of them is a failure of authority design: an objective in the wrong place, a constraint that did not arrive, a bound nobody versioned, a behaviour nobody specified. That is the honest summary of the decentralised-fab vision and the reason this page has spent its length on registers, envelopes, feeds and logs rather than on architectures of agents. The fabs that will devolve most in ten years are the ones treating decision authority as an engineering artefact today.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Authority envelope
The versioned, reviewable document that states which decisions a scope may make alone, within what numeric bounds, under what preconditions, consuming what budget, and with what revert. The central artefact of decentralised autonomy — without one, local authority is a habit rather than a delegation.
Decision rights
The written allocation of who or what may make each recurring decision in the fab. Distinct from capability: a system can be able to decide and have no right to. Recorded in a decision-rights register, which is the first artefact of any devolution programme.
Constraint feed
The live push of global facts a devolved decision must respect — queue-time remaining, dedication and qualification state, reticle location, remaining budgets — carrying a timestamp and a validity. Its freshness sets the ceiling on how far authority can be devolved.
Constraint span
How far a decision couples beyond the scope making it. A run-to-run offset spans one chamber; a reticle move spans every bay with lots on that layer. Span, not difficulty, is what determines whether a decision can be devolved at all.
Lease
A bounded, expiring grant of a contended shared resource issued by an arbiter. Chosen over an assignment because it degrades safely: it returns automatically on expiry, so a scope that loses contact with the arbiter cannot hold the resource indefinitely.
Arbiter
The small central component that grants and refuses leases on scarce shared resources and enforces the invariants no local scope may cross. Deliberately not a planner — the moment it optimises, it becomes a central scheduler and the fab has re-centralised through the side door.
Global invariant
A hard limit enforced centrally that no devolved decision may violate — an unbroken queue-time link, a concurrent-hold budget, routing only to qualified chambers. Invariants belong in the arbiter as rules, never in agents as penalty terms, because a rule can be shown to a customer.
Q-time (queue time)
The maximum permitted interval between two linked process steps, after which wafers must be reworked or scrapped. The canonical cross-scope constraint in a wafer fab: it couples two steps that no single bay owns, which is why it is the constraint most often missing from a devolved decision.
Escalation rate
The share of devolved decisions that fall outside the envelope and route to a human. Read as a trend, not a defect count — a rise means conditions have moved beyond the envelope’s validity, and a fall to zero usually means the bounds are too wide.
Reconciliation
Writing every devolved action back into the system of record with its inputs, the envelope version in force and any arbiter grant, as part of the action path rather than as a later reporting job. What turns a customer change-control audit from a project into an export.
Degraded mode (partition behaviour)
The specified behaviour of each devolved decision when the central plane is unreachable. Necessarily per decision rather than global: conservative actions such as holds may continue under a cached envelope, while anything consuming a shared resource must stop.
PCN (process change notification)
The contractual notice a fab owes its customers before a qualified process changes. Relevant to autonomy because the rule that decides is part of the process — changing an envelope or an arbitration policy on a qualified flow is a process change, not a configuration edit.

Frequently asked questions

The questions engineering, production-control and quality teams ask most often when they start moving decision authority around.

What is decentralised autonomy in a wafer fab?

It is the deliberate placement of decision authority at the smallest scope that holds every constraint the decision touches — a chamber, a tool, a bay — rather than at a central scheduler. The emphasis is on authority rather than computation: a tool running a large model locally has decentralised computation, while a tool permitted to place a hold without asking has decentralised authority. Only the second changes how the fab behaves, and only the second requires a written envelope, a constraint feed and an audit trail.

Which fab decisions should never be decentralised?

Any decision whose constraints are held by no single local scope. Reticle movement is the clearest case, because a reticle is a physical singleton wanted by multiple bays; hot-lot priority is another, because the lot passes through queues owned by bays that never coordinate. Wafer starts and product mix are commercial decisions wearing operational clothes and correctly stay central permanently. These are not waiting for better models — they need arbitration, which allocates a shared resource centrally while the decisions using it stay local.

How is this different from edge AI or on-tool inference?

Edge AI is about where computation runs; decentralised autonomy is about where authority sits. A fab can deploy on-tool inference everywhere and remain at rung 2 of the ladder on this page, because every model output still terminates in an alarm a human acts on. The questions that separate the two are documentary rather than technical: is there a written statement of what the local system may do alone, is there a numeric bound on it, and does an action taken under it reach the system of record.

Do we need a multi-agent platform to start?

No, and buying one first is the most reliable way to spend a year without devolving anything. The first two artefacts are a decision-rights register and one authority envelope, both of which are documents produced by a fortnight of argument between industrial engineering, process and quality. Most fabs discover in the process that one or two decisions are already devolvable with the systems they have. Platform choices become tractable at rung 4, when several devolved decisions start contending for the same scarce resources.

How do queue-time constraints limit decentralisation?

Queue-time windows link two process steps that no single bay owns, so a bay choosing what to run next cannot respect them unless the remaining time on each linked lot is fed to it live. Without that feed, a locally optimal choice quietly breaks a downstream link and the cost appears later as rework or scrap that nobody attributes to a dispatch decision. The practical rule for a devolution programme is blunt: if you cannot feed queue-time remaining to the edge, do not devolve lot selection at all.

What does the published research actually claim about multi-agent fab control?

Less than the marketing around it suggests, and it is worth reading precisely. Holonic and multi-agent manufacturing control architectures have a literature going back two decades, including work on how to benchmark such systems against one another. Recent analyses measure the coordination gap between joint and modular learning for scheduling with transport resources — that is, they quantify what is lost when decisions are learned separately. Learned fab scheduling is largely evaluated in simulation. None of this is a report of negotiating agents running a production fab.

How does autonomy interact with customer qualification and change control?

In a qualified process the rule that decides is part of the process, so changing an autonomous decision rule is a process change with the notification obligations that follow, not a configuration edit. Practically this means three things: the envelope goes through the same review forum as a recipe change, every action carries the envelope version it ran under, and the record must let someone reconstruct months later what the system was permitted to do and why it acted. Fabs serving automotive and medical customers should assume the strictest reading.

What is an authority envelope, concretely?

A short, versioned document naming one decision, the scope permitted to make it, the numeric bounds, the preconditions that suspend it, the budget it consumes and the revert. A worked example: this bay may hold chambers on this toolset for up to four hours, at most two concurrently, never when any lot in the bay has under ninety minutes of queue time remaining, with every hold written to the MES within thirty seconds and the static rule one switch away. The test is whether a new shift supervisor could apply it at 3am without calling anyone.

How do we stop local agents optimising against the whole line?

By putting the global objective in the arbiter as hard invariants rather than in the agents as reward terms. Bidding may decide who gets a scarce resource; it must never decide whether a queue-time link may be broken or whether a lot may route to an unqualified chamber. Then monitor at the level the damage occurs: segment queue-time breach rate and bottleneck starvation rather than bay throughput. Local optimisation failures are individually defensible and collectively expensive, which is exactly why they need a structural rather than a behavioural fix.

How long does it take to move from rung 2 to rung 3?

About 90 days when scoped to a single decision in a single bay with a named owner, and several years when scoped to a function. The work is envelope drafting, constraint plumbing, reconciliation and attribution rather than model development, because at rung 2 the local models already exist. The dominant cost is usually not engineering but the qualification and change-control path, which is why the first decision should be one whose scope you fully control — sampling and engineering holds qualify, lot dispatch usually does not.

Is a fully autonomous fab realistic by 2035?

A fab in which an enumerated set of decisions executes under distributed policy, degrades gracefully when the central plane is unreachable and is auditable to customer standard is realistic — some fabs will be close within a decade. A fab in which no human owns a decision class is not, and not because of AI capability: physics keeps some couplings covalent, qualification keeps the deciding rule inside the process, and commercial decisions about what to build remain commercial. The useful framing is that future-readiness here is almost entirely present-readiness.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for semiconductor, manufacturing and energy operators — trace-data pipelines, control-loop integration, scheduling and decision support wired into the MES, APC and dispatch layer, with the decision rights, escalation paths and audit trails designed alongside the models rather than retrofitted after them.

  • · Equipment-integration work against SECS/GEM and SEMI EDA interfaces
  • · Decision-rights and escalation design run jointly with process, industrial-engineering and quality teams
  • · Integration-first delivery: bounded write-back, reconciliation into the MES record, tested revert
  • · 33 cited sources on this page

Sources

  1. GlobalFoundriesDigital manufacturing — factory control tower, AI centre of excellence, material handling (opens in a new tab)
  2. GlobalFoundriesQuality and certifications (opens in a new tab)
  3. GlobalFoundriesNewsroom (opens in a new tab)
  4. BoschDresden semiconductor plant — connected production, AI in manufacturing, digital twin (opens in a new tab)
  5. World Economic ForumGlobal Lighthouse Network (opens in a new tab)
  6. FlexcitonThe decisions that nobody owns: closing the gap between dispatching and planning (opens in a new tab)
  7. FlexcitonTechnical insight: solving the queue timer conundrum (opens in a new tab)
  8. FlexcitonSolving the Q-timer conundrum: improving yield without sacrificing capacity (opens in a new tab)
  9. Flexciton2025 front-end fab insights report (opens in a new tab)
  10. FlexcitonThe road to fab autonomy (opens in a new tab)
  11. arXiv (preprint)Industry 4.0: contributions of holonic manufacturing control architectures and future challenges (opens in a new tab)
  12. arXiv (preprint)Proposition of an implementation framework enabling benchmarking of holonic manufacturing systems (opens in a new tab)
  13. arXiv (preprint)An analysis of the coordination gap between joint and modular learning for job shop scheduling with transportation resources (opens in a new tab)
  14. arXiv (preprint)Multi-agent decision transformers for dynamic dispatching in material handling systems (opens in a new tab)
  15. arXiv (preprint)Coordinating from memory: graph-structured experience reuse for multi-agent adaptation in dynamic manufacturing (opens in a new tab)
  16. arXiv (preprint)Hybrid ASP-based multi-objective scheduling of semiconductor manufacturing processes (opens in a new tab)
  17. arXiv (preprint)Semiconductor fab scheduling with self-supervised and reinforcement learning (opens in a new tab)
  18. arXiv (preprint)On scheduling a photolithography process containing cluster tools (opens in a new tab)
  19. arXiv (preprint)Flocking behavior: an innovative inspiration for the optimization of production plants (opens in a new tab)
  20. arXiv (preprint)Applications of federated learning in manufacturing (opens in a new tab)
  21. SEMIStandards programme (opens in a new tab)
  22. SEMIIndustry association home (opens in a new tab)
  23. imecResearch and expertise (opens in a new tab)
  24. imecExpertise areas (opens in a new tab)
  25. ASMLTechnology (opens in a new tab)
  26. Lam ResearchDeposition, etch and clean systems (opens in a new tab)
  27. KLAProcess control, inspection and metrology (opens in a new tab)
  28. Applied MaterialsSemiconductor equipment and automation software (opens in a new tab)
  29. TSMCCompany and technology news (opens in a new tab)
  30. IntelNewsroom (opens in a new tab)
  31. Micron TechnologyNewsroom (opens in a new tab)
  32. Semiconductor EngineeringManufacturing and process coverage (opens in a new tab)
  33. McKinsey & CompanySemiconductors insights (opens in a new tab)

Find out where your decision authority actually sits — then move one decision

We run the assessment with your process, industrial-engineering, production-control and quality leads, walk your real decision list through the authority map on this page, and leave you with a marked-up register plus a costed 90-day plan for the first devolvable decision. You keep the register and the plan whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.