Manufacturing (Non-Automotive)Readiness & Transformation Roadmap
AI readiness for ESG in manufacturing: the evidence chain behind every disclosed number
AI readiness for ESG in manufacturing is the condition in which a plant's environmental, social and governance data is metered, allocated to production and traceable to source — so a model can join the calculation path without lowering the assurance grade of the disclosed number. It is an evidence problem before it is a modelling one.

Key takeaways
- An ESG number is worth what its weakest input can survive. Grade every input A to D — metered primary, activity data times a published factor, supplier-specific secondary, or modelled proxy — and the readiness question stops being philosophical: what share of the disclosed figure is grade A, and can you prove it?
- AI's legitimate first job in manufacturing ESG is to move data up the grade ladder and to label what stays modelled. Extraction from supplier documents, anomaly detection on meter series, factor matching and bounded gap-filling all do that. Generative drafting of the disclosure narrative does not, and is the stage-2 distraction.
- The binding constraint in almost every plant is allocation, not measurement. Sub-meters usually exist; the join between the 15-minute interval series and the MES production order does not, so no figure can be cut by line, order or SKU — which is exactly the resolution CBAM, ESPR and customer PCF requests demand.
- Assurance readiness is measurable as trace time: pick a disclosed figure at random and time how long it takes to reach the meter reading or supplier document behind it. Weeks means the number is rebuilt rather than traced. Minutes means you have a lineage record, and limited assurance becomes a sampling exercise.
- Freeze the methodology and the factor set per reporting period, and treat any change as a labelled restatement. Silently recomputing history when a factor library updates is the single most common way a manufacturer turns a defensible number into an audit finding.
Abbreviations used on this page
- CSRD
- Corporate Sustainability Reporting Directive (EU)
- ESRS
- European Sustainability Reporting Standards (E1 is the climate standard)
- CBAM
- Carbon Border Adjustment Mechanism (EU import carbon regime)
- ESPR
- Ecodesign for Sustainable Products Regulation (EU)
- DPP
- Digital Product Passport (the data carrier ESPR introduces)
- PCF
- Product carbon footprint — kg CO₂e per unit of product
- LCA
- Life-cycle assessment
- EPD
- Environmental product declaration (a verified supplier LCA summary)
- MES
- Manufacturing execution system — the production-order system of record
- EMS
- Energy management system (the ISO 50001 construct and its software)
- ERP
- Enterprise resource planning
- BOM
- Bill of materials
Free · 8 questions · ~3 minutes
Score your ESG evidence chain
Eight questions, one at a time, about three minutes. They ask what actually produces your disclosed figures — the meters, the factors, the allocation rules and the controls — and place you on the five-stage evidence ladder. Your result names the weakest of the four dimensions, which is the one that caps the grade of every number you publish.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised readiness report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the specific evidence gaps that cap your grade, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Invoice-grade
ESG numbers are assembled once a year from utility invoices, the purchase ledger and supplier emails, at group or site level, in spreadsheets.
Your next movePick one site and one utility. Get interval metering into a system that keeps history, and freeze a written methodology for that one number.
Stage 2 · Metered
Site-level primary data exists — sub-meters, an EMS, weighbridge tickets — but it lives in the energy team's tools and is reconciled to the disclosure by hand.
Your next moveJoin one interval meter to one line's MES production orders and agree the allocation rule for shared utilities. One line, not one site.
Stage 3 · Allocated
Environmental data is joined to production data, so energy, emissions, water and waste can be cut by line, production order and SKU, and rebuilt from raw.
Your next moveMake lineage a stored artefact. Every disclosed figure carries its grade, its factor version and a pointer to the raw record that produced it.
Stage 4 · Assured
Every disclosed figure carries its evidence grade, factor version and lineage, so assurance is a sampling exercise — and any model in the calculation path is registered, bounded and disclosed.
Your next movePut one assured number into a live operating decision — scheduling, a setpoint, a supplier award — with a holdout so the change is attributable.
Stage 5 · Steering
The same assured dataset drives operating decisions — scheduling against carbon intensity, setpoints, supplier awards, product design — with models recommending and, inside stated bounds, acting.
Your next moveSeparate the steering copy from the disclosed copy with a scheduled reconciliation, and version the decision policy with the same rigour as the model.
0 / 24
Evidence quality
— / 6
Allocation and traceability
— / 6
Assurance and control
— / 6
Decision use
— / 6
Your score maps to a stage on the evidence ladder. The dimension breakdown matters more than the total: the lowest dimension caps the grade of every figure you disclose, and it is where the next investment belongs regardless of how strong the other three look. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the evidence ladder. The dimension breakdown matters more than the total: the lowest dimension caps the grade of every figure you disclose, and it is where the next investment belongs regardless of how strong the other three look.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this tested against your actual reporting close?
We sit with your reporting controller and your plant energy engineer, sample real disclosed figures, time the trace on each one, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan and the sample results either way.
How the score maps to a stage
- 0–5 — Stage 1, Invoice-grade. ESG numbers are assembled once a year from utility invoices, the purchase ledger and supplier emails, at group or site level, in spreadsheets.
- 6–11 — Stage 2, Metered. Site-level primary data exists — sub-meters, an EMS, weighbridge tickets — but it lives in the energy team's tools and is reconciled to the disclosure by hand.
- 12–16 — Stage 3, Allocated. Environmental data is joined to production data, so energy, emissions, water and waste can be cut by line, production order and SKU, and rebuilt from raw.
- 17–21 — Stage 4, Assured. Every disclosed figure carries its evidence grade, factor version and lineage, so assurance is a sampling exercise — and any model in the calculation path is registered, bounded and disclosed.
- 22–24 — Stage 5, Steering. The same assured dataset drives operating decisions — scheduling against carbon intensity, setpoints, supplier awards, product design — with models recommending and, inside stated bounds, acting.
What AI readiness for ESG means in a manufacturing plant
A definition, the four evidence grades every ESG input falls into, and the chain a number has to survive between a meter and an assurance opinion.
AI readiness for ESG in manufacturing is the condition in which environmental, social and governance data is metered, allocated to production and traceable to source, so that a model can join the calculation path without lowering the assurance grade of the disclosed number. It is deliberately a narrower question than general AI readiness. The models involved are unexceptional — extraction, classification, anomaly detection, gap-filling — and the difficulty sits entirely in what they are being asked to touch: figures that will be published, verified by a third party, quoted to customers and, for some product categories, declared to a customs authority.
The organising idea on this page is the evidence grade. Every input to an ESG figure sits in one of four grades, from a metered primary reading down to a modelled proxy, and a composite figure is worth what its weakest material input can survive. That single framing settles most of the arguments a readiness review runs into. It tells you which data to fix first, what a model may legitimately do to each class of input, and what must never happen — a modelled value presented as if it had been measured. It also converts an abstract readiness question into an arithmetic one: what share of this figure is grade A, and can you prove it?
| Grade | What it is | Typical manufacturing source | What it survives | What AI may legitimately do |
|---|---|---|---|---|
| A · Metered primary | A physical reading tied to a timestamp and an asset | Sub-meters, flow meters, weighbridge and waste tickets, DCS and PLC tags | Limited and reasonable assurance; CBAM verification; a customer's own audit | Detect faults, outliers and drift; reconcile meters against each other — never generate the value |
| B · Activity × published factor | A measured mass, volume or run time multiplied by a documented factor | MES consumption records, BOM quantities, ERP goods receipts | Limited assurance where the factor is versioned, dated and sourced | Match a material to the closest available factor and keep the register current — a human confirms each match |
| C · Supplier-specific secondary | A supplier's own declared figure for a purchased item | EPDs, supplier PCF sheets, questionnaire returns, contractual disclosures | Limited assurance for Scope 3 category 1 where the document is retained and current | Extract from PDFs, normalise units, flag boundary mismatches — a human verifies before the value is used |
| D · Modelled or proxy | An estimate: spend-based factors, sector averages, imputed gaps | Purchase ledger, industry databases, interpolation across missing periods | Disclosure only, and only where it is explicitly labelled as estimated | Fill gaps with an explicit uncertainty band and a label — never present the result as measured |
Grades are not a maturity score in disguise. A mature manufacturer still discloses grade D figures — a small, immaterial purchased category is not worth a supplier engagement programme — and an immature one occasionally has excellent grade A metering on the one utility that dominates its cost base. What changes with maturity is whether the mix is known, whether it is disclosed, and whether the grade travels with the number when it moves between systems. That last property is what makes the difference between a chain a model can safely join and one it cannot, because a model that cannot see the grade of its inputs will happily average a meter reading with a spend-based estimate and produce something that looks like a measurement.
The ESG evidence chain, and where a model is allowed to touch it
Three lanes, one destination. The top lane is where primary data is produced and — in most plants — where it stops. The middle lane is where the largest share of the footprint comes from and where the grade is lowest. The bottom lane is where the methodology is set and where the whole chain is finally tested. The dashed red path is the failure mode this page exists to prevent.
- Data & feeds
- System-of-record action
- AI / model
- Human in the loop
- Where value leaks
The process, in words
- Site lane: meters and process instruments produce interval readings, a historian or energy management system retains them raw, and the allocation step joins that series to MES production orders using written rules for shared utilities. What comes out is kWh and kg CO₂e per production order — grade A data at the resolution CBAM, product passports and customer requests all ask for. In most plants this lane stops at the historian.
- Chain lane: purchase and BOM records say who supplied what; supplier documents — EPDs, PCF sheets, questionnaire returns — arrive as PDFs in a dozen layouts. An extraction model reads them and proposes structured entries, a named human verifies each one, and the result lands in a factor register where every factor carries a grade, a source and a date. Gaps that remain are filled with an explicitly labelled estimate and an uncertainty band.
- Methodology lane: the boundary, the allocation rules and the factor set are frozen at the start of the reporting period and constrain both other lanes — the allocation rule governs the site join, the gap-fill bounds govern what the model is permitted to invent. The disclosed figure carries the grade mix forward, and the assurance opinion is formed by sampling it back to raw.
- The failure path, dashed in red: a modelled value loses its label somewhere between the model and the disclosure, is aggregated with metered data as if it were equivalent, and is published as measured. Nothing about that figure is wrong until someone asks how it was produced — at which point it cannot be traced, and it is restated.
Step-by-step insights
- Meters and instruments — retention is the decision that matters
- The metering itself is rarely the constraint; retention is. Many energy management systems are configured to keep raw interval data for months and monthly rollups forever, because that is what a utility-cost use case needs. The moment you want to allocate to production orders, rollups are useless: a production order lasts hours, not months, and no monthly total can be decomposed into one. Check the retention policy before the meter list. A plant with fewer meters and full interval history is closer to grade A allocation than a heavily instrumented plant that throws the detail away.
- The allocation join — one rule per shared utility, written down
- This is the single highest-leverage piece of work on the page and it is mostly adjudication rather than engineering. Which meter serves which line. How compressed air is apportioned across three lines sharing a compressor — run hours, nameplate, or measured flow. Where changeover and clean-down energy goes, given it belongs to the product but sits in no production order. Whether waste is allocated by mass or by cost. Each rule needs to be decided once, written down, and applied consistently, because an assurer will ask and a customer comparing two suppliers' footprints will notice if the rules differ. The join itself lives at the production-order level, which is ISA-95 level 3 — the layer the MES already owns.
- Supplier documents — extraction is where AI genuinely pays
- A mid-size manufacturer with a few thousand purchased line items may receive supplier environmental data in twenty formats: EPDs following different product category rules, PCF sheets with different boundaries, questionnaire returns in spreadsheets, and a long tail of PDFs. Extracting these by hand is the reason Scope 3 stays grade D. A model that reads the document, proposes a structured entry and flags a boundary mismatch — cradle-to-gate versus cradle-to-grave, mass versus unit basis — turns weeks of analyst time into a review queue. The control that makes it admissible is the verification step: a named human confirms each entry, and the original document is retained and linked, so the assurer samples the document rather than the model.
- Gap-fill — bounded, labelled, and never silently promoted
- Every real inventory has holes: a supplier that will not respond, a month of missing sub-meter data, a new material with no factor. Filling them is legitimate and expected — what matters is that the fill is bounded by a rule set in the methodology register, carries an uncertainty band, and keeps its grade D label everywhere downstream. The failure is not the estimate; it is the estimate that arrives at the disclosure indistinguishable from a meter reading. Set a ceiling on the modelled share of any material figure and alert when it is approached, because gap-fill expands quietly: each individually reasonable fill moves the mix one notch, and nobody is watching the aggregate.
- The methodology register — freeze the period, log the change
- The methodology register holds the boundary, the consolidation approach, the allocation rules and the factor set, and its defining property is that it is frozen for the duration of a reporting period. Factor libraries update; grid emission factors are restated by their publishers; a supplier issues a revised EPD. None of those may change a period already in flight. New values apply prospectively, and if a change is material enough to warrant restating history, that restatement is explicit and logged. Manufacturers who let libraries update in place discover it the following year, when last year's published figure no longer reproduces and no one can explain the delta.
- The assurance opinion — a sample, not a rebuild
- The whole chain exists to make one interaction cheap. An assurer names a figure and asks how it was produced; either the answer is a lineage record returned in minutes, or it is three weeks of reconstruction. That difference is what separates a stage-4 manufacturer from a stage-3 one, and it is directly measurable as trace time. It is also the reason to build the lineage store while building the pipeline: retrofitting lineage onto a calculation that already runs means reconstructing provenance for figures whose inputs have since changed, which is strictly harder than recording it as you go.
Real today, and in production
Extraction of supplier EPDs and PCF sheets into a structured, human-verified register. Anomaly and drift detection on meter series, which catches a stuck sensor in a day rather than at quarter end. Matching BOM materials to the nearest available factor with a confidence score. Bounded gap-filling with an explicit uncertainty band. Drafting disclosure narrative from a locked, already-computed dataset. Every one of these is ordinary engineering with an ordinary control around it.
Claimed, and true only under conditions
Real-time Scope 3, continuous product footprints and automated supplier engagement are all achievable — where suppliers actually exchange primary data, which for most manufacturers is a minority of spend today. The honest version of the claim is that the machinery works and the inputs do not exist yet. Buy it as a destination, sequence it behind a supplier data programme, and do not let a demonstration on a curated dataset set the expectation for your own long tail.
Not credible, and a live compliance risk
A model that infers your Scope 3 from your purchase ledger and presents the output as a measured inventory. Assurance evidence generated rather than retrieved. A product footprint produced without a stated boundary or allocation rule. Anything that removes the grade label from a value on its way to a disclosure. These are not conservative objections about model quality — they are the specific artefacts that turn a reporting exercise into a misstatement, and they are the reason the register and the labelling exist.
What the evidence ladder releases, stage by stage
Value stays close to flat through the first two stages — the data exists but cannot be cut, so nothing outside the reporting function can use it — and inflects at allocation, when a figure first becomes something a customer, a planner or a designer can act on. This is why programmes measured in meters installed rather than figures allocated report activity without results.
Usable, defensible ESG value released by stage
- Stage 1 · Invoice-grade — 22% of operators. ESG numbers are assembled once a year from utility invoices, the purchase ledger and supplier emails, at group or site level, in spreadsheets.
- Stage 2 · Metered — 37% of operators. Site-level primary data exists — sub-meters, an EMS, weighbridge tickets — but it lives in the energy team's tools and is reconciled to the disclosure by hand.
- Stage 3 · Allocated — 27% of operators. Environmental data is joined to production data, so energy, emissions, water and waste can be cut by line, production order and SKU, and rebuilt from raw.
- Stage 4 · Assured — 11% of operators. Every disclosed figure carries its evidence grade, factor version and lineage, so assurance is a sampling exercise — and any model in the calculation path is registered, bounded and disclosed.
- Stage 5 · Steering — 3% of operators. The same assured dataset drives operating decisions — scheduling against carbon intensity, setpoints, supplier awards, product design — with models recommending and, inside stated bounds, acting.
Curve shape: logistic, plotted from the stage data above. Distribution: Shape consistent with acatech's Industrie 4.0 Maturity Index stage progression.
The five stages of an ESG evidence chain
For each stage: what it looks like on a real plant, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps manufacturers there, and what leaving costs.
The five stages describe how far an ESG number has travelled from an invoice to an operating decision, and each one is defined by what the number can survive rather than by how much technology sits behind it. The ladder runs invoice-grade, metered, allocated, assured, steering. It is written for practitioners: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
| Stage | Smallest unit available | Typical grade mix | What the figure survives |
|---|---|---|---|
| 1 · Invoice-grade | Site, financial year | Mostly D, some B | An internal report; a first-year assurance engagement with findings |
| 2 · Metered | Site, month | A for site energy, D for purchased inputs | An energy-cost conversation; still no answer to a customer's product question |
| 3 · Allocated | Line, shift, production order | A for energy, B–C for materials, D labelled | A customer PCF request, with follow-up questions answered |
| 4 · Assured | Production order, with lineage | Known and disclosed mix | A sampling assurance engagement; a CBAM verification; a restatement handled cleanly |
| 5 · Steering | Batch or decision, near real time | Known mix, reconciled between views | All of the above, plus use as a live constraint without corrupting the disclosed copy |
Select a stage
Every stage's full detail is present in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Invoice-grade
22% of operators sit here
ESG numbers are assembled once a year from utility invoices, the purchase ledger and supplier emails, at group or site level, in spreadsheets.
Stage 1 is not a failure of intent. Most manufacturers arrive here because the first reporting obligation was met the way every first obligation is met — by a small team, in a spreadsheet, against a deadline. The numbers are usually defensible in the narrow sense that each one traces to an invoice. What is missing is resolution and repeatability: the figure exists for the group and the year, and reproducing it next year means repeating the same manual assembly rather than re-running anything.
The tell is the unit of account. At stage 1 the atom is a bill. Electricity is what the utility charged, gas is what the meter reader recorded at the boundary, and the emissions attached to purchased materials are the ledger spend multiplied by a sector-average factor. None of those can be divided by a product, a line or a shift, because the underlying record was never associated with production in the first place. Every downstream question — what does this SKU emit, which line is the intensity outlier, what did the efficiency project actually save — is unanswerable by construction.
This is a cheap stage to leave and an expensive one to stay in, and the expense is not the reporting effort. It is that the ESG number never becomes an operational number, so it never earns operational attention, so the metering and allocation work that would fix it never gets funded. The loop closes on itself. Manufacturers who sit here for three reporting cycles usually find that the workbook has grown, the assurance scope has widened, and the underlying data is exactly where it was.
In practice
The February scramble
A building-products manufacturer with eleven plants closes its sustainability data every February. Two analysts email each plant controller for twelve utility invoices, a waste contractor summary and a water bill, retype the figures into a master workbook, and apply a factor set downloaded the previous spring. Scope 3 is the purchase ledger sliced by commodity code and multiplied by spend-based factors. The pack is good, the assurer's findings are manageable, and not one figure in it can be attributed to a production line.
What it looks like
- Scope 1 and 2 come from utility invoices; Scope 3 from spend times an average factor
- The smallest unit anyone can produce is a site and a financial year
- The methodology lives in the workbook and in the head of whoever built it
- No model touches the numbers, because there is no dataset for one to touch
Diagnostic signals you can check this week
- Ask for one plant's figure for one month. If the answer requires the analyst who built the workbook, you are here
- Ask which factor set version was used last year and this year. A screenshot or a filename is the stage-1 answer
- Count the spreadsheets between a meter and a disclosed figure. More than two is diagnostic
- Ask what operating decision would change if the number moved 10%. If none, the number is a reporting artefact
Anti-pattern · Buying an ESG platform to fix a data problem
The instinctive move is to procure a reporting platform, on the theory that the software will supply the rigour. It cannot: the platform ingests exactly the invoices the spreadsheet ingested, and adds a workflow, an audit log of who typed what, and a licence fee. What it does not add is a meter, a production-order join, or a graded factor register. Buy the platform after the evidence chain exists and it will save real effort; buy it before, and you have paid to industrialise the estimate.
What holds you here
There is no primary data and no production context, so every figure is an estimate assembled by hand and nothing can be automated without automating the estimate.
Highest-leverage next move
Pick one site and one utility. Get interval metering into a system that keeps history, and freeze a written methodology for that one number.
Cost of leaving
- Effort
- 3–6 months
- Team
- One plant energy engineer and one reporting controller, part-time
- Risk
- Low — nothing currently disclosed depends on the new work yet
- To next stage
- 3–6 months
If this is you, the next step is
A two-week review: which meters exist, which are missing, what one line would cost to instrument.
Stage 2
Metered
37% of operators sit here
Site-level primary data exists — sub-meters, an EMS, weighbridge tickets — but it lives in the energy team's tools and is reconciled to the disclosure by hand.
Stage 2 is the most populated stage and the most misread. There is genuine primary data: half-hourly electricity, gas by meter, often steam and compressed air, sometimes water and effluent, retained in an energy management system bought for ISO 50001 or for a utility-cost programme. The plant energy engineer can tell you the baseload, the weekend draw and which compressor is drifting. By any reasonable definition the measurement problem is solved.
What is not solved is the join. The interval series is indexed by meter and timestamp; production is indexed by order, product and shift in the MES. Nobody has written the rule that connects them, so the ESG number continues to be built from monthly totals — the one shape the interval data can be reduced to without needing production context. The energy team's rich dataset and the reporting team's thin dataset coexist for years, and each team assumes the other has already solved the part it can see.
The organisational consequence matters more than the technical one. Because no figure can be cut by product or order, no commercial conversation can use it. Sales cannot answer a customer's product carbon footprint request, procurement cannot price a low-carbon variant, and the plant cannot rank its lines by intensity. The ESG dataset stays inside the sustainability function, which is why it stays underfunded, which is why the join never gets built.
In practice
Two teams, two datasets, one plant
At a speciality chemicals site the energy manager has four years of half-hourly electricity by substation and can show you the exact hour a dryer was left running over a bank holiday. The group reporting team, three floors up, discloses annual site electricity from the utility portal. When a customer asks for the cradle-to-gate footprint of one product family, neither dataset answers: one has no product, the other has no resolution. The work to connect them is two weeks and has never been anyone's objective.
What it looks like
- Interval data from electricity, gas, steam, compressed-air and water meters is collected and retained
- The energy team can answer questions the reporting team cannot, and vice versa
- Reporting still consumes monthly totals, not the interval series
- AI, where present, drafts narrative or classifies invoices — nothing in the calculation path
Diagnostic signals you can check this week
- Ask the energy team and the reporting team for the same site's annual electricity. Compare the two numbers and, more revealingly, the two methods
- Ask whether any interval series has ever been joined to a production order. Not 'could it' — has it
- Check whether the EMS retains raw interval data or only monthly rollups. Rollups cap you here permanently
- Ask who would notice if a sub-meter flatlined for a fortnight. Silence is the stage-2 answer
Anti-pattern · Adding more meters instead of using the ones you have
When the ESG number still looks coarse, the reflex is to extend metering — more sub-meters, more points, a bigger capital request. It is the wrong order. An unallocated meter adds a series nobody can attribute; a meter joined to production changes what the business can sell and schedule. Do the join on the coverage you already have, discover which three missing points actually block a product-level figure, and buy those. The metering plan written after the first allocation is a fraction of the one written before it.
What holds you here
Metered data is not joined to production, so no figure can be cut by line, order or product — and every obligation that matters now asks for exactly that cut.
Highest-leverage next move
Join one interval meter to one line's MES production orders and agree the allocation rule for shared utilities. One line, not one site.
Cost of leaving
- Effort
- 4–8 months
- Team
- One data engineer, the plant energy engineer, an MES or IT contact
- Risk
- Medium — the first join changes numbers people have already disclosed
- To next stage
- 4–8 months
If this is you, the next step is
The stage 2→3 move on a single line, typically inside a quarter.
Stage 3
Allocated
27% of operators sit here
Environmental data is joined to production data, so energy, emissions, water and waste can be cut by line, production order and SKU, and rebuilt from raw.
Stage 3 is the first stage where the ESG dataset can answer a commercial question. Once metered consumption is allocated to production orders, cradle-to-gate footprints become arithmetic rather than a project: the BOM supplies the purchased inputs, the factor register supplies their emissions, the allocated energy supplies the conversion step, and a per-unit figure falls out. The same join answers the intensity question by line and shift, which is the number a plant manager can actually act on.
The work is unglamorous and mostly about rules. Which meter serves which line. How compressed air is apportioned when three lines share a compressor — by run hours, by nameplate, by measured flow. What happens to the energy consumed during changeover and cleaning, which belongs to the product but is not in any production order. Whether waste is allocated by mass or by cost. Every one of these is a decision that has to be written down, defended once, and then applied consistently, because an assurer will ask and a customer comparing two suppliers' footprints will care.
This is also where AI first belongs in the chain, and where its role should be narrowly drawn. Supplier EPDs and PCF sheets arrive as PDFs in a dozen layouts; extraction into a structured factor register with a human verifying each entry is a genuine order-of-magnitude saving. Meter series drift, stick and flatline; anomaly detection catches that before a quarterly review does. Materials on a BOM need matching to the closest available factor; a model proposes, a human confirms. What must not happen is a model producing a value that is then disclosed as if it had been measured.
In practice
The first product footprint that survived a customer
An industrial pump manufacturer joined half-hourly electricity from three cells to MES production orders, apportioned compressed air by measured flow, and matched every line on the BOM of one high-runner family to a factor with a stated source. The resulting cradle-to-gate figure went to a customer who came back with two questions — the allocation rule for the shared paint line, and the date of the aluminium factor. Both had answers. That exchange, not the number itself, is what stage 3 buys.
What it looks like
- kWh and kg CO₂e are available per production order, not only per site per month
- A written allocation rule exists for shared utilities — compressed air, chilled water, HVAC
- Emission factors sit in a register with a source and a date, not in a spreadsheet column
- Model-assisted work has started where it belongs: extraction, anomaly detection, factor matching
Diagnostic signals you can check this week
- Ask for kWh per unit for one SKU last month. A number in minutes is stage 3; a project plan is stage 2
- Ask to see the allocation rule for a shared utility in writing. Verbal consensus is not a rule
- Check whether the factor register records a source and a date per factor, and who last changed each one
- Ask what happens to a disclosed figure when a factor library updates. 'It just changes' is the finding waiting to happen
Anti-pattern · Chasing resolution before grade
Allocation makes finer cuts possible, and the temptation is to push resolution everywhere — per batch, per machine, per minute — while the inputs are still mostly modelled. A per-batch figure built on a spend-based Scope 3 factor is not more accurate than the annual one; it is the same estimate with more decimal places and more surface area for an assurer to sample. Raise the grade of the largest contributors first, then raise the resolution. Precision that outruns evidence is the fastest route to a restatement.
What holds you here
The chain from raw reading to disclosed figure is reproducible but not evidenced: the trace lives in scripts and people's heads, not in a record an assurer can sample.
Highest-leverage next move
Make lineage a stored artefact. Every disclosed figure carries its grade, its factor version and a pointer to the raw record that produced it.
Cost of leaving
- Effort
- 6–12 months
- Team
- Data engineer, plant energy engineer, LCA or sustainability analyst, MES owner
- Risk
- Medium — the first allocated figures will disagree with previously disclosed ones
- To next stage
- 6–12 months
If this is you, the next step is
A working session on shared utilities, changeover and waste — the three rules that get argued about.
Stage 4
Assured
11% of operators sit here
Every disclosed figure carries its evidence grade, factor version and lineage, so assurance is a sampling exercise — and any model in the calculation path is registered, bounded and disclosed.
Stage 4 changes what assurance costs and what it means. Below it, an assurance engagement is an evidence hunt: the assurer names a figure, the team goes looking, and three weeks disappear into reconstructing a calculation that was never designed to be reconstructed. At stage 4 the same request is a query. The figure carries a lineage record — which raw readings, which factor version, which allocation rule, which reviewer — and the engagement becomes what it is supposed to be, a sample of a controlled process rather than a rebuild of an uncontrolled one.
The discipline that makes this work is period freezing. The methodology, the boundary and the factor set are fixed at the start of a reporting period and cannot change inside it. When a factor library publishes an update, the new values apply prospectively; if the change is material enough to warrant restating history, that restatement is explicit, labelled and logged. Manufacturers who let a factor library update recompute prior periods in place discover the problem the following year, when last year's disclosed figure no longer reproduces and nobody can say why.
This is also the stage where AI has to be governed rather than merely used. Any model contributing to a disclosed number needs a register entry: what it does, which version, what inputs, who reviews its output, what the override rate is, and how its contribution is labelled in the disclosure. That is not a bureaucratic flourish — it is the same control an assurer applies to any other estimation technique, and it is increasingly what the AI-specific frameworks expect. The manufacturers who find this easy are the ones who wrote the register while building the pipeline rather than in the month before the audit.
In practice
The twenty-figure dry run
Before its first assured reporting cycle, a packaging manufacturer asked its internal audit team to pick twenty disclosed figures at random and trace each to source. Fourteen took under five minutes from the lineage store. Four took a day, all of them Scope 3 items where a supplier document existed as an email attachment rather than a register entry. Two could not be traced at all and were restated. The exercise cost a week and removed every material finding from the external engagement that followed.
What it looks like
- A sampled figure traces to a meter reading or supplier document in minutes, on demand
- The factor set and methodology are frozen per reporting period, with a change log
- A written restatement policy exists and has been used at least once
- Model-derived fields are flagged, version-pinned and reviewed, with the override log retained
Diagnostic signals you can check this week
- Pick a disclosed figure at random and time the trace to raw. Median trace time is the stage-4 metric
- Ask whether last year's figure still reproduces exactly from this year's system. If not, factors are updating in place
- Ask to see the model register entry for any AI in the reporting chain, including the override log
- Ask when the restatement policy was last used. A policy never exercised is a document, not a control
Anti-pattern · Treating assurance as the destination
Reaching assurance-grade data is a large achievement and a natural place to stop, which is exactly the trap. An assured dataset that changes no operating decision is a cost centre with excellent documentation, defended annually against a finance function that can see its price and not its return. The transition out of stage 4 is not more control — it is putting one assured number into a live decision, with a holdout, so somebody outside the sustainability function has a reason to care whether it is right.
What holds you here
The assured dataset is still an output. Nothing operational changes because of it, so the apparatus is a cost centre defended annually rather than a capability anyone fights for.
Highest-leverage next move
Put one assured number into a live operating decision — scheduling, a setpoint, a supplier award — with a holdout so the change is attributable.
Cost of leaving
- Effort
- 12–18 months
- Team
- Reporting controller, data engineer, internal audit partner, model owner
- Risk
- Higher — the control framework, not the pipeline, becomes the binding constraint
- To next stage
- 12–18 months
If this is you, the next step is
We sample twenty figures against your own systems and report trace time and gaps.
Stage 5
Steering
3% of operators sit here
The same assured dataset drives operating decisions — scheduling against carbon intensity, setpoints, supplier awards, product design — with models recommending and, inside stated bounds, acting.
Stage 5 is narrower than the phrase suggests and should stay that way. It is not an autonomous sustainability function; it is a small, enumerated set of decisions in which an environmental figure is a live constraint or objective. Scheduling energy-intensive batches against a forecast grid carbon intensity qualifies. Setpoint optimisation on a dryer or a kiln where energy is the dominant cost qualifies. Choosing between two qualified suppliers on a blended price-and-footprint score qualifies. Anything touching product safety, permitted emissions limits or a regulatory declaration stays under human decision indefinitely, and that is the correct answer rather than a stage not yet reached.
The engineering is largely done by the time an operator arrives here; what is new is a control problem with a specific shape. A number that steers operations and is also disclosed has become both a management KPI and a reported figure, and the incentive to shade it now exists inside the plant rather than only at group level. The structural answer is separation with reconciliation: one calculation engine, a steering view that can be provisional and fast, a disclosure view that is frozen and slow, and a scheduled reconciliation whose divergences are investigated rather than explained.
Regression is the standing risk. Grid factors change, product mixes shift, a new line arrives with no allocation rule, and a decision policy tuned on last year's conditions quietly stops being valid. The leading indicator is the same one autonomy programmes use everywhere: the rate at which decisions fall outside their stated bounds and escalate. When it rises, the world has moved outside the policy before an incident says so.
In practice
The batch that moved to Tuesday
A food manufacturer with a large refrigerated drying load schedules its most energy-intensive campaigns against a day-ahead grid carbon-intensity forecast, inside bounds set by the production planner: never at the cost of a customer commitment, never more than a stated shift in the published plan, always reversible by one action. Roughly one campaign in six is moved. The reported saving is computed against a holdout of comparable campaigns left on the original schedule, and the steering figure reconciles monthly to the disclosed one.
What it looks like
- At least one scheduled decision carries an environmental constraint or objective, not just a report
- Design and procurement see product-level footprints at the point of decision
- The steering copy and the disclosed copy of every figure reconcile on a stated cadence
- The decision policy — what may execute unattended, within what bounds — is versioned and reviewed
Diagnostic signals you can check this week
- Name a decision in the last quarter that changed because of an environmental figure, and the holdout it was measured against
- Ask whether the steering figure and the disclosed figure have ever diverged, and what happened when they did
- Check whether the decision policy is versioned and reviewed, or edited in a settings screen
- Ask what the escalation rate has done over the last six months. Nobody watching it is the regression signal
Anti-pattern · Letting the steering number become the reported number
Once a provisional figure is good enough to schedule against, the temptation is to disclose it — it is fresher, it is already in front of people, and maintaining two views feels like duplication. It is not duplication; it is the control. The steering view is allowed to be fast and revisable, which is precisely what a disclosed figure must not be. Collapse them and the first time a provisional value is restated after publication, the whole dataset's credibility is spent.
What holds you here
Sustaining it is a control problem: a figure that steers operations and is also disclosed is both a target and a measure, and the policy behind it drifts out of validity without anyone noticing.
Highest-leverage next move
Separate the steering copy from the disclosed copy with a scheduled reconciliation, and version the decision policy with the same rigour as the model.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team, production planning, plus a standing reporting and controls forum
- Risk
- Concentrated — low frequency, high consequence, and now visible to a regulator
If this is you, the next step is
We take one decision and test its bounds, its reconciliation and its rollback against a real scenario.
Where manufacturers actually sit on the ladder
The distribution across the five stages, and why the metered-to-allocated step is the largest single transition loss.
Most non-automotive manufacturers are at stage 2: they have real primary data on site energy and no way to attribute it to a product. The distribution below is weighted heavily toward that middle, and the gap between stage 2 and stage 3 is the largest single drop on the ladder — not because allocation is technically hard, but because it sits between two functions and belongs to neither. The energy team owns the meters, the MES team owns the production orders, and the join is nobody's objective until a customer or a customs declaration makes it one.
Illustrative distribution of manufacturers across the evidence ladder
Illustrative, not measured: a model-derived distribution synthesised from CDP's disclosure data on primary-data availability, acatech's Industrie 4.0 Maturity Index stage progression and the WEF Global Lighthouse Network's reporting on scaled deployments. Treat the shape as the argument, not the individual percentages.
Share of manufacturers
- 22% — 1 · Invoice-grade
- 37% — 2 · Metered (the plateau)
- 27% — 3 · Allocated
- 11% — 4 · Assured
- 3% — 5 · Steering
Source: Illustrative distribution, synthesised from CDP, acatech and World Economic Forum research
The imbalance is structural rather than cultural. CDP reports (opens in a new tab) that supply-chain Scope 3 emissions are on average 26 times a company's direct operational emissions, and that corporates are roughly twice as likely to measure operational emissions as supply-chain ones. For a manufacturer, that ratio maps almost exactly onto the evidence-grade picture: the smaller part of the footprint is grade A and the larger part is grade D. Any readiness programme that starts with the metered part is starting where the data already is rather than where the footprint is — which is defensible as a first step and indefensible as a destination. It is worth reading that alongside the IEA's tracking of industrial energy demand and emissions (opens in a new tab), which is a reminder that the absolute quantities involved are large enough that the resolution argument is not an accounting nicety.
The external anchors for the ladder itself are worth naming. acatech's Industrie 4.0 Maturity Index (opens in a new tab) describes the same shape for production data — computerisation, connectivity, visibility, transparency, predictive capacity, adaptability — and the ESG ladder is essentially that progression applied to environmental data with an assurance constraint bolted on. The World Economic Forum's Global Lighthouse Network (opens in a new tab) documents manufacturers that have scaled digital deployments across multiple sites, and its recurring finding — that scaling is an organisational problem long before it is a technical one — is precisely what the stage 2 plateau looks like from the inside. NIST's manufacturing programmes (opens in a new tab) and MHI's annual industry survey (opens in a new tab) track the same adoption-versus-impact gap across the wider industrial base.
What each obligation actually demands from the plant
CSRD and ESRS, CBAM, ESPR and the product passport, ISSB, ISO 50001 and the social and governance strands — the number each one wants, the grade it demands, and the system it has to come from.
Each reporting obligation resolves, in the plant, to a specific number at a specific resolution from a specific system — and they do not all ask for the same thing. That distinction is the whole planning problem. A manufacturer that reads CSRD, CBAM, ESPR and a customer questionnaire as four versions of 'report your emissions' will build four workbooks. A manufacturer that reads them as four cuts of one allocated dataset will build the allocation once and serve all four from it, which is the difference between a permanent reporting function and a capability.
| Obligation or driver | The number it demands | Grade demanded | System of record | Answerable from stage |
|---|---|---|---|---|
| ESRS E1 under CSRD | Scope 1, 2 and 3 by category, energy mix, targets and transition plan | A for Scope 1–2, B–C for material Scope 3 | EMS or historian, ERP, factor register | 3 |
| CBAM definitive regime | Embedded emissions per tonne of covered good, per consignment, verified | A for direct, B for precursors | MES, EMS, ERP and the customs declaration | 3 |
| ESPR and the digital product passport | Product-level attributes and footprint per item or batch | A–B at product resolution | PLM, MES, factor register | 4 |
| Customer PCF requests | Cradle-to-gate kg CO₂e per unit with a stated boundary | B, with grade A energy | PLM and MES, joined to the BOM | 3 |
| ISSB IFRS S1 and S2 | Financially material sustainability risk, on investor-grade controls | A–B with lineage | Group consolidation, on the same source data | 4 |
| ISO 50001 and ISO 14001 | Energy performance indicators, significant energy uses, objectives | A | EMS, normalised for output and weather | 2 |
| ESRS S1 — own workforce | Headcount, incident and injury rates, training hours, pay gap | A from HR and EHS systems | HRIS and EHS incident system | 2 |
| EU AI Act and AI management systems | An inventory of AI systems, risk classification, human oversight evidence | Register, not measurement | Model register and change control | 4 |
Three rows carry most of the sequencing weight. The CBAM definitive regime (opens in a new tab) has applied since 1 January 2026 across six goods sectors — cement, iron and steel, aluminium, fertilisers, electricity and hydrogen — and requires embedded-emissions data per consignment rather than per year, which is a resolution demand no site-and-year figure can be reshaped into. The Ecodesign for Sustainable Products Regulation (opens in a new tab) introduces the digital product passport, pushing the same requirement down to the individual item or batch. And CSRD reporting under the ESRS (opens in a new tab) brings assurance into the picture, which is what converts a data-quality preference into a control requirement. The scope of who must report has been actively debated and narrowed since the directive came into force; the resolution the underlying obligations demand has not moved.
A digital identity card for products, components, and materials, which will store relevant information to support products' sustainability, promote their circularity and strengthen legal compliance.
The accounting standards underneath all of this are stable and worth reading once rather than paraphrased forever. The GHG Protocol Corporate Standard (opens in a new tab) defines the organisational inventory and its five accounting principles; the Product Standard (opens in a new tab) and ISO 14067 define a product footprint and, critically, the allocation rules a manufacturer will otherwise argue about internally; and the Scope 3 calculation guidance (opens in a new tab) is explicit that supplier-specific data is preferred over averages, which is the formal statement of the grade ladder. The ISSB standards (opens in a new tab) layer investor-grade control expectations on the same numbers. ISO 14064-1 for organisational inventories, ISO 14064-3 for verification, ISO 50001 for energy management and ISO 14001 for environmental management are all catalogued at the ISO standards catalogue (opens in a new tab); none of them prohibits a model in the calculation path, and all of them require that the path be describable.
The social and governance strands are usually treated as a separate problem and should not be. ESRS S1 asks for workforce data — headcount, incident rates, training hours, pay gaps — which lives in HR and EHS systems that have their own grade problem: incident narratives are free text, contractor headcount is often reconstructed from site access logs, and training records span three systems. The temptation to point a language model at all of it is strong and the constraint is sharp, because worker data carries obligations that energy data does not. The governance strand closes the loop: under the EU AI Act's risk-based framework (opens in a new tab) and management-system standards such as ISO/IEC 42001, the models you deploy in the ESG chain are themselves in scope for inventory, classification and oversight evidence. NIST's AI Risk Management Framework (opens in a new tab) is the most practical free reference for structuring that register, and the human-rights and worker-data side is treated in depth in our manufacturing AI human-rights governance guide.
What the transitions look like in public
Three publicly reported manufacturing programmes, read against the evidence ladder. None is an Atomic Loops engagement — every figure is quoted from the operator's own published material.
Very few manufacturers publish the detail of how AI touches their ESG numbers, and that absence is itself informative: what they do publish is the evidence chain — the unit of account, the resolution, the reporting perimeter — which is exactly the thing that has to exist before any model matters. Read the three below for their data architecture rather than their targets. Each one shows a different rung of the ladder made visible in public disclosure.
Three manufacturers, read against the ladder
Outcomes and figures as published by the operators themselves; verify against the linked source before reusing them, as we have not independently audited them. Card images are generated industry scenes from our own library, not photographs of these companies, and imply no endorsement.
SiemensIndustrial technology and electronics · multi-site global estate24
- Challenge
- Reporting environmental performance consistently across a very large, heterogeneous plant estate, where each site had its own metering history, its own systems and its own definitions of the same quantity.
- Approach
- Siemens organises its sustainability work under a published framework it calls DEGREE, with targets set to 2030, and consolidates performance into a single Sustainability Statement covering strategy, governance, operationalisation and performance across its material topics.
- Reported outcome
- Siemens publishes a Sustainability Statement 2025 that it describes as delivering 'clear, comprehensive insights into our ESG strategy, governance frameworks, operationalization, and performance across all material sustainability topics', and states that more than 90% of its business enables customers to achieve a positive sustainability impact.
- What it shows about the curveThe visible artefact is a single consolidated statement across a heterogeneous estate. That is a stage-4 signature: it is only producible when every site's figures share definitions, boundaries and a controlled consolidation path, which is a data-governance achievement long before it is a reporting one.
HolcimBuilding materials · cement, aggregates and ready-mix concrete34
- Challenge
- In cement and concrete the product is the emission, and the unit of account is the tonne. A site-and-year figure is commercially useless when customers, specifiers and — under CBAM — customs authorities are asking about a specific material at a specific intensity.
- Approach
- Holcim reports environmental performance in per-tonne intensity terms alongside the commercial share of its lower-carbon product ranges, tying the environmental unit of account directly to the sales unit of account.
- Reported outcome
- Holcim publishes Scope 1 intensity of 502 kg CO₂ per tonne of cementitious material, down 11% against its 2020 baseline, against a 2030 target of under 400 kg; it reports ECOPlanet cement at a 36% share of cement net sales and ECOPact ready-mix at 31% of ready-mix net sales, with a target of over 50% of net sales from the two ranges by 2030, and freshwater withdrawal of 179 litres per tonne, 25% below its 2020 baseline.
- What it shows about the curveWhen the environmental figure and the revenue figure share a denominator, allocation stops being a reporting cost and becomes a commercial instrument. That alignment is what funds the stage 3 to 4 transition in materials businesses, and it is why CBAM-exposed sectors are ahead of the rest of manufacturing on product-level data.
HenkelAdhesives, consumer brands and industrial chemicals · global plant network34
- Challenge
- Moving sustainability data out of a standalone reporting cycle and into the same control environment, calendar and perimeter as financial reporting, across a plant network producing both industrial and consumer goods.
- Approach
- Henkel publishes under what it calls its 2030+ Sustainability Ambition Framework, covering the three ESG dimensions it names Regenerative Planet, Thriving Communities and Trusted Partner, and reports through a Sustainable Impact Report and a separate Sustainability Indicators set.
- Reported outcome
- Henkel publishes its Sustainability Statement inside its Annual Report — the report is listed as 'Annual Report 2025 incl. Sustainability Statement' — alongside a Sustainable Impact Report 2025 and Sustainability Indicators 2025.
- What it shows about the curvePutting the sustainability statement inside the annual report moves the data into the audited perimeter, on the financial calendar, under the same controls. That is the clearest external marker of stage 4 available to an outsider: the number now has to close when the books close.
The common thread is that none of these manufacturers is publicly differentiated by its models. What each publishes is a property of its evidence chain: a consolidated statement across a heterogeneous estate, an environmental unit of account matched to the commercial one, or a sustainability statement inside the audited annual report. Those are the conditions under which a model becomes useful and safe. A manufacturer at stage 2 that deploys the same model gets a faster route to a number nobody can trace.
The four dimensions that set your stage
Readiness is not one number. Four dimensions gate each other, and the lowest one caps the grade of everything you publish.
Readiness is scored on four dimensions — evidence quality, allocation and traceability, assurance and control, and decision use — and the lowest of them is the real stage, because each gates the others. Excellent metering with no allocation produces a dataset only the energy team can use. Perfect allocation over spend-based factors produces an auditable pipeline of estimates. Rigorous control over data nobody acts on produces an expensive annual artefact. The dimensions are not a scorecard for its own sake; they are the four ways a chain can fail while looking healthy from every other angle.
Evidence quality
What grade your inputs actually are, and whether the grade is recorded rather than remembered. The binding question is what share of a material figure is grade A or B, and whether anyone could tell from the disclosure. This dimension is usually strong on site energy and weak everywhere else, which mirrors the 26-to-1 ratio CDP reports (opens in a new tab) between supply-chain and operational emissions.
Allocation and traceability
Whether a figure can be cut to the resolution an obligation demands, and whether it can be traced back to raw. This is the dimension that separates stage 2 from stage 3, and it is overwhelmingly the lowest-scoring one in manufacturing — not for technical reasons but because the join between meters and production orders sits between two functions and belongs to neither.
Assurance and control
Whether the methodology and factor set are versioned and frozen per period, whether restatement is a policy rather than an accident, and whether any model in the path is registered, reviewed and labelled. Manufacturers consistently overestimate this dimension because the controls exist as documents; the test is whether the restatement policy has ever been used.
Decision use
Whether an environmental figure changes an operating decision, and whether the change can be attributed. This is the dimension that determines funding. A dataset used only for disclosure is defended annually against its cost; a dataset that schedules a batch, sets a setpoint or wins a supplier award has an operational owner whose targets improve when it is right.
Diagnosing the real constraint
Plot your primary-data share against your allocation and traceability. The quadrant names the next investment — and three of the four common answers are not 'get better data'.
Metered but unallocated
- Good primary data trapped inside the energy function
- The most common position, and the highest-leverage one
- Fix: join one meter to one line's production orders — not more meters
Assurance-ready
- Grade and resolution both in place
- Constraint has moved to decision use and control
- Fix: put one assured figure into a live decision with a holdout
Spreadsheet ESG
- Neither grade nor resolution
- Normal at stage 1, and cheap to leave
- Fix: meter one utility at one site and freeze one written methodology
Precisely wrong
- An auditable pipeline built on modelled inputs
- The most dangerous quadrant, because it feels assured
- Fix: raise the grade of the largest contributors before raising resolution further
The bottom-right quadrant deserves the warning it gets. A manufacturer that has invested in pipelines, lineage and controls over a Scope 3 inventory built from spend-based factors has produced something that looks exactly like readiness from a governance review and is, in evidence terms, an annual report of estimates with excellent version history. It is also the position most likely to be reached by buying software first, because a platform can supply structure and cannot supply grade. The test is simple and unforgiving: pick your three largest emission lines and ask what physical measurement or supplier document sits at the bottom of each.