Manufacturing (Automotive)Regulations, Compliance & Governance
AI compliance for carbon reduction targets in automotive manufacturing: making a modelled emissions number audit-grade
AI compliance for carbon reduction targets is the discipline of keeping machine-generated emissions figures fit for regulated use. In automotive manufacturing, models that gap-fill supplier product carbon footprints, classify purchase lines or forecast plant energy now feed assured CSRD statements, CBAM declarations and customer PCF requests, so every figure needs a provenance class, a method version and a reproducible trail.

Key takeaways
- No regulation forbids estimating emissions with a model. The GHG Protocol's Scope 3 standard expects secondary data, and ESRS E1 asks you to disclose how much of your Scope 3 rests on primary data rather than to eliminate estimates. What fails assurance is an estimate that cannot be told apart from a measurement.
- Treat a gap-fill model as a measurement instrument, not as analytics. Automotive already qualifies gauges through measurement systems analysis — bias, repeatability, stated uncertainty, a calibration record. A carbon estimator that arrives without those artefacts is an uncalibrated gauge feeding a legally required disclosure.
- One figure serves six consumers with completely different tolerances. An ESRS E1 statement accepts a labelled estimate; a CBAM declaration on imported steel or aluminium runs on defined rules for actual versus default values; a type-approval CO2 figure comes from a regulated test procedure and no model may substitute for it.
- The binding constraint is the ledger, not the model. A carbon ledger where every row carries value, unit, provenance class, method version, factor-set version, uncertainty and owner turns a five-week assurance scramble into an export — and it is the same artefact that answers a customer PCF request and a CBAM verifier.
- A reduction claim without a counterfactual is not a reduction. AI-driven paint-shop or compressor savings are physically real, but unless a baseline and a comparable untreated period or line exist, the saving cannot be separated from weather, mix and volume — and it will not survive the first challenge from an assurance provider.
Abbreviations used on this page
- PCF
- Product carbon footprint — cradle-to-gate CO2e for one part or material
- CSRD
- Corporate Sustainability Reporting Directive (EU)
- ESRS E1
- European Sustainability Reporting Standard E1, the climate-change standard
- CBAM
- Carbon Border Adjustment Mechanism (EU)
- GHG Protocol
- Greenhouse Gas Protocol — the Scope 1/2/3 accounting standards
- LCA
- Life-cycle assessment
- SBTi
- Science Based Targets initiative
- EF
- Emission factor — kg CO2e per unit of activity
- BoM
- Bill of materials
- MES
- Manufacturing execution system
- EMS
- Energy management system (ISO 50001) or environmental management system (ISO 14001)
- MSA
- Measurement systems analysis — the automotive discipline for qualifying a gauge
Free · 8 questions · ~3 minutes
Score your carbon numbers, not your ambition
Eight questions, one at a time, about three minutes. They ask where your emissions figures actually come from, how tightly the method is controlled, what evidence exists behind them, and whether they change any decision. Answer them and we build your personalised report — your stage on the provenance ladder, your score on each of the four dimensions, and the specific gap standing between you and an assurable number — and send it to your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised provenance report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the gaps most likely to surface in a limited-assurance sample, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Spend-proxy
Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
Your next movePick one commodity family — aluminium castings, cold-rolled steel, wiring harnesses — and rebuild its figure on mass and process rather than spend.
Stage 2 · Modelled
Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
Your next moveAdd provenance class, method version, factor-set version and uncertainty to every row of the carbon ledger — before improving a single model.
Stage 3 · Attributed
Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
Your next moveRun the three control tests: back-test the estimator against arriving primary data, rehearse a figure reproduction against the clock, and force a factor-set update through end to end.
Stage 4 · Verified
An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
Your next moveWire the verified figure into the decisions that create it: supplier award scoring, make-or-buy, material substitution and plant energy dispatch.
Stage 5 · Steered
The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.
Your next moveTreat the decision policy as a versioned, reviewable artefact with the same rigour as the estimator, and monitor the escalation rate as its expiry signal.
0 / 24
Activity data provenance
— / 6
Method and factor control
— / 6
Assurance evidence
— / 6
Target and decision linkage
— / 6
Your score maps to a stage on the provenance ladder. The dimension breakdown matters more than the total: the lowest dimension is what an assurance provider will find first, and it is where the next investment belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the provenance ladder. The dimension breakdown matters more than the total: the lowest dimension is what an assurance provider will find first, and it is where the next investment belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this read against your actual reporting calendar?
We walk your sustainability, purchasing and manufacturing IT leads through the dimension scores, map them onto your next CSRD cycle, CBAM declaration window and largest customer PCF request, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–5 — Stage 1, Spend-proxy. Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
- 6–11 — Stage 2, Modelled. Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
- 12–16 — Stage 3, Attributed. Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
- 17–21 — Stage 4, Verified. An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
- 22–24 — Stage 5, Steered. The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.
What AI compliance for carbon reduction targets means in automotive
A definition, the three very different things people mean by 'AI and carbon', and the path one figure travels from a meter or a purchase line to a regulated disclosure.
AI compliance for carbon reduction targets is the discipline of keeping machine-generated emissions figures fit for regulated use. In an automotive manufacturer that means three specific controls: every figure knows how it was produced, every method and factor version is recorded against the figures it changed, and every reported reduction can be separated from volume, mix and price movement. The models are the easy part. The controls are the work.
The phrase covers three activities that are constantly conflated and carry completely different risk. The first is AI used to reduce emissions — paint-shop oven control, compressed-air leak detection, press-line load shifting, logistics consolidation. These change physical reality and their compliance exposure is limited to how the saving is evidenced. The second is AI used inside the carbon account — gap-filling supplier PCFs, mapping purchase lines to activity data, extracting figures from supplier declarations, forecasting the trajectory against a target. This is where almost all the regulatory exposure sits, because its output is a number in a legally required disclosure. The third is the footprint of the AI estate itself, which is now large enough that the IEA tracks it as a distinct load (opens in a new tab) and large enough to be asked about in a Scope 2 or Scope 3 breakdown.
Nothing in the regulatory frame forbids the second activity. The GHG Protocol's Corporate Value Chain (Scope 3) Standard (opens in a new tab) explicitly anticipates secondary data and estimation across all fifteen categories; ESRS E1 under the CSRD (opens in a new tab) asks a reporter to disclose the extent to which Scope 3 rests on primary data obtained from the value chain, not to eliminate estimates. The failure mode is narrower and more mundane: an estimate that cannot be told apart from a measurement, or one that cannot be reproduced when it is sampled. That is a control finding, and control findings are what turn a reporting exercise into a restatement.
What a verifier can reproduce, against maturity
The curve is not linear. Through stages 1 and 2 — where most manufacturers sit — the reported total may be perfectly reasonable while the reproducible share stays close to flat, because provenance was never captured. The inflection happens at stage 3, when the ledger starts recording how each figure was produced rather than only what it was.
Share of reported figures reproducible from the trail by stage
- Stage 1 · Spend-proxy — 21% of operators. Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
- Stage 2 · Modelled — 38% of operators. Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
- Stage 3 · Attributed — 26% of operators. Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
- Stage 4 · Verified — 12% of operators. An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
- Stage 5 · Steered — 3% of operators. The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.
Curve shape: logistic, plotted from the stage data above. Distribution: Framed against the evidence expectations in IAASB's ISSA 5000.
How one carbon figure reaches a regulated disclosure
Three provenance classes feed one ledger, and the ledger feeds every consumer. The stage is determined by what travels with the number: at stages 1-2 a modelled value can leave as a bare figure and the restatement path opens; from stage 3 the provenance class, method version and uncertainty travel with it, which is what makes the verifier's reconstruction possible at all.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- The primary lane is the only one nobody argues about. Plant meters and the energy management system produce kWh, gas and compressed-air consumption by shop; reconciled against utility invoices, this becomes Scope 1 and Scope 2 activity data with a physical audit trail behind every figure.
- The exchanged lane is the one that grows. A part-level PCF request goes out under a common data model, and the supplier returns a declaration that states its method, its boundary and its version. When that arrives, it supersedes whatever the model had estimated — and the supersession itself is logged, because the delta is the interesting part.
- The modelled lane is where the volume is and where the exposure is. ERP purchase lines and the bill of materials feed a gap-fill estimator that produces a figure for every part no supplier has declared. Written into the ledger with a provenance flag, a model version and an uncertainty band, this is an ordinary, defensible practice. Written out as a bare number, it becomes an unlabelled estimate, and when a customer or an assurance provider challenges it, the only available answer is a restatement.
- Everything converges on one ledger and diverges to consumers with different tolerances — an ESRS E1 statement under limited assurance, a CBAM declaration on imported steel and aluminium, a contractual PCF answer to a customer. The verifier's test is the same in all three: reconstruct the value from the method, the factor set and the source record.
Step-by-step insights
- Meters and the invoice reconciliation — the cheapest credibility on the page
- Scope 1 and Scope 2 are the easiest figures to get right and the ones most often left half-done. A plant with sub-metering by shop can attribute consumption to the paint shop, the body shop and compressed-air generation; a plant with a single incoming supply can only attribute to the site. The reconciliation against utility invoices is what converts telemetry into evidence, because the invoice is a third-party record with a counterparty. Manufacturers that skip the reconciliation end up defending their own meters, which is a much harder conversation than pointing at a bill.
- The PCF request — a data problem that is really a contract problem
- Suppliers do not withhold part-level footprints out of obstruction; most tier-twos genuinely cannot produce one, and those that can want to know what happens to the number. Both objections are commercial, not technical. The programmes that move fastest write the PCF obligation into the award, provide a machine-readable path so the answer is not a bespoke spreadsheet each time, and state plainly how the figure will be used. The sector has already solved the adjacent problem — exchanging sensitive supplier data across tiers under a shared trust and assessment framework — so the request should ride on that infrastructure rather than on a new portal nobody outside your top twenty suppliers will ever log into.
- Supersession beats overwriting
- When a primary declaration arrives for a part that was previously modelled, the instinct is to overwrite the estimate. Do not: write a new row that supersedes the old one, keep both, and record the delta. Three things then become possible that are otherwise impossible. Previously published figures remain reconstructable, which is the whole point of the trail. The differences between arriving primary data and prior estimates accumulate into a free, continuously refreshed validation set for the estimator. And the restatement conversation, when it comes, has arithmetic behind it rather than an apology.
- The gap-fill estimator is an instrument, not a report
- A model that produces a footprint for a part nobody has declared is performing a measurement by inference, and automotive already has a vocabulary for that. Measurement systems analysis asks whether a gauge is biased, how much it varies on repeat measurement, and what tolerance the result should be quoted with. Ask the same three questions of the estimator, refresh the answers against arriving primary PCFs each cycle, and publish them as a model card. It costs a few weeks, it is the single most effective thing you can do to make modelled figures survivable, and it uses a discipline the plant already runs on the shop floor.
- One ledger, many consumers — and never a second copy
- The strongest structural rule on this diagram is that the disclosure, the CBAM declaration, the customer answer and the internal decision all read the same rows. The moment a second copy exists — a reporting extract that gets hand-adjusted, a procurement sheet with its own factors — the two diverge, and the divergence is discovered by an outsider rather than by you. Serve every consumer by exporting a view of the ledger, and make the export mechanical enough that nobody is tempted to fix a number on the way out.
- Why the restatement path is drawn as a consequence of the customer answer
- The disclosure is sampled once a year by a provider who works to a defined standard. A customer PCF answer is challenged unpredictably, by procurement teams comparing your figure against a competitor's or against their own model, sometimes years later and usually with commercial pressure attached. In practice the first serious challenge to a bare, unlabelled estimate tends to arrive down that channel, not from the auditor — which is why the modelled lane deserves the same discipline as the disclosure even though the disclosure gets all the governance attention.
The five stages of the provenance ladder
For each stage: what the ledger actually looks like, the signals a reviewer can check in an afternoon, the anti-pattern that traps manufacturers there, and what leaving costs in team terms.
The ladder below measures one thing: how much a reported carbon figure carries with it. It runs from Spend-proxy, where the number has no physical unit in it at all, through Modelled, Attributed and Verified, to Steered, where the same verified figure that leaves the building in a disclosure also decides which supplier wins an award. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Spend-proxy
21% of operators sit here
Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
Stage 1 is not the absence of a carbon number — most automotive manufacturers have had one for years. It is the absence of any resolution beneath it. A spend-based Scope 3 line is computed by taking what you paid a supplier and multiplying it by a sector-average factor, which means the figure moves when procurement negotiates a discount and does not move when the supplier switches to renewable electricity. It is directionally useful for a first inventory and actively misleading as a target-tracking instrument.
The tell is what happens when someone asks a follow-up question. Ask which of your aluminium castings suppliers has the highest footprint per kilogram and the honest answer at stage 1 is that the question cannot be answered from the account, because the account never had a kilogram in it. Ask what your reduction programme achieved last year and the answer is a number that also reflects volume, mix, currency and commodity prices. Every finance director eventually asks a version of this question, and stage 1 has no answer that survives it.
AI at this stage is nearly always doing one job: classifying purchase lines into spend categories so the average factors can be applied. That is genuine, useful work and it is also the highest-risk work on the page, because it feels like accounting automation while it is in fact making thousands of unreviewed judgements that determine which factor gets applied to which euro. When the classifier is wrong it is wrong systematically, and the error lands in a legally required disclosure with no marker on it.
In practice
The discount that cut emissions
A tier-one supplier's sustainability lead presented a 4% year-on-year reduction in purchased-goods emissions to the board. Procurement had renegotiated steel contracts during a soft market; tonnage was flat. The account, built on spend times an average factor, recorded the price fall as a decarbonisation. Nobody had lied and nobody had checked, because the account had no tonnes in it to check against.
What it looks like
- Scope 3 category 1 is calculated by multiplying spend by an average factor
- No supplier has been asked for a part-level PCF
- The carbon figure is rebuilt in a spreadsheet each reporting cycle
- Nobody can say which parts, plants or suppliers drive the total
Diagnostic signals you can check this week
- Ask for the footprint of one part number. If the answer requires new analysis, you are here
- Check whether the Scope 3 category 1 calculation takes any input other than currency
- Ask how many supplier-declared PCFs are held in a system rather than in an inbox
- Ask what would change in the reported number if a supplier switched to renewable power. If the answer is nothing, the account cannot see decarbonisation
Anti-pattern · Buying an ESG platform before agreeing a boundary
The instinctive fix is a carbon-accounting platform, selected on connector count and dashboard quality. It reliably produces a faster version of the same spend-proxy number, because the constraint was never calculation speed — it was that no boundary, allocation rule or provenance convention had been agreed. Write the boundary down first, on one commodity, and the platform requirement stops being a guess.
What holds you here
The account has no physical unit in it, so no part, plant or supplier decision can be attributed and no reduction can be distinguished from a price movement.
Highest-leverage next move
Pick one commodity family — aluminium castings, cold-rolled steel, wiring harnesses — and rebuild its figure on mass and process rather than spend.
Cost of leaving
- Effort
- 2-4 months
- Team
- One data engineer, one sustainability analyst, part-time procurement support
- Risk
- Low — the work is additive and no disclosure depends on it yet
- To next stage
- 2-4 months
If this is you, the next step is
A 2-week exercise: pick a commodity family, classify every line by data provenance, size the gap.
Stage 2
Modelled
38% of operators sit here
Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
Stage 2 is the most dangerous stage on this ladder, precisely because it looks like a large step forward — and technically it is. The account now speaks in kilograms and process routes. A model trained on the PCFs suppliers have declared can produce a credible estimate for the several thousand parts where none exists, and the resulting total is far more useful for engineering decisions than anything stage 1 produced.
What is missing is the marker. In almost every stage-2 estate the modelled value and the supplier-declared value are written into the same field, and the only way to tell them apart is to ask the analyst who built the extract. That is fine until the figure leaves the building. Under limited assurance, the first thing a provider does is sample figures and trace them to source; a sampled figure that resolves to a model with no version, no back-test and no uncertainty band is not a finding about the model, it is a finding about the control environment.
The second thing missing is version discipline. Emission-factor sets are updated, global warming potential sets change between assessment reports, allocation rules get refined, and the model itself is retrained. At stage 2 none of these events is recorded against the figures they changed, so when this year's number differs from last year's, nobody can decompose the difference into real-world change and method change. That decomposition is exactly what a restatement conversation turns on.
In practice
The estimate that became a customer commitment
A supplier answered an OEM's part-level PCF request with a figure produced by its gap-fill model, sent as a plain number in a spreadsheet cell. Eighteen months later the OEM's own assurance provider sampled that part. The supplier could not reproduce the figure: the model had been retrained twice, the factor set had been updated once, and no version had been recorded against the response. The number was not wrong. It was simply unreconstructable, which for a contractual commitment is the same problem.
What it looks like
- Part-level footprints exist, built from BoM mass and process routes
- A model gap-fills the parts no supplier has declared
- Estimated and declared figures sit in the same column with no marker
- Method and factor-set versions are not recorded against individual figures
Diagnostic signals you can check this week
- Open the carbon dataset and look for a column that says how each figure was produced. Usually there is not one
- Pick five reported figures at random and ask which model version and factor set produced them
- Check whether a retrained gap-fill model triggers any recalculation of previously reported figures
- Ask whether any modelled figure carries an uncertainty band. At stage 2 it almost never does
Anti-pattern · Improving the model instead of labelling the output
When a modelled figure is challenged, the reflex is to make the model better — more features, more training PCFs, a tighter error metric. Accuracy is not what failed. A moderately accurate estimate labelled as an estimate, with a version and an uncertainty band, passes review; an excellent estimate presented as a fact does not. Spend the next quarter on the ledger schema and provenance flags, then revisit accuracy when you can measure what an accuracy point is worth in disclosed tonnes.
What holds you here
Estimates and measurements are indistinguishable in the ledger, so no sampled figure can be reproduced and no restatement can be decomposed.
Highest-leverage next move
Add provenance class, method version, factor-set version and uncertainty to every row of the carbon ledger — before improving a single model.
Cost of leaving
- Effort
- 4-8 months
- Team
- One data engineer, one LCA practitioner, a named ledger owner
- Risk
- Medium — retrofitting provenance onto already-published figures forces a restatement conversation
- To next stage
- 4-8 months
If this is you, the next step is
The stage 2 to 3 move is the most common engagement on this ladder. Typically 90 days.
Stage 3
Attributed
26% of operators sit here
Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
Stage 3 is the first stage where the carbon number survives contact with someone who did not build it. The change is not analytical, it is structural: the ledger stops holding numbers and starts holding statements. A row no longer says 4.81 kg CO2e; it says 4.81 kg CO2e, modelled, estimator v7, factor set 2026.1, cradle-to-gate, plus or minus 22% at the 80th percentile, owned by the commodity lead, superseding the value published on the previous cycle.
The discipline that gets you here is closer to quality engineering than to data science, which is convenient, because automotive manufacturers already have it. A part characteristic that matters is not simply measured — the gauge is qualified through measurement systems analysis, the study is recorded, the calibration has a date, and the result is quoted with a tolerance. Apply that vocabulary to a gap-fill estimator and every question an assurance provider asks has an existing answer shape. The organisations that struggle at this stage are usually the ones treating the carbon model as reporting software rather than as instrumentation.
The constraint that emerges is supplier data. Once provenance is visible, the primary-data share stops being a rhetorical claim and becomes a number on a slide, and it is almost always lower than anyone expected. This is the point at which the programme's centre of gravity moves out of the sustainability team and into purchasing, because the only way to raise it is to make a part-level PCF a condition of the award and to give suppliers a machine-readable path to send one.
In practice
The supersession that took four minutes
A castings supplier returned a primary PCF for a structural part eleven months after the modelled figure had been published. In a stage-3 estate the arriving declaration was validated against the declared method, written as a new ledger row superseding the old one, and flagged to the reporting owner with the delta and its effect on the commodity total. The whole event took four minutes of human attention. In the same estate two years earlier it would have been an email nobody actioned until the next reporting cycle, by which time the published figure had already been assured.
What it looks like
- Each ledger row records value, unit, provenance class, method version and factor-set version
- Modelled rows carry an uncertainty band derived from a back-test
- An arriving supplier PCF supersedes the modelled figure and the change is logged
- The primary-data share is a tracked, reportable metric rather than an estimate
Diagnostic signals you can check this week
- Pick any reported figure and ask the system — not a person — how it was produced
- Check whether an arriving supplier PCF creates a superseding row or overwrites the old value
- Ask for the primary-data share by commodity. If the number exists and is uncomfortable, you are here
- Look for a back-test that compares last year's modelled estimates against primary data that arrived afterwards
Anti-pattern · Declaring the ledger done and moving to dashboards
Provenance discipline is unglamorous, and the temptation once it exists is to spend the next two quarters on visualisation for the executive committee. The gap that opens is control testing: nobody has yet proven that the model's estimates are unbiased, that the reproduction path actually works under time pressure, or that a factor-set update propagates. Those three tests are what convert a well-structured ledger into an assurable one, and they cost far less than the dashboard.
What holds you here
Provenance is visible but unproven — no back-test, no reproduction rehearsal and no tested propagation of a factor-set update.
Highest-leverage next move
Run the three control tests: back-test the estimator against arriving primary data, rehearse a figure reproduction against the clock, and force a factor-set update through end to end.
Cost of leaving
- Effort
- 6-12 months
- Team
- Ledger owner, LCA practitioner, ML engineer, a purchasing counterpart
- Risk
- Medium — raising the primary-data share depends on supplier leverage you may not have on every commodity
- To next stage
- 6-12 months
If this is you, the next step is
Bias study, back-test against arriving primary data, stated uncertainty, model card. Four weeks.
Stage 4
Verified
12% of operators sit here
An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
Stage 4 is defined from the outside in. The question is no longer whether your team understands how the number was produced but whether a reasonably diligent stranger can rebuild it from artefacts you already keep. That is the practical content of a limited-assurance engagement and it is the same test a CBAM verifier, an OEM's supplier-audit team and a customer's procurement quality function each apply in their own vocabulary.
The work at this stage is control design, and it is mostly borrowed. Change control for the estimator is the change control you already run on MES software. The bias and spread study is a gauge study with a different unit. The reproduction rehearsal is an audit rehearsal. The evidence pack is a control plan. Manufacturers that already hold IATF 16949 certification and run ISO 14001 environmental management have every one of these muscles; what is usually missing is the decision to apply them to a system the sustainability team owns and manufacturing quality has never looked at.
The characteristic stage-4 discovery is that the models were never the weak point. Reproduction rehearsals overwhelmingly fail on inputs: a factor set someone downloaded and stored locally, a supplier declaration that exists only as a PDF attachment, an allocation decision made in a meeting and never written down. Fixing those is dull, cheap and the difference between an assurance engagement that costs five weeks of the team's year and one that costs five days.
In practice
The five-figure rehearsal
Before its first assured cycle, a manufacturer picked five published figures at random and gave a two-person team the verifier's standard questions with the original analysts excluded from the room. Three reconstructed inside the hour. One took a day because the emission factor had come from a spreadsheet on a shared drive with no version. One could not be reconstructed at all: the supplier declaration behind it had been received verbally and confirmed in an email thread. The two failures cost about three weeks to fix, and they were found in a rehearsal instead of in an engagement.
What it looks like
- Model changes go through change control with a documented impact on reported figures
- A sampled figure is reconstructed from source in under an hour, without the analyst who built it
- The gap-fill estimator has a current bias and spread study against arriving primary data
- Recalculation and restatement policies exist and have been exercised at least once
Diagnostic signals you can check this week
- Run a blind reproduction on three sampled figures with the original analyst absent, and time it
- Ask when the estimator's bias study was last refreshed and against how many arriving primary PCFs
- Check whether a factor-set update has ever been pushed through to previously reported figures
- Ask to see the evidence pack from the last cycle. If it was assembled for the audit rather than generated by the system, the control is manual
Anti-pattern · Automating the reduction claim before the baseline exists
Stage 4 makes the account credible, which is exactly when someone proposes wiring optimisation savings straight into the reported reduction. Energy models estimate a counterfactual, and a counterfactual is not a measurement. Without a documented baseline period and a comparable untreated line, shift or building, the claimed saving cannot be separated from volume, mix and weather. The first challenge to an unsupported reduction claim tends to discredit the whole account, including the parts that were sound.
What holds you here
The account is defensible but passive — it reports the past accurately and does not yet change a sourcing, engineering or dispatch decision.
Highest-leverage next move
Wire the verified figure into the decisions that create it: supplier award scoring, make-or-buy, material substitution and plant energy dispatch.
Cost of leaving
- Effort
- 12-18 months
- Team
- Ledger owner, model owner, internal audit or quality partner, external assurance liaison
- Risk
- Higher — control gaps found late become restatements, and restatements are visible to the market
- To next stage
- 12-18 months
If this is you, the next step is
We sample five published figures and time the reconstruction against a verifier's question set.
Stage 5
Steered
3% of operators sit here
The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.
Stage 5 is narrower than the marketing language around it suggests. It is not an autonomous carbon-management system. It is the specific arrangement in which one ledger serves both the outward-facing disclosure and the inward-facing decision, so that the number a purchasing manager sees when scoring a quotation is the number that will appear, months later, in an assured statement. That single-source property is the whole point: the moment the decision runs on a different figure from the disclosure, both become negotiable.
The hard part here is governance rather than engineering. If carbon intensity carries weight in an award decision, the weighting is a commercial policy with legal consequences and it has to be versioned, reviewed and defensible to a supplier who lost on it. If an engineering change is approved partly on a footprint delta, the estimator behind that delta has just become part of the product development process and inherits its documentation expectations. The organisations that do this well tend to route the policy through the same forum that handles customer-specific requirements, because it is the forum that already knows how to version a commitment.
Sustaining stage 5 is the part most often underestimated. Factor sets move, suppliers change plants and power contracts, regulations amend their scope and reporting dates, and a threshold set against 2026 conditions quietly stops being valid. The leading indicator worth watching is the rate at which decisions fall outside the policy's stated bounds and route to a human: when it rises, the world has moved and the policy needs a review before an incident forces one.
In practice
The award that turned on a declared figure
A manufacturer scoring two aluminium castings suppliers found the quotations within a per-part cost of each other and materially apart on declared cradle-to-gate footprint. The award went to the higher-cost supplier on a documented weighting, and the losing supplier asked for the basis. Because both figures were supplier-declared under a common method with recorded versions, the answer was a two-page extract from the ledger rather than a negotiation. The same rows fed the following year's disclosure unchanged.
What it looks like
- Carbon intensity is a scored criterion in supplier award decisions, with the data source named
- Engineering change proposals carry a footprint delta computed from the same ledger
- Plant energy dispatch and the reported Scope 2 figure resolve to one set of meters
- The decision policy — thresholds, weights, escalation — is versioned and reviewed like code
Diagnostic signals you can check this week
- Ask a purchasing manager where the carbon figure on their scoring sheet comes from. At stage 5 they name the ledger
- Check whether the reported Scope 2 figure and the plant's energy-dispatch model read the same meters
- Ask whether the award-scoring weights are versioned, dated and reviewed by a named forum
- Track the share of decisions escalating outside policy bounds. A rising trend means the thresholds have expired
Anti-pattern · Treating the decision policy as a spreadsheet setting
Weights and thresholds get tuned in a configuration sheet with no version history, no review record and no note of who changed what. It works until a supplier challenges an award, a regulator asks how a threshold was set, or an internal auditor asks which policy version applied on a given date. Version the policy, minute the reviews, keep the trail — the artefact that will be examined is the policy, not the model.
What holds you here
Sustaining the arrangement is a governance problem — thresholds expire, factor sets move and the policy must be reviewed on a cadence, not on an incident.
Highest-leverage next move
Treat the decision policy as a versioned, reviewable artefact with the same rigour as the estimator, and monitor the escalation rate as its expiry signal.
Cost of leaving
- Effort
- Continuous
- Team
- Ledger and model owners plus a standing forum spanning purchasing, engineering, quality and sustainability
- Risk
- Concentrated — low frequency, high consequence, commercial and regulatory in nature
If this is you, the next step is
We run a real scoring decision against the policy, the trail and a supplier challenge scenario.
Where automotive manufacturers actually sit on the ladder
The distribution across the five stages, why the stage 2 to 3 move is the largest single loss, and the regulatory dates that make it urgent rather than worthy.
Most automotive manufacturers are at stage 2 — part-level figures exist, a model fills the gaps, and the ledger cannot tell the two apart. The distribution below is weighted heavily toward that position: a large majority have moved off pure spend proxies for their main commodities, and a small minority have a figure that a stranger could reproduce from the trail. The drop between stage 2 and stage 3 is the largest single transition loss on the ladder, and it is not a modelling gap. It is a schema gap.
Distribution of automotive manufacturers across the provenance ladder
Illustrative distribution. Stage 2 is both the mode and the plateau: a modelled account that reads well internally and cannot survive being sampled. Percentages are a model-derived reading of the published evidence below, not a survey result.
Share of manufacturers
- 21% — 1 · Spend-proxy
- 38% — 2 · Modelled (the plateau)
- 26% — 3 · Attributed
- 12% — 4 · Verified
- 3% — 5 · Steered
Source: Illustrative, synthesised from CDP supply-chain reporting and the EU disclosure frame
The reason the plateau exists is that stage 2 is genuinely comfortable. Internally the account looks sophisticated: it speaks in kilograms, it responds to engineering questions, and the year-on-year movement is explicable to a management committee. Nothing exposes the missing provenance until an outsider samples a figure — and by then the number has been published. Manufacturers that have been through one assured cycle almost never describe the experience as an accuracy problem; they describe it as five weeks of people reconstructing things from inboxes.
The urgency is calendar-driven rather than philosophical. The CBAM definitive regime (opens in a new tab) began on 1 January 2026 for covered goods including iron, steel and aluminium, which is most of a vehicle's structural mass by weight. The EU CO2 emission performance standards for cars and vans (opens in a new tab) set a 55% reduction for new cars by 2030 and 100% from 2035 against the 2021 baseline, with excess-emissions premiums attached. The Batteries Regulation (opens in a new tab) adds a carbon footprint declaration and, from February 2027, a battery passport. And the CSRD's scope and timing were themselves amended by the 2025 simplification package, so the reporting perimeter your programme was designed against in 2024 is not necessarily the one that applies now — a reason to build provenance that is robust to scope changes rather than tuned to one deadline.
Sector context is worth reading alongside the regulation rather than instead of it. ACEA (opens in a new tab) publishes the European manufacturers' position and the fleet data behind it, the European Environment Agency (opens in a new tab) monitors registered new-vehicle CO2 across the EU, and the Science Based Targets initiative (opens in a new tab) sets the validation rules that decide whether a corporate target counts as science-based at all. None of these bodies asks you to stop estimating. All of them assume you can say how you estimated.
One carbon number, six consumers — and each accepts different evidence
The same figure lands in a sustainability statement, a customs declaration, a customer's contract file and a type-approval dossier. Only one of them will not tolerate a model at all.
A carbon figure produced once is consumed at least six times, and the six consumers have completely different tolerances for a modelled value. This is the single most useful thing to internalise about compliance in this domain, because it dissolves the unhelpful question — is AI allowed in carbon accounting? — into six answerable ones. An ESRS E1 statement expects estimates and asks you to quantify how many. A CBAM declaration runs on rules about actual versus default values at the installation level. A type-approval CO2 figure comes from a prescribed test procedure and admits no model at all.
| Consumer of the figure | What it is | Where a modelled value stands | Evidence it asks for | Who checks it |
|---|---|---|---|---|
| ESRS E1 sustainability statement | Gross Scope 1, 2 and 3, intensity, targets and transition plan under the CSRD | Expected — the standard asks you to disclose the extent to which Scope 3 rests on primary data from the value chain | Method, boundary, factor sources, the primary-data share and a traceable path from figure to source | Assurance provider, limited assurance |
| CBAM declaration on imported goods | Embedded emissions in imported iron, steel, aluminium and other covered goods | Constrained — the regulation governs when actual installation data must be used and when default values are permitted | Installation-level data where required, supported by verification | Accredited verifier and the national competent authority |
| Customer part-level PCF response | A contractual answer to an OEM or tier-one request, part number by part number | Permitted if declared, and increasingly required to state its data-quality basis | Declared method, boundary, calculation version and a data-quality rating | Customer procurement quality, then their assurance provider |
| EV battery carbon footprint declaration | A declaration for batteries placed on the EU market under the Batteries Regulation | Bounded by the methodology set in the delegated act, calculated per battery model and manufacturing plant | Plant- and model-specific calculation following the prescribed method | Conformity assessment route defined by the Regulation |
| Science-based target progress | Progress against a validated near-term or net-zero target | Permitted, but a method or model change must be handled as a recalculation, never presented as progress | A stable baseline, a written recalculation policy and consistent boundaries over time | The target-validation body, and your own board |
| Type-approval CO2 for the vehicle | The regulated CO2 value for a vehicle type, feeding fleet compliance | Not applicable — the value comes from a prescribed test procedure, not from an estimate | Type-approval test evidence under the applicable vehicle regulation | Type-approval authority and technical service |
The asymmetry in that table is the design constraint. Because one ledger serves all six, its schema has to satisfy the strictest consumer that will ever read a given row — and which consumer that is depends on the commodity, not on the reporting team's preference. A modelled footprint for a plastic clip will only ever inform an ESRS E1 total. A modelled footprint for an imported aluminium casting may end up adjacent to a customs declaration, where the CBAM rules on actual versus default values (opens in a new tab) apply. The practical consequence is that provenance discipline should be introduced commodity by commodity in order of regulatory exposure, starting with the CBAM-covered materials, rather than uniformly across a bill of materials with ten thousand lines.
Jan 2026
CBAM definitive regime begins for covered goods including iron, steel and aluminium
European Commission
Feb 2027
Battery passport required for batteries placed on the EU market
European Commission
2035
100% CO2 reduction required for new cars registered in the EU, against 2021
European Commission
Underneath the consumers sits the estate that produces the numbers. Automotive carbon data comes from six operating domains, each with its own system of record, its own AI use and its own regulated destination. The map below is how we scope provenance work with manufacturers: pick the domain where the regulated consumer is tightest and the system of record is already yours, and start there.
| Domain | Where AI is actually used | System of record | The figure it moves | Regulated consumer |
|---|---|---|---|---|
| Purchased materials and parts | Gap-filling missing PCFs, mapping purchase lines to activity data, extracting figures from supplier declarations | ERP purchasing, PLM and BoM, supplier PCF exchange | Scope 3 category 1 — purchased goods and services | ESRS E1, customer PCF requests, CBAM on imported goods |
| Plant energy — press, body, paint, utilities | Oven and booth set-point control, compressed-air leak detection, load shifting against tariff and grid carbon intensity | EMS, MES, building management, sub-metering | Scope 1 and Scope 2, energy per vehicle produced | ESRS E1, ISO 50001 energy review |
| Inbound and outbound logistics | Load consolidation, mode shift, empty-run reduction, milk-run redesign | TMS and carrier EDI | Scope 3 categories 4 and 9 — transport and distribution | ESRS E1, transport stage of a customer PCF |
| Battery and cell supply | Cell supplier data validation, plant-level energy attribution, chemistry and sourcing scenario modelling | Cell supplier declarations, cell plant EMS | Battery carbon footprint per kWh of capacity | Batteries Regulation declaration and battery passport |
| Product engineering | Footprint deltas on material substitution, gauge reduction and design alternatives | PLM, CAD and BoM, LCA tooling | Cradle-to-gate footprint per vehicle programme | ESRS E1 transition plan, ecodesign and product passport |
| Use phase and fleet | Fleet-mix and demand scenario modelling behind the transition plan | Sales, homologation and fleet systems | Scope 3 category 11 — use of sold products | Fleet CO2 standards and type approval, ESRS E1 |
Two rows in that map deserve a warning. Use-phase emissions dominate an OEM's total footprint by a wide margin, which makes category 11 tempting to model aggressively — but the underlying per-vehicle CO2 figure is a type-approval output governed by the vehicle regulation regime (opens in a new tab) and its prescribed test procedures. Model the fleet mix and the lifetime mileage assumption, document both, and leave the regulated value alone. The second warning is the battery row: the Batteries Regulation (opens in a new tab) requires a calculation tied to a specific battery model made at a specific plant, so an estimate built from a cell chemistry average is structurally the wrong shape no matter how accurate it is. Precision does not fix a boundary error.
The materials rows are where the sector bodies help most. European Aluminium (opens in a new tab) and worldsteel (opens in a new tab) publish the methodological groundwork for primary and secondary metal footprints, which matters because the difference between primary and recycled aluminium is large enough to swamp most other levers in a body-in-white. If your gap-fill model cannot distinguish the two for a given part, the estimate is not merely uncertain — it is uncertain in a direction that procurement decisions actively change.
What the upper stages look like in public
Three publicly reported programmes, read against the provenance ladder. None is an Atomic Loops engagement — each links to the manufacturer's own published material.
The clearest public evidence for the ladder is in what large manufacturers chose to publish and what they chose to contractualise. In each case below, the distinguishing move was not a better model — it was that a carbon figure acquired a defined method, a named owner and a consequence. Read the outcomes as reported by the companies themselves; we have not independently audited them, and the linked pages are the manufacturers' own.
Three programmes read against the ladder
Outcomes as reported by the manufacturers in their own published material. Card images are generated industry scenes from our illustration library, not photographs of these companies' facilities, and imply no endorsement or involvement.
BMW GroupPremium OEM · Munich · multi-brand vehicle and motorcycle production35
- Challenge
- The large majority of a premium vehicle's cradle-to-gate footprint sits in materials and components the manufacturer does not meter — steel, aluminium and battery cells produced in other companies' plants, on other companies' grids.
- Approach
- BMW Group publishes life-cycle CO2e reduction targets expressed per vehicle, describes contracting for CO2-reduced steel and secondary aluminium, and states that CO2 performance is applied as a criterion in supplier selection. It is also a founding participant in the automotive data-space work built to exchange part-level product carbon footprints between tiers.
- Reported outcome
- As reported by the company, life-cycle CO2e per vehicle is tracked and published as a target metric, and carbon performance forms part of how suppliers are selected rather than a separate reporting exercise.
- What it shows about the curveThis is the stage 5 signature: the same figure that appears in the disclosure also carries weight in an award decision. That only works if the underlying data is declared under a common method, which is why the data-space work and the contractual requirement arrived together rather than in sequence.
Volkswagen GroupMulti-brand OEM group · Wolfsburg · plants across Europe, the Americas and Asia24
- Challenge
- Producing one comparable decarbonisation figure across many brands and plants, each with its own systems, grids and reporting history, without the number becoming a negotiation between business units.
- Approach
- Volkswagen Group publishes a group-level decarbonisation measure expressed per vehicle across the life cycle, sets CO2 requirements for suppliers within its purchasing process, and participates in the cross-industry work on part-level footprint exchange.
- Reported outcome
- As reported by the company, a per-vehicle life-cycle decarbonisation measure is published and tracked at group level alongside supplier CO2 requirements applied through purchasing.
- What it shows about the curveA group-level index is only meaningful if the boundary and allocation rules are frozen across brands and years. The hard work behind a headline figure like this is method stability, not calculation — which is exactly what the Attributed and Verified stages are made of.
Toyota Motor CorporationGlobal OEM · Toyota City · manufacturing across every major region24
- Challenge
- Reducing and evidencing plant CO2 across a very large global manufacturing footprint on grids whose carbon intensity varies by an order of magnitude between regions.
- Approach
- Toyota's published environmental programme sets long-range challenges covering both new-vehicle CO2 and plant CO2, and the company reports plant energy reduction and low-carbon power procurement by region within its environmental reporting.
- Reported outcome
- As reported by the company, plant CO2 reduction progress is published within a long-running environmental programme, with regional detail behind the group figure.
- What it shows about the curvePlant Scope 1 and 2 is the cheapest primary data a manufacturer will ever hold — it is metered, invoice-reconciled and inside your own fence. The contrast between how well-evidenced that lane is and how thin the upstream lane is, in the same account, is precisely what the primary-data share metric exposes.
What none of these programmes did was start with a model. In each, the sequence was the same: fix the boundary, obtain declared data for the commodities that matter, then apply computation to the remainder. That ordering is the whole argument of this page in three public examples — and it is the opposite of the order most internal programmes follow, which is to build the estimator first because it is the part a data team can deliver without anyone else's cooperation.
The four dimensions that set your stage
Provenance is not one number. Four dimensions gate each other, and the lowest one is the stage an assurance provider will find.
Your stage on the ladder is set by the lowest of four dimensions, not by the average. Activity data provenance, method and factor control, assurance evidence, and target and decision linkage each gate the others: an immaculately version-controlled method applied to spend-based inputs still produces a figure with no physical unit in it, and a rich primary-data estate with no reproduction path still fails a sample. The assessment on this page scores all four separately for exactly this reason.
Activity data provenance
Where each number physically comes from — a meter, a supplier declaration, or a model — and whether the ledger records which. The binding question is your primary-data share by commodity, and it only becomes a real number once provenance is captured. The exchange infrastructure now exists: Catena-X (opens in a new tab) defines a common automotive data model for passing part-level PCFs between tiers, the VDA (opens in a new tab) coordinates much of the German-language guidance around it, and the trust framework suppliers are already assessed against for data exchange is TISAX (opens in a new tab). None of that raises your primary-data share on its own; a clause in the award does.
Method and factor control
Whether the boundary, allocation rules, factor sets and global warming potential set behind a figure are recorded against that figure and versioned when they change. The GHG Protocol Product Standard (opens in a new tab) is the reference for the product-level rules, and ISO 14067 and ISO 14064-1 sit behind most sector guidance. The failure here is rarely choosing the wrong method; it is changing method silently, so that a year-on-year movement cannot be decomposed into real change and method change.
Assurance evidence
Whether a stranger can rebuild a sampled figure from artefacts you already keep — the practical content of a limited-assurance engagement under IAASB's ISSA 5000 (opens in a new tab). Automotive has an advantage here that most sectors do not: manufacturers already run certified management systems with documented control plans and audit rehearsals under IATF 16949 (opens in a new tab) and ISO 14001. The muscle exists. It has usually just never been pointed at the carbon ledger.
Target and decision linkage
Whether a reported reduction can be separated from volume, mix, price and weather, and whether the figure changes any decision inside the business. A claim with a documented baseline, a named lever and a comparable untreated line or period is evidence; a year-on-year difference is arithmetic. This is also the dimension that determines whether the ledger's controls ever get funded, because a figure that only leaves the building is a cost centre.
Diagnosing the real constraint
Plot your primary-data share against your provenance discipline. Three of the four quadrants have a next move that is not 'improve the model' — and the most dangerous position is the one that feels the most advanced.
Rich data, no trail
- Good inputs, nothing recording how each figure was produced
- The most dangerous quadrant — it reads as the most advanced
- Fix: retrofit the ledger schema before publishing another cycle
Assurance-ready
- Declared inputs and a reproducible trail
- Constraint moves to control testing and decision linkage
- Fix: rehearse a sample, then wire the figure into an award decision
Spend-proxy reporting
- Neither foundation in place
- Common at stage 1, and honest about it
- Fix: rebuild one commodity on mass and process, not spend
Honest but coarse
- Every figure labelled; most of them modelled
- Highest-leverage position on the matrix
- Fix: buy primary data with award clauses, not with a better estimator
The counter-intuitive result is that 'honest but coarse' is a stronger starting position than 'rich data, no trail'. A manufacturer whose figures are mostly modelled but fully labelled has a working control environment and a data-acquisition problem, which purchasing can solve with a clause. A manufacturer with excellent primary data and no provenance schema has a data asset and no control environment, and every cycle published in that state adds to the volume of figures that may later need restating.