Redefining Technology

Manufacturing (Automotive)Regulations, Compliance & Governance

AI compliance for carbon reduction targets in automotive manufacturing: making a modelled emissions number audit-grade

AI compliance for carbon reduction targets is the discipline of keeping machine-generated emissions figures fit for regulated use. In automotive manufacturing, models that gap-fill supplier product carbon footprints, classify purchase lines or forecast plant energy now feed assured CSRD statements, CBAM declarations and customer PCF requests, so every figure needs a provenance class, a method version and a reproducible trail.

Illustrative scene of an automotive plant energy and emissions control room, with carbon accounting overlays across body shop, paint shop and supplier data feeds
Manufacturing (Automotive) · Regulations, Compliance & Governance

Key takeaways

  1. No regulation forbids estimating emissions with a model. The GHG Protocol's Scope 3 standard expects secondary data, and ESRS E1 asks you to disclose how much of your Scope 3 rests on primary data rather than to eliminate estimates. What fails assurance is an estimate that cannot be told apart from a measurement.
  2. Treat a gap-fill model as a measurement instrument, not as analytics. Automotive already qualifies gauges through measurement systems analysis — bias, repeatability, stated uncertainty, a calibration record. A carbon estimator that arrives without those artefacts is an uncalibrated gauge feeding a legally required disclosure.
  3. One figure serves six consumers with completely different tolerances. An ESRS E1 statement accepts a labelled estimate; a CBAM declaration on imported steel or aluminium runs on defined rules for actual versus default values; a type-approval CO2 figure comes from a regulated test procedure and no model may substitute for it.
  4. The binding constraint is the ledger, not the model. A carbon ledger where every row carries value, unit, provenance class, method version, factor-set version, uncertainty and owner turns a five-week assurance scramble into an export — and it is the same artefact that answers a customer PCF request and a CBAM verifier.
  5. A reduction claim without a counterfactual is not a reduction. AI-driven paint-shop or compressor savings are physically real, but unless a baseline and a comparable untreated period or line exist, the saving cannot be separated from weather, mix and volume — and it will not survive the first challenge from an assurance provider.

Abbreviations used on this page

PCF
Product carbon footprint — cradle-to-gate CO2e for one part or material
CSRD
Corporate Sustainability Reporting Directive (EU)
ESRS E1
European Sustainability Reporting Standard E1, the climate-change standard
CBAM
Carbon Border Adjustment Mechanism (EU)
GHG Protocol
Greenhouse Gas Protocol — the Scope 1/2/3 accounting standards
LCA
Life-cycle assessment
SBTi
Science Based Targets initiative
EF
Emission factor — kg CO2e per unit of activity
BoM
Bill of materials
MES
Manufacturing execution system
EMS
Energy management system (ISO 50001) or environmental management system (ISO 14001)
MSA
Measurement systems analysis — the automotive discipline for qualifying a gauge

Free · 8 questions · ~3 minutes

Score your carbon numbers, not your ambition

Eight questions, one at a time, about three minutes. They ask where your emissions figures actually come from, how tightly the method is controlled, what evidence exists behind them, and whether they change any decision. Answer them and we build your personalised report — your stage on the provenance ladder, your score on each of the four dimensions, and the specific gap standing between you and an assurable number — and send it to your inbox.

0 of 8 answered

Question 1 of 8Activity data provenance

How is your largest Scope 3 category — purchased goods and services — actually calculated today?

Spend-based figures move with commodity prices and currency, so they cannot show a decarbonisation and cannot be attributed to a supplier decision.

How the score maps to a stage
  • 0–5 — Stage 1, Spend-proxy. Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
  • 6–11 — Stage 2, Modelled. Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
  • 12–16 — Stage 3, Attributed. Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
  • 17–21 — Stage 4, Verified. An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
  • 22–24 — Stage 5, Steered. The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.

What AI compliance for carbon reduction targets means in automotive

A definition, the three very different things people mean by 'AI and carbon', and the path one figure travels from a meter or a purchase line to a regulated disclosure.

AI compliance for carbon reduction targets is the discipline of keeping machine-generated emissions figures fit for regulated use. In an automotive manufacturer that means three specific controls: every figure knows how it was produced, every method and factor version is recorded against the figures it changed, and every reported reduction can be separated from volume, mix and price movement. The models are the easy part. The controls are the work.

The phrase covers three activities that are constantly conflated and carry completely different risk. The first is AI used to reduce emissions — paint-shop oven control, compressed-air leak detection, press-line load shifting, logistics consolidation. These change physical reality and their compliance exposure is limited to how the saving is evidenced. The second is AI used inside the carbon account — gap-filling supplier PCFs, mapping purchase lines to activity data, extracting figures from supplier declarations, forecasting the trajectory against a target. This is where almost all the regulatory exposure sits, because its output is a number in a legally required disclosure. The third is the footprint of the AI estate itself, which is now large enough that the IEA tracks it as a distinct load (opens in a new tab) and large enough to be asked about in a Scope 2 or Scope 3 breakdown.

Nothing in the regulatory frame forbids the second activity. The GHG Protocol's Corporate Value Chain (Scope 3) Standard (opens in a new tab) explicitly anticipates secondary data and estimation across all fifteen categories; ESRS E1 under the CSRD (opens in a new tab) asks a reporter to disclose the extent to which Scope 3 rests on primary data obtained from the value chain, not to eliminate estimates. The failure mode is narrower and more mundane: an estimate that cannot be told apart from a measurement, or one that cannot be reproduced when it is sampled. That is a control finding, and control findings are what turn a reporting exercise into a restatement.

What a verifier can reproduce, against maturity

The curve is not linear. Through stages 1 and 2 — where most manufacturers sit — the reported total may be perfectly reasonable while the reproducible share stays close to flat, because provenance was never captured. The inflection happens at stage 3, when the ledger starts recording how each figure was produced rather than only what it was.

Share of reported figures reproducible from the trail by stage

  • Stage 1 · Spend-proxy — 21% of operators. Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.
  • Stage 2 · Modelled — 38% of operators. Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.
  • Stage 3 · Attributed — 26% of operators. Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.
  • Stage 4 · Verified — 12% of operators. An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.
  • Stage 5 · Steered — 3% of operators. The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.

Curve shape: logistic, plotted from the stage data above. Distribution: Framed against the evidence expectations in IAASB's ISSA 5000.

How one carbon figure reaches a regulated disclosure

Three provenance classes feed one ledger, and the ledger feeds every consumer. The stage is determined by what travels with the number: at stages 1-2 a modelled value can leave as a bare figure and the restatement path opens; from stage 3 the provenance class, method version and uncertainty travel with it, which is what makes the verifier's reconstruction possible at all.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • The primary lane is the only one nobody argues about. Plant meters and the energy management system produce kWh, gas and compressed-air consumption by shop; reconciled against utility invoices, this becomes Scope 1 and Scope 2 activity data with a physical audit trail behind every figure.
  • The exchanged lane is the one that grows. A part-level PCF request goes out under a common data model, and the supplier returns a declaration that states its method, its boundary and its version. When that arrives, it supersedes whatever the model had estimated — and the supersession itself is logged, because the delta is the interesting part.
  • The modelled lane is where the volume is and where the exposure is. ERP purchase lines and the bill of materials feed a gap-fill estimator that produces a figure for every part no supplier has declared. Written into the ledger with a provenance flag, a model version and an uncertainty band, this is an ordinary, defensible practice. Written out as a bare number, it becomes an unlabelled estimate, and when a customer or an assurance provider challenges it, the only available answer is a restatement.
  • Everything converges on one ledger and diverges to consumers with different tolerances — an ESRS E1 statement under limited assurance, a CBAM declaration on imported steel and aluminium, a contractual PCF answer to a customer. The verifier's test is the same in all three: reconstruct the value from the method, the factor set and the source record.
Step-by-step insights
Meters and the invoice reconciliation — the cheapest credibility on the page
Scope 1 and Scope 2 are the easiest figures to get right and the ones most often left half-done. A plant with sub-metering by shop can attribute consumption to the paint shop, the body shop and compressed-air generation; a plant with a single incoming supply can only attribute to the site. The reconciliation against utility invoices is what converts telemetry into evidence, because the invoice is a third-party record with a counterparty. Manufacturers that skip the reconciliation end up defending their own meters, which is a much harder conversation than pointing at a bill.
The PCF request — a data problem that is really a contract problem
Suppliers do not withhold part-level footprints out of obstruction; most tier-twos genuinely cannot produce one, and those that can want to know what happens to the number. Both objections are commercial, not technical. The programmes that move fastest write the PCF obligation into the award, provide a machine-readable path so the answer is not a bespoke spreadsheet each time, and state plainly how the figure will be used. The sector has already solved the adjacent problem — exchanging sensitive supplier data across tiers under a shared trust and assessment framework — so the request should ride on that infrastructure rather than on a new portal nobody outside your top twenty suppliers will ever log into.
Supersession beats overwriting
When a primary declaration arrives for a part that was previously modelled, the instinct is to overwrite the estimate. Do not: write a new row that supersedes the old one, keep both, and record the delta. Three things then become possible that are otherwise impossible. Previously published figures remain reconstructable, which is the whole point of the trail. The differences between arriving primary data and prior estimates accumulate into a free, continuously refreshed validation set for the estimator. And the restatement conversation, when it comes, has arithmetic behind it rather than an apology.
The gap-fill estimator is an instrument, not a report
A model that produces a footprint for a part nobody has declared is performing a measurement by inference, and automotive already has a vocabulary for that. Measurement systems analysis asks whether a gauge is biased, how much it varies on repeat measurement, and what tolerance the result should be quoted with. Ask the same three questions of the estimator, refresh the answers against arriving primary PCFs each cycle, and publish them as a model card. It costs a few weeks, it is the single most effective thing you can do to make modelled figures survivable, and it uses a discipline the plant already runs on the shop floor.
One ledger, many consumers — and never a second copy
The strongest structural rule on this diagram is that the disclosure, the CBAM declaration, the customer answer and the internal decision all read the same rows. The moment a second copy exists — a reporting extract that gets hand-adjusted, a procurement sheet with its own factors — the two diverge, and the divergence is discovered by an outsider rather than by you. Serve every consumer by exporting a view of the ledger, and make the export mechanical enough that nobody is tempted to fix a number on the way out.
Why the restatement path is drawn as a consequence of the customer answer
The disclosure is sampled once a year by a provider who works to a defined standard. A customer PCF answer is challenged unpredictably, by procurement teams comparing your figure against a competitor's or against their own model, sometimes years later and usually with commercial pressure attached. In practice the first serious challenge to a bare, unlabelled estimate tends to arrive down that channel, not from the auditor — which is why the modelled lane deserves the same discipline as the disclosure even though the disclosure gets all the governance attention.

The five stages of the provenance ladder

For each stage: what the ledger actually looks like, the signals a reviewer can check in an afternoon, the anti-pattern that traps manufacturers there, and what leaving costs in team terms.

The ladder below measures one thing: how much a reported carbon figure carries with it. It runs from Spend-proxy, where the number has no physical unit in it at all, through Modelled, Attributed and Verified, to Steered, where the same verified figure that leaves the building in a disclosure also decides which supplier wins an award. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Spend-proxy

21% of operators sit here

Carbon figures are derived from purchase value and industry-average factors, so nothing can be attributed to a part, a plant or a decision.

Stage 1 is not the absence of a carbon number — most automotive manufacturers have had one for years. It is the absence of any resolution beneath it. A spend-based Scope 3 line is computed by taking what you paid a supplier and multiplying it by a sector-average factor, which means the figure moves when procurement negotiates a discount and does not move when the supplier switches to renewable electricity. It is directionally useful for a first inventory and actively misleading as a target-tracking instrument.

The tell is what happens when someone asks a follow-up question. Ask which of your aluminium castings suppliers has the highest footprint per kilogram and the honest answer at stage 1 is that the question cannot be answered from the account, because the account never had a kilogram in it. Ask what your reduction programme achieved last year and the answer is a number that also reflects volume, mix, currency and commodity prices. Every finance director eventually asks a version of this question, and stage 1 has no answer that survives it.

AI at this stage is nearly always doing one job: classifying purchase lines into spend categories so the average factors can be applied. That is genuine, useful work and it is also the highest-risk work on the page, because it feels like accounting automation while it is in fact making thousands of unreviewed judgements that determine which factor gets applied to which euro. When the classifier is wrong it is wrong systematically, and the error lands in a legally required disclosure with no marker on it.

In practice

The discount that cut emissions

A tier-one supplier's sustainability lead presented a 4% year-on-year reduction in purchased-goods emissions to the board. Procurement had renegotiated steel contracts during a soft market; tonnage was flat. The account, built on spend times an average factor, recorded the price fall as a decarbonisation. Nobody had lied and nobody had checked, because the account had no tonnes in it to check against.

What it looks like

  • Scope 3 category 1 is calculated by multiplying spend by an average factor
  • No supplier has been asked for a part-level PCF
  • The carbon figure is rebuilt in a spreadsheet each reporting cycle
  • Nobody can say which parts, plants or suppliers drive the total

Diagnostic signals you can check this week

  • Ask for the footprint of one part number. If the answer requires new analysis, you are here
  • Check whether the Scope 3 category 1 calculation takes any input other than currency
  • Ask how many supplier-declared PCFs are held in a system rather than in an inbox
  • Ask what would change in the reported number if a supplier switched to renewable power. If the answer is nothing, the account cannot see decarbonisation

Anti-pattern · Buying an ESG platform before agreeing a boundary

The instinctive fix is a carbon-accounting platform, selected on connector count and dashboard quality. It reliably produces a faster version of the same spend-proxy number, because the constraint was never calculation speed — it was that no boundary, allocation rule or provenance convention had been agreed. Write the boundary down first, on one commodity, and the platform requirement stops being a guess.

What holds you here

The account has no physical unit in it, so no part, plant or supplier decision can be attributed and no reduction can be distinguished from a price movement.

Highest-leverage next move

Pick one commodity family — aluminium castings, cold-rolled steel, wiring harnesses — and rebuild its figure on mass and process rather than spend.

Cost of leaving

Effort
2-4 months
Team
One data engineer, one sustainability analyst, part-time procurement support
Risk
Low — the work is additive and no disclosure depends on it yet
To next stage
2-4 months

If this is you, the next step is

A 2-week exercise: pick a commodity family, classify every line by data provenance, size the gap.

Run a provenance census on one commodity

Stage 2

Modelled

38% of operators sit here

Activity data and models now produce part-level estimates, but the ledger cannot distinguish an estimate from a measurement.

Stage 2 is the most dangerous stage on this ladder, precisely because it looks like a large step forward — and technically it is. The account now speaks in kilograms and process routes. A model trained on the PCFs suppliers have declared can produce a credible estimate for the several thousand parts where none exists, and the resulting total is far more useful for engineering decisions than anything stage 1 produced.

What is missing is the marker. In almost every stage-2 estate the modelled value and the supplier-declared value are written into the same field, and the only way to tell them apart is to ask the analyst who built the extract. That is fine until the figure leaves the building. Under limited assurance, the first thing a provider does is sample figures and trace them to source; a sampled figure that resolves to a model with no version, no back-test and no uncertainty band is not a finding about the model, it is a finding about the control environment.

The second thing missing is version discipline. Emission-factor sets are updated, global warming potential sets change between assessment reports, allocation rules get refined, and the model itself is retrained. At stage 2 none of these events is recorded against the figures they changed, so when this year's number differs from last year's, nobody can decompose the difference into real-world change and method change. That decomposition is exactly what a restatement conversation turns on.

In practice

The estimate that became a customer commitment

A supplier answered an OEM's part-level PCF request with a figure produced by its gap-fill model, sent as a plain number in a spreadsheet cell. Eighteen months later the OEM's own assurance provider sampled that part. The supplier could not reproduce the figure: the model had been retrained twice, the factor set had been updated once, and no version had been recorded against the response. The number was not wrong. It was simply unreconstructable, which for a contractual commitment is the same problem.

What it looks like

  • Part-level footprints exist, built from BoM mass and process routes
  • A model gap-fills the parts no supplier has declared
  • Estimated and declared figures sit in the same column with no marker
  • Method and factor-set versions are not recorded against individual figures

Diagnostic signals you can check this week

  • Open the carbon dataset and look for a column that says how each figure was produced. Usually there is not one
  • Pick five reported figures at random and ask which model version and factor set produced them
  • Check whether a retrained gap-fill model triggers any recalculation of previously reported figures
  • Ask whether any modelled figure carries an uncertainty band. At stage 2 it almost never does

Anti-pattern · Improving the model instead of labelling the output

When a modelled figure is challenged, the reflex is to make the model better — more features, more training PCFs, a tighter error metric. Accuracy is not what failed. A moderately accurate estimate labelled as an estimate, with a version and an uncertainty band, passes review; an excellent estimate presented as a fact does not. Spend the next quarter on the ledger schema and provenance flags, then revisit accuracy when you can measure what an accuracy point is worth in disclosed tonnes.

What holds you here

Estimates and measurements are indistinguishable in the ledger, so no sampled figure can be reproduced and no restatement can be decomposed.

Highest-leverage next move

Add provenance class, method version, factor-set version and uncertainty to every row of the carbon ledger — before improving a single model.

Cost of leaving

Effort
4-8 months
Team
One data engineer, one LCA practitioner, a named ledger owner
Risk
Medium — retrofitting provenance onto already-published figures forces a restatement conversation
To next stage
4-8 months

If this is you, the next step is

The stage 2 to 3 move is the most common engagement on this ladder. Typically 90 days.

Design the provenance schema

Stage 3

Attributed

26% of operators sit here

Every figure carries its provenance, method version and uncertainty, and a supplier's primary declaration automatically supersedes the model's estimate.

Stage 3 is the first stage where the carbon number survives contact with someone who did not build it. The change is not analytical, it is structural: the ledger stops holding numbers and starts holding statements. A row no longer says 4.81 kg CO2e; it says 4.81 kg CO2e, modelled, estimator v7, factor set 2026.1, cradle-to-gate, plus or minus 22% at the 80th percentile, owned by the commodity lead, superseding the value published on the previous cycle.

The discipline that gets you here is closer to quality engineering than to data science, which is convenient, because automotive manufacturers already have it. A part characteristic that matters is not simply measured — the gauge is qualified through measurement systems analysis, the study is recorded, the calibration has a date, and the result is quoted with a tolerance. Apply that vocabulary to a gap-fill estimator and every question an assurance provider asks has an existing answer shape. The organisations that struggle at this stage are usually the ones treating the carbon model as reporting software rather than as instrumentation.

The constraint that emerges is supplier data. Once provenance is visible, the primary-data share stops being a rhetorical claim and becomes a number on a slide, and it is almost always lower than anyone expected. This is the point at which the programme's centre of gravity moves out of the sustainability team and into purchasing, because the only way to raise it is to make a part-level PCF a condition of the award and to give suppliers a machine-readable path to send one.

In practice

The supersession that took four minutes

A castings supplier returned a primary PCF for a structural part eleven months after the modelled figure had been published. In a stage-3 estate the arriving declaration was validated against the declared method, written as a new ledger row superseding the old one, and flagged to the reporting owner with the delta and its effect on the commodity total. The whole event took four minutes of human attention. In the same estate two years earlier it would have been an email nobody actioned until the next reporting cycle, by which time the published figure had already been assured.

What it looks like

  • Each ledger row records value, unit, provenance class, method version and factor-set version
  • Modelled rows carry an uncertainty band derived from a back-test
  • An arriving supplier PCF supersedes the modelled figure and the change is logged
  • The primary-data share is a tracked, reportable metric rather than an estimate

Diagnostic signals you can check this week

  • Pick any reported figure and ask the system — not a person — how it was produced
  • Check whether an arriving supplier PCF creates a superseding row or overwrites the old value
  • Ask for the primary-data share by commodity. If the number exists and is uncomfortable, you are here
  • Look for a back-test that compares last year's modelled estimates against primary data that arrived afterwards

Anti-pattern · Declaring the ledger done and moving to dashboards

Provenance discipline is unglamorous, and the temptation once it exists is to spend the next two quarters on visualisation for the executive committee. The gap that opens is control testing: nobody has yet proven that the model's estimates are unbiased, that the reproduction path actually works under time pressure, or that a factor-set update propagates. Those three tests are what convert a well-structured ledger into an assurable one, and they cost far less than the dashboard.

What holds you here

Provenance is visible but unproven — no back-test, no reproduction rehearsal and no tested propagation of a factor-set update.

Highest-leverage next move

Run the three control tests: back-test the estimator against arriving primary data, rehearse a figure reproduction against the clock, and force a factor-set update through end to end.

Cost of leaving

Effort
6-12 months
Team
Ledger owner, LCA practitioner, ML engineer, a purchasing counterpart
Risk
Medium — raising the primary-data share depends on supplier leverage you may not have on every commodity
To next stage
6-12 months

If this is you, the next step is

Bias study, back-test against arriving primary data, stated uncertainty, model card. Four weeks.

Qualify your gap-fill model like a gauge

Stage 4

Verified

12% of operators sit here

An external verifier can reproduce a sampled figure from the trail alone, and the controls around models, factors and boundaries are tested rather than described.

Stage 4 is defined from the outside in. The question is no longer whether your team understands how the number was produced but whether a reasonably diligent stranger can rebuild it from artefacts you already keep. That is the practical content of a limited-assurance engagement and it is the same test a CBAM verifier, an OEM's supplier-audit team and a customer's procurement quality function each apply in their own vocabulary.

The work at this stage is control design, and it is mostly borrowed. Change control for the estimator is the change control you already run on MES software. The bias and spread study is a gauge study with a different unit. The reproduction rehearsal is an audit rehearsal. The evidence pack is a control plan. Manufacturers that already hold IATF 16949 certification and run ISO 14001 environmental management have every one of these muscles; what is usually missing is the decision to apply them to a system the sustainability team owns and manufacturing quality has never looked at.

The characteristic stage-4 discovery is that the models were never the weak point. Reproduction rehearsals overwhelmingly fail on inputs: a factor set someone downloaded and stored locally, a supplier declaration that exists only as a PDF attachment, an allocation decision made in a meeting and never written down. Fixing those is dull, cheap and the difference between an assurance engagement that costs five weeks of the team's year and one that costs five days.

In practice

The five-figure rehearsal

Before its first assured cycle, a manufacturer picked five published figures at random and gave a two-person team the verifier's standard questions with the original analysts excluded from the room. Three reconstructed inside the hour. One took a day because the emission factor had come from a spreadsheet on a shared drive with no version. One could not be reconstructed at all: the supplier declaration behind it had been received verbally and confirmed in an email thread. The two failures cost about three weeks to fix, and they were found in a rehearsal instead of in an engagement.

What it looks like

  • Model changes go through change control with a documented impact on reported figures
  • A sampled figure is reconstructed from source in under an hour, without the analyst who built it
  • The gap-fill estimator has a current bias and spread study against arriving primary data
  • Recalculation and restatement policies exist and have been exercised at least once

Diagnostic signals you can check this week

  • Run a blind reproduction on three sampled figures with the original analyst absent, and time it
  • Ask when the estimator's bias study was last refreshed and against how many arriving primary PCFs
  • Check whether a factor-set update has ever been pushed through to previously reported figures
  • Ask to see the evidence pack from the last cycle. If it was assembled for the audit rather than generated by the system, the control is manual

Anti-pattern · Automating the reduction claim before the baseline exists

Stage 4 makes the account credible, which is exactly when someone proposes wiring optimisation savings straight into the reported reduction. Energy models estimate a counterfactual, and a counterfactual is not a measurement. Without a documented baseline period and a comparable untreated line, shift or building, the claimed saving cannot be separated from volume, mix and weather. The first challenge to an unsupported reduction claim tends to discredit the whole account, including the parts that were sound.

What holds you here

The account is defensible but passive — it reports the past accurately and does not yet change a sourcing, engineering or dispatch decision.

Highest-leverage next move

Wire the verified figure into the decisions that create it: supplier award scoring, make-or-buy, material substitution and plant energy dispatch.

Cost of leaving

Effort
12-18 months
Team
Ledger owner, model owner, internal audit or quality partner, external assurance liaison
Risk
Higher — control gaps found late become restatements, and restatements are visible to the market
To next stage
12-18 months

If this is you, the next step is

We sample five published figures and time the reconstruction against a verifier's question set.

Rehearse a limited-assurance sample

Stage 5

Steered

3% of operators sit here

The same verified figure that leaves the building in a disclosure also steers sourcing, engineering and energy decisions inside it, under a versioned policy.

Stage 5 is narrower than the marketing language around it suggests. It is not an autonomous carbon-management system. It is the specific arrangement in which one ledger serves both the outward-facing disclosure and the inward-facing decision, so that the number a purchasing manager sees when scoring a quotation is the number that will appear, months later, in an assured statement. That single-source property is the whole point: the moment the decision runs on a different figure from the disclosure, both become negotiable.

The hard part here is governance rather than engineering. If carbon intensity carries weight in an award decision, the weighting is a commercial policy with legal consequences and it has to be versioned, reviewed and defensible to a supplier who lost on it. If an engineering change is approved partly on a footprint delta, the estimator behind that delta has just become part of the product development process and inherits its documentation expectations. The organisations that do this well tend to route the policy through the same forum that handles customer-specific requirements, because it is the forum that already knows how to version a commitment.

Sustaining stage 5 is the part most often underestimated. Factor sets move, suppliers change plants and power contracts, regulations amend their scope and reporting dates, and a threshold set against 2026 conditions quietly stops being valid. The leading indicator worth watching is the rate at which decisions fall outside the policy's stated bounds and route to a human: when it rises, the world has moved and the policy needs a review before an incident forces one.

In practice

The award that turned on a declared figure

A manufacturer scoring two aluminium castings suppliers found the quotations within a per-part cost of each other and materially apart on declared cradle-to-gate footprint. The award went to the higher-cost supplier on a documented weighting, and the losing supplier asked for the basis. Because both figures were supplier-declared under a common method with recorded versions, the answer was a two-page extract from the ledger rather than a negotiation. The same rows fed the following year's disclosure unchanged.

What it looks like

  • Carbon intensity is a scored criterion in supplier award decisions, with the data source named
  • Engineering change proposals carry a footprint delta computed from the same ledger
  • Plant energy dispatch and the reported Scope 2 figure resolve to one set of meters
  • The decision policy — thresholds, weights, escalation — is versioned and reviewed like code

Diagnostic signals you can check this week

  • Ask a purchasing manager where the carbon figure on their scoring sheet comes from. At stage 5 they name the ledger
  • Check whether the reported Scope 2 figure and the plant's energy-dispatch model read the same meters
  • Ask whether the award-scoring weights are versioned, dated and reviewed by a named forum
  • Track the share of decisions escalating outside policy bounds. A rising trend means the thresholds have expired

Anti-pattern · Treating the decision policy as a spreadsheet setting

Weights and thresholds get tuned in a configuration sheet with no version history, no review record and no note of who changed what. It works until a supplier challenges an award, a regulator asks how a threshold was set, or an internal auditor asks which policy version applied on a given date. Version the policy, minute the reviews, keep the trail — the artefact that will be examined is the policy, not the model.

What holds you here

Sustaining the arrangement is a governance problem — thresholds expire, factor sets move and the policy must be reviewed on a cadence, not on an incident.

Highest-leverage next move

Treat the decision policy as a versioned, reviewable artefact with the same rigour as the estimator, and monitor the escalation rate as its expiry signal.

Cost of leaving

Effort
Continuous
Team
Ledger and model owners plus a standing forum spanning purchasing, engineering, quality and sustainability
Risk
Concentrated — low frequency, high consequence, commercial and regulatory in nature

If this is you, the next step is

We run a real scoring decision against the policy, the trail and a supplier challenge scenario.

Stress-test a carbon-weighted award policy

Where automotive manufacturers actually sit on the ladder

The distribution across the five stages, why the stage 2 to 3 move is the largest single loss, and the regulatory dates that make it urgent rather than worthy.

Most automotive manufacturers are at stage 2 — part-level figures exist, a model fills the gaps, and the ledger cannot tell the two apart. The distribution below is weighted heavily toward that position: a large majority have moved off pure spend proxies for their main commodities, and a small minority have a figure that a stranger could reproduce from the trail. The drop between stage 2 and stage 3 is the largest single transition loss on the ladder, and it is not a modelling gap. It is a schema gap.

Distribution of automotive manufacturers across the provenance ladder

Illustrative distribution. Stage 2 is both the mode and the plateau: a modelled account that reads well internally and cannot survive being sampled. Percentages are a model-derived reading of the published evidence below, not a survey result.

Share of manufacturers

  • 21% — 1 · Spend-proxy
  • 38% — 2 · Modelled (the plateau)
  • 26% — 3 · Attributed
  • 12% — 4 · Verified
  • 3% — 5 · Steered

Source: Illustrative, synthesised from CDP supply-chain reporting and the EU disclosure frame

The reason the plateau exists is that stage 2 is genuinely comfortable. Internally the account looks sophisticated: it speaks in kilograms, it responds to engineering questions, and the year-on-year movement is explicable to a management committee. Nothing exposes the missing provenance until an outsider samples a figure — and by then the number has been published. Manufacturers that have been through one assured cycle almost never describe the experience as an accuracy problem; they describe it as five weeks of people reconstructing things from inboxes.

The urgency is calendar-driven rather than philosophical. The CBAM definitive regime (opens in a new tab) began on 1 January 2026 for covered goods including iron, steel and aluminium, which is most of a vehicle's structural mass by weight. The EU CO2 emission performance standards for cars and vans (opens in a new tab) set a 55% reduction for new cars by 2030 and 100% from 2035 against the 2021 baseline, with excess-emissions premiums attached. The Batteries Regulation (opens in a new tab) adds a carbon footprint declaration and, from February 2027, a battery passport. And the CSRD's scope and timing were themselves amended by the 2025 simplification package, so the reporting perimeter your programme was designed against in 2024 is not necessarily the one that applies now — a reason to build provenance that is robust to scope changes rather than tuned to one deadline.

Sector context is worth reading alongside the regulation rather than instead of it. ACEA (opens in a new tab) publishes the European manufacturers' position and the fleet data behind it, the European Environment Agency (opens in a new tab) monitors registered new-vehicle CO2 across the EU, and the Science Based Targets initiative (opens in a new tab) sets the validation rules that decide whether a corporate target counts as science-based at all. None of these bodies asks you to stop estimating. All of them assume you can say how you estimated.

One carbon number, six consumers — and each accepts different evidence

The same figure lands in a sustainability statement, a customs declaration, a customer's contract file and a type-approval dossier. Only one of them will not tolerate a model at all.

A carbon figure produced once is consumed at least six times, and the six consumers have completely different tolerances for a modelled value. This is the single most useful thing to internalise about compliance in this domain, because it dissolves the unhelpful question — is AI allowed in carbon accounting? — into six answerable ones. An ESRS E1 statement expects estimates and asks you to quantify how many. A CBAM declaration runs on rules about actual versus default values at the installation level. A type-approval CO2 figure comes from a prescribed test procedure and admits no model at all.

Consumer of the figureWhat it isWhere a modelled value standsEvidence it asks forWho checks it
ESRS E1 sustainability statementGross Scope 1, 2 and 3, intensity, targets and transition plan under the CSRDExpected — the standard asks you to disclose the extent to which Scope 3 rests on primary data from the value chainMethod, boundary, factor sources, the primary-data share and a traceable path from figure to sourceAssurance provider, limited assurance
CBAM declaration on imported goodsEmbedded emissions in imported iron, steel, aluminium and other covered goodsConstrained — the regulation governs when actual installation data must be used and when default values are permittedInstallation-level data where required, supported by verificationAccredited verifier and the national competent authority
Customer part-level PCF responseA contractual answer to an OEM or tier-one request, part number by part numberPermitted if declared, and increasingly required to state its data-quality basisDeclared method, boundary, calculation version and a data-quality ratingCustomer procurement quality, then their assurance provider
EV battery carbon footprint declarationA declaration for batteries placed on the EU market under the Batteries RegulationBounded by the methodology set in the delegated act, calculated per battery model and manufacturing plantPlant- and model-specific calculation following the prescribed methodConformity assessment route defined by the Regulation
Science-based target progressProgress against a validated near-term or net-zero targetPermitted, but a method or model change must be handled as a recalculation, never presented as progressA stable baseline, a written recalculation policy and consistent boundaries over timeThe target-validation body, and your own board
Type-approval CO2 for the vehicleThe regulated CO2 value for a vehicle type, feeding fleet complianceNot applicable — the value comes from a prescribed test procedure, not from an estimateType-approval test evidence under the applicable vehicle regulationType-approval authority and technical service
The six consumers of an automotive carbon figure and the evidence each one asks for. Read your own estate down the right-hand column: the consumer with the tightest tolerance sets the design of the ledger, not the loosest.

The asymmetry in that table is the design constraint. Because one ledger serves all six, its schema has to satisfy the strictest consumer that will ever read a given row — and which consumer that is depends on the commodity, not on the reporting team's preference. A modelled footprint for a plastic clip will only ever inform an ESRS E1 total. A modelled footprint for an imported aluminium casting may end up adjacent to a customs declaration, where the CBAM rules on actual versus default values (opens in a new tab) apply. The practical consequence is that provenance discipline should be introduced commodity by commodity in order of regulatory exposure, starting with the CBAM-covered materials, rather than uniformly across a bill of materials with ten thousand lines.

Jan 2026

CBAM definitive regime begins for covered goods including iron, steel and aluminium

European Commission

Feb 2027

Battery passport required for batteries placed on the EU market

European Commission

2035

100% CO2 reduction required for new cars registered in the EU, against 2021

European Commission

Underneath the consumers sits the estate that produces the numbers. Automotive carbon data comes from six operating domains, each with its own system of record, its own AI use and its own regulated destination. The map below is how we scope provenance work with manufacturers: pick the domain where the regulated consumer is tightest and the system of record is already yours, and start there.

DomainWhere AI is actually usedSystem of recordThe figure it movesRegulated consumer
Purchased materials and partsGap-filling missing PCFs, mapping purchase lines to activity data, extracting figures from supplier declarationsERP purchasing, PLM and BoM, supplier PCF exchangeScope 3 category 1 — purchased goods and servicesESRS E1, customer PCF requests, CBAM on imported goods
Plant energy — press, body, paint, utilitiesOven and booth set-point control, compressed-air leak detection, load shifting against tariff and grid carbon intensityEMS, MES, building management, sub-meteringScope 1 and Scope 2, energy per vehicle producedESRS E1, ISO 50001 energy review
Inbound and outbound logisticsLoad consolidation, mode shift, empty-run reduction, milk-run redesignTMS and carrier EDIScope 3 categories 4 and 9 — transport and distributionESRS E1, transport stage of a customer PCF
Battery and cell supplyCell supplier data validation, plant-level energy attribution, chemistry and sourcing scenario modellingCell supplier declarations, cell plant EMSBattery carbon footprint per kWh of capacityBatteries Regulation declaration and battery passport
Product engineeringFootprint deltas on material substitution, gauge reduction and design alternativesPLM, CAD and BoM, LCA toolingCradle-to-gate footprint per vehicle programmeESRS E1 transition plan, ecodesign and product passport
Use phase and fleetFleet-mix and demand scenario modelling behind the transition planSales, homologation and fleet systemsScope 3 category 11 — use of sold productsFleet CO2 standards and type approval, ESRS E1
The automotive carbon estate. 'Where AI is actually used' lists deployments we see in production today, not aspirations — and each row's regulated consumer sets the evidence standard for everything in it.

Two rows in that map deserve a warning. Use-phase emissions dominate an OEM's total footprint by a wide margin, which makes category 11 tempting to model aggressively — but the underlying per-vehicle CO2 figure is a type-approval output governed by the vehicle regulation regime (opens in a new tab) and its prescribed test procedures. Model the fleet mix and the lifetime mileage assumption, document both, and leave the regulated value alone. The second warning is the battery row: the Batteries Regulation (opens in a new tab) requires a calculation tied to a specific battery model made at a specific plant, so an estimate built from a cell chemistry average is structurally the wrong shape no matter how accurate it is. Precision does not fix a boundary error.

The materials rows are where the sector bodies help most. European Aluminium (opens in a new tab) and worldsteel (opens in a new tab) publish the methodological groundwork for primary and secondary metal footprints, which matters because the difference between primary and recycled aluminium is large enough to swamp most other levers in a body-in-white. If your gap-fill model cannot distinguish the two for a given part, the estimate is not merely uncertain — it is uncertain in a direction that procurement decisions actively change.

What the upper stages look like in public

Three publicly reported programmes, read against the provenance ladder. None is an Atomic Loops engagement — each links to the manufacturer's own published material.

The clearest public evidence for the ladder is in what large manufacturers chose to publish and what they chose to contractualise. In each case below, the distinguishing move was not a better model — it was that a carbon figure acquired a defined method, a named owner and a consequence. Read the outcomes as reported by the companies themselves; we have not independently audited them, and the linked pages are the manufacturers' own.

Three programmes read against the ladder

Outcomes as reported by the manufacturers in their own published material. Card images are generated industry scenes from our illustration library, not photographs of these companies' facilities, and imply no endorsement or involvement.

Illustrative scene: an automotive engineering studio with a vehicle under a data overlay and a server rack alongsideBMW GroupPremium OEM · Munich · multi-brand vehicle and motorcycle production35
Challenge
The large majority of a premium vehicle's cradle-to-gate footprint sits in materials and components the manufacturer does not meter — steel, aluminium and battery cells produced in other companies' plants, on other companies' grids.
Approach
BMW Group publishes life-cycle CO2e reduction targets expressed per vehicle, describes contracting for CO2-reduced steel and secondary aluminium, and states that CO2 performance is applied as a criterion in supplier selection. It is also a founding participant in the automotive data-space work built to exchange part-level product carbon footprints between tiers.
Reported outcome
As reported by the company, life-cycle CO2e per vehicle is tracked and published as a target metric, and carbon performance forms part of how suppliers are selected rather than a separate reporting exercise.
What it shows about the curveThis is the stage 5 signature: the same figure that appears in the disclosure also carries weight in an award decision. That only works if the underlying data is declared under a common method, which is why the data-space work and the contractual requirement arrived together rather than in sequence.

BMW Group PressClub (opens in a new tab)

Illustrative scene: a vehicle assembly environment with energy and emissions data displayed alongside the lineVolkswagen GroupMulti-brand OEM group · Wolfsburg · plants across Europe, the Americas and Asia24
Challenge
Producing one comparable decarbonisation figure across many brands and plants, each with its own systems, grids and reporting history, without the number becoming a negotiation between business units.
Approach
Volkswagen Group publishes a group-level decarbonisation measure expressed per vehicle across the life cycle, sets CO2 requirements for suppliers within its purchasing process, and participates in the cross-industry work on part-level footprint exchange.
Reported outcome
As reported by the company, a per-vehicle life-cycle decarbonisation measure is published and tracked at group level alongside supplier CO2 requirements applied through purchasing.
What it shows about the curveA group-level index is only meaningful if the boundary and allocation rules are frozen across brands and years. The hard work behind a headline figure like this is method stability, not calculation — which is exactly what the Attributed and Verified stages are made of.

Volkswagen Group — sustainability (opens in a new tab)

Illustrative scene: a plant utilities and energy monitoring area with production equipment visible beyondToyota Motor CorporationGlobal OEM · Toyota City · manufacturing across every major region24
Challenge
Reducing and evidencing plant CO2 across a very large global manufacturing footprint on grids whose carbon intensity varies by an order of magnitude between regions.
Approach
Toyota's published environmental programme sets long-range challenges covering both new-vehicle CO2 and plant CO2, and the company reports plant energy reduction and low-carbon power procurement by region within its environmental reporting.
Reported outcome
As reported by the company, plant CO2 reduction progress is published within a long-running environmental programme, with regional detail behind the group figure.
What it shows about the curvePlant Scope 1 and 2 is the cheapest primary data a manufacturer will ever hold — it is metered, invoice-reconciled and inside your own fence. The contrast between how well-evidenced that lane is and how thin the upstream lane is, in the same account, is precisely what the primary-data share metric exposes.

Toyota — sustainability (opens in a new tab)

What none of these programmes did was start with a model. In each, the sequence was the same: fix the boundary, obtain declared data for the commodities that matter, then apply computation to the remainder. That ordering is the whole argument of this page in three public examples — and it is the opposite of the order most internal programmes follow, which is to build the estimator first because it is the part a data team can deliver without anyone else's cooperation.

The four dimensions that set your stage

Provenance is not one number. Four dimensions gate each other, and the lowest one is the stage an assurance provider will find.

Your stage on the ladder is set by the lowest of four dimensions, not by the average. Activity data provenance, method and factor control, assurance evidence, and target and decision linkage each gate the others: an immaculately version-controlled method applied to spend-based inputs still produces a figure with no physical unit in it, and a rich primary-data estate with no reproduction path still fails a sample. The assessment on this page scores all four separately for exactly this reason.

  • Activity data provenance

    Where each number physically comes from — a meter, a supplier declaration, or a model — and whether the ledger records which. The binding question is your primary-data share by commodity, and it only becomes a real number once provenance is captured. The exchange infrastructure now exists: Catena-X (opens in a new tab) defines a common automotive data model for passing part-level PCFs between tiers, the VDA (opens in a new tab) coordinates much of the German-language guidance around it, and the trust framework suppliers are already assessed against for data exchange is TISAX (opens in a new tab). None of that raises your primary-data share on its own; a clause in the award does.

  • Method and factor control

    Whether the boundary, allocation rules, factor sets and global warming potential set behind a figure are recorded against that figure and versioned when they change. The GHG Protocol Product Standard (opens in a new tab) is the reference for the product-level rules, and ISO 14067 and ISO 14064-1 sit behind most sector guidance. The failure here is rarely choosing the wrong method; it is changing method silently, so that a year-on-year movement cannot be decomposed into real change and method change.

  • Assurance evidence

    Whether a stranger can rebuild a sampled figure from artefacts you already keep — the practical content of a limited-assurance engagement under IAASB's ISSA 5000 (opens in a new tab). Automotive has an advantage here that most sectors do not: manufacturers already run certified management systems with documented control plans and audit rehearsals under IATF 16949 (opens in a new tab) and ISO 14001. The muscle exists. It has usually just never been pointed at the carbon ledger.

  • Target and decision linkage

    Whether a reported reduction can be separated from volume, mix, price and weather, and whether the figure changes any decision inside the business. A claim with a documented baseline, a named lever and a comparable untreated line or period is evidence; a year-on-year difference is arithmetic. This is also the dimension that determines whether the ledger's controls ever get funded, because a figure that only leaves the building is a cost centre.

Diagnosing the real constraint

Plot your primary-data share against your provenance discipline. Three of the four quadrants have a next move that is not 'improve the model' — and the most dangerous position is the one that feels the most advanced.

Rich data, no trail

  • Good inputs, nothing recording how each figure was produced
  • The most dangerous quadrant — it reads as the most advanced
  • Fix: retrofit the ledger schema before publishing another cycle

Assurance-ready

  • Declared inputs and a reproducible trail
  • Constraint moves to control testing and decision linkage
  • Fix: rehearse a sample, then wire the figure into an award decision

Spend-proxy reporting

  • Neither foundation in place
  • Common at stage 1, and honest about it
  • Fix: rebuild one commodity on mass and process, not spend

Honest but coarse

  • Every figure labelled; most of them modelled
  • Highest-leverage position on the matrix
  • Fix: buy primary data with award clauses, not with a better estimator
Primary data share — top: Metered and supplier-declared, bottom: Spend factors and industry averages
Provenance discipline — left: One value per line, no method recorded, right: Class, version and uncertainty on every row

The counter-intuitive result is that 'honest but coarse' is a stronger starting position than 'rich data, no trail'. A manufacturer whose figures are mostly modelled but fully labelled has a working control environment and a data-acquisition problem, which purchasing can solve with a clause. A manufacturer with excellent primary data and no provenance schema has a data asset and no control environment, and every cycle published in that state adds to the volume of figures that may later need restating.

The carbon-ledger reference architecture, layer by layer

What actually has to exist for each stage, which layer manufacturers habitually skip, and why the estimation layer is the least important part of the stack.

A stage-4 carbon capability needs six layers, and the order in which you build them decides whether the programme compounds or produces an expensive spreadsheet. The architecture below is deliberately unfashionable: nothing in it is vendor-specific, every layer is defined by what it must guarantee rather than by what product provides it, and the layer everyone wants to build first — estimation — sits fourth for a reason.

Layers required by stage

Each layer is annotated with the stage that first requires it. A programme trying to reach stage 3 without the method registry and the ledger is building a stage-2 estimate with better tooling.

  1. Source systems

    Stage 1+

    • ERP purchasing and BoMPart masters, mass, purchase lines, plant of origin
    • Plant metering and EMSSub-metered by shop, reconciled to utility invoices
    • Supplier declarationsHowever they arrive — portal, message or attachment
  2. Activity data layer

    Stage 2+

    • Unit normalisationKilograms, kWh, tonne-kilometres — never currency
    • Boundary rulesCradle-to-gate scope written once and referenced everywhere
    • Completeness checksWhich parts and sites have no figure at all
  3. Method and factor registry

    Stage 3+

    • Factor setsVersioned, dated, licence recorded, never a local download
    • Allocation and GWP rulesRecorded per calculation rather than assumed
    • Recalculation policyWhat triggers a baseline restatement, and who signs it
  4. Estimation layer

    Stage 2+

    • Gap-fill estimatorModel card, bias and spread study, stated uncertainty
    • Purchase-line classifierConfidence threshold with a human review queue
    • Declaration extractionStructured fields, with the source document retained
  5. Carbon ledger

    Stage 3+

    • Figure rowsValue, unit, provenance class, versions, owner, timestamp
    • Supersession chainA primary declaration replaces an estimate; both are kept
    • UncertaintyA band per row, aggregated per commodity and per scope
  6. Disclosure, exchange and assurance

    Stage 4+

    • Regulated exportsESRS E1, CBAM declaration, battery footprint declaration
    • Customer PCF responsesPart-level, with method and data quality declared
    • Reproduction and evidenceGenerated by the system, not assembled for the audit

Pipeline described

  1. Source systems (stage 1+) — ERP purchasing and BoM: Part masters, mass, purchase lines, plant of origin; Plant metering and EMS: Sub-metered by shop, reconciled to utility invoices; Supplier declarations: However they arrive — portal, message or attachment
  2. Activity data layer (stage 2+) — Unit normalisation: Kilograms, kWh, tonne-kilometres — never currency; Boundary rules: Cradle-to-gate scope written once and referenced everywhere; Completeness checks: Which parts and sites have no figure at all
  3. Method and factor registry (stage 3+) — Factor sets: Versioned, dated, licence recorded, never a local download; Allocation and GWP rules: Recorded per calculation rather than assumed; Recalculation policy: What triggers a baseline restatement, and who signs it
  4. Estimation layer (stage 2+) — Gap-fill estimator: Model card, bias and spread study, stated uncertainty; Purchase-line classifier: Confidence threshold with a human review queue; Declaration extraction: Structured fields, with the source document retained
  5. Carbon ledger (stage 3+) — Figure rows: Value, unit, provenance class, versions, owner, timestamp; Supersession chain: A primary declaration replaces an estimate; both are kept; Uncertainty: A band per row, aggregated per commodity and per scope
  6. Disclosure, exchange and assurance (stage 4+) — Regulated exports: ESRS E1, CBAM declaration, battery footprint declaration; Customer PCF responses: Part-level, with method and data quality declared; Reproduction and evidence: Generated by the system, not assembled for the audit
Step-by-step insights
Source systems — the part of the stack you cannot buy
Almost every difficult question in a carbon programme resolves to a source-system property: does the ERP record the plant a part was made in, does the BoM carry a mass, is the site sub-metered by shop or only at the incoming supply. These determine what resolution the account can ever reach, and no amount of downstream software raises them. Audit the source systems before selecting anything, because the shortest path to a better figure is frequently a field that already exists and is not populated.
Activity data — the layer that kills the spend proxy
The single most consequential rule in this architecture is that the activity layer holds physical units and never currency. Once a purchase line has become kilograms of a specified alloy from a named plant, the figure stops moving with commodity prices and starts moving with the world. It also becomes checkable: tonnes bought should reconcile to tonnes consumed and scrapped, and where they do not, the discrepancy is usually a data problem worth finding before an auditor finds it.
The method and factor registry — versioning is the whole feature
Factor sets are updated, GWP values differ between assessment reports, and allocation conventions get refined. A registry that stores each set with a version, a date and a licence, and that records which version produced which figure, is what makes year-on-year movement decomposable into real change and method change. Without it, the honest answer to 'why did the number move' is 'several things changed and we cannot separate them', which is the sentence that starts a restatement.
The estimation layer sits fourth on purpose
It is the layer teams want to build first, because it is the one a data function can deliver without anyone else's cooperation, and it is the layer that adds least on its own. An estimator without normalised activity data is fitting to noise; an estimator without a factor registry cannot say what it computed against; an estimator without a ledger has nowhere to write its provenance. Build it fourth and it takes weeks. Build it first and it gets rebuilt.
The carbon ledger — a general ledger, not a data warehouse
The mental model that works is accounting, not analytics. Entries are immutable, corrections are made by superseding rather than overwriting, every entry names an owner, and the trial balance — the total — is derived from entries rather than maintained separately. Manufacturers who model this on their financial ledger get to reuse a set of conventions that have been stress-tested for a century, and they get a finance function that already understands what is being proposed.
Disclosure and exchange — export views, never second copies
Each regulated consumer wants a different shape: ESRS E1 by scope and category, CBAM by good and installation, a customer request by part number, a battery declaration by model and plant. All four should be views over the same ledger rows, generated mechanically. The moment any of them becomes a maintained spreadsheet that someone adjusts before submission, the estate has two versions of the truth and the divergence will be discovered externally.

One governance note about the estimation layer. Models used for corporate carbon accounting are not, on their face, in the high-risk categories of the EU AI Act (opens in a new tab), and this page does not argue otherwise. But the documentation those categories demand — data governance, logging, human oversight, performance monitoring — is almost exactly what a sustainability assurance provider asks for, and the NIST AI Risk Management Framework (opens in a new tab) gives the same artefacts a neutral, internationally recognised shape. Building the model card, the change log and the monitoring record once means you have already answered two audiences: the one that asks about your AI governance and the one that asks about your carbon controls.

The layer most often skipped is the method and factor registry, and skipping it is what makes the second reporting cycle harder than the first. A programme with a registry answers a change in factor set by rerunning; a programme without one answers it by re-deriving, and re-derivation is where quietly inconsistent methods enter an account that used to be consistent.

A 90-day plan: making the aluminium castings footprint verifiable

The stage 2 to 3 move made concrete on one automotive commodity — high-pressure die-cast aluminium structural parts, which are both a large Scope 3 line and a CBAM-covered material. Contains no model development.

Moving one stage takes about 90 days when it is scoped to a single commodity, and several years when it is scoped to a reporting perimeter. To make that concrete, the plan below runs the transition on one commodity most vehicle manufacturers and tier-ones share: high-pressure die-cast aluminium structural parts. It is chosen deliberately. Aluminium's footprint varies by an order of magnitude between primary and recycled metal and between grids, it is heavy enough to matter in a body-in-white, and as an imported good it sits inside the CBAM perimeter — so the same work serves the disclosure and the customs declaration. The gap-fill model already exists at most stage-2 manufacturers, so the quarter contains no model development at all.

Stage 2 to stage 3 on one commodity, in one quarter

One commodity family, one owner, one boundary. If a phase needs more than its window, narrow the scope — fewer part numbers, one plant — rather than extending the plan.

  1. Days 1-20

    Freeze the boundary and take a provenance census

    Pull twelve months of purchase lines and part masters for the commodity. Write the boundary down once — cradle-to-gate, what is included, how scrap and recycled content are allocated — and get the LCA practitioner and the commodity buyer to sign the same page. Then classify every line by how its figure was produced: supplier-declared, modelled, or spend-proxy. The output is a census by tonnage and by spend, and it is almost always more uncomfortable than expected.

    A written boundary and a provenance census by tonnage

  2. Days 21-45

    Buy the primary data with a clause, not a portal

    Rank suppliers by tonnage and request part-level footprints from the top of the list, using the common automotive data model rather than a bespoke template. Attach the request to the commercial relationship: a PCF obligation in the next award, a stated method, a stated deadline, and a plain description of what the figure will be used for. Offer the smaller suppliers a calculation route rather than only a form, because most tier-twos are not withholding a number — they do not have one.

    Declared PCFs covering the majority of commodity tonnage

  3. Days 46-70

    Qualify the estimator like a gauge

    Back-test the gap-fill model against every primary PCF that arrived in phase two and against any that arrived over the previous year. Compute signed bias and spread, split by primary versus recycled input where the data allows. Publish a one-page model card: what it estimates, on what inputs, its bias, its spread, its uncertainty band and who owns it. Then write the provenance class, model version, factor-set version and uncertainty into every remaining ledger row.

    A calibrated estimator with a stated error, and no unlabelled figures

  4. Days 71-90

    Rehearse the reproduction against the clock

    Sample five published figures across the commodity — at least one declared, one modelled and one superseded. Give a two-person team the verifier's standard question set with the original analysts excluded, and time each reconstruction from value to method version to factor set to source record. Fix anything that takes over an hour. Package the result as the evidence pack for the commodity and name its owner.

    Five figures reproduced end to end, an evidence pack, a named owner

The order matters

  1. Provenance before precision

    A moderately accurate estimate that is labelled, versioned and bounded passes review. An excellent estimate presented as a fact does not. Improve the model after the ledger can carry its provenance, when you can price an accuracy point in disclosed tonnes rather than in error percentage.

  2. Boundary before benchmarking

    Comparing your figure against a competitor's, a sector average or last year's before the boundary is written down produces a difference nobody can attribute. Two footprints for the same casting can differ by a factor of three on allocation rules alone, entirely legitimately.

  3. Award clauses before better estimators

    Raising the primary-data share is a commercial act, not a technical one. One clause in a sourcing decision moves the number further than a quarter of model work, and it moves it in the direction the disclosure standard actually asks about.

Instrumenting it: the metrics, the thresholds and the readiness checklist

Where each control metric actually comes from — the formula, the system that produces it, and the stage at which it starts measuring something real. All telemetry, no self-report.

A control you cannot name a source system for is an intention. Every metric below reduces to counts and timestamps that the ledger, the model registry and the supplier exchange already record — the instrumentation work is joining them, not creating them. Read the 'honest from' column as a warning: a primary-data share computed before provenance exists is a guess with a decimal point on it.

MetricFormula / readSourceCadenceHonest from
Primary-data shareTonnage covered by supplier-declared PCF ÷ total purchased tonnageCarbon ledger + ERP purchasingMonthly, by commodityStage 3
Provenance coverageLedger rows carrying class, method version and factor version ÷ all rowsCarbon ledgerWeeklyStage 3
Gap-fill biasMean signed (modelled − primary) ÷ primary, on parts where primary later arrivedSupersession logPer reporting cycleStage 3
Gap-fill spreadInterquartile range of the same ratio, split by material classSupersession logPer reporting cycleStage 3
Supersession lagDays from a primary declaration arriving to the ledger reflecting itExchange inbox + ledger timestampsMonthlyStage 3
Reproduction timeMinutes to rebuild a sampled figure to source, analyst excludedRehearsal logQuarterly sample of fiveStage 4
Factor-set currencyAge of the factor set in use, per commodityMethod and factor registryQuarterlyStage 4
Restatement ratePreviously reported figures restated ÷ figures reportedLedger supersession chainPer cycleStage 4
Attributed reductionTonnes claimed against a documented baseline and counterfactual ÷ total tonnes claimedReduction registerPer cycleStage 4
AI estate energykWh attributed to training and inference for the carbon stackCloud billing or DC sub-meteringMonthlyStage 4
Instrumentation build sheet for a carbon-provenance programme. 'Honest from' is the stage at which the metric first measures something real rather than something asserted.

Two of those rows deserve comment. Gap-fill bias is the most valuable metric on the page and the cheapest to compute, because arriving primary declarations continuously refresh the validation set for free — every supersession is a labelled example the estimator was tested on in production. And AI estate energy is the row that surprises people: the carbon-accounting stack is itself an emitting asset, and the IEA's work on energy and AI (opens in a new tab) has made data-centre load a mainstream question rather than a footnote. It is a small number for most manufacturers and an awkward one to be asked about without an answer.

The stage transitions themselves are verified by four observable measures, each readable from the same systems.

MeasureStage 2Stage 3Stage 4How to read it
Provenance coverageNear zeroEvery row classedClassed and versionedShare of ledger rows carrying class and versions
Estimator qualificationNoneBack-testedBias and spread trackedDate and sample size of the most recent study
Reproduction timeDays, or not at allWithin a dayUnder an hourTimed rehearsal with the original analyst absent
Reduction attributionYear-on-year differenceBaseline and lever namedCounterfactual heldWhether an untreated line, shift or period exists
Verification measures for each transition on the provenance ladder. All four are readable from the ledger and the registry rather than from a self-assessment.

Limited-assurance readiness checklist

Eight controls. If you cannot tick all eight, the account is still at stage 2 regardless of how good the underlying figures are. Tick as you go — this list works without JavaScript.

0 of 8 ticked

Tick honestly — the blank list is data too

Most stage-2 manufacturers can genuinely tick one or two of these rather than none. If nothing applies yet, do not start with tooling: run the 90-day plan above on one commodity. Six of these eight controls fall out of doing that once, and the other two are policy documents that take an afternoon each.

Where this is heading is more part-level data, not less. The Ecodesign for Sustainable Products Regulation (opens in a new tab) extends product-level information requirements and digital product passports well beyond batteries over time, and every extension pushes the same demand down the supply chain: a figure per part, with a method attached. A manufacturer with a provenance-carrying ledger answers each new requirement with an export. A manufacturer without one answers it with another project.

Failure modes that turn a good carbon programme into a restatement

Provenance is not monotonic. Four regressions account for almost all of it, and each one happens quietly enough that the account looks healthy until it is sampled.

Manufacturers regress on this ladder, usually without noticing, because the conditions that supported a stage stopped holding while the reporting kept working. That is the specific danger of a carbon account: unlike a production line, it produces plausible output right up to the moment somebody checks it, and the checking happens annually at best. Four regressions account for almost all of what we see.

Likelihood: highImpact: high

The estimator is retrained and nothing is recalculated

A routine model refresh — new training PCFs, a better feature — silently changes thousands of figures. Because the ledger records a value rather than a version, nobody can identify which reported numbers moved for model reasons, and this year's total differs from last year's for reasons that cannot be decomposed in front of an auditor.

PreventionEvery model release carries a quantified impact assessment on affected figures, and the ledger records the version so the affected set is a query rather than an investigation.

Likelihood: highImpact: medium

A supplier declaration arrives and dies in an inbox

Primary data is the scarcest asset in the account, and it typically arrives as an attachment to an email addressed to whoever chased it. Where the exchange path stops at a person, declarations sit unprocessed for months while the modelled figure they should have superseded is published and assured.

PreventionThe exchange path writes into the ledger directly, and supersession lag — days from arrival to reflection — is monitored as a standing metric with an owner.

Likelihood: mediumImpact: high

A price movement is reported as a decarbonisation

Any account with spend-derived lines will occasionally record a commodity price fall, a currency move or a mix shift as a reduction. Reported once with confidence, it is very hard to withdraw, and the correction lands in the following cycle as an unexplained increase.

PreventionA reduction register that requires a baseline, a named lever and a counterfactual before a claim can be published; anything without all three is reported as movement, not as reduction.

Likelihood: mediumImpact: medium

The boundary drifts as commodities are added

Each new commodity brought into the account arrives with its own allocation convention, usually inherited from whichever consultant or supplier tool produced the first figures. Two years later the account is internally inconsistent in ways that only surface when someone compares two commodities' intensities and gets an implausible ratio.

PreventionOne boundary document referenced by every figure, and a review gate that a new commodity cannot pass without being mapped onto it explicitly.

The common thread is that all four are cheap to prevent and expensive to discover. A model impact assessment is an afternoon; an unexplained year-on-year movement is a board conversation. A supersession-lag metric is a query; a superseded figure published under assurance is a restatement. This is the ordinary economics of control design, and automotive quality functions have argued it successfully for decades — which is the strongest available argument for putting the carbon ledger in front of people who already think this way.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Product carbon footprint (PCF)
The greenhouse gas emissions attributable to one part, material or product over a defined boundary — in automotive supply chains almost always cradle-to-gate, expressed as kg CO2e per part or per kilogram.
Primary and secondary data
Primary data is measured or supplier-declared for the specific activity being accounted for; secondary data is drawn from databases, sector averages or proxies. ESRS E1 asks a reporter to disclose how much of Scope 3 rests on primary data from the value chain.
Gap-fill estimate
A figure produced by a model for a part, site or activity where no measured or declared value exists. Legitimate and expected in a Scope 3 account — provided it is labelled as such and carries a version and an uncertainty band.
Provenance class
The field on a ledger row that records how its figure was produced: metered, supplier-declared, modelled or default. Its absence is the single most common finding in a first limited-assurance engagement.
Carbon ledger
The system of record for emissions figures, built on accounting rather than analytics conventions: immutable entries, corrections by supersession rather than overwriting, a named owner per row, and totals derived from entries.
Emission factor set
The versioned collection of kg CO2e per unit values applied to activity data. Factor sets are updated and licensed, so the version used must be recorded against each figure or year-on-year movements become undecomposable.
Cradle-to-gate boundary
The scope covering raw material extraction through to the point a product leaves the producing facility, excluding use and end of life. The most common boundary for automotive part-level footprints, and the one most often left unwritten.
Allocation
The rule that divides a facility's emissions between co-products, or between primary and recycled material streams. Two defensible allocation choices can move the same casting's footprint by a factor of several, which is why the rule belongs in the boundary document.
Recalculation policy
The written rule stating what triggers restating a baseline or a previously reported figure — a method change, a factor-set update, an acquisition, or a material correction — and who approves it. Without one, method changes and real changes arrive in the same number.
Limited assurance
The assurance level applied to sustainability statements under the CSRD, expressed as a conclusion that nothing has come to the practitioner's attention suggesting material misstatement. In practice it means sampling figures and tracing them to source.
Embedded emissions
Under CBAM, the greenhouse gases attributable to the production of an imported good such as steel or aluminium. The regulation governs when installation-level actual data must be used and when default values are permitted.
Counterfactual baseline
The documented statement of what consumption or emissions would have been without an intervention, against which a saving is claimed. Usually a comparable untreated line, shift, building or period — the thing that separates an optimisation result from a coincidence.

Frequently asked questions

The questions sustainability, purchasing and manufacturing IT leads ask most often when AI starts producing numbers that end up in a disclosure.

Does using AI to estimate emissions breach the CSRD or the GHG Protocol?

No. The GHG Protocol's Scope 3 standard explicitly anticipates secondary data and estimation, and ESRS E1 asks reporters to disclose how much of Scope 3 rests on primary data rather than to eliminate estimates. What fails review is an estimate that cannot be distinguished from a measurement, or one that cannot be reproduced when sampled. Label every modelled figure with its provenance class, model version, factor-set version and uncertainty band, and the same estimate becomes ordinary practice rather than a control finding.

What is the difference between primary, secondary and modelled data in a carbon account?

Primary data is measured or declared for the specific activity — a meter reading, a supplier's part-level footprint. Secondary data comes from databases or sector averages applied to your activity. A modelled figure is secondary data with an inference layer on top, where an estimator predicts a value from part attributes or spend categories. All three are legitimate inputs. The distinction matters because assurance providers, CBAM verifiers and customers each apply different tolerances to them, and because the primary-data share is itself a disclosable metric.

Can a modelled product carbon footprint be used in a CBAM declaration?

Only within the rules the regulation sets. CBAM governs when actual, installation-level data must be used for imported goods such as steel and aluminium and when default values are permitted, and declarations are subject to verification. That is a much tighter regime than an ESRS E1 total, where a labelled estimate is expected. The practical consequence is to sequence provenance work by regulatory exposure rather than by spend: get CBAM-covered commodities onto declared installation data first, and treat the rest of the bill of materials as a second wave.

How do we qualify an emissions gap-fill model the way we qualify a gauge?

Ask the three questions measurement systems analysis asks of any instrument. Is it biased — what is the mean signed difference between its estimates and primary declarations that arrived afterwards? How much does it vary — what is the spread of that difference, split by material class? What tolerance should its output be quoted with — what uncertainty band do you publish? Refresh the study each reporting cycle using arriving supplier declarations as the validation set, and record the answers on a one-page model card with a named owner.

What does ESRS E1 actually require about primary data?

It requires transparency about it rather than a minimum level of it. The standard expects gross Scope 1, 2 and 3 emissions, an intensity metric, targets and a transition plan, and it asks reporters to be explicit about the extent to which value-chain figures rest on primary data obtained from suppliers rather than on secondary sources. That framing is helpful: it converts an unanswerable question about whether estimation is allowed into a measurable one about what proportion of your account is declared, which you can then improve deliberately.

Does the EU AI Act apply to carbon-accounting models?

Models used for corporate carbon accounting do not, on their face, fall into the AI Act's high-risk categories, and nothing on this page argues that they do. The useful observation is different: the documentation those categories demand — data governance, logging, human oversight, performance monitoring — is almost exactly what a sustainability assurance provider asks for. Building a model card, a change log and a monitoring record once satisfies both audiences, and the NIST AI Risk Management Framework gives those artefacts a recognised shape if you would rather not invent one.

How do we count the emissions of the AI systems themselves?

As Scope 2 where the compute runs in your own data centre, and as Scope 3 category 1 where it runs on a cloud provider, using whatever consumption or emissions reporting that provider makes available. The number is usually small relative to a manufacturing footprint, but it is increasingly asked about, and having no answer reads badly in an account otherwise built on provenance. Instrument it the same way as everything else: a named source, a stated method and a monthly cadence.

How do we prove an AI-driven energy saving in the paint shop actually reduced emissions?

With a counterfactual, not with the optimiser's own estimate. The paint shop is typically a plant's largest energy consumer, so the savings are real and worth claiming — but consumption also moves with ambient conditions, colour mix, line rate and downtime. Establish a baseline period, keep a comparable untreated booth, line or shift where the process allows, and normalise for the drivers you cannot hold constant. A saving claimed from the model's predicted counterfactual alone will not survive the first challenge.

What does part-level PCF exchange change about supplier data?

It changes the unit of the conversation from a company-level questionnaire to a part number, and it standardises the shape of the answer so a supplier is not building a bespoke spreadsheet for each customer. That matters most for tier-twos and tier-threes, who receive similar requests from several customers at once. What it does not change is the incentive: suppliers respond to award criteria and contractual obligations, not to portals. Treat the data model as the delivery mechanism and the sourcing clause as the reason anyone uses it.

Who should own the carbon ledger — sustainability, finance or manufacturing IT?

Split it the way you split any regulated reporting system. Sustainability owns the method: boundary, allocation, factor selection and the disclosure narrative. Manufacturing IT owns the pipeline and the ledger's operational behaviour, including model versioning and the exchange path. Finance owns the control environment, because the artefacts — immutable entries, supersession, restatement policy, evidence trails — are the ones finance already runs on. Programmes with only a sustainability owner stall at stage 2 because nobody operates the pipeline; programmes with only an IT owner produce a fast, methodologically incoherent account.

How long does it take to get from a spend proxy to a verifiable figure?

About 90 days per commodity when scoped tightly, and two to three years for a full reporting perimeter. The quarter is spent on the boundary document, a provenance census, a supplier data request attached to a commercial lever, a bias study on the existing estimator and a timed reproduction rehearsal — none of which is model development. The perimeter takes years because raising the primary-data share depends on supplier leverage that arrives with sourcing cycles, and those cycles are not something a programme can accelerate on its own.

What happens when a supplier's primary declaration contradicts our modelled estimate?

Treat it as a supersession, not a correction. Write the declaration as a new ledger row that supersedes the estimate, keep both, and record the delta — the difference is a labelled test result for your estimator, produced in production and at no cost. If the delta is large, check the boundary before checking the model: differences of a factor of two or three between two defensible cradle-to-gate calculations usually come from allocation rules or recycled-content assumptions rather than from either party being wrong.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for automotive manufacturers and their supplier networks — energy and process optimisation, vision inspection, supplier data extraction and decision support running against live plant and ERP data, integrated into the systems that already hold the record rather than delivered as dashboards.

  • · Production deployments across body, paint, assembly and supplier data operations
  • · Carbon-ledger and provenance work run jointly with sustainability and manufacturing IT teams
  • · Integration-first delivery: ERP and EMS write-back, evidence trails, reproducible reruns
  • · 26 cited sources on this page

Sources

  1. European CommissionCorporate sustainability reporting (CSRD) (opens in a new tab)
  2. European CommissionCarbon Border Adjustment Mechanism (CBAM) (opens in a new tab)
  3. European CommissionCO2 emission performance standards for cars and vans (opens in a new tab)
  4. European CommissionBatteries and accumulators — the EU Batteries Regulation (opens in a new tab)
  5. European CommissionEcodesign for Sustainable Products Regulation (opens in a new tab)
  6. European CommissionRegulatory framework for AI (the EU AI Act) (opens in a new tab)
  7. GHG ProtocolCorporate Value Chain (Scope 3) Standard (opens in a new tab)
  8. GHG ProtocolProduct Life Cycle Accounting and Reporting Standard (opens in a new tab)
  9. IAASBISSA 5000 — general requirements for sustainability assurance engagements (opens in a new tab)
  10. Science Based Targets initiativeScience-based target setting and validation (opens in a new tab)
  11. CDPSupply chain programme and disclosure data (opens in a new tab)
  12. Catena-X Automotive NetworkAutomotive data ecosystem and standards (opens in a new tab)
  13. ACEAEuropean automobile manufacturers' association (opens in a new tab)
  14. VDAGerman association of the automotive industry (opens in a new tab)
  15. ENX AssociationTISAX — trusted information security assessment exchange (opens in a new tab)
  16. IATF Global OversightIATF 16949 automotive quality management oversight (opens in a new tab)
  17. UNECEVehicle regulations and type-approval regime (landing page; bot-walled to automated checks) (opens in a new tab)
  18. European AluminiumAluminium environmental methodology and sector data (opens in a new tab)
  19. worldsteelSteel life-cycle and sustainability methodology (opens in a new tab)
  20. European Environment AgencyMonitoring CO2 emissions from new vehicles in Europe (opens in a new tab)
  21. International Energy AgencyEnergy and AI (opens in a new tab)
  22. International Energy AgencyData centres and data transmission networks (opens in a new tab)
  23. NISTAI Risk Management Framework (opens in a new tab)
  24. BMW GroupPressClub — corporate and sustainability announcements (opens in a new tab)
  25. Volkswagen GroupSustainability reporting and decarbonisation programme (opens in a new tab)
  26. Toyota Motor CorporationSustainability and environmental reporting (opens in a new tab)

Find out whether your carbon numbers would survive being sampled

We run the assessment with your sustainability, purchasing and manufacturing IT leads, sample published figures and time the reconstruction, then leave you with a costed 90-day plan for your weakest dimension. You keep the plan whether or not we build it.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.