Redefining Technology

Manufacturing (Automotive)Regulations, Compliance & Governance

AI audit readiness in automotive manufacturing: the evidence an auditor will ask you to produce

AI audit readiness is the ability to produce, on demand, the evidence that an AI-influenced process is under control: the control plan entry, the measurement-system analysis, the model documentation, and the traceable record of why one specific part passed on one specific date.

Illustrative scene: an automotive body-in-white line with robotic inspection cells and quality engineers reviewing on-screen inspection data
Manufacturing (Automotive) · Regulations, Compliance & Governance

Key takeaways

  1. An auditor does not audit your model. They audit your process, and the model is now part of it — which means the control plan, the PFMEA, the measurement-system analysis and the work instruction all have to say so before the audit, not during it.
  2. If a model influences an accept or reject decision, it is a measurement system. Attribute agreement analysis against a reference sample is the evidence that satisfies MSA expectations, and it has to be re-run on every model version, not once at installation.
  3. The single question that separates readiness from paperwork is reconstruction: show why this specific part passed on this specific date. Answering it needs the part identity, the model version, the input artefact, the score and the threshold — all retained together and joinable.
  4. Adding a model to an already-approved process is a process change. Where the customer's requirements demand notification or resubmission, an undeclared model turns a good inspection into a PPAP problem, regardless of how well it performs.
  5. Audit readiness is a rehearsed capability, not a document set. Plants that pick a random part quarterly and time the reconstruction find their gaps as internal findings; plants that do not find them in front of a certification body.

Abbreviations used on this page

IATF
International Automotive Task Force — publisher of IATF 16949, the automotive quality management standard
VDA
Verband der Automobilindustrie — the German automotive association behind VDA 6.3 process audits
APQP
Advanced product quality planning
PPAP
Production part approval process
FMEA
Failure mode and effects analysis (PFMEA when applied to a process)
MSA
Measurement systems analysis
GR&R
Gauge repeatability and reproducibility — the variation study inside an MSA
SPC
Statistical process control
CSR
Customer-specific requirements — the OEM rules layered on top of IATF 16949
QMS
Quality management system
MES
Manufacturing execution system
TISAX
Trusted Information Security Assessment Exchange — the automotive information-security assessment scheme

Free · 8 questions · ~3 minutes

Score your plant on audit readiness

Eight questions, one at a time, about three minutes. Answer them and we build your personalised readiness report — your stage on the Unprepared to Continuously auditable ladder, your score on each of the four dimensions, and the specific evidence gap standing between you and the next stage — and send it to your inbox. Your result doubles as the agenda for your next internal audit.

0 of 8 answered

Question 1 of 8Evidence completeness

If an auditor picked one AI-influenced inspection today, what could you put in front of them within the hour?

Audits are conducted in the particular. The gap between what you can describe and what you can produce is the finding.

How the score maps to a stage
  • 0–5 — Stage 1, Unprepared. AI already influences decisions on the floor, but nothing in the quality system records that it exists.
  • 6–11 — Stage 2, Documented. The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
  • 12–16 — Stage 3, Traceable. Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
  • 17–21 — Stage 4, Rehearsed. The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
  • 22–24 — Stage 5, Continuously auditable. The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.

What AI audit readiness means in an automotive plant

A definition, the question every audit regime reduces to, and the evidence chain that has to answer it before the auditor arrives.

AI audit readiness is the ability to produce, on demand, the evidence that an AI-influenced process is under control. In an automotive plant that is a concrete list rather than a posture: the control plan entry that names the AI station and its reaction plan, the process FMEA rows covering how the model fails, the measurement-system study that characterises it as a gauge, the model documentation that states what it is for and how it was validated, and the per-part record that shows why this specific body passed on this specific date and shift.

Read from the auditor's side of the table, the framing changes. An auditor is not there to assess your model. They are there to assess your process — and the model has become part of it, which means it inherits every expectation the process already carried under IATF 16949 (opens in a new tab), the IATF Rules for the certification scheme (opens in a new tab), and whatever your customers add through their own requirements. Nothing about that is AI-specific. What is AI-specific is how easily a model slips into a process without any of those documents noticing, because retraining does not look like a process change to the people doing it.

Effort to answer an auditor, as readiness matures

The curve is not linear. Through the Unprepared and Documented stages, answering a specific question about a specific part is a manual assembly job whose cost barely falls — better documents do not make per-part evidence appear. The inflection is at Traceable, when the record is emitted at decision time; from there, rehearsal and automation drive the marginal cost of an audit answer towards zero. Illustrative, and consistent with the evidence-lifecycle expectations set out in the NIST AI Risk Management Framework.

Evidence available without preparation by stage

  • Stage 1 · Unprepared — 22% of operators. AI already influences decisions on the floor, but nothing in the quality system records that it exists.
  • Stage 2 · Documented — 38% of operators. The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
  • Stage 3 · Traceable — 24% of operators. Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
  • Stage 4 · Rehearsed — 12% of operators. The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
  • Stage 5 · Continuously auditable — 4% of operators. The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.

Curve shape: logistic, plotted from the stage data above. Distribution: Illustrative — framed against the NIST AI Risk Management Framework.

  • Is the process defined?

    Every regime opens here. The control plan, the work instruction and the PFMEA have to describe the process that is actually running, including the part the model plays in it. A station whose documents describe the pre-AI method fails this question before any discussion of model quality begins.

  • Is it capable and in control?

    For a conventional gauge this is MSA and capability. For a model making accept and reject calls it is an agreement study against reference parts, a stated performance floor, and evidence that the floor is monitored rather than assumed. The AIAG core-tools material (opens in a new tab) is the vocabulary your auditor will use for it.

  • Can you prove it for this part?

    The particular, not the general. Serial number in, evidence out: station, timestamp, model version, input artefact, score, threshold, decision, disposition. This is the question that separates a tidy document file from a readiness capability, and it is the one most plants cannot answer inside an audit slot.

  • What did you do when it was wrong?

    Models are wrong sometimes; that is not the finding. The finding is a nonconformance handled as a station fault, with no link to the model version, no containment scoped by deployment window, and no update to the FMEA that claimed to have considered the failure mode.

How an auditor's question gets answered — or does not

The same question, asked three ways down the plant. The top lane is where most plants are for their newest AI station: the question routes to a person, who routes it to another person, and the answer is a recollection. The middle lane answers from the quality system. The bottom lane answers without preparation, because the evidence was emitted when the decision was made.

  • Human in the loop
  • Where value leaks
  • System-of-record action
  • Data & feeds
  • AI / model

The process, in words

  • In the unprepared lane the question routes to people rather than to systems. Quality escalates to engineering, engineering reconstructs from memory and a screenshot, and because neither the model version nor the input artefact was retained, the answer cannot be evidenced. The finding that follows is written against document control and process change, not against the technology — which is why it pulls every other AI-assisted station into scope.
  • In the documented and traceable lane the auditor's question opens the control plan at that station, the control plan resolves to the decision records for that part, the record carries the model version, and the version resolves to the model file holding intended use, validation and change history. A quality engineer answers in the room, from the system, without calling the person who built the model.
  • In the continuously auditable lane nothing is assembled, because the evidence was emitted when the decision was made. Each inspection writes its own record, each model version is archived and re-runnable, and the auditor is given a filtered read-only view instead of a prepared pack. The same pipeline watches the stated performance floor, and a breach raises a nonconformance without anyone deciding to raise one.
Step-by-step insights
The screenshot answer — why it fails on category, not on content
A screenshot of the station screen usually shows the correct outcome, and it still fails, because an audit asks for a record rather than an observation. A record has provenance: it was written by the system at the time of the event, it is retained under a schedule, and it cannot be produced after the fact. The distinction matters most when the answer is favourable — a plant that can only demonstrate good results without records has demonstrated that it has no records, which is the finding regardless of the result.
The control plan as the entry point, not the summary
Auditors open the control plan at the station because it is the document that binds a characteristic to a method, a frequency, a sample size and a reaction plan. When a model takes over a characteristic, all five of those change: the method is now a classifier, the frequency is often every part rather than a sample, the sample size concept may disappear entirely, and the reaction plan needs a branch for the model being unavailable or out of its envelope. A control plan that merely mentions 'vision system' has recorded the equipment and lost the process.
The decision record is the join, and it needs a real key
The single most common structural failure is that decisions and identities live in different systems with no shared key. The station historian has timestamps and outcomes; the MES has serial numbers and build events; the deployment log has model versions and dates. Joining them by time is an interpretation, and interpretations do not survive an audit or a warranty investigation. Read the part identity at the station and write it into the decision record — that one field converts three systems into one answer.
Version pinning: naming a version is not the same as keeping it
Most plants that record a model version cannot re-run it, because the artefact was overwritten, the environment moved, or the preprocessing lived in code that has since changed. Pinning means the exact artefact, the exact preprocessing and the exact threshold configuration are retained together and can be executed against the retained input. That is what turns 'the record says version 7' into 'here is version 7 producing the same score on the same image', which is the difference between a claim and a demonstration.
The read-only auditor view — less exposure, not more
Plants hesitate to give an auditor a live view, on the theory that a prepared pack is safer. In practice the pack is the risk: it is assembled under time pressure, it invites questions about what was left out, and it proves nothing about the underlying system. A filtered, read-only, observed view scoped to one station and one date range shows the record as it exists, answers follow-up questions in seconds, and demonstrates the control everyone is actually trying to evidence.
The performance floor closes the loop back to the reaction plan
A stated floor — maximum escape rate on a safety characteristic, minimum agreement against reference parts, a false-reject ceiling the line can absorb — is what makes drift an event rather than a slow disappointment. Wire the breach into the same nonconformance route a gauge out of calibration takes, and the AI station stops being a special case. That is the whole trick of AI audit readiness: not new machinery for a new technology, but making the new technology reach the machinery that already exists.

The five stages in detail

For each stage: what it looks like on the floor, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving costs.

Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a plant, the diagnostic signals are checks you can run against your own quality system this week, and the anti-pattern is the specific mistake most often made trying to leave that stage. The ladder runs Unprepared, Documented, Traceable, Rehearsed, Continuously auditable — and the expensive gap is between Documented and Traceable, because that is where general description has to become per-part evidence.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Unprepared

22% of operators sit here

AI already influences decisions on the floor, but nothing in the quality system records that it exists.

Stage 1 is not the absence of AI — it is the absence of AI from the quality system. The station works, the operators like it, the scrap numbers improved, and the control plan on the wall still describes an ultrasonic spot check performed every fiftieth body. Nothing in the QMS knows the model is there, so nothing in the QMS can answer a question about it.

The pattern that produces this is ordinary and blameless. A vision or process model is installed as an engineering improvement, usually under a continuous-improvement budget rather than a programme change, and the change lands on the floor faster than the document control cycle that would have carried it. Quality is told about the result rather than consulted about the method. Six months later nobody can say precisely which version has been running since May.

The cost is asymmetric and it arrives all at once. Day to day the gap is invisible; in the audit room it produces the worst possible category of finding, because the auditor is not disputing the technology — they are observing that the documented process and the actual process are different things. That is a system finding, not a station finding, and it puts every other AI-influenced station in the plant into scope on the same day.

In practice

The station that was not on the control plan

A body-shop team installed a camera-based surface inspection ahead of the sealer booth and switched the manual visual check to a spot audit. Scrap fell, the line ran better, and everybody was pleased. Eleven months later a surveillance auditor walked the line, saw a screen showing pass and fail decisions, and asked where in the control plan that characteristic was recorded. It was not. The finding was written against document control and process change, and the corrective action pulled in four other AI-assisted stations nobody had thought to mention.

What it looks like

  • No register of which decisions a model touches
  • Control plans and work instructions still describe the pre-AI method
  • Model behaviour is explained verbally, from memory, by whoever built it
  • Station logs roll off long before an audit could reach back to them

Diagnostic signals you can check this week

  • Ask for a list of every decision on site that a model influences — if the answer takes more than a day to produce, you are here
  • Compare the control plan for one AI-assisted station against what the operator actually does at that station
  • Ask which model version has been running since a specific date three months ago, and watch who gets called
  • Check the retention on the station's decision logs against your customer's record-retention requirement

Anti-pattern · Writing an AI policy instead of a control plan entry

The instinctive response to an AI-shaped compliance worry is a governance document: a policy, a committee, an ethics statement. None of these is what an auditor asks for at a station. The artefacts that discharge the obligation are the ones the plant already uses — control plan, PFMEA, work instruction, MSA record, change record — and one station properly written into those four is worth more in an audit than a twenty-page policy nobody on the floor has read.

What holds you here

Nobody has an authoritative list of where AI already influences a decision, so the gap cannot be sized, let alone closed.

Highest-leverage next move

Build the register — station, decision, part families, owner, model version — then write the highest-exposure station into the control plan and PFMEA.

Cost of leaving

Effort
4–8 weeks
Team
One quality engineer, one controls or ML engineer, part-time
Risk
Low — the work is documentation and inventory, nothing on the line changes
To next stage
1–3 months

If this is you, the next step is

A two-week exercise: the register, the gaps, and the order to close them.

Inventory your AI-influenced decisions

Stage 2

Documented

38% of operators sit here

The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.

Stage 2 is the stage most plants reach as soon as somebody senior asks the question, and it is a genuine improvement: the documented process and the running process agree. The control plan names the station, the characteristic and the reaction plan; the PFMEA has rows for false accept and false reject; a validation report sits in the quality file with a date and a signature on it.

What is missing is the per-part evidence. Documentation describes the process in general, and audits are conducted in the particular. The auditor picks a serial number off a rack, or a date off a shipping record, and asks what happened to that part. At stage 2 the honest answer is that the plant can describe what should have happened, and would need engineering support and a working day to establish what did.

This stage is stable and it is where most of the risk sits, because it passes casual inspection. The documents are correct, the file is tidy, and the weakness only surfaces under a specific request — which is exactly what a VDA 6.3 process auditor or an escalated customer visit is designed to make. Plants often discover their stage-2 status during a customer complaint rather than during an audit, when a containment question needs a model version and nobody can supply one.

In practice

The validation report with no serial numbers behind it

A powertrain plant documented its leak-test anomaly model properly: control plan revision, PFMEA update, a validation report showing agreement against a labelled reference set, and a reaction plan for out-of-limit behaviour. Then a customer raised a field concern on parts built across a two-week window. The question was simple — which of those parts had been judged by the model, and on which version. Reconstructing it took three people the better part of a week, because the decisions were in the station historian, the versions were in an engineering deployment spreadsheet, and the two shared no key.

What it looks like

  • Control plans name the AI stations, their characteristics and reaction plans
  • PFMEAs carry model failure modes, not just mechanical ones
  • A validation report exists for the model version in service
  • Answering a question about one part still means emailing engineering

Diagnostic signals you can check this week

  • Pick a serial number from six months ago and ask for the model version that judged it — time the answer
  • Check whether the deployment record and the decision log can be joined without human interpretation
  • Ask whether the validation report names the exact model artefact that is running today
  • Look for the reaction plan being exercised in real records, not only described in the control plan

Anti-pattern · Treating the validation report as permanent

A validation report is a statement about one model version, one dataset and one operating envelope. Plants at stage 2 routinely treat it as a property of the station: the report goes in the file, the model is retrained twice over the next year, and the file is never touched. An auditor who compares the artefact hash or the version in the report against the version on the station finds the gap in about ninety seconds, and it converts a documentation strength into a change-control finding.

What holds you here

Decisions, model versions and input artefacts live in three systems with no shared key, so per-part evidence is a manual archaeology exercise.

Highest-leverage next move

Emit one decision record per inspection — part identity, station, timestamp, model version, input reference, score, threshold, outcome — and retain it to the customer's schedule.

Cost of leaving

Effort
3–5 months
Team
Quality engineer, MES or data engineer, ML engineer for versioning
Risk
Medium — the decision log has to become a retained record, which pulls in IT retention and storage
To next stage
3–6 months

If this is you, the next step is

We wire the decision record and version pinning on one station, then hand you the pattern.

Make one station reconstructable

Stage 3

Traceable

24% of operators sit here

Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.

Stage 3 is where the evidence stops being assembled and starts being emitted. Every inspection writes its own record at the moment of decision, and that record carries enough to rebuild the decision later: what part, what station, what model version, what input, what score, what threshold, what outcome, what the operator did next. The quality system holds the link from the control plan's station identifier down to those records, so the walk from document to data takes one step.

The operational payoff shows up long before the audit. Containment gets cheap and precise. When a defect escapes, the question 'which parts were judged by the version that was running between these two dates' has an answer in minutes, and the containment is scoped to that population instead of to the whole build window. Plants routinely find that this alone pays for the work, because over-broad containment is one of the most expensive reflexes in automotive quality.

The constraint that emerges here is a governance one. Traceability makes it possible to answer any question about the past, but it does not by itself tell you whether the process is still inside its stated limits today, or whether anybody would notice if it were not. That is the difference between being able to satisfy an auditor when asked and being able to demonstrate control without being asked — which is what the next stage buys.

In practice

The containment that fitted on one screen

A stamping supplier had a customer complaint on a crack that the vision model should have caught. Because every inspection carried the model version and the retained image, quality pulled the exact population judged by that version in the affected shift pattern — a few hundred parts rather than the eleven days of production the customer had asked to be sorted. The 8D containment section named the version, the window and the population, and the corrective action attached the re-validated version. The customer closed it without a special audit.

What it looks like

  • One decision record per inspection, joinable to the serial number or VIN
  • Model versions are pinned, archived and re-runnable
  • Nonconformances can be scoped by model version and time window
  • Reconstruction is a query, not a project

Diagnostic signals you can check this week

  • Time a reconstruction from a randomly chosen serial number — under an hour is stage 3
  • Check that an archived model version can actually be re-run, not merely named
  • Confirm that nonconformance records reference model version and time window as standard fields
  • Ask whether the retained input artefact is the one the decision was made on, or a later re-capture

Anti-pattern · Retaining everything instead of retaining the right thing

The reflex at stage 3 is to keep every frame from every camera forever, which fails on cost and then quietly fails on completeness when someone trims the retention without changing the schedule. The record that matters is small: identity, version, input reference, score, threshold, outcome. Define a retention class per artefact type, tie it to the customer requirement and the product's service life, and sample-test the oldest retained record every quarter to prove the schedule is real.

What holds you here

The plant can answer any question it is asked, but nothing proves the process stayed inside its limits between audits.

Highest-leverage next move

State a performance floor for each model, put monitoring and an alarm behind it, and rehearse the audit question quarterly on a randomly chosen part.

Cost of leaving

Effort
4–8 months
Team
Quality systems owner, data engineer, ML engineer, IT for retention
Risk
Medium — storage and retention decisions become contractual, and they are hard to reverse
To next stage
6–9 months

If this is you, the next step is

One session to define the record, the retention class and the join to the QMS.

Design your evidence record

Stage 4

Rehearsed

12% of operators sit here

The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.

Stage 4 changes who finds the problem. The plant runs the audit on itself: an internal auditor picks a part at random, asks the questions a certification body would ask, and times how long the answers take. What comes out is a findings list — a control plan that lags a model change by three weeks, a station whose reaction plan has never been exercised, a retention class that expired earlier than the customer requires. Those are cheap findings when you raise them and expensive ones when somebody else does.

The other half of the stage is the performance floor. A model in production without a stated floor — a maximum escape rate on safety characteristics, a minimum agreement against reference parts, a false-reject ceiling the line can absorb — has no definition of 'still working', which means drift is only detectable in hindsight through scrap and warranty. Writing the floor down converts the model from an engineering asset into a controlled process with a reaction plan, which is the language every audit regime already speaks.

What makes this stage durable is that rehearsal has a cadence rather than a trigger. Plants that rehearse only when an audit is booked get good at the audit they expect; plants that rehearse quarterly on a random part get good at the audit they do not. The second group is the one that survives a customer escalation, which arrives without a date in the calendar.

In practice

The quarterly drill that found the silent rollback

During a routine reconstruction drill, a quality engineer pulled a part from the previous quarter and found the model version on the record did not match anything in the change log. The explanation was mundane: a station had been re-imaged after a controller fault and had come back with the previous model version, running for nine days before the next deployment overwrote it. Nothing had failed, no defect had escaped, and no external auditor would have looked. The drill turned it into an internal finding and a permanent change to the station recovery procedure.

What it looks like

  • Every model has a stated performance floor with monitoring and an alarm
  • Internal audit covers AI-influenced stations on a schedule, not on request
  • Reconstruction drills are timed and the time is trending down
  • Findings from rehearsal are tracked and closed like external findings

Diagnostic signals you can check this week

  • Ask for the date and result of the last reconstruction drill, and who chose the part
  • Check that every production model has a written performance floor with a named owner and an alarm
  • Look at whether rehearsal findings are in the same tracker as external findings, with the same closure discipline
  • Confirm that internal audit schedules AI stations across shifts, not only on days when engineering is on site

Anti-pattern · Rehearsing with the part you already understand

The drill only works if the part is chosen adversarially. Teams naturally reach for a part from a station they trust, in a period they remember, which tests recall rather than the system. Have somebody outside the AI team pick — quality, an internal auditor, or a random draw from the shipping record — and require the reconstruction to be produced without calling the person who built the model. That single rule is what converts a demonstration into a test.

What holds you here

Evidence is produced by the effort of a good team rather than by the system, so quality depends on which people are available that week.

Highest-leverage next move

Make the evidence a by-product: emit at decision time, bind it to the QMS record automatically, and give auditors a read-only view instead of a prepared pack.

Cost of leaving

Effort
6–12 months
Team
Quality manager, internal audit, ML engineer, plant IT
Risk
Medium — rehearsal generates real findings, and the closure workload is genuine
To next stage
9–18 months

If this is you, the next step is

We play the auditor for a day and leave you the findings list and the evidence gaps.

Run a mock audit on an AI station

Stage 5

Continuously auditable

4% of operators sit here

The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.

Stage 5 removes the preparation window. Nothing is assembled for an audit because nothing needs to be: the control plan resolves to the station, the station resolves to its decision records, the records resolve to a pinned model artefact and its validation, and the change history is the deployment log rather than a document describing it. The plant's answer to 'show me' is a query somebody runs while the auditor watches.

The distinguishing behaviour is that the quality system initiates. A performance-floor breach opens a nonconformance the way a gauge going out of calibration does, with the same reaction plan discipline. A model deployment cannot complete without its validation record, its MSA refresh and its control-plan revision, because the workflow refuses to proceed. This is unglamorous plumbing and it is the only version of AI governance that survives a busy quarter, because it does not depend on anybody remembering.

Sustaining it is a change-control problem rather than a technical one. Customer-specific requirements move, retention obligations lengthen with service life, new part families arrive that were never in the validation envelope, and each of those quietly invalidates a piece of the arrangement. Plants at stage 5 treat the audit-readiness configuration itself as a controlled document with an owner and a review cycle — otherwise the system that made evidence free slowly stops matching what the evidence is for.

In practice

The audit that was conducted from a read-only screen

At a plant running several AI-assisted inspections, the surveillance audit for those stations was conducted from a filtered view: the auditor picked a serial number, and the record came back with station, timestamp, model version, retained input, score, threshold, outcome and the operator's disposition, with links to the control plan revision and the validation record for that version. The quality manager's preparation for that portion of the audit had been to check that the view still worked. Nothing was assembled, because nothing had been disassembled.

What it looks like

  • Evidence is emitted at decision time and bound to the QMS record automatically
  • Auditors are given a filtered, read-only view rather than a prepared pack
  • Model change, control plan revision and customer notification are one workflow
  • Performance-floor breaches open a nonconformance without human initiation

Diagnostic signals you can check this week

  • Ask what preparation was done for the last audit of an AI station — 'none' is the stage-5 answer
  • Check whether a deployment can complete without its validation and MSA artefacts attached
  • Confirm that a performance-floor breach raises a nonconformance without a person deciding to raise one
  • Ask who owns the readiness configuration itself, and when it was last reviewed against customer requirements

Anti-pattern · Letting the read-only view drift from the customer's requirement

Once evidence is free, attention moves elsewhere, and the configuration ages quietly. A customer lengthens a retention period, a new part family enters the line outside the validated envelope, a CSR adds a notification duty for process change — and the automated pack keeps producing yesterday's correct answer. Put the readiness configuration under document control with a named owner and an annual review against every active customer's requirements, and sample-test the oldest retained record each quarter.

What holds you here

The readiness configuration ages against moving customer-specific requirements, so a correct system slowly stops answering the current question.

Highest-leverage next move

Put the readiness configuration under document control: a named owner, an annual review against every active CSR, and a quarterly test of the oldest retained record.

Cost of leaving

Effort
Continuous
Team
Quality systems owner, platform engineer, standing change-control forum
Risk
Concentrated — low frequency, high consequence, and always tied to a customer requirement change

If this is you, the next step is

We attack the pipeline with real audit scenarios and report what it cannot answer.

Stress-test a continuous evidence pipeline

Where automotive plants sit — and the four dimensions that set the stage

The distribution across the ladder, why Documented is the plateau, and the four dimensions that gate each other.

Most automotive plants are at Documented. Once someone senior asks the question, control plans and FMEAs get updated quickly — that work is well understood and it is measured in weeks. What does not follow automatically is the per-part evidence, because that requires a change to how the station writes records rather than a change to what a document says. The distribution below is illustrative rather than surveyed, and it is drawn to match a pattern any quality manager will recognise: a large documented middle, a thin top, and a tail that has not started.

Illustrative distribution of automotive plants across the readiness ladder

Illustrative, not surveyed. The shape reflects the structural argument on this page: documentation is cheap and follows a single management request, while per-part traceability requires the station to emit records and therefore lags. Read the shape, not the decimals.

Share of plants

  • 22% — 1 · Unprepared
  • 38% — 2 · Documented (the plateau)
  • 24% — 3 · Traceable
  • 12% — 4 · Rehearsed
  • 4% — 5 · Continuously auditable

Source: Illustrative distribution, framed against the NIST AI Risk Management Framework's evidence expectations

The pressure on the plateau is coming from two directions at once. Customers are adding questions about AI-influenced processes to existing audits rather than creating new ones, which means the evidence has to sit inside the quality system your IATF 16949 certification (opens in a new tab) already covers. At the same time the horizontal frame is filling in: ISO/IEC 42001 (opens in a new tab) gives an AI management system something to certify against, the NIST AI Risk Management Framework (opens in a new tab) supplies the vocabulary customer questionnaires increasingly borrow, and the EU AI Act (opens in a new tab) reaches AI embedded in type-approved products on its own timetable. None of these replaces the plant audit. All of them make the plant audit ask more.

  • Evidence completeness

    Whether the artefacts an auditor asks for exist at all: control plan entry, PFMEA rows, agreement study, validation record, change history, competence records. This is the cheapest dimension to fix and the one most often assumed to be someone else's file.

  • Process documentation

    Whether the documents describe the process that is actually running today, including after the last retraining. A correct document that is three model versions out of date is a change-control finding, which is a heavier category than a missing document.

  • Traceability and reconstruction

    Whether the plant can go from one part to the decision that judged it. This is the dimension that separates Documented from Traceable, and it is overwhelmingly the lowest-scoring dimension in readiness reviews, because it is the only one that cannot be fixed by writing something.

  • Audit rehearsal

    Whether the plant tests itself on a cadence rather than before a date. Rehearsal converts unknown gaps into internal findings, which are cheap; the alternative is discovering them in a surveillance audit or a customer escalation, which is not.

Diagnosing the real constraint

Plot your process documentation against your traceability. The quadrant names the next investment — and only one of the four answers is 'write another document'.

Traceable but undeclared

  • The data exists; the quality system does not know the model does
  • Common where AI arrived as an engineering improvement
  • Fix: control plan revision, PFMEA rows, work instruction — weeks, not quarters

Audit-ready

  • Documents match the floor and the floor emits records
  • Constraint moves to rehearsal and change control
  • Fix: quarterly reconstruction drills and a performance floor with an alarm

Unprepared

  • Neither the document nor the record exists
  • Every AI station in the plant is in scope on the same day
  • Fix: inventory first, then the highest-exposure station end to end

Paper compliance

  • The file is tidy and nothing can be proved for a specific part
  • The most dangerous quadrant — it passes a document review
  • Fix: emit the decision record before improving another document
Traceability and reconstruction — top: Any part reconstructable on demand, bottom: Cannot rebuild a past decision
Process documentation — left: Control plan silent on the model, right: Control plan, PFMEA and MSA all cover it

What each audit asks, and the evidence that answers it

First the regimes that actually visit an automotive plant. Then the map that matters: what is asked, what satisfies it, who owns it, and where it lives.

Four kinds of audit reach an automotive plant, and AI changes what each one asks without changing who asks it. The certification body arrives on the IATF cycle and audits the quality management system. A customer sends a VDA 6.3 (opens in a new tab) process auditor and audits the process, element by element, at the station. Customer-specific requirements (opens in a new tab) arrive continuously through supplier quality and add duties the standard does not — notification thresholds for process change, retention periods, run-at-rate conditions. And your own internal audit programme is meant to find all of it first. A fifth, TISAX (opens in a new tab), is not a quality audit at all, but it reaches the AI development environment through the customer data your training sets contain. All of it sits on the ISO 9001 (opens in a new tab) management-system spine, which is why none of these audits needs an AI clause to reach a model: they already reach every process.

AuditWho runs itRhythmWhat it asks once a model is in the loop
IATF 16949 certification auditAn IATF-recognised certification body under the IATF RulesThree-year certification cycleWhether the quality management system covers the AI-influenced process at all: control plan, PFMEA, MSA, competence, change control, records and retention
IATF 16949 surveillance auditThe same certification bodyBetween certification audits, on the scheme's scheduleWhether anything changed since the last visit without the system noticing — new stations, new model versions, uncontrolled documents, drifting performance
VDA 6.3 process auditThe customer, or a qualified VDA 6.3 auditor on their behalfPer programme, per escalation, or on a customer scheduleThe process at the station: is the AI inspection capable, is the reaction plan real, is the operator competent, is the evidence available where the work happens
Customer-specific requirementsCustomer supplier-quality engineersContinuous, plus programme milestonesWhatever that customer's own rules add: notification duties for process change, approval before a model may replace a documented method, retention periods, run-at-rate conditions
PPAP or part submission reviewCustomer programme qualityPer part, per changeWhether the submitted package still describes the process actually running — a model added after approval is a process change until proved otherwise
Internal audit programmeYour own trained internal auditorsPlanned coverage of processes, shifts and clausesEverything above, first. Internal audit is where an AI station should fail, at a cost you control
Layered process auditSupervisors through to plant leadershipDaily to monthly, by layerWhether the AI station is being run today the way the work instruction says, including who may override a reject and what they record when they do
TISAX assessmentAn ENX-accredited audit providerOn the label's validity cycleInformation security around the data the models are built on — customer drawings, prototype images and quality data pull the AI development environment into assessment scope
ISO/IEC 42001 certificationAn accredited certification body, voluntarily engagedOptional today, appearing in customer questionnairesWhether an AI management system exists around the models: policy, impact assessment, lifecycle controls, monitoring and improvement
The audit regimes that visit an automotive plant, and what each one asks once a model influences a decision. The right-hand column is the one to read twice — it is where AI-related findings actually originate.

The map below is the centre of this page. Each row is a question an auditor asks in words a plant recognises, the artefact that satisfies it, the person who has to own that artefact, and the system it has to live in. The column that fails in practice is 'who owns it'. Model documentation drafted by a data team and stored in a code repository is not owned by anyone the auditor will speak to; the moment the process owner cannot produce it without help, the plant is at stage 1 for that station regardless of how good the documentation is.

What the auditor asksWhat satisfies itWho owns itWhere it lives
Is this process defined and controlled?Control plan revision naming the AI station, the characteristic it judges, the method, the frequency and the reaction plan for the model being unavailable or out of envelopeProcess owner (manufacturing quality engineer)Control plan register in the QMS, with the station identifier that also appears in the MES
Has the risk of this process been assessed?PFMEA rows for model-specific failure modes — false accept, false reject, drift, changed lighting or fixture, wrong version deployed — with detection and controls that are realCross-functional FMEA team, chaired by qualityFMEA workbook in the QMS, revision-linked to the control plan
Is the thing making the decision a capable measurement system?Attribute agreement analysis against a reference set of known-good and known-bad parts, with acceptance criteria agreed before the study and recorded against the model versionMetrology or quality lab, with the process ownerMSA study record plus an entry in the gauge and equipment register
Was the part approved with this process?PPAP package whose control plan, MSA and capability evidence describe the process actually running; a documented process change if the model arrived after approvalProgramme quality (APQP lead)Internal PPAP file and the customer's submission portal
What is this AI system for, and what is it not for?Intended-use statement: part families, characteristics, operating envelope, explicitly excluded conditions, and what the model must never be used to decideAI product owner, countersigned by the process ownerModel file in the model register, referenced from the control plan
Where did the training data come from?Provenance record: source stations, date ranges, part numbers, how images or signals were selected, exclusions, and the labelling procedure with labeller qualificationData owner in engineeringDataset register with content hashes, linked to the model version
How do you know it works?Validation report against a held-out labelled sample, broken down by defect class and part family, with acceptance criteria set before the run and the exact model artefact identifiedQuality engineeringValidation record attached to the model version in the model register
What performance must it hold, and who watches?A written performance floor — escape ceiling on safety characteristics, minimum agreement, false-reject ceiling — with monitoring, an alarm and a named responderProcess owner, with engineering on callMonitoring system, with the floor and reaction restated in the control plan
What changed, when, and who approved it?Change record per model version: what changed, the dataset it was trained on, the validation result, the MSA refresh, the approver, the effective timestamp and the stations affectedChange control boardChange record in the QMS, keyed to the MES or station deployment log
Why did this part pass on this date?One decision record per inspection: part identity, station, timestamp, model version, input artefact reference, score, threshold, outcome and the operator's dispositionQuality systems owner (MES and data)Quality data store, retained to the customer's schedule and joinable to the build record
What did you do when it was wrong?Nonconformance and 8D treating the model as part of the process: containment scoped by model version and time window, root cause, corrective action, and updates to the PFMEA and control planQuality managerNonconformance and 8D system, cross-referenced to the model version
Who is competent to run this?Training records for operators, team leaders and quality staff covering the AI station, override authority, escalation, and what to do when the model is unavailableArea manager with training and HRCompetence matrix and training records, auditable per shift
How long do you keep all of it?A retention schedule per artefact type, aligned to the customer requirement and the product's service life, with evidence that the oldest retained record is still retrievableQuality systems ownerRecords retention schedule, tested quarterly against live storage
The clause-to-evidence map for an AI-influenced process in an automotive plant. Ownership sits with a person in the quality chain in every row — that is deliberate, and it is the column most plants get wrong.

Two properties of this map decide how much work it represents. First, nothing in the left-hand column is new. Every one of those questions has been asked of automotive processes for decades; what changes is that the answer now has to include a model version and a retained input. Second, every row has a home in a system the plant already runs — QMS, MES, gauge register, training records — which means audit readiness is an integration exercise rather than a new platform. Plants that respond by buying an AI governance tool disconnected from the QMS end up with a fourteenth system and the same finding.

How to use the map

  1. Fill the ownership column first

    Go down the map and write a name against every row for one station. Rows with no name, or with a name outside the quality chain, are your real gaps — an artefact nobody in the audit room owns is an artefact that will not be produced in the audit room.

  2. Then fill 'where it lives' with a system, not a folder

    A path on a shared drive is a folder; the QMS, the gauge register, the MES and the model register are systems with access control, versioning and retention. Anywhere the answer is a folder, the artefact is one reorganisation away from being unfindable.

  3. Close the row that fails soonest, not the row that is easiest

    Order the remaining gaps by which audit reaches them first. A missing control plan entry is found by any auditor who walks the line; a retention gap is found only when someone asks for an old part. Sequence accordingly, and record the reasoning — an auditor will accept a prioritised, dated plan far more readily than a silent gap.

The quality artefacts AI now touches

APQP, PPAP, the control plan, the FMEA and the MSA were all written for processes with human or mechanical decision-makers. Here is what changes in each when a model joins.

Every core quality artefact in an automotive plant changes when a model influences a decision, and none of them needs to be replaced. The AIAG core tools (opens in a new tab) — advanced product quality planning, the production part approval process (opens in a new tab), the AIAG and VDA FMEA method (opens in a new tab), measurement systems analysis (opens in a new tab) and statistical process control — already carry the concepts a model needs: a defined method, a characterised measurement system, an assessed failure mode, an approved part submission. What changes is what has to be written into each one, and what an auditor will find if it is not.

ArtefactWhat it says without AIWhat must change with a model in the loopThe finding when it does not
APQP phase gatesProcess design is signed off before build, with the measurement method agreed at the gateThe model is a process design element: its intended use, data plan, validation plan and acceptance criteria belong at the gate, not after launchA model introduced post-launch with no design record, so the launch package describes a process that no longer exists
PFMEAFailure modes of the equipment and the operator, with detection controls and rankingsModel failure modes as first-class rows: false accept, false reject, drift, changed lighting or fixture, wrong version deployed, model unavailableA PFMEA that ranks detection highly on the strength of an inspection whose own failure modes were never assessed
Control planCharacteristic, method, sample size, frequency, reaction planThe method becomes the model and its threshold; frequency often becomes every part; the reaction plan needs branches for out-of-envelope and unavailableA control plan describing an ultrasonic or manual check that stopped being performed months ago — documented process versus running process
MSAGR&R on the gauge, appraiser variation, bias and linearity where applicableAttribute agreement analysis on the model against reference parts, re-run per version, with acceptance criteria recorded before the studyAn accept or reject decision made by an uncharacterised measurement system — usually the single most damaging AI finding available
SPC and capabilityControl charts on the characteristic, capability indices from measured valuesCharting the model's own behaviour — score distributions, reject rate by shift and part family — alongside the characteristic it judgesDrift discovered through scrap and warranty months later, with no chart that would have shown it earlier
PPAP packageDesign records, control plan, MSA, capability, part submission warrantThe package must describe the process that is running; adding a model to an approved process is a process change until the customer's rules say otherwiseA submission whose control plan and MSA no longer match the floor — a document-integrity problem across every part family on that line
Work instructionsHow the operator performs and records the checkWhat the screen shows, what the operator may override, what they must record when they do, and what happens when the model is offlineOverrides happening with no authority defined and no reason captured, so the record cannot explain its own exceptions
Reaction planWhat to do when the characteristic is out of limitsAdditional branches: model out of envelope, monitoring alarm, version mismatch at the station, retained-input capture failingA station that keeps running and keeps deciding when its own preconditions have failed
Gauge and equipment registerCalibration status, interval, responsible personThe model as a registered decision-making instrument: version, agreement study date, next refresh, ownerNo equivalent of a calibration interval, so nothing triggers a re-study when the world changes around the model
Competence matrixWho is trained and authorised on the processWho may run, override, escalate and disable the AI station, per shift, with dated evidenceAn auditor finding an authorised override performed by someone with no record of training on the station
Records retentionHow long quality records are kept, by typeRetention classes for decision records, retained inputs and model artefacts, tied to the customer requirement and service lifeThe evidence existing at the time of the decision and not at the time of the question
What changes in each quality artefact once a model influences the decision, and the finding that follows when it does not. Nothing here is a new document — it is the existing document carrying a new kind of decision-maker.

The measurement-systems row is the one most worth dwelling on, because it is the least intuitive and the most consequential. A vision model that outputs pass or fail is an attribute gauge, and the established way to characterise an attribute gauge is an agreement study: a set of reference parts whose true state is known and agreed, judged repeatedly, with agreement measured against the reference and between repeats. That gives you the same shape of evidence a GR&R gives for a variable gauge — a number, an acceptance criterion, and a date. Nothing about the model being a neural network changes the applicability of the method; what changes is that the study has to be repeated on every version, because a retrained model is a different gauge.

  • The reference set has to be curated, not sampled

    A reference set drawn at random from production will contain almost no defects, which makes agreement look excellent and means nothing. Curate it: known-good and known-bad parts across the defect classes the control plan cares about, with the true state agreed by more than one qualified person and recorded. Keep the physical parts where you can, and the retained inputs where you cannot.

  • Agreement must be reported by defect class, not in aggregate

    An overall agreement figure hides the class that matters. A model with excellent aggregate agreement can be systematically poor on the one defect type that maps to a safety characteristic, and that is precisely the breakdown an auditor asks for when the characteristic is customer-designated. Report the matrix, not the headline.

  • The study belongs to the version, not the station

    Attach the agreement study to the model version in the register, the same way a calibration certificate attaches to a specific gauge. When the version changes, the study's status becomes 'due', and the deployment workflow should refuse to complete without a current one. This is the mechanism that keeps documentation from ageing invisibly.

  • Re-study when the world changes, not only when the model does

    New lighting, a new fixture, a new supplier's surface finish, a new part family: each of these moves the input distribution without touching the model. Define the triggers in the control plan alongside the interval, and treat an unplanned trigger the way you would treat a gauge dropped on the floor.

  • Keep the acceptance criterion out of the analyst's hands

    Agree the acceptance criterion before the study runs and record it with the plan. A criterion chosen after the result is a finding waiting to be made, and it is trivially detectable — the auditor simply asks when the number was decided and by whom.

In order to make the process audit- and certification-proof, development at the Neckarsulm location was carried out in close coordination with the German Association for Quality (DGQ), the Fraunhofer Institute for Industrial Engineering (IAO), and the Fraunhofer Institute for Manufacturing Engineering and Automation (IPA).

That sentence is worth reading twice, because of what it implies about sequencing. Audi did not build an inspection model and then ask how to certify it; the audit and certification question was carried alongside the development, with the German Association for Quality (opens in a new tab) and two Fraunhofer institutes (opens in a new tab) involved in how the process would be evidenced. The same release notes that there are no independent certifications for AI applications of this kind, which is exactly why the evidence has to be built into the quality system the plant already has certified rather than deferred to a scheme that does not yet exist.

What audit-grade AI looks like in public

Three publicly reported programmes, read against the readiness ladder. None is an Atomic Loops engagement — each links to the manufacturer's own published material.

The clearest public evidence for this page's argument is in what large manufacturers said about their own AI quality programmes — not the accuracy claims, but the surrounding sentences about how the process would be evidenced, who was involved, and how it would scale across sites. In each case below the differentiator is structural: what the model was allowed to decide, what recorded the decision, and what had to be true before the method could replace a documented one.

Three programmes read against the ladder

Outcomes as reported by the manufacturers themselves. The card images are illustrative industry scenes from our media library, not photographs of the named manufacturers' facilities, and no endorsement is implied. Verify figures against the linked source before reusing them; we have not independently audited them.

Illustrative scene: a robotic sensor head scanning an automotive body structure while quality engineers watch analysis screensAudiPremium OEM · Neckarsulm body shop, rolled out across Group sites24
Challenge
Resistance spot welding quality had been monitored by ultrasonic checks on a random sample — roughly 5,000 spot welds per vehicle inspected by production staff — a documented sampling method that could never reach the whole population.
Approach
Audi moved the assessment to AI analysis of the welding process data, and reports developing it in close coordination with the German Association for Quality (DGQ) and the Fraunhofer IAO and IPA institutes specifically so the process would be, in Audi's words, audit- and certification-proof.
Reported outcome
Audi reports analysing around 1.5 million spot welds on 300 vehicles each shift at Neckarsulm, with employees redirected to investigating flagged anomalies, and the technical infrastructure being installed at three further Volkswagen Group locations.
What it shows about the curveReplacing a documented sampling method with a model is a process change before it is an improvement. Carrying the audit question alongside the development — with an external quality body in the room — is what lets the change land in the control plan rather than in a finding.

Audi — AI for quality control of spot welds (opens in a new tab)

Illustrative scene: quality specialists reviewing a vehicle-specific digital inspection overlay beside a body on the lineBMW GroupGlobal OEM · Plant Regensburg, around 1,400 vehicles a day34
Challenge
Final inspection at a plant building roughly 1,400 vehicles a day — one every 57 seconds — has to cover an enormous configuration space, where a single generic inspection catalogue is either too long for the takt or too short for the variant in front of the inspector.
Approach
BMW Group reports an AI system, developed at Regensburg with a start-up partner, that analyses vehicle configuration and live production data to generate a vehicle-specific inspection scope, ordered and delivered to trained specialists through a smartphone app with standardised coding for findings.
Reported outcome
BMW Group publicly describes the system generating customised inspection specifications per vehicle and organising the inspection intelligently, with findings recorded through standardised coding — including optional voice capture with transcription.
What it shows about the curveWhen a model decides what gets inspected, the inspection scope itself becomes a controlled characteristic. Standardised coding of findings is the quiet part that matters for readiness: it is what makes the resulting records comparable, queryable and reconstructable later.

BMW Group — artificial intelligence as a quality booster (opens in a new tab)

Illustrative scene: production engineers reviewing a networked vehicle inspection dashboard on the assembly floorVolkswagen GroupMulti-brand group · computer vision scaled across plants23
Challenge
Computer-vision quality applications were being developed plant by plant — label verification at one site, press-shop crack detection at another — with each deployment carrying its own evidence practices into a separately certified quality system.
Approach
Volkswagen Group reports developing computer-vision applications with a dedicated expert team and rolling them out across the Group through its industrial cloud platform, so solutions built at one location become available to others rather than being rebuilt.
Reported outcome
The Group describes applications including label content and placement verification at Porsche Leipzig and machine-learning detection of fine cracks and defects in press-shop components at Audi Ingolstadt, with Group-wide rollout via the shared platform.
What it shows about the curveScaling a model across sites scales the audit obligation with it. Every receiving plant is separately certified, so the evidence pattern — control plan wording, agreement study, decision record — has to travel with the model, or the tenth deployment starts its documentation from zero.

Volkswagen Group — computer vision in production (opens in a new tab)

Read together, the three make one argument. Audi shows the sequencing: the audit question travels with the development, not after it. BMW shows the scope question: once a model decides what to inspect, the inspection plan is itself a controlled output and the coding of findings determines whether anything can be reconstructed. Volkswagen shows the multiplication problem: a model that scales across plants multiplies the number of separately certified quality systems that must each carry the same evidence. None of these is an argument about model quality, and that is the point.

The model documentation an auditor will accept

Five elements, each answering a question the plant already knows how to ask — and the plausible substitute that fails for each one.

An auditor will accept model documentation that reads like process documentation: specific, dated, owned, and tied to the exact artefact running on the station. Five elements carry almost all of it — intended use, training-data provenance, validation evidence, a performance floor, and change history — and each one exists to answer a question a quality auditor has been asking about processes for decades. What makes the difference is not length. A two-page model file with a version, a signature and a hash is worth more than a forty-page technical report describing an architecture nobody can map to the station.

  • Intended use — the scope statement

    What the model decides, on which part families and characteristics, inside what operating envelope, and what it must never be used to decide. The exclusions carry the weight: a model validated on one supplier's surface finish and silently applied to another's is the most common way a good model produces a bad decision, and the intended-use statement is the artefact that makes that misuse visible.

  • Training-data provenance — where the evidence came from

    Source stations, date ranges, part numbers, selection and exclusion criteria, and the labelling procedure with the qualification of whoever applied it. Automotive auditors are unusually comfortable here, because it is the same question they ask about any reference standard: who says this is correct, and on what authority. Record the answer with hashes so the dataset behind a version cannot silently become a different dataset.

  • Validation evidence — the proof, broken down

    Results against a held-out labelled sample, reported by defect class and part family rather than in aggregate, with the acceptance criteria agreed and recorded before the run. Name the exact model artefact the results belong to. A validation report that does not identify its own artefact cannot be matched to the station, which turns the plant's strongest document into an unusable one.

  • Performance floor — the definition of still working

    A written limit the model must hold: an escape ceiling on safety characteristics, a minimum agreement against reference parts, a false-reject ceiling the line can absorb. Without it, 'the model is working' is an opinion, drift is only visible in scrap and warranty, and there is nothing for a reaction plan to react to.

  • Change history — what changed, when, approved by whom

    One record per version: what changed, the dataset, the validation result, the agreement study, the approver, the effective timestamp and the stations affected. This is the element auditors probe hardest, because it is the one that reveals whether the other four are current or historical.

ElementEvidence that satisfiesThe substitute that failsReview trigger
Intended useA signed scope statement naming part families, characteristics, envelope and explicit exclusions, referenced from the control planA project brief or a slide describing the use case in marketing termsNew part family, new supplier material, new line speed, any scope extension request
Data provenanceDataset register entry with sources, date ranges, selection rules, exclusions, labelling procedure and content hashesA folder of images with a filename convention and no record of who labelled themEvery retraining; any change to the labelling procedure or the people applying it
ValidationHeld-out results by defect class and part family, acceptance criteria dated before the run, artefact identified by version and hashAn aggregate accuracy figure quoted in an email or a project closure reportEvery version; any change to thresholds, preprocessing or camera configuration
Measurement-system studyAttribute agreement analysis against curated reference parts, recorded in the gauge register with acceptance criteriaThe validation report re-used as though it were an MSA — a different question with a different sampleEvery version, and on the interval recorded in the control plan
Performance floorA written limit per characteristic with monitoring, alarm threshold, named responder and a reaction plan branchA dashboard someone checks, with no stated limit and no alarmQuarterly, and after any nonconformance attributed to the model
Monitoring and reactionAlarm routing to a named role per shift, with evidence the route has been exercisedAlerts landing in a shared mailbox or a chat channel nobody ownsShift-pattern changes, on-call changes, and after every missed alarm
Change historyA change record per version in the QMS, keyed to the deployment log, with approver and effective timestampA commit history in a code repository, which records the code change and not the process changeEvery deployment, including rollbacks and station re-images
Human oversightDefined override authority, the reason captured as structured data, and competence records per shiftAn informal understanding that operators can 'always call it' with no record of when they didAny change to authority, staffing model or work instruction
Retirement and rollbackA documented path to revert to the previous method, tested, with the control plan branch that covers running without the modelAn assumption that the old method could be restored if needed, never exercisedAnnually, and before any major model change
Model documentation: the evidence that satisfies each element, the plausible substitute that fails, and what should trigger a review. The middle column is where most plants actually are.

It is worth being clear about how the horizontal AI frameworks relate to this pack. ISO/IEC 42001 (opens in a new tab) describes a management system for AI, and its lifecycle controls map cleanly onto the elements above — it is a useful scaffold, and increasingly a line in customer questionnaires, but it is not what the plant auditor opens at the station. The NIST AI Risk Management Framework (opens in a new tab) is the most practical vocabulary for the governance conversation. The EU AI Act (opens in a new tab) matters most where AI ends up as a safety component of a type-approved product rather than as a plant inspection aid. Build the pack for the plant audit first: it satisfies the immediate obligation and it is the substrate every one of those frameworks then asks you to organise.

Reconstruction: why this part passed on this date

The evidence architecture behind a single answer, the minimum record that makes it possible, and the drill that proves it works.

Reconstruction is the ability to take one part identity and return the decision that judged it, with everything needed to explain that decision: the station, the timestamp and shift, the model version, the input the decision was made on, the score, the threshold applied, the outcome and what the operator did next. It is the single hardest requirement on this page and the one that cannot be satisfied by writing a document, because the record either was written at the time or it was not. Everything else on this page can be added retrospectively; this cannot.

The evidence architecture, layer by layer

Each layer is annotated with the readiness stage that first requires it. A plant attempting Traceable without the evidence-capture layer is running a Documented process with extra software.

  1. Station and sensing

    Stage 1+

    • Sensor rigCamera, lighting and fixture state recorded, not assumed
    • Station controllerPLC and HMI events with synchronised timestamps
    • Part identitySerial or VIN read at the station, never inferred from time
  2. Decision layer

    Stage 2+

    • Served modelOne pinned artefact per deployment, hash recorded
    • Threshold and policyThe accept or reject rule, versioned separately from the model
    • Operator overrideThe override, the authority and the reason, captured as data
  3. Evidence capture

    Stage 3+

    • Decision recordOne row per inspection, written at decision time
    • Input artefact storeThe image or signal judged, retained by class and schedule
    • Model fileIntended use, provenance, validation, approvals, change history
  4. Quality-system binding

    Stage 3+

    • Control plan linkStation identifier in the control plan resolves to the records
    • Nonconformance linkNonconformance and 8D records carry model version and window
    • Change controlThe deployment log is the change record, not a copy of it
  5. Retention and audit access

    Stage 4+

    • Retention schedulePer artefact class, matched to customer and service life
    • Reconstruction servicePart identity in, evidence pack out, timed
    • Auditor viewRead-only, scoped, observed — no prepared pack

Pipeline described

  1. Station and sensing (stage 1+) — Sensor rig: Camera, lighting and fixture state recorded, not assumed; Station controller: PLC and HMI events with synchronised timestamps; Part identity: Serial or VIN read at the station, never inferred from time
  2. Decision layer (stage 2+) — Served model: One pinned artefact per deployment, hash recorded; Threshold and policy: The accept or reject rule, versioned separately from the model; Operator override: The override, the authority and the reason, captured as data
  3. Evidence capture (stage 3+) — Decision record: One row per inspection, written at decision time; Input artefact store: The image or signal judged, retained by class and schedule; Model file: Intended use, provenance, validation, approvals, change history
  4. Quality-system binding (stage 3+) — Control plan link: Station identifier in the control plan resolves to the records; Nonconformance link: Nonconformance and 8D records carry model version and window; Change control: The deployment log is the change record, not a copy of it
  5. Retention and audit access (stage 4+) — Retention schedule: Per artefact class, matched to customer and service life; Reconstruction service: Part identity in, evidence pack out, timed; Auditor view: Read-only, scoped, observed — no prepared pack
Step-by-step insights
Station and sensing — record the conditions, not only the result
The most under-recorded facts at an AI station are the physical ones: which lighting programme was active, whether the fixture had been adjusted, whether a lens had been cleaned that shift. These are exactly the variables that move an input distribution without touching a line of code, and when an escape is investigated six months later they are the difference between a root cause and a shrug. Capture fixture and lighting state alongside the decision; it is a handful of fields and it is the cheapest insurance on the station.
Decision layer — version the threshold separately from the model
Thresholds get tuned far more often than models get retrained, frequently by someone trying to reduce false rejects during a difficult shift. If the threshold is not versioned separately and recorded per decision, the record cannot explain why two identical scores produced different outcomes on different days. Keep the accept or reject rule as its own controlled artefact, with its own change record, and write the applied threshold into every decision row.
Evidence capture — write at decision time or not at all
Records assembled later are reconstructions of a reconstruction. Writing the record at the moment of the decision costs almost nothing at the station and eliminates an entire class of dispute, because the record's existence is itself evidence about the process. This is also what makes retention meaningful: a schedule can only govern records that were created, and a schedule over records that are generated on demand is not a schedule at all.
Quality-system binding — the join that makes the audit short
The binding layer is what turns three systems into one answer: the control plan's station identifier must be the same identifier the decision records carry, and the nonconformance system must have model version and time window as real fields rather than free text in a description. Plants that skip this end up with excellent data and an auditor who watches a person copy values between screens, which reads as fragility even when the data is complete.
Retention and access — prove the oldest record, not the newest
Every plant can retrieve last week. The test that matters is the oldest record the schedule claims to hold: pull it quarterly, confirm the input artefact is still readable, and confirm the archived model version still runs. Storage migrations, format changes and quiet cost-saving trims are the usual causes of a schedule that exists on paper and not on disk, and they are only ever discovered by someone deliberately looking.
FieldExampleWhy it is asked forIf it is missing
Part identityVIN or serial, read at the stationThe audit is always about a specific partNothing can be joined to the build record; every answer becomes an inference from timestamps
Station and fixtureBS-042, fixture rev CLocates the decision in the documented processThe control plan cannot be tied to the record, so document and data never meet
Timestamp and shift2026-02-11 03:14, night shiftTies the decision to staffing, conditions and the change logChange windows cannot be correlated; containment cannot be bounded by time
Model identifier and versionsurface-defect v7.2.1Establishes which gauge made the callThe validation and agreement evidence cannot be matched to the decision
Model artefact hashsha256:9f4c…Proves the version label refers to the artefact you still holdA version number that cannot be verified is a claim rather than a record
Input referenceObject key for the retained image or signal windowLets the decision be re-run and reviewed by a personThe decision cannot be re-examined, only described
Score or confidence0.94Shows how close the decision was to the boundaryMarginal decisions cannot be distinguished from clear ones in an investigation
Threshold applied0.90, policy v3Explains why that score produced that outcomeTwo identical scores with different outcomes cannot be explained
OutcomePass, or reject with classThe decision itself, in the control plan's vocabularyRecords cannot be aggregated into a control chart or a capability view
Operator action and reasonOverride to pass, reason code 04Distinguishes the model's decision from the plant's decisionOverrides become invisible, and the model gets credited or blamed for human calls
Downstream dispositionReworked, scrapped, releasedCloses the loop to what actually happened to the partThe nonconformance trail stops at the station and never reaches the part
Retention classSafety characteristic — 15 yearsGoverns how long each artefact must surviveRetention is applied uniformly, which is either too expensive or too short
The minimum record for one AI-influenced decision. Every field earns its place by answering something an auditor or a warranty investigation will ask; the right-hand column is what happens without it.

The reconstruction drill

  1. Have somebody else pick the part

    Quality, internal audit or a random draw from the shipping record — never the AI team, and never a part anybody in the room remembers. The point is to test the system, not the recall of the person who built it.

  2. Start the clock and use only the systems

    No phone calls to the model author. If a step requires a person outside the quality chain, that is a finding, and it should be written down as one at the moment it happens.

  3. Resolve the station and the control plan revision in force

    Confirm the control plan revision that governed that date names the AI inspection and its reaction plan. A revision that post-dates the part is itself informative — it means the process was running undocumented at the time.

  4. Pull the decision record and the retained input

    Check that the input artefact still opens, and that its reference resolves without a manual search. A broken reference discovered in a drill costs nothing; the same break discovered during a warranty investigation costs a great deal.

  5. Re-run the archived model version on the retained input

    The score should reproduce. If it does not, you have discovered that either the artefact, the preprocessing or the threshold configuration was not fully pinned — the most valuable finding a drill can produce.

  6. Assemble the pack, stop the clock, log the gaps

    Control plan revision, PFMEA rows, agreement study for that version, validation record, change record, decision record, retained input, operator disposition. Record the elapsed time and raise every gap as an internal finding with an owner and a date.

When the model was wrong

Models are wrong sometimes; that is not the finding. The finding is a nonconformance handled as a station fault, with no link back to the process the model is part of.

A model being wrong is a nonconformance in the process, and it should travel the route every other process nonconformance travels: containment, root cause, corrective action, and an update to the documents that claimed the failure mode was controlled. What makes AI nonconformances distinctive is the containment step. A conventional process fault is bounded by time and station; a model fault is bounded by model version, threshold policy and input conditions — which is why the decision record has to carry all three, and why a plant that cannot bound the population ends up sorting the whole build window.

FailureHow it shows upContainment scoped byRecords it must produce
Escape — a defect passedCustomer complaint, end-of-line audit, or a warranty pattern months laterModel version, threshold policy and the date window they were in force, filtered to affected part familiesNonconformance and 8D naming the version; the population list; re-validation of the corrected version; PFMEA and control plan updates
Overkill — good parts rejectedReject rate steps up on one shift or one part family; rework queues buildVersion and threshold window, plus the input condition that changed — lighting, fixture, materialReject-rate evidence by class; the input-condition finding; a documented threshold or model change through change control
Silent driftNothing at all, until a performance-floor alarm or a scrap trendThe interval since the last agreement study, because the population is everything judged in itThe floor breach record; a fresh agreement study; a decision on whether the interval or the trigger set was wrong
Wrong version deployedStation re-image, failed rollout, or a rollback nobody recordedThe exact window between the unintended deployment and its correction, from the deployment logChange-control finding; the corrected deployment record; a procedure change for station recovery
Out-of-envelope inputA new part family, new supplier finish or new line speed enters without a scope reviewThe part families outside the intended-use statement and the dates they ranScope-extension record; validation on the new family; updated intended-use statement and control plan
The ways a model is wrong, how each shows up, what bounds the containment, and the records each must produce. Escapes and overkill are the two obvious cases; the other three are the ones that surprise plants.
Likelihood: highImpact: high

The nonconformance is written against the station, not the process

A reject that should have been a pass gets logged as an equipment fault, closed by a controls engineer, and never reaches the model version. The 8D reads correctly and contains nothing that would let anyone find the same failure again, because the field that identifies the actual decision-maker was never populated.

PreventionMake model version and threshold policy mandatory fields on any nonconformance raised at an AI-influenced station.

Likelihood: highImpact: high

Containment is scoped by build date because version data is missing

Without a version window the only defensible boundary is time, so containment expands to the whole build period. The cost lands in sorting, expedited freight and customer confidence — and the same event repeats the next time, because nothing about the record has changed.

PreventionRetain model version and threshold on every decision record; rehearse a version-bounded containment query quarterly.

Likelihood: mediumImpact: high

The corrective action retrains the model and stops there

The new version fixes the observed case, and the PFMEA that ranked detection highly on the old inspection is left unchanged, the control plan still describes the previous threshold, and the agreement study on file belongs to a version no longer running. The plant has corrected the defect and left the documentation defect in place.

PreventionMake PFMEA and control-plan review a required closure step for any 8D involving a model, with the same sign-off as a tooling change.

Likelihood: mediumImpact: high

Overrides absorb the problem until nobody can see it

When a model becomes noisy, operators override to keep the line moving. Without a captured reason code the overrides look like normal operation, the reject-rate signal disappears, and the underlying drift runs for months. The eventual escape is then investigated against a period whose records show a well-behaved model.

PreventionCapture override reason codes as structured data and chart override rate alongside reject rate on the same review.

Likelihood: lowImpact: high

The performance floor is loosened instead of the cause being found

A breach is resolved by moving the limit, usually informally, usually to avoid stopping the line. Nothing in the record distinguishes this from a genuine re-baselining, and the floor slowly ceases to mean anything — at which point the monitoring around it is decoration.

PreventionTreat the floor as a controlled document: changes need the same approval as a control-plan revision, with the rationale recorded.

There is a cultural point underneath the mechanics. Plants that hide model errors accumulate exactly the evidence an auditor is most suspicious of: an inspection process with no recorded failures. A model that has never been wrong in the record either has not been running long, or is not being recorded. Well-run AI stations produce a visible, boring stream of small nonconformances with tidy containment and closed corrective actions — and that stream is the single most persuasive artefact a plant can put in front of a process auditor.

A 90-day plan: one AI inspection station, audit-ready

The Documented to Traceable transition on one concrete problem — a body-shop surface-inspection station that was never written into the control plan — ahead of the next surveillance audit. Contains no model development.

Bringing one station to audit-ready takes about 90 days, and bringing a plant takes years — which is why the plan below is scoped to one station and one decision. The problem it solves is the most common one in this domain: an AI surface-inspection station on the body-in-white line that was installed as an engineering improvement, works well, has quietly replaced a documented manual check, and appears nowhere in the control plan. The quarter contains no model development at all. Every hour goes into documentation, characterisation, records and rehearsal.

Documented to Traceable on one inspection station, in one quarter

One station, one owner, one part family. If a phase needs longer than its window, narrow the scope — fewer characteristics, one shift pattern — rather than extending the plan.

  1. Days 1–15

    Inventory and pick the station an auditor would open first

    Walk the plant and list every decision a model influences: station, characteristic, part families, who owns it, which model version is running and since when. Rank by audit exposure — customer-designated or safety characteristics first, then how recently the model changed. Pick one station, and name the manufacturing quality engineer who owns it as the single accountable person.

    A one-page AI decision register and one named station with one named owner

  2. Days 16–45

    Write the model into the quality system

    Revise the control plan so it names the AI inspection, the characteristic it judges, the method, the frequency and a reaction plan with branches for out-of-envelope and unavailable. Add model failure modes to the PFMEA. Rewrite the work instruction to cover the screen, the override authority and the reason codes. Update the competence matrix and train the shifts. Route it all through normal document control.

    Control plan, PFMEA, work instruction and training records that match the floor

  3. Days 46–70

    Characterise the model and emit the record

    Curate a reference set of known-good and known-bad parts across the defect classes in the control plan, agree acceptance criteria in writing, and run an attribute agreement analysis against the version in service; record it in the gauge register. In parallel, make the station write one decision record per inspection — part identity, station, timestamp, version, input reference, score, threshold, outcome, operator action — and set the retention class from the customer requirement.

    An agreement study against the running version, and decision records accumulating with a retention class

  4. Days 71–90

    Set the floor, then rehearse the audit

    Write the performance floor, wire the alarm to a named responder per shift, and exercise it once deliberately. Then run the reconstruction drill: have quality pick a part from at least six months back, time the assembly of the full pack without calling the model author, and raise every gap as an internal finding with an owner and a date.

    A timed reconstruction, a monitored floor, and a findings list you raised yourself

The order matters

  1. Documents before data, but only just

    The control plan revision is the cheapest and highest-value artefact, and it can be done in days — so do it first. But do not spend the whole quarter on documents: a plant with perfect paperwork and no decision record is in the paper-compliance quadrant, which passes a document review and fails a process audit.

  2. Agreement study before performance floor

    The floor should be set from measured agreement on reference parts, not from a target somebody would like to hit. Running the study first means the floor is defensible when an auditor asks where the number came from, which is the second question after 'what is your floor'.

  3. One station before one plant

    Resist rolling the pattern out until the first station has survived a drill. The drill is what reveals the parts of the pattern that do not work — a reference that will not resolve, a version that will not re-run, an owner who is not in the quality chain — and fixing those once is far cheaper than fixing them in twelve places.

  4. Raise your own findings, in your own tracker

    Everything the drill exposes goes into the same system your external findings go into, with an owner and a date. An auditor who sees a plant finding and closing its own AI-related gaps reads a functioning quality system; the same gaps with no record read as an unmanaged process.

Audit rehearsal as a standing practice

Readiness decays between audits. A drill calendar keeps it honest — and turns the findings you would have received into findings you raised yourself.

Audit rehearsal is the practice of asking yourself the audit's questions on a schedule, on parts and stations you did not choose, and treating the answers as findings. It works because readiness is not a state a plant reaches but a condition that decays: models are retrained, thresholds are tuned, stations are re-imaged, part families arrive, retention schedules quietly change, and people who knew things leave. Every one of those events moves the plant back down the ladder without anybody deciding to move it.

DrillFrequencyWho runs itPass condition
Reconstruction drillQuarterlyInternal audit picks the part; quality assemblesFull evidence pack for a part at least six months old, assembled in under 30 minutes without calling the model author
Control-plan walkQuarterly, rotating stationsManufacturing quality engineer with the area supervisorThe control plan, the work instruction and what the operator actually does agree on method, override authority and reaction plan
Agreement study refreshOn the interval in the control plan, and on every model versionMetrology or quality labA current study against curated reference parts, reported by defect class, within the recorded acceptance criteria
Change-history auditQuarterlyChange control boardEvery version that ran in the period has a change record, a validation result and an approver; no unexplained versions on any station
Escape simulationTwice a yearQuality managerA version-bounded containment population produced in under an hour from a hypothetical escape date
Floor-breach exerciseTwice a yearProcess owner with on-call engineeringThe alarm reaches the named responder on the shift it is triggered, and a nonconformance is opened by the route it should take
Retention testQuarterlyQuality systems ownerThe oldest record the schedule claims is retrievable, its input artefact opens, and its model version still runs
A rehearsal calendar for AI-influenced processes. Each drill has an owner and a pass condition, so the outcome is a finding or a pass rather than an impression.

Would your AI-influenced processes survive an audit next week?

Tick honestly — this is a diagnostic, not a scorecard, and the blank boxes are the useful part. It works with JavaScript disabled.

0 of 8 ticked

Tick honestly — the blank list is data too

Almost no plant genuinely ticks zero; most can claim one or two once they look. If none apply yet, do not start with tooling. Start with the register: one page listing every decision a model influences, with an owner beside each. It usually takes a fortnight and it changes the conversation, because you cannot size a gap you have not enumerated.

One last framing, because it decides how this work gets funded. Audit readiness is usually proposed as risk reduction, which competes badly against capacity and cost projects. It funds better as containment economics: the version-bounded population, produced in an hour, that turns an eleven-day sort into a few hundred parts. Plants that have been through one such event never need the argument again — and the plants that have not are the ones for whom this page is most urgent, because they will meet the argument at the worst possible moment.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Control plan
The document binding each characteristic to a method, sample size, frequency and reaction plan. When a model takes over a characteristic, all of those change, and the control plan is the first document an auditor opens at the station.
Attribute agreement analysis
The measurement-systems study appropriate to a pass or fail decision: repeated judgements of curated reference parts, with agreement measured against the known reference and between repeats. The equivalent of a GR&R for an attribute gauge, and the correct evidence for an AI inspection.
Reference set
The curated collection of known-good and known-bad parts, covering the defect classes the control plan cares about, whose true state is agreed and recorded. Used for agreement studies and re-validation; deliberately not a random production sample.
Intended-use statement
The scope record for a model: which part families and characteristics it may judge, the operating envelope, and the conditions explicitly excluded. The artefact that makes silent scope creep visible before it becomes a nonconformance.
Model file
The plant-facing documentation set for one model: intended use, data provenance, validation evidence, performance floor and change history, owned in the quality chain and referenced from the control plan rather than stored only in a code repository.
Performance floor
The written limit a model must hold in production — an escape ceiling on safety characteristics, a minimum agreement against reference parts, a false-reject ceiling — with monitoring, an alarm and a named responder behind it.
Decision record
The row written at the moment of an inspection carrying part identity, station, timestamp, model version, input reference, score, threshold, outcome and operator disposition. The artefact that makes per-part reconstruction possible.
Reconstruction time
The elapsed time to assemble the full evidence pack for one randomly chosen historical part, without help from whoever built the model. The most honest single measure of readiness, because it cannot be improved by writing another document.
Version pinning
Retaining the exact model artefact, preprocessing and threshold configuration together so an archived version can be re-run against a retained input. The difference between naming a version in a record and being able to demonstrate it.
Escape
A defect the inspection should have caught and passed. In an AI-influenced process the containment population is bounded by model version, threshold policy and the date window they were in force — not by the whole build period.
Overkill
Good parts rejected by the inspection. Usually the first visible symptom of a changed input condition — new lighting, fixture or material — and the failure mode that operator overrides most often mask.
Layered process audit
The routine, multi-level check that the process is being run today as documented, conducted by supervisors up through plant leadership. For AI stations it is where override behaviour and screen-versus-instruction mismatches surface first.

Frequently asked questions

The questions plant quality teams ask most often when an auditor is due and a model is in the process.

Does IATF 16949 say anything specific about artificial intelligence?

Not by name, and that is the point. The standard is written about processes, measurement systems, change control and records, and a model that influences an accept or reject decision falls inside all of those clauses without needing an AI-specific one. In practice the AI-related findings that plants receive are ordinary categories: documented process differing from the running process, an uncharacterised measurement system, an uncontrolled change, or records that cannot be produced for a specific part. Prepare against those categories rather than waiting for an AI clause.

Do we have to run an MSA on a vision model?

If it makes or influences the accept or reject decision, treat it as a measurement system and produce the equivalent evidence. For a pass or fail output that means an attribute agreement analysis against a curated reference set of known-good and known-bad parts, with acceptance criteria agreed before the study and results reported by defect class rather than in aggregate. The study belongs to the model version, not to the station, so it becomes due again on every retraining, threshold change or significant change to lighting, fixture or material.

Is adding an AI inspection to an approved process a PPAP change?

Assume yes until your customer's requirements say otherwise. A submitted package describes a process — its method, its measurement system and its control plan — and replacing a documented manual or sampling method with a model changes all three. Some customers require notification, some require resubmission, and thresholds differ by customer and by characteristic. Check the customer-specific requirements before the model goes live, record the decision and its basis, and keep that record with the PPAP file so the reasoning is available at the next audit.

How long should we retain AI decision records and the images behind them?

Set retention by artefact class and drive it from the customer requirement and the product's service life rather than from storage cost. Decision records are small and should typically match your quality-record retention for that characteristic. Retained inputs are large, so define classes: keep everything for safety-related characteristics, and sample or window the rest. Whatever the schedule says, test it quarterly by retrieving the oldest record it claims to hold and confirming the input still opens and the archived model version still runs.

What does an auditor actually want to see for model documentation?

Something that reads like process documentation: intended use, data provenance, validation evidence, a performance floor and a change history, each dated, owned by someone in the quality chain and tied to the exact artefact running on the station. Length is not the signal. A two-page model file with a version identifier, a hash, an approver and a date carries more weight than a long technical report describing an architecture that cannot be matched to the station, because the auditor's question is about control rather than about method.

Who should own AI audit readiness — quality, engineering or IT?

Quality owns it, with engineering and IT supplying the mechanism. The process owner, normally a manufacturing quality engineer, has to be able to produce the evidence without help, because they are the person the auditor speaks to at the station. Engineering owns version pinning, the archived artefacts and the ability to re-run them. IT owns retention and access. Ownership sitting in a data team is the most common structural failure on this page: excellent documentation that nobody in the audit room can produce.

How do we handle a model that is retrained every few weeks?

Bring it into change control and reduce the per-change cost rather than reducing the frequency. Each version needs a change record, a validation result, a refreshed agreement study and, where the change affects the documented method or threshold, a control-plan revision. If that is too heavy for a fortnightly cadence, the honest answer is usually that the cadence is faster than the process warrants. Slow the release train to a rhythm the quality system can absorb, and batch improvements into fewer, better-evidenced versions.

What is the minimum record for one AI-influenced inspection?

Part identity, station and fixture, timestamp with shift, model identifier and version, model artefact hash, input reference, score, threshold applied, outcome, operator action with reason code, downstream disposition and a retention class. Twelve fields, written at the moment of the decision. That set answers the reconstruction question, bounds a containment by version and window, and lets the model's own behaviour be charted alongside the characteristic it judges. Anything missing from it turns a later answer into an inference.

Does a VDA 6.3 process audit treat AI differently from a certification audit?

It goes deeper at the station and lighter on the system. A certification body audits whether your quality management system covers the AI-influenced process; a VDA 6.3 auditor stands at the station and probes the process itself — is the inspection capable, is the reaction plan real and exercised, is the operator competent and does the work instruction match what they do, is the evidence available where the work happens. Plants that pass the first and fail the second usually have correct documents and no per-part records.

What happens in an audit if we cannot reconstruct a specific part?

The outcome depends far more on whether you knew than on the gap itself. A plant that raised the gap internally, has an owner, a plan and a date is demonstrating a functioning quality system, and the finding is usually proportionate. A plant discovering it live is demonstrating that its process is not under control, and the finding tends to widen to every AI-influenced station in the plant because the auditor has no basis to assume the others are different. Rehearsal is what converts the first case into the normal one.

Does the EU AI Act change what a plant auditor asks?

Not directly, and not yet in most plants. An inspection model that judges parts is a manufacturing process aid rather than a safety component of a type-approved vehicle, and the plant audit you face this year is still IATF 16949, VDA 6.3 and your customers' own requirements. Where the Act bites hardest is AI embedded in the product and reaching the road, on its own timetable through type-approval legislation. The practical response is the same either way: build the evidence into the quality system you already have certified.

How does TISAX relate to AI audit readiness?

Through the data rather than the decision. TISAX assesses information security, and the training sets behind plant models routinely contain customer drawings, prototype images and quality data covered by customer confidentiality requirements — which pulls the AI development environment, its storage and its access controls into assessment scope. It is a frequent and avoidable finding. Treat the dataset register, the model store and the annotation tooling as in-scope systems from the start, with the same access control and logging as any other environment holding customer data.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for manufacturing, logistics and energy operators — vision inspection, process prediction and decision support running against live plant data, integrated into the MES and quality systems with the record-keeping, change control and rollback that audited estates demand.

  • · Production AI delivered into MES and quality systems at automotive plants
  • · Evidence-first delivery: decision logging, model versioning and reconstruction built in
  • · Readiness reviews run jointly with plant quality, metrology and engineering teams
  • · 20 cited sources on this page

Sources

  1. International Automotive Task ForceIATF Global Oversight (opens in a new tab)
  2. International Automotive Task ForceIATF 16949:2016 — automotive quality management (opens in a new tab)
  3. International Automotive Task ForceRules for achieving and maintaining IATF recognition, 5th edition (opens in a new tab)
  4. International Automotive Task ForceCustomer-specific requirements (opens in a new tab)
  5. VDA QMCQuality management and VDA 6.3 process audits (opens in a new tab)
  6. Verband der AutomobilindustrieVDA QMC publications webshop (opens in a new tab)
  7. AIAGAutomotive quality core tools (opens in a new tab)
  8. AIAGAIAG & VDA FMEA Handbook (opens in a new tab)
  9. AIAGMeasurement Systems Analysis (MSA) reference manual (opens in a new tab)
  10. AIAGProduction Part Approval Process (PPAP) reference manual (opens in a new tab)
  11. ISOISO 9001:2015 — quality management systems (opens in a new tab)
  12. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  13. ENX AssociationTISAX — Trusted Information Security Assessment Exchange (opens in a new tab)
  14. NISTAI Risk Management Framework (opens in a new tab)
  15. European CommissionRegulatory framework for AI (EU AI Act) (opens in a new tab)
  16. AudiAI for quality control of spot welds (opens in a new tab)
  17. BMW GroupArtificial intelligence as a quality booster (opens in a new tab)
  18. Volkswagen GroupComputer vision in Group production (opens in a new tab)
  19. Deutsche Gesellschaft für QualitätGerman Association for Quality (DGQ) (opens in a new tab)
  20. FraunhoferFraunhofer-Gesellschaft (opens in a new tab)

Find out what an auditor would find — before one does

We run the assessment with your quality, metrology and engineering leads, walk the clause-to-evidence map against one real station, and leave you with a costed 90-day plan for your weakest dimension. You keep the plan and the marked-up map whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.