Manufacturing (Automotive)Regulations, Compliance & Governance
AI audit readiness in automotive manufacturing: the evidence an auditor will ask you to produce
AI audit readiness is the ability to produce, on demand, the evidence that an AI-influenced process is under control: the control plan entry, the measurement-system analysis, the model documentation, and the traceable record of why one specific part passed on one specific date.

Key takeaways
- An auditor does not audit your model. They audit your process, and the model is now part of it — which means the control plan, the PFMEA, the measurement-system analysis and the work instruction all have to say so before the audit, not during it.
- If a model influences an accept or reject decision, it is a measurement system. Attribute agreement analysis against a reference sample is the evidence that satisfies MSA expectations, and it has to be re-run on every model version, not once at installation.
- The single question that separates readiness from paperwork is reconstruction: show why this specific part passed on this specific date. Answering it needs the part identity, the model version, the input artefact, the score and the threshold — all retained together and joinable.
- Adding a model to an already-approved process is a process change. Where the customer's requirements demand notification or resubmission, an undeclared model turns a good inspection into a PPAP problem, regardless of how well it performs.
- Audit readiness is a rehearsed capability, not a document set. Plants that pick a random part quarterly and time the reconstruction find their gaps as internal findings; plants that do not find them in front of a certification body.
Abbreviations used on this page
- IATF
- International Automotive Task Force — publisher of IATF 16949, the automotive quality management standard
- VDA
- Verband der Automobilindustrie — the German automotive association behind VDA 6.3 process audits
- APQP
- Advanced product quality planning
- PPAP
- Production part approval process
- FMEA
- Failure mode and effects analysis (PFMEA when applied to a process)
- MSA
- Measurement systems analysis
- GR&R
- Gauge repeatability and reproducibility — the variation study inside an MSA
- SPC
- Statistical process control
- CSR
- Customer-specific requirements — the OEM rules layered on top of IATF 16949
- QMS
- Quality management system
- MES
- Manufacturing execution system
- TISAX
- Trusted Information Security Assessment Exchange — the automotive information-security assessment scheme
Free · 8 questions · ~3 minutes
Score your plant on audit readiness
Eight questions, one at a time, about three minutes. Answer them and we build your personalised readiness report — your stage on the Unprepared to Continuously auditable ladder, your score on each of the four dimensions, and the specific evidence gap standing between you and the next stage — and send it to your inbox. Your result doubles as the agenda for your next internal audit.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised readiness report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the evidence gaps behind each one, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Unprepared
AI already influences decisions on the floor, but nothing in the quality system records that it exists.
Your next moveBuild the register — station, decision, part families, owner, model version — then write the highest-exposure station into the control plan and PFMEA.
Stage 2 · Documented
The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
Your next moveEmit one decision record per inspection — part identity, station, timestamp, model version, input reference, score, threshold, outcome — and retain it to the customer's schedule.
Stage 3 · Traceable
Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
Your next moveState a performance floor for each model, put monitoring and an alarm behind it, and rehearse the audit question quarterly on a randomly chosen part.
Stage 4 · Rehearsed
The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
Your next moveMake the evidence a by-product: emit at decision time, bind it to the QMS record automatically, and give auditors a read-only view instead of a prepared pack.
Stage 5 · Continuously auditable
The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.
Your next movePut the readiness configuration under document control: a named owner, an annual review against every active CSR, and a quarterly test of the oldest retained record.
0 / 24
Evidence completeness
— / 6
Process documentation
— / 6
Traceability & reconstruction
— / 6
Audit rehearsal
— / 6
Your score maps to a stage on the readiness ladder. The dimension breakdown matters more than the total: the lowest dimension is the one an auditor reaches first, and it is where the next week of work belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the readiness ladder. The dimension breakdown matters more than the total: the lowest dimension is the one an auditor reaches first, and it is where the next week of work belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this read against your actual audit calendar?
We walk your quality, metrology and engineering leads through the dimension scores, map them onto your next surveillance audit, VDA 6.3 visit and customer milestones, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–5 — Stage 1, Unprepared. AI already influences decisions on the floor, but nothing in the quality system records that it exists.
- 6–11 — Stage 2, Documented. The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
- 12–16 — Stage 3, Traceable. Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
- 17–21 — Stage 4, Rehearsed. The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
- 22–24 — Stage 5, Continuously auditable. The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.
What AI audit readiness means in an automotive plant
A definition, the question every audit regime reduces to, and the evidence chain that has to answer it before the auditor arrives.
AI audit readiness is the ability to produce, on demand, the evidence that an AI-influenced process is under control. In an automotive plant that is a concrete list rather than a posture: the control plan entry that names the AI station and its reaction plan, the process FMEA rows covering how the model fails, the measurement-system study that characterises it as a gauge, the model documentation that states what it is for and how it was validated, and the per-part record that shows why this specific body passed on this specific date and shift.
Read from the auditor's side of the table, the framing changes. An auditor is not there to assess your model. They are there to assess your process — and the model has become part of it, which means it inherits every expectation the process already carried under IATF 16949 (opens in a new tab), the IATF Rules for the certification scheme (opens in a new tab), and whatever your customers add through their own requirements. Nothing about that is AI-specific. What is AI-specific is how easily a model slips into a process without any of those documents noticing, because retraining does not look like a process change to the people doing it.
Effort to answer an auditor, as readiness matures
The curve is not linear. Through the Unprepared and Documented stages, answering a specific question about a specific part is a manual assembly job whose cost barely falls — better documents do not make per-part evidence appear. The inflection is at Traceable, when the record is emitted at decision time; from there, rehearsal and automation drive the marginal cost of an audit answer towards zero. Illustrative, and consistent with the evidence-lifecycle expectations set out in the NIST AI Risk Management Framework.
Evidence available without preparation by stage
- Stage 1 · Unprepared — 22% of operators. AI already influences decisions on the floor, but nothing in the quality system records that it exists.
- Stage 2 · Documented — 38% of operators. The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
- Stage 3 · Traceable — 24% of operators. Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
- Stage 4 · Rehearsed — 12% of operators. The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
- Stage 5 · Continuously auditable — 4% of operators. The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.
Curve shape: logistic, plotted from the stage data above. Distribution: Illustrative — framed against the NIST AI Risk Management Framework.
Is the process defined?
Every regime opens here. The control plan, the work instruction and the PFMEA have to describe the process that is actually running, including the part the model plays in it. A station whose documents describe the pre-AI method fails this question before any discussion of model quality begins.
Is it capable and in control?
For a conventional gauge this is MSA and capability. For a model making accept and reject calls it is an agreement study against reference parts, a stated performance floor, and evidence that the floor is monitored rather than assumed. The AIAG core-tools material (opens in a new tab) is the vocabulary your auditor will use for it.
Can you prove it for this part?
The particular, not the general. Serial number in, evidence out: station, timestamp, model version, input artefact, score, threshold, decision, disposition. This is the question that separates a tidy document file from a readiness capability, and it is the one most plants cannot answer inside an audit slot.
What did you do when it was wrong?
Models are wrong sometimes; that is not the finding. The finding is a nonconformance handled as a station fault, with no link to the model version, no containment scoped by deployment window, and no update to the FMEA that claimed to have considered the failure mode.
How an auditor's question gets answered — or does not
The same question, asked three ways down the plant. The top lane is where most plants are for their newest AI station: the question routes to a person, who routes it to another person, and the answer is a recollection. The middle lane answers from the quality system. The bottom lane answers without preparation, because the evidence was emitted when the decision was made.
- Human in the loop
- Where value leaks
- System-of-record action
- Data & feeds
- AI / model
The process, in words
- In the unprepared lane the question routes to people rather than to systems. Quality escalates to engineering, engineering reconstructs from memory and a screenshot, and because neither the model version nor the input artefact was retained, the answer cannot be evidenced. The finding that follows is written against document control and process change, not against the technology — which is why it pulls every other AI-assisted station into scope.
- In the documented and traceable lane the auditor's question opens the control plan at that station, the control plan resolves to the decision records for that part, the record carries the model version, and the version resolves to the model file holding intended use, validation and change history. A quality engineer answers in the room, from the system, without calling the person who built the model.
- In the continuously auditable lane nothing is assembled, because the evidence was emitted when the decision was made. Each inspection writes its own record, each model version is archived and re-runnable, and the auditor is given a filtered read-only view instead of a prepared pack. The same pipeline watches the stated performance floor, and a breach raises a nonconformance without anyone deciding to raise one.
Step-by-step insights
- The screenshot answer — why it fails on category, not on content
- A screenshot of the station screen usually shows the correct outcome, and it still fails, because an audit asks for a record rather than an observation. A record has provenance: it was written by the system at the time of the event, it is retained under a schedule, and it cannot be produced after the fact. The distinction matters most when the answer is favourable — a plant that can only demonstrate good results without records has demonstrated that it has no records, which is the finding regardless of the result.
- The control plan as the entry point, not the summary
- Auditors open the control plan at the station because it is the document that binds a characteristic to a method, a frequency, a sample size and a reaction plan. When a model takes over a characteristic, all five of those change: the method is now a classifier, the frequency is often every part rather than a sample, the sample size concept may disappear entirely, and the reaction plan needs a branch for the model being unavailable or out of its envelope. A control plan that merely mentions 'vision system' has recorded the equipment and lost the process.
- The decision record is the join, and it needs a real key
- The single most common structural failure is that decisions and identities live in different systems with no shared key. The station historian has timestamps and outcomes; the MES has serial numbers and build events; the deployment log has model versions and dates. Joining them by time is an interpretation, and interpretations do not survive an audit or a warranty investigation. Read the part identity at the station and write it into the decision record — that one field converts three systems into one answer.
- Version pinning: naming a version is not the same as keeping it
- Most plants that record a model version cannot re-run it, because the artefact was overwritten, the environment moved, or the preprocessing lived in code that has since changed. Pinning means the exact artefact, the exact preprocessing and the exact threshold configuration are retained together and can be executed against the retained input. That is what turns 'the record says version 7' into 'here is version 7 producing the same score on the same image', which is the difference between a claim and a demonstration.
- The read-only auditor view — less exposure, not more
- Plants hesitate to give an auditor a live view, on the theory that a prepared pack is safer. In practice the pack is the risk: it is assembled under time pressure, it invites questions about what was left out, and it proves nothing about the underlying system. A filtered, read-only, observed view scoped to one station and one date range shows the record as it exists, answers follow-up questions in seconds, and demonstrates the control everyone is actually trying to evidence.
- The performance floor closes the loop back to the reaction plan
- A stated floor — maximum escape rate on a safety characteristic, minimum agreement against reference parts, a false-reject ceiling the line can absorb — is what makes drift an event rather than a slow disappointment. Wire the breach into the same nonconformance route a gauge out of calibration takes, and the AI station stops being a special case. That is the whole trick of AI audit readiness: not new machinery for a new technology, but making the new technology reach the machinery that already exists.
The five stages in detail
For each stage: what it looks like on the floor, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving costs.
Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions in a plant, the diagnostic signals are checks you can run against your own quality system this week, and the anti-pattern is the specific mistake most often made trying to leave that stage. The ladder runs Unprepared, Documented, Traceable, Rehearsed, Continuously auditable — and the expensive gap is between Documented and Traceable, because that is where general description has to become per-part evidence.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Unprepared
22% of operators sit here
AI already influences decisions on the floor, but nothing in the quality system records that it exists.
Stage 1 is not the absence of AI — it is the absence of AI from the quality system. The station works, the operators like it, the scrap numbers improved, and the control plan on the wall still describes an ultrasonic spot check performed every fiftieth body. Nothing in the QMS knows the model is there, so nothing in the QMS can answer a question about it.
The pattern that produces this is ordinary and blameless. A vision or process model is installed as an engineering improvement, usually under a continuous-improvement budget rather than a programme change, and the change lands on the floor faster than the document control cycle that would have carried it. Quality is told about the result rather than consulted about the method. Six months later nobody can say precisely which version has been running since May.
The cost is asymmetric and it arrives all at once. Day to day the gap is invisible; in the audit room it produces the worst possible category of finding, because the auditor is not disputing the technology — they are observing that the documented process and the actual process are different things. That is a system finding, not a station finding, and it puts every other AI-influenced station in the plant into scope on the same day.
In practice
The station that was not on the control plan
A body-shop team installed a camera-based surface inspection ahead of the sealer booth and switched the manual visual check to a spot audit. Scrap fell, the line ran better, and everybody was pleased. Eleven months later a surveillance auditor walked the line, saw a screen showing pass and fail decisions, and asked where in the control plan that characteristic was recorded. It was not. The finding was written against document control and process change, and the corrective action pulled in four other AI-assisted stations nobody had thought to mention.
What it looks like
- No register of which decisions a model touches
- Control plans and work instructions still describe the pre-AI method
- Model behaviour is explained verbally, from memory, by whoever built it
- Station logs roll off long before an audit could reach back to them
Diagnostic signals you can check this week
- Ask for a list of every decision on site that a model influences — if the answer takes more than a day to produce, you are here
- Compare the control plan for one AI-assisted station against what the operator actually does at that station
- Ask which model version has been running since a specific date three months ago, and watch who gets called
- Check the retention on the station's decision logs against your customer's record-retention requirement
Anti-pattern · Writing an AI policy instead of a control plan entry
The instinctive response to an AI-shaped compliance worry is a governance document: a policy, a committee, an ethics statement. None of these is what an auditor asks for at a station. The artefacts that discharge the obligation are the ones the plant already uses — control plan, PFMEA, work instruction, MSA record, change record — and one station properly written into those four is worth more in an audit than a twenty-page policy nobody on the floor has read.
What holds you here
Nobody has an authoritative list of where AI already influences a decision, so the gap cannot be sized, let alone closed.
Highest-leverage next move
Build the register — station, decision, part families, owner, model version — then write the highest-exposure station into the control plan and PFMEA.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- One quality engineer, one controls or ML engineer, part-time
- Risk
- Low — the work is documentation and inventory, nothing on the line changes
- To next stage
- 1–3 months
If this is you, the next step is
A two-week exercise: the register, the gaps, and the order to close them.
Stage 2
Documented
38% of operators sit here
The AI-influenced processes are described correctly in the quality system, but the evidence for any individual part still has to be assembled by hand.
Stage 2 is the stage most plants reach as soon as somebody senior asks the question, and it is a genuine improvement: the documented process and the running process agree. The control plan names the station, the characteristic and the reaction plan; the PFMEA has rows for false accept and false reject; a validation report sits in the quality file with a date and a signature on it.
What is missing is the per-part evidence. Documentation describes the process in general, and audits are conducted in the particular. The auditor picks a serial number off a rack, or a date off a shipping record, and asks what happened to that part. At stage 2 the honest answer is that the plant can describe what should have happened, and would need engineering support and a working day to establish what did.
This stage is stable and it is where most of the risk sits, because it passes casual inspection. The documents are correct, the file is tidy, and the weakness only surfaces under a specific request — which is exactly what a VDA 6.3 process auditor or an escalated customer visit is designed to make. Plants often discover their stage-2 status during a customer complaint rather than during an audit, when a containment question needs a model version and nobody can supply one.
In practice
The validation report with no serial numbers behind it
A powertrain plant documented its leak-test anomaly model properly: control plan revision, PFMEA update, a validation report showing agreement against a labelled reference set, and a reaction plan for out-of-limit behaviour. Then a customer raised a field concern on parts built across a two-week window. The question was simple — which of those parts had been judged by the model, and on which version. Reconstructing it took three people the better part of a week, because the decisions were in the station historian, the versions were in an engineering deployment spreadsheet, and the two shared no key.
What it looks like
- Control plans name the AI stations, their characteristics and reaction plans
- PFMEAs carry model failure modes, not just mechanical ones
- A validation report exists for the model version in service
- Answering a question about one part still means emailing engineering
Diagnostic signals you can check this week
- Pick a serial number from six months ago and ask for the model version that judged it — time the answer
- Check whether the deployment record and the decision log can be joined without human interpretation
- Ask whether the validation report names the exact model artefact that is running today
- Look for the reaction plan being exercised in real records, not only described in the control plan
Anti-pattern · Treating the validation report as permanent
A validation report is a statement about one model version, one dataset and one operating envelope. Plants at stage 2 routinely treat it as a property of the station: the report goes in the file, the model is retrained twice over the next year, and the file is never touched. An auditor who compares the artefact hash or the version in the report against the version on the station finds the gap in about ninety seconds, and it converts a documentation strength into a change-control finding.
What holds you here
Decisions, model versions and input artefacts live in three systems with no shared key, so per-part evidence is a manual archaeology exercise.
Highest-leverage next move
Emit one decision record per inspection — part identity, station, timestamp, model version, input reference, score, threshold, outcome — and retain it to the customer's schedule.
Cost of leaving
- Effort
- 3–5 months
- Team
- Quality engineer, MES or data engineer, ML engineer for versioning
- Risk
- Medium — the decision log has to become a retained record, which pulls in IT retention and storage
- To next stage
- 3–6 months
If this is you, the next step is
We wire the decision record and version pinning on one station, then hand you the pattern.
Stage 3
Traceable
24% of operators sit here
Any individual part can be reconstructed on demand: which model version judged it, on what input, against what threshold, with what result.
Stage 3 is where the evidence stops being assembled and starts being emitted. Every inspection writes its own record at the moment of decision, and that record carries enough to rebuild the decision later: what part, what station, what model version, what input, what score, what threshold, what outcome, what the operator did next. The quality system holds the link from the control plan's station identifier down to those records, so the walk from document to data takes one step.
The operational payoff shows up long before the audit. Containment gets cheap and precise. When a defect escapes, the question 'which parts were judged by the version that was running between these two dates' has an answer in minutes, and the containment is scoped to that population instead of to the whole build window. Plants routinely find that this alone pays for the work, because over-broad containment is one of the most expensive reflexes in automotive quality.
The constraint that emerges here is a governance one. Traceability makes it possible to answer any question about the past, but it does not by itself tell you whether the process is still inside its stated limits today, or whether anybody would notice if it were not. That is the difference between being able to satisfy an auditor when asked and being able to demonstrate control without being asked — which is what the next stage buys.
In practice
The containment that fitted on one screen
A stamping supplier had a customer complaint on a crack that the vision model should have caught. Because every inspection carried the model version and the retained image, quality pulled the exact population judged by that version in the affected shift pattern — a few hundred parts rather than the eleven days of production the customer had asked to be sorted. The 8D containment section named the version, the window and the population, and the corrective action attached the re-validated version. The customer closed it without a special audit.
What it looks like
- One decision record per inspection, joinable to the serial number or VIN
- Model versions are pinned, archived and re-runnable
- Nonconformances can be scoped by model version and time window
- Reconstruction is a query, not a project
Diagnostic signals you can check this week
- Time a reconstruction from a randomly chosen serial number — under an hour is stage 3
- Check that an archived model version can actually be re-run, not merely named
- Confirm that nonconformance records reference model version and time window as standard fields
- Ask whether the retained input artefact is the one the decision was made on, or a later re-capture
Anti-pattern · Retaining everything instead of retaining the right thing
The reflex at stage 3 is to keep every frame from every camera forever, which fails on cost and then quietly fails on completeness when someone trims the retention without changing the schedule. The record that matters is small: identity, version, input reference, score, threshold, outcome. Define a retention class per artefact type, tie it to the customer requirement and the product's service life, and sample-test the oldest retained record every quarter to prove the schedule is real.
What holds you here
The plant can answer any question it is asked, but nothing proves the process stayed inside its limits between audits.
Highest-leverage next move
State a performance floor for each model, put monitoring and an alarm behind it, and rehearse the audit question quarterly on a randomly chosen part.
Cost of leaving
- Effort
- 4–8 months
- Team
- Quality systems owner, data engineer, ML engineer, IT for retention
- Risk
- Medium — storage and retention decisions become contractual, and they are hard to reverse
- To next stage
- 6–9 months
If this is you, the next step is
One session to define the record, the retention class and the join to the QMS.
Stage 4
Rehearsed
12% of operators sit here
The plant practises the audit before the auditor arrives: performance floors are monitored, reconstruction is timed, and the gaps are raised as internal findings.
Stage 4 changes who finds the problem. The plant runs the audit on itself: an internal auditor picks a part at random, asks the questions a certification body would ask, and times how long the answers take. What comes out is a findings list — a control plan that lags a model change by three weeks, a station whose reaction plan has never been exercised, a retention class that expired earlier than the customer requires. Those are cheap findings when you raise them and expensive ones when somebody else does.
The other half of the stage is the performance floor. A model in production without a stated floor — a maximum escape rate on safety characteristics, a minimum agreement against reference parts, a false-reject ceiling the line can absorb — has no definition of 'still working', which means drift is only detectable in hindsight through scrap and warranty. Writing the floor down converts the model from an engineering asset into a controlled process with a reaction plan, which is the language every audit regime already speaks.
What makes this stage durable is that rehearsal has a cadence rather than a trigger. Plants that rehearse only when an audit is booked get good at the audit they expect; plants that rehearse quarterly on a random part get good at the audit they do not. The second group is the one that survives a customer escalation, which arrives without a date in the calendar.
In practice
The quarterly drill that found the silent rollback
During a routine reconstruction drill, a quality engineer pulled a part from the previous quarter and found the model version on the record did not match anything in the change log. The explanation was mundane: a station had been re-imaged after a controller fault and had come back with the previous model version, running for nine days before the next deployment overwrote it. Nothing had failed, no defect had escaped, and no external auditor would have looked. The drill turned it into an internal finding and a permanent change to the station recovery procedure.
What it looks like
- Every model has a stated performance floor with monitoring and an alarm
- Internal audit covers AI-influenced stations on a schedule, not on request
- Reconstruction drills are timed and the time is trending down
- Findings from rehearsal are tracked and closed like external findings
Diagnostic signals you can check this week
- Ask for the date and result of the last reconstruction drill, and who chose the part
- Check that every production model has a written performance floor with a named owner and an alarm
- Look at whether rehearsal findings are in the same tracker as external findings, with the same closure discipline
- Confirm that internal audit schedules AI stations across shifts, not only on days when engineering is on site
Anti-pattern · Rehearsing with the part you already understand
The drill only works if the part is chosen adversarially. Teams naturally reach for a part from a station they trust, in a period they remember, which tests recall rather than the system. Have somebody outside the AI team pick — quality, an internal auditor, or a random draw from the shipping record — and require the reconstruction to be produced without calling the person who built the model. That single rule is what converts a demonstration into a test.
What holds you here
Evidence is produced by the effort of a good team rather than by the system, so quality depends on which people are available that week.
Highest-leverage next move
Make the evidence a by-product: emit at decision time, bind it to the QMS record automatically, and give auditors a read-only view instead of a prepared pack.
Cost of leaving
- Effort
- 6–12 months
- Team
- Quality manager, internal audit, ML engineer, plant IT
- Risk
- Medium — rehearsal generates real findings, and the closure workload is genuine
- To next stage
- 9–18 months
If this is you, the next step is
We play the auditor for a day and leave you the findings list and the evidence gaps.
Stage 5
Continuously auditable
4% of operators sit here
The evidence is a by-product of running the process, so any AI-influenced decision is answerable at any moment without preparation.
Stage 5 removes the preparation window. Nothing is assembled for an audit because nothing needs to be: the control plan resolves to the station, the station resolves to its decision records, the records resolve to a pinned model artefact and its validation, and the change history is the deployment log rather than a document describing it. The plant's answer to 'show me' is a query somebody runs while the auditor watches.
The distinguishing behaviour is that the quality system initiates. A performance-floor breach opens a nonconformance the way a gauge going out of calibration does, with the same reaction plan discipline. A model deployment cannot complete without its validation record, its MSA refresh and its control-plan revision, because the workflow refuses to proceed. This is unglamorous plumbing and it is the only version of AI governance that survives a busy quarter, because it does not depend on anybody remembering.
Sustaining it is a change-control problem rather than a technical one. Customer-specific requirements move, retention obligations lengthen with service life, new part families arrive that were never in the validation envelope, and each of those quietly invalidates a piece of the arrangement. Plants at stage 5 treat the audit-readiness configuration itself as a controlled document with an owner and a review cycle — otherwise the system that made evidence free slowly stops matching what the evidence is for.
In practice
The audit that was conducted from a read-only screen
At a plant running several AI-assisted inspections, the surveillance audit for those stations was conducted from a filtered view: the auditor picked a serial number, and the record came back with station, timestamp, model version, retained input, score, threshold, outcome and the operator's disposition, with links to the control plan revision and the validation record for that version. The quality manager's preparation for that portion of the audit had been to check that the view still worked. Nothing was assembled, because nothing had been disassembled.
What it looks like
- Evidence is emitted at decision time and bound to the QMS record automatically
- Auditors are given a filtered, read-only view rather than a prepared pack
- Model change, control plan revision and customer notification are one workflow
- Performance-floor breaches open a nonconformance without human initiation
Diagnostic signals you can check this week
- Ask what preparation was done for the last audit of an AI station — 'none' is the stage-5 answer
- Check whether a deployment can complete without its validation and MSA artefacts attached
- Confirm that a performance-floor breach raises a nonconformance without a person deciding to raise one
- Ask who owns the readiness configuration itself, and when it was last reviewed against customer requirements
Anti-pattern · Letting the read-only view drift from the customer's requirement
Once evidence is free, attention moves elsewhere, and the configuration ages quietly. A customer lengthens a retention period, a new part family enters the line outside the validated envelope, a CSR adds a notification duty for process change — and the automated pack keeps producing yesterday's correct answer. Put the readiness configuration under document control with a named owner and an annual review against every active customer's requirements, and sample-test the oldest retained record each quarter.
What holds you here
The readiness configuration ages against moving customer-specific requirements, so a correct system slowly stops answering the current question.
Highest-leverage next move
Put the readiness configuration under document control: a named owner, an annual review against every active CSR, and a quarterly test of the oldest retained record.
Cost of leaving
- Effort
- Continuous
- Team
- Quality systems owner, platform engineer, standing change-control forum
- Risk
- Concentrated — low frequency, high consequence, and always tied to a customer requirement change
If this is you, the next step is
We attack the pipeline with real audit scenarios and report what it cannot answer.
Where automotive plants sit — and the four dimensions that set the stage
The distribution across the ladder, why Documented is the plateau, and the four dimensions that gate each other.
Most automotive plants are at Documented. Once someone senior asks the question, control plans and FMEAs get updated quickly — that work is well understood and it is measured in weeks. What does not follow automatically is the per-part evidence, because that requires a change to how the station writes records rather than a change to what a document says. The distribution below is illustrative rather than surveyed, and it is drawn to match a pattern any quality manager will recognise: a large documented middle, a thin top, and a tail that has not started.
Illustrative distribution of automotive plants across the readiness ladder
Illustrative, not surveyed. The shape reflects the structural argument on this page: documentation is cheap and follows a single management request, while per-part traceability requires the station to emit records and therefore lags. Read the shape, not the decimals.
Share of plants
- 22% — 1 · Unprepared
- 38% — 2 · Documented (the plateau)
- 24% — 3 · Traceable
- 12% — 4 · Rehearsed
- 4% — 5 · Continuously auditable
The pressure on the plateau is coming from two directions at once. Customers are adding questions about AI-influenced processes to existing audits rather than creating new ones, which means the evidence has to sit inside the quality system your IATF 16949 certification (opens in a new tab) already covers. At the same time the horizontal frame is filling in: ISO/IEC 42001 (opens in a new tab) gives an AI management system something to certify against, the NIST AI Risk Management Framework (opens in a new tab) supplies the vocabulary customer questionnaires increasingly borrow, and the EU AI Act (opens in a new tab) reaches AI embedded in type-approved products on its own timetable. None of these replaces the plant audit. All of them make the plant audit ask more.
Evidence completeness
Whether the artefacts an auditor asks for exist at all: control plan entry, PFMEA rows, agreement study, validation record, change history, competence records. This is the cheapest dimension to fix and the one most often assumed to be someone else's file.
Process documentation
Whether the documents describe the process that is actually running today, including after the last retraining. A correct document that is three model versions out of date is a change-control finding, which is a heavier category than a missing document.
Traceability and reconstruction
Whether the plant can go from one part to the decision that judged it. This is the dimension that separates Documented from Traceable, and it is overwhelmingly the lowest-scoring dimension in readiness reviews, because it is the only one that cannot be fixed by writing something.
Audit rehearsal
Whether the plant tests itself on a cadence rather than before a date. Rehearsal converts unknown gaps into internal findings, which are cheap; the alternative is discovering them in a surveillance audit or a customer escalation, which is not.
Diagnosing the real constraint
Plot your process documentation against your traceability. The quadrant names the next investment — and only one of the four answers is 'write another document'.
Traceable but undeclared
- The data exists; the quality system does not know the model does
- Common where AI arrived as an engineering improvement
- Fix: control plan revision, PFMEA rows, work instruction — weeks, not quarters
Audit-ready
- Documents match the floor and the floor emits records
- Constraint moves to rehearsal and change control
- Fix: quarterly reconstruction drills and a performance floor with an alarm
Unprepared
- Neither the document nor the record exists
- Every AI station in the plant is in scope on the same day
- Fix: inventory first, then the highest-exposure station end to end
Paper compliance
- The file is tidy and nothing can be proved for a specific part
- The most dangerous quadrant — it passes a document review
- Fix: emit the decision record before improving another document
What each audit asks, and the evidence that answers it
First the regimes that actually visit an automotive plant. Then the map that matters: what is asked, what satisfies it, who owns it, and where it lives.
Four kinds of audit reach an automotive plant, and AI changes what each one asks without changing who asks it. The certification body arrives on the IATF cycle and audits the quality management system. A customer sends a VDA 6.3 (opens in a new tab) process auditor and audits the process, element by element, at the station. Customer-specific requirements (opens in a new tab) arrive continuously through supplier quality and add duties the standard does not — notification thresholds for process change, retention periods, run-at-rate conditions. And your own internal audit programme is meant to find all of it first. A fifth, TISAX (opens in a new tab), is not a quality audit at all, but it reaches the AI development environment through the customer data your training sets contain. All of it sits on the ISO 9001 (opens in a new tab) management-system spine, which is why none of these audits needs an AI clause to reach a model: they already reach every process.
| Audit | Who runs it | Rhythm | What it asks once a model is in the loop |
|---|---|---|---|
| IATF 16949 certification audit | An IATF-recognised certification body under the IATF Rules | Three-year certification cycle | Whether the quality management system covers the AI-influenced process at all: control plan, PFMEA, MSA, competence, change control, records and retention |
| IATF 16949 surveillance audit | The same certification body | Between certification audits, on the scheme's schedule | Whether anything changed since the last visit without the system noticing — new stations, new model versions, uncontrolled documents, drifting performance |
| VDA 6.3 process audit | The customer, or a qualified VDA 6.3 auditor on their behalf | Per programme, per escalation, or on a customer schedule | The process at the station: is the AI inspection capable, is the reaction plan real, is the operator competent, is the evidence available where the work happens |
| Customer-specific requirements | Customer supplier-quality engineers | Continuous, plus programme milestones | Whatever that customer's own rules add: notification duties for process change, approval before a model may replace a documented method, retention periods, run-at-rate conditions |
| PPAP or part submission review | Customer programme quality | Per part, per change | Whether the submitted package still describes the process actually running — a model added after approval is a process change until proved otherwise |
| Internal audit programme | Your own trained internal auditors | Planned coverage of processes, shifts and clauses | Everything above, first. Internal audit is where an AI station should fail, at a cost you control |
| Layered process audit | Supervisors through to plant leadership | Daily to monthly, by layer | Whether the AI station is being run today the way the work instruction says, including who may override a reject and what they record when they do |
| TISAX assessment | An ENX-accredited audit provider | On the label's validity cycle | Information security around the data the models are built on — customer drawings, prototype images and quality data pull the AI development environment into assessment scope |
| ISO/IEC 42001 certification | An accredited certification body, voluntarily engaged | Optional today, appearing in customer questionnaires | Whether an AI management system exists around the models: policy, impact assessment, lifecycle controls, monitoring and improvement |
The map below is the centre of this page. Each row is a question an auditor asks in words a plant recognises, the artefact that satisfies it, the person who has to own that artefact, and the system it has to live in. The column that fails in practice is 'who owns it'. Model documentation drafted by a data team and stored in a code repository is not owned by anyone the auditor will speak to; the moment the process owner cannot produce it without help, the plant is at stage 1 for that station regardless of how good the documentation is.
| What the auditor asks | What satisfies it | Who owns it | Where it lives |
|---|---|---|---|
| Is this process defined and controlled? | Control plan revision naming the AI station, the characteristic it judges, the method, the frequency and the reaction plan for the model being unavailable or out of envelope | Process owner (manufacturing quality engineer) | Control plan register in the QMS, with the station identifier that also appears in the MES |
| Has the risk of this process been assessed? | PFMEA rows for model-specific failure modes — false accept, false reject, drift, changed lighting or fixture, wrong version deployed — with detection and controls that are real | Cross-functional FMEA team, chaired by quality | FMEA workbook in the QMS, revision-linked to the control plan |
| Is the thing making the decision a capable measurement system? | Attribute agreement analysis against a reference set of known-good and known-bad parts, with acceptance criteria agreed before the study and recorded against the model version | Metrology or quality lab, with the process owner | MSA study record plus an entry in the gauge and equipment register |
| Was the part approved with this process? | PPAP package whose control plan, MSA and capability evidence describe the process actually running; a documented process change if the model arrived after approval | Programme quality (APQP lead) | Internal PPAP file and the customer's submission portal |
| What is this AI system for, and what is it not for? | Intended-use statement: part families, characteristics, operating envelope, explicitly excluded conditions, and what the model must never be used to decide | AI product owner, countersigned by the process owner | Model file in the model register, referenced from the control plan |
| Where did the training data come from? | Provenance record: source stations, date ranges, part numbers, how images or signals were selected, exclusions, and the labelling procedure with labeller qualification | Data owner in engineering | Dataset register with content hashes, linked to the model version |
| How do you know it works? | Validation report against a held-out labelled sample, broken down by defect class and part family, with acceptance criteria set before the run and the exact model artefact identified | Quality engineering | Validation record attached to the model version in the model register |
| What performance must it hold, and who watches? | A written performance floor — escape ceiling on safety characteristics, minimum agreement, false-reject ceiling — with monitoring, an alarm and a named responder | Process owner, with engineering on call | Monitoring system, with the floor and reaction restated in the control plan |
| What changed, when, and who approved it? | Change record per model version: what changed, the dataset it was trained on, the validation result, the MSA refresh, the approver, the effective timestamp and the stations affected | Change control board | Change record in the QMS, keyed to the MES or station deployment log |
| Why did this part pass on this date? | One decision record per inspection: part identity, station, timestamp, model version, input artefact reference, score, threshold, outcome and the operator's disposition | Quality systems owner (MES and data) | Quality data store, retained to the customer's schedule and joinable to the build record |
| What did you do when it was wrong? | Nonconformance and 8D treating the model as part of the process: containment scoped by model version and time window, root cause, corrective action, and updates to the PFMEA and control plan | Quality manager | Nonconformance and 8D system, cross-referenced to the model version |
| Who is competent to run this? | Training records for operators, team leaders and quality staff covering the AI station, override authority, escalation, and what to do when the model is unavailable | Area manager with training and HR | Competence matrix and training records, auditable per shift |
| How long do you keep all of it? | A retention schedule per artefact type, aligned to the customer requirement and the product's service life, with evidence that the oldest retained record is still retrievable | Quality systems owner | Records retention schedule, tested quarterly against live storage |
Two properties of this map decide how much work it represents. First, nothing in the left-hand column is new. Every one of those questions has been asked of automotive processes for decades; what changes is that the answer now has to include a model version and a retained input. Second, every row has a home in a system the plant already runs — QMS, MES, gauge register, training records — which means audit readiness is an integration exercise rather than a new platform. Plants that respond by buying an AI governance tool disconnected from the QMS end up with a fourteenth system and the same finding.
How to use the map
Fill the ownership column first
Go down the map and write a name against every row for one station. Rows with no name, or with a name outside the quality chain, are your real gaps — an artefact nobody in the audit room owns is an artefact that will not be produced in the audit room.
Then fill 'where it lives' with a system, not a folder
A path on a shared drive is a folder; the QMS, the gauge register, the MES and the model register are systems with access control, versioning and retention. Anywhere the answer is a folder, the artefact is one reorganisation away from being unfindable.
Close the row that fails soonest, not the row that is easiest
Order the remaining gaps by which audit reaches them first. A missing control plan entry is found by any auditor who walks the line; a retention gap is found only when someone asks for an old part. Sequence accordingly, and record the reasoning — an auditor will accept a prioritised, dated plan far more readily than a silent gap.
The quality artefacts AI now touches
APQP, PPAP, the control plan, the FMEA and the MSA were all written for processes with human or mechanical decision-makers. Here is what changes in each when a model joins.
Every core quality artefact in an automotive plant changes when a model influences a decision, and none of them needs to be replaced. The AIAG core tools (opens in a new tab) — advanced product quality planning, the production part approval process (opens in a new tab), the AIAG and VDA FMEA method (opens in a new tab), measurement systems analysis (opens in a new tab) and statistical process control — already carry the concepts a model needs: a defined method, a characterised measurement system, an assessed failure mode, an approved part submission. What changes is what has to be written into each one, and what an auditor will find if it is not.
| Artefact | What it says without AI | What must change with a model in the loop | The finding when it does not |
|---|---|---|---|
| APQP phase gates | Process design is signed off before build, with the measurement method agreed at the gate | The model is a process design element: its intended use, data plan, validation plan and acceptance criteria belong at the gate, not after launch | A model introduced post-launch with no design record, so the launch package describes a process that no longer exists |
| PFMEA | Failure modes of the equipment and the operator, with detection controls and rankings | Model failure modes as first-class rows: false accept, false reject, drift, changed lighting or fixture, wrong version deployed, model unavailable | A PFMEA that ranks detection highly on the strength of an inspection whose own failure modes were never assessed |
| Control plan | Characteristic, method, sample size, frequency, reaction plan | The method becomes the model and its threshold; frequency often becomes every part; the reaction plan needs branches for out-of-envelope and unavailable | A control plan describing an ultrasonic or manual check that stopped being performed months ago — documented process versus running process |
| MSA | GR&R on the gauge, appraiser variation, bias and linearity where applicable | Attribute agreement analysis on the model against reference parts, re-run per version, with acceptance criteria recorded before the study | An accept or reject decision made by an uncharacterised measurement system — usually the single most damaging AI finding available |
| SPC and capability | Control charts on the characteristic, capability indices from measured values | Charting the model's own behaviour — score distributions, reject rate by shift and part family — alongside the characteristic it judges | Drift discovered through scrap and warranty months later, with no chart that would have shown it earlier |
| PPAP package | Design records, control plan, MSA, capability, part submission warrant | The package must describe the process that is running; adding a model to an approved process is a process change until the customer's rules say otherwise | A submission whose control plan and MSA no longer match the floor — a document-integrity problem across every part family on that line |
| Work instructions | How the operator performs and records the check | What the screen shows, what the operator may override, what they must record when they do, and what happens when the model is offline | Overrides happening with no authority defined and no reason captured, so the record cannot explain its own exceptions |
| Reaction plan | What to do when the characteristic is out of limits | Additional branches: model out of envelope, monitoring alarm, version mismatch at the station, retained-input capture failing | A station that keeps running and keeps deciding when its own preconditions have failed |
| Gauge and equipment register | Calibration status, interval, responsible person | The model as a registered decision-making instrument: version, agreement study date, next refresh, owner | No equivalent of a calibration interval, so nothing triggers a re-study when the world changes around the model |
| Competence matrix | Who is trained and authorised on the process | Who may run, override, escalate and disable the AI station, per shift, with dated evidence | An auditor finding an authorised override performed by someone with no record of training on the station |
| Records retention | How long quality records are kept, by type | Retention classes for decision records, retained inputs and model artefacts, tied to the customer requirement and service life | The evidence existing at the time of the decision and not at the time of the question |
The measurement-systems row is the one most worth dwelling on, because it is the least intuitive and the most consequential. A vision model that outputs pass or fail is an attribute gauge, and the established way to characterise an attribute gauge is an agreement study: a set of reference parts whose true state is known and agreed, judged repeatedly, with agreement measured against the reference and between repeats. That gives you the same shape of evidence a GR&R gives for a variable gauge — a number, an acceptance criterion, and a date. Nothing about the model being a neural network changes the applicability of the method; what changes is that the study has to be repeated on every version, because a retrained model is a different gauge.
The reference set has to be curated, not sampled
A reference set drawn at random from production will contain almost no defects, which makes agreement look excellent and means nothing. Curate it: known-good and known-bad parts across the defect classes the control plan cares about, with the true state agreed by more than one qualified person and recorded. Keep the physical parts where you can, and the retained inputs where you cannot.
Agreement must be reported by defect class, not in aggregate
An overall agreement figure hides the class that matters. A model with excellent aggregate agreement can be systematically poor on the one defect type that maps to a safety characteristic, and that is precisely the breakdown an auditor asks for when the characteristic is customer-designated. Report the matrix, not the headline.
The study belongs to the version, not the station
Attach the agreement study to the model version in the register, the same way a calibration certificate attaches to a specific gauge. When the version changes, the study's status becomes 'due', and the deployment workflow should refuse to complete without a current one. This is the mechanism that keeps documentation from ageing invisibly.
Re-study when the world changes, not only when the model does
New lighting, a new fixture, a new supplier's surface finish, a new part family: each of these moves the input distribution without touching the model. Define the triggers in the control plan alongside the interval, and treat an unplanned trigger the way you would treat a gauge dropped on the floor.
Keep the acceptance criterion out of the analyst's hands
Agree the acceptance criterion before the study runs and record it with the plan. A criterion chosen after the result is a finding waiting to be made, and it is trivially detectable — the auditor simply asks when the number was decided and by whom.
In order to make the process audit- and certification-proof, development at the Neckarsulm location was carried out in close coordination with the German Association for Quality (DGQ), the Fraunhofer Institute for Industrial Engineering (IAO), and the Fraunhofer Institute for Manufacturing Engineering and Automation (IPA).
That sentence is worth reading twice, because of what it implies about sequencing. Audi did not build an inspection model and then ask how to certify it; the audit and certification question was carried alongside the development, with the German Association for Quality (opens in a new tab) and two Fraunhofer institutes (opens in a new tab) involved in how the process would be evidenced. The same release notes that there are no independent certifications for AI applications of this kind, which is exactly why the evidence has to be built into the quality system the plant already has certified rather than deferred to a scheme that does not yet exist.
What audit-grade AI looks like in public
Three publicly reported programmes, read against the readiness ladder. None is an Atomic Loops engagement — each links to the manufacturer's own published material.
The clearest public evidence for this page's argument is in what large manufacturers said about their own AI quality programmes — not the accuracy claims, but the surrounding sentences about how the process would be evidenced, who was involved, and how it would scale across sites. In each case below the differentiator is structural: what the model was allowed to decide, what recorded the decision, and what had to be true before the method could replace a documented one.
Three programmes read against the ladder
Outcomes as reported by the manufacturers themselves. The card images are illustrative industry scenes from our media library, not photographs of the named manufacturers' facilities, and no endorsement is implied. Verify figures against the linked source before reusing them; we have not independently audited them.
AudiPremium OEM · Neckarsulm body shop, rolled out across Group sites24
- Challenge
- Resistance spot welding quality had been monitored by ultrasonic checks on a random sample — roughly 5,000 spot welds per vehicle inspected by production staff — a documented sampling method that could never reach the whole population.
- Approach
- Audi moved the assessment to AI analysis of the welding process data, and reports developing it in close coordination with the German Association for Quality (DGQ) and the Fraunhofer IAO and IPA institutes specifically so the process would be, in Audi's words, audit- and certification-proof.
- Reported outcome
- Audi reports analysing around 1.5 million spot welds on 300 vehicles each shift at Neckarsulm, with employees redirected to investigating flagged anomalies, and the technical infrastructure being installed at three further Volkswagen Group locations.
- What it shows about the curveReplacing a documented sampling method with a model is a process change before it is an improvement. Carrying the audit question alongside the development — with an external quality body in the room — is what lets the change land in the control plan rather than in a finding.
Audi — AI for quality control of spot welds (opens in a new tab)
BMW GroupGlobal OEM · Plant Regensburg, around 1,400 vehicles a day34
- Challenge
- Final inspection at a plant building roughly 1,400 vehicles a day — one every 57 seconds — has to cover an enormous configuration space, where a single generic inspection catalogue is either too long for the takt or too short for the variant in front of the inspector.
- Approach
- BMW Group reports an AI system, developed at Regensburg with a start-up partner, that analyses vehicle configuration and live production data to generate a vehicle-specific inspection scope, ordered and delivered to trained specialists through a smartphone app with standardised coding for findings.
- Reported outcome
- BMW Group publicly describes the system generating customised inspection specifications per vehicle and organising the inspection intelligently, with findings recorded through standardised coding — including optional voice capture with transcription.
- What it shows about the curveWhen a model decides what gets inspected, the inspection scope itself becomes a controlled characteristic. Standardised coding of findings is the quiet part that matters for readiness: it is what makes the resulting records comparable, queryable and reconstructable later.
BMW Group — artificial intelligence as a quality booster (opens in a new tab)
Volkswagen GroupMulti-brand group · computer vision scaled across plants23
- Challenge
- Computer-vision quality applications were being developed plant by plant — label verification at one site, press-shop crack detection at another — with each deployment carrying its own evidence practices into a separately certified quality system.
- Approach
- Volkswagen Group reports developing computer-vision applications with a dedicated expert team and rolling them out across the Group through its industrial cloud platform, so solutions built at one location become available to others rather than being rebuilt.
- Reported outcome
- The Group describes applications including label content and placement verification at Porsche Leipzig and machine-learning detection of fine cracks and defects in press-shop components at Audi Ingolstadt, with Group-wide rollout via the shared platform.
- What it shows about the curveScaling a model across sites scales the audit obligation with it. Every receiving plant is separately certified, so the evidence pattern — control plan wording, agreement study, decision record — has to travel with the model, or the tenth deployment starts its documentation from zero.
Volkswagen Group — computer vision in production (opens in a new tab)
Read together, the three make one argument. Audi shows the sequencing: the audit question travels with the development, not after it. BMW shows the scope question: once a model decides what to inspect, the inspection plan is itself a controlled output and the coding of findings determines whether anything can be reconstructed. Volkswagen shows the multiplication problem: a model that scales across plants multiplies the number of separately certified quality systems that must each carry the same evidence. None of these is an argument about model quality, and that is the point.