Construction & InfrastructureRegulations, Compliance & Governance
Project AI NIST compliance in construction and infrastructure: running the AI RMF project by project
Project AI NIST compliance is the practice of running the NIST AI Risk Management Framework at the level of one construction or infrastructure project rather than the enterprise. It produces a written project profile — which AI is in use, what each use is for, how it is measured, what happens when it is wrong — that survives handover.

Key takeaways
- The NIST AI RMF is voluntary, non-certifiable and use-case agnostic. There is no such thing as a NIST-certified project — what the framework produces is a written profile and a record, and on a construction project that record is the deliverable, not a badge.
- The framework assumes a persistent organisation. Construction delivers through temporary ones, so GOVERN lives in the firm while MAP, MEASURE and MANAGE happen inside a project that will not exist in three years. That mismatch, not the framework's content, is what makes project AI compliance hard.
- The AI RMF Core carries 4 functions, 19 categories and 72 subcategories. A usable project profile triages that down to the handful that apply per AI use and names the existing project artefact — the PEP, the RACI, the risk register, the method statement — that already carries each one.
- MEASURE is where projects fail. Reliance is placed on the vendor's published benchmark rather than on a number measured against this project's own held-out records and seeded defects, and the gap only becomes visible when something is missed.
- The evidence pack is worth more at handover than during delivery. The asset owner inherits every consequence of the project's AI and none of the project's memory, so a profile that is not a contracted information deliverable is a profile that dies at practical completion.
Abbreviations used on this page
- AI RMF
- NIST Artificial Intelligence Risk Management Framework (NIST AI 100-1, January 2023)
- GAI
- Generative AI — the subject of NIST's separate Generative AI Profile, NIST AI 600-1
- TEVV
- Test, evaluation, verification and validation — NIST's term for the evidence behind a performance claim
- PEP
- Project execution plan
- BEP
- BIM execution plan (the ISO 19650 information-delivery plan)
- CDE
- Common data environment (the ISO 19650 project information store)
- JV
- Joint venture — the temporary company two or more contractors form to deliver one project
- EWN
- Early warning notice (the NEC contract mechanism for raising an emerging risk)
- NCR
- Non-conformance report
- RFI
- Request for information (the formal query from site to the design team)
- CDM
- Construction (Design and Management) Regulations 2015
- O&M
- Operation and maintenance information — what the asset owner receives at completion
Free · 8 questions · ~3 minutes
Score one project against the AI RMF ladder
Eight questions about one live project, one at a time, about three minutes. Answer them and we build your personalised profile report — your rung on the ladder, your score on each of the four dimensions, and the specific artefact standing between you and the next rung — and send it to your inbox. Score the project, not the firm: the answers differ, and the project's answer is the one a reviewer will ask for.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised project profile report is ready
Tell us where to send it. Your rung appears on screen straight away, and the full report — dimension scores, the artefacts you are missing in the order they should be produced, and a 90-day plan for the weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Referenced
Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
Your next moveWrite the register: every AI use on one live project, each with a named owner and a one-paragraph intended purpose.
Stage 2 · Profiled
Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
Your next moveBuild the project test set — held-out project records plus seeded defects — and set one written threshold per AI use.
Stage 3 · Measured
Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
Your next moveWire the AI rows into the project's existing risk machinery — risk register, early warning notice, deviation record — and rehearse one end to end.
Stage 4 · Managed
Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
Your next moveMake the evidence pack a deliverable: name it in the information-delivery plan and hand it over with the asset information.
Stage 5 · Transferable
Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.
Your next moveClose the loop: make portfolio review of returned project profiles a scheduled item that changes the policy, the template and the tool catalogue.
0 / 24
Register and intended purpose
— / 6
Measurement on project data
— / 6
Response and escalation
— / 6
Handover and durability
— / 6
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the project's position, and on most projects it is measurement or handover rather than the register everyone starts with. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the project's position, and on most projects it is measurement or handover rather than the register everyone starts with.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want the missing artefacts drafted against your own project?
We take your live project — its packages, its systems, its contract form — and produce the register, the intended-purpose statements and the measurement plan for the three highest-consequence AI uses on it, mapped to the AI RMF subcategories each one answers. No obligation, and you keep the artefacts either way.
How the score maps to a stage
- 0–4 — Stage 1, Referenced. Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
- 5–10 — Stage 2, Profiled. Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
- 11–16 — Stage 3, Measured. Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
- 17–21 — Stage 4, Managed. Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
- 22–24 — Stage 5, Transferable. Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.
What project-level NIST AI RMF compliance actually is
A definition, the four functions, and the diagram that shows which of them belong to the firm and which belong to a project that will not exist in three years.
Project-level NIST AI RMF compliance is the practice of running the NIST AI Risk Management Framework (opens in a new tab) inside a single construction or infrastructure project, and producing artefacts that a reviewer could ask that project for. Not the firm's policy — the project's register, its intended-purpose statements, its test records, its reliance decisions and its deviation history. The framework is voluntary, non-sector-specific and use-case agnostic by design, which means it tells you what has to be true and leaves you to name the artefact that makes it true. On a project, that artefact almost always already exists.
Two things about the framework matter before anything else. First, it is not certifiable: there is no NIST audit, no certificate and no accredited body issuing one, and any supplier claiming to be 'NIST AI RMF certified' is describing a self-assessment. Certification is the domain of ISO/IEC 42001 (opens in a new tab), which is a different instrument for a different purpose. Second, the framework is organised into four functions — GOVERN, MAP, MEASURE and MANAGE — deliberately echoing the structure of NIST's older and far more widely deployed Cybersecurity Framework (opens in a new tab), which many construction IT and information-security teams already work to. Of the four, GOVERN is explicitly cross-cutting and the other three are context-bound. That distinction is the whole of this page's argument, because in construction GOVERN belongs to a company that persists and MAP, MEASURE and MANAGE belong to a project organisation that does not.
The published Core carries 4 functions, 19 categories and 72 subcategories, with a companion Playbook (opens in a new tab) of suggested actions and a set of crosswalks (opens in a new tab) to other regimes. Read end to end that is a formidable document and a poor project instrument — nobody mobilising a viaduct package is going to work through 72 subcategories per tool. The usable move is triage: for each AI use on the project, identify the handful of subcategories that genuinely bite, and name the existing project artefact that already carries each one. The diagram below is the shape that triage takes.
Where the four AI RMF functions attach to a construction project
Three organisations, one framework. The enterprise persists and owns GOVERN. The project is a temporary organisation and owns MAP, MEASURE and MANAGE. The asset owner sets requirements before the project exists and inherits its consequences afterwards. Compliance fails at the two boundaries, not in the middle lane.
- Human in the loop
- Data & feeds
- AI / model
- System-of-record action
- Where value leaks
The process, in words
- The enterprise lane is GOVERN and only GOVERN: a board-owned AI policy with a risk appetite, an approved tool catalogue that says what may be used and where, and a profile template that says what every project must fill in. None of these is project-specific and none of them, on their own, produces a single piece of project evidence.
- The project lane is where compliance actually happens, in order. At mobilisation the project writes its AI register (MAP 1). Each row gets an intended-purpose statement and a consequence class taken from the project's own risk matrix (MAP 2–5). Each high-consequence row gets tested against held-out project records and seeded defects (MEASURE). The test result produces a written reliance decision — what the output may and may not be used for — and everything outside that decision routes to the project's existing deviation and early-warning machinery (MANAGE).
- The asset owner's lane brackets the project at both ends. Before the project exists, the employer's requirements set which AI is mandated, permitted or excluded and what information must be delivered. After it ends, the operating duty holder inherits an asset whose design, inspection and commissioning records were partly AI-assisted, and has no access to the project's reasoning except through the evidence pack.
- The two dashed returns are what turns a project exercise into an organisational capability: escalations from the project's deviation route reach portfolio review, and the completed evidence pack goes back to the enterprise so the next mobilisation starts from a filled-in profile rather than a blank template.
Step-by-step insights
- Why GOVERN in the enterprise lane is necessary and not sufficient
- NIST places GOVERN at the centre of the framework deliberately: it is the function that infuses the other three, covering culture, accountability, policy and third-party risk. In a firm with a stable product this works cleanly, because the same people who set the policy also run the systems it governs. In construction the two are separated by a contract, a site and often a joint-venture boundary. The result is a predictable failure: the enterprise completes GOVERN honestly, the projects are told they are covered, and no project artefact is ever created. The policy is a floor, not a profile. Treat it as an input to the project lane — a source of constraints — and judge compliance by what exists on the project.
- The register is written at mobilisation because that is the only cheap moment
- Mobilisation is the one point in a project's life when everybody expects to fill things in: the PEP is being written, the RACI is being agreed, the CDE is being configured, method statements are being drafted. Adding a register row per AI use costs minutes then and days later, because by month six the tools have accumulated informally and finding them means interviewing package leads. The register also has to catch three arrival routes that nobody thinks of as procurement: embedded features inside tools already in use, client-specified platforms named in the employer's requirements, and individual subscriptions bought on expenses. Two of those three the project did not choose, and both are on its risk.
- Intended purpose and consequence class do the triage the framework cannot
- 72 subcategories is unusable per tool; two sentences per tool is not. The intended-purpose statement says what this output is for on this project — 'flags probable clashes for engineer review, does not approve designs'. The consequence class says what happens if it is wrong, expressed on the project's own existing risk matrix rather than an invented AI-specific scale. Those two fields together determine everything downstream: how much test evidence the use has to carry, whether a DPIA is needed, whether a second human check stays in the method statement, and which subcategories are worth writing against at all. Projects that skip straight to controls end up applying the same weight to a document summariser and a lifting-plan checker.
- MEASURE on project data — the step that is almost always skipped
- The single most common gap in the middle lane is that no number in the register was produced by the project. Vendor benchmarks are computed on the vendor's corpus, under the vendor's definition of a correct answer, and averaged across customers whose sites, standards and design conventions differ from yours. They are useful for shortlisting and useless for reliance. A project test set is a records exercise, not a research one: pull known-bad cases from your own NCR log, RFI log, snagging list and design-review tracker, hold them out, add a controlled set of deliberately seeded defects, and score. NIST's term for this whole family of activity is TEVV, and the framework is explicit that its parameters must be aligned to actual deployment conditions.
- The reliance decision is the artefact a reviewer will ask for
- Registers describe and tests measure, but neither says what the project decided. The reliance decision closes that gap in one dated line: given this measured performance, this output may be used for X and may not be used for Y, and here is the residual control. It is the construction equivalent of a design-check certificate — an explicit statement of what has been accepted, by whom, on what evidence. It is also the artefact that protects individuals: without it, an engineer who acted on a flagged clash is carrying a judgement nobody recorded, and a project manager who overrode one is carrying a different one.
- Both boundaries fail silently, which is why the dashed edges matter
- Nothing on a project announces that the handover boundary has failed. Practical completion happens, the team demobilises, the CDE is archived, and the omission surfaces years later as an unanswerable query. The mobilisation boundary fails the same way: the next project starts, a blank register is issued, and the firm pays the stage-1 cost again while believing it has an AI capability. The fix in both cases is contractual rather than cultural — name the evidence pack in the information-delivery plan so it is a deliverable with a date, and make portfolio review of returned profiles a scheduled item with an owner. Neither is expensive; both are invisible until they are missing.
The ladder: from Referenced to Transferable
Five rungs describing how far the framework has actually travelled into a project — what each looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps projects there, and what leaving costs.
The ladder measures how far the framework has travelled from a document into a project's own machinery. It runs Referenced → Profiled → Measured → Managed → Transferable, and the order is not arbitrary: each rung is the function that becomes possible once the previous one has produced an artefact. You cannot measure what is not registered, you cannot rehearse a response to a failure you cannot detect, and you cannot hand over a record that was never assembled.
Read the rungs as descriptions of a project rather than of a firm. A contractor with a mature enterprise policy will have projects spread across three rungs simultaneously, and the useful question is always which rung the scheme in front of you is on. Each panel below is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run this week against your own CDE and risk register, and the anti-pattern is the specific mistake most often made trying to leave that rung.
Defensible reliance released against position on the ladder
The curve is not linear. Referenced and Profiled release very little defensible reliance — a register describes untested confidence — and the inflection is at Measured, when the first number the project can actually interrogate exists. Transferable adds little inside the project and everything after it.
Defensible reliance on project AI by stage
- Stage 1 · Referenced — 36% of operators. Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
- Stage 2 · Profiled — 30% of operators. Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
- Stage 3 · Measured — 20% of operators. Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
- Stage 4 · Managed — 11% of operators. Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
- Stage 5 · Transferable — 3% of operators. Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.
Curve shape: logistic, plotted from the stage data above. Distribution: Structured on the NIST AI RMF Core functions (NIST AI 100-1).
Select a rung
Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Referenced
36% of operators sit here
Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
Referenced is not ignorance. It is usually the opposite: someone in the business has read the framework properly, written a good paragraph about it, and won work with it. The gap is structural rather than intellectual — the paragraph was written by a bid team in a head office, and the thing it describes has to happen inside a project organisation that did not exist when the bid was submitted and that answers to a different director.
What makes this stage hard to see from the inside is that everything about it looks compliant. There is a policy. There is a named framework. There may even be a slide in the induction pack. What is missing is any artefact on the project that a reviewer could ask for: no register, no intended-purpose statement, no measurement plan, no owner. Ask for the project's AI risk position and you get a firm-level document plus an assurance that the project follows it.
The cost of staying here is not a fine — the AI RMF carries no penalty, because it is voluntary. The cost is that reliance accumulates invisibly. Site teams start acting on outputs they have never tested, design teams start submitting drawings a tool helped produce, and commercial teams start valuing progress from an analytic nobody has calibrated. None of that is recorded, so when a question arrives — from a client, an insurer, a coroner or a regulator — the project has to reconstruct from memory what it was relying on, and memory has already demobilised.
In practice
The tender answer nobody read again
A tier-one contractor won a place on a five-year highways framework partly on a digital submission that stated its AI use 'aligns with the NIST AI Risk Management Framework'. Eighteen months into the first scheme, a walk of the project turned up eleven distinct AI tools in daily use: two progress-analytics platforms, a document-summarising assistant licensed by the design team, an embedded scheduling optimiser in the planning software, telematics anomaly alerts on hired plant, and six individual generative subscriptions bought on expenses. The only artefact anyone could produce was the tender paragraph.
What it looks like
- The framework is cited in a PQQ or tender response; no project artefact exists
- Nobody on the project team can say how many AI tools are running on it
- AI arrives by individual subscription, embedded vendor feature and client-specified platform, untracked
- The words 'AI' and 'machine learning' do not appear in the PEP, the risk register or the RACI
Diagnostic signals you can check this week
- Search the project execution plan for 'AI' and for 'machine learning', and count the hits. Zero is the common result
- Ask the document controller how many documents in the CDE were AI-assisted. If the question is novel, you are here
- Ask the project director who owns AI risk on this project. A firm-level name is not a project answer
- Ask the bid team what was promised about AI in the winning submission, then ask the project team whether they have seen it
Anti-pattern · Adopting the framework at the enterprise and calling the projects done
The instinctive move is a corporate programme: map the whole AI RMF Core at group level, publish a policy, brief the business. It is genuinely useful work and it does not move a single project off this stage, because the framework's MAP, MEASURE and MANAGE functions are all context-bound — intended purpose, affected people and consequence all differ per project, per package and per contract. A group-level map of 72 subcategories with no project register underneath it is an enterprise artefact describing an empty set. Write the register on one live project first; the corporate map then has something real to generalise from.
What holds you here
There is no project-level record, so none of the framework's functions has anything to attach to.
Highest-leverage next move
Write the register: every AI use on one live project, each with a named owner and a one-paragraph intended purpose.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- A digital lead and a project manager, part-time, plus an hour each from six package leads
- Risk
- Low — the output is a register and a template; nothing about delivery changes yet
- To next stage
- 1–3 months
If this is you, the next step is
A two-week walk of one live project: every AI use found, named and owned.
Stage 2
Profiled
30% of operators sit here
Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
Profiled is the first stage where the framework is doing work rather than being cited. The register is a modest artefact — often a single spreadsheet or a CDE-hosted table with a dozen rows — and it changes the conversation immediately, because for the first time the project can answer the question 'what are you relying on?' with a list rather than an anecdote.
Two entries on that list surprise almost every project. The first is embedded AI: the scheduling optimiser inside the planning tool, the anomaly detection in the plant telematics portal, the auto-classification in the document system. Nobody procured these as AI and nobody thinks of them as AI, so they arrive under the radar and end up carrying real weight. The second is client-specified AI — a platform named in the employer's requirements that the project must use and cannot change. Both belong on the register precisely because the project did not choose them.
The stage's characteristic weakness is that the register describes and does not test. Intended purpose is written from the vendor's description, consequence class is assigned by feel, and the accuracy figure everyone repeats came from a sales deck. That is still a large improvement on stage 1, and it is also a trap: a well-formatted register can make a project feel assured while every number in it is somebody else's. MAP without MEASURE produces confident documentation of untested reliance.
In practice
The register that found the optimiser
A JV delivering a water-treatment upgrade ran a two-week register exercise across its packages. Ten of the twelve rows were expected. The eleventh was the sequencing optimiser built into the planning software, which had been quietly re-ordering resource-levelled activities since mobilisation and whose output fed the programme submitted monthly to the client. Nobody had classified it, nobody owned it, and its intended purpose — as far as the project could establish — was whatever the planner assumed it did. It went onto the register with a consequence class of 'commercial and programme', an owner, and a note that its outputs had never been compared with a hand-levelled baseline.
What it looks like
- A project AI register exists, dated at mobilisation and owned by a named individual
- Each row carries a one-paragraph intended-purpose statement written in project language
- Each row carries a consequence class taken from the project's own risk matrix
- Embedded and client-specified AI is on the register alongside the tools the project bought
Diagnostic signals you can check this week
- Open the register and count rows that describe AI the project did not buy. If the answer is zero, the sweep missed the embedded layer
- Pick a row and read its intended-purpose statement aloud to the package manager. If they disagree with it, the statement was written by the vendor
- Check whether consequence class uses the project's existing risk matrix or an invented scale. An invented scale will not survive a risk meeting
- Ask when the register was last updated. If it was mobilisation and the project is now at stage four, it is a snapshot, not a register
Anti-pattern · Grading the register instead of testing it
Once a register exists, the natural next move is to make it better: add columns, add a risk score, add a RAG status, circulate a template across the portfolio. This is administratively satisfying and epistemically empty, because every field in it is still an assertion. A three-column register with one measured number in it is worth more than a fifteen-column register with none. Take the highest-consequence row, build a small test set from the project's own records, and measure it. The register's real value is that it tells you which single row to measure first.
What holds you here
The register describes each AI use but nothing measures it, so reliance still rests on the vendor's numbers.
Highest-leverage next move
Build the project test set — held-out project records plus seeded defects — and set one written threshold per AI use.
Cost of leaving
- Effort
- 2–4 months
- Team
- A profile owner (usually the digital or design manager), a package manager per row, a data-protection lead for the rows that touch people
- Risk
- Medium — the first honest test can contradict a claim the project has already relied on
- To next stage
- 3–6 months
If this is you, the next step is
We assemble it from your own NCR, RFI and inspection records — typically three weeks.
Stage 3
Measured
20% of operators sit here
Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
Measured is where the framework starts paying for itself, and the reason is unglamorous: a number measured on this project is arguable in a way a vendor benchmark is not. When a design manager says a checking tool has a 4% miss rate on the project's own seeded non-conformances, the conversation moves immediately to what to do about the 4% — sampling rate, second check, which packages. When the same manager says the vendor reports 96% accuracy, the conversation stops, because nobody in the room can interrogate a number produced on somebody else's data.
The test set is the artefact that defines this stage, and building it is a records exercise rather than a data-science one. Every project already holds a corpus of known-bad examples: the NCR log, the RFI log, the snagging list, the rejected inspection records, the design-review comment tracker. Seeding a held-out set from those, plus a controlled set of defects the team introduces deliberately, gives a recall figure that means something in the project's own units. NIST's term for this whole activity is TEVV, and the AI RMF is explicit that its parameters have to be aligned to real deployment conditions rather than to laboratory ones.
The stage's limitation is that measurement without a rehearsed response is a report, not a control. Projects at this stage can tell you precisely how often a tool is wrong and have no agreed answer for the morning it is wrong about something that matters. The miss gets handled by whoever notices, competently and invisibly, and no record is created — which means the third recurrence of the same failure mode still looks like a first.
In practice
The 4% that changed a sampling rate
On a hospital fit-out, the design team ran an automated standards checker over mechanical and electrical submissions. The project built a held-out set of 120 records — 80 drawn from its own historic NCRs and RFIs, 40 with defects seeded deliberately by two engineers who then swapped sets — and measured recall at 96%, with the misses concentrated in clashes involving builder's work openings. The number itself was reassuring. What changed practice was the concentration: the project kept a manual second check on that one category, dropped the blanket 100% manual re-check everywhere else, and wrote both decisions into the register as a reliance decision with a date.
What it looks like
- A held-out test set exists per high-consequence AI use, built from the project's own records
- Each use has a written threshold agreed with the package manager who relies on it
- Test runs are dated, sampled and recorded, and repeated when the tool or the site changes
- An override log records what site actually did with the output, not just what was produced
Diagnostic signals you can check this week
- Ask to see a test record with a date, a sample size and a result. A screenshot of a vendor dashboard is not one
- Ask where the test cases came from. If the answer is 'the vendor supplied them', the measurement is circular
- Check whether the threshold is written down and whether the person who relies on the output agreed it
- Look at the override log. A tool with no overrides at all usually means nobody is checking, not that nothing is wrong
Anti-pattern · Measuring the model instead of the decision
Teams that get good at measurement often start optimising the wrong quantity: overall accuracy across all cases, reported monthly, trending gently upward. On a project, overall accuracy is close to meaningless because the consequences are wildly asymmetric — a missed clash in a riser is not a missed clash in a plant room, and a false positive costs an hour while a false negative can cost a package. Measure per consequence class, report recall on the categories that hurt, and accept a worse headline number in exchange for a defensible one. The AI RMF's own framing is that measurement has to be tied to context and intended use, not to a general benchmark.
What holds you here
Measurement exists but no response is rehearsed, so a bad output is handled by improvisation and leaves no record.
Highest-leverage next move
Wire the AI rows into the project's existing risk machinery — risk register, early warning notice, deviation record — and rehearse one end to end.
Cost of leaving
- Effort
- 4–8 months
- Team
- Profile owner, a QA or design manager per measured use, two engineers for a week to seed and score the test set
- Risk
- Medium — measurement can invalidate reliance the project has already priced into its programme
- To next stage
- 4–8 months
If this is you, the next step is
We connect the thresholds to your risk register, EWN route and deviation records.
Stage 4
Managed
11% of operators sit here
Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
Managed is the stage at which AI stops being a digital-team topic and becomes an ordinary project risk. The signature is boring in the best way: AI items appear on the same weekly risk agenda as temporary works, ground conditions and long-lead procurement, discussed by the same people in the same vocabulary, with the same expectations about owners and dates. Nothing about the framework requires a parallel governance structure, and projects that build one usually find it starved of attention within two months.
The characteristic capability here is rehearsal. Construction is unusually good at this — it already drills fire, rescue from height, confined-space entry and concrete-pour contingency — and the same instinct transfers directly. Running one deliberate exercise, in which a tool is declared unavailable or wrong on a live package and the team executes the fallback, converts a paragraph in the method statement into a known quantity. It also surfaces the thing paper never does: how many downstream artefacts had already absorbed the AI's output before anyone noticed.
The stage's blind spot is time. Everything works while the project exists, because the people who wrote the profile are still in the office and the reasoning is still in their heads. Nothing at this stage has been built to survive demobilisation. The evidence exists but has never been assembled; the register is current but has no consignee; the reasoning behind a reliance decision lives with a design manager who moves to the next scheme in March. A project can be genuinely excellent at MANAGE and still hand over nothing at all.
In practice
The exercised fallback that found six documents
A rail systems project ran a half-day exercise on its document-classification assistant: the tool was declared unavailable at 09:00 and the team executed the manual route in the method statement. The manual route worked. What the exercise found was downstream — six technical submissions already in the client's approval queue carried classification metadata the assistant had produced, and none was flagged as machine-assisted. The fix was a metadata field and a CDE naming rule, agreed the same week. Neither would have surfaced from a document review, because on paper the fallback was already compliant.
What it looks like
- AI rows sit in the project risk register with owners, dates and movement between reviews
- A named non-AI fallback exists in the method statement and has been exercised at least once
- A deviation record and an early warning notice have actually been raised on an AI output
- Model or version changes from suppliers trigger a re-test rather than a silent update
Diagnostic signals you can check this week
- Open the last four project risk-register versions and look for AI rows that moved. Static rows mean the register is decorative
- Ask when the fallback was last exercised, and whether the exercise found anything. 'It has never needed to be' is not an answer
- Ask what happened the last time a supplier pushed a model update. If nobody noticed, change control does not cover the AI layer
- Ask who receives the AI evidence at practical completion. At this stage the honest answer is usually nobody
Anti-pattern · Building a parallel AI governance forum
The temptation at this stage is to formalise: a monthly AI risk board, its own minutes, its own escalation path, its own reporting line to the enterprise. On a project it starves. Attendance decays through the second month, the commercial team never joins, and the forum's decisions arrive too late to affect a package that was let six weeks earlier. Everything the framework asks for at MANAGE maps onto machinery projects already run daily — the risk register, the early warning notice, the deviation record, the method statement, the change-control process. Use them. The measure of success is that AI risk becomes indistinguishable in process terms from any other project risk.
What holds you here
Everything works while the project exists; nothing has been built to survive handover, demobilisation or the next mobilisation.
Highest-leverage next move
Make the evidence pack a deliverable: name it in the information-delivery plan and hand it over with the asset information.
Cost of leaving
- Effort
- 6–12 months
- Team
- Project risk manager, profile owner, package managers, plus commercial input where supplier change control is involved
- Risk
- Medium — exercising a fallback on a live package costs real hours and occasionally exposes uncomfortable downstream dependency
- To next stage
- 9–18 months
If this is you, the next step is
We draft the information-delivery requirement and the pack's contents list.
Stage 5
Transferable
3% of operators sit here
Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.
Transferable is a narrow stage and deliberately so. It does not mean the project has automated anything, and it does not mean the AI is more capable than at stage 4. It means one specific thing: that the record of what was relied on, how it was tested and what went wrong has been made into an object that crosses two organisational boundaries — the handover to a permanent asset owner, and the mobilisation of the next temporary organisation.
The handover boundary is the one with statutory weight. In the UK, the Building Safety Act regime made the durability of building information an explicit duty for higher-risk buildings, and the CDM regulations have long required a health and safety file to pass to the client at completion. Neither instrument names AI, and neither needs to: the question a duty holder will ask in year seven is what the design or inspection record was based on, and 'a tool the contractor used, which nobody has any record of' is a bad answer regardless of which regime is asking. The AI evidence pack is the cheapest available insurance against that question, and it costs almost nothing if it is assembled continuously rather than reconstructed at completion.
The mobilisation boundary is the one with commercial weight. A firm whose projects each start from a blank register pays the same setup cost every time and learns nothing across schemes. A firm whose next mobilisation starts from the last project's returned profile — the same tools, the same consequence classes, the same test sets, the same known failure categories — compresses stage 1 and 2 into a fortnight. This is also the only mechanism by which GOVERN becomes real rather than declarative: the enterprise's policy improves because projects hand something back, not because the policy was rewritten.
In practice
The pack that answered a year-three question
A design-and-build contractor handed a school over with an AI evidence pack comprising the project register, seven intended-purpose statements, four test-run records, two reliance decisions and a deviation history with three closed entries. In year three the operator queried an as-built discrepancy in a services zone. The pack showed that the automated clash review had run at RIBA stage 4, that builder's work openings had been a known weak category with a manual second check retained, and that the specific zone had been manually verified on a dated record. The query closed in a fortnight rather than becoming a claim.
What it looks like
- The AI evidence pack is named in the information-delivery plan as a contracted deliverable
- The asset owner receives the register, the test records, the reliance decisions and the deviation history
- Mobilisation on the next project starts from a returned profile, not from a blank template
- Portfolio review of returned profiles changes the policy, the template and the approved tool catalogue
Diagnostic signals you can check this week
- Look in the information-delivery plan for the AI evidence pack by name. If it is not a named deliverable, it will not be produced
- Ask the asset owner's information manager whether they have ever received one, and what they did with it
- Compare two consecutive mobilisations: did the second start from the first's register or from a blank template?
- Ask what changed in the corporate policy or the approved tool catalogue as a result of the last project's returned profile
Anti-pattern · Treating the pack as a completion task
The pack looks like a document, so it gets scheduled like one: a task in the handover programme, assigned in the final quarter, to be assembled from whatever can still be found. By then the design manager who wrote the reliance decisions has gone, the test set lives on a demobilised laptop, and half the deviation history was handled verbally. What is produced is a folder that satisfies the deliverable and answers no future question. The pack has to accrete — every register update, test run, reliance decision and deviation filed into it on the day it happens — so that completion is a transmittal rather than an archaeology exercise.
What holds you here
Sustaining transferability is a portfolio discipline — the constraint becomes whether the enterprise actually reads what its projects hand back.
Highest-leverage next move
Close the loop: make portfolio review of returned project profiles a scheduled item that changes the policy, the template and the tool catalogue.
Cost of leaving
- Effort
- Continuous
- Team
- Project information manager plus a standing portfolio reviewer in the enterprise
- Risk
- Concentrated — low frequency, long latency; the failure surfaces years after the team has gone
If this is you, the next step is
We pose a year-three query and see whether the pack can answer it.
Where construction and infrastructure projects actually sit
The distribution across the ladder, and why the Profiled → Measured step loses more projects than any other.
Most projects are at Referenced or Profiled. The distribution is heavily weighted toward the two rungs that produce documents rather than measurements: a large majority of live schemes can point to a firm-level policy, a growing minority hold a project register, and only a small fraction have a number about an AI tool that was produced on their own data. Fewer still have handed anything to an asset owner.
Distribution of construction and infrastructure projects across the ladder
Referenced is the mode. The drop from Profiled to Measured is the largest single loss on the ladder, because it is the first rung that requires the project to produce evidence rather than description.
Share of projects
- 36% — 1 · Referenced (the mode)
- 30% — 2 · Profiled
- 20% — 3 · Measured
- 11% — 4 · Managed
- 3% — 5 · Transferable
The shape of the loss is specific to project delivery. In a persistent organisation the Profiled → Measured step is a resourcing decision: someone is asked to build a test set and does. On a project it is a scheduling problem, because the window in which measurement is cheap — before reliance is priced into the programme — is also the window in which the team is busiest. NIST published the framework in January 2023 (opens in a new tab) with an accompanying roadmap (opens in a new tab) precisely because measurement science for AI was the acknowledged weak point, and construction inherits that weakness on top of its own.
It is worth being clear about what a distribution like this does and does not tell you. It is a synthesis, labelled illustrative above, and the honest comparison for any given project is not against the market but against schemes of similar contract form, client type and package mix. A design-and-build hospital under a two-stage contract has a different achievable rung from a five-year highways framework lot, because the framework lot has repeat mobilisations to learn from and the hospital does not.
The project profile: every function against the artefact that carries it
The page's centrepiece — the AI RMF Core crosswalked to construction project artefacts, with the owner on a project and the evidence each one leaves behind.
A project profile is the AI RMF Core answered in project artefacts rather than in new documents. That is the single most important design decision available to a project team: for almost every category in the framework, construction already runs an instrument that satisfies it — the project execution plan, the RACI, the risk register, the method statement, the DPIA, the deviation record, the early warning notice, the information-delivery plan. Creating a parallel AI governance apparatus is both more expensive and less durable, because the parallel apparatus has no natural attendance and no contractual home.
The table below is the crosswalk we use to write a profile. It is organised by AI RMF category rather than subcategory — the subcategory detail belongs in the Playbook (opens in a new tab), which is where you go once you know which categories bite for a given use. Read it as a checklist of artefacts, not of clauses: if the right-hand column is empty for a row that matters to your consequence class, that is the piece of work. Two rows — GOVERN 6 and MANAGE 3 — reach outside the project into supplier paper, and the contractual machinery behind them is a subject of its own, covered by the vendor-governance page in this cell.
| Function · category | What it asks of the project | Project artefact that carries it | Owner on a project | Evidence it leaves behind |
|---|---|---|---|---|
| GOVERN 1 · Policies and procedures | That documented policies and procedures for AI risk are in place and applied | The firm's AI policy, referenced by name and version in the PEP, with project deviations recorded | Digital or technical director; project director countersigns | A PEP section naming the policy version in force at mobilisation |
| GOVERN 2 · Accountability structures | That roles and lines of responsibility for AI risk are documented and staffed | The project RACI, extended with a named profile owner for each AI use | Project director | A RACI row per AI use naming an individual, not a team |
| GOVERN 4 · Risk culture | That teams are committed to identifying and communicating AI risk | The standing AI item on the weekly project risk meeting, and the induction slide | Project manager | Risk meeting minutes carrying AI items alongside temporary works |
| GOVERN 6 · Third-party risk | That policies address AI risks arising from supplier software and data | The AI schedule in the subcontract or supply agreement, and the register's supplier column | Commercial manager | Signed schedules plus a register row per supplier-brought tool |
| MAP 1 · Context established | That the context of use, and who is in it, is understood and recorded | The project AI register: each use, its package, its site, its data and its users | Profile owner | A register with one row per AI use, dated at mobilisation and updated at gates |
| MAP 2 · Categorisation | That the AI system is categorised by what it does and how it is used | The intended-purpose statement — one paragraph per use, in project language | Package or design manager | A signed intended-purpose statement per AI use |
| MAP 3 · Capabilities and expected benefits | That targeted usage, goals and expected benefits are understood | The reliance decision: what the output may and may not be used for | Project director with the profile owner | A dated reliance decision recorded against the register row |
| MAP 4 · Risks and benefits mapped | That risks are mapped for all components, including third-party software and data | Consequence classification against the project's existing risk matrix | Project risk manager | A consequence class per AI use, on the project's own scale |
| MAP 5 · Impacts characterised | That impacts on individuals, groups and communities are characterised | The DPIA, plus the workforce and neighbour consultation record where site monitoring is involved | Data protection lead with the HSE manager | DPIA reference and consultation minutes attached to the register row |
| MEASURE 1 · Methods and metrics | That appropriate methods and metrics are identified, documented and applied | The measurement plan: metric, threshold, sample size, cadence, who reads it | Profile owner | A versioned measurement plan per AI use |
| MEASURE 2 · Trustworthy characteristics evaluated | That the system is evaluated for validity, reliability, safety, privacy and bias in context | The project test set: held-out project records plus deliberately seeded defects | Design or QA manager | Test-run records with dates, sample sizes and results by category |
| MEASURE 3 · Risks tracked over time | That mechanisms exist to track identified risks as conditions change | The monthly AI line in the project performance report | Project controls | A trend by month, not a single snapshot |
| MEASURE 4 · Feedback on measurement | That feedback on the efficacy of measurement is gathered and assessed | The override log — what site actually did with each output | Package manager | Override and acceptance rates by package and by month |
| MANAGE 1 · Risks prioritised and responded to | That AI risks are prioritised, responded to and managed on a documented basis | AI rows in the project risk register, with owners, dates and review movement | Project risk manager | Risk register entries that move between reviews |
| MANAGE 2 · Benefits maximised, impacts minimised | That strategies exist to sustain value while limiting negative impact | The named non-AI fallback retained in the method statement, and its exercise record | Package manager | A method statement naming the manual route, plus a dated exercise record |
| MANAGE 3 · Third-party risks managed | That risks from third-party systems and data are managed in service | The model-change notice route in the supply agreement, and the re-test trigger | Commercial manager with the profile owner | Change notices received, and the re-test record each one triggered |
| MANAGE 4 · Treatments documented and communicated | That treatments are documented, monitored and incidents communicated | The deviation record and the early warning notice | Project manager | Deviation records with dates, owners and closure |
Three triage rules make this table usable in a mobilisation week rather than a quarter. They are the difference between a profile a project actually maintains and a matrix that is completed once and never opened.
Triage by consequence, not by tool count
A project with eleven AI uses does not need eleven full profiles. Two or three will carry safety, structural or statutory consequence and earn every row of the table; the rest carry commercial or programme consequence and earn a register row, an intended-purpose statement and a threshold. Applying uniform depth is the fastest way to guarantee the deep rows are done badly, because the effort was spent on a document summariser.
Name the existing artefact before you create a new one
For every row, ask what the project already produces that could carry it. Almost always something does. The DPIA already exists for site monitoring; the method statement already names manual checks; the early warning notice already exists as a contractual mechanism under NEC. Creating an AI-specific parallel is how governance ends up unattended — and an artefact with a contractual home outlives one with only a policy home.
Write the owner as a person, at mobilisation
Every row on the crosswalk has an owner column for a reason. 'The digital team' is not an owner on a project; a named design manager is. Assign at mobilisation while the RACI is being agreed, and re-assign explicitly at every demobilisation — the framework's accountability requirement is the one construction is best equipped to meet and most likely to let lapse quietly when someone rolls off.
One extension is worth knowing about. Where generative tools are in use — bid text, specification drafting, document summarisation, RFI responses — NIST publishes a separate Generative AI Profile (NIST AI 600-1) (opens in a new tab), which enumerates twelve risks specific to or exacerbated by generative systems and sets out suggested actions against the same four functions. Two of the twelve matter disproportionately on projects: confabulation, which NIST defines as confidently stated but erroneous content, and information integrity. On a project both attach to the same artefact — a document in the CDE that reads as authoritative because it is well-formatted. The practical response is a provenance field on generated documents and a named human check before issue, both of which sit naturally in the information-delivery plan.
The temporary-organisation problem, and how much evidence a use has to carry
Why a framework written for persistent organisations strains against project delivery — and the 2×2 that decides how deep the measurement has to go.
The AI RMF strains against construction because it assumes an organisation that persists, and construction delivers through organisations that dissolve. A joint venture is incorporated to build one thing and wound up afterwards. A framework lot runs for five years with a delivery team that turns over twice. A design-and-build contract hands a permanent asset from a temporary builder to a permanent owner and then removes every person who understood the decisions. None of that is a defect in the framework; it is a translation problem, and it produces four failure points that are specific to project delivery.
Ownership dissolves faster than the risk does
A model's consequences persist in the asset for decades; the person who accepted its output has moved on within eighteen months. This is not new to construction — it is the reason design-check certificates and the CDM health and safety file exist. Apply the same instinct: every reliance decision names an individual and a date, and demobilisation triggers explicit re-assignment rather than silent lapse. The CDM 2015 regime (opens in a new tab) already establishes that duty holders and their information obligations are handed on rather than extinguished at completion.
Joint ventures have two policies and one project
On a JV, each parent brings its own AI policy, approved tool catalogue and risk appetite, and they will not agree. Resolving that at parent level takes longer than the project has. The workable move is to settle it once, at mobilisation, in the JV's own PEP: which parent's catalogue governs, whose DPIA template applies, where the register lives, and who owns each row. Two hours of argument at mobilisation replaces a recurring one at every package let.
The client specifies AI the project cannot refuse
Employer's requirements increasingly name platforms — a progress-monitoring system, a common data environment with built-in classification, an asset-tagging tool. The project cannot change them and still carries their consequences on site. Register them explicitly as client-specified, record the reliance decision anyway, and raise the residual risk through the contract's own early-warning mechanism rather than absorbing it silently.
Handover is a cliff, not a taper
On the day of practical completion, the project's memory ends. Every reliance decision that was not written down, every test that was not recorded, every deviation that was handled verbally becomes unrecoverable. In the UK the building safety regime (opens in a new tab) has already pushed information durability up the agenda for higher-risk buildings, and the Building Safety Regulator (opens in a new tab) is explicit that information must transfer with the asset. An AI evidence pack is the cheapest way to make the project's reasoning transferable, and it is only cheap if it accretes rather than being assembled at the end.
The second half of the translation problem is depth. Once the register exists, the recurring question is not whether to test but how hard, and the answer is a function of two variables: what happens if the output is wrong, and where the number you are relying on came from. The matrix below is how we set that on a project.
How much test evidence does this AI output have to carry?
Plot each register row. The vertical axis is the consequence class you assigned during MAP; the horizontal is the provenance of the performance number you are actually relying on. Only one quadrant is genuinely dangerous, and one is merely expensive.
Unearned reliance
- High consequence, someone else's number
- The quadrant every serious incident starts in
- Fix: build a project test set before the next package relies on it
Warranted reliance
- High consequence, measured on your own records
- Thresholds written, second check retained where recall is weak
- Fix: nothing — re-test when the tool, the site or the standard changes
Proportionate
- Low consequence, vendor benchmark
- Register row, intended purpose, a threshold and periodic sampling
- Fix: sample quarterly and move on — depth here buys nothing
Over-assured
- Low consequence, fully measured
- TEVV effort spent where it cannot pay back
- Fix: move the effort to the top-left row that is still untested
It is the joint responsibility of all AI actors to determine whether AI technology is an appropriate or necessary tool for a given context or purpose, and how to use it responsibly.
That sentence is the reason the matrix has two axes rather than one. Appropriateness is contextual: the same clash-detection tool is warranted on a warehouse and unearned on a hospital riser, not because the model changed but because the consequence class did. It is also why a project cannot inherit its depth decision from a corporate policy — the policy does not know which package you are letting next month.
What the enterprise-and-project split looks like in public
Two publicly reported programmes — one contractor, one asset owner — read against the ladder. Neither is an Atomic Loops engagement; each links to the organisation's own published material.
The clearest public evidence for the enterprise-and-project split is in what large organisations have chosen to build and to require. A contractor that builds AI centrally still has to profile each arrival on each project; an asset owner that mandates digital delivery is writing the requirements that make a project's evidence pack exist. The two cases below sit on opposite sides of the same boundary, and both are read here against the ladder rather than presented as endorsements.
Two reference points read against the ladder
Outcomes as published by the organisations themselves; we have not independently audited them — verify figures against the linked source before reusing them. Both card images are generated industry scenes from this page's image library, not photographs of the named organisations.
SkanskaGlobal contractor and project developer · Nordics, Europe, US13
- Challenge
- Delivering AI capability to a business that executes through hundreds of separate project organisations, where a tool approved centrally still arrives on each project with a different intended purpose, a different consequence class and a different client's requirements attached.
- Approach
- Skanska USA Building established a Digital Transformation and Solutions Team uniting its data, emerging technology and AI capabilities, and built internal tools — the Sidekick suite and Skanska Metriks cost modelling — on the firm's own accumulated project record rather than licensing general-purpose products.
- Reported outcome
- Skanska reports that the Sidekick suite scaled from an initial 2024 pilot to more than 1,000 employee users supporting work across 500-plus projects, with the tools intended to surface safety and operational risks earlier and reduce administrative burden.
- What it shows about the curveBuilding centrally is a strong GOVERN position and does not discharge MAP or MEASURE. A tool used across 500 projects has 500 intended-purpose statements to write and 500 consequence classes to assign, because what it is for and what breaks if it is wrong differ by contract, package and client — which is exactly why the profile is per project and not per tool.
Skanska — Digital Transformation and Solutions Team press release (opens in a new tab)
National HighwaysStrategic road network operator · England · asset owner and client24
- Challenge
- Operating a permanent network that is built and modified by a rotating population of temporary supplier project organisations, each of which demobilises and takes its reasoning with it while the operator retains the asset for decades.
- Approach
- National Highways publishes its Digital Roads programme, setting out how digital design, construction and operation — including connected data and digital twins of the network — are to be delivered by its supply chain, so requirements are stated by the permanent organisation rather than negotiated per scheme.
- Reported outcome
- The published programme positions digital delivery and data as network-level requirements on suppliers across design, construction and operation, making information transfer an explicit expectation of the client rather than a project-team preference.
- What it shows about the curveThe rung above Managed is bought by the client, not by the contractor. Where the asset owner names the information it will receive, the evidence pack becomes a deliverable with a date and a recipient; where it does not, even a well-run project hands over nothing, because nobody asked and nobody would have read it.
Read together, the two cases describe the boundary this page is about. The contractor case shows that enterprise capability, however good, stops at the project gate: the framework's context-bound functions have to be executed by whoever is standing on the site. The asset-owner case shows the other side: the durability of a project's AI record is largely determined by what the permanent organisation asked for before the project began. A project team that wants to reach Transferable and has a silent client will have to write the requirement itself, into the information-delivery plan, and get it agreed.